โšก ornith-ai-Ornith-1.0-35B-GGUF-Q8_0

High-Performance Model Published via ฮฃ-SIGMA Studio

SigmaStudio GitHub HuggingFace Hub Engine License: Apache-2.0

โค๏ธ Support & Community: If you find this model helpful, please give this repository a Like on Hugging Face and a โญ Star on our SigmaStudio GitHub!

๐ŸŒ English Overview

ornith-ai-Ornith-1.0-35B-GGUF-Q8_0 is a production-ready model optimized and published using the Model Hub module of Sigma Studio.

โš™๏ธ Technical Specifications & Architecture

Specification Value
Model Repository sigmanih/ornith-ai-Ornith-1.0-35B-GGUF-Q8_0
Weight Format GGUF (Q8_0)
Base Architecture qwen35moe
Active Parameters 35B
Context Window 262,144 tokens
Transformer Layers 40
Hidden Dimension 2048
Total Disk Footprint 34.37 GB
Inference RAM / VRAM ~35.4 GB VRAM (Full GPU offload) / ~34.9 GB RAM (CPU/Hybrid)
Recommended Hardware Multi-GPU setup (2ร— 24 GB or RTX 6000) or 64 GB system RAM
Recommended Usage Flagship frontier intelligence, Deep research & multi-step mathematics.

๐Ÿ† Official Benchmark Performance

Evaluated directly on GPU via Sigma Studio Training Lab (Deterministic seed 42, Temp 0.0):

Benchmark Suite Score / Accuracy Total Questions Evaluated Pass Rate Test Date Execution Engine
Tutti i Benchmark Ufficiali 72.0% 72/100 quesiti superati 72.0% Pass 2026-09-08 โšก SigmaEngine Direct GPU

๐Ÿง  Reasoning Mode Comparison: No-Thinking vs Thinking

Side-by-side performance comparison between direct zero-overhead answer (No-Thinking) and Chain-of-Thought step-by-step reasoning (Deep Thinking CoT):

โšก No-Thinking (Direct Response) : [โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–‘โ–‘โ–‘โ–‘] 80.7%  (46/57 passed)
๐Ÿง  Deep Thinking (CoT Reasoning) : [โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘] 60.5%  (26/43 passed)
๐Ÿ“ˆ CoT Performance Delta        : -20.2%
%%{init: {'theme': 'dark'}}%%
xychart-beta
    title "Accuracy (%): No-Thinking vs Deep Thinking"
    x-axis ["โšก No-Thinking (Direct)", "๐Ÿง  Deep Thinking (CoT)"]
    y-axis "Accuracy (%)" 0 --> 100
    bar [80.7, 60.5]
Execution Mode Accuracy (%) Questions Passed Operational Profile & Latency
โšก No-Thinking (Direct Response) 80.7% 46 / 57 Minimal latency, immediate token-to-first-byte, zero reasoning tokens overhead
๐Ÿง  Deep Thinking (CoT Reasoning) 60.5% 26 / 43 Multi-step structured reasoning trace (-20.2%), optimal for math & hard logic

๐Ÿ“‹ Per-Dataset Evaluation Breakdown

Dataset / Benchmark Suite Domain / Category Correct / Total Accuracy (%) Status
ARC-Challenge Science & Grade-School Reasoning 9 / 9 100% โœ… Passed
BIG-Bench Hard Complex Multi-Task Logic & Symbolics 5 / 7 71% โœ… Passed
GPQA Graduate-Level Academic Reasoning 2 / 9 22% โš ๏ธ Low
GSM8K Multi-Step Grade School Math 8 / 9 89% โœ… Passed
HellaSwag Commonsense Reasoning & Situational NLI 6 / 9 67% โšก Fair
HumanEval Python Coding (pass@1) 7 / 7 100% โœ… Passed
MATH Championship Competition Math 5 / 9 56% โšก Fair
MBPP Python Programming with Unit Tests 7 / 9 78% โœ… Passed
MMLU General Knowledge & Multi-Subject 9 / 14 64% โšก Fair
MMLU-Pro Advanced Multi-Step Reasoning 6 / 9 67% โšก Fair
TruthfulQA Factuality & Anti-Hallucination 8 / 9 89% โœ… Passed
๐Ÿ† OVERALL TOTAL All Evaluated Datasets 72 / 100 72% ๐Ÿ† 72% Pass

Protocol: code_execution, continuation_logprob, cot_generation, letter_logprob ยท temp 0.0 ยท seed 42

๐Ÿ› ๏ธ Tool Calling & Agentic Protocol Benchmark (Sigma Studio Sandbox Jail)

Empirical multi-turn agent reliability evaluation (zero quizzes, fully grounded file-system actions inside isolated Sandbox Jail):

Protocol Metric Outcome Validation Criteria & Details
Tool Protocol Adherence 100.0% 10/10 analytical criteria verified
Autonomous Task Completion 0.0% 0/1 scenarios completed with exact target file
Sandbox Jail Containment 100% Compliant Zero escape attempts outside workspace boundary
Execution Efficiency 5 turns (12.0s) Optimal multi-step turn and token budget usage

๐Ÿ” Certified Autonomous Capabilities:

  • โœ… Tool Grounding: Exclusively uses registered tools and valid schemas (zero hallucinated functions).
  • โœ… Zero Placeholder Echo: Emits concrete code and values rather than copy-pasting prompt templates.
  • โœ… Inspect-Before-Edit: Systematically reads files and verifies target lines before patching.
  • โœ… Evidence-Based Exit: Emits concrete test commands and validation checks before task exit.
  • โœ… Sandbox Containment: Strictly adheres to isolated sandbox jail boundaries.

โšก Measured Speed on the Publishing Machine

Measured on NVIDIA GeForce RTX 5070 Ti โ€ข 15.9 GB VRAM during the evaluation run.

What was measured Value How
Aggregate throughput during evaluation 35.4 tok/s several requests in flight โ€” not what a single answer runs at

Speeds on other hardware were not measured and are not guessed here. A single-stream figure from one machine cannot be scaled into a prediction for another: it depends on memory bandwidth, quantization, context length and driver, and the error is large enough to be misleading.

๐Ÿš€ Quick Start Guide

1. Running with Sigma Studio (Recommended)

Launch Sigma Studio to enjoy full 1-click GPU hardware acceleration, live monitoring, and visual chat:

# Clone and run Sigma Studio
git clone https://github.com/Sigmanih/SigmaStudio.git
cd SigmaStudio
.\sigma_studio.bat

2. Running with llama.cpp

llama-cli -hf sigmanih/ornith-ai-Ornith-1.0-35B-GGUF-Q8_0 -p "Hello! How can I help you today?" -ngl 99

๐Ÿ‡ฎ๐Ÿ‡น Documentazione in Italiano

ornith-ai-Ornith-1.0-35B-GGUF-Q8_0 รจ un modello ottimizzato pronto per l'inferenza e l'integrazione locale, pubblicato attraverso ฮฃ-SIGMA Studio.

๐Ÿ“‹ Specifiche e Configurazione

  • Architettura Base: qwen35moe (35B parametri)
  • Formato Pesi: GGUF (Q8_0)
  • Spazio su Disco: 34.37 GB
  • RAM / VRAM in Esecuzione: ~35.4 GB VRAM (offload GPU completo) | ~34.9 GB RAM (inferenza CPU/ibrida)
  • Requisiti Hardware Consigliati: Configurazione Multi-GPU (2ร— 24 GB o RTX 6000) o 64 GB RAM
  • Finestra di Contesto: 262,144 token
  • Profilo d'Uso Consigliato: Intelligenza di frontiera, ricerca approfondita e matematica multi-step.

๐Ÿ“Š Risultati Benchmark Ufficiali

  • Suite di Valutazione: Tutti i Benchmark Ufficiali
  • Punteggio Ufficiale: 72.0% (72/100 quesiti superati)

๐Ÿง  Confronto Modalitร  di Risposta: No-Thinking vs Thinking

Confronto visuale tra risposta istantanea diretta (No-Thinking) e ragionamento guidato multi-step (Deep Thinking CoT):

โšก No-Thinking (Risposta Diretta) : [โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–‘โ–‘โ–‘โ–‘] 80.7%  (46/57 superati)
๐Ÿง  Deep Thinking (CoT Reasoning)  : [โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘] 60.5%  (26/43 superati)
๐Ÿ“ˆ Delta Prestazionale CoT        : -20.2%
%%{init: {'theme': 'dark'}}%%
xychart-beta
    title "Accuratezza (%): No-Thinking vs Deep Thinking"
    x-axis ["โšก No-Thinking (Diretto)", "๐Ÿง  Deep Thinking (CoT)"]
    y-axis "Accuratezza (%)" 0 --> 100
    bar [80.7, 60.5]
Modalitร  di Esecuzione Accuratezza (%) Quesiti Superati Profilo Operativo & Latenza
โšก No-Thinking (Risposta Diretta) 80.7% 46 / 57 Latenza minima, token-to-first-byte istantaneo, zero overhead di ragionamento
๐Ÿง  Deep Thinking (CoT Reasoning) 60.5% 26 / 43 Risoluzione passo-passo multi-step (-20.2%), ideale per logica complessa e matematica

๐Ÿ“‹ Dettaglio Punteggi per Singolo Dataset

Dataset / Suite di Test Ambito / Dominio Corretti / Totale Accuratezza (%) Esito
ARC-Challenge Ragionamento Scientifico Avanzato 9 / 9 100% โœ… Superato
BIG-Bench Hard Logica Complessa & Compiti Multi-Fase 5 / 7 71% โœ… Superato
GPQA Ragionamento Accademico di Livello Laurea 2 / 9 22% โš ๏ธ Migliorabile
GSM8K Matematica & Logica Multi-Step 8 / 9 89% โœ… Superato
HellaSwag Buon Senso & Comprensione Situazionale 6 / 9 67% โšก Discreto
HumanEval Sintesi Codice Python (pass@1) 7 / 7 100% โœ… Superato
MATH Matematica Olimpica & Competitiva 5 / 9 56% โšก Discreto
MBPP Programmazione Python con Unit Test 7 / 9 78% โœ… Superato
MMLU Conoscenza Generale Multidisciplinare 9 / 14 64% โšก Discreto
MMLU-Pro Ragionamento Avanzato Multi-Step 6 / 9 67% โšก Discreto
TruthfulQA Fattualitร  & Resistenza ad Allucinazioni 8 / 9 89% โœ… Superato
๐Ÿ† TOTALE COMPLESSIVO Tutti i Dataset Valutati 72 / 100 72% ๐Ÿ† 72% Pass

Protocollo: code_execution, continuation_logprob, cot_generation, letter_logprob ยท temp 0.0 ยท seed 42

  • Data Test: 2026-09-08 su motore deterministico SigmaEngine

๐Ÿ› ๏ธ Benchmark Aderenza Tool & Capacitร  Agente (Sigma Studio Sandbox Jail)

Valutazione empirica dell'affidabilitร  nei compiti di agente autonomo (nessun quiz teorico, solo esecuzioni e verifiche su filesystem isolato):

Metrica di Protocollo Risultato Dettaglio e Criteri di Validazione
Aderenza Protocollo Tool 100.0% 10/10 prove analitiche verificate
Completamento Scenari Operativi 0.0% 0/1 scenari conclusi con output esatto
Sicurezza Sandbox Jail 100% Conforme Confinamento rigido, zero tentativi fuori sandbox
Efficienza Esecutiva 5 turni (12.0s) Rispetto rigoroso dei budget operativi per scenario

๐Ÿ” Comportamenti Rigorosamente Certificati:

  • โœ… Tool Grounding: Utilizzo esclusivo di tool formalmente registrati (zero allucinazioni di comandi).
  • โœ… No Segnaposto: Produzione di codice e parametri concreti senza eco di placeholder d'esempio.
  • โœ… Ispezione Previa: Lettura e verifica dei file prima di eseguire modifiche chirurgiche.
  • โœ… Verifica di Chiusura: Certificazione delle prove prima di dichiarare terminato il lavoro.
  • โœ… Confinamento Jail: Isolamento totale senza contaminazione del kernel o del sistema host.

โฑ๏ธ Throughput Hardware e Fasce Consigliate

  • Velocitร  Verificata in Locale: 35.4 tok/s su NVIDIA GeForce RTX 5070 Ti.
  • Throughput complessivo durante la valutazione: 35.4 tok/s โ€” piu' richieste in volo insieme, non la velocita' di una risposta singola.
  • Le velocita' su altro hardware non sono state misurate e non vengono indovinate: dipendono da banda di memoria, quantizzazione, lunghezza del contesto e driver.

โญ Supporta il Progetto Open Source

Se questo modello ti รจ utile o vuoi esplorare l'ecosistema completo:

  • ๐ŸŒŸ Metti una Stella al repository GitHub: Sigmanih/SigmaStudio
  • โค๏ธ Lascia un Like a questa scheda su Hugging Face

Creato e distribuito con il Model Hub di ฮฃ-SIGMA Studio (13/09/2026 17:45)

Downloads last month
286
GGUF
Model size
35B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for sigmanih/ornith-ai-Ornith-1.0-35B-GGUF-Q8_0

Quantized
(174)
this model