Image-Text-to-Text
GGUF
English
llama.cpp
qwen
qwen3.6
qwen-vl
multimodal
vision
ornith
turboquant
tq3_4s
conversational
Instructions to use YTan2000/Ornith-1.0-35B-TQ3_4S with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use YTan2000/Ornith-1.0-35B-TQ3_4S with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf YTan2000/Ornith-1.0-35B-TQ3_4S:F16 # Run inference directly in the terminal: llama cli -hf YTan2000/Ornith-1.0-35B-TQ3_4S:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf YTan2000/Ornith-1.0-35B-TQ3_4S:F16 # Run inference directly in the terminal: llama cli -hf YTan2000/Ornith-1.0-35B-TQ3_4S:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf YTan2000/Ornith-1.0-35B-TQ3_4S:F16 # Run inference directly in the terminal: ./llama-cli -hf YTan2000/Ornith-1.0-35B-TQ3_4S:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf YTan2000/Ornith-1.0-35B-TQ3_4S:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf YTan2000/Ornith-1.0-35B-TQ3_4S:F16
Use Docker
docker model run hf.co/YTan2000/Ornith-1.0-35B-TQ3_4S:F16
- LM Studio
- Jan
- vLLM
How to use YTan2000/Ornith-1.0-35B-TQ3_4S with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "YTan2000/Ornith-1.0-35B-TQ3_4S" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "YTan2000/Ornith-1.0-35B-TQ3_4S", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/YTan2000/Ornith-1.0-35B-TQ3_4S:F16
- Ollama
How to use YTan2000/Ornith-1.0-35B-TQ3_4S with Ollama:
ollama run hf.co/YTan2000/Ornith-1.0-35B-TQ3_4S:F16
- Unsloth Desktop
- Pi
How to use YTan2000/Ornith-1.0-35B-TQ3_4S with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf YTan2000/Ornith-1.0-35B-TQ3_4S:F16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "YTan2000/Ornith-1.0-35B-TQ3_4S:F16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use YTan2000/Ornith-1.0-35B-TQ3_4S with Docker Model Runner:
docker model run hf.co/YTan2000/Ornith-1.0-35B-TQ3_4S:F16
- Lemonade
How to use YTan2000/Ornith-1.0-35B-TQ3_4S with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull YTan2000/Ornith-1.0-35B-TQ3_4S:F16
Run and chat with the model
lemonade run user.Ornith-1.0-35B-TQ3_4S-F16
List all available models
lemonade list
- Hermes Agent
How to use YTan2000/Ornith-1.0-35B-TQ3_4S with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf YTan2000/Ornith-1.0-35B-TQ3_4S:F16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default YTan2000/Ornith-1.0-35B-TQ3_4S:F16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use YTan2000/Ornith-1.0-35B-TQ3_4S with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf YTan2000/Ornith-1.0-35B-TQ3_4S:F16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "YTan2000/Ornith-1.0-35B-TQ3_4S:F16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
File size: 3,538 Bytes
531b14a | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 | ---
license: apache-2.0
language:
- en
library_name: gguf
pipeline_tag: text-generation
model_name: Ornith-1.0-35B-TQ3_4S
tags:
- gguf
- llama.cpp
- qwen
- qwen3.6
- ornith
- turboquant
- tq3_4s
base_model:
- deepreinforce-ai/Ornith-1.0-35B
model-index:
- name: Ornith-1.0-35B-TQ3_4S
results: []
---
# Ornith-1.0-35B-TQ3_4S

`Ornith-1.0-35B-TQ3_4S` is a compact TurboQuant GGUF build of [deepreinforce-ai/Ornith-1.0-35B](https://huggingface.co/deepreinforce-ai/Ornith-1.0-35B).
## Required Runtime
This model uses the custom `TQ3_4S` tensor type and requires [turbo-tan/llama.cpp-tq3](https://github.com/turbo-tan/llama.cpp-tq3).
> Stock `llama.cpp` builds without TurboQuant support cannot load this model.
This is the standard 35B model, not an MTP release. Do not add draft-MTP speculative-decoding flags.
## Files
- [`Ornith-1.0-35B-TQ3_4S.gguf`](Ornith-1.0-35B-TQ3_4S.gguf) - main model, 13.30 GB (12.39 GiB)
- [`thumbnail.png`](thumbnail.png) - model card banner
- [`benchmark.png`](benchmark.png) - benchmark comparison card
- [`benchmark_notes.md`](benchmark_notes.md) - compact benchmark provenance and comparison
## Build the Required Runtime
```bash
git clone https://github.com/turbo-tan/llama.cpp-tq3
cd llama.cpp-tq3
cmake -S . -B build \
-DCMAKE_BUILD_TYPE=Release \
-DGGML_CUDA=ON \
-DGGML_CUDA_FA_ALL_QUANTS=OFF \
-DGGML_CUDA_GRAPHS=ON
cmake --build build --target llama-server -j
```
For an RTX 3090, `-DCMAKE_CUDA_ARCHITECTURES=86` may be added explicitly. Use the architecture matching your GPU on other systems.
## Recommended Runtime
Validated on an NVIDIA GeForce RTX 3090 Founders Edition with 24 GiB VRAM:
```bash
./build/bin/llama-server \
-m Ornith-1.0-35B-TQ3_4S.gguf \
--alias Ornith-1.0-35B-TQ3_4S \
--host 127.0.0.1 --port 8080 \
-c 32768 -np 1 -ngl 99 -fa on \
-ctk q8_0 -ctv tq3_0 \
--reasoning off --jinja
```
Runtime notes:
- `-fa on` enables flash attention at runtime.
- The validated CUDA build uses `GGML_CUDA_FA_ALL_QUANTS=OFF`.
- `-ngl 99` fully offloads the model on supported GPUs. Avoid partial offload when comparing the published speed.
- Reduce context from `32768` if the available VRAM is lower than 24 GiB.
## Quick Smoke Test
```bash
curl -s http://127.0.0.1:8080/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"Ornith-1.0-35B-TQ3_4S","messages":[{"role":"user","content":"Write ONLY the word ok."}],"max_tokens":16,"temperature":0.0}'
```
Expected assistant content:
```text
ok
```
## Benchmark Summary
Local BenchLoop and Hard86 comparison on an RTX 3090 FE using the TurboTan runtime and the launch settings above:

| Metric | Result |
|---|---:|
| 35B field overall | 95.75 |
| Hard86 | 81.4% |
| EasyCode | 100.0% |
| Toolcall | 88.3% |
| Data extract | 86.5% |
| Instruct follow | 65.5% |
| Reason math | 73.3% |
| Generation speed | 146.3 tok/s |
| Size | 13.0 GB reported; 12.39 GiB file |
The field score uses `0.85 * task_score + 0.15 * size_factor`, with size normalized to the smallest displayed 35B model. Hard86 is weighted at `2x`, EasyCode at `0.5x`, and the remaining benchmark categories at `1x`.
## Validation Notes
- Benchmark results are local measurements, not claims from the parent model repository.
## License
Use is subject to the [base model](https://huggingface.co/deepreinforce-ai/Ornith-1.0-35B) license and the licenses of [turbo-tan/llama.cpp-tq3](https://github.com/turbo-tan/llama.cpp-tq3).
|