How to use from
SGLang
Install from pip and serve model
# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
    --model-path "FreedomAISVR/Ornith-1.0-35B-NVFP4-GGUF" \
    --host 0.0.0.0 \
    --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "FreedomAISVR/Ornith-1.0-35B-NVFP4-GGUF",
		"prompt": "Once upon a time,",
		"max_tokens": 512,
		"temperature": 0.5
	}'
Use Docker images
docker run --gpus all \
    --shm-size 32g \
    -p 30000:30000 \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    --env "HF_TOKEN=<secret>" \
    --ipc=host \
    lmsysorg/sglang:latest \
    python3 -m sglang.launch_server \
        --model-path "FreedomAISVR/Ornith-1.0-35B-NVFP4-GGUF" \
        --host 0.0.0.0 \
        --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "FreedomAISVR/Ornith-1.0-35B-NVFP4-GGUF",
		"prompt": "Once upon a time,",
		"max_tokens": 512,
		"temperature": 0.5
	}'
Quick Links

Ornith 1.0 35B β€” NVFP4 GGUF

NVFP4 quantization of deepreinforce-ai/Ornith-1.0-35B, a 35B parameter Qwen3.5 MoE coding agent with 256 experts (8 active per token).

About the Model

Ornith-1.0-35B is the lightweight member of the Ornith family, designed for efficient single-GPU deployment.

  • State-of-the-Art Coding Agents: Post-trained on top of Qwen 3.5, achieving state-of-the-art performance among open-source models
  • Self-Improving Training Framework: Ornith-1.0 employs RL to learn to generate not only solution rollouts, but also the scaffold that drives those rollouts
  • 35B total parameters with 8B active per token (256 experts, 8 active)
  • 40-layer MoE architecture with sliding + full attention hybrid
  • 262K context window
  • MIT License β€” globally accessible, no regional limitations

Architecture

  • Text model: Qwen3.5 MoE β€” 40 layers, 2048 hidden, 256 experts (8 active/token)
  • Vocabulary: 248,320 tokens

Quantization

Quantized from the BF16 safetensors using llama.cpp (build 537).

NVFP4 (NVIDIA FP4) uses 4-bit floating point quantization optimized for NVIDIA Blackwell GPUs.

Files

File Size Description
ornith-1.0-35b-nvfp4.gguf ~18.4 GB NVFP4 quantized model

Usage

llama-server \
  -m ornith-1.0-35b-nvfp4.gguf \
  -ngl 99 \
  --host 0.0.0.0 \
  --port 8080

Hardware Requirements

  • Minimum: 20 GB VRAM
  • Recommended: 24+ GB VRAM for full GPU offload

License

MIT

Downloads last month
459
GGUF
Model size
35B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for FreedomAISVR/Ornith-1.0-35B-NVFP4-GGUF

Quantized
(178)
this model