Instructions to use neko-legends/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use neko-legends/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf neko-legends/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP:NVFP4 # Run inference directly in the terminal: llama cli -hf neko-legends/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP:NVFP4
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf neko-legends/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP:NVFP4 # Run inference directly in the terminal: llama cli -hf neko-legends/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP:NVFP4
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf neko-legends/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP:NVFP4 # Run inference directly in the terminal: ./llama-cli -hf neko-legends/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP:NVFP4
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf neko-legends/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP:NVFP4 # Run inference directly in the terminal: ./build/bin/llama-cli -hf neko-legends/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP:NVFP4
Use Docker
docker model run hf.co/neko-legends/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP:NVFP4
- LM Studio
- Jan
- vLLM
How to use neko-legends/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "neko-legends/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "neko-legends/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/neko-legends/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP:NVFP4
- Ollama
How to use neko-legends/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP with Ollama:
ollama run hf.co/neko-legends/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP:NVFP4
- Unsloth Desktop
- Pi
How to use neko-legends/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf neko-legends/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP:NVFP4
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "neko-legends/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP:NVFP4" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use neko-legends/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP with Docker Model Runner:
docker model run hf.co/neko-legends/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP:NVFP4
- Lemonade
How to use neko-legends/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull neko-legends/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP:NVFP4
Run and chat with the model
lemonade run user.Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP-NVFP4
List all available models
lemonade list
- Hermes Agent
How to use neko-legends/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf neko-legends/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP:NVFP4
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default neko-legends/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP:NVFP4
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use neko-legends/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf neko-legends/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP:NVFP4
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "neko-legends/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP:NVFP4" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
This repo publishes one recommended AEON-trunk MTP artifact. It was validated for text serving only; the original safetensors family is multimodal, but this GGUF card does not claim vision or multimodal serving support.
Quick Start
Use ornith-1.0-35b-aeon-ultimate-uncensored-nvfp4-gguf-mtp.gguf. It is the AEON NVFP4 GGUF with grafted compatible MTP block.
Run with a current CUDA 13.x llama.cpp build and enable --spec-type draft-mtp or the tuned draft-mtp,ngram-mod profile.
On the tested RTX 5090 machine, llama.cpp initialized MTP at full 262k context and reported BLACKWELL_NATIVE_FP4 = 1.
RTX 5090 Snapshot
RTX 5090, Windows, llama.cpp-b9267-cuda13.1, context 262144, generation 1024 tokens, temperature=0.6.
| Runtime | Prompt | Decode tok/s | Full-wall tok/s | Prompt prefill |
|---|---|---|---|---|
| Base native GGUF | 10k | 133.0 | 106.0 | 1.9s |
| AEON-trunk MTP GGUF | 10k | 131.5 | 101.5 | 2.2s |
| Base native GGUF | 200k | 82.1 | 18.9 | 41.0s |
| AEON-trunk MTP GGUF | 200k | 86.0 | 15.9 | 52.1s |
Tuning note: for the 10k prompt, draft-mtp with --spec-draft-n-max 2 reached 133.7 decode tok/s and 104.0 full-wall tok/s. The chart uses the single temp=0.6 draft-mtp,ngram-mod profile for both prompt sizes.
Files
| File | Size | Notes |
|---|---|---|
ornith-1.0-35b-aeon-ultimate-uncensored-nvfp4-gguf-mtp.gguf |
23.4 GB (21.80 GiB) | Recommended AEON Ultimate Uncensored NVFP4 trunk/body GGUF with grafted compatible MTP block |
images/aeon-ornith-windows-docker-vs-gguf.png |
RTX 5090 Windows benchmark comparison chart |
Which File Should I Use?
Use ornith-1.0-35b-aeon-ultimate-uncensored-nvfp4-gguf-mtp.gguf for the AEON Ultimate Uncensored NVFP4 GGUF with MTP serving support in llama.cpp. This repository intentionally publishes only the AEON-trunk MTP artifact.
This is a text-generation GGUF. The original safetensors model family is multimodal, but this GGUF file was validated for text serving only.
MTP Provenance
AEON's compressed-tensors checkpoint advertises mtp_num_hidden_layers = 1 in config metadata, but the downloaded model.safetensors contained no mtp, nextn, or model.layers.40 tensor names. A direct conversion with MTP metadata failed in llama.cpp because blk.40.attn_norm.weight and the rest of the MTP block were absent.
The recommended MTP file in this repo was therefore built as a graft:
- Base/trunk/body: local base-only GGUF intermediate converted from
AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4, not published in this repo - MTP block donor: 20
blk.40.*tensors froms-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF - Result metadata:
qwen35moe.block_count = 41,qwen35moe.nextn_predict_layers = 1,GGUF.tensor_count = 993
Local validation confirmed llama.cpp initializes draft-mtp successfully at full 262k context and reports BLACKWELL_NATIVE_FP4 = 1 on RTX 5090.
SHA256 for ornith-1.0-35b-aeon-ultimate-uncensored-nvfp4-gguf-mtp.gguf:
3F0545EE14ED3B01A18E794945E33FFE6876F9A3C3787316A652C6CFDE4BDDE3
Example llama.cpp Command
$LlamaServer = Join-Path "<path-to-llama.cpp-build-folder>" "llama-server.exe"
$Model = Join-Path "<path-to-model-folder>" "ornith-1.0-35b-aeon-ultimate-uncensored-nvfp4-gguf-mtp.gguf"
& $LlamaServer `
--model "$Model" `
--alias aeon-ornith-1.0-35b-nvfp4-aeon-mtp `
--host 127.0.0.1 `
--port 39199 `
--device CUDA0 `
--gpu-layers all `
--gpu-layers-draft all `
--ctx-size 262144 `
--cache-type-k q4_0 `
--cache-type-v q4_0 `
--cache-type-k-draft q4_0 `
--cache-type-v-draft q4_0 `
--flash-attn on `
--parallel 1 `
--cont-batching `
--jinja `
--metrics `
--slots `
--spec-type draft-mtp `
--spec-draft-n-max 2 `
--spec-draft-p-min 0.0
For very long prompts, draft-mtp,ngram-mod with --spec-draft-n-max 3 was the better measured high-context profile in this run.
RTX 5090 Windows Benchmark Details
| Runtime | Prompt target | Prompt tokens | Decode tok/s | Prompt prefill | Full-wall tok/s | Wall time |
|---|---|---|---|---|---|---|
| Base native GGUF | 10k | 8,905 | 133.0 | 1.9s | 106.0 | 9.7s |
| AEON-trunk MTP GGUF | 10k | 8,905 | 131.5 | 2.2s | 101.5 | 10.1s |
| AEON-trunk MTP tuned n_max=2 | 10k | 8,905 | 133.7 | 2.1s | 104.0 | 9.8s |
| Base native GGUF | 200k | 174,588 | 82.1 | 41.0s | 18.9 | 54.1s |
| AEON-trunk MTP GGUF | 200k | 174,588 | 86.0 | 52.1s | 15.9 | 64.5s |
Censorship Smoke Test
A short local smoke test against ornith-1.0-35b-aeon-ultimate-uncensored-nvfp4-gguf-mtp.gguf on 2026-06-28 asked for neutral factual summaries of politically sensitive history/current-affairs topics. The model returned direct factual answers with no refusal or evasion markers detected. This is a small smoke test, not a formal safety or truthfulness evaluation.
Source And Credits
- AEON NVFP4 source:
AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4 - MTP block donor:
s-batman/Ornith-1.0-35B-NVFP4-MTP-GGUF - Base lineage:
deepreinforce-ai/Ornith-1.0-35B - Local benchmark/launcher work:
neko-legends/nvidia-local-llm-profiles
Responsible Use
This is an uncensored/abliterated model family. You are responsible for downstream usage, deployment policy, and any application-level safeguards. Older llama.cpp builds may not load current GGUF/NVFP4 files correctly.
- Downloads last month
- 252
4-bit
Model tree for neko-legends/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP
Base model
ornith-ai/Ornith-1.0-35B