Neko Legends local inference release

Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP

RTX 5090 validated

A text-generation GGUF package for recent llama.cpp builds: AEON Ultimate Uncensored NVFP4 trunk/body weights with a compatible MTP block, ready for Blackwell native FP4 local serving.

FormatGGUF
QuantNVFP4
Spec decodedraft-mtp
Validated ctx262k
Target stackllama.cpp
Artifact23.4 GB

This repo publishes one recommended AEON-trunk MTP artifact. It was validated for text serving only; the original safetensors family is multimodal, but this GGUF card does not claim vision or multimodal serving support.

Quick Start

Download

Use ornith-1.0-35b-aeon-ultimate-uncensored-nvfp4-gguf-mtp.gguf. It is the AEON NVFP4 GGUF with grafted compatible MTP block.

Serve

Run with a current CUDA 13.x llama.cpp build and enable --spec-type draft-mtp or the tuned draft-mtp,ngram-mod profile.

Expect

On the tested RTX 5090 machine, llama.cpp initialized MTP at full 262k context and reported BLACKWELL_NATIVE_FP4 = 1.

RTX 5090 Snapshot

RTX 5090, Windows, llama.cpp-b9267-cuda13.1, context 262144, generation 1024 tokens, temperature=0.6.

Runtime Prompt Decode tok/s Full-wall tok/s Prompt prefill
Base native GGUF 10k 133.0 106.0 1.9s
AEON-trunk MTP GGUF 10k 131.5 101.5 2.2s
Base native GGUF 200k 82.1 18.9 41.0s
AEON-trunk MTP GGUF 200k 86.0 15.9 52.1s

Tuning note: for the 10k prompt, draft-mtp with --spec-draft-n-max 2 reached 133.7 decode tok/s and 104.0 full-wall tok/s. The chart uses the single temp=0.6 draft-mtp,ngram-mod profile for both prompt sizes.

Windows native GGUF and MTP benchmark chart
AEON Ornith Ultimate Uncensored NVFP4 Windows Docker vs native GGUF benchmark chart

Files

File Size Notes
ornith-1.0-35b-aeon-ultimate-uncensored-nvfp4-gguf-mtp.gguf 23.4 GB (21.80 GiB) Recommended AEON Ultimate Uncensored NVFP4 trunk/body GGUF with grafted compatible MTP block
images/aeon-ornith-windows-docker-vs-gguf.png RTX 5090 Windows benchmark comparison chart

Which File Should I Use?

Use ornith-1.0-35b-aeon-ultimate-uncensored-nvfp4-gguf-mtp.gguf for the AEON Ultimate Uncensored NVFP4 GGUF with MTP serving support in llama.cpp. This repository intentionally publishes only the AEON-trunk MTP artifact.

This is a text-generation GGUF. The original safetensors model family is multimodal, but this GGUF file was validated for text serving only.

MTP Provenance

AEON's compressed-tensors checkpoint advertises mtp_num_hidden_layers = 1 in config metadata, but the downloaded model.safetensors contained no mtp, nextn, or model.layers.40 tensor names. A direct conversion with MTP metadata failed in llama.cpp because blk.40.attn_norm.weight and the rest of the MTP block were absent.

The recommended MTP file in this repo was therefore built as a graft:

Local validation confirmed llama.cpp initializes draft-mtp successfully at full 262k context and reports BLACKWELL_NATIVE_FP4 = 1 on RTX 5090.

SHA256 for ornith-1.0-35b-aeon-ultimate-uncensored-nvfp4-gguf-mtp.gguf:

3F0545EE14ED3B01A18E794945E33FFE6876F9A3C3787316A652C6CFDE4BDDE3

Example llama.cpp Command

$LlamaServer = Join-Path "<path-to-llama.cpp-build-folder>" "llama-server.exe"
$Model = Join-Path "<path-to-model-folder>" "ornith-1.0-35b-aeon-ultimate-uncensored-nvfp4-gguf-mtp.gguf"

& $LlamaServer `
  --model "$Model" `
  --alias aeon-ornith-1.0-35b-nvfp4-aeon-mtp `
  --host 127.0.0.1 `
  --port 39199 `
  --device CUDA0 `
  --gpu-layers all `
  --gpu-layers-draft all `
  --ctx-size 262144 `
  --cache-type-k q4_0 `
  --cache-type-v q4_0 `
  --cache-type-k-draft q4_0 `
  --cache-type-v-draft q4_0 `
  --flash-attn on `
  --parallel 1 `
  --cont-batching `
  --jinja `
  --metrics `
  --slots `
  --spec-type draft-mtp `
  --spec-draft-n-max 2 `
  --spec-draft-p-min 0.0

For very long prompts, draft-mtp,ngram-mod with --spec-draft-n-max 3 was the better measured high-context profile in this run.

RTX 5090 Windows Benchmark Details

Runtime Prompt target Prompt tokens Decode tok/s Prompt prefill Full-wall tok/s Wall time
Base native GGUF 10k 8,905 133.0 1.9s 106.0 9.7s
AEON-trunk MTP GGUF 10k 8,905 131.5 2.2s 101.5 10.1s
AEON-trunk MTP tuned n_max=2 10k 8,905 133.7 2.1s 104.0 9.8s
Base native GGUF 200k 174,588 82.1 41.0s 18.9 54.1s
AEON-trunk MTP GGUF 200k 174,588 86.0 52.1s 15.9 64.5s

Censorship Smoke Test

A short local smoke test against ornith-1.0-35b-aeon-ultimate-uncensored-nvfp4-gguf-mtp.gguf on 2026-06-28 asked for neutral factual summaries of politically sensitive history/current-affairs topics. The model returned direct factual answers with no refusal or evasion markers detected. This is a small smoke test, not a formal safety or truthfulness evaluation.

Source And Credits

Responsible Use

This is an uncensored/abliterated model family. You are responsible for downstream usage, deployment policy, and any application-level safeguards. Older llama.cpp builds may not load current GGUF/NVFP4 files correctly.

Downloads last month
252
GGUF
Model size
36B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for neko-legends/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP