YuYu1015-Ornith-1.0-35B-abliterated-NVFP4

English | 繁體中文

NVFP4 (NVIDIA FP4) quant of YuYu1015/YuYu1015-Ornith-1.0-35B-abliterated (the BF16 source) · also available as GGUF

Support me on Ko-fi


English

NVFP4 quant of the abliterated (uncensored) Qwen3.5 35B MoE reasoning model, produced with NVIDIA TensorRT Model Optimizer (nvfp4_mlp_only, MSE calibration on reasoning data). For high-throughput low-precision inference on NVIDIA Blackwell GPUs.

Item Value
Format NVFP4 (E2M1 + 2-level scale, group size 16), modelopt
Size ~23.7 GB (3 shards)
Quantized routed experts + shared-expert MLP → NVFP4
Kept BF16 attention, GatedDeltaNet linear-attn, MoE router, shared_expert_gate, lm_head, embeddings
KV-cache not quantized (BF16)

This mixed-precision recipe follows NVIDIA's higher-accuracy guidance for FP4 PTQ — the routed experts (the bulk of the weights) go to FP4 while every precision-sensitive tensor stays in BF16, so reasoning quality is preserved.

Requirements

NVFP4 needs Blackwell (sm_120 / sm_100, e.g. RTX PRO 6000 / B200) and a runtime with modelopt-NVFP4 support (SGLang or vLLM). It will not load with plain transformers (use a serving runtime).

Usage

SGLang:

python -m sglang.launch_server --model-path YuYu1015/YuYu1015-Ornith-1.0-35B-abliterated-NVFP4 --quantization modelopt_fp4 --trust-remote-code

vLLM:

vllm serve YuYu1015/YuYu1015-Ornith-1.0-35B-abliterated-NVFP4 --trust-remote-code

Recommended Sampling Parameters

Reasoning model (emits <think>…</think>). Official Qwen3.5 settings:

temperature 1.0 · top_p 0.95 · top_k 20 · min_p 0.0 · presence_penalty 0.0

Safety Warning

This model has safety filtering removed (abliterated) and may generate sensitive or inappropriate content. Users are solely responsible for all consequences and legal liability, and must ensure usage complies with local laws and ethical standards.

Credits


繁體中文

YuYu1015/YuYu1015-Ornith-1.0-35B-abliterated(BF16 來源)的 NVFP4(NVIDIA FP4)量化版本;另有 GGUF 版本

NVIDIA TensorRT Model Optimizer(nvfp4_mlp_only、推理資料 MSE 校準)量化的 abliterated(去審查)Qwen3.5 35B MoE 推理模型,供 NVIDIA Blackwell GPU 高吞吐低精度推理。

項目 數值
格式 NVFP4(E2M1 + 兩級縮放,group size 16),modelopt
大小 ~23.7 GB(3 shards)
量化 routed experts + shared-expert MLP → NVFP4
保 BF16 attention、GatedDeltaNet linear-attn、MoE router、shared_expert_gatelm_head、embedding
KV-cache 不量化(BF16)

此混合精度配方遵循 NVIDIA 對 FP4 PTQ 的高精度建議 —— 只把 routed experts(佔大多數參數)壓到 FP4,所有對精度敏感的張量全保 BF16,以保留推理能力。

需求

NVFP4 需 Blackwell(sm_120 / sm_100,如 RTX PRO 6000 / B200) 及支援 modelopt-NVFP4 的 runtime(SGLangvLLM)。無法用純 transformers 載入(請用推理 runtime)。

使用方式

SGLang:

python -m sglang.launch_server --model-path YuYu1015/YuYu1015-Ornith-1.0-35B-abliterated-NVFP4 --quantization modelopt_fp4 --trust-remote-code

vLLM:

vllm serve YuYu1015/YuYu1015-Ornith-1.0-35B-abliterated-NVFP4 --trust-remote-code

建議取樣參數

推理模型(輸出 <think>…</think>)。Qwen3.5 官方設定:

temperature 1.0 · top_p 0.95 · top_k 20 · min_p 0.0 · presence_penalty 0.0

安全警告

此模型已移除安全過濾(abliterated),可能產生敏感或不當內容。使用者須自行承擔所有風險與法律責任,並確保使用方式符合當地法規與倫理標準。

致謝

Downloads last month
271
Safetensors
Model size
19B params
Tensor type
BF16
·
F8_E4M3
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for YuYu1015/YuYu1015-Ornith-1.0-35B-abliterated-NVFP4

Quantized
(2)
this model