YuYu1015-Ornith-1.0-35B-abliterated-NVFP4
NVFP4 (NVIDIA FP4) quant of YuYu1015/YuYu1015-Ornith-1.0-35B-abliterated (the BF16 source) · also available as GGUF
English
NVFP4 quant of the abliterated (uncensored) Qwen3.5 35B MoE reasoning model, produced with NVIDIA TensorRT Model Optimizer (nvfp4_mlp_only, MSE calibration on reasoning data). For high-throughput low-precision inference on NVIDIA Blackwell GPUs.
| Item | Value |
|---|---|
| Format | NVFP4 (E2M1 + 2-level scale, group size 16), modelopt |
| Size | ~23.7 GB (3 shards) |
| Quantized | routed experts + shared-expert MLP → NVFP4 |
| Kept BF16 | attention, GatedDeltaNet linear-attn, MoE router, shared_expert_gate, lm_head, embeddings |
| KV-cache | not quantized (BF16) |
This mixed-precision recipe follows NVIDIA's higher-accuracy guidance for FP4 PTQ — the routed experts (the bulk of the weights) go to FP4 while every precision-sensitive tensor stays in BF16, so reasoning quality is preserved.
- BF16 source: -abliterated · GGUF: -GGUF
Requirements
NVFP4 needs Blackwell (sm_120 / sm_100, e.g. RTX PRO 6000 / B200) and a runtime with modelopt-NVFP4 support (SGLang or vLLM). It will not load with plain transformers (use a serving runtime).
Usage
SGLang:
python -m sglang.launch_server --model-path YuYu1015/YuYu1015-Ornith-1.0-35B-abliterated-NVFP4 --quantization modelopt_fp4 --trust-remote-code
vLLM:
vllm serve YuYu1015/YuYu1015-Ornith-1.0-35B-abliterated-NVFP4 --trust-remote-code
Recommended Sampling Parameters
Reasoning model (emits <think>…</think>). Official Qwen3.5 settings:
temperature 1.0 · top_p 0.95 · top_k 20 · min_p 0.0 · presence_penalty 0.0
Safety Warning
This model has safety filtering removed (abliterated) and may generate sensitive or inappropriate content. Users are solely responsible for all consequences and legal liability, and must ensure usage complies with local laws and ethical standards.
Credits
- BF16 source: YuYu1015/YuYu1015-Ornith-1.0-35B-abliterated
- Base model: deepreinforce-ai/Ornith-1.0-35B
- Author: YuYu1015
繁體中文
YuYu1015/YuYu1015-Ornith-1.0-35B-abliterated(BF16 來源)的 NVFP4(NVIDIA FP4)量化版本;另有 GGUF 版本。
以 NVIDIA TensorRT Model Optimizer(nvfp4_mlp_only、推理資料 MSE 校準)量化的 abliterated(去審查)Qwen3.5 35B MoE 推理模型,供 NVIDIA Blackwell GPU 高吞吐低精度推理。
| 項目 | 數值 |
|---|---|
| 格式 | NVFP4(E2M1 + 兩級縮放,group size 16),modelopt |
| 大小 | ~23.7 GB(3 shards) |
| 量化 | routed experts + shared-expert MLP → NVFP4 |
| 保 BF16 | attention、GatedDeltaNet linear-attn、MoE router、shared_expert_gate、lm_head、embedding |
| KV-cache | 不量化(BF16) |
此混合精度配方遵循 NVIDIA 對 FP4 PTQ 的高精度建議 —— 只把 routed experts(佔大多數參數)壓到 FP4,所有對精度敏感的張量全保 BF16,以保留推理能力。
- BF16 來源: -abliterated · GGUF: -GGUF
需求
NVFP4 需 Blackwell(sm_120 / sm_100,如 RTX PRO 6000 / B200) 及支援 modelopt-NVFP4 的 runtime(SGLang 或 vLLM)。無法用純 transformers 載入(請用推理 runtime)。
使用方式
SGLang:
python -m sglang.launch_server --model-path YuYu1015/YuYu1015-Ornith-1.0-35B-abliterated-NVFP4 --quantization modelopt_fp4 --trust-remote-code
vLLM:
vllm serve YuYu1015/YuYu1015-Ornith-1.0-35B-abliterated-NVFP4 --trust-remote-code
建議取樣參數
推理模型(輸出 <think>…</think>)。Qwen3.5 官方設定:
temperature 1.0 · top_p 0.95 · top_k 20 · min_p 0.0 · presence_penalty 0.0
安全警告
此模型已移除安全過濾(abliterated),可能產生敏感或不當內容。使用者須自行承擔所有風險與法律責任,並確保使用方式符合當地法規與倫理標準。
致謝
- BF16 來源: YuYu1015/YuYu1015-Ornith-1.0-35B-abliterated
- 基礎模型: deepreinforce-ai/Ornith-1.0-35B
- 作者: YuYu1015
- Downloads last month
- 271
Model tree for YuYu1015/YuYu1015-Ornith-1.0-35B-abliterated-NVFP4
Base model
ornith-ai/Ornith-1.0-35B