Image-Text-to-Text
NInfer
nvfp4
fp8
qwen3_5
abliterated
multimodal
mtp
dflash2
speculative-decoding
blackwell
sm_120a
Instructions to use Dragoy/ThinkingCap-Qwen3.8-27B-abliterated-NVFP4-NInfer with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NInfer
How to use Dragoy/ThinkingCap-Qwen3.8-27B-abliterated-NVFP4-NInfer with NInfer:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Download transplant-ninfer-runtime-report.json from Dragoy/ThinkingCap-Qwen3.8-27B-abliterated-NVFP4-NInfer: direct link, hf CLI and curl.
- Browser
- Download file 8.36 kB
-
https://huggingface.co/Dragoy/ThinkingCap-Qwen3.8-27B-abliterated-NVFP4-NInfer/resolve/main/transplant-ninfer-runtime-report.json
- Command line
-
hf download hf://Dragoy/ThinkingCap-Qwen3.8-27B-abliterated-NVFP4-NInfer/transplant-ninfer-runtime-report.json
-
curl -L -o transplant-ninfer-runtime-report.json https://huggingface.co/Dragoy/ThinkingCap-Qwen3.8-27B-abliterated-NVFP4-NInfer/resolve/main/transplant-ninfer-runtime-report.json
8.36 kB
| { | |
| "status": "runtime_passed", | |
| "cases": { | |
| "text": { | |
| "stdout": "Hello!", | |
| "stderr": "e render/preprocess 350 us\ngenerate vision 0 us\ngenerate text prefill 35.5 ms\ngenerate decode 28.5 ms\ngenerate total 64.7 ms\nsummary sampling greedy (temperature 0)\nsummary finish reason stop-token\nsummary prompt tokens 18\nsummary reused prompt tokens 0\nsummary generated tokens 3\nsummary model elapsed 64.0 ms\nsummary prefill speed 506.8 tok/s\nsummary decode speed 70.1 tok/s\nsummary throughput (overall) 46.8 tok/s\nsummary device 0\nsummary max context 8192\nsummary KV capacity policy explicit\nsummary KV capacity 8192\nsummary KV page groups 128 / 128\nsummary gpu weights used 19.0 GiB / 19.0 GiB\nsummary gpu sequence used 416.7 MiB / 416.7 MiB\nsummary kv cache dtype fp8-e4m3-row256\nsummary kv cache payload 258.0 MiB\nsummary gpu workspace peak 152.6 MiB / 152.6 MiB\nsummary runtime reservation 581.3 MiB\nsummary free after weights 75.4 GiB\nsummary free after startup 74.9 GiB\nsummary KV capacity headroom 0 B\nsummary planned slack 74.9 GiB\nsummary CUDA Graph allowance 12.0 MiB\nsummary planned device total 19.5 GiB\n" | |
| }, | |
| "arithmetic": { | |
| "stdout": "7 plus 9 equals **16**.", | |
| "stderr": "e render/preprocess 348 us\ngenerate vision 0 us\ngenerate text prefill 36.6 ms\ngenerate decode 128 ms\ngenerate total 166 ms\nsummary sampling greedy (temperature 0)\nsummary finish reason stop-token\nsummary prompt tokens 20\nsummary reused prompt tokens 0\nsummary generated tokens 10\nsummary model elapsed 165 ms\nsummary prefill speed 546.5 tok/s\nsummary decode speed 70.2 tok/s\nsummary throughput (overall) 60.7 tok/s\nsummary device 0\nsummary max context 8192\nsummary KV capacity policy explicit\nsummary KV capacity 8192\nsummary KV page groups 128 / 128\nsummary gpu weights used 19.0 GiB / 19.0 GiB\nsummary gpu sequence used 416.7 MiB / 416.7 MiB\nsummary kv cache dtype fp8-e4m3-row256\nsummary kv cache payload 258.0 MiB\nsummary gpu workspace peak 152.6 MiB / 152.6 MiB\nsummary runtime reservation 581.3 MiB\nsummary free after weights 75.4 GiB\nsummary free after startup 74.9 GiB\nsummary KV capacity headroom 0 B\nsummary planned slack 74.9 GiB\nsummary CUDA Graph allowance 12.0 MiB\nsummary planned device total 19.5 GiB\n" | |
| }, | |
| "vision": { | |
| "stdout": "NIFER VISION 731;3;左侧", | |
| "stderr": " render/preprocess 8.84 ms\ngenerate vision 14.5 ms\ngenerate text prefill 60.2 ms\ngenerate decode 186 ms\ngenerate total 270 ms\nsummary sampling greedy (temperature 0)\nsummary finish reason stop-token\nsummary prompt tokens 428\nsummary reused prompt tokens 0\nsummary generated tokens 14\nsummary model elapsed 261 ms\nsummary prefill speed 7.11k tok/s\nsummary decode speed 69.9 tok/s\nsummary throughput (overall) 53.7 tok/s\nsummary device 0\nsummary max context 8192\nsummary KV capacity policy explicit\nsummary KV capacity 8192\nsummary KV page groups 128 / 128\nsummary gpu weights used 19.3 GiB / 19.3 GiB\nsummary gpu sequence used 416.7 MiB / 416.7 MiB\nsummary kv cache dtype fp8-e4m3-row256\nsummary kv cache payload 258.0 MiB\nsummary gpu workspace peak 413.3 MiB / 413.3 MiB\nsummary runtime reservation 842.0 MiB\nsummary free after weights 75.2 GiB\nsummary free after startup 74.4 GiB\nsummary KV capacity headroom 0 B\nsummary planned slack 74.3 GiB\nsummary CUDA Graph allowance 12.0 MiB\nsummary planned device total 20.1 GiB\n" | |
| }, | |
| "mtp": { | |
| "stdout": "Mars is the fourth planet from the Sun and is often referred to as the Red Planet due to its reddish appearance.", | |
| "stderr": " stop-token\nsummary prompt tokens 19\nsummary reused prompt tokens 0\nsummary generated tokens 26\nsummary model elapsed 174 ms\nsummary prefill speed 495.6 tok/s\nsummary decode speed 184.3 tok/s\nsummary throughput (overall) 149.4 tok/s\nsummary device 0\nsummary max context 8192\nsummary KV capacity policy explicit\nsummary KV capacity 8192\nsummary KV page groups 128 / 128\nsummary gpu weights used 19.7 GiB / 19.7 GiB\nsummary gpu sequence used 441.8 MiB / 441.8 MiB\nsummary kv cache dtype fp8-e4m3-row256\nsummary kv cache payload 274.3 MiB\nsummary gpu workspace peak 152.6 MiB / 152.6 MiB\nsummary runtime reservation 676.4 MiB\nsummary free after weights 74.7 GiB\nsummary free after startup 74.1 GiB\nsummary KV capacity headroom 0 B\nsummary planned slack 74.0 GiB\nsummary CUDA Graph allowance 82.0 MiB\nsummary planned device total 20.4 GiB\nsummary mtp draft window 3\nsummary mtp rounds 8\nsummary mtp fallback steps 0\nsummary mtp drafted tokens 24\nsummary mtp accepted tokens 18\nsummary mtp acceptance rate 75.0%\nsummary mtp acceptance length 3.25 tok/round\nsummary mtp accepted by pos 8,6,4\n" | |
| }, | |
| "dflash2": { | |
| "stdout": "Titan is the largest moon of Saturn and the second-largest moon in our solar system.", | |
| "stderr": "oken\nsummary prompt tokens 19\nsummary reused prompt tokens 0\nsummary generated tokens 18\nsummary model elapsed 145 ms\nsummary prefill speed 529.3 tok/s\nsummary decode speed 156.2 tok/s\nsummary throughput (overall) 124.4 tok/s\nsummary device 0\nsummary max context 8192\nsummary KV capacity policy explicit\nsummary KV capacity 8192\nsummary KV page groups 128 / 128\nsummary gpu weights used 21.4 GiB / 21.4 GiB\nsummary gpu sequence used 524.2 MiB / 524.2 MiB\nsummary kv cache dtype fp8-e4m3-row256\nsummary kv cache payload 258.0 MiB\nsummary gpu workspace peak 152.6 MiB / 152.6 MiB\nsummary runtime reservation 964.7 MiB\nsummary free after weights 73.0 GiB\nsummary free after startup 72.4 GiB\nsummary KV capacity headroom 0 B\nsummary planned slack 72.1 GiB\nsummary CUDA Graph allowance 288.0 MiB\nsummary planned device total 22.3 GiB\nsummary dflash2 draft window 7\nsummary dflash2 rounds 6\nsummary dflash2 fallback steps 0\nsummary dflash2 drafted tokens 42\nsummary dflash2 accepted tokens 12\nsummary dflash2 acceptance rate 28.6%\nsummary dflash2 acceptance length 3.00 tok/round\nsummary dflash2 accepted by pos 4,3,3,2,0,0,0\n" | |
| } | |
| }, | |
| "artifact": "qwen3_8_27b_thinkingcap_heretic_nvfp4.ninfer", | |
| "mode": "nonthinking" | |
| } |