Behemoth-T1-123B โ NVFP4 (FP4 weights + FP4 activations, Blackwell)
NVFP4 of tacodevs/Behemoth-T1-123B: FP4 (e2m1) weights and activations in groups of 16 with FP8 e4m3 scales
(compressed-tensors nvfp4-pack-quantized, scheme NVFP4, static minmax activation scales). Built 2026-09-04 with
llm-compressor 0.13.0 (QuantizationModifier, RTN + calibrated scales, no GPTQ), 128 character-card roleplay
conversations (Gryphe/Sonnet3.5-Charcard-Roleplay, 8192 max tokens).
Target hardware: Blackwell (SM100/SM120 โ B200, RTX PRO 6000) where vLLM uses native FP4 tensor cores; on Hopper vLLM falls back to emulation (correct but not faster). Quality numbers (KL vs BF16) are recorded in the model card once measured.
- Downloads last month
- 30
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support
Model tree for tacodevs/Behemoth-T1-123B-NVFP4
Base model
mistralai/Mistral-Large-Instruct-2411 Finetuned
TheDrummer/Behemoth-R1-123B-v2 Adapter
tacodevs/Behemoth-T1-123B