Behemoth-T1-123B โ€” NVFP4 (FP4 weights + FP4 activations, Blackwell)

NVFP4 of tacodevs/Behemoth-T1-123B: FP4 (e2m1) weights and activations in groups of 16 with FP8 e4m3 scales (compressed-tensors nvfp4-pack-quantized, scheme NVFP4, static minmax activation scales). Built 2026-09-04 with llm-compressor 0.13.0 (QuantizationModifier, RTN + calibrated scales, no GPTQ), 128 character-card roleplay conversations (Gryphe/Sonnet3.5-Charcard-Roleplay, 8192 max tokens).

Target hardware: Blackwell (SM100/SM120 โ€” B200, RTX PRO 6000) where vLLM uses native FP4 tensor cores; on Hopper vLLM falls back to emulation (correct but not faster). Quality numbers (KL vs BF16) are recorded in the model card once measured.

Downloads last month
30
Safetensors
Model size
69B params
Tensor type
BF16
ยท
U8
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for tacodevs/Behemoth-T1-123B-NVFP4