tacodevs's picture
Add files using upload-large-folder tool
a9b4d1c verified
|
Raw History Blame Contribute Delete
1.01 kB
metadata
license: other
license_name: mistral-ai-research-license
license_link: https://mistral.ai/licenses/MNPL-0.1.md
base_model:
  - tacodevs/Behemoth-T1-123B
tags:
  - mistral
  - mistral-large
  - 123b
  - roleplay
  - nvfp4
  - fp4
  - w4a4
  - quantized
  - llm-compressor
  - compressed-tensors
  - blackwell

Behemoth-T1-123B — NVFP4 (FP4 weights + FP4 activations, Blackwell)

NVFP4 of tacodevs/Behemoth-T1-123B: FP4 (e2m1) weights and activations in groups of 16 with FP8 e4m3 scales (compressed-tensors nvfp4-pack-quantized, scheme NVFP4, static minmax activation scales). Built 2026-09-04 with llm-compressor 0.13.0 (QuantizationModifier, RTN + calibrated scales, no GPTQ), 128 character-card roleplay conversations (Gryphe/Sonnet3.5-Charcard-Roleplay, 8192 max tokens).

Target hardware: Blackwell (SM100/SM120 — B200, RTX PRO 6000) where vLLM uses native FP4 tensor cores; on Hopper vLLM falls back to emulation (correct but not faster). Quality numbers (KL vs BF16) are recorded in the model card once measured.