YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Qwen3.8-Flash-Next bf16 reference logits
Ground truth for Splash quality checks, from
dev/benchmarks/qwen4exp/reference_logits.py (full-precision checkpoint,
float32 on CPU). Each <name>.reference.bin is bf16 logits [tokens x 248320]
for <name>.passage.json.
- code: 1,495-token code prompt; short: 65 tokens; long4k: first 4,096 tokens of the long prompt (crosses the 2,048-token sparse-indexer budget).
- code.plan.bin: the reference run with the current package's quantization
simulated (
--quant experts:4,router:8,shared:4,attn:8,gdn_in:8,gdn_ab:8, gdn_out:8,ple:4,head:8,embed:8): 90.2% same pick, KL 0.122.
Score an engine dump:
SPLASH_DUMP_PREFILL_LOGITS=/tmp/x.bin build/engine-tests/generate-sample
build/splash.metallib ~/models/qwen38-flash-next-splash 1 ""
python dev/benchmarks/qwen4exp/compare_logits.py --passages code.passage.json
--name code --reference code.reference.bin --engine /tmp/x.bin [--begin N --end M]
(Use ~/vllm-mlx-env-3.12/bin/python for the reference and compare scripts.)
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support