Commit History

remove single-card section pending further evaluation
dd8bb26
verified

Capicua25x commited on

single-card section: only path D (fp8 KV, 104k window) shown; llama.cpp mentioned briefly as usable-but-not-recommended, not a parallel option
532de84
verified

Capicua25x commited on

fix: stray single tildes were rendering as unintended strikethrough on HF markdown β€” replaced approx-marks with β‰ˆ throughout prose
b498936
verified

Capicua25x commited on

16GB correction: the 27B does not fit in ANY runtime at usable context (Q4 weights leave no KV room) β€” recommend a smaller model, not offload hopes
8779eaf
verified

Capicua25x commited on

Single-card section: honest positioning vs llama.cpp β€” chat rig -> llama.cpp GGUF; small multi-user API serving -> this path
f5d5ac5
verified

Capicua25x commited on

Single-card max-context variant: fp8 KV doubles the pool to 113.6K -> 104k window on one 32GB card, essay throughput improves; gsm8k smoke noted
167c266
verified

Capicua25x commited on

Single-card section: TP1 on 1x R9700 measured (v4) β€” serve command, table, 4-user ceiling; explicit 16GB does-not-fit note
f14a2bc
verified

Capicua25x commited on

Honest v4 table for THIS model (arm C): equal to FP8 within noise at realistic 6k context; FP8 edge is compute-bound shapes only
a9f717d
verified

Capicua25x commited on

Measurement update (bench v4): replay confound, honest absolutes; relative rc gains stand
868ea0e
verified

Capicua25x commited on

card: Terminal-Bench hard lands β€” local B 0.341 vs cloud ref 0.273, same 3600s clock
c8a0b18
verified

Capicua25x commited on

card: MMLU-Pro reference lands (bf16 cloud 0.804 vs local B 0.817)
440110c
verified

Capicua25x commited on

card: τ² airline 0.86 / retail 0.82 + MMLU-Pro 0.817 (n=1120) land β€” B's class-A row complete
832a6da
verified

Capicua25x commited on

card: τ² B lands official β€” 0.94 (107/114, thinking ON, full clean run; reference parity)
20b810a
verified

Capicua25x commited on

card: Thinking-mode recipe β€” chat_template_kwargs.enable_thinking, reasoning_effort silently ignored by vLLM (τ² 0.90β†’0.63 trap), preflight check
ae401c0
verified

Capicua25x commited on

card: RETRACT τ² B 0.65 β€” harness passed an inert reasoning knob, cell measured no-think; think-ON rerun in flight
8fa7a9e
verified

Capicua25x commited on

card: τ² B lands at 0.65 (from-scratch clean pass, post-campaign window)
ed70191
verified

Capicua25x commited on

card: drop the GSM8K seed-repeat row (noise-band note covers it)
21dafc0
verified

Capicua25x commited on

card: retract τ² B cell pending from-scratch rerun in a clean window (B back to ⏳)
92081df
verified

Capicua25x commited on

card: HLE judged (sol) β€” ref 0.30 / C 0.25 / B 0.275, triangle within noise
e38fce2
verified

Capicua25x commited on

card: τ²-telecom B cell lands (0.63 first clean pass, load caveat footnote); pin quick start + links to rc10
639451b
verified

Capicua25x commited on

card: community credit in the rc10 banner (concurrency here, single-stream insight theirs)
4be0b6f
verified

Capicua25x commited on

card: rc10 performance-upgrade callout (+12-28% this model, +34% FP8 arm)
767304d
verified

Capicua25x commited on

card: GSM8K think newest cells (rc10)
96eaab2
verified

Capicua25x commited on

card: C throughput refreshed on rc10
974f394
verified

Capicua25x commited on

card: B throughput refreshed on in-tree tuned GEMM configs
f57b928
verified

Capicua25x commited on

card: AA-LCR shows both seed runs (0.77Β·0.81)
850d73f
verified

Capicua25x commited on

card: canonical benchmark labels
bd02f9a
verified

Capicua25x commited on

card: IFEval cells promoted to newest runs (C 0.95Β·0.91); rule text updated to most-recent-run
be0ca0c
verified

Capicua25x commited on

card: B LCR corrected to 0.81 (B-final-tagged run; earlier 0.77 was the pre-final arm)
88739e6
verified

Capicua25x commited on

card: B first-run GPQA-D 0.85 and AIME25 0.97 published
2bdf06c
verified

Capicua25x commited on

card: link the public benchmark harness (Capicua25x/modelbench)
8517cfa
verified

Capicua25x commited on

card: simplified to the release format β€” quick start, paired performance (short/6k), first-run accuracy table, recipe
8711186
verified

Capicua25x commited on

Prove the reference's HLE empties are endpoint failures, not refusals
ff88997
verified

Capicua25x commited on

Quality: land HLE + Terminal-Bench, use the reference's HLE re-run, flag TB as time-limited
017478f
verified

Capicua25x commited on

Correct the correction: the old 6k figures were mislabelled, not unsourced
16605fb
verified

Capicua25x commited on

Rebuild the throughput tables from traceable logs; withdraw the untraceable 6k figures
28d93c9
verified

Capicua25x commited on

Correct the fp8-KV throughput column: two confounds, and the block-aligned numbers
1acd12e
verified

Capicua25x commited on

Point serving at rc7 and document that fp8 KV needs the fp8-query path
f9b6cb6
verified

Capicua25x commited on

Add calibrated KV-cache scales (51 F32 scalars, weights unchanged)
dcaf079
verified

Capicua25x commited on

tau2 retail lands at 0.767: this build is behind on 2 of 3 domains, -9 items net β€” replaces the 'domain-split' framing. GPQA reference repaired to 48/60 after one of three errored docs recovered (margin +7). Withdraw the 'genuine non-answers' reading of the reference's HLE empties: no finish_reason was recorded for that run
eab3544
verified

Capicua25x commited on

name who serves the bf16 reference (AkashML via OpenRouter, pinned) in the results-table headers as well as the definitions, and scope the empty-response observation to this measurement rather than the provider generally
8e0551d
verified

Capicua25x commited on

rebalance the engineering section: state what this port added (~1,230 lines, two major version lines forward-ported, a W4xA8 kernel where the borrowed one was W8A8) alongside the credit, instead of burying it under the lineage
ea9f5f4
verified

Capicua25x commited on

add 'What it took to run this on RDNA4': the 0.26.1 port, the FP8-WMMA kernel and its shape dispatch, the spec-decode verify gate, the RDNA4 dispatch changes, and what targeting Quark actually required β€” with the lineage credited and forward-ported work marked as such
943dc70
verified

Capicua25x commited on

add a links block (GitHub branch + kernel file, Docker Hub, base model, port notes) and correct the intro: this build is level with the reference overall, not above it on every cell
a201fa3
verified

Capicua25x commited on

correct the bf16 source size (55.6 GB, not ~54) and show how 22.3 GB reconciles β€” packed U8 + retained BF16 + scales β€” since 4-bit 27B naively suggests ~14 GB
bbdc3eb
verified

Capicua25x commited on

say what each column is: the bf16 reference is a hosted third-party endpoint (AkashML via OpenRouter), not a local run, and its empty-response failures depress it and flatter everything measured against it
ffd9c80
verified

Capicua25x commited on

declare 4-bit explicitly and explain HF's auto '8-bit' tag and '16B' size: both come from MXFP4 living in U8, and both affect every MXFP4 repo including AMD's own
790d28c
verified

Capicua25x commited on

explain the '16B' sidebar figure: HF sums safetensors elements and MXFP4 packs two 4-bit weights per U8 byte, so 12.05B packed bytes = 22.7B params; the model is ~27.8B logical, unchanged from upstream
923bef0
verified

Capicua25x commited on

state benchmark provenance: cells ran on rc5 + two bind-mounted kernel files, byte-identical (sha256-verified) to the ones in the rc6 image the card tells you to pull
2bcad71
verified

Capicua25x commited on

GPQA: annotate that the reference left 3 of 60 docs unanswered β€” its answered rate is 0.8246, so this build's +8 margin is ~+5.5, not 8
9118d20
verified

Capicua25x commited on