Instructions to use yunicro/MiDM-9B-q35-e1-bx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use yunicro/MiDM-9B-q35-e1-bx with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
|
Download benchmarks/v0.2.1/README.md from yunicro/MiDM-9B-q35-e1-bx: direct link, hf CLI and curl.
- Browser
- Download file 10.3 kB
-
https://huggingface.co/yunicro/MiDM-9B-q35-e1-bx/resolve/main/benchmarks/v0.2.1/README.md
- Command line
-
hf download hf://yunicro/MiDM-9B-q35-e1-bx/benchmarks/v0.2.1/README.md
-
curl -L -o README.md https://huggingface.co/yunicro/MiDM-9B-q35-e1-bx/resolve/main/benchmarks/v0.2.1/README.md
10.3 kB
| # Benchmark coverage, architecture and parameter-performance atlas | |
| **v0.2.1 documentation and analysis update. Model weights are unchanged from v0.2.0.** All figures below are derived from completed runs or identified published outcomes. No new model inference, external paid model calls or financial results are implied. | |
| ## What fills the previous blanks | |
| |Benchmark / metric|MiDM 4B|MiDM 9B|Comparison and status| | |
| |---|---:|---:|---| | |
| |Typed Decisions test (N=2000)|1577/2000 (78.85%)|1586/2000 (79.30%)|Matched local measurement| | |
| |Kev transfer test (N=764)|610/764 (79.84%)|631/764 (82.59%)|Matched local measurement| | |
| |JevBench public accuracy (N=231)|165/231 (71.43%)|178/231 (77.06%)|Matched local measurement| | |
| |JevBench hard accuracy (N=111)|51/111 (45.95%)|60/111 (54.05%)|Matched local measurement| | |
| |DeepSWE held-out 38, best-of-4|Not measured|Not measured|Released CLM head replay: **31/38 = 81.58%**. MiDM needs the same raw state/action texts; cached CLM embeddings cannot be used as MiDM inputs.| | |
| |Terminal-Bench verifier, reported 30 tasks|Not measured|Not measured|CLM author reports 87.6%; matching candidate pool, head and full scoring protocol not found. Not a local result and not necessarily task-level integer accuracy.| | |
| |Official JevBench composite|Not measured|Not measured|Public accuracy cannot replace sealed decisions, distribution calibration, controlled latency and cost axes.| | |
| A blank score is not zero. Missing measurements remain explicit and are not plotted as zero or copied from another model. The [scoped upstream inventory](upstream_inventory.json) rechecks official GitHub and Hugging Face releases; it does not claim an exhaustive search of the internet. A fresh Terminal-Bench agent rollout would be a different experiment from the released best-of-N verifier comparison. | |
| ## Architecture: what is being compared | |
|  | |
| |System|Input interaction and output|Learned components|Parameter disclosure| | |
| |---|---|---|---| | |
| |Released CLM 8B|State and each candidate encoded independently; two projection MLPs; scaled cosine similarity|Frozen Qwen3-8B encoder plus learned state/action heads|Nominal 8B encoder; task-specific head differs from reference head| | |
| |Local CLM-style 4B control|Independent encoded state/action features; projection heads|Local reimplementation, Qwen3-4B frozen|Nominal 4B; **not** released CLM 8B| | |
| |MiDM 4B / 9B|State, question and option lines in one causal sequence; shared linear pointer at each option end; softmax over options|Frozen NF4 base + rank-16 LoRA + pointer|Nominal 4B / 9B base; exact adapter/head counts below| | |
| |Qwen3 4B reasoning|Generates a token sequence and parses a selected answer|Existing instruction/reasoning model|Nominal 4B; more inference work than one-pass selection| | |
| |Jev / Kev|Typed decision interface; benchmark-author outcomes available|Exact architecture not established by this local audit|Kev nominal 4B/8B labels; Jev size unknown here| | |
| |Solar Pro4|Generative API answer parsed to a choice|Proprietary service; internals not audited|Unknown here; omitted from parameter-axis plots| | |
| **Causal attention matters:** a later option can see preceding options, but an earlier option-end state cannot see later options. MiDM is not a bidirectional all-options cross-encoder. Option-order sensitivity and context truncation therefore remain relevant. CLM can cache state/action embeddings independently; MiDM option scores generally require the joint sequence to be recomputed. Neither parameter count alone nor adapter file size measures serving latency. | |
| |Package|Nominal base|LoRA parameters (exact)|Pointer parameters (exact)|Trainable components total| | |
| |---|---:|---:|---:|---:| | |
| |MiDM 4B|4B|30,474,240|2,561|30,476,801| | |
| |MiDM 9B|9B|40,108,032|4,097|40,112,129| | |
| Counts are computed from shipped safetensors shapes, not estimated from file bytes. Base labels are nominal and are not exact instantiated text/vision/LM-head parameter counts. The base model is downloaded separately; a small adapter does not make the full model equally small. | |
| ## Parameters versus performance | |
|  | |
| The primary figure contains only the matched 4B/9B run. With 2.25 times the nominal base parameters, 9B improves Typed Decisions by **0.45 pp**, Kev transfer test by **2.75 pp**, JevBench public by **5.63 pp**, and its hard tier by **8.11 pp**. These retrospective point differences do not establish a general scaling law or statistical superiority. | |
|  | |
| The contextual plot keeps historical local and benchmark-published rows visibly separate. Historical bases, precision, epoch count, training mix and execution paths differ. It is not a matched architecture or size ablation. Systems with unknown parameter counts are excluded from this x-axis, rather than assigned a guessed size. | |
| ## Same public items, different inference protocols | |
|  | |
| |System|Nominal B|Public correct / 231|Public accuracy|Hard accuracy|Evidence class| | |
| |---|---:|---:|---:|---:|---| | |
| |MiDM 4B (matched)|4|165/231|71.43%|45.95%|local matched BF16/NF4| | |
| |MiDM 9B (matched)|9|178/231|77.06%|54.05%|local matched BF16/NF4| | |
| |CLM-style 4B control|4|95/231|41.13%|36.04%|historical local; configuration differs| | |
| |Frozen joint pointer 4B|4|109/231|47.19%|36.94%|historical local; configuration differs| | |
| |MiDM Qwen3 4B|4|159/231|68.83%|39.64%|historical local; configuration differs| | |
| |MiDM Qwen3 8B (2ep)|8|163/231|70.56%|42.34%|historical local; configuration differs| | |
| |MiDM Qwen3.5 2B|2|147/231|63.64%|37.84%|historical local; configuration differs| | |
| |Kev 4B (published)|4|153/231|66.23%|36.94%|benchmark-author per-item outcomes, JevBench v1.2 snapshot| | |
| |Kev 8B (published)|8|165/231|71.43%|45.05%|benchmark-author per-item outcomes, JevBench v1.2 snapshot| | |
| |Jev 1.13 (published)|Unknown|200/231|86.58%|72.97%|benchmark-author per-item outcomes, JevBench v1.2 snapshot| | |
| |Jev 1.13 Free (local API run)|Unknown|194/231|83.98%|67.57%|local API responses; different prompting/compute from pointer models| | |
| |Solar Pro4 (local API run)|Unknown|220/231|95.24%|90.99%|local API responses; different prompting/compute from pointer models| | |
| |Qwen3 4B reasoning|4|189/231|81.82%|65.77%|local reasoning run; different token budget and output path| | |
| Jev published 200/231 and the local Jev Free run 194/231 are distinct service/protocol observations; neither overwrites the other. Solar 220/231 uses a generative API. They are useful capability references, not equal-compute or equal-cost experiments. The incomplete Space Bunny cohort (162/231 in the inspected snapshot) is excluded from the full-cohort graph; its 159/162 must not be ranked against complete 231-item results. This update makes no fresh calls to these model APIs. | |
| ## Architecture control and limits | |
|  | |
| The Qwen3-4B control series holds the nominal base family/size and original task mix, but training recipes and precision details still differ. Kev transfer accuracy rises from 41.10% (independent frozen heads) to 62.96% (frozen causal pointer) to 71.60% (QLoRA pointer). This supports candidate-context interaction as a useful design direction, but does not isolate every causal contribution. It must not be confused with a direct comparison against the released CLM 8B head. | |
| ## SQL/Python and deployment tradeoffs | |
| |Same stored candidate pool|4B|9B|Interpretation| | |
| |---|---:|---:|---| | |
| |Spider test, RAG5 + dev-selected blend|78.34%|77.27%|Regression| | |
| |SQL holdout, RAG5 + dev-selected blend|81.66%|81.77%|Small point increase| | |
| |Python hard80, RAG0|28.75%|25.00%|Regression| | |
| These are best-of-N selection results, not direct code generation. [Full SQL/Python protocol](https://github.com/YeoHoonYun/midm-decision-models/blob/v0.2.1/docs/evaluations/20261005/SQL_PYTHON_9B.md). The learned router reaches 79.65% public accuracy, below local reasoning alone at 81.82%; it is not promoted as the default. Long-context two-process pilots reduce throughput despite fitting in VRAM. [Serving and routing measurements](https://github.com/YeoHoonYun/midm-decision-models/blob/v0.2.1/docs/evaluations/20261005/README.md). | |
| ## Interpretation | |
| 9B improves several bounded-choice benchmarks and especially the public hard tier, but it does not dominate the generative comparison models or improve every application. The 4B model remains a smaller one-pass option scorer; the 9B model trades a larger base for higher decision accuracy on these suites. Unknown proprietary sizes, different inference budgets, seen public data and historical recipe differences prevent a universal performance-per-parameter ranking. | |
| All outcomes are retrospective. Public-item accuracy is not an official composite score; trainable parameters are not full-model parameters; nominal size is not VRAM; batched throughput is not single-request latency. No financial inputs, probabilities, generated reports, raw benchmark examples or credentials are included. | |
| ## Sources and reproduction | |
| [Machine-readable aggregates and source hashes](atlas.json) 路 [Plotting script](plot_atlas.py) 路 [Upstream asset inventory](upstream_inventory.json). Run `python plot_atlas.py` with Python, NumPy and Matplotlib to reproduce PNG/SVG plots from `atlas.json`. | |
| - [CLM source and verifier protocol](https://github.com/Contrastive-LM/CLM); [projection implementation](https://github.com/Contrastive-LM/CLM/blob/main/src/clm/heads.py). | |
| - [Released CLM DeepSWE heads](https://huggingface.co/Contrastive-LM/deepswe-clm-heads-8k). | |
| - [JevBench v1.2 per-task source](https://github.com/fstandhartinger/jevbench/blob/main/results/v1.2/jevbench-v1.2-per-task.json). The archived local snapshot, not an assumed current leaderboard, underlies these figures. | |
| - [MiDM primary evaluation](https://github.com/YeoHoonYun/midm-decision-models/blob/v0.2.1/docs/evaluations/20261005/README.md), [four-benchmark rescore](https://github.com/YeoHoonYun/midm-decision-models/blob/v0.2.1/docs/evaluations/20261005/four-benchmarks/README.md), and previously published local aggregate controls in this repository. | |