Instructions to use yunicro/MiDM-9B-q35-e1-bx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use yunicro/MiDM-9B-q35-e1-bx with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
v0.2.1 benchmark/architecture atlas; preserve model weights
Browse files- .gitattributes +5 -0
- README.md +84 -71
- benchmark_manifest_v0.2.1.json +25 -0
- benchmarks/v0.2.1/README.md +102 -0
- benchmarks/v0.2.1/architecture.png +3 -0
- benchmarks/v0.2.1/architecture.svg +353 -0
- benchmarks/v0.2.1/architecture_control.png +3 -0
- benchmarks/v0.2.1/architecture_control.svg +294 -0
- benchmarks/v0.2.1/atlas.json +587 -0
- benchmarks/v0.2.1/context_size_performance.png +3 -0
- benchmarks/v0.2.1/context_size_performance.svg +719 -0
- benchmarks/v0.2.1/matched_size_performance.png +3 -0
- benchmarks/v0.2.1/matched_size_performance.svg +774 -0
- benchmarks/v0.2.1/plot_atlas.py +94 -0
- benchmarks/v0.2.1/public_comparison.png +3 -0
- benchmarks/v0.2.1/public_comparison.svg +804 -0
- benchmarks/v0.2.1/upstream_inventory.json +80 -0
- manifest.json +33 -18
.gitattributes
CHANGED
|
@@ -33,3 +33,8 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
+
benchmarks/v0.2.1/architecture.png filter=lfs diff=lfs merge=lfs -text
|
| 37 |
+
benchmarks/v0.2.1/architecture_control.png filter=lfs diff=lfs merge=lfs -text
|
| 38 |
+
benchmarks/v0.2.1/context_size_performance.png filter=lfs diff=lfs merge=lfs -text
|
| 39 |
+
benchmarks/v0.2.1/matched_size_performance.png filter=lfs diff=lfs merge=lfs -text
|
| 40 |
+
benchmarks/v0.2.1/public_comparison.png filter=lfs diff=lfs merge=lfs -text
|
README.md
CHANGED
|
@@ -1,71 +1,84 @@
|
|
| 1 |
-
---
|
| 2 |
-
license: apache-2.0
|
| 3 |
-
base_model: Qwen/Qwen3.5-9B-Base
|
| 4 |
-
library_name: peft
|
| 5 |
-
pipeline_tag: text-classification
|
| 6 |
-
tags: [midm, decision-model, typed-decisions, pointer-head, lora, qlora]
|
| 7 |
-
---
|
| 8 |
-
|
| 9 |
-
# MiDM-9B-q35-e1-bx
|
| 10 |
-
|
| 11 |
-
|
| 12 |
-
|
| 13 |
-
|
| 14 |
-
|
| 15 |
-
|
| 16 |
-
|
| 17 |
-
|
| 18 |
-
|
| 19 |
-
|
| 20 |
-
|
| 21 |
-
|
| 22 |
-
|
| 23 |
-
|
| 24 |
-
|
| 25 |
-
|
| 26 |
-
|
| 27 |
-
|
| 28 |
-
|
| 29 |
-
|
| 30 |
-
|
| 31 |
-
|
| 32 |
-
|
| 33 |
-
|
| 34 |
-
|
| 35 |
-
|
| 36 |
-
|
| 37 |
-
|
| 38 |
-
|
| 39 |
-
|
| 40 |
-
|
| 41 |
-
|
| 42 |
-
|
| 43 |
-
|
| 44 |
-
|
| 45 |
-
|
| 46 |
-
|
| 47 |
-
|
| 48 |
-
|
| 49 |
-
|
| 50 |
-
|
| 51 |
-
|
| 52 |
-
-
|
| 53 |
-
|
| 54 |
-
|
| 55 |
-
|
| 56 |
-
|
| 57 |
-
|
| 58 |
-
|
| 59 |
-
|
| 60 |
-
|
| 61 |
-
|
| 62 |
-
|
| 63 |
-
|
| 64 |
-
|
| 65 |
-
|
| 66 |
-
|
| 67 |
-
|
| 68 |
-
|
| 69 |
-
|
| 70 |
-
|
| 71 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
base_model: Qwen/Qwen3.5-9B-Base
|
| 4 |
+
library_name: peft
|
| 5 |
+
pipeline_tag: text-classification
|
| 6 |
+
tags: [midm, decision-model, typed-decisions, pointer-head, lora, qlora]
|
| 7 |
+
---
|
| 8 |
+
|
| 9 |
+
# MiDM-9B-q35-e1-bx
|
| 10 |
+
|
| 11 |
+
## v0.2.1 benchmark and architecture analysis
|
| 12 |
+
|
| 13 |
+
Documentation update; this repository's existing weights and inference code are unchanged. [Full benchmark atlas](benchmarks/v0.2.1/README.md) · [GitHub release](https://github.com/YeoHoonYun/midm-decision-models/releases/tag/v0.2.1) · [Version DOI](https://doi.org/10.5281/zenodo.23164651).
|
| 14 |
+
|
| 15 |
+

|
| 16 |
+
|
| 17 |
+

|
| 18 |
+
|
| 19 |
+
The atlas separates matched measurements from historical/published/API comparisons. JevBench public accuracy is not its official composite. Missing MiDM DeepSWE and local Terminal-Bench measurements remain explicit; their candidate pools were not found. Unknown proprietary parameter counts are not guessed. SQL/Python regressions remain documented.
|
| 20 |
+
|
| 21 |
+

|
| 22 |
+
|
| 23 |
+
|
| 24 |
+
Version **v0.2.0**. This repository contains the LoRA adapter, a linear pointer head and a self-contained Python loader. It scores the provided options jointly in one forward pass; it does not generate SQL, Python or explanations. The API accepts `choice`, yes/no `noul`, and ordinal `score` questions. Outputs are option probabilities, not guaranteed calibrated confidence.
|
| 25 |
+
|
| 26 |
+
## Use
|
| 27 |
+
|
| 28 |
+
Install `requirements.txt`, download this repository at revision `v0.2.0`, and run from that directory:
|
| 29 |
+
|
| 30 |
+
```python
|
| 31 |
+
from midm import MiDM
|
| 32 |
+
model = MiDM.from_pretrained(".") # downloads the base separately; NF4 by default
|
| 33 |
+
result = model.predict(
|
| 34 |
+
state="A receipt is required for a refund. The customer has no receipt.",
|
| 35 |
+
questions={"action": {"type": "choice", "instructions": "Choose the next action.",
|
| 36 |
+
"criteria": {"approve": "Approve the refund.", "review": "Request a receipt or review."}}},
|
| 37 |
+
)
|
| 38 |
+
print(result)
|
| 39 |
+
```
|
| 40 |
+
|
| 41 |
+
A compatible NVIDIA GPU and the listed dependencies are required. The configured maximum input is 4096 tokens; individual options are capped at 192 tokens. Long states are truncated while retaining the head and tail. All inference can run offline once the package and base are cached. Apache-2.0 applies to this adapter and code; base model and training data retain their own terms. This repository includes no training data.
|
| 42 |
+
|
| 43 |
+
# MiDM v0.2.0 — 9B checkpoint and completed evaluation results
|
| 44 |
+
|
| 45 |
+
Adds the Qwen3.5-9B-Base LoRA adapter and learned pointer head, self-contained loader and verified 4096-token inference configuration. The original 4B release remains available. This model scores supplied options; it does not generate SQL, Python or report prose. The report writer remains Qwen3 8B.
|
| 46 |
+
|
| 47 |
+
## Matched primary evaluation
|
| 48 |
+
|
| 49 |
+
NF4/BF16 on RTX 3090, max input 4096, batch token budget 4096, TTA=1; same items and ordering.
|
| 50 |
+
|
| 51 |
+
|Suite|N|4B|9B|
|
| 52 |
+
|---|---:|---:|---:|
|
| 53 |
+
|Typed Decisions test|2000|78.85%|79.30%|
|
| 54 |
+
|Kev transfer test|764|79.84%|82.59%|
|
| 55 |
+
|JevBench original|72|91.67%|97.22%|
|
| 56 |
+
|JevBench easy|48|100.00%|100.00%|
|
| 57 |
+
|JevBench hard|111|45.95%|54.05%|
|
| 58 |
+
|
| 59 |
+
These are retrospective hard-label accuracies, not official JevBench full/sealed composite scores. The separate 600-question Typed Decisions holdout is a development/model-selection split; it is not the public 2000-question test. Earlier development results used their recorded configurations and must not be mixed into this matched table.
|
| 60 |
+
|
| 61 |
+
## Additional completed tests and limitations
|
| 62 |
+
|
| 63 |
+
- CLM released DeepSWE verifier replay: 31/38 (81.58%) reproduced. The 13 mixed-outcome tasks give exact one-sided p=0.062 against random selection. Numerical reproduction succeeded; significance at 5% was not established.
|
| 64 |
+
- MiDM DeepSWE and local Terminal-Bench remain unmeasured: matching raw candidate traces/evaluation artifacts were not found. No official benchmark score is substituted.
|
| 65 |
+
- SQL/Python best-of-N: 4B → 9B Spider 78.34% → 77.27%, SQL holdout 81.66% → 81.77%, Python hard80 28.75% → 25.00%. Stored pools, dev-selected blend weights; no consistent gain.
|
| 66 |
+
- Learned routing: JevBench public 231 accuracy 77.06% MiDM, 79.65% router, 81.82% local Qwen3-4B reasoning alone. Router not promoted; no non-inferiority or latency-saving claim.
|
| 67 |
+
- Concurrency: two short-input 4B processes nearly doubled throughput; 2048-token 4B and 4096-token 9B lost throughput. Small synthetic single-trial pilots, not endurance or minimum-memory certification.
|
| 68 |
+
- Same-date report integration: 4B/9B candidate-order agreement 1/4 versus 3/4, 19/19 schema-valid sections each, one unknown fact ID each. No semantic-quality or financial-outcome superiority established; only aggregate diagnostics are released.
|
| 69 |
+
|
| 70 |
+
Training: QLoRA NF4, rank16/alpha32, one epoch, seed0, learning rate2e-4, training context1024; td_train, kev_train, breadth_v1, pp6_new. Inference context4096 matches evaluation. The earlier 1024-token release candidate was not published; package verification uses the final loader. Base weights are separate and retain their own license, as do training datasets. No raw questions, answers, financial data, reports or credentials are included.
|
| 71 |
+
|
| 72 |
+
Model: https://huggingface.co/yunicro/MiDM-9B-q35-e1-bx/tree/v0.2.0
|
| 73 |
+
|
| 74 |
+
Code and full evidence: https://github.com/YeoHoonYun/midm-decision-models/releases/tag/v0.2.0
|
| 75 |
+
|
| 76 |
+
Zenodo versioning retains concept DOI https://doi.org/10.5281/zenodo.23084583 . The v0.1.0 version DOI does not archive this new release.
|
| 77 |
+
|
| 78 |
+
Version DOI: https://doi.org/10.5281/zenodo.23162744
|
| 79 |
+
|
| 80 |
+
## Packaged-model verification
|
| 81 |
+
|
| 82 |
+
The final package reproduces the original development evaluation exactly: **513/600 (85.50%)**, identical choices on 600/600 questions and maximum absolute probability difference 0.0. The original protocol uses NF4/BF16, 4096-token context, 4096 padded tokens per batch, attention masks and TTA=1. Adapter weights, pointer head and encoded inputs match the source exactly. This is package reproducibility, not a new independent test.
|
| 83 |
+
|
| 84 |
+
A separate single-question, unmasked check gave **509/600 (84.83%)** and failed its original tolerance gate. It is retained rather than hidden. That verification bypassed the public API batching/mask path. The matched check changes batching and mask together, so their individual effects were not isolated. Do not assume identical predictions across serving configurations.
|
benchmark_manifest_v0.2.1.json
ADDED
|
@@ -0,0 +1,25 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"documentation_version": "0.2.1",
|
| 3 |
+
"weights_unchanged": {
|
| 4 |
+
"adapter_model.safetensors": "7231ef52a482760601bd6e9038ec9aae40034d0f63c6fb5de677b5f81e6dd5c2",
|
| 5 |
+
"pointer_head.safetensors": "f534d840c1b521c04292e763db200e6ab10df12fb4ebdabc78de734c07d88e4a"
|
| 6 |
+
},
|
| 7 |
+
"files": {
|
| 8 |
+
"manifest.json": "71ca79f112c3d0a3e26b213dda6465a2be8aa6f77fc795c93a8bac377c83d605",
|
| 9 |
+
"README.md": "b392455d328c9af67f162394beeef4d1c230d5d5652fda4e94023d51af15f75d",
|
| 10 |
+
"benchmarks/v0.2.1/architecture.png": "cc5bfb7fc225b1f974837e36ef43717af271fe1d56281f6faa8790f4d1762280",
|
| 11 |
+
"benchmarks/v0.2.1/architecture.svg": "0d77746badfff06e4b07f686753830f565568e98ad6b88dfa3e6e179c413bba7",
|
| 12 |
+
"benchmarks/v0.2.1/architecture_control.png": "18e6afac826bda17ed80a7f0252f63c69109aeeca41026efbc50d01845b52746",
|
| 13 |
+
"benchmarks/v0.2.1/architecture_control.svg": "d837c9733a0e3af6c4b7c2078a940af18f3d74c6bb5871ff2231f4525c89b661",
|
| 14 |
+
"benchmarks/v0.2.1/atlas.json": "3d82078bac0024ac3fc4a814580a7fd11528c6f659acda99099e8a3703e70c4a",
|
| 15 |
+
"benchmarks/v0.2.1/context_size_performance.png": "62364f4bdfd3a00358bf2631a03d550b86d75b62b0682e936a7782d30eda5588",
|
| 16 |
+
"benchmarks/v0.2.1/context_size_performance.svg": "ef50c39dc7e19f0f289064ae41baf77889ca8fa48f656a4f92f4b9d8cd06c7c3",
|
| 17 |
+
"benchmarks/v0.2.1/matched_size_performance.png": "6e5b233ccfca43b71bd4f2884d38881751628db14369a035eac97ac5995b8729",
|
| 18 |
+
"benchmarks/v0.2.1/matched_size_performance.svg": "ea0fbc1d8c4658254c8be3a4e37094262877c404b4357e4d1e65fb2e4c07d83b",
|
| 19 |
+
"benchmarks/v0.2.1/plot_atlas.py": "2081011cfd8cef64f0f16887f5bf5df4fb1ba71c9a2da31e640e141ee7b03254",
|
| 20 |
+
"benchmarks/v0.2.1/public_comparison.png": "2ce2d9502232d5683e840afa3e97a3b203665b31b9dec6fa63d4fc3b1576aa87",
|
| 21 |
+
"benchmarks/v0.2.1/public_comparison.svg": "921ef3941ca10f90acce98bed502e95caaa143a6eb6ad3a706765cd36d3dcbfc",
|
| 22 |
+
"benchmarks/v0.2.1/README.md": "0768c6ea91f2ce2e218e268d4e7c3fb9a943e6070464976f0be57bd4555fe2b7",
|
| 23 |
+
"benchmarks/v0.2.1/upstream_inventory.json": "2515bc93354d78242e0df50aef7130ed67d19190d53384daf1fcd0746e3c1186"
|
| 24 |
+
}
|
| 25 |
+
}
|
benchmarks/v0.2.1/README.md
ADDED
|
@@ -0,0 +1,102 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Benchmark coverage, architecture and parameter-performance atlas
|
| 2 |
+
|
| 3 |
+
**v0.2.1 documentation and analysis update. Model weights are unchanged from v0.2.0.** All figures below are derived from completed runs or identified published outcomes. No new model inference, external paid model calls or financial results are implied.
|
| 4 |
+
|
| 5 |
+
## What fills the previous blanks
|
| 6 |
+
|
| 7 |
+
|Benchmark / metric|MiDM 4B|MiDM 9B|Comparison and status|
|
| 8 |
+
|---|---:|---:|---|
|
| 9 |
+
|Typed Decisions test (N=2000)|1577/2000 (78.85%)|1586/2000 (79.30%)|Matched local measurement|
|
| 10 |
+
|Kev transfer test (N=764)|610/764 (79.84%)|631/764 (82.59%)|Matched local measurement|
|
| 11 |
+
|JevBench public accuracy (N=231)|165/231 (71.43%)|178/231 (77.06%)|Matched local measurement|
|
| 12 |
+
|JevBench hard accuracy (N=111)|51/111 (45.95%)|60/111 (54.05%)|Matched local measurement|
|
| 13 |
+
|DeepSWE held-out 38, best-of-4|Not measured|Not measured|Released CLM head replay: **31/38 = 81.58%**. MiDM needs the same raw state/action texts; cached CLM embeddings cannot be used as MiDM inputs.|
|
| 14 |
+
|Terminal-Bench verifier, reported 30 tasks|Not measured|Not measured|CLM author reports 87.6%; matching candidate pool, head and full scoring protocol not found. Not a local result and not necessarily task-level integer accuracy.|
|
| 15 |
+
|Official JevBench composite|Not measured|Not measured|Public accuracy cannot replace sealed decisions, distribution calibration, controlled latency and cost axes.|
|
| 16 |
+
|
| 17 |
+
A blank score is not zero. Missing measurements remain explicit and are not plotted as zero or copied from another model. The [scoped upstream inventory](upstream_inventory.json) rechecks official GitHub and Hugging Face releases; it does not claim an exhaustive search of the internet. A fresh Terminal-Bench agent rollout would be a different experiment from the released best-of-N verifier comparison.
|
| 18 |
+
|
| 19 |
+
## Architecture: what is being compared
|
| 20 |
+
|
| 21 |
+

|
| 22 |
+
|
| 23 |
+
|System|Input interaction and output|Learned components|Parameter disclosure|
|
| 24 |
+
|---|---|---|---|
|
| 25 |
+
|Released CLM 8B|State and each candidate encoded independently; two projection MLPs; scaled cosine similarity|Frozen Qwen3-8B encoder plus learned state/action heads|Nominal 8B encoder; task-specific head differs from reference head|
|
| 26 |
+
|Local CLM-style 4B control|Independent encoded state/action features; projection heads|Local reimplementation, Qwen3-4B frozen|Nominal 4B; **not** released CLM 8B|
|
| 27 |
+
|MiDM 4B / 9B|State, question and option lines in one causal sequence; shared linear pointer at each option end; softmax over options|Frozen NF4 base + rank-16 LoRA + pointer|Nominal 4B / 9B base; exact adapter/head counts below|
|
| 28 |
+
|Qwen3 4B reasoning|Generates a token sequence and parses a selected answer|Existing instruction/reasoning model|Nominal 4B; more inference work than one-pass selection|
|
| 29 |
+
|Jev / Kev|Typed decision interface; benchmark-author outcomes available|Exact architecture not established by this local audit|Kev nominal 4B/8B labels; Jev size unknown here|
|
| 30 |
+
|Solar Pro4|Generative API answer parsed to a choice|Proprietary service; internals not audited|Unknown here; omitted from parameter-axis plots|
|
| 31 |
+
|
| 32 |
+
**Causal attention matters:** a later option can see preceding options, but an earlier option-end state cannot see later options. MiDM is not a bidirectional all-options cross-encoder. Option-order sensitivity and context truncation therefore remain relevant. CLM can cache state/action embeddings independently; MiDM option scores generally require the joint sequence to be recomputed. Neither parameter count alone nor adapter file size measures serving latency.
|
| 33 |
+
|
| 34 |
+
|Package|Nominal base|LoRA parameters (exact)|Pointer parameters (exact)|Trainable components total|
|
| 35 |
+
|---|---:|---:|---:|---:|
|
| 36 |
+
|MiDM 4B|4B|30,474,240|2,561|30,476,801|
|
| 37 |
+
|MiDM 9B|9B|40,108,032|4,097|40,112,129|
|
| 38 |
+
|
| 39 |
+
Counts are computed from shipped safetensors shapes, not estimated from file bytes. Base labels are nominal and are not exact instantiated text/vision/LM-head parameter counts. The base model is downloaded separately; a small adapter does not make the full model equally small.
|
| 40 |
+
|
| 41 |
+
## Parameters versus performance
|
| 42 |
+
|
| 43 |
+

|
| 44 |
+
|
| 45 |
+
The primary figure contains only the matched 4B/9B run. With 2.25 times the nominal base parameters, 9B improves Typed Decisions by **0.45 pp**, Kev transfer test by **2.75 pp**, JevBench public by **5.63 pp**, and its hard tier by **8.11 pp**. These retrospective point differences do not establish a general scaling law or statistical superiority.
|
| 46 |
+
|
| 47 |
+

|
| 48 |
+
|
| 49 |
+
The contextual plot keeps historical local and benchmark-published rows visibly separate. Historical bases, precision, epoch count, training mix and execution paths differ. It is not a matched architecture or size ablation. Systems with unknown parameter counts are excluded from this x-axis, rather than assigned a guessed size.
|
| 50 |
+
|
| 51 |
+
## Same public items, different inference protocols
|
| 52 |
+
|
| 53 |
+

|
| 54 |
+
|
| 55 |
+
|System|Nominal B|Public correct / 231|Public accuracy|Hard accuracy|Evidence class|
|
| 56 |
+
|---|---:|---:|---:|---:|---|
|
| 57 |
+
|MiDM 4B (matched)|4|165/231|71.43%|45.95%|local matched BF16/NF4|
|
| 58 |
+
|MiDM 9B (matched)|9|178/231|77.06%|54.05%|local matched BF16/NF4|
|
| 59 |
+
|CLM-style 4B control|4|95/231|41.13%|36.04%|historical local; configuration differs|
|
| 60 |
+
|Frozen joint pointer 4B|4|109/231|47.19%|36.94%|historical local; configuration differs|
|
| 61 |
+
|MiDM Qwen3 4B|4|159/231|68.83%|39.64%|historical local; configuration differs|
|
| 62 |
+
|MiDM Qwen3 8B (2ep)|8|163/231|70.56%|42.34%|historical local; configuration differs|
|
| 63 |
+
|MiDM Qwen3.5 2B|2|147/231|63.64%|37.84%|historical local; configuration differs|
|
| 64 |
+
|Kev 4B (published)|4|153/231|66.23%|36.94%|benchmark-author per-item outcomes, JevBench v1.2 snapshot|
|
| 65 |
+
|Kev 8B (published)|8|165/231|71.43%|45.05%|benchmark-author per-item outcomes, JevBench v1.2 snapshot|
|
| 66 |
+
|Jev 1.13 (published)|Unknown|200/231|86.58%|72.97%|benchmark-author per-item outcomes, JevBench v1.2 snapshot|
|
| 67 |
+
|Jev 1.13 Free (local API run)|Unknown|194/231|83.98%|67.57%|local API responses; different prompting/compute from pointer models|
|
| 68 |
+
|Solar Pro4 (local API run)|Unknown|220/231|95.24%|90.99%|local API responses; different prompting/compute from pointer models|
|
| 69 |
+
|Qwen3 4B reasoning|4|189/231|81.82%|65.77%|local reasoning run; different token budget and output path|
|
| 70 |
+
|
| 71 |
+
Jev published 200/231 and the local Jev Free run 194/231 are distinct service/protocol observations; neither overwrites the other. Solar 220/231 uses a generative API. They are useful capability references, not equal-compute or equal-cost experiments. The incomplete Space Bunny cohort (162/231 in the inspected snapshot) is excluded from the full-cohort graph; its 159/162 must not be ranked against complete 231-item results. This update makes no fresh calls to these model APIs.
|
| 72 |
+
|
| 73 |
+
## Architecture control and limits
|
| 74 |
+
|
| 75 |
+

|
| 76 |
+
|
| 77 |
+
The Qwen3-4B control series holds the nominal base family/size and original task mix, but training recipes and precision details still differ. Kev transfer accuracy rises from 41.10% (independent frozen heads) to 62.96% (frozen causal pointer) to 71.60% (QLoRA pointer). This supports candidate-context interaction as a useful design direction, but does not isolate every causal contribution. It must not be confused with a direct comparison against the released CLM 8B head.
|
| 78 |
+
|
| 79 |
+
## SQL/Python and deployment tradeoffs
|
| 80 |
+
|
| 81 |
+
|Same stored candidate pool|4B|9B|Interpretation|
|
| 82 |
+
|---|---:|---:|---|
|
| 83 |
+
|Spider test, RAG5 + dev-selected blend|78.34%|77.27%|Regression|
|
| 84 |
+
|SQL holdout, RAG5 + dev-selected blend|81.66%|81.77%|Small point increase|
|
| 85 |
+
|Python hard80, RAG0|28.75%|25.00%|Regression|
|
| 86 |
+
|
| 87 |
+
These are best-of-N selection results, not direct code generation. [Full SQL/Python protocol](https://github.com/YeoHoonYun/midm-decision-models/blob/v0.2.1/docs/evaluations/20261005/SQL_PYTHON_9B.md). The learned router reaches 79.65% public accuracy, below local reasoning alone at 81.82%; it is not promoted as the default. Long-context two-process pilots reduce throughput despite fitting in VRAM. [Serving and routing measurements](https://github.com/YeoHoonYun/midm-decision-models/blob/v0.2.1/docs/evaluations/20261005/README.md).
|
| 88 |
+
|
| 89 |
+
## Interpretation
|
| 90 |
+
|
| 91 |
+
9B improves several bounded-choice benchmarks and especially the public hard tier, but it does not dominate the generative comparison models or improve every application. The 4B model remains a smaller one-pass option scorer; the 9B model trades a larger base for higher decision accuracy on these suites. Unknown proprietary sizes, different inference budgets, seen public data and historical recipe differences prevent a universal performance-per-parameter ranking.
|
| 92 |
+
|
| 93 |
+
All outcomes are retrospective. Public-item accuracy is not an official composite score; trainable parameters are not full-model parameters; nominal size is not VRAM; batched throughput is not single-request latency. No financial inputs, probabilities, generated reports, raw benchmark examples or credentials are included.
|
| 94 |
+
|
| 95 |
+
## Sources and reproduction
|
| 96 |
+
|
| 97 |
+
[Machine-readable aggregates and source hashes](atlas.json) · [Plotting script](plot_atlas.py) · [Upstream asset inventory](upstream_inventory.json). Run `python plot_atlas.py` with Python, NumPy and Matplotlib to reproduce PNG/SVG plots from `atlas.json`.
|
| 98 |
+
|
| 99 |
+
- [CLM source and verifier protocol](https://github.com/Contrastive-LM/CLM); [projection implementation](https://github.com/Contrastive-LM/CLM/blob/main/src/clm/heads.py).
|
| 100 |
+
- [Released CLM DeepSWE heads](https://huggingface.co/Contrastive-LM/deepswe-clm-heads-8k).
|
| 101 |
+
- [JevBench v1.2 per-task source](https://github.com/fstandhartinger/jevbench/blob/main/results/v1.2/jevbench-v1.2-per-task.json). The archived local snapshot, not an assumed current leaderboard, underlies these figures.
|
| 102 |
+
- [MiDM primary evaluation](https://github.com/YeoHoonYun/midm-decision-models/blob/v0.2.1/docs/evaluations/20261005/README.md), [four-benchmark rescore](https://github.com/YeoHoonYun/midm-decision-models/blob/v0.2.1/docs/evaluations/20261005/four-benchmarks/README.md), and previously published local aggregate controls in this repository.
|
benchmarks/v0.2.1/architecture.png
ADDED
|
Git LFS Details
|
benchmarks/v0.2.1/architecture.svg
ADDED
|
|
benchmarks/v0.2.1/architecture_control.png
ADDED
|
Git LFS Details
|
benchmarks/v0.2.1/architecture_control.svg
ADDED
|
|
benchmarks/v0.2.1/atlas.json
ADDED
|
@@ -0,0 +1,587 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"created": "2026-10-05T23:16:39.058397",
|
| 3 |
+
"metric": "hard-label argmax accuracy; public JevBench accuracy is NOT official composite",
|
| 4 |
+
"parameter_axis": "nominal advertised base-model billions, not exact active/total count, trainable count or memory",
|
| 5 |
+
"models": [
|
| 6 |
+
{
|
| 7 |
+
"id": "midm_4b",
|
| 8 |
+
"name": "MiDM 4B (matched)",
|
| 9 |
+
"parameters_b_nominal": 4,
|
| 10 |
+
"family": "MiDM pointer + QLoRA",
|
| 11 |
+
"base": "Qwen3.5-4B-Base",
|
| 12 |
+
"provenance": "local matched BF16/NF4",
|
| 13 |
+
"metrics": {
|
| 14 |
+
"td_test": {
|
| 15 |
+
"n": 2000,
|
| 16 |
+
"correct": 1577,
|
| 17 |
+
"accuracy": 0.7885
|
| 18 |
+
},
|
| 19 |
+
"kevT_dev": {
|
| 20 |
+
"n": 764,
|
| 21 |
+
"correct": 604,
|
| 22 |
+
"accuracy": 0.7905759162303665
|
| 23 |
+
},
|
| 24 |
+
"kevT_test": {
|
| 25 |
+
"n": 764,
|
| 26 |
+
"correct": 610,
|
| 27 |
+
"accuracy": 0.7984293193717278
|
| 28 |
+
},
|
| 29 |
+
"jb_original": {
|
| 30 |
+
"n": 72,
|
| 31 |
+
"correct": 66,
|
| 32 |
+
"accuracy": 0.9166666666666666
|
| 33 |
+
},
|
| 34 |
+
"jb_easy": {
|
| 35 |
+
"n": 48,
|
| 36 |
+
"correct": 48,
|
| 37 |
+
"accuracy": 1.0
|
| 38 |
+
},
|
| 39 |
+
"jb_hard": {
|
| 40 |
+
"n": 111,
|
| 41 |
+
"correct": 51,
|
| 42 |
+
"accuracy": 0.4594594594594595
|
| 43 |
+
},
|
| 44 |
+
"jb_public": {
|
| 45 |
+
"n": 231,
|
| 46 |
+
"correct": 165,
|
| 47 |
+
"accuracy": 0.7142857142857143
|
| 48 |
+
}
|
| 49 |
+
}
|
| 50 |
+
},
|
| 51 |
+
{
|
| 52 |
+
"id": "midm_9b",
|
| 53 |
+
"name": "MiDM 9B (matched)",
|
| 54 |
+
"parameters_b_nominal": 9,
|
| 55 |
+
"family": "MiDM pointer + QLoRA",
|
| 56 |
+
"base": "Qwen3.5-9B-Base",
|
| 57 |
+
"provenance": "local matched BF16/NF4",
|
| 58 |
+
"metrics": {
|
| 59 |
+
"td_test": {
|
| 60 |
+
"n": 2000,
|
| 61 |
+
"correct": 1586,
|
| 62 |
+
"accuracy": 0.793
|
| 63 |
+
},
|
| 64 |
+
"kevT_dev": {
|
| 65 |
+
"n": 764,
|
| 66 |
+
"correct": 604,
|
| 67 |
+
"accuracy": 0.7905759162303665
|
| 68 |
+
},
|
| 69 |
+
"kevT_test": {
|
| 70 |
+
"n": 764,
|
| 71 |
+
"correct": 631,
|
| 72 |
+
"accuracy": 0.8259162303664922
|
| 73 |
+
},
|
| 74 |
+
"jb_original": {
|
| 75 |
+
"n": 72,
|
| 76 |
+
"correct": 70,
|
| 77 |
+
"accuracy": 0.9722222222222222
|
| 78 |
+
},
|
| 79 |
+
"jb_easy": {
|
| 80 |
+
"n": 48,
|
| 81 |
+
"correct": 48,
|
| 82 |
+
"accuracy": 1.0
|
| 83 |
+
},
|
| 84 |
+
"jb_hard": {
|
| 85 |
+
"n": 111,
|
| 86 |
+
"correct": 60,
|
| 87 |
+
"accuracy": 0.5405405405405406
|
| 88 |
+
},
|
| 89 |
+
"jb_public": {
|
| 90 |
+
"n": 231,
|
| 91 |
+
"correct": 178,
|
| 92 |
+
"accuracy": 0.7705627705627706
|
| 93 |
+
}
|
| 94 |
+
}
|
| 95 |
+
},
|
| 96 |
+
{
|
| 97 |
+
"id": "clm_style4",
|
| 98 |
+
"name": "CLM-style 4B control",
|
| 99 |
+
"parameters_b_nominal": 4,
|
| 100 |
+
"family": "Frozen independent encoders + projection heads",
|
| 101 |
+
"base": "Qwen/Qwen3-4B",
|
| 102 |
+
"provenance": "historical local; configuration differs",
|
| 103 |
+
"metrics": {
|
| 104 |
+
"td_test": {
|
| 105 |
+
"n": 2000,
|
| 106 |
+
"correct": 1518,
|
| 107 |
+
"accuracy": 0.759
|
| 108 |
+
},
|
| 109 |
+
"kevT_dev": {
|
| 110 |
+
"n": 764,
|
| 111 |
+
"correct": 280,
|
| 112 |
+
"accuracy": 0.36649214659685864
|
| 113 |
+
},
|
| 114 |
+
"kevT_test": {
|
| 115 |
+
"n": 764,
|
| 116 |
+
"correct": 314,
|
| 117 |
+
"accuracy": 0.4109947643979058
|
| 118 |
+
},
|
| 119 |
+
"jb_original": {
|
| 120 |
+
"n": 72,
|
| 121 |
+
"correct": 29,
|
| 122 |
+
"accuracy": 0.4027777777777778
|
| 123 |
+
},
|
| 124 |
+
"jb_easy": {
|
| 125 |
+
"n": 48,
|
| 126 |
+
"correct": 26,
|
| 127 |
+
"accuracy": 0.5416666666666666
|
| 128 |
+
},
|
| 129 |
+
"jb_hard": {
|
| 130 |
+
"n": 111,
|
| 131 |
+
"correct": 40,
|
| 132 |
+
"accuracy": 0.36036036036036034
|
| 133 |
+
},
|
| 134 |
+
"jb_public": {
|
| 135 |
+
"n": 231,
|
| 136 |
+
"correct": 95,
|
| 137 |
+
"accuracy": 0.41125541125541126
|
| 138 |
+
}
|
| 139 |
+
}
|
| 140 |
+
},
|
| 141 |
+
{
|
| 142 |
+
"id": "pointer_frozen4",
|
| 143 |
+
"name": "Frozen joint pointer 4B",
|
| 144 |
+
"parameters_b_nominal": 4,
|
| 145 |
+
"family": "Frozen causal encoder + pointer",
|
| 146 |
+
"base": "Qwen/Qwen3-4B",
|
| 147 |
+
"provenance": "historical local; configuration differs",
|
| 148 |
+
"metrics": {
|
| 149 |
+
"td_test": {
|
| 150 |
+
"n": 2000,
|
| 151 |
+
"correct": 1160,
|
| 152 |
+
"accuracy": 0.58
|
| 153 |
+
},
|
| 154 |
+
"kevT_dev": {
|
| 155 |
+
"n": 764,
|
| 156 |
+
"correct": 439,
|
| 157 |
+
"accuracy": 0.574607329842932
|
| 158 |
+
},
|
| 159 |
+
"kevT_test": {
|
| 160 |
+
"n": 764,
|
| 161 |
+
"correct": 481,
|
| 162 |
+
"accuracy": 0.6295811518324608
|
| 163 |
+
},
|
| 164 |
+
"jb_original": {
|
| 165 |
+
"n": 72,
|
| 166 |
+
"correct": 26,
|
| 167 |
+
"accuracy": 0.3611111111111111
|
| 168 |
+
},
|
| 169 |
+
"jb_easy": {
|
| 170 |
+
"n": 48,
|
| 171 |
+
"correct": 42,
|
| 172 |
+
"accuracy": 0.875
|
| 173 |
+
},
|
| 174 |
+
"jb_hard": {
|
| 175 |
+
"n": 111,
|
| 176 |
+
"correct": 41,
|
| 177 |
+
"accuracy": 0.36936936936936937
|
| 178 |
+
},
|
| 179 |
+
"jb_public": {
|
| 180 |
+
"n": 231,
|
| 181 |
+
"correct": 109,
|
| 182 |
+
"accuracy": 0.47186147186147187
|
| 183 |
+
}
|
| 184 |
+
}
|
| 185 |
+
},
|
| 186 |
+
{
|
| 187 |
+
"id": "midm_q3_4",
|
| 188 |
+
"name": "MiDM Qwen3 4B",
|
| 189 |
+
"parameters_b_nominal": 4,
|
| 190 |
+
"family": "MiDM pointer + QLoRA",
|
| 191 |
+
"base": "Qwen/Qwen3-4B",
|
| 192 |
+
"provenance": "historical local; configuration differs",
|
| 193 |
+
"metrics": {
|
| 194 |
+
"td_test": {
|
| 195 |
+
"n": 2000,
|
| 196 |
+
"correct": 1578,
|
| 197 |
+
"accuracy": 0.789
|
| 198 |
+
},
|
| 199 |
+
"kevT_dev": {
|
| 200 |
+
"n": 764,
|
| 201 |
+
"correct": 535,
|
| 202 |
+
"accuracy": 0.7002617801047121
|
| 203 |
+
},
|
| 204 |
+
"kevT_test": {
|
| 205 |
+
"n": 764,
|
| 206 |
+
"correct": 547,
|
| 207 |
+
"accuracy": 0.7159685863874345
|
| 208 |
+
},
|
| 209 |
+
"jb_original": {
|
| 210 |
+
"n": 72,
|
| 211 |
+
"correct": 67,
|
| 212 |
+
"accuracy": 0.9305555555555556
|
| 213 |
+
},
|
| 214 |
+
"jb_easy": {
|
| 215 |
+
"n": 48,
|
| 216 |
+
"correct": 48,
|
| 217 |
+
"accuracy": 1.0
|
| 218 |
+
},
|
| 219 |
+
"jb_hard": {
|
| 220 |
+
"n": 111,
|
| 221 |
+
"correct": 44,
|
| 222 |
+
"accuracy": 0.3963963963963964
|
| 223 |
+
},
|
| 224 |
+
"jb_public": {
|
| 225 |
+
"n": 231,
|
| 226 |
+
"correct": 159,
|
| 227 |
+
"accuracy": 0.6883116883116883
|
| 228 |
+
}
|
| 229 |
+
}
|
| 230 |
+
},
|
| 231 |
+
{
|
| 232 |
+
"id": "midm_q3_8",
|
| 233 |
+
"name": "MiDM Qwen3 8B (2ep)",
|
| 234 |
+
"parameters_b_nominal": 8,
|
| 235 |
+
"family": "MiDM pointer + QLoRA",
|
| 236 |
+
"base": "Qwen/Qwen3-8B",
|
| 237 |
+
"provenance": "historical local; configuration differs",
|
| 238 |
+
"metrics": {
|
| 239 |
+
"td_test": {
|
| 240 |
+
"n": 2000,
|
| 241 |
+
"correct": 1610,
|
| 242 |
+
"accuracy": 0.805
|
| 243 |
+
},
|
| 244 |
+
"kevT_dev": {
|
| 245 |
+
"n": 764,
|
| 246 |
+
"correct": 557,
|
| 247 |
+
"accuracy": 0.7290575916230366
|
| 248 |
+
},
|
| 249 |
+
"kevT_test": {
|
| 250 |
+
"n": 764,
|
| 251 |
+
"correct": 590,
|
| 252 |
+
"accuracy": 0.7722513089005235
|
| 253 |
+
},
|
| 254 |
+
"jb_original": {
|
| 255 |
+
"n": 72,
|
| 256 |
+
"correct": 68,
|
| 257 |
+
"accuracy": 0.9444444444444444
|
| 258 |
+
},
|
| 259 |
+
"jb_easy": {
|
| 260 |
+
"n": 48,
|
| 261 |
+
"correct": 48,
|
| 262 |
+
"accuracy": 1.0
|
| 263 |
+
},
|
| 264 |
+
"jb_hard": {
|
| 265 |
+
"n": 111,
|
| 266 |
+
"correct": 47,
|
| 267 |
+
"accuracy": 0.42342342342342343
|
| 268 |
+
},
|
| 269 |
+
"jb_public": {
|
| 270 |
+
"n": 231,
|
| 271 |
+
"correct": 163,
|
| 272 |
+
"accuracy": 0.7056277056277056
|
| 273 |
+
}
|
| 274 |
+
}
|
| 275 |
+
},
|
| 276 |
+
{
|
| 277 |
+
"id": "midm_q35_2",
|
| 278 |
+
"name": "MiDM Qwen3.5 2B",
|
| 279 |
+
"parameters_b_nominal": 2,
|
| 280 |
+
"family": "MiDM pointer + QLoRA",
|
| 281 |
+
"base": "Qwen/Qwen3.5-2B-Base",
|
| 282 |
+
"provenance": "historical local; configuration differs",
|
| 283 |
+
"metrics": {
|
| 284 |
+
"td_test": {
|
| 285 |
+
"n": 2000,
|
| 286 |
+
"correct": 1562,
|
| 287 |
+
"accuracy": 0.781
|
| 288 |
+
},
|
| 289 |
+
"kevT_dev": {
|
| 290 |
+
"n": 764,
|
| 291 |
+
"correct": 488,
|
| 292 |
+
"accuracy": 0.6387434554973822
|
| 293 |
+
},
|
| 294 |
+
"kevT_test": {
|
| 295 |
+
"n": 764,
|
| 296 |
+
"correct": 506,
|
| 297 |
+
"accuracy": 0.662303664921466
|
| 298 |
+
},
|
| 299 |
+
"jb_original": {
|
| 300 |
+
"n": 72,
|
| 301 |
+
"correct": 57,
|
| 302 |
+
"accuracy": 0.7916666666666666
|
| 303 |
+
},
|
| 304 |
+
"jb_easy": {
|
| 305 |
+
"n": 48,
|
| 306 |
+
"correct": 48,
|
| 307 |
+
"accuracy": 1.0
|
| 308 |
+
},
|
| 309 |
+
"jb_hard": {
|
| 310 |
+
"n": 111,
|
| 311 |
+
"correct": 42,
|
| 312 |
+
"accuracy": 0.3783783783783784
|
| 313 |
+
},
|
| 314 |
+
"jb_public": {
|
| 315 |
+
"n": 231,
|
| 316 |
+
"correct": 147,
|
| 317 |
+
"accuracy": 0.6363636363636364
|
| 318 |
+
}
|
| 319 |
+
}
|
| 320 |
+
},
|
| 321 |
+
{
|
| 322 |
+
"id": "kev-4b",
|
| 323 |
+
"name": "Kev 4B (published)",
|
| 324 |
+
"parameters_b_nominal": 4,
|
| 325 |
+
"family": "External typed-decision model; internals not audited",
|
| 326 |
+
"base": null,
|
| 327 |
+
"provenance": "benchmark-author per-item outcomes, JevBench v1.2 snapshot",
|
| 328 |
+
"metrics": {
|
| 329 |
+
"jb_public": {
|
| 330 |
+
"n": 231,
|
| 331 |
+
"correct": 153,
|
| 332 |
+
"accuracy": 0.6623376623376623
|
| 333 |
+
},
|
| 334 |
+
"jb_original": {
|
| 335 |
+
"n": 72,
|
| 336 |
+
"correct": 64,
|
| 337 |
+
"accuracy": 0.8888888888888888
|
| 338 |
+
},
|
| 339 |
+
"jb_easy": {
|
| 340 |
+
"n": 48,
|
| 341 |
+
"correct": 48,
|
| 342 |
+
"accuracy": 1.0
|
| 343 |
+
},
|
| 344 |
+
"jb_hard": {
|
| 345 |
+
"n": 111,
|
| 346 |
+
"correct": 41,
|
| 347 |
+
"accuracy": 0.36936936936936937
|
| 348 |
+
}
|
| 349 |
+
}
|
| 350 |
+
},
|
| 351 |
+
{
|
| 352 |
+
"id": "kev-8b",
|
| 353 |
+
"name": "Kev 8B (published)",
|
| 354 |
+
"parameters_b_nominal": 8,
|
| 355 |
+
"family": "External typed-decision model; internals not audited",
|
| 356 |
+
"base": null,
|
| 357 |
+
"provenance": "benchmark-author per-item outcomes, JevBench v1.2 snapshot",
|
| 358 |
+
"metrics": {
|
| 359 |
+
"jb_public": {
|
| 360 |
+
"n": 231,
|
| 361 |
+
"correct": 165,
|
| 362 |
+
"accuracy": 0.7142857142857143
|
| 363 |
+
},
|
| 364 |
+
"jb_original": {
|
| 365 |
+
"n": 72,
|
| 366 |
+
"correct": 67,
|
| 367 |
+
"accuracy": 0.9305555555555556
|
| 368 |
+
},
|
| 369 |
+
"jb_easy": {
|
| 370 |
+
"n": 48,
|
| 371 |
+
"correct": 48,
|
| 372 |
+
"accuracy": 1.0
|
| 373 |
+
},
|
| 374 |
+
"jb_hard": {
|
| 375 |
+
"n": 111,
|
| 376 |
+
"correct": 50,
|
| 377 |
+
"accuracy": 0.45045045045045046
|
| 378 |
+
}
|
| 379 |
+
}
|
| 380 |
+
},
|
| 381 |
+
{
|
| 382 |
+
"id": "jev-1.13.0",
|
| 383 |
+
"name": "Jev 1.13 (published)",
|
| 384 |
+
"parameters_b_nominal": null,
|
| 385 |
+
"family": "External typed-decision model; internals not audited",
|
| 386 |
+
"base": null,
|
| 387 |
+
"provenance": "benchmark-author per-item outcomes, JevBench v1.2 snapshot",
|
| 388 |
+
"metrics": {
|
| 389 |
+
"jb_public": {
|
| 390 |
+
"n": 231,
|
| 391 |
+
"correct": 200,
|
| 392 |
+
"accuracy": 0.8658008658008658
|
| 393 |
+
},
|
| 394 |
+
"jb_original": {
|
| 395 |
+
"n": 72,
|
| 396 |
+
"correct": 71,
|
| 397 |
+
"accuracy": 0.9861111111111112
|
| 398 |
+
},
|
| 399 |
+
"jb_easy": {
|
| 400 |
+
"n": 48,
|
| 401 |
+
"correct": 48,
|
| 402 |
+
"accuracy": 1.0
|
| 403 |
+
},
|
| 404 |
+
"jb_hard": {
|
| 405 |
+
"n": 111,
|
| 406 |
+
"correct": 81,
|
| 407 |
+
"accuracy": 0.7297297297297297
|
| 408 |
+
}
|
| 409 |
+
}
|
| 410 |
+
},
|
| 411 |
+
{
|
| 412 |
+
"id": "jev_free",
|
| 413 |
+
"name": "Jev 1.13 Free (local API run)",
|
| 414 |
+
"parameters_b_nominal": null,
|
| 415 |
+
"family": "Typed API",
|
| 416 |
+
"base": null,
|
| 417 |
+
"provenance": "local API responses; different prompting/compute from pointer models",
|
| 418 |
+
"metrics": {
|
| 419 |
+
"jb_easy": {
|
| 420 |
+
"n": 48,
|
| 421 |
+
"correct": 48,
|
| 422 |
+
"accuracy": 1.0
|
| 423 |
+
},
|
| 424 |
+
"jb_original": {
|
| 425 |
+
"n": 72,
|
| 426 |
+
"correct": 71,
|
| 427 |
+
"accuracy": 0.9861111111111112
|
| 428 |
+
},
|
| 429 |
+
"jb_hard": {
|
| 430 |
+
"n": 111,
|
| 431 |
+
"correct": 75,
|
| 432 |
+
"accuracy": 0.6756756756756757
|
| 433 |
+
},
|
| 434 |
+
"jb_public": {
|
| 435 |
+
"n": 231,
|
| 436 |
+
"correct": 194,
|
| 437 |
+
"accuracy": 0.8398268398268398
|
| 438 |
+
}
|
| 439 |
+
}
|
| 440 |
+
},
|
| 441 |
+
{
|
| 442 |
+
"id": "solar",
|
| 443 |
+
"name": "Solar Pro4 (local API run)",
|
| 444 |
+
"parameters_b_nominal": null,
|
| 445 |
+
"family": "Generative API",
|
| 446 |
+
"base": null,
|
| 447 |
+
"provenance": "local API responses; different prompting/compute from pointer models",
|
| 448 |
+
"metrics": {
|
| 449 |
+
"jb_easy": {
|
| 450 |
+
"n": 48,
|
| 451 |
+
"correct": 48,
|
| 452 |
+
"accuracy": 1.0
|
| 453 |
+
},
|
| 454 |
+
"jb_original": {
|
| 455 |
+
"n": 72,
|
| 456 |
+
"correct": 71,
|
| 457 |
+
"accuracy": 0.9861111111111112
|
| 458 |
+
},
|
| 459 |
+
"jb_hard": {
|
| 460 |
+
"n": 111,
|
| 461 |
+
"correct": 101,
|
| 462 |
+
"accuracy": 0.9099099099099099
|
| 463 |
+
},
|
| 464 |
+
"jb_public": {
|
| 465 |
+
"n": 231,
|
| 466 |
+
"correct": 220,
|
| 467 |
+
"accuracy": 0.9523809523809523
|
| 468 |
+
}
|
| 469 |
+
}
|
| 470 |
+
},
|
| 471 |
+
{
|
| 472 |
+
"id": "qwen3_4_reason",
|
| 473 |
+
"name": "Qwen3 4B reasoning",
|
| 474 |
+
"parameters_b_nominal": 4,
|
| 475 |
+
"family": "Generative local reasoning",
|
| 476 |
+
"base": "Qwen3-4B",
|
| 477 |
+
"provenance": "local reasoning run; different token budget and output path",
|
| 478 |
+
"metrics": {
|
| 479 |
+
"jb_easy": {
|
| 480 |
+
"n": 48,
|
| 481 |
+
"correct": 48,
|
| 482 |
+
"accuracy": 1.0
|
| 483 |
+
},
|
| 484 |
+
"jb_original": {
|
| 485 |
+
"n": 72,
|
| 486 |
+
"correct": 68,
|
| 487 |
+
"accuracy": 0.9444444444444444
|
| 488 |
+
},
|
| 489 |
+
"jb_hard": {
|
| 490 |
+
"n": 111,
|
| 491 |
+
"correct": 73,
|
| 492 |
+
"accuracy": 0.6576576576576577
|
| 493 |
+
},
|
| 494 |
+
"jb_public": {
|
| 495 |
+
"n": 231,
|
| 496 |
+
"correct": 189,
|
| 497 |
+
"accuracy": 0.8181818181818182
|
| 498 |
+
}
|
| 499 |
+
}
|
| 500 |
+
}
|
| 501 |
+
],
|
| 502 |
+
"trainable_parameters": {
|
| 503 |
+
"4": {
|
| 504 |
+
"nominal_base_b": 4,
|
| 505 |
+
"lora_parameters": 30474240,
|
| 506 |
+
"pointer_parameters": 2561,
|
| 507 |
+
"rank": 16,
|
| 508 |
+
"alpha": 32,
|
| 509 |
+
"base": "Qwen/Qwen3.5-4B-Base",
|
| 510 |
+
"artifact_hashes": {
|
| 511 |
+
"adapter_model.safetensors": "7b165d1d0d57a29e9c9ae39f639f77edc1ef8d2084174eb2afcc364eb8f00055",
|
| 512 |
+
"pointer_head.safetensors": "7e9e1f0f94d00ad5835d7829c22ef74827f4a95dcbf335e9ceac3a6d60bd9d50",
|
| 513 |
+
"adapter_config.json": "50ba2db8ff05e6d6a1f8bdecd8005a3f76e8cf8e926ec84e3d7e5666d07ffc11"
|
| 514 |
+
}
|
| 515 |
+
},
|
| 516 |
+
"9": {
|
| 517 |
+
"nominal_base_b": 9,
|
| 518 |
+
"lora_parameters": 40108032,
|
| 519 |
+
"pointer_parameters": 4097,
|
| 520 |
+
"rank": 16,
|
| 521 |
+
"alpha": 32,
|
| 522 |
+
"base": "Qwen/Qwen3.5-9B-Base",
|
| 523 |
+
"artifact_hashes": {
|
| 524 |
+
"adapter_model.safetensors": "7231ef52a482760601bd6e9038ec9aae40034d0f63c6fb5de677b5f81e6dd5c2",
|
| 525 |
+
"pointer_head.safetensors": "f534d840c1b521c04292e763db200e6ab10df12fb4ebdabc78de734c07d88e4a",
|
| 526 |
+
"adapter_config.json": "96f5ae3d77c1c0ef156344c0664cda89bd3a1449ca2f71d398744a2dc073d2d5"
|
| 527 |
+
}
|
| 528 |
+
}
|
| 529 |
+
},
|
| 530 |
+
"matched_protocol": {
|
| 531 |
+
"dtype": "NF4/BF16",
|
| 532 |
+
"gpu": "RTX 3090",
|
| 533 |
+
"max_input_tokens": 4096,
|
| 534 |
+
"batch_token_budget": 4096,
|
| 535 |
+
"tta": 1
|
| 536 |
+
},
|
| 537 |
+
"sources": [
|
| 538 |
+
{
|
| 539 |
+
"source": "experiments/113_clm_reproduction_20260928/results/comparison_20261005/metrics.json",
|
| 540 |
+
"sha256": "8a7665f6b490259213cf1ea9a82651ec1e470fa71177a93d84a1138dfa00675a"
|
| 541 |
+
},
|
| 542 |
+
{
|
| 543 |
+
"source": "experiments/113_clm_reproduction_20260928/results/multi/FINAL_qwen3-4b_e40.json",
|
| 544 |
+
"sha256": "5b1f9a0fe1cc8ddb606bad84d55bf3e8474d1e77af4b59ffff5d86d7d3f8a967"
|
| 545 |
+
},
|
| 546 |
+
{
|
| 547 |
+
"source": "experiments/113_clm_reproduction_20260928/results/pointer/FINAL_H_q3_4b_headonly.json",
|
| 548 |
+
"sha256": "be7765ceaaf1aa71fb8210832da40ceba48f776e3ed0067df934bae1f81b81a1"
|
| 549 |
+
},
|
| 550 |
+
{
|
| 551 |
+
"source": "experiments/113_clm_reproduction_20260928/results/pointer/FINAL_qwen3-4b.json",
|
| 552 |
+
"sha256": "853d842cbe2d96bfeaad0b3879863e89878655dae41615b0848ef226fd069ab8"
|
| 553 |
+
},
|
| 554 |
+
{
|
| 555 |
+
"source": "experiments/113_clm_reproduction_20260928/results/pointer/FINAL_B_qwen3-8b-2ep.json",
|
| 556 |
+
"sha256": "bc80e97bd672812f442db984852484ba84ba2769b4ccb27a1844c660c659f29f"
|
| 557 |
+
},
|
| 558 |
+
{
|
| 559 |
+
"source": "experiments/113_clm_reproduction_20260928/results/pointer/FINAL_P_q35_2b_e1_bf16.json",
|
| 560 |
+
"sha256": "2ef9456fb45f8cfe73e8d6ce5a8824285d532a33e441f45df59e2ae1157d267a"
|
| 561 |
+
},
|
| 562 |
+
{
|
| 563 |
+
"source": "experiments/113_clm_reproduction_20260928/results/jev_paired/jev_paired.json",
|
| 564 |
+
"sha256": "cc82c78affa57db07373cdc173ddc4630a3d5efc94c3d089c8cecc08ae463a04"
|
| 565 |
+
},
|
| 566 |
+
{
|
| 567 |
+
"source": "experiments/116_api_vs_midm_20261003/free_model_catalog/comparison_latest.json",
|
| 568 |
+
"sha256": "4fa24e256c0dd910deddca300bfd9d093e0fbbb54b7a1c12033409d51b703c7d"
|
| 569 |
+
},
|
| 570 |
+
{
|
| 571 |
+
"source": "experiments/116_api_vs_midm_20261003/router_pilot_20261005/RESULTS.json",
|
| 572 |
+
"sha256": "2f78ad47cf8ee4f01e6e88dc346de8792c159af17eb9550a0e4dcf891a0fa044"
|
| 573 |
+
},
|
| 574 |
+
{
|
| 575 |
+
"source": "research_topics/clm_decision_heads_novelty_20260928/huggingface/MiDM-4B-q35-e1-bx/adapter_config.json",
|
| 576 |
+
"sha256": "50ba2db8ff05e6d6a1f8bdecd8005a3f76e8cf8e926ec84e3d7e5666d07ffc11"
|
| 577 |
+
},
|
| 578 |
+
{
|
| 579 |
+
"source": "research_topics/clm_decision_heads_novelty_20260928/huggingface/MiDM-9B-q35-e1-bx/adapter_config.json",
|
| 580 |
+
"sha256": "96f5ae3d77c1c0ef156344c0664cda89bd3a1449ca2f71d398744a2dc073d2d5"
|
| 581 |
+
},
|
| 582 |
+
{
|
| 583 |
+
"source": "experiments/113_clm_reproduction_20260928/results/benchmark_atlas_20261005/upstream_inventory.json",
|
| 584 |
+
"sha256": "516c617f6ee0ddf4253cd816a85a3f9514d7c2e2bb9b46950009506327e8563a"
|
| 585 |
+
}
|
| 586 |
+
]
|
| 587 |
+
}
|
benchmarks/v0.2.1/context_size_performance.png
ADDED
|
Git LFS Details
|
benchmarks/v0.2.1/context_size_performance.svg
ADDED
|
|
benchmarks/v0.2.1/matched_size_performance.png
ADDED
|
Git LFS Details
|
benchmarks/v0.2.1/matched_size_performance.svg
ADDED
|
|
benchmarks/v0.2.1/plot_atlas.py
ADDED
|
@@ -0,0 +1,94 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Reproduce the public figures from aggregate atlas.json; no network or model needed."""
|
| 2 |
+
from pathlib import Path
|
| 3 |
+
import json, textwrap
|
| 4 |
+
import numpy as np
|
| 5 |
+
import matplotlib
|
| 6 |
+
matplotlib.use('Agg')
|
| 7 |
+
import matplotlib.pyplot as plt
|
| 8 |
+
from matplotlib.patches import FancyBboxPatch, FancyArrowPatch
|
| 9 |
+
|
| 10 |
+
ROOT=Path(__file__).resolve().parent
|
| 11 |
+
D=json.loads((ROOT/'atlas.json').read_text(encoding='utf-8'))
|
| 12 |
+
MODELS={r['id']:r for r in D['models']}
|
| 13 |
+
TEAL='#007f86';BLUE='#395a9e';ORANGE='#c56a23';PURPLE='#8667a7';GRAY='#657580'
|
| 14 |
+
plt.rcParams.update({'font.family':'DejaVu Sans','font.size':10,'axes.spines.top':False,'axes.spines.right':False,'axes.labelcolor':'#283a47','text.color':'#283a47','axes.titleweight':'bold','svg.fonttype':'none','savefig.facecolor':'white'})
|
| 15 |
+
|
| 16 |
+
def save(fig,name):
|
| 17 |
+
for ext in ['png','svg']:
|
| 18 |
+
p=ROOT/(name+'.'+ext)
|
| 19 |
+
fig.savefig(p,dpi=170,bbox_inches='tight',**({'metadata':{'Date':None}} if ext=='svg' else {}))
|
| 20 |
+
if ext=='svg':p.write_bytes(('\n'.join(line.rstrip() for line in p.read_text(encoding='utf-8').splitlines())+'\n').encode('utf-8'))
|
| 21 |
+
plt.close(fig)
|
| 22 |
+
|
| 23 |
+
def matched():
|
| 24 |
+
fig,axes=plt.subplots(2,2,figsize=(12,8));fig.suptitle('MiDM: nominal base parameters vs matched decision accuracy',fontsize=17,y=.98)
|
| 25 |
+
for ax,(key,title) in zip(axes.flat,[('td_test','Typed Decisions test · N=2000'),('kevT_test','Kev transfer test · N=764'),('jb_public','JevBench public · N=231'),('jb_hard','JevBench hard · N=111')]):
|
| 26 |
+
y=[100*MODELS[m]['metrics'][key]['accuracy'] for m in ['midm_4b','midm_9b']]
|
| 27 |
+
ax.plot([4,9],y,color=TEAL,linewidth=1.5,linestyle='--',zorder=2);ax.scatter([4,9],y,c=[BLUE,TEAL],s=95,zorder=3)
|
| 28 |
+
for x,v in zip([4,9],y):ax.annotate(f'{v:.2f}%',(x,v),xytext=(0,11),textcoords='offset points',ha='center',weight='bold')
|
| 29 |
+
ax.set(title=title,xlim=(2.7,10.3),ylim=(0,103),xticks=[4,9],xlabel='Nominal base parameters (billions)',ylabel='Hard-label accuracy (%)');ax.grid(axis='y',alpha=.16)
|
| 30 |
+
ax.text(.5,.1,f'9B minus 4B: +{y[1]-y[0]:.2f} percentage points',ha='center',transform=ax.transAxes,fontsize=10)
|
| 31 |
+
fig.subplots_adjust(left=.085,right=.97,bottom=.13,top=.89,wspace=.24,hspace=.45)
|
| 32 |
+
fig.text(.08,.035,'Same fixed items, NF4/BF16, RTX 3090, context 4096, TTA=1. Retrospective point estimates.\nJevBench public accuracy is not the official composite. Connecting lines are not fitted scaling laws.',fontsize=10,color=GRAY)
|
| 33 |
+
save(fig,'matched_size_performance')
|
| 34 |
+
|
| 35 |
+
def context():
|
| 36 |
+
ids=['midm_q35_2','clm_style4','pointer_frozen4','midm_q3_4','midm_q3_8','midm_4b','midm_9b','kev-4b','kev-8b','qwen3_4_reason']
|
| 37 |
+
label={'midm_q35_2':'MiDM Q3.5 2B','clm_style4':'CLM-style 4B','pointer_frozen4':'Frozen pointer 4B','midm_q3_4':'MiDM Q3 4B','midm_q3_8':'MiDM Q3 8B','midm_4b':'MiDM Q3.5 4B','midm_9b':'MiDM Q3.5 9B','kev-4b':'Kev 4B [published]','kev-8b':'Kev 8B [published]','qwen3_4_reason':'Qwen3 4B reasoning'}
|
| 38 |
+
offsets=[{'midm_q35_2':(.8,58),'clm_style4':(4.3,36),'pointer_frozen4':(4.3,48),'midm_q3_4':(2.2,77),'midm_q3_8':(6.2,66),'midm_4b':(4.3,73),'midm_9b':(7.1,85),'kev-4b':(1,67),'kev-8b':(6,79),'qwen3_4_reason':(4.3,88)},
|
| 39 |
+
{'midm_q35_2':(.8,49),'clm_style4':(2.4,23),'pointer_frozen4':(4.3,30),'midm_q3_4':(4.3,42),'midm_q3_8':(6.4,36),'midm_4b':(4.3,52),'midm_9b':(7,62),'kev-4b':(1,32),'kev-8b':(6.2,49),'qwen3_4_reason':(4.3,71)}]
|
| 40 |
+
fig,axes=plt.subplots(1,2,figsize=(16,7));fig.suptitle('Parameter context: local measurements and published comparison rows',fontsize=17,y=.97)
|
| 41 |
+
for j,(ax,key,title) in enumerate(zip(axes,['jb_public','jb_hard'],['Same public 231 items','Public hard tier · 111 items'])):
|
| 42 |
+
for id in ids:
|
| 43 |
+
r=MODELS[id];x=r['parameters_b_nominal'];y=r['metrics'][key]['accuracy']*100
|
| 44 |
+
published=id.startswith('kev-');primary=id in ['midm_4b','midm_9b'];reason=id=='qwen3_4_reason'
|
| 45 |
+
color=PURPLE if published else ORANGE if reason else TEAL if primary else GRAY
|
| 46 |
+
marker='s' if published else 'D' if reason else 'o' if primary else '^'
|
| 47 |
+
ax.scatter(x,y,s=75,marker=marker,c=color,zorder=3)
|
| 48 |
+
ax.annotate(label[id],(x,y),xytext=offsets[j][id],textcoords='data',ha='left',fontsize=8.5,color=color,arrowprops={'arrowstyle':'-','color':color,'lw':.5})
|
| 49 |
+
ax.set(title=title,xlabel='Nominal base parameters (billions)',ylabel='Hard-label accuracy (%)',xlim=(.5,10.7),ylim=(0,103),xticks=[2,4,8,9]);ax.grid(alpha=.12)
|
| 50 |
+
fig.subplots_adjust(left=.06,right=.97,bottom=.2,top=.86,wspace=.20)
|
| 51 |
+
fig.text(.06,.04,'Teal circles: matched MiDM 4B/9B. Gray triangles: historical local settings. Purple squares: benchmark-published Kev.\nOrange diamonds: local generative reasoning. Bases, precision, data, epochs and compute differ across these groups.\nUnknown-size Jev and Solar are excluded here, not assigned an estimated parameter count. No architecture-only or scaling-law claim.',fontsize=10,color=GRAY)
|
| 52 |
+
save(fig,'context_size_performance')
|
| 53 |
+
|
| 54 |
+
def comparisons():
|
| 55 |
+
rows=sorted(D['models'],key=lambda r:r['metrics']['jb_public']['accuracy']);labels=[r['name'] for r in rows]
|
| 56 |
+
fig,axes=plt.subplots(1,2,figsize=(15,8.5),sharey=True);fig.suptitle('Public JevBench cohort: capability comparison, not equal compute',fontsize=17,y=.97)
|
| 57 |
+
for ax,key,title in zip(axes,['jb_public','jb_hard'],['Public accuracy · 231 items','Hard accuracy · 111 items']):
|
| 58 |
+
vals=[r['metrics'][key]['accuracy']*100 for r in rows]
|
| 59 |
+
colors=[TEAL if r['id'] in ['midm_4b','midm_9b'] else PURPLE if 'published' in r['name'] else ORANGE if r['id'] in ['solar','jev_free','qwen3_4_reason'] else GRAY for r in rows]
|
| 60 |
+
ax.barh(range(len(rows)),vals,color=colors,height=.66)
|
| 61 |
+
for i,v in enumerate(vals):ax.text(v+1,i,f'{v:.2f}%',va='center',fontsize=9)
|
| 62 |
+
ax.set(title=title,xlim=(0,112),xlabel='Hard-label accuracy (%)',yticks=range(len(rows)),yticklabels=labels);ax.grid(axis='x',alpha=.14);ax.set_axisbelow(True)
|
| 63 |
+
axes[0].tick_params(axis='y',labelsize=9)
|
| 64 |
+
fig.subplots_adjust(left=.25,right=.96,bottom=.15,top=.88,wspace=.15)
|
| 65 |
+
fig.text(.05,.045,'Teal: matched MiDM. Gray: historical local. Purple: benchmark-published outcomes. Orange: local API/reasoning runs.\nAll rows cover the same 231 public items, but service versions, prompts, training exposure and computation differ.\nJev published and Jev Free measured are separate observations. Incomplete cohorts are excluded. Not the official composite.',fontsize=10,color=GRAY)
|
| 66 |
+
save(fig,'public_comparison')
|
| 67 |
+
|
| 68 |
+
def architecture():
|
| 69 |
+
fig,ax=plt.subplots(figsize=(16,7));ax.set(xlim=(0,16),ylim=(0,7));ax.axis('off');fig.suptitle('Architecture comparison: how a bounded decision is produced',fontsize=18,y=.96)
|
| 70 |
+
lanes=[(5.3,'CLM reference 8B',PURPLE,['State and candidates\nencoded separately','Frozen Qwen3-8B\nlast-token embeddings','Two projection MLPs\nstate / action','Scaled cosine score\nrank / softmax']),
|
| 71 |
+
(3.35,'MiDM 4B / 9B',TEAL,['State + question +\nordered option lines','Causal base in NF4\n+ rank-16 LoRA','Hidden state at\neach option end','Shared linear pointer\nsoftmax over options']),
|
| 72 |
+
(1.4,'Generative reasoning',ORANGE,['State + question +\ncandidate prompt','Autoregressive decoder\nreasoning / answer tokens','Parse generated\nanswer or tool output','Typed choice\nwith parser checks'])]
|
| 73 |
+
for y,title,color,boxes in lanes:
|
| 74 |
+
ax.text(.15,y+.62,title,fontsize=12,weight='bold',color=color)
|
| 75 |
+
for j,text in enumerate(boxes):
|
| 76 |
+
x=.15+j*4.02;box=FancyBboxPatch((x,y-.45),3.55,.82,boxstyle='round,pad=0.08',edgecolor=color,facecolor=color+'12',linewidth=1.4);ax.add_patch(box);ax.text(x+1.775,y-.035,text,ha='center',va='center',fontsize=10.5)
|
| 77 |
+
if j<3:ax.add_patch(FancyArrowPatch((x+3.65,y-.035),(x+3.9,y-.035),arrowstyle='-|>',mutation_scale=14,color=color))
|
| 78 |
+
note={'CLM reference 8B':'Independent embedding reuse is possible; released task-specific heads are separate artifacts.','MiDM 4B / 9B':'One causal pass: later options see earlier options, not vice versa. This is not bidirectional attention.','Generative reasoning':'More decoding work; a higher accuracy does not imply equivalent latency, cost or parameter disclosure.'}[title]
|
| 79 |
+
ax.text(.15,y-.8,note,color=GRAY,fontsize=9.5)
|
| 80 |
+
fig.subplots_adjust(left=.025,right=.985,bottom=.03,top=.91);save(fig,'architecture')
|
| 81 |
+
|
| 82 |
+
def control():
|
| 83 |
+
ids=['clm_style4','pointer_frozen4','midm_q3_4'];names=['Independent frozen\nprojection heads','Frozen causal\npointer head','Causal pointer\n+ QLoRA'];x=np.arange(3);fig,ax=plt.subplots(figsize=(11,6.5))
|
| 84 |
+
for delta,key,name,color in [(-.18,'td_test','Typed Decisions test (2000)',BLUE),(.18,'kevT_test','Kev transfer test (764)',TEAL)]:
|
| 85 |
+
vals=[100*MODELS[i]['metrics'][key]['accuracy'] for i in ids];bars=ax.bar(x+delta,vals,.34,label=name,color=color)
|
| 86 |
+
ax.bar_label(bars,labels=[f'{v:.2f}%' for v in vals],padding=4,fontsize=11)
|
| 87 |
+
ax.set(title='Historical Qwen3-4B controls: input interaction and adaptation',ylabel='Hard-label accuracy (%)',ylim=(0,104),xticks=x,xticklabels=names);ax.grid(axis='y',alpha=.15);ax.set_axisbelow(True);ax.legend(loc='upper left',frameon=False)
|
| 88 |
+
fig.subplots_adjust(left=.08,right=.97,bottom=.24,top=.88)
|
| 89 |
+
fig.text(.08,.045,'Same nominal base family/size and original task mix; training recipes and recorded precision are not fully matched.\nThis is a historical design control, not an architecture-only causal estimate and not the released CLM 8B.\nPrimary Qwen3.5 4B/9B comparisons are shown separately.',fontsize=10,color=GRAY)
|
| 90 |
+
save(fig,'architecture_control')
|
| 91 |
+
|
| 92 |
+
if __name__=='__main__':
|
| 93 |
+
for f in [matched,context,comparisons,architecture,control]:f()
|
| 94 |
+
print('Wrote five figures as PNG and SVG from atlas.json.')
|
benchmarks/v0.2.1/public_comparison.png
ADDED
|
Git LFS Details
|
benchmarks/v0.2.1/public_comparison.svg
ADDED
|
|
benchmarks/v0.2.1/upstream_inventory.json
ADDED
|
@@ -0,0 +1,80 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"checked_at": "2026-10-05T23:11:58.401173",
|
| 3 |
+
"sources": [
|
| 4 |
+
{
|
| 5 |
+
"url": "https://huggingface.co/api/models?author=Contrastive-LM&limit=100",
|
| 6 |
+
"sha": null,
|
| 7 |
+
"ids": [
|
| 8 |
+
"Contrastive-LM/CLM-v0.1-8B",
|
| 9 |
+
"Contrastive-LM/deepswe-clm-heads-8k"
|
| 10 |
+
],
|
| 11 |
+
"file_count": 0,
|
| 12 |
+
"evaluation_related_files": []
|
| 13 |
+
},
|
| 14 |
+
{
|
| 15 |
+
"url": "https://huggingface.co/api/models/Contrastive-LM/CLM-v0.1-8B",
|
| 16 |
+
"sha": "e939398d4556fcd9400c76fa8c5a513202f42b0a",
|
| 17 |
+
"ids": null,
|
| 18 |
+
"file_count": 6,
|
| 19 |
+
"evaluation_related_files": []
|
| 20 |
+
},
|
| 21 |
+
{
|
| 22 |
+
"url": "https://huggingface.co/api/models/Contrastive-LM/deepswe-clm-heads-8k",
|
| 23 |
+
"sha": "c60876f3fdf7a75dc58d33b776e469b7e903d0ee",
|
| 24 |
+
"ids": null,
|
| 25 |
+
"file_count": 9,
|
| 26 |
+
"evaluation_related_files": [
|
| 27 |
+
"heldout_tasks.json",
|
| 28 |
+
"split_summary.json",
|
| 29 |
+
"task_split.json",
|
| 30 |
+
"verification.json"
|
| 31 |
+
]
|
| 32 |
+
},
|
| 33 |
+
{
|
| 34 |
+
"url": "https://huggingface.co/api/datasets?author=Contrastive-LM&limit=100",
|
| 35 |
+
"sha": null,
|
| 36 |
+
"ids": [
|
| 37 |
+
"Contrastive-LM/deepswe-clm-train-embeddings-8k",
|
| 38 |
+
"Contrastive-LM/CLM-v0.1-Pretrain-Nemotron"
|
| 39 |
+
],
|
| 40 |
+
"file_count": 0,
|
| 41 |
+
"evaluation_related_files": []
|
| 42 |
+
},
|
| 43 |
+
{
|
| 44 |
+
"url": "https://huggingface.co/api/datasets/Contrastive-LM/deepswe-clm-train-embeddings-8k",
|
| 45 |
+
"sha": "4e42267114a2b28dc1d4fab85a2fe2790abc16d2",
|
| 46 |
+
"ids": null,
|
| 47 |
+
"file_count": 18,
|
| 48 |
+
"evaluation_related_files": []
|
| 49 |
+
},
|
| 50 |
+
{
|
| 51 |
+
"url": "https://huggingface.co/api/datasets/Contrastive-LM/CLM-v0.1-Pretrain-Nemotron",
|
| 52 |
+
"sha": "05ca7d05c043350326c5e86c589b9542962747b4",
|
| 53 |
+
"ids": null,
|
| 54 |
+
"file_count": 1876,
|
| 55 |
+
"evaluation_related_files": []
|
| 56 |
+
},
|
| 57 |
+
{
|
| 58 |
+
"url": "https://api.github.com/repos/Contrastive-LM/CLM/git/trees/main?recursive=1",
|
| 59 |
+
"sha": "bb42c6c5bf914fd449bed2f6ca65be80602cb1f7",
|
| 60 |
+
"ids": null,
|
| 61 |
+
"file_count": 46,
|
| 62 |
+
"evaluation_related_files": [
|
| 63 |
+
"evaluation/bon_eval.py"
|
| 64 |
+
]
|
| 65 |
+
}
|
| 66 |
+
],
|
| 67 |
+
"local_deepswe_columns": [
|
| 68 |
+
"trajectory_id",
|
| 69 |
+
"step_idx",
|
| 70 |
+
"task_id",
|
| 71 |
+
"model",
|
| 72 |
+
"config",
|
| 73 |
+
"reward",
|
| 74 |
+
"fold",
|
| 75 |
+
"prm_score",
|
| 76 |
+
"state_embedding",
|
| 77 |
+
"action_embedding"
|
| 78 |
+
],
|
| 79 |
+
"finding": "No matching raw DeepSWE candidate text or Terminal-Bench 30-task verifier pool found in these inspected releases. This is a scoped search, not proof that no such assets exist elsewhere."
|
| 80 |
+
}
|
manifest.json
CHANGED
|
@@ -1,18 +1,33 @@
|
|
| 1 |
-
{
|
| 2 |
-
"version": "0.2.
|
| 3 |
-
"model": "MiDM-9B-q35-e1-bx",
|
| 4 |
-
"files": {
|
| 5 |
-
"adapter_config.json": "96f5ae3d77c1c0ef156344c0664cda89bd3a1449ca2f71d398744a2dc073d2d5",
|
| 6 |
-
"adapter_model.safetensors": "7231ef52a482760601bd6e9038ec9aae40034d0f63c6fb5de677b5f81e6dd5c2",
|
| 7 |
-
"evaluation.json": "a3684eef4f8a678a4d821319eff403f83e1dd45db6b432f63e956f8a8d20f746",
|
| 8 |
-
"LICENSE": "cfc7749b96f63bd31c3c42b5c471bf756814053e847c10f3eb003417bc523d30",
|
| 9 |
-
"midm.py": "2a3985c09b38802d240acec138b507358ee4bd30e413c82dd3215ee7ac0681fd",
|
| 10 |
-
"midm_config.json": "fe75b3e6b3eab1e9f3a9e9ec9475d4c5551ca9d57473320838a51c1affd8b077",
|
| 11 |
-
"NOTICE": "8c1e780a6af2a8c5402c719d95bd3be4ebd5f05c7fce0f9def1cff665c4ca21d",
|
| 12 |
-
"package_validation.json": "efb81406f22b909db9a77fbce5c3b3fe80af8201c5cda6abff11d31077957470",
|
| 13 |
-
"pointer_head.safetensors": "f534d840c1b521c04292e763db200e6ab10df12fb4ebdabc78de734c07d88e4a",
|
| 14 |
-
"README.md": "
|
| 15 |
-
"requirements.txt": "550b986e177232471ed3dee7efacb0d8b0dacf0cbc7044b9b81a58dc38d808c2",
|
| 16 |
-
"VERIFICATION.md": "4459ba20a9f9447df6ff205b5dfce778a18a1d64ca25614dec500a828f2e7b89"
|
| 17 |
-
|
| 18 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"version": "0.2.1",
|
| 3 |
+
"model": "MiDM-9B-q35-e1-bx",
|
| 4 |
+
"files": {
|
| 5 |
+
"adapter_config.json": "96f5ae3d77c1c0ef156344c0664cda89bd3a1449ca2f71d398744a2dc073d2d5",
|
| 6 |
+
"adapter_model.safetensors": "7231ef52a482760601bd6e9038ec9aae40034d0f63c6fb5de677b5f81e6dd5c2",
|
| 7 |
+
"evaluation.json": "a3684eef4f8a678a4d821319eff403f83e1dd45db6b432f63e956f8a8d20f746",
|
| 8 |
+
"LICENSE": "cfc7749b96f63bd31c3c42b5c471bf756814053e847c10f3eb003417bc523d30",
|
| 9 |
+
"midm.py": "2a3985c09b38802d240acec138b507358ee4bd30e413c82dd3215ee7ac0681fd",
|
| 10 |
+
"midm_config.json": "fe75b3e6b3eab1e9f3a9e9ec9475d4c5551ca9d57473320838a51c1affd8b077",
|
| 11 |
+
"NOTICE": "8c1e780a6af2a8c5402c719d95bd3be4ebd5f05c7fce0f9def1cff665c4ca21d",
|
| 12 |
+
"package_validation.json": "efb81406f22b909db9a77fbce5c3b3fe80af8201c5cda6abff11d31077957470",
|
| 13 |
+
"pointer_head.safetensors": "f534d840c1b521c04292e763db200e6ab10df12fb4ebdabc78de734c07d88e4a",
|
| 14 |
+
"README.md": "b392455d328c9af67f162394beeef4d1c230d5d5652fda4e94023d51af15f75d",
|
| 15 |
+
"requirements.txt": "550b986e177232471ed3dee7efacb0d8b0dacf0cbc7044b9b81a58dc38d808c2",
|
| 16 |
+
"VERIFICATION.md": "4459ba20a9f9447df6ff205b5dfce778a18a1d64ca25614dec500a828f2e7b89",
|
| 17 |
+
"benchmarks/v0.2.1/architecture.png": "cc5bfb7fc225b1f974837e36ef43717af271fe1d56281f6faa8790f4d1762280",
|
| 18 |
+
"benchmarks/v0.2.1/architecture.svg": "0d77746badfff06e4b07f686753830f565568e98ad6b88dfa3e6e179c413bba7",
|
| 19 |
+
"benchmarks/v0.2.1/architecture_control.png": "18e6afac826bda17ed80a7f0252f63c69109aeeca41026efbc50d01845b52746",
|
| 20 |
+
"benchmarks/v0.2.1/architecture_control.svg": "d837c9733a0e3af6c4b7c2078a940af18f3d74c6bb5871ff2231f4525c89b661",
|
| 21 |
+
"benchmarks/v0.2.1/atlas.json": "3d82078bac0024ac3fc4a814580a7fd11528c6f659acda99099e8a3703e70c4a",
|
| 22 |
+
"benchmarks/v0.2.1/context_size_performance.png": "62364f4bdfd3a00358bf2631a03d550b86d75b62b0682e936a7782d30eda5588",
|
| 23 |
+
"benchmarks/v0.2.1/context_size_performance.svg": "ef50c39dc7e19f0f289064ae41baf77889ca8fa48f656a4f92f4b9d8cd06c7c3",
|
| 24 |
+
"benchmarks/v0.2.1/matched_size_performance.png": "6e5b233ccfca43b71bd4f2884d38881751628db14369a035eac97ac5995b8729",
|
| 25 |
+
"benchmarks/v0.2.1/matched_size_performance.svg": "ea0fbc1d8c4658254c8be3a4e37094262877c404b4357e4d1e65fb2e4c07d83b",
|
| 26 |
+
"benchmarks/v0.2.1/plot_atlas.py": "2081011cfd8cef64f0f16887f5bf5df4fb1ba71c9a2da31e640e141ee7b03254",
|
| 27 |
+
"benchmarks/v0.2.1/public_comparison.png": "2ce2d9502232d5683e840afa3e97a3b203665b31b9dec6fa63d4fc3b1576aa87",
|
| 28 |
+
"benchmarks/v0.2.1/public_comparison.svg": "921ef3941ca10f90acce98bed502e95caaa143a6eb6ad3a706765cd36d3dcbfc",
|
| 29 |
+
"benchmarks/v0.2.1/README.md": "0768c6ea91f2ce2e218e268d4e7c3fb9a943e6070464976f0be57bd4555fe2b7",
|
| 30 |
+
"benchmarks/v0.2.1/upstream_inventory.json": "2515bc93354d78242e0df50aef7130ed67d19190d53384daf1fcd0746e3c1186"
|
| 31 |
+
},
|
| 32 |
+
"weights_version": "0.2.0"
|
| 33 |
+
}
|