Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string

MiDM-9B-q35-e1-bx

v0.2.2 financial validation and report update

Documentation only; weights remain v0.2.0. No improved financial checkpoint is released. ScenarioView report · Validation notes. MiDM baseline action accuracy was 54.37%; tested residual variants did not justify replacement. These are reused retrospective KOSPI decisions, not live signals or general benchmark gains.

v0.2.1 benchmark and architecture analysis

Documentation update; this repository's existing weights and inference code are unchanged. Full benchmark atlas · GitHub release · Version DOI.

Matched size and performance

Architecture comparison

The atlas separates matched measurements from historical/published/API comparisons. JevBench public accuracy is not its official composite. Missing MiDM DeepSWE and local Terminal-Bench measurements remain explicit; their candidate pools were not found. Unknown proprietary parameter counts are not guessed. SQL/Python regressions remain documented.

Public cohort comparison

Version v0.2.0. This repository contains the LoRA adapter, a linear pointer head and a self-contained Python loader. It scores the provided options jointly in one forward pass; it does not generate SQL, Python or explanations. The API accepts choice, yes/no noul, and ordinal score questions. Outputs are option probabilities, not guaranteed calibrated confidence.

Use

Install requirements.txt, download this repository at revision v0.2.0, and run from that directory:

from midm import MiDM
model = MiDM.from_pretrained(".")  # downloads the base separately; NF4 by default
result = model.predict(
    state="A receipt is required for a refund. The customer has no receipt.",
    questions={"action": {"type": "choice", "instructions": "Choose the next action.",
        "criteria": {"approve": "Approve the refund.", "review": "Request a receipt or review."}}},
)
print(result)

A compatible NVIDIA GPU and the listed dependencies are required. The configured maximum input is 4096 tokens; individual options are capped at 192 tokens. Long states are truncated while retaining the head and tail. All inference can run offline once the package and base are cached. Apache-2.0 applies to this adapter and code; base model and training data retain their own terms. This repository includes no training data.

MiDM v0.2.0 — 9B checkpoint and completed evaluation results

Adds the Qwen3.5-9B-Base LoRA adapter and learned pointer head, self-contained loader and verified 4096-token inference configuration. The original 4B release remains available. This model scores supplied options; it does not generate SQL, Python or report prose. The report writer remains Qwen3 8B.

Matched primary evaluation

NF4/BF16 on RTX 3090, max input 4096, batch token budget 4096, TTA=1; same items and ordering.

Suite N 4B 9B
Typed Decisions test 2000 78.85% 79.30%
Kev transfer test 764 79.84% 82.59%
JevBench original 72 91.67% 97.22%
JevBench easy 48 100.00% 100.00%
JevBench hard 111 45.95% 54.05%

These are retrospective hard-label accuracies, not official JevBench full/sealed composite scores. The separate 600-question Typed Decisions holdout is a development/model-selection split; it is not the public 2000-question test. Earlier development results used their recorded configurations and must not be mixed into this matched table.

Additional completed tests and limitations

  • CLM released DeepSWE verifier replay: 31/38 (81.58%) reproduced. The 13 mixed-outcome tasks give exact one-sided p=0.062 against random selection. Numerical reproduction succeeded; significance at 5% was not established.
  • MiDM DeepSWE and local Terminal-Bench remain unmeasured: matching raw candidate traces/evaluation artifacts were not found. No official benchmark score is substituted.
  • SQL/Python best-of-N: 4B → 9B Spider 78.34% → 77.27%, SQL holdout 81.66% → 81.77%, Python hard80 28.75% → 25.00%. Stored pools, dev-selected blend weights; no consistent gain.
  • Learned routing: JevBench public 231 accuracy 77.06% MiDM, 79.65% router, 81.82% local Qwen3-4B reasoning alone. Router not promoted; no non-inferiority or latency-saving claim.
  • Concurrency: two short-input 4B processes nearly doubled throughput; 2048-token 4B and 4096-token 9B lost throughput. Small synthetic single-trial pilots, not endurance or minimum-memory certification.
  • Same-date report integration: 4B/9B candidate-order agreement 1/4 versus 3/4, 19/19 schema-valid sections each, one unknown fact ID each. No semantic-quality or financial-outcome superiority established; only aggregate diagnostics are released.

Training: QLoRA NF4, rank16/alpha32, one epoch, seed0, learning rate2e-4, training context1024; td_train, kev_train, breadth_v1, pp6_new. Inference context4096 matches evaluation. The earlier 1024-token release candidate was not published; package verification uses the final loader. Base weights are separate and retain their own license, as do training datasets. No raw questions, answers, financial data, reports or credentials are included.

Model: https://huggingface.co/yunicro/MiDM-9B-q35-e1-bx/tree/v0.2.0

Code and full evidence: https://github.com/YeoHoonYun/midm-decision-models/releases/tag/v0.2.0

Zenodo versioning retains concept DOI https://doi.org/10.5281/zenodo.23084583 . The v0.1.0 version DOI does not archive this new release.

Version DOI: https://doi.org/10.5281/zenodo.23162744

Packaged-model verification

The final package reproduces the original development evaluation exactly: 513/600 (85.50%), identical choices on 600/600 questions and maximum absolute probability difference 0.0. The original protocol uses NF4/BF16, 4096-token context, 4096 padded tokens per batch, attention masks and TTA=1. Adapter weights, pointer head and encoded inputs match the source exactly. This is package reproducibility, not a new independent test.

A separate single-question, unmasked check gave 509/600 (84.83%) and failed its original tolerance gate. It is retained rather than hidden. That verification bypassed the public API batching/mask path. The matched check changes batching and mask together, so their individual effects were not isolated. Do not assume identical predictions across serving configurations.

Application case: scheduled internal financial reports

ScenarioView application case ? Live reports. Fixed editions at 18:30 KST after the Korean session and 09:00 KST after the US session replace repeated data checks. Each edition generates the integrated report set once from stored internal data using local MiDM. Generator scripts stay local; weights are unchanged.

Internal-only multi-asset report application

Auto by country · Historical validation including B3 · 한국어 reports · English reports · Architecture · Report manifest. New local MiDM inference covers KOSPI, S&P 500, and five internally ranked market-cap company representatives per market. Generation uses only stored internal inputs and local model execution, with zero external data/model calls. Publishing is separate. Source dates and ranking caveats are explicit; this is not proven asset-specific trading alpha. Generation/scheduling/publishing scripts are not uploaded. Model weights remain unchanged.

Downloads last month
73
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for yunicro/MiDM-9B-q35-e1-bx

Adapter
(57)
this model