yunicro commited on
Commit
5e7e925
·
verified ·
1 Parent(s): bb03572

v0.2.2 report and financial validation; unchanged model weights

Browse files
README.md CHANGED
@@ -8,6 +8,10 @@ tags: [midm, decision-model, typed-decisions, pointer-head, lora, qlora]
8
 
9
  # MiDM-9B-q35-e1-bx
10
 
 
 
 
 
11
  ## v0.2.1 benchmark and architecture analysis
12
 
13
  Documentation update; this repository's existing weights and inference code are unchanged. [Full benchmark atlas](benchmarks/v0.2.1/README.md) · [GitHub release](https://github.com/YeoHoonYun/midm-decision-models/releases/tag/v0.2.1) · [Version DOI](https://doi.org/10.5281/zenodo.23164651).
 
8
 
9
  # MiDM-9B-q35-e1-bx
10
 
11
+ ## v0.2.2 financial validation and report update
12
+
13
+ **Documentation only; weights remain v0.2.0. No improved financial checkpoint is released.** [ScenarioView report](https://yeohoonyun.github.io/midm-decision-models/) · [Validation notes](financial_validation/v0.2.2/README.md). MiDM baseline action accuracy was 54.37%; tested residual variants did not justify replacement. These are reused retrospective KOSPI decisions, not live signals or general benchmark gains.
14
+
15
  ## v0.2.1 benchmark and architecture analysis
16
 
17
  Documentation update; this repository's existing weights and inference code are unchanged. [Full benchmark atlas](benchmarks/v0.2.1/README.md) · [GitHub release](https://github.com/YeoHoonYun/midm-decision-models/releases/tag/v0.2.1) · [Version DOI](https://doi.org/10.5281/zenodo.23164651).
financial_validation/v0.2.2/README.md ADDED
@@ -0,0 +1,15 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # v0.2.2 — ScenarioView report and financial validation
2
+
3
+ Documentation/report release dated 2026-10-07. Model weights, tokenizer and loader are unchanged from v0.2.0; this is NOT an improved-accuracy model release.
4
+
5
+ - Publishes a mobile-friendly ScenarioView report using the previous local report layout styles and section order.
6
+ - Includes public macro observations with observation dates and source attribution, RT definitions, aggregate action metrics, and section-level provenance.
7
+ - Frozen MiDM action accuracy: 54.37%; B3+MiDM stacking: 50.93%; development-selected horizon/position residual: 47.49%. No candidate promoted.
8
+ - Evaluation: KOSPI, 2024-01-02 to 2026-08-20, 756 hypothetical decisions / 129 sampled dates, previously seen retrospective data; not independent trades or portfolio returns.
9
+ - No validated economic leading/coincident/lagging status. New prompt context packs are prepared locally but have not been inference-tested.
10
+ - No private source prices, per-date predictions, generated private narratives, checkpoint files or credentials published.
11
+ - The report is issued today but does not contain a freshly inferred today-market signal. No automatic daily schedule is configured.
12
+
13
+ Report: https://yeohoonyun.github.io/midm-decision-models/
14
+
15
+ The prior Zenodo DOI continues to refer to v0.2.1; no new Zenodo deposit is claimed.
financial_validation/v0.2.2/evaluation.json ADDED
@@ -0,0 +1,132 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "documentation_version": "0.2.2",
3
+ "weights_version": "0.2.0",
4
+ "report_date": "2026-10-07",
5
+ "market": "KOSPI",
6
+ "evaluation_period": [
7
+ "2024-01-02",
8
+ "2026-08-20"
9
+ ],
10
+ "unique_dates": 129,
11
+ "hypothetical_decisions": 756,
12
+ "robust_heads": {
13
+ "raw_MiDM": {
14
+ "n": 756,
15
+ "accuracy": 0.5436507936507936,
16
+ "balanced_accuracy": 0.4923740904551291,
17
+ "utility_regret_bp": 202.6737499060071,
18
+ "long_fraction": 0.8716931216931217
19
+ },
20
+ "always_long": {
21
+ "n": 756,
22
+ "accuracy": 0.5687830687830688,
23
+ "balanced_accuracy": 0.5,
24
+ "utility_regret_bp": 200.93760231754584,
25
+ "long_fraction": 1.0
26
+ },
27
+ "numeric_residual": {
28
+ "n": 756,
29
+ "accuracy": 0.4880952380952381,
30
+ "balanced_accuracy": 0.4828577543158796,
31
+ "utility_regret_bp": 254.72531339383275,
32
+ "long_fraction": 0.5357142857142857
33
+ },
34
+ "horizon_residual": {
35
+ "n": 756,
36
+ "accuracy": 0.4748677248677249,
37
+ "balanced_accuracy": 0.47753602511057214,
38
+ "utility_regret_bp": 264.7359928251329,
39
+ "long_fraction": 0.4775132275132275
40
+ },
41
+ "latent_residual": {
42
+ "n": 756,
43
+ "accuracy": 0.5052910052910053,
44
+ "balanced_accuracy": 0.4998287915537167,
45
+ "utility_regret_bp": 246.82435842755385,
46
+ "long_fraction": 0.5396825396825397
47
+ },
48
+ "random_features": {
49
+ "n": 756,
50
+ "accuracy": 0.48412698412698413,
51
+ "balanced_accuracy": 0.5172064488514767,
52
+ "utility_regret_bp": 237.01935194499768,
53
+ "long_fraction": 0.2619047619047619
54
+ }
55
+ },
56
+ "fusion": {
57
+ "raw_MiDM": {
58
+ "n": 756,
59
+ "accuracy": 0.5436507936507936,
60
+ "balanced_accuracy": 0.4923740904551291,
61
+ "long_fraction": 0.8716931216931217,
62
+ "utility_regret_bp": 202.6737499060071
63
+ },
64
+ "B3_logistic": {
65
+ "n": 756,
66
+ "accuracy": 0.5066137566137566,
67
+ "balanced_accuracy": 0.4972820659152518,
68
+ "long_fraction": 0.5674603174603174,
69
+ "utility_regret_bp": 263.8134549025119
70
+ },
71
+ "MiDM_calibrated": {
72
+ "n": 756,
73
+ "accuracy": 0.548941798941799,
74
+ "balanced_accuracy": 0.4951704950777571,
75
+ "long_fraction": 0.8902116402116402,
76
+ "utility_regret_bp": 201.03404296303668
77
+ },
78
+ "stacked": {
79
+ "n": 756,
80
+ "accuracy": 0.5092592592592593,
81
+ "balanced_accuracy": 0.5010914538450564,
82
+ "long_fraction": 0.5595238095238095,
83
+ "utility_regret_bp": 260.755558602786
84
+ },
85
+ "mix_0": {
86
+ "n": 756,
87
+ "accuracy": 0.5066137566137566,
88
+ "balanced_accuracy": 0.4972820659152518,
89
+ "long_fraction": 0.5674603174603174,
90
+ "utility_regret_bp": 263.8134549025119
91
+ },
92
+ "mix_0.25": {
93
+ "n": 756,
94
+ "accuracy": 0.5079365079365079,
95
+ "balanced_accuracy": 0.49844485661292626,
96
+ "long_fraction": 0.5687830687830688,
97
+ "utility_regret_bp": 263.4556812103133
98
+ },
99
+ "mix_0.5": {
100
+ "n": 756,
101
+ "accuracy": 0.5079365079365079,
102
+ "balanced_accuracy": 0.49844485661292626,
103
+ "long_fraction": 0.5687830687830688,
104
+ "utility_regret_bp": 263.4556812103133
105
+ },
106
+ "mix_0.75": {
107
+ "n": 756,
108
+ "accuracy": 0.5092592592592593,
109
+ "balanced_accuracy": 0.4984947924097589,
110
+ "long_fraction": 0.578042328042328,
111
+ "utility_regret_bp": 263.84772576633515
112
+ },
113
+ "mix_1": {
114
+ "n": 756,
115
+ "accuracy": 0.548941798941799,
116
+ "balanced_accuracy": 0.4951704950777571,
117
+ "long_fraction": 0.8902116402116402,
118
+ "utility_regret_bp": 201.03404296303668
119
+ },
120
+ "always_long": {
121
+ "n": 756,
122
+ "accuracy": 0.5687830687830688,
123
+ "balanced_accuracy": 0.5,
124
+ "long_fraction": 1.0,
125
+ "utility_regret_bp": 200.93760231754584
126
+ }
127
+ },
128
+ "promotion": false,
129
+ "private_rows_included": false,
130
+ "scope": "Retrospective action decisions, not RT case accuracy or portfolio returns",
131
+ "prompt_enrichment_inference": "not run"
132
+ }