kimsan0622 commited on
Commit
0ec5aa8
·
verified ·
1 Parent(s): db1776c

Publish camera-ready metadata reproduction and validation scope

Browse files
Files changed (2) hide show
  1. README.md +5 -1
  2. manifest.json +4 -3
README.md CHANGED
@@ -128,8 +128,12 @@ The API returns a wrapper object with `id`, `object`, `created`, `model`, and `r
128
 
129
  ## Diagnose your own model
130
 
131
- Use the [complete KoEVD own-model workflow](https://github.com/KETI-NLP/KoEVD/blob/main/docs/OWN_MODEL_GUIDE.md) to generate your model’s responses, measure Task 1–5 behavior, and produce HTML/JSON/CSV reports of metadata-specific strengths, weaknesses, strategies and source-matched tool-selection gaps. The current GitHub source adds this pipeline; the original v0.1.0 code tag does not include it. A subsequent [live B200 pipeline test](https://github.com/KETI-NLP/KoEVD/blob/main/docs/PIPELINE_VALIDATION_RESULTS.md) completed 40/959 actual generation requests with unmodified code. It also found batch-dependent strategy outputs and concerning safety judgments; execution success does not establish measurement accuracy.
132
 
133
  ## Observed batch sensitivity — 2026-09-15
134
 
135
  The [live pipeline test](https://github.com/KETI-NLP/KoEVD/blob/main/docs/PIPELINE_VALIDATION_RESULTS.md) completed inference, but strategy sets and raw outputs differed at batch sizes 2 and 4 for 5/191 responses (2.62%). One change was `e_humour` to `unsafe`; unsafe-marker counts were 22 versus 23. The cause is unresolved and no runtime patch is established. Record batch size, input ordering, precision and environment; exact batch invariance is not guaranteed. The unsafe marker is a measurement proxy and should not automatically replace the separate safety classifier output.
 
 
 
 
 
128
 
129
  ## Diagnose your own model
130
 
131
+ Use the [complete KoEVD own-model workflow](https://github.com/KETI-NLP/KoEVD/blob/main/docs/OWN_MODEL_GUIDE.md) to generate your model’s responses, measure Task 1–5 behavior, and produce HTML/JSON/CSV reports of metadata-specific strengths, weaknesses, strategies and source-matched tool-selection gaps. The current GitHub source adds this pipeline; the original v0.1.0 code tag does not include it. A subsequent [live B200 pipeline test](https://github.com/KETI-NLP/KoEVD/blob/main/docs/PIPELINE_VALIDATION_RESULTS.md) completed 40 smoke requests and 959 expanded generation requests with unmodified code. It also found batch-dependent strategy outputs and concerning safety judgments; execution success does not establish measurement accuracy.
132
 
133
  ## Observed batch sensitivity — 2026-09-15
134
 
135
  The [live pipeline test](https://github.com/KETI-NLP/KoEVD/blob/main/docs/PIPELINE_VALIDATION_RESULTS.md) completed inference, but strategy sets and raw outputs differed at batch sizes 2 and 4 for 5/191 responses (2.62%). One change was `e_humour` to `unsafe`; unsafe-marker counts were 22 versus 23. The cause is unresolved and no runtime patch is established. Record batch size, input ordering, precision and environment; exact batch invariance is not guaranteed. The unsafe marker is a measurement proxy and should not automatically replace the separate safety classifier output.
136
+
137
+ ## Later reproduction scope — 2026-09-15
138
+
139
+ A separate B200 check reproduced the assistant-safety classifier’s historical raw outputs and labels on all 532 original validation inputs at batches 2, 4, and 32, retaining macro/micro F1 0.9981/0.9981 and the same single error. It did not rerun response-strategy validation F1. On 191 new generated responses, suspicious all-safe safety outputs and 5/191 batch-dependent strategy label sets remain unresolved; no new human gold or new-response F1 was established. See the [reproduction scope](https://github.com/KETI-NLP/KoEVD/blob/main/docs/REPRODUCIBILITY.md). Earlier operational-test statements above describe their original 12/96-candidate sample, not this subsequent check.
manifest.json CHANGED
@@ -24,8 +24,8 @@
24
  },
25
  {
26
  "path": "README.md",
27
- "bytes": 8805,
28
- "sha256": "7f9cb712b7b9b039308ea7b8ab96ed50fa43a80f2f9ee02b6e3dbca39f702866"
29
  },
30
  {
31
  "path": "adapter_config.json",
@@ -87,5 +87,6 @@
87
  "bytes": 665,
88
  "sha256": "47bfa3e7727312946b29ac10d6dd0672d63cf7815b2a160b9523872040d2e536"
89
  }
90
- ]
 
91
  }
 
24
  },
25
  {
26
  "path": "README.md",
27
+ "bytes": 9555,
28
+ "sha256": "b0ef79a9e60705f27cfa1f847ecc5825fcbfff07d470a0b8fe09a9396cef5742"
29
  },
30
  {
31
  "path": "adapter_config.json",
 
87
  "bytes": 665,
88
  "sha256": "47bfa3e7727312946b29ac10d6dd0672d63cf7815b2a160b9523872040d2e536"
89
  }
90
+ ],
91
+ "documentation_update": "2026-09-15 camera-ready clarification; canonical data and model weights unchanged"
92
  }