thegovind commited on
Commit
8f77fd0
·
verified ·
1 Parent(s): 7c0a3b8

Card and licence corrections: audit scope, reproduction wording, MiMo lineage

Browse files
Files changed (2) hide show
  1. LICENSE.md +5 -6
  2. README.md +4 -4
LICENSE.md CHANGED
@@ -1,19 +1,18 @@
1
  blink model weights: non-commercial research licence
2
 
3
- These weights are a modification of Qwen/Qwen3.5-9B (Apache License 2.0, Copyright 2026 Alibaba Cloud;
4
- the full licence text is in LICENSE-Qwen). The modification (fine-tuning with LoRA adapters merged into
5
- the weights) was made by thegovind.
6
 
7
  The fine-tuning data included third-party datasets released under different terms, among them
8
  non-commercial (CC BY-NC 3.0 / 4.0) and share-alike (CC BY-SA 3.0 / 4.0) licences, and sources that
9
  state no licence (listed in README.md). Whether and how those terms apply to trained weights is unsettled.
10
 
11
  The author of the modification permits you to use, copy and run it for non-commercial research and
12
- evaluation only, provided this notice and LICENSE-Qwen are kept with every copy. No licence for commercial
13
  use is granted, and nothing here grants rights in any third-party data. You remain responsible for
14
- complying with the Apache License 2.0 for the Qwen base weights and with the terms of the upstream datasets.
15
 
16
  THE WEIGHTS ARE PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND.
17
 
18
  This is a personal research release. It is not an official product of any company, and it is not
19
- affiliated with TypeSafe AI, Alibaba Cloud or the Qwen team.
 
1
  blink model weights: non-commercial research licence
2
 
3
+ These weights are a modification of XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B (its model card declares MIT; see LICENSE-MiMo.md). It is based on Qwen/Qwen3.5-9B (Apache License 2.0; see LICENSE-Qwen).
4
+ The modification (fine-tuning with LoRA adapters merged into the weights) was made by thegovind.
 
5
 
6
  The fine-tuning data included third-party datasets released under different terms, among them
7
  non-commercial (CC BY-NC 3.0 / 4.0) and share-alike (CC BY-SA 3.0 / 4.0) licences, and sources that
8
  state no licence (listed in README.md). Whether and how those terms apply to trained weights is unsettled.
9
 
10
  The author of the modification permits you to use, copy and run it for non-commercial research and
11
+ evaluation only, provided LICENSE.md, LICENSE-MiMo.md and LICENSE-Qwen are kept with every copy. No licence for commercial
12
  use is granted, and nothing here grants rights in any third-party data. You remain responsible for
13
+ complying with MiMo's declared MIT terms and the Apache License 2.0 for the Qwen base weights and with the terms of the upstream datasets.
14
 
15
  THE WEIGHTS ARE PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND.
16
 
17
  This is a personal research release. It is not an official product of any company, and it is not
18
+ affiliated with TypeSafe AI, Xiaomi, Alibaba Cloud or the Qwen team.
README.md CHANGED
@@ -260,11 +260,11 @@ These are source-repository licences; they don't settle rights in every underlyi
260
 
261
  - **Scorer parity.** `blink.py` and the lab scorer agree within 1e-7 on this checkpoint (p99 |Δp| 1.4e-8); the published graft matches the evaluated adapter on JevBench's 231 public items (0 argmax changes, max |Δp| 0.0). The kit's per-request timer on 1,000 random suite requests served one at a time measured 68.3 ms median via HTTP vs 66.9 ms in-process. From a fresh Hub download and install at the pinned revision, `serve.py` with JevBench's stock adapter returned easy 48/48, standard 70/72 and hard 77/111, matching the evaluation's answers (0 argmax changes; max |Δp| about 0.03).
262
  - **Selection.** The prompt format came from earlier DI-S reads; this MiMo run used the sample as a pre-registered gate. The full suite followed that gate; it scored 56.60 on the 129,422 requests outside DI-S, which were not used for selection.
263
- - **Training overlap.** Public train splits also used by the 0.1 index: ContractNLI, iSarcasmEval, VAST, Amazon ESCI, Humicroedit, ChessBench (searchless_chess training positions; none of the 5,000 test positions), GSM8K (train split; solution-checking items). We also used ANLI and BANKING77 train splits; they're in the 0.1 suite but outside its index, and both are in the 0.2 panel. No identical suite test row was used; the overlap audit below covers shared passages.
264
  - **Partitions.** Public-source data included training and development partitions.
265
- - **Final-mixture audit.** Rechecked every question row (including teacher-written rows) against the complete 0.1 suite (132,422 requests) and JevBench's 231 public items. The checks used normalised text of at least 30 characters and 13-word passages. No public JevBench item matched; no chess position is shared.
266
- - **Suite overlap.** 16 BANKING77/VAST training rows share a 13-word passage with 31 suite requests: 23 of VAST's 3,006 and 8 of BANKING77's 3,080. Two VAST training posts are near-duplicates of a test post; none of these texts is identical. Dropping those requests leaves the index at 56.53 (VAST 0.7805 → 0.7803); BANKING77 is outside the index.
267
- - **Audit limits.** Semantic or pretraining overlap can't be ruled out; private JevBench items weren't available to check.
268
  - **Generated reasoning.** Our programs computed the labels for CRUXEval-style code and CLadder-style causal questions; no items from those benchmarks were used. We didn't reuse the suite's GSM8K distractors.
269
  - **Teacher documents.** We kept Qwen3.8-27B's documents only if a fresh blind solve by that same teacher agreed with the answer. That's an agreement filter, not independent verification.
270
  - The repo ships no benchmark items, GPQA text, JevBench items or teacher traces.
 
260
 
261
  - **Scorer parity.** `blink.py` and the lab scorer agree within 1e-7 on this checkpoint (p99 |Δp| 1.4e-8); the published graft matches the evaluated adapter on JevBench's 231 public items (0 argmax changes, max |Δp| 0.0). The kit's per-request timer on 1,000 random suite requests served one at a time measured 68.3 ms median via HTTP vs 66.9 ms in-process. From a fresh Hub download and install at the pinned revision, `serve.py` with JevBench's stock adapter returned easy 48/48, standard 70/72 and hard 77/111, matching the evaluation's answers (0 argmax changes; max |Δp| about 0.03).
262
  - **Selection.** The prompt format came from earlier DI-S reads; this MiMo run used the sample as a pre-registered gate. The full suite followed that gate; it scored 56.60 on the 129,422 requests outside DI-S, which were not used for selection.
263
+ - **Training overlap.** Public train splits also used by the 0.1 index: ContractNLI, iSarcasmEval, VAST, Amazon ESCI, Humicroedit, ChessBench (searchless_chess training positions; none of the 5,000 test positions), GSM8K (train split; solution-checking items). We also used ANLI and BANKING77 train splits; they're in the 0.1 suite but outside its index, and both are in the 0.2 panel. The audit below reports what was checked and any shared passages.
264
  - **Partitions.** Public-source data included training and development partitions.
265
+ - **Final-mixture audit.** Rechecked every question row (including teacher-written rows) against the complete 0.1 suite (132,422 requests) and JevBench's 231 public items. The checks looked for exact matches of normalised strings of at least 30 characters in any field and shared 13-word passages in each row's question text (instructions, state.question, state.code). Strings or passages seen in 20 or more suite requests were treated as prompt templates and ignored. No public JevBench item matched under these checks; a separate position check found no shared chess positions.
266
+ - **Suite overlap.** 16 BANKING77/VAST training rows share a 13-word passage with 31 suite requests: 23 of VAST's 3,006 and 8 of BANKING77's 3,080. Two VAST training posts are near-duplicates of a test post, but these rows had no exact normalised-text match under the audit. Dropping those requests leaves the index at 56.53 (VAST 0.7805 → 0.7803); BANKING77 is outside the index.
267
+ - **Audit limits.** The 13-word passage check didn't search long-document bodies or option text. Semantic or pretraining overlap can't be ruled out, and private JevBench items weren't available to check.
268
  - **Generated reasoning.** Our programs computed the labels for CRUXEval-style code and CLadder-style causal questions; no items from those benchmarks were used. We didn't reuse the suite's GSM8K distractors.
269
  - **Teacher documents.** We kept Qwen3.8-27B's documents only if a fresh blind solve by that same teacher agreed with the answer. That's an agreement filter, not independent verification.
270
  - The repo ships no benchmark items, GPQA text, JevBench items or teacher traces.