- Qwen2.5-7B-Instruct — Converge Collective v2
- Stage 1 — v1 (9 source models)
- Stage 2 — v1 → v2 (5 source checkpoints)
- V&V board — v2 vs Qwen2.5-7B-Instruct (one pinned harness, identical items)
- 7B-class comparison, by axis and tier
- Tier A — Math: GSM8K (HELM Lite v1.13.0; 1,000 problems, 5-shot, greedy,
stop=none, final-number exact match) - Tier A — Knowledge: MMLU (HELM MMLU v1.13.0; 14,042 questions, 57 subjects, 5-shot, single-letter answer, macro average)
- Tier B — Code: HumanEval+ and MBPP+ (EvalPlus leaderboard snapshot, 2026-09-27; greedy pass@1)
- Tier C — axes with no matched public reference
- Tier A — Math: GSM8K (HELM Lite v1.13.0; 1,000 problems, 5-shot, greedy,
- Honest framing
- Limitations
- License
- Corrections
- Stage 1 — v1 (9 source models)
Qwen2.5-7B-Instruct — Converge Collective v2
The second state of the converge collective: v1 grown by Optitransfer's converge model-merging platform. The platform has three parts:
- crdt-merge: the convergent merge layer.
- ACFA: accountable aggregation with verifiable receipts (arXiv:2607.10305).
- E4: the trust engine.
These are covered by UK patent applications GB2607132.4 and GB2608127.3. This release was built with the platform's merge path; ACFA receipts and E4 trust scoring were not applied to it.
v2 is a 7.6B-parameter, training-free weight-space merge. It loads and prompts exactly like Qwen2.5-7B-Instruct: same architecture, tokenizer, chat template and configuration files.
- Pre-registered headline: v2 improves on 2 of 12 benches (both calibrated), regresses on 2 (MATH Level 5 and MMLU-Pro), and shows no measurable difference on 8. Nothing is netted.
- Strongest result: on GSM8K under HELM's own protocol, v2 scores 87.0 vs 83.3 for Qwen2.5-7B-Instruct on the same 1,000 prompts. That is +3.7 points, paired 95% CI [+1.5, +5.8], exact McNemar p = 0.001. The harness reproduces HELM's published score for the base model (83.3 vs 83.0).
- Weights:
model.safetensorsat revisioneb3683bf, sha2561c0d2060b805878a3c70126befd0d2ec128543acfac423cdf86ed85a11c01169. Every result below is pinned to these bytes. - Comparison with other 7B-class models, axis by axis: calibrated HELM reference rows for GSM8K and MMLU, and EvalPlus leaderboard rows for code (see below).
Stage 1 — v1 (9 source models)
| # | Source model |
|---|---|
| 1 | Qwen/Qwen2.5-7B-Instruct (base model) |
| 2 | mistralai/Mistral-7B-Instruct-v0.3 |
| 3 | microsoft/Phi-3-mini-4k-instruct |
| 4 | microsoft/phi-2 |
| 5 | HuggingFaceTB/SmolLM2-1.7B-Instruct |
| 6 | ibm-granite/granite-3.0-2b-instruct |
| 7 | EleutherAI/pythia-2.8b |
| 8 | EleutherAI/pythia-1.4b |
| 9 | facebook/opt-2.7b |
Stage 2 — v1 → v2 (5 source checkpoints)
| # | Source checkpoint | Note |
|---|---|---|
| 1 | Qwen/Qwen2.5-Coder-7B-Instruct | |
| 2 | Qwen/Qwen2.5-7B | pre-trained base model |
| 3 | unsloth/Qwen2.5-7B-Instruct | re-upload; matched Qwen2.5-7B-Instruct on every tensor compared |
| 4 | unsloth/Qwen2.5-Coder-7B-Instruct | re-upload; byte-identical to Qwen/Qwen2.5-Coder-7B-Instruct |
| 5 | Qwen/Qwen2.5-7B-Instruct-1M | its long-context configuration is not used; v2 keeps the 32k configuration |
We do not attribute any result below to any particular source model.
V&V board — v2 vs Qwen2.5-7B-Instruct (one pinned harness, identical items)
Both models ran on the same engine, with the same prompts, items and decoding. Scores are in percent, on each benchmark's primary metric.
- "v2 ✓ / seed ✓" gives the two discordant counts: items only v2 got right, and items only the base model got right.
- Reading rule, fixed before the run: a difference is an improvement or a regression only if the paired-bootstrap 95% CI excludes 0 and the exact McNemar p < 0.05. Anything else is "no measurable difference".
- ✓ Calibrated means our harness reproduces the published score for Qwen2.5-7B-Instruct under that exact protocol.
- Not calibrated means no published reference with a matching protocol exists. The absolute score is specific to this harness. The paired difference is valid on every row.
| Axis | Benchmark (protocol) | n | Qwen2.5-7B-Instruct | v2 | Δ | 95% CI | p | v2 ✓ / seed ✓ | Reading | Calibrated |
|---|---|---|---|---|---|---|---|---|---|---|
| Math | GSM8K — HELM Lite replay (5-shot, greedy, final-number match) | 1,000 | 83.3 | 87.0 | +3.7 | [+1.5, +5.8] | 0.0011 | 80 / 43 | improvement | ✓ HELM Lite v1.13.0: 83.0 (ours 83.3; tol. ±2.38) |
| Knowledge | MMLU — HELM replay (5-shot, single letter; macro over 57 subjects) ¹ | 14,042 | 72.92 | 73.40 | +0.47 | [+0.11, +1.03] (items) ¹ | 0.016 (items) ¹ | 576 / 496 | improvement ¹ | ✓ HELM MMLU v1.13.0: 72.87 (ours 72.92; tol. ±1.0) |
| Math | MATH Level 5 — 0-shot CoT, math-verify | 1,324 | 54.31 | 48.56 | −5.74 | [−8.23, −3.32] | 6.0×10⁻⁶ | 101 / 177 | regression | not calibrated |
| Knowledge | MMLU-Pro — 5-shot CoT, exact match | 12,032 | 56.77 | 54.88 | −1.89 | [−2.64, −1.16] | 6.5×10⁻⁷ | 920 / 1,147 | regression | not calibrated |
| Math | GSM8K — lm-eval 5-shot, flexible extract | 1,319 | 87.49 | 88.32 | +0.83 | [−0.83, +2.43] | 0.375 | 69 / 58 | no measurable difference | not calibrated |
| Math | GSM8K — lm-eval 0-shot CoT, flexible extract | 1,319 | 78.77 | 79.45 | +0.68 | [−0.76, +2.12] | 0.426 | 55 / 46 | no measurable difference | not calibrated |
| Math | MATH-500 — 0-shot CoT, math-verify | 500 | 74.8 | 73.6 | −1.2 | [−4.2, +1.8] | 0.519 | 27 / 33 | no measurable difference | not calibrated |
| Math | AIME 2024+2025 — avg@8, T 0.6 (exploratory) | 60 | 9.58 | 8.13 | −1.46 | [−3.96, +1.04] | — | — | no measurable difference | not calibrated |
| Code | HumanEval+ — EvalPlus, greedy pass@1 | 164 | 75.00 | 78.05 | +3.05 | [−1.83, +7.93] | 0.332 | 11 / 6 | no measurable difference | not calibrated ² |
| Code | MBPP+ — EvalPlus, greedy pass@1 | 378 | 68.78 | 70.63 | +1.85 | [−1.59, +5.29] | 0.382 | 27 / 20 | no measurable difference | not calibrated ² |
| Instruction following | IFEval — prompt-level strict | 541 | 71.53 | 73.01 | +1.48 | [−1.85, +4.62] | 0.422 | 42 / 34 | no measurable difference ³ | not calibrated |
| Reasoning | ARC-Challenge — 25-shot, acc_norm | 1,172 | 66.98 | 65.87 | −1.11 | [−2.56, +0.34] | 0.177 | 33 / 46 | no measurable difference | not calibrated |
Scoreline (pre-registered): 2 improvements / 2 regressions / 8 no measurable difference. Every benchmark on the V&V board is reported. Regressions are not netted against improvements.
¹ MMLU. The scores are the macro average over 57 subjects. That is the calibrated aggregate on HELM's scale, and its Δ is +0.47. The pre-registered paired test runs on the 14,042 items: 72.11 vs 71.54 item-level exact match, +0.57, CI [+0.11, +1.03], p = 0.016.
Part of the gain is answer format. The base model gave 44 answers that are not one of the option letters A–D; v2 gave 3. v2 answered 23 of those 44 items correctly, which accounts for 0.16 of the 0.57 points. Two further checks are not significant:
- On the 13,998 items where both models gave a valid option letter, the difference is +0.41 (p = 0.084).
- A post-hoc subject-stratified test of the macro average gives CI [−0.07, +1.02].
Treat the MMLU improvement as small and format-sensitive.
² Code. HumanEval+ and MBPP+ were scored with EvalPlus 0.3.1's own evaluator, the same tool as the EvalPlus leaderboard. That leaderboard has no Qwen2.5-7B-Instruct row, so the harness cannot be calibrated against it.
³ IFEval. lm-eval 0.4.13's IFEval scorer is not fully deterministic: one checker substitutes a random letter, and language detection is unseeded. Rescoring moves the prompt-strict difference within about +0.9 to +1.8 points. The reading is "no measurable difference" either way.
How robust the GSM8K-HELM gain is. Removing every item where either model hit the 400-token cap leaves 968 items; there the gain is +3.6 points (p = 0.001). Mean output length is essentially unchanged: 198.0 tokens for v2 vs 199.3 for the base model.
The gain is tied to HELM's protocol. Under lm-eval's 5-shot and 0-shot CoT GSM8K protocols, the flexible-extract primary metric shows no measurable difference. The 5-shot strict-match scorer shows +2.73, mostly from formatting (see below). Do not read the HELM result as a general "better at math" claim: MATH Level 5 regresses.
Every other banked metric (secondary; many tests, no multiplicity correction; no claim rests on these)
| Benchmark — metric | Qwen2.5-7B-Instruct | v2 | Δ | 95% CI | p |
|---|---|---|---|---|---|
| MMLU (HELM) — item-level exact match | 71.54 | 72.11 | +0.57 | [+0.11, +1.03] | 0.016 |
GSM8K 5-shot — strict match (#### N format) |
83.32 | 86.05 | +2.73 | [+0.83, +4.62] | 0.006 |
| GSM8K 5-shot — numeric | 89.54 | 88.78 | −0.76 | [−2.27, +0.68] | 0.373 |
| GSM8K 5-shot — math-verify | 89.23 | 88.78 | −0.45 | [−1.97, +1.06] | 0.627 |
| GSM8K 0-shot CoT — numeric | 82.41 | 82.87 | +0.45 | — | — |
| GSM8K 0-shot CoT — math-verify | 82.03 | 83.02 | +0.99 | [−0.53, +2.50] | 0.237 |
| HumanEval — base tests | 81.71 | 84.76 | +3.05 | [−1.22, +7.93] | 0.302 |
| MBPP — base tests | 79.10 | 82.80 | +3.70 | [+0.26, +7.14] | 0.049 |
| MATH Level 5 — boxed-answer match | 50.68 | 45.32 | −5.36 | [−7.78, −2.95] | 1.5×10⁻⁵ |
| MATH-500 — boxed-answer match | 69.0 | 67.2 | −1.8 | [−4.8, +1.2] | 0.306 |
| ARC-Challenge — acc (unnormalised) | 64.33 | 62.54 | −1.79 | [−3.24, −0.34] | 0.022 |
| IFEval — prompt-level loose | 73.57 | 75.79 | +2.22 | [−0.92, +5.36] | 0.207 |
| IFEval — instruction-level strict / loose | 78.90 / 80.70 | 80.58 / 82.49 | +1.68 / +1.80 | — | — |
| AIME 2024 / AIME 2025 — avg@8 (exploratory) | 12.08 / 7.08 | 8.75 / 7.50 | −3.33 / +0.42 | — | — |
| AIME 2024+2025 — pass@8 (exploratory) | 28.33 | 25.00 | −3.33 | — | 0.727 |
Two notes on these rows:
- The GSM8K 5-shot strict-match gain is mostly formatting. The base model left 59 answers without a
#### Nline; v2 left 30. - The GSM8K 0-shot CoT strict-match scorer is excluded. Its answer pattern never matches either model's output, so it scores 0 for both.
Majority voting on GSM8K. 32 samples at T = 0.7 per problem, for both models. No paired test was run on these rows.
| Protocol — scorer | maj@8 seed | maj@8 v2 | Δ | maj@32 seed | maj@32 v2 | Δ |
|---|---|---|---|---|---|---|
| 5-shot — flexible | 91.43 | 92.19 | +0.76 | 92.27 | 92.65 | +0.38 |
| 5-shot — strict | 91.51 | 92.12 | +0.61 | 92.12 | 92.57 | +0.45 |
| 5-shot — numeric | 92.87 | 92.27 | −0.61 | 93.48 | 92.87 | −0.61 |
| 5-shot — math-verify | 92.72 | 92.19 | −0.53 | 93.48 | 92.72 | −0.76 |
| 0-shot CoT — flexible | 81.27 | 80.67 | −0.61 | 81.05 | 81.12 | +0.08 |
| 0-shot CoT — numeric | 84.99 | 84.15 | −0.83 | 84.76 | 84.69 | −0.08 |
| 0-shot CoT — math-verify | 85.37 | 84.15 | −1.21 | 84.99 | 84.61 | −0.38 |
Voting lifts both models by similar amounts. On the numeric and math-verify scorers the base model's vote is higher.
7B-class comparison, by axis and tier
Model names, sizes and scores appear exactly as each source table publishes them. The tested claim on every axis is the paired difference between v2 and Qwen2.5-7B-Instruct. Other models' rows are reference points, not a ranking.
| Tier | What it means | Axes |
|---|---|---|
| A: calibrated, protocol-matched | Our harness replays the reference's exact prompts and reproduces its published score for Qwen2.5-7B-Instruct. The other rows ran under the identical run specification. | Math (GSM8K-HELM), Knowledge (MMLU-HELM) |
| B: same tool, not calibrated | We used the leaderboard's own evaluator, but the leaderboard has no Qwen2.5-7B-Instruct row to calibrate against. | Code (HumanEval+, MBPP+) |
| C: no matched public reference | We found no published table that runs other models under our protocol. The 7B comparison is Qwen2.5-7B-Instruct on identical items: the board above. | Math (MATH-500, MATH L5, AIME, GSM8K via lm-eval), Knowledge (MMLU-Pro), Instruction following (IFEval), Reasoning (ARC-C) |
Tier A — Math: GSM8K (HELM Lite v1.13.0; 1,000 problems, 5-shot, greedy, stop=none, final-number exact match)
| Model | Size | GSM8K | Source |
|---|---|---|---|
| Converge Collective v2 | 7B | 87.0 | our replay |
| Qwen2.5-7B-Instruct | 7B | 83.3 | our replay, same harness as v2 |
| Qwen2.5 Instruct Turbo (7B) | 7B | 83.0 | HELM published row |
| Llama 3.1 Instruct Turbo (8B) | 8B | 79.8 | HELM published row |
| Gemma 2 Instruct (9B) | 9B | 76.2 | HELM published row |
| Mistral Instruct v0.3 (7B) | 7B | 53.8 | HELM published row |
These are all the table's 6–10B rows that share the stop=none run specification. Eight other 6–9B rows ran with HELM's default stop sequence, so they are not protocol-matched and are excluded: Qwen1.5 (7B), Gemma (7B), Llama 3 (8B), Mistral v0.1 (7B), Yi (6B), Llama 2 (7B), Falcon (7B) and OLMo (7B).
Tier A — Knowledge: MMLU (HELM MMLU v1.13.0; 14,042 questions, 57 subjects, 5-shot, single-letter answer, macro average)
| Model | Size | MMLU | Source |
|---|---|---|---|
| Phi-3 (7B) (HELM label; microsoft/phi-3-small-8k-instruct) | 7B | 75.68 | HELM published row |
| Converge Collective v2 | 7B | 73.40 | our replay |
| Qwen2.5-7B-Instruct | 7B | 72.92 | our replay, same harness as v2 |
| Qwen2.5 Instruct Turbo (7B) | 7B | 72.87 | HELM published row |
| Mistral Instruct v0.3 (7B) | 7B | 59.91 | HELM published row |
| Llama 3.1 Instruct Turbo (8B) | 8B | 56.06 | HELM published row |
These are all the 6–10B rows whose 57 per-subject run specifications are identical to the Qwen2.5-7B-Instruct run we replay. Nine other 6–9B rows use different prompt instructions (0 of 57 subjects match), so they are excluded: Gemma 2 (9B), Llama 3 (8B), Gemma (7B), Yi (6B), Qwen1.5 (7B), Mistral v0.1 (7B), OLMo 1.7 (7B), Llama 2 (7B) and OLMo (7B).
About the HELM rows. They are frozen results published by the HELM project, obtained through Together AI endpoints. The exception is Phi-3 (7B), which ran on HELM's Hugging Face deployment. Our numbers come from a self-run replay of HELM's exact request prompts, using the published weights in bf16 on vLLM. The replay reproduces HELM's aggregate for Qwen2.5-7B-Instruct. Per item, it agrees with HELM's own record on 939 of 1,000 GSM8K items and 13,822 of 14,042 MMLU items, so serving differences exist. HELM has not evaluated, listed or endorsed this model.
Tier B — Code: HumanEval+ and MBPP+ (EvalPlus leaderboard snapshot, 2026-09-27; greedy pass@1)
Rows shown:
- the five highest-scoring 6–9B rows on each benchmark, with ties included;
- every general-purpose 7–8B chat model that EvalPlus evaluated with its instruction prompt (EvalPlus's
promptedflag); - Qwen2.5-Coder-7B-Instruct, a v2 source model. Its figure is vendor-reported; it is not on the snapshot.
gemma-1.1-2b-it is excluded: the snapshot lists it at 7.0B, but it is a 2B model. The complete 6–9B list follows below.
| Model | Size | HumanEval+ | MBPP+ | Source |
|---|---|---|---|---|
| Qwen2.5-Coder-7B-Instruct (code specialist) | 7B | 84.1 | 71.7 | vendor report, arXiv:2409.12186 Table 16 |
| CodeQwen1.5-7B-Chat | 7B | 78.7 | 69.0 | EvalPlus leaderboard |
| Converge Collective v2 | 7B | 78.05 | 70.63 | our run, EvalPlus 0.3.1 |
| OpenCoder-8B-Instruct | 8B | 77.4 | 71.4 | EvalPlus leaderboard |
| Qwen2.5-7B-Instruct | 7B | 75.00 | 68.78 | our run, same harness as v2 |
| Artigenz-Coder-DS-6.7B | 6.7B | 72.6 | 69.6 | EvalPlus leaderboard |
| OpenCodeInterpreter-DS-6.7B | 6.7B | 72.0 | 66.4 | EvalPlus leaderboard |
| DeepSeek-Coder-6.7B-instruct | 6.7B | 71.3 | 65.6 | EvalPlus leaderboard |
| DeepSeek-Coder-7B-instruct-v1.5 | 7B | 71.3 | 62.2 | EvalPlus leaderboard |
| Magicoder-S-DS-6.7B | 6.7B | 71.3 | 69.0 | EvalPlus leaderboard |
| OpenChat-3.5-7B-0106 | 7B | 67.7 | 54.5 | EvalPlus leaderboard |
| Llama3.1-8B-instruct | 8B | 62.8 | 55.6 | EvalPlus leaderboard |
| Llama3-8B-instruct | 8B | 56.7 | 54.8 | EvalPlus leaderboard |
| Mistral-7B-Instruct-v0.2 | 7B | 36.0 | 37.0 | EvalPlus leaderboard |
| gemma-1.1-7b-it | 7B | 35.4 | 45.0 | EvalPlus leaderboard |
| xDAN-L1-Chat-RL-v1-7B | 7B | 32.9 | 41.3 | EvalPlus leaderboard |
| gemma-7b-it | 7B | 25.0 | 36.8 | EvalPlus leaderboard |
v2 does not differ measurably from its base model on either code benchmark. It is not a code specialist.
All 43 rows of the EvalPlus snapshot with size 6–9B
| Model (EvalPlus label) | Size (B) | HumanEval+ | MBPP+ | EvalPlus prompted |
|---|---|---|---|---|
| CodeQwen1.5-7B-Chat | 7 | 78.7 | 69.0 | yes |
| OpenCoder-8B-Instruct | 8 | 77.4 | 71.4 | yes |
| Artigenz-Coder-DS-6.7B | 6.7 | 72.6 | 69.6 | yes |
| OpenCodeInterpreter-DS-6.7B | 6.7 | 72.0 | 66.4 | yes |
| DeepSeek-Coder-6.7B-instruct | 6.7 | 71.3 | 65.6 | yes |
| DeepSeek-Coder-7B-instruct-v1.5 | 7 | 71.3 | 62.2 | yes |
| Magicoder-S-DS-6.7B | 6.7 | 71.3 | 69.0 | yes |
| WaveCoder-Ultra-6.7B | 7 | 69.5 | 63.5 | yes |
| Magicoder-S-CL-7B | 7 | 67.7 | 60.1 | yes |
| OpenChat-3.5-7B-0106 | 7 | 67.7 | 54.5 | yes |
| speechless-coder-ds-6.7B | 6.7 | 65.9 | 64.4 | yes |
| Llama3.1-8B-instruct | 8 | 62.8 | 55.6 | yes |
| Code-290k-6.7B-Instruct | 6.7 | 59.7 | — | yes |
| Llama3-8B-instruct | 8 | 56.7 | 54.8 | yes |
| codegemma-7b-it | 7 | 51.8 | 56.9 | yes |
| speechless-starcoder2-7b | 7 | 51.8 | 56.3 | yes |
| speechless-coding-7B-16k-tora | 7 | 50.6 | 50.6 | yes |
| CodeQwen1.5-7B | 7 | 45.7 | 60.8 | no |
| WizardCoder-Python-7B-V1.0 | 7 | 45.1 | 49.5 | yes |
| Mistral-codealpaca-7B | 7 | 42.1 | — | no |
| MistralHermes-CodePro-7B-v1 | 7 | 42.1 | 46.4 | yes |
| codegemma-7b | 7 | 41.5 | 52.4 | no |
| speechless-code-mistral-7B-v1.0 | 7 | 41.5 | 48.7 | yes |
| DeepSeek-Coder-6.7B-base | 6.7 | 39.6 | 58.7 | no |
| Mistral-7B-Instruct-v0.2 | 7 | 36.0 | 37.0 | yes |
| CodeLlama-7B | 7 | 35.4 | 46.8 | no |
| gemma-1.1-7b-it | 7 | 35.4 | 45.0 | yes |
| xDAN-L1-Chat-RL-v1-7B | 7 | 32.9 | 41.3 | yes |
| StarCoder2-7B | 7 | 29.9 | — | no |
| Llama3-8B-base | 8 | 29.3 | 51.6 | no |
| gemma-7b | 7 | 28.7 | 43.4 | no |
| CodeGen-6B | 6 | 25.6 | 42.9 | no |
| gemma-7b-it | 7 | 25.0 | 36.8 | yes |
| CodeT5+-6B | 6 | 24.4 | 41.5 | no |
| Mistral-7B | 7 | 23.8 | 42.1 | no |
| Zephyr β-7B | 7 | 23.2 | 34.7 | no |
| StarCoderBase-7B | 7 | 21.3 | — | no |
| CodeGen2-7B | 7 | 17.7 | — | no |
| gemma-1.1-2b-it | 7 | 17.7 | 23.3 | yes |
| InCoder-6.7B | 6.7 | 12.2 | — | no |
| Vicuna-7B | 7 | 11.6 | — | no |
| GPT-J-6B | 6 | 11.0 | — | no |
| StableLM-7B | 7 | 2.4 | — | no |
Tier C — axes with no matched public reference
| Axis | Benchmarks on this board | 7B comparison available |
|---|---|---|
| Math | MATH-500, MATH Level 5, AIME 2024+2025, GSM8K (lm-eval 5-shot and 0-shot CoT) | Qwen2.5-7B-Instruct on identical items (board above) |
| Knowledge | MMLU-Pro | Qwen2.5-7B-Instruct on identical items (board above) |
| Instruction following | IFEval | Qwen2.5-7B-Instruct on identical items (board above) |
| Reasoning | ARC-Challenge | Qwen2.5-7B-Instruct on identical items (board above) |
We make no cross-model claim on these axes.
Honest framing
- A static merge is a trade, not a lift. v2 gains on GSM8K under HELM's protocol and, slightly, on MMLU. It loses on MATH Level 5 and MMLU-Pro. Eight benches show no measurable difference.
- What is tested. The tested claims are paired differences against Qwen2.5-7B-Instruct on identical items. Other models' HELM and EvalPlus rows are reference points, not a ranking.
- Self-reported results. The evaluation record, 60 files of results, raw per-item outputs and exports, is committed by Merkle root
0220d670dec55f8494a322cf2e5b247e8b426c27bc710e4b4bf757d105f7efb9(RFC 6962, SHA-256). The record is unsigned and not yet public. - How the paired statistics were produced. The board's pinned code computed the paired test for GSM8K-HELM, MATH-500, HumanEval+ and AIME. For the other rows, we computed the same test after the run from the banked per-item outputs, with the same method and seed. Exploratory runs outside the board are not reported here.
- Contamination. v2 was not built by optimising on any of these test sets. The source models' training data is third-party, and we cannot certify it free of benchmark contamination. v2 contains weights from source models that Qwen2.5-7B-Instruct does not, so pairing does not cancel exposure specific to them.
Research output under the converge two-tier charter — the people's tier.
Evaluation setup
Models: we evaluated two models on identical items.
- v2 at revision
eb3683bfa089b173e59ef129e3a6670ec4ae2ac3. - Qwen/Qwen2.5-7B-Instruct at revision
a09a35458c702b33eeacc393d103063234e8bc28, as the control.
Before the run, the served weight bytes of both were checked against their sha256.
- v2 at revision
Engine: vLLM 0.30.0, dtype bfloat16,
max_model_len32,768, seed 1234, on one NVIDIA A100-SXM4-80GB. HumanEval+ and MBPP+ ran through EvalPlus's own vLLM provider withmax_model_len2,048. Software: torch 2.13.0, transformers 5.17.0, lm-eval 0.4.13, EvalPlus 0.3.1, math-verify 0.9.0. The build was checked against these pins before any run.Prompt format:
- GSM8K-HELM and MMLU-HELM: HELM's own prompts, sent as one user turn.
- GSM8K 0-shot CoT, MATH, IFEval and AIME: the chat template.
- GSM8K 5-shot, ARC-Challenge and MMLU-Pro: plain completion, chat template off.
- HumanEval+ and MBPP+: EvalPlus's own prompting.
Where the chat template was used, no system prompt was passed, so the template's default applied.
Decoding: greedy (temperature 0, top_p 1, repetition penalty 1.0), except:
- AIME: 8 samples at T 0.6, top_p 0.95, up to 16,384 new tokens;
- majority-vote rows: 32 samples at T 0.7.
The repository's
generation_config.jsonsampling defaults were disabled.Generation budgets (max new tokens):
- GSM8K-HELM: 400
- GSM8K 5-shot and 0-shot CoT: 1,024
- HumanEval+ and MBPP+: 768
- MATH: 4,096
- IFEval: 1,280
- MMLU-Pro: 2,048
- MMLU-HELM: 1
- ARC-Challenge: loglikelihood scoring, no generation
Data:
- GSM8K and MMLU replay HELM's banked request files.
- Pinned Hugging Face dataset revisions: openai/gsm8k, HuggingFaceH4/MATH-500, EleutherAI/hendrycks_math, allenai/ai2_arc, google/IFEval, Maxwell-Jia/AIME_2024, math-ai/aime25 and TIGER-Lab/MMLU-Pro.
- HumanEval+ and MBPP+: EvalPlus's releases.
Statistics:
- Paired bootstrap, 10,000 resamples, 95% percentile CI.
- Exact two-sided McNemar test on the same items.
- The reading rule was fixed in the evaluation code before the run.
Checks:
- Every suite completed with full item coverage, identical item sets for both models, a passing decoding audit and no voided blocks.
- Every headline number was recomputed independently from the raw per-item outputs.
Evaluation date: 2026-09-28.
Behaviour differences worth knowing
Format compliance. Malformed or unparsed answers, base model → v2:
- MMLU, answers that are not an option letter A–D: 44 → 3
- GSM8K-HELM: 9 → 3
- GSM8K 5-shot answers without a
#### Nline: 59 → 30 - MMLU-Pro: 48 → 79 (the one suite where v2 does worse)
Less answer-repetition on MMLU-Pro. About 12% of the base model's responses repeat "the answer is (X)" five or more times after answering, often until the 2,048-token cap. For v2 it is under 2%. Total generated tokens: 5.23M for the base model vs 2.39M for v2.
More generation-cap hits on hard math:
- MATH Level 5: 33 vs 22 at 4,096 tokens
- MATH-500: 11 vs 4
- AIME: 21 vs 4 of 480 samples at 16,384 tokens
The MATH Level 5 regression persists on the items both models finished: −5.9 points on 1,274 items.
Item-level churn. On GSM8K-HELM, v2 fixed 80 items the base model got wrong, and got 43 items wrong that the base model got right. We have not tested whether that churn is stable across runs.
Quick start
v2 loads and prompts exactly like Qwen2.5-7B-Instruct. The board's greedy scores were produced in bfloat16.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "Optitransfer/Qwen2.5-7B-Instruct-converge-collective-v2"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(
repo, dtype=torch.bfloat16, device_map="auto" # use torch_dtype= on older transformers
)
messages = [{"role": "user", "content": "A train travels 180 km in 2.5 hours. What is its average speed in km/h?"}]
text = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tok(text, return_tensors="pt").to(model.device)
# Greedy, as evaluated. The repo's generation_config samples by default (T=0.7, top_p=0.8, top_k=20, repetition_penalty=1.05).
out = model.generate(**inputs, max_new_tokens=512, do_sample=False, repetition_penalty=1.0)
print(tok.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
from vllm import LLM, SamplingParams
llm = LLM(
model="Optitransfer/Qwen2.5-7B-Instruct-converge-collective-v2",
dtype="bfloat16",
max_model_len=32768,
generation_config="vllm", # ignore the repo's sampling defaults, as in our evaluation
)
params = SamplingParams(temperature=0.0, max_tokens=512)
out = llm.chat([{"role": "user", "content": "Explain the difference between a list and a tuple in Python."}], params)
print(out[0].outputs[0].text)
Model details
| Architecture | Qwen2ForCausalLM (Qwen2.5-7B-Instruct's), 28 layers, hidden size 3,584 |
| Parameters | 7,615,616,512 |
| Weights | one model.safetensors, 15,231,271,784 bytes, stored as float16 |
| Configured context | 32,768 tokens (max_position_embeddings; RoPE θ = 1,000,000; no rope scaling) |
| Unchanged from Qwen2.5-7B-Instruct | config.json, generation_config.json, tokenizer.json, tokenizer_config.json, all byte-identical |
| Training | none: no fine-tuning, no gradient training, no training data |
- dtype note.
config.jsondeclarestorch_dtype: bfloat16, unchanged from the base model, while the stored tensors are float16. All reported results loaded the weights in bfloat16. Float16 inference was not evaluated. - File hashes (sha256):
config.json:7463bb0ea78315365e6c6b74de4e73bbcc8359dfb0c5a737584e077d42c0b03cgeneration_config.json:3a8f9087e486054c8a4a08dae2e5a3ba62e23da212b5b8c08bc42cb983c3459ftokenizer.json:c0382117ea329cdf097041132f6d735924b697924d6f6fc3945713e96ce87539tokenizer_config.json:5b5d4f65d0acd3b2d56a35b56d374a36cbc1c8fa5cf3b3febbbfabf22f359583
- If the revision you load differs from
eb3683bf, check thatmodel.safetensorsstill has the sha256 above. A card-only commit changes the revision hash, not the weights.
Limitations
- Not a long-context model. It is configured for 32,768 tokens, and no long-context evaluation was run.
tokenizer_config.jsoninheritsmodel_max_length: 131072from the base model; that value is not v2's configured context. - Weaker than the base model on MATH Level 5 and MMLU-Pro (see the board).
- Not a code specialist. HumanEval+ and MBPP+ show no measurable difference from the base model.
- Identifies as Qwen. When no system message is given, the inherited chat template inserts "You are Qwen, created by Alibaba Cloud. You are a helpful assistant." The model may therefore describe itself as Qwen.
- Default decoding samples. The repository's generation settings sample (T 0.7, top_p 0.8, top_k 20, repetition penalty 1.05). The board's scores are greedy, except the AIME and majority-vote rows.
- Not evaluated for safety, bias or toxicity. It was evaluated in English only. It may produce incorrect, biased or harmful content; do not use it for high-stakes decisions without independent evaluation.
License
This model descends, through v1, from facebook/opt-2.7b. That model is distributed under the OPT-175B License Agreement, which limits use to non-commercial research. This model is therefore made available for non-commercial research use only, subject to three sets of terms:
- the OPT-175B License Agreement;
- the Apache-2.0 licenses of the Qwen, Mistral, SmolLM2, Granite and Pythia sources;
- the MIT licenses of microsoft/phi-2 and microsoft/Phi-3-mini-4k-instruct.
No rights are granted beyond those licenses. OPT-175B is licensed under the OPT-175B license, Copyright (c) Meta Platforms, Inc. All Rights Reserved.
model.safetensors is modified relative to every source model. config.json, generation_config.json, tokenizer.json and tokenizer_config.json are Qwen2.5-7B-Instruct's files, unmodified.
Not affiliated with or endorsed by Alibaba Cloud, the Qwen team, the HELM project or EvalPlus. "Qwen" is used only to identify the model this one is derived from.
Corrections
Earlier revisions of this card, up to and including eb3683bf, contained the following, now withdrawn; they should not be relied on:
- an
apache-2.0license declaration; - build-process statements;
- a lineage count;
- an evaluation table that is not hash-linked to these weights.
The repository's receipt.json belongs to v1 and does not describe this model. The weights are unchanged from revision eb3683bf.
- Downloads last month
- 1,181