Qwen2.5-7B-Instruct — Converge Collective v2

The second state of the converge collective: v1 grown by Optitransfer's converge model-merging platform. The platform has three parts:

  • crdt-merge: the convergent merge layer.
  • ACFA: accountable aggregation with verifiable receipts (arXiv:2607.10305).
  • E4: the trust engine.

These are covered by UK patent applications GB2607132.4 and GB2608127.3. This release was built with the platform's merge path; ACFA receipts and E4 trust scoring were not applied to it.

v2 is a 7.6B-parameter, training-free weight-space merge. It loads and prompts exactly like Qwen2.5-7B-Instruct: same architecture, tokenizer, chat template and configuration files.

  • Pre-registered headline: v2 improves on 2 of 12 benches (both calibrated), regresses on 2 (MATH Level 5 and MMLU-Pro), and shows no measurable difference on 8. Nothing is netted.
  • Strongest result: on GSM8K under HELM's own protocol, v2 scores 87.0 vs 83.3 for Qwen2.5-7B-Instruct on the same 1,000 prompts. That is +3.7 points, paired 95% CI [+1.5, +5.8], exact McNemar p = 0.001. The harness reproduces HELM's published score for the base model (83.3 vs 83.0).
  • Weights: model.safetensors at revision eb3683bf, sha256 1c0d2060b805878a3c70126befd0d2ec128543acfac423cdf86ed85a11c01169. Every result below is pinned to these bytes.
  • Comparison with other 7B-class models, axis by axis: calibrated HELM reference rows for GSM8K and MMLU, and EvalPlus leaderboard rows for code (see below).

Stage 1 — v1 (9 source models)

# Source model
1 Qwen/Qwen2.5-7B-Instruct (base model)
2 mistralai/Mistral-7B-Instruct-v0.3
3 microsoft/Phi-3-mini-4k-instruct
4 microsoft/phi-2
5 HuggingFaceTB/SmolLM2-1.7B-Instruct
6 ibm-granite/granite-3.0-2b-instruct
7 EleutherAI/pythia-2.8b
8 EleutherAI/pythia-1.4b
9 facebook/opt-2.7b

Stage 2 — v1 → v2 (5 source checkpoints)

# Source checkpoint Note
1 Qwen/Qwen2.5-Coder-7B-Instruct
2 Qwen/Qwen2.5-7B pre-trained base model
3 unsloth/Qwen2.5-7B-Instruct re-upload; matched Qwen2.5-7B-Instruct on every tensor compared
4 unsloth/Qwen2.5-Coder-7B-Instruct re-upload; byte-identical to Qwen/Qwen2.5-Coder-7B-Instruct
5 Qwen/Qwen2.5-7B-Instruct-1M its long-context configuration is not used; v2 keeps the 32k configuration

We do not attribute any result below to any particular source model.

V&V board — v2 vs Qwen2.5-7B-Instruct (one pinned harness, identical items)

Both models ran on the same engine, with the same prompts, items and decoding. Scores are in percent, on each benchmark's primary metric.

  • "v2 ✓ / seed ✓" gives the two discordant counts: items only v2 got right, and items only the base model got right.
  • Reading rule, fixed before the run: a difference is an improvement or a regression only if the paired-bootstrap 95% CI excludes 0 and the exact McNemar p < 0.05. Anything else is "no measurable difference".
  • ✓ Calibrated means our harness reproduces the published score for Qwen2.5-7B-Instruct under that exact protocol.
  • Not calibrated means no published reference with a matching protocol exists. The absolute score is specific to this harness. The paired difference is valid on every row.
Axis Benchmark (protocol) n Qwen2.5-7B-Instruct v2 Δ 95% CI p v2 ✓ / seed ✓ Reading Calibrated
Math GSM8K — HELM Lite replay (5-shot, greedy, final-number match) 1,000 83.3 87.0 +3.7 [+1.5, +5.8] 0.0011 80 / 43 improvement ✓ HELM Lite v1.13.0: 83.0 (ours 83.3; tol. ±2.38)
Knowledge MMLU — HELM replay (5-shot, single letter; macro over 57 subjects) ¹ 14,042 72.92 73.40 +0.47 [+0.11, +1.03] (items) ¹ 0.016 (items) ¹ 576 / 496 improvement ¹ ✓ HELM MMLU v1.13.0: 72.87 (ours 72.92; tol. ±1.0)
Math MATH Level 5 — 0-shot CoT, math-verify 1,324 54.31 48.56 −5.74 [−8.23, −3.32] 6.0×10⁻⁶ 101 / 177 regression not calibrated
Knowledge MMLU-Pro — 5-shot CoT, exact match 12,032 56.77 54.88 −1.89 [−2.64, −1.16] 6.5×10⁻⁷ 920 / 1,147 regression not calibrated
Math GSM8K — lm-eval 5-shot, flexible extract 1,319 87.49 88.32 +0.83 [−0.83, +2.43] 0.375 69 / 58 no measurable difference not calibrated
Math GSM8K — lm-eval 0-shot CoT, flexible extract 1,319 78.77 79.45 +0.68 [−0.76, +2.12] 0.426 55 / 46 no measurable difference not calibrated
Math MATH-500 — 0-shot CoT, math-verify 500 74.8 73.6 −1.2 [−4.2, +1.8] 0.519 27 / 33 no measurable difference not calibrated
Math AIME 2024+2025 — avg@8, T 0.6 (exploratory) 60 9.58 8.13 −1.46 [−3.96, +1.04] — — no measurable difference not calibrated
Code HumanEval+ — EvalPlus, greedy pass@1 164 75.00 78.05 +3.05 [−1.83, +7.93] 0.332 11 / 6 no measurable difference not calibrated ²
Code MBPP+ — EvalPlus, greedy pass@1 378 68.78 70.63 +1.85 [−1.59, +5.29] 0.382 27 / 20 no measurable difference not calibrated ²
Instruction following IFEval — prompt-level strict 541 71.53 73.01 +1.48 [−1.85, +4.62] 0.422 42 / 34 no measurable difference ³ not calibrated
Reasoning ARC-Challenge — 25-shot, acc_norm 1,172 66.98 65.87 −1.11 [−2.56, +0.34] 0.177 33 / 46 no measurable difference not calibrated

Scoreline (pre-registered): 2 improvements / 2 regressions / 8 no measurable difference. Every benchmark on the V&V board is reported. Regressions are not netted against improvements.

¹ MMLU. The scores are the macro average over 57 subjects. That is the calibrated aggregate on HELM's scale, and its Δ is +0.47. The pre-registered paired test runs on the 14,042 items: 72.11 vs 71.54 item-level exact match, +0.57, CI [+0.11, +1.03], p = 0.016.

Part of the gain is answer format. The base model gave 44 answers that are not one of the option letters A–D; v2 gave 3. v2 answered 23 of those 44 items correctly, which accounts for 0.16 of the 0.57 points. Two further checks are not significant:

  • On the 13,998 items where both models gave a valid option letter, the difference is +0.41 (p = 0.084).
  • A post-hoc subject-stratified test of the macro average gives CI [−0.07, +1.02].

Treat the MMLU improvement as small and format-sensitive.

² Code. HumanEval+ and MBPP+ were scored with EvalPlus 0.3.1's own evaluator, the same tool as the EvalPlus leaderboard. That leaderboard has no Qwen2.5-7B-Instruct row, so the harness cannot be calibrated against it.

³ IFEval. lm-eval 0.4.13's IFEval scorer is not fully deterministic: one checker substitutes a random letter, and language detection is unseeded. Rescoring moves the prompt-strict difference within about +0.9 to +1.8 points. The reading is "no measurable difference" either way.

How robust the GSM8K-HELM gain is. Removing every item where either model hit the 400-token cap leaves 968 items; there the gain is +3.6 points (p = 0.001). Mean output length is essentially unchanged: 198.0 tokens for v2 vs 199.3 for the base model.

The gain is tied to HELM's protocol. Under lm-eval's 5-shot and 0-shot CoT GSM8K protocols, the flexible-extract primary metric shows no measurable difference. The 5-shot strict-match scorer shows +2.73, mostly from formatting (see below). Do not read the HELM result as a general "better at math" claim: MATH Level 5 regresses.

Every other banked metric (secondary; many tests, no multiplicity correction; no claim rests on these)
Benchmark — metric Qwen2.5-7B-Instruct v2 Δ 95% CI p
MMLU (HELM) — item-level exact match 71.54 72.11 +0.57 [+0.11, +1.03] 0.016
GSM8K 5-shot — strict match (#### N format) 83.32 86.05 +2.73 [+0.83, +4.62] 0.006
GSM8K 5-shot — numeric 89.54 88.78 −0.76 [−2.27, +0.68] 0.373
GSM8K 5-shot — math-verify 89.23 88.78 −0.45 [−1.97, +1.06] 0.627
GSM8K 0-shot CoT — numeric 82.41 82.87 +0.45 — —
GSM8K 0-shot CoT — math-verify 82.03 83.02 +0.99 [−0.53, +2.50] 0.237
HumanEval — base tests 81.71 84.76 +3.05 [−1.22, +7.93] 0.302
MBPP — base tests 79.10 82.80 +3.70 [+0.26, +7.14] 0.049
MATH Level 5 — boxed-answer match 50.68 45.32 −5.36 [−7.78, −2.95] 1.5×10⁻⁵
MATH-500 — boxed-answer match 69.0 67.2 −1.8 [−4.8, +1.2] 0.306
ARC-Challenge — acc (unnormalised) 64.33 62.54 −1.79 [−3.24, −0.34] 0.022
IFEval — prompt-level loose 73.57 75.79 +2.22 [−0.92, +5.36] 0.207
IFEval — instruction-level strict / loose 78.90 / 80.70 80.58 / 82.49 +1.68 / +1.80 — —
AIME 2024 / AIME 2025 — avg@8 (exploratory) 12.08 / 7.08 8.75 / 7.50 −3.33 / +0.42 — —
AIME 2024+2025 — pass@8 (exploratory) 28.33 25.00 −3.33 — 0.727

Two notes on these rows:

  • The GSM8K 5-shot strict-match gain is mostly formatting. The base model left 59 answers without a #### N line; v2 left 30.
  • The GSM8K 0-shot CoT strict-match scorer is excluded. Its answer pattern never matches either model's output, so it scores 0 for both.

Majority voting on GSM8K. 32 samples at T = 0.7 per problem, for both models. No paired test was run on these rows.

Protocol — scorer maj@8 seed maj@8 v2 Δ maj@32 seed maj@32 v2 Δ
5-shot — flexible 91.43 92.19 +0.76 92.27 92.65 +0.38
5-shot — strict 91.51 92.12 +0.61 92.12 92.57 +0.45
5-shot — numeric 92.87 92.27 −0.61 93.48 92.87 −0.61
5-shot — math-verify 92.72 92.19 −0.53 93.48 92.72 −0.76
0-shot CoT — flexible 81.27 80.67 −0.61 81.05 81.12 +0.08
0-shot CoT — numeric 84.99 84.15 −0.83 84.76 84.69 −0.08
0-shot CoT — math-verify 85.37 84.15 −1.21 84.99 84.61 −0.38

Voting lifts both models by similar amounts. On the numeric and math-verify scorers the base model's vote is higher.

7B-class comparison, by axis and tier

Model names, sizes and scores appear exactly as each source table publishes them. The tested claim on every axis is the paired difference between v2 and Qwen2.5-7B-Instruct. Other models' rows are reference points, not a ranking.

Tier What it means Axes
A: calibrated, protocol-matched Our harness replays the reference's exact prompts and reproduces its published score for Qwen2.5-7B-Instruct. The other rows ran under the identical run specification. Math (GSM8K-HELM), Knowledge (MMLU-HELM)
B: same tool, not calibrated We used the leaderboard's own evaluator, but the leaderboard has no Qwen2.5-7B-Instruct row to calibrate against. Code (HumanEval+, MBPP+)
C: no matched public reference We found no published table that runs other models under our protocol. The 7B comparison is Qwen2.5-7B-Instruct on identical items: the board above. Math (MATH-500, MATH L5, AIME, GSM8K via lm-eval), Knowledge (MMLU-Pro), Instruction following (IFEval), Reasoning (ARC-C)

Tier A — Math: GSM8K (HELM Lite v1.13.0; 1,000 problems, 5-shot, greedy, stop=none, final-number exact match)

Model Size GSM8K Source
Converge Collective v2 7B 87.0 our replay
Qwen2.5-7B-Instruct 7B 83.3 our replay, same harness as v2
Qwen2.5 Instruct Turbo (7B) 7B 83.0 HELM published row
Llama 3.1 Instruct Turbo (8B) 8B 79.8 HELM published row
Gemma 2 Instruct (9B) 9B 76.2 HELM published row
Mistral Instruct v0.3 (7B) 7B 53.8 HELM published row

These are all the table's 6–10B rows that share the stop=none run specification. Eight other 6–9B rows ran with HELM's default stop sequence, so they are not protocol-matched and are excluded: Qwen1.5 (7B), Gemma (7B), Llama 3 (8B), Mistral v0.1 (7B), Yi (6B), Llama 2 (7B), Falcon (7B) and OLMo (7B).

Tier A — Knowledge: MMLU (HELM MMLU v1.13.0; 14,042 questions, 57 subjects, 5-shot, single-letter answer, macro average)

Model Size MMLU Source
Phi-3 (7B) (HELM label; microsoft/phi-3-small-8k-instruct) 7B 75.68 HELM published row
Converge Collective v2 7B 73.40 our replay
Qwen2.5-7B-Instruct 7B 72.92 our replay, same harness as v2
Qwen2.5 Instruct Turbo (7B) 7B 72.87 HELM published row
Mistral Instruct v0.3 (7B) 7B 59.91 HELM published row
Llama 3.1 Instruct Turbo (8B) 8B 56.06 HELM published row

These are all the 6–10B rows whose 57 per-subject run specifications are identical to the Qwen2.5-7B-Instruct run we replay. Nine other 6–9B rows use different prompt instructions (0 of 57 subjects match), so they are excluded: Gemma 2 (9B), Llama 3 (8B), Gemma (7B), Yi (6B), Qwen1.5 (7B), Mistral v0.1 (7B), OLMo 1.7 (7B), Llama 2 (7B) and OLMo (7B).

About the HELM rows. They are frozen results published by the HELM project, obtained through Together AI endpoints. The exception is Phi-3 (7B), which ran on HELM's Hugging Face deployment. Our numbers come from a self-run replay of HELM's exact request prompts, using the published weights in bf16 on vLLM. The replay reproduces HELM's aggregate for Qwen2.5-7B-Instruct. Per item, it agrees with HELM's own record on 939 of 1,000 GSM8K items and 13,822 of 14,042 MMLU items, so serving differences exist. HELM has not evaluated, listed or endorsed this model.

Tier B — Code: HumanEval+ and MBPP+ (EvalPlus leaderboard snapshot, 2026-09-27; greedy pass@1)

Rows shown:

  • the five highest-scoring 6–9B rows on each benchmark, with ties included;
  • every general-purpose 7–8B chat model that EvalPlus evaluated with its instruction prompt (EvalPlus's prompted flag);
  • Qwen2.5-Coder-7B-Instruct, a v2 source model. Its figure is vendor-reported; it is not on the snapshot.

gemma-1.1-2b-it is excluded: the snapshot lists it at 7.0B, but it is a 2B model. The complete 6–9B list follows below.

Model Size HumanEval+ MBPP+ Source
Qwen2.5-Coder-7B-Instruct (code specialist) 7B 84.1 71.7 vendor report, arXiv:2409.12186 Table 16
CodeQwen1.5-7B-Chat 7B 78.7 69.0 EvalPlus leaderboard
Converge Collective v2 7B 78.05 70.63 our run, EvalPlus 0.3.1
OpenCoder-8B-Instruct 8B 77.4 71.4 EvalPlus leaderboard
Qwen2.5-7B-Instruct 7B 75.00 68.78 our run, same harness as v2
Artigenz-Coder-DS-6.7B 6.7B 72.6 69.6 EvalPlus leaderboard
OpenCodeInterpreter-DS-6.7B 6.7B 72.0 66.4 EvalPlus leaderboard
DeepSeek-Coder-6.7B-instruct 6.7B 71.3 65.6 EvalPlus leaderboard
DeepSeek-Coder-7B-instruct-v1.5 7B 71.3 62.2 EvalPlus leaderboard
Magicoder-S-DS-6.7B 6.7B 71.3 69.0 EvalPlus leaderboard
OpenChat-3.5-7B-0106 7B 67.7 54.5 EvalPlus leaderboard
Llama3.1-8B-instruct 8B 62.8 55.6 EvalPlus leaderboard
Llama3-8B-instruct 8B 56.7 54.8 EvalPlus leaderboard
Mistral-7B-Instruct-v0.2 7B 36.0 37.0 EvalPlus leaderboard
gemma-1.1-7b-it 7B 35.4 45.0 EvalPlus leaderboard
xDAN-L1-Chat-RL-v1-7B 7B 32.9 41.3 EvalPlus leaderboard
gemma-7b-it 7B 25.0 36.8 EvalPlus leaderboard

v2 does not differ measurably from its base model on either code benchmark. It is not a code specialist.

All 43 rows of the EvalPlus snapshot with size 6–9B
Model (EvalPlus label) Size (B) HumanEval+ MBPP+ EvalPlus prompted
CodeQwen1.5-7B-Chat 7 78.7 69.0 yes
OpenCoder-8B-Instruct 8 77.4 71.4 yes
Artigenz-Coder-DS-6.7B 6.7 72.6 69.6 yes
OpenCodeInterpreter-DS-6.7B 6.7 72.0 66.4 yes
DeepSeek-Coder-6.7B-instruct 6.7 71.3 65.6 yes
DeepSeek-Coder-7B-instruct-v1.5 7 71.3 62.2 yes
Magicoder-S-DS-6.7B 6.7 71.3 69.0 yes
WaveCoder-Ultra-6.7B 7 69.5 63.5 yes
Magicoder-S-CL-7B 7 67.7 60.1 yes
OpenChat-3.5-7B-0106 7 67.7 54.5 yes
speechless-coder-ds-6.7B 6.7 65.9 64.4 yes
Llama3.1-8B-instruct 8 62.8 55.6 yes
Code-290k-6.7B-Instruct 6.7 59.7 — yes
Llama3-8B-instruct 8 56.7 54.8 yes
codegemma-7b-it 7 51.8 56.9 yes
speechless-starcoder2-7b 7 51.8 56.3 yes
speechless-coding-7B-16k-tora 7 50.6 50.6 yes
CodeQwen1.5-7B 7 45.7 60.8 no
WizardCoder-Python-7B-V1.0 7 45.1 49.5 yes
Mistral-codealpaca-7B 7 42.1 — no
MistralHermes-CodePro-7B-v1 7 42.1 46.4 yes
codegemma-7b 7 41.5 52.4 no
speechless-code-mistral-7B-v1.0 7 41.5 48.7 yes
DeepSeek-Coder-6.7B-base 6.7 39.6 58.7 no
Mistral-7B-Instruct-v0.2 7 36.0 37.0 yes
CodeLlama-7B 7 35.4 46.8 no
gemma-1.1-7b-it 7 35.4 45.0 yes
xDAN-L1-Chat-RL-v1-7B 7 32.9 41.3 yes
StarCoder2-7B 7 29.9 — no
Llama3-8B-base 8 29.3 51.6 no
gemma-7b 7 28.7 43.4 no
CodeGen-6B 6 25.6 42.9 no
gemma-7b-it 7 25.0 36.8 yes
CodeT5+-6B 6 24.4 41.5 no
Mistral-7B 7 23.8 42.1 no
Zephyr β-7B 7 23.2 34.7 no
StarCoderBase-7B 7 21.3 — no
CodeGen2-7B 7 17.7 — no
gemma-1.1-2b-it 7 17.7 23.3 yes
InCoder-6.7B 6.7 12.2 — no
Vicuna-7B 7 11.6 — no
GPT-J-6B 6 11.0 — no
StableLM-7B 7 2.4 — no

Tier C — axes with no matched public reference

Axis Benchmarks on this board 7B comparison available
Math MATH-500, MATH Level 5, AIME 2024+2025, GSM8K (lm-eval 5-shot and 0-shot CoT) Qwen2.5-7B-Instruct on identical items (board above)
Knowledge MMLU-Pro Qwen2.5-7B-Instruct on identical items (board above)
Instruction following IFEval Qwen2.5-7B-Instruct on identical items (board above)
Reasoning ARC-Challenge Qwen2.5-7B-Instruct on identical items (board above)

We make no cross-model claim on these axes.

Honest framing

  • A static merge is a trade, not a lift. v2 gains on GSM8K under HELM's protocol and, slightly, on MMLU. It loses on MATH Level 5 and MMLU-Pro. Eight benches show no measurable difference.
  • What is tested. The tested claims are paired differences against Qwen2.5-7B-Instruct on identical items. Other models' HELM and EvalPlus rows are reference points, not a ranking.
  • Self-reported results. The evaluation record, 60 files of results, raw per-item outputs and exports, is committed by Merkle root 0220d670dec55f8494a322cf2e5b247e8b426c27bc710e4b4bf757d105f7efb9 (RFC 6962, SHA-256). The record is unsigned and not yet public.
  • How the paired statistics were produced. The board's pinned code computed the paired test for GSM8K-HELM, MATH-500, HumanEval+ and AIME. For the other rows, we computed the same test after the run from the banked per-item outputs, with the same method and seed. Exploratory runs outside the board are not reported here.
  • Contamination. v2 was not built by optimising on any of these test sets. The source models' training data is third-party, and we cannot certify it free of benchmark contamination. v2 contains weights from source models that Qwen2.5-7B-Instruct does not, so pairing does not cancel exposure specific to them.

Research output under the converge two-tier charter — the people's tier.

Evaluation setup
  • Models: we evaluated two models on identical items.

    • v2 at revision eb3683bfa089b173e59ef129e3a6670ec4ae2ac3.
    • Qwen/Qwen2.5-7B-Instruct at revision a09a35458c702b33eeacc393d103063234e8bc28, as the control.

    Before the run, the served weight bytes of both were checked against their sha256.

  • Engine: vLLM 0.30.0, dtype bfloat16, max_model_len 32,768, seed 1234, on one NVIDIA A100-SXM4-80GB. HumanEval+ and MBPP+ ran through EvalPlus's own vLLM provider with max_model_len 2,048. Software: torch 2.13.0, transformers 5.17.0, lm-eval 0.4.13, EvalPlus 0.3.1, math-verify 0.9.0. The build was checked against these pins before any run.

  • Prompt format:

    • GSM8K-HELM and MMLU-HELM: HELM's own prompts, sent as one user turn.
    • GSM8K 0-shot CoT, MATH, IFEval and AIME: the chat template.
    • GSM8K 5-shot, ARC-Challenge and MMLU-Pro: plain completion, chat template off.
    • HumanEval+ and MBPP+: EvalPlus's own prompting.

    Where the chat template was used, no system prompt was passed, so the template's default applied.

  • Decoding: greedy (temperature 0, top_p 1, repetition penalty 1.0), except:

    • AIME: 8 samples at T 0.6, top_p 0.95, up to 16,384 new tokens;
    • majority-vote rows: 32 samples at T 0.7.

    The repository's generation_config.json sampling defaults were disabled.

  • Generation budgets (max new tokens):

    • GSM8K-HELM: 400
    • GSM8K 5-shot and 0-shot CoT: 1,024
    • HumanEval+ and MBPP+: 768
    • MATH: 4,096
    • IFEval: 1,280
    • MMLU-Pro: 2,048
    • MMLU-HELM: 1
    • ARC-Challenge: loglikelihood scoring, no generation
  • Data:

    • GSM8K and MMLU replay HELM's banked request files.
    • Pinned Hugging Face dataset revisions: openai/gsm8k, HuggingFaceH4/MATH-500, EleutherAI/hendrycks_math, allenai/ai2_arc, google/IFEval, Maxwell-Jia/AIME_2024, math-ai/aime25 and TIGER-Lab/MMLU-Pro.
    • HumanEval+ and MBPP+: EvalPlus's releases.
  • Statistics:

    • Paired bootstrap, 10,000 resamples, 95% percentile CI.
    • Exact two-sided McNemar test on the same items.
    • The reading rule was fixed in the evaluation code before the run.
  • Checks:

    • Every suite completed with full item coverage, identical item sets for both models, a passing decoding audit and no voided blocks.
    • Every headline number was recomputed independently from the raw per-item outputs.
  • Evaluation date: 2026-09-28.

Behaviour differences worth knowing
  • Format compliance. Malformed or unparsed answers, base model → v2:

    • MMLU, answers that are not an option letter A–D: 44 → 3
    • GSM8K-HELM: 9 → 3
    • GSM8K 5-shot answers without a #### N line: 59 → 30
    • MMLU-Pro: 48 → 79 (the one suite where v2 does worse)
  • Less answer-repetition on MMLU-Pro. About 12% of the base model's responses repeat "the answer is (X)" five or more times after answering, often until the 2,048-token cap. For v2 it is under 2%. Total generated tokens: 5.23M for the base model vs 2.39M for v2.

  • More generation-cap hits on hard math:

    • MATH Level 5: 33 vs 22 at 4,096 tokens
    • MATH-500: 11 vs 4
    • AIME: 21 vs 4 of 480 samples at 16,384 tokens

    The MATH Level 5 regression persists on the items both models finished: −5.9 points on 1,274 items.

  • Item-level churn. On GSM8K-HELM, v2 fixed 80 items the base model got wrong, and got 43 items wrong that the base model got right. We have not tested whether that churn is stable across runs.

Quick start

v2 loads and prompts exactly like Qwen2.5-7B-Instruct. The board's greedy scores were produced in bfloat16.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "Optitransfer/Qwen2.5-7B-Instruct-converge-collective-v2"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(
    repo, dtype=torch.bfloat16, device_map="auto"  # use torch_dtype= on older transformers
)

messages = [{"role": "user", "content": "A train travels 180 km in 2.5 hours. What is its average speed in km/h?"}]
text = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tok(text, return_tensors="pt").to(model.device)

# Greedy, as evaluated. The repo's generation_config samples by default (T=0.7, top_p=0.8, top_k=20, repetition_penalty=1.05).
out = model.generate(**inputs, max_new_tokens=512, do_sample=False, repetition_penalty=1.0)
print(tok.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
from vllm import LLM, SamplingParams

llm = LLM(
    model="Optitransfer/Qwen2.5-7B-Instruct-converge-collective-v2",
    dtype="bfloat16",
    max_model_len=32768,
    generation_config="vllm",  # ignore the repo's sampling defaults, as in our evaluation
)
params = SamplingParams(temperature=0.0, max_tokens=512)
out = llm.chat([{"role": "user", "content": "Explain the difference between a list and a tuple in Python."}], params)
print(out[0].outputs[0].text)
Model details
Architecture Qwen2ForCausalLM (Qwen2.5-7B-Instruct's), 28 layers, hidden size 3,584
Parameters 7,615,616,512
Weights one model.safetensors, 15,231,271,784 bytes, stored as float16
Configured context 32,768 tokens (max_position_embeddings; RoPE θ = 1,000,000; no rope scaling)
Unchanged from Qwen2.5-7B-Instruct config.json, generation_config.json, tokenizer.json, tokenizer_config.json, all byte-identical
Training none: no fine-tuning, no gradient training, no training data
  • dtype note. config.json declares torch_dtype: bfloat16, unchanged from the base model, while the stored tensors are float16. All reported results loaded the weights in bfloat16. Float16 inference was not evaluated.
  • File hashes (sha256):
    • config.json: 7463bb0ea78315365e6c6b74de4e73bbcc8359dfb0c5a737584e077d42c0b03c
    • generation_config.json: 3a8f9087e486054c8a4a08dae2e5a3ba62e23da212b5b8c08bc42cb983c3459f
    • tokenizer.json: c0382117ea329cdf097041132f6d735924b697924d6f6fc3945713e96ce87539
    • tokenizer_config.json: 5b5d4f65d0acd3b2d56a35b56d374a36cbc1c8fa5cf3b3febbbfabf22f359583
  • If the revision you load differs from eb3683bf, check that model.safetensors still has the sha256 above. A card-only commit changes the revision hash, not the weights.

Limitations

  • Not a long-context model. It is configured for 32,768 tokens, and no long-context evaluation was run. tokenizer_config.json inherits model_max_length: 131072 from the base model; that value is not v2's configured context.
  • Weaker than the base model on MATH Level 5 and MMLU-Pro (see the board).
  • Not a code specialist. HumanEval+ and MBPP+ show no measurable difference from the base model.
  • Identifies as Qwen. When no system message is given, the inherited chat template inserts "You are Qwen, created by Alibaba Cloud. You are a helpful assistant." The model may therefore describe itself as Qwen.
  • Default decoding samples. The repository's generation settings sample (T 0.7, top_p 0.8, top_k 20, repetition penalty 1.05). The board's scores are greedy, except the AIME and majority-vote rows.
  • Not evaluated for safety, bias or toxicity. It was evaluated in English only. It may produce incorrect, biased or harmful content; do not use it for high-stakes decisions without independent evaluation.

License

This model descends, through v1, from facebook/opt-2.7b. That model is distributed under the OPT-175B License Agreement, which limits use to non-commercial research. This model is therefore made available for non-commercial research use only, subject to three sets of terms:

  • the OPT-175B License Agreement;
  • the Apache-2.0 licenses of the Qwen, Mistral, SmolLM2, Granite and Pythia sources;
  • the MIT licenses of microsoft/phi-2 and microsoft/Phi-3-mini-4k-instruct.

No rights are granted beyond those licenses. OPT-175B is licensed under the OPT-175B license, Copyright (c) Meta Platforms, Inc. All Rights Reserved.

model.safetensors is modified relative to every source model. config.json, generation_config.json, tokenizer.json and tokenizer_config.json are Qwen2.5-7B-Instruct's files, unmodified.

Not affiliated with or endorsed by Alibaba Cloud, the Qwen team, the HELM project or EvalPlus. "Qwen" is used only to identify the model this one is derived from.

Corrections

Earlier revisions of this card, up to and including eb3683bf, contained the following, now withdrawn; they should not be relied on:

  • an apache-2.0 license declaration;
  • build-process statements;
  • a lineage count;
  • an evaluation table that is not hash-linked to these weights.

The repository's receipt.json belongs to v1 and does not describe this model. The weights are unchanged from revision eb3683bf.

Downloads last month
1,181
Safetensors
Model size
8B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Optitransfer/Qwen2.5-7B-Instruct-converge-collective-v2

Papers for Optitransfer/Qwen2.5-7B-Instruct-converge-collective-v2