Qwen3.5-KETI-HAECHI-27B / reports /DETAILED_SCORE_REPORT.md
byunggill's picture
Upload folder using huggingface_hub
f15d5c0 verified
|
Raw
History Blame Contribute Delete
14.1 kB

Qwen3.5-27B Official Base vs v28cl + 9B Heritage Recipe

Comparison generated: 2026-09-02 (Asia/Seoul).

Contract

  • Official base: /data/workspace/post-training/outputs/qwen35-27b-official (untouched official Qwen/Qwen3.5-27B weights).
  • Tuned: /data/workspace/post-training/outputs/qwen35-27b-v28cl-h400-hs100-scale2500-step2500-hf.
  • All deltas are tuned minus official base. Percentage metrics use absolute percentage points.
  • OpenCompass uses the same API, thinking, stop-token, parser, and dataset configuration for both models.
  • Tau3 uses four trials per task and identical BM25/tool simulation settings.
  • The tuned lineage follows the successful 9B heritage recipe: v28cl parent, specialist/task-vector transfer, H400/H300 acquisition, then H400/HS100 scale-up.

Executive summary

Benchmark Official Base Tuned Delta Direction
General-VL macro (6) 45.44% 77.60% +32.16 pp higher
Heritage macro (3) 38.06% 40.21% +2.15 pp higher
OCR macro (3) 19.99% 22.09% +2.11 pp higher
heritage_simple 0.00% 1.95% +1.95 pp higher
H400 direct exact 0.92% 38.07% +37.16 pp higher
H400 hard choice 26.45% 80.43% +53.98 pp higher
core_average 63.49% 64.80% +1.31 pp higher
extra_registered_average 57.26% 57.61% +0.35 pp higher
korean_hf_average 17.18% 16.77% -0.41 pp higher
Overall Acc 32.77% 32.11% -0.66 pp higher
Weighted overall (n=96) 69.79% 71.88% +2.08 pp higher
Weighted overall (n=1500) 41.67% 40.00% -1.67 pp higher

General multimodal

Benchmark Official Base Tuned Delta Direction
General-VL macro (6) 45.44% 77.60% +32.16 pp higher
MMBench_DEV_EN_V11 30.42% 90.63% +60.22 pp higher
MMStar 39.40% 77.33% +37.93 pp higher
MMStar_KO 45.20% 72.33% +27.13 pp higher
KRETA 48.54% 86.15% +37.60 pp higher
MMMU_Pro_10c 44.28% 61.85% +17.57 pp higher
HallusionBench aAcc 64.77% 77.29% +12.51 pp higher

Mammoth exact accuracy

Benchmark Official Base Tuned Delta Direction
Heritage macro (3) 38.06% 40.21% +2.15 pp higher
OCR macro (3) 19.99% 22.09% +2.11 pp higher
heritage_multi_attr 22.01% 28.26% +6.25 pp higher
heritage_reverse_mt 92.19% 90.43% -1.76 pp higher
heritage_simple 0.00% 1.95% +1.95 pp higher
ocr_font 19.66% 19.66% +0.00 pp higher
ocr_outdoor 36.13% 41.41% +5.27 pp higher
ocr_public_exec 4.17% 5.21% +1.04 pp higher

Mammoth target containment

Benchmark Official Base Tuned Delta Direction
Heritage macro (3) 44.51% 43.53% -0.98 pp higher
OCR macro (3) 32.47% 32.68% +0.22 pp higher
heritage_multi_attr 35.42% 35.55% +0.13 pp higher
heritage_reverse_mt 97.07% 92.58% -4.49 pp higher
heritage_simple 1.04% 2.47% +1.43 pp higher
ocr_font 20.96% 20.70% -0.26 pp higher
ocr_outdoor 70.70% 70.70% +0.00 pp higher
ocr_public_exec 5.73% 6.64% +0.91 pp higher

Mammoth normalized similarity

Benchmark Official Base Tuned Delta Direction
Heritage macro (3) 45.68% 50.25% +4.57 pp higher
OCR macro (3) 35.00% 38.59% +3.59 pp higher
heritage_multi_attr 39.14% 44.36% +5.22 pp higher
heritage_reverse_mt 92.20% 90.43% -1.77 pp higher
heritage_simple 5.70% 15.95% +10.25 pp higher
ocr_font 23.66% 24.12% +0.46 pp higher
ocr_outdoor 59.73% 64.96% +5.22 pp higher
ocr_public_exec 21.61% 26.71% +5.09 pp higher

Mammoth OCR CER

Benchmark Official Base Tuned Delta Direction
Heritage macro (3) 2.4569 0.6781 -1.7788 lower
OCR macro (3) 1.5337 1.1192 -0.4146 lower
heritage_multi_attr 0.7850 0.6776 -0.1074 lower
heritage_reverse_mt 0.1016 0.0957 -0.0059 lower
heritage_simple 6.4840 1.2608 -5.2231 lower
ocr_font 1.4261 1.0366 -0.3895 lower
ocr_outdoor 1.4172 0.9583 -0.4589 lower
ocr_public_exec 1.7579 1.3625 -0.3954 lower

Mammoth OCR line recall

Benchmark Official Base Tuned Delta Direction
Heritage macro (3) 44.51% 43.53% -0.98 pp higher
OCR macro (3) 32.47% 32.68% +0.22 pp higher
heritage_multi_attr 35.42% 35.55% +0.13 pp higher
heritage_reverse_mt 97.07% 92.58% -4.49 pp higher
heritage_simple 1.04% 2.47% +1.43 pp higher
ocr_font 20.96% 20.70% -0.26 pp higher
ocr_outdoor 70.70% 70.70% +0.00 pp higher
ocr_public_exec 5.73% 6.64% +0.91 pp higher

Cultural heritage title-clean

Benchmark Official Base Tuned Delta Direction
overall / exact 0.17% 2.64% +2.47 pp higher
overall / containment 1.11% 3.05% +1.94 pp higher
overall / similarity 7.25% 11.11% +3.86 pp higher
overall / identity-any exact 0.31% 4.19% +3.89 pp higher
identity / exact 0.46% 6.89% +6.43 pp higher
identity / containment 0.70% 7.28% +6.58 pp higher
identity / similarity 13.16% 19.93% +6.77 pp higher
identity / identity-any exact 0.62% 7.81% +7.19 pp higher
view_caption / exact 0.00% 0.26% +0.26 pp higher
view_caption / containment 1.34% 0.69% -0.65 pp higher
view_caption / similarity 3.96% 6.19% +2.24 pp higher
view_caption / identity-any exact 0.00% 0.40% +0.40 pp higher
national_treasure / exact 0.00% 9.92% +9.92 pp higher
national_treasure / containment 0.62% 10.12% +9.50 pp higher
national_treasure / similarity 11.47% 22.91% +11.44 pp higher
national_treasure / identity-any exact 0.00% 12.10% +12.10 pp higher
other / exact 0.19% 1.51% +1.31 pp higher
other / containment 1.19% 1.96% +0.77 pp higher
other / similarity 6.60% 9.28% +2.68 pp higher
other / identity-any exact 0.37% 2.49% +2.11 pp higher

H400/HS100 target gates

Benchmark Official Base Tuned Delta Direction
H400 direct exact 0.92% 38.07% +37.16 pp higher
H400 hard choice 26.45% 80.43% +53.98 pp higher
H400 knowledge image 11.45% 25.30% +13.85 pp higher
H400 knowledge text 5.97% 6.28% +0.31 pp higher
HS100 train exact 0.00% 12.00% +12.00 pp higher
HS100 unseen exact 0.50% 7.43% +6.93 pp higher
HS100 Simple exact 0.00% 0.98% +0.98 pp higher

OpenCompass Core

Benchmark Official Base Tuned Delta Direction
core_average 63.49% 64.80% +1.31 pp higher
IFEval 85.58% 85.21% -0.37 pp higher
aime2024 30.00% 26.67% -3.33 pp higher
aime2025 60.00% 73.33% +13.33 pp higher
math_prm800k_500 86.80% 89.00% +2.20 pp higher
bbh 47.56% 46.84% -0.72 pp higher
GPQA_diamond 82.83% 84.85% +2.02 pp higher
mmlu_pro 85.55% 85.57% +0.02 pp higher
openai_humaneval 96.95% 96.34% -0.61 pp higher
lcb_code_generation 74.50% 76.00% +1.50 pp higher
leval 14.88% 15.00% +0.12 pp higher
longbench 6.81% 6.89% +0.08 pp higher
LongBenchv2 54.47% 57.46% +2.99 pp higher
keti_long_ctx_gutenberg 99.38% 99.25% -0.13 pp higher

OpenCompass Extra

Benchmark Official Base Tuned Delta Direction
extra_registered_average 57.26% 57.61% +0.35 pp higher
extra_registered_mc_average 88.46% 89.37% +0.91 pp higher
extra_registered_math_average 53.57% 54.12% +0.55 pp higher
ARC-c 73.56% 80.34% +6.78 pp higher
ARC-e 83.07% 85.01% +1.94 pp higher
BoolQ 90.55% 90.24% -0.31 pp higher
COPA 100.00% 100.00% +0.00 pp higher
commonsense_qa 88.45% 87.96% -0.49 pp higher
hellaswag 94.28% 94.12% -0.16 pp higher
openbookqa 89.80% 90.00% +0.20 pp higher
piqa 95.43% 95.81% +0.38 pp higher
siqa 76.51% 76.77% +0.26 pp higher
winogrande 92.90% 93.45% +0.55 pp higher
mmlu 92.00% 92.13% +0.13 pp higher
mmlu-stem 94.90% 94.71% -0.19 pp higher
mmlu-humanities 91.56% 91.64% +0.08 pp higher
mmlu-social-science 91.96% 92.64% +0.68 pp higher
mmlu-other 88.27% 88.41% +0.14 pp higher
gsm8k 9.86% 10.69% +0.83 pp higher
math 97.28% 97.56% +0.28 pp higher
mbpp 42.40% 43.00% +0.60 pp higher
lambada 0.00% 0.00% +0.00 pp higher
drop 91.89% 92.00% +0.11 pp higher
nq 32.52% 32.66% +0.14 pp higher

OpenCompass Korean

Benchmark Official Base Tuned Delta Direction
korean_hf_average 17.18% 16.77% -0.41 pp higher
kmmlu 1.97% 1.33% -0.64 pp higher
csatqa 19.23% 18.70% -0.53 pp higher
haerae 19.21% 19.21% +0.00 pp higher
k2_eval 17.36% 16.67% -0.69 pp higher
kobest 46.49% 46.61% +0.12 pp higher
kobest_boolq 53.43% 53.43% +0.00 pp higher
kobest_copa 49.60% 50.00% +0.40 pp higher
kobest_hellaswag 22.20% 22.40% +0.20 pp higher
kobest_sentineg 50.00% 50.00% +0.00 pp higher
kobest_wic 57.21% 57.21% +0.00 pp higher
kobalt 9.71% 9.71% +0.00 pp higher
kr_clinical_qa 6.31% 5.18% -1.13 pp higher

BFCL V4

Benchmark Official Base Tuned Delta Direction
Overall Acc 32.77% 32.11% -0.66 pp higher
Non-Live AST Acc 89.60% 89.44% -0.16 pp higher
Non-Live Simple AST 79.42% 79.25% -0.17 pp higher
Non-Live Multiple AST 95.50% 95.50% +0.00 pp higher
Non-Live Parallel AST 91.50% 91.00% -0.50 pp higher
Non-Live Parallel Multiple AST 92.00% 92.00% +0.00 pp higher
Multi Turn Acc 66.25% 64.50% -1.75 pp higher
Multi Turn Base 76.50% 76.00% -0.50 pp higher
Multi Turn Miss Func 66.00% 64.00% -2.00 pp higher
Multi Turn Miss Param 53.00% 52.00% -1.00 pp higher
Multi Turn Long Context 69.50% 66.00% -3.50 pp higher

ToolSandbox

Benchmark Official Base Tuned Delta Direction
ToolSandbox similarity (n=32) 0.6077 0.6043 -0.0034 higher

Tau2

Benchmark Official Base Tuned Delta Direction
Tau2 airline (n=32) 65.62% 71.88% +6.25 pp higher
Tau2 retail (n=32) 56.25% 65.62% +9.38 pp higher
Tau2 telecom (n=32) 87.50% 78.12% -9.38 pp higher
Weighted overall (n=96) 69.79% 71.88% +2.08 pp higher
Domain macro 69.79% 71.88% +2.08 pp higher

Tau3

Benchmark Official Base Tuned Delta Direction
Tau3 airline (n=200) 67.00% 64.50% -2.50 pp higher
Tau3 retail (n=456) 50.66% 44.08% -6.58 pp higher
Tau3 telecom (n=456) 50.66% 51.54% +0.88 pp higher
Tau3 banking knowledge (n=388) 7.47% 9.02% +1.55 pp higher
Weighted overall (n=1500) 41.67% 40.00% -1.67 pp higher
Domain macro 43.95% 42.28% -1.66 pp higher

Largest improvements

  • General multimodal / MMBench_DEV_EN_V11: +60.22 pp
  • H400/HS100 target gates / H400 hard choice: +53.98 pp
  • General multimodal / MMStar: +37.93 pp
  • General multimodal / KRETA: +37.60 pp
  • H400/HS100 target gates / H400 direct exact: +37.16 pp
  • General multimodal / General-VL macro (6): +32.16 pp
  • General multimodal / MMStar_KO: +27.13 pp
  • General multimodal / MMMU_Pro_10c: +17.57 pp
  • H400/HS100 target gates / H400 knowledge image: +13.85 pp
  • OpenCompass Core / aime2025: +13.33 pp

Largest regressions

  • Tau2 / Tau2 telecom (n=32): -9.38 pp
  • Tau3 / Tau3 retail (n=456): -6.58 pp
  • Mammoth OCR line recall / heritage_reverse_mt: -4.49 pp
  • Mammoth target containment / heritage_reverse_mt: -4.49 pp
  • BFCL V4 / Multi Turn Long Context: -3.50 pp
  • OpenCompass Core / aime2024: -3.33 pp
  • Tau3 / Tau3 airline (n=200): -2.50 pp
  • BFCL V4 / Multi Turn Miss Func: -2.00 pp
  • Mammoth normalized similarity / heritage_reverse_mt: -1.77 pp
  • Mammoth exact accuracy / heritage_reverse_mt: -1.76 pp

Reading the result

  • H400 direct exact measures open-ended canonical-name retrieval; H400 hard measures closed-set choice. A large gap between them indicates recognition/discrimination is stronger than exact name generation.
  • HS100 train, unseen, and Simple distinguish memorization, held-out view generalization, and transfer to a separate simple-identification population. They are not interchangeable with Mammoth Heritage Simple.
  • Mammoth exact is strict. Containment and normalized similarity show whether an answer is usable but differs in annotation, spacing, aliases, or added explanation.
  • OCR CER is the only lower-is-better metric in this report; all other deltas are better when positive.
  • OpenCompass, General-VL, BFCL, ToolSandbox, Tau2, and Tau3 are retention gates. Improvements in heritage should be interpreted together with these broad-capability deltas.

Source locations

  • Official OpenCompass: /data/workspace/eval_models/test/outputs/*q35-27b-official-20260901r2full
  • Tuned OpenCompass: /data/workspace/eval_models/test/outputs/*q35-27b-v28cl-hs25-20260902r1
  • Official Tau3: /data/workspace/eval_models/tau3_axes/q35-27b-official-20260901r2full
  • Tuned Tau3: /data/workspace/eval_models/tau3_axes/q35-27b-v28cl-hs25-20260902r1-full