Qwen3.5-KETI-HAECHI-27B / reports /BENCHMARK_GUIDE.md
byunggill's picture
Upload folder using huggingface_hub
f15d5c0 verified
|
Raw History Blame Contribute Delete
4.59 kB
# Benchmark Guide
## Reading the numbers
All percentage deltas in the model report are **Tuned minus that model's own official Base**, in absolute percentage points. Qwen and HCX raw scores must not be mixed into a single architecture ranking because their official checkpoints, chat templates, vision policies and pretraining histories differ.
| Metric | Meaning | Direction | Important caveat |
|---|---|:---:|---|
| Exact / normalized exact | After evaluator normalization, prediction and canonical answer are identical | Higher | Correct content with extra prose can fail |
| Containment | Canonical target occurs inside normalized output | Higher | More permissive than exact; hallucinated extra claims may still pass |
| Format | Required answer wrapper/choice/short-answer form was parsed | Higher | Format success is not semantic correctness |
| Similarity | Evaluator's normalized string/semantic similarity in [0,1] | Higher | Near-name answers can score partially without being canonical |
| CER | Character edit distance divided by reference length | Lower | Can exceed 1 when insertion is heavy |
| Line recall | Fraction of reference OCR lines recovered | Higher | Does not penalize all extra text |
| Accuracy / aAcc | Correct rows divided by evaluated rows; Hallusion aAcc balances its paired structure | Higher | Compare only matched protocols |
| pass@1 | First generated program passes hidden tests | Higher | Sensitive to code extraction and runtime |
| ToolSandbox similarity | Milestone completion minus minefield penalties, aggregated by scenario | Higher | Continuous, not plain accuracy |
| Tau reward | End-to-end task reward from actions, state checks and communication | Higher | Domain/trial weighting must match |
## Benchmark families
### General vision-language
- **MMBench DEV EN v1.1**: English visual perception, relation, logic and knowledge multiple choice.
- **MMStar / MMStar-KO**: leakage-reduced multimodal core skills in English and Korean.
- **KRETA**: Korean-centric document, chart, scene and reasoning evaluation.
- **MMMU-Pro 10c**: college-level multidisciplinary visual problems with ten choices.
- **HallusionBench**: visual faithfulness and hallucination resistance; report `aAcc`.
- **MathVista MINI**: visual mathematical reasoning. It is available in the HCX comparison; the Qwen run retained predictions but not an authoritative scored row.
### Mammoth OCR and heritage
- **Heritage Multi** asks attributes, **Reverse** selects the matching image, and **Simple** freely names a heritage object.
- **OCR Font**, **Outdoor** and **Public** cover rendered fonts, signs/scenes and administrative documents.
- Exact, containment, similarity, CER and line recall should be read together. Exact alone confounds recognition and output policy.
### Cultural Focus and H400/HS100
- **Cultural Focus identity** is exact official-title recognition; **view caption** is strict caption/year matching. `overall` aggregates both and is therefore much harder than identity alone.
- **H400 Direct** is open-vocabulary held-out naming; **H400 Hard** is same-family multiple choice; **Knowledge Image/Text** test factual knowledge.
- **HS100 Train** measures memorization/fit, **Unseen** measures new-view transfer, and **Simple** is only the 102-row intersection with the original Simple style. These are not aliases for full Mammoth Heritage Simple (768 rows).
### OpenCompass language
- **Core**: IFEval, AIME 2024/2025, PRM800K math, BBH, GPQA, MMLU-Pro, HumanEval, LiveCodeBench and long-context suites.
- **Extra**: ARC, BoolQ, COPA, CommonsenseQA, HellaSwag, OpenBookQA, PIQA, SIQA, WinoGrande, MMLU, GSM8K, MATH, MBPP, LAMBADA, DROP and Natural Questions.
- **Korean**: KMMLU, CSATQA, HAE-RAE, K2-Eval, KoBEST, KOBALT and Korean clinical QA.
- The first row in each score table is the configured aggregate, not another dataset.
### Tool use and agents
- **BFCL V4** parses function calls into ASTs and tests simple, multiple, parallel and multi-turn calls, including missing-function/parameter and long-context cases.
- **ToolSandbox** runs phone-like stateful scenarios with distractor tools, milestone goals and minefields.
- **Tau2** evaluates airline, retail and telecom interactions. **Tau3** adds revised environments and banking knowledge with repeated trials.
- BFCL's run directory retained score CSVs and diagnostic logs but not the raw generation corpus; paired output examples are therefore provided for Tau2/Tau3, while BFCL is documented with its category scores and observed empty-response diagnostics in the main report.