# SwissLegalEvals — Results Report

**Generated:** 2026-07-28 · **Models completed:** 15 / 15 · **Profile:** `default` (full, uncapped — 55,861 samples/model)

Benchmarks: **SLDS** (landmark-decision summarization), **SwiLTra-Bench** (legal translation: `sdst`, `slt`, `sscprt`), **LEXam** (open questions `lexam_oq` + multiple-choice `lexam_mcq_*`).

---

## Scoring key

| Family | What it measures | Metric | Scale |
|--------|------------------|--------|-------|
| `slds` | Landmark decision summarization | LLM-judge (DeepSeek-V4-Pro) | 0–100 |
| `sdst` | Decision-summary translation (text level) | LLM-judge (gpt-4o-mini) | 0–100 |
| `slt` | Law translation (paragraph level) | LLM-judge (gpt-4o-mini) | 0–100 |
| `sscprt` | Supreme-court press-release translation | LLM-judge (gpt-4o-mini) | 0–100 |
| `lexam_oq` | Legal open questions | LLM-judge (DeepSeek-R1) | 0–100 |
| `lexam_mcq_{4,8,16}_idk` | Legal MCQ with "I don't know" option | `trad_score` (plain accuracy) | 0–1 (shown ×100) |

`trad_score` is plain accuracy (1 if correct, else 0; picking IDK scores 0 like any wrong answer). The separate `idk_score` is the calibration metric (+1 correct, 0 for IDK, −1 for a wrong/unparseable answer) — that is the one that rewards abstaining when uncertain and penalizes confident wrong answers. The LEXam MCQ table reports `idk_score (trad_score)`.

`translation_avg` = mean of `sdst`, `slt`, `sscprt`. `mcq_avg` = mean of the three MCQ `trad_score`s (×100). `overall` = mean of `slds`, `lexam_oq`, `translation_avg`, `mcq_avg`.

---

## Overall ranking

Composite over the four task groups. **Hy-MT2-30B ran translation only** (SwiLTra-Bench), so its overall is translation-only and not directly comparable.

| Rank | Model | Overall | Translation | SLDS | LEXam OQ | MCQ avg |
|------|-------|--------:|------------:|-----:|---------:|--------:|
| —  | Hy-MT2-30B-A3B *(translation only)* | (61.94) | 61.94 | – | – | – |
| 1  | NVIDIA-Nemotron-3-Ultra-550B | **59.03** | 63.42 | 55.13 | 66.45 | 51.14 |
| 2  | Kimi-K2.6 | **58.64** | 65.34 | 49.81 | 63.87 | 55.55 |
| 3  | DeepSeek-V4-Pro | **58.29** | 64.22 | 50.49 | 63.64 | 54.81 |
| 4  | DeepSeek-V4-Flash | **57.29** | 61.28 | 50.97 | 63.06 | 53.87 |
| 5  | MiniMax-M3 | **56.91** | 63.98 | 49.55 | 63.39 | 50.71 |
| 6  | GLM-5.2 | **55.46** | 66.24 | 52.27 | 58.77 | 44.55 |
| 7  | gemma-4-31B-it | **54.77** | 57.18 | 48.04 | 57.82 | 56.03 |
| 8  | gpt-oss-120b | **51.44** | 58.48 | 46.27 | 55.61 | 45.40 |
| 9  | Mistral-Medium-3.5-128B | **48.11** | 58.69 | 33.97 | 52.81 | 46.97 |
| 10 | Qwen3.5-35B-A3B | **42.77** | 42.53 | 46.73 | 43.77 | 38.06 |
| 11 | Apertus-v1.5-70B | **40.22** | 65.21 | 37.99 | 43.67 | 14.02 |
| 12 | Llama-3.3-70B-Instruct | **33.00** | 60.77 | 16.10 | 40.20 | 14.92 |
| 13 | Olmo-3.1-32B-Think | **26.04** | 20.79 | 18.39 | 38.52 | 26.45 |
| 14 | LFM2.5-8B-A1B | **24.75** | 37.75 | 11.08 | 24.51 | 25.66 |

---

## Translation (SwiLTra-Bench) — all 15 models

| Rank | Model | Translation avg | sdst | slt | sscprt |
|------|-------|----------------:|-----:|----:|-------:|
| 1  | GLM-5.2 | **66.24** | 71.01 | 67.25 | 60.47 |
| 2  | Kimi-K2.6 | **65.34** | 70.11 | 66.34 | 59.55 |
| 3  | Apertus-v1.5-70B | **65.21** | 71.68 | 67.06 | 56.88 |
| 4  | DeepSeek-V4-Pro | **64.22** | 69.22 | 62.23 | 61.22 |
| 5  | MiniMax-M3 | **63.98** | 66.99 | 65.83 | 59.11 |
| 6  | NVIDIA-Nemotron-3-Ultra-550B | **63.42** | 69.56 | 63.66 | 57.02 |
| 7  | Hy-MT2-30B-A3B | **61.94** | 66.60 | 63.45 | 55.76 |
| 8  | DeepSeek-V4-Flash | **61.28** | 65.70 | 63.92 | 54.23 |
| 9  | Llama-3.3-70B-Instruct | **60.77** | 68.12 | 62.94 | 51.26 |
| 10 | Mistral-Medium-3.5-128B | **58.69** | 65.48 | 58.47 | 52.12 |
| 11 | gpt-oss-120b | **58.48** | 63.80 | 59.47 | 52.17 |
| 12 | gemma-4-31B-it | **57.18** | 63.47 | 48.93 | 59.16 |
| 13 | Qwen3.5-35B-A3B | **42.53** | 45.50 | 30.61 | 51.48 |
| 14 | LFM2.5-8B-A1B | **37.75** | 41.97 | 40.98 | 30.30 |
| 15 | Olmo-3.1-32B-Think | **20.79** | 23.97 | 28.80 | 9.61 |

---

## LEXam multiple-choice (`idk_score (trad_score)` ×100)

| Model | mcq_4 | mcq_8 | mcq_16 | avg |
|-------|------:|------:|-------:|----:|
| DeepSeek-V4-Pro | 32.18 (62.42) | 12.72 (54.63) | -4.33 (47.37) | **13.52 (54.81)** |
| gemma-4-31B-it | 33.66 (66.40) | 11.25 (55.45) | -7.54 (46.23) | **12.45 (56.03)** |
| Kimi-K2.6 | 37.85 (67.85) | 10.80 (54.84) | -11.46 (43.97) | **12.40 (55.55)** |
| DeepSeek-V4-Flash | 30.49 (62.05) | 9.07 (52.95) | -6.00 (46.60) | **11.19 (53.87)** |
| NVIDIA-Nemotron-3-Ultra-550B | 33.59 (63.06) | 7.16 (48.67) | -15.29 (41.68) | **8.49 (51.14)** |
| MiniMax-M3 | 28.34 (62.42) | 7.70 (50.06) | -20.21 (39.66) | **5.28 (50.71)** |
| Mistral-Medium-3.5-128B | 19.16 (57.54) | -5.04 (46.02) | -24.81 (37.35) | **-3.56 (46.97)** |
| gpt-oss-120b | 18.71 (58.82) | -16.08 (41.58) | -28.18 (35.79) | **-8.51 (45.40)** |
| GLM-5.2 | 2.88 (50.76) | -4.90 (44.90) | -23.98 (38.01) | **-8.67 (44.55)** |
| Qwen3.5-35B-A3B | 8.49 (52.91) | -12.58 (32.00) | -41.37 (29.28) | **-15.15 (38.06)** |
| Olmo-3.1-32B-Think | 1.76 (32.32) | -27.87 (25.30) | -51.35 (21.75) | **-25.82 (26.45)** |
| LFM2.5-8B-A1B | -9.97 (32.89) | -35.61 (26.57) | -61.34 (17.52) | **-35.64 (25.66)** |
| Apertus-v1.5-70B | -48.21 (24.54) | -74.82 (11.62) | -88.17 (5.91) | **-70.40 (14.02)** |
| Llama-3.3-70B-Instruct | -48.22 (25.89) | -75.51 (12.22) | -86.71 (6.65) | **-70.15 (14.92)** |

(Hy-MT2-30B did not run LEXam.)

---

## Notes & caveats

- **Different scales:** judge families are 0–100; MCQ is a 0–1 `trad_score` (plain accuracy — IDK counts as wrong). The composite rescales MCQ to 0–100 — treat `overall` as an indicative aggregate, not a normalized benchmark score.
- **Judge self-preference:** SLDS is judged by DeepSeek-V4-Pro, so DeepSeek models' SLDS scores may carry mild self-preference bias.
- **MCQ difficulty scales with option count:** every model drops from mcq_4 → mcq_16, as expected.
- **Top tier is tight:** Nemotron, Kimi, both DeepSeek variants and MiniMax-M3 sit within ~2.5 points overall. Nemotron leads SLDS and LEXam OQ; GLM-5.2 leads translation.
- **GLM-5.2 is the best translator** (66.24 avg, top on `slt`) but Apertus posts the highest `sdst` (71.68). GLM's overall is dragged down by weak MCQ and the lowest LEXam-OQ of that tier — a strong-translation / weaker-exam profile.
- **Apertus-v1.5-70B** is a near-frontier translator (65.21, third overall; best `sdst`) whose MCQ collapses to 14.02, same failure mode as Llama-3.3-70B.
- **Mistral-Medium-3.5-128B** lands mid-pack at 48.11: solid translation (58.69) and MCQ (46.97) but weak SLDS (33.97).
- **MiniMax-M3** lands mid-frontier: strong translation (63.98, near the top) but middling MCQ.
- **Llama-3.3-70B-Instruct** is a translation specialist: it scores 60.77 on translation but only 16.10 on SLDS and 14.92 on MCQ.
- **Small open models lag hard** on summarization (`slds`) and press-release translation (`sscprt`) — LFM2.5-8B and Olmo are near-floor on `slds`/`sscprt`.

## Provenance

- Source: `results/results/**/results_*.json` (latest run per model).
- Regenerate: `./.venv/bin/python -m swiss_legal_evals.aggregate` then `... -m swiss_legal_evals.plot`.
- Per-family bar charts: `plots/*.html`. Tidy tables: `results/summary_long.csv`, `results/summary_family_mean.csv`, `results/summary.csv`.

### Not included
- **qwen3.5-397b** — paused (never run to completion).
