Buckets:
SwissLegalEvals — Results Report
Generated: 2026-07-28 · Models completed: 15 / 15 · Profile: default (full, uncapped — 55,861 samples/model)
Benchmarks: SLDS (landmark-decision summarization), SwiLTra-Bench (legal translation: sdst, slt, sscprt), LEXam (open questions lexam_oq + multiple-choice lexam_mcq_*).
Scoring key
| Family | What it measures | Metric | Scale |
|---|---|---|---|
slds |
Landmark decision summarization | LLM-judge (DeepSeek-V4-Pro) | 0–100 |
sdst |
Decision-summary translation (text level) | LLM-judge (gpt-4o-mini) | 0–100 |
slt |
Law translation (paragraph level) | LLM-judge (gpt-4o-mini) | 0–100 |
sscprt |
Supreme-court press-release translation | LLM-judge (gpt-4o-mini) | 0–100 |
lexam_oq |
Legal open questions | LLM-judge (DeepSeek-R1) | 0–100 |
lexam_mcq_{4,8,16}_idk |
Legal MCQ with "I don't know" option | trad_score (plain accuracy) |
0–1 (shown ×100) |
trad_score is plain accuracy (1 if correct, else 0; picking IDK scores 0 like any wrong answer). The separate idk_score is the calibration metric (+1 correct, 0 for IDK, −1 for a wrong/unparseable answer) — that is the one that rewards abstaining when uncertain and penalizes confident wrong answers. The LEXam MCQ table reports idk_score (trad_score).
translation_avg = mean of sdst, slt, sscprt. mcq_avg = mean of the three MCQ trad_scores (×100). overall = mean of slds, lexam_oq, translation_avg, mcq_avg.
Overall ranking
Composite over the four task groups. Hy-MT2-30B ran translation only (SwiLTra-Bench), so its overall is translation-only and not directly comparable.
| Rank | Model | Overall | Translation | SLDS | LEXam OQ | MCQ avg |
|---|---|---|---|---|---|---|
| — | Hy-MT2-30B-A3B (translation only) | (61.94) | 61.94 | – | – | – |
| 1 | NVIDIA-Nemotron-3-Ultra-550B | 59.03 | 63.42 | 55.13 | 66.45 | 51.14 |
| 2 | Kimi-K2.6 | 58.64 | 65.34 | 49.81 | 63.87 | 55.55 |
| 3 | DeepSeek-V4-Pro | 58.29 | 64.22 | 50.49 | 63.64 | 54.81 |
| 4 | DeepSeek-V4-Flash | 57.29 | 61.28 | 50.97 | 63.06 | 53.87 |
| 5 | MiniMax-M3 | 56.91 | 63.98 | 49.55 | 63.39 | 50.71 |
| 6 | GLM-5.2 | 55.46 | 66.24 | 52.27 | 58.77 | 44.55 |
| 7 | gemma-4-31B-it | 54.77 | 57.18 | 48.04 | 57.82 | 56.03 |
| 8 | gpt-oss-120b | 51.44 | 58.48 | 46.27 | 55.61 | 45.40 |
| 9 | Mistral-Medium-3.5-128B | 48.11 | 58.69 | 33.97 | 52.81 | 46.97 |
| 10 | Qwen3.5-35B-A3B | 42.77 | 42.53 | 46.73 | 43.77 | 38.06 |
| 11 | Apertus-v1.5-70B | 40.22 | 65.21 | 37.99 | 43.67 | 14.02 |
| 12 | Llama-3.3-70B-Instruct | 33.00 | 60.77 | 16.10 | 40.20 | 14.92 |
| 13 | Olmo-3.1-32B-Think | 26.04 | 20.79 | 18.39 | 38.52 | 26.45 |
| 14 | LFM2.5-8B-A1B | 24.75 | 37.75 | 11.08 | 24.51 | 25.66 |
Translation (SwiLTra-Bench) — all 15 models
| Rank | Model | Translation avg | sdst | slt | sscprt |
|---|---|---|---|---|---|
| 1 | GLM-5.2 | 66.24 | 71.01 | 67.25 | 60.47 |
| 2 | Kimi-K2.6 | 65.34 | 70.11 | 66.34 | 59.55 |
| 3 | Apertus-v1.5-70B | 65.21 | 71.68 | 67.06 | 56.88 |
| 4 | DeepSeek-V4-Pro | 64.22 | 69.22 | 62.23 | 61.22 |
| 5 | MiniMax-M3 | 63.98 | 66.99 | 65.83 | 59.11 |
| 6 | NVIDIA-Nemotron-3-Ultra-550B | 63.42 | 69.56 | 63.66 | 57.02 |
| 7 | Hy-MT2-30B-A3B | 61.94 | 66.60 | 63.45 | 55.76 |
| 8 | DeepSeek-V4-Flash | 61.28 | 65.70 | 63.92 | 54.23 |
| 9 | Llama-3.3-70B-Instruct | 60.77 | 68.12 | 62.94 | 51.26 |
| 10 | Mistral-Medium-3.5-128B | 58.69 | 65.48 | 58.47 | 52.12 |
| 11 | gpt-oss-120b | 58.48 | 63.80 | 59.47 | 52.17 |
| 12 | gemma-4-31B-it | 57.18 | 63.47 | 48.93 | 59.16 |
| 13 | Qwen3.5-35B-A3B | 42.53 | 45.50 | 30.61 | 51.48 |
| 14 | LFM2.5-8B-A1B | 37.75 | 41.97 | 40.98 | 30.30 |
| 15 | Olmo-3.1-32B-Think | 20.79 | 23.97 | 28.80 | 9.61 |
LEXam multiple-choice (idk_score (trad_score) ×100)
| Model | mcq_4 | mcq_8 | mcq_16 | avg |
|---|---|---|---|---|
| DeepSeek-V4-Pro | 32.18 (62.42) | 12.72 (54.63) | -4.33 (47.37) | 13.52 (54.81) |
| gemma-4-31B-it | 33.66 (66.40) | 11.25 (55.45) | -7.54 (46.23) | 12.45 (56.03) |
| Kimi-K2.6 | 37.85 (67.85) | 10.80 (54.84) | -11.46 (43.97) | 12.40 (55.55) |
| DeepSeek-V4-Flash | 30.49 (62.05) | 9.07 (52.95) | -6.00 (46.60) | 11.19 (53.87) |
| NVIDIA-Nemotron-3-Ultra-550B | 33.59 (63.06) | 7.16 (48.67) | -15.29 (41.68) | 8.49 (51.14) |
| MiniMax-M3 | 28.34 (62.42) | 7.70 (50.06) | -20.21 (39.66) | 5.28 (50.71) |
| Mistral-Medium-3.5-128B | 19.16 (57.54) | -5.04 (46.02) | -24.81 (37.35) | -3.56 (46.97) |
| gpt-oss-120b | 18.71 (58.82) | -16.08 (41.58) | -28.18 (35.79) | -8.51 (45.40) |
| GLM-5.2 | 2.88 (50.76) | -4.90 (44.90) | -23.98 (38.01) | -8.67 (44.55) |
| Qwen3.5-35B-A3B | 8.49 (52.91) | -12.58 (32.00) | -41.37 (29.28) | -15.15 (38.06) |
| Olmo-3.1-32B-Think | 1.76 (32.32) | -27.87 (25.30) | -51.35 (21.75) | -25.82 (26.45) |
| LFM2.5-8B-A1B | -9.97 (32.89) | -35.61 (26.57) | -61.34 (17.52) | -35.64 (25.66) |
| Apertus-v1.5-70B | -48.21 (24.54) | -74.82 (11.62) | -88.17 (5.91) | -70.40 (14.02) |
| Llama-3.3-70B-Instruct | -48.22 (25.89) | -75.51 (12.22) | -86.71 (6.65) | -70.15 (14.92) |
(Hy-MT2-30B did not run LEXam.)
Notes & caveats
- Different scales: judge families are 0–100; MCQ is a 0–1
trad_score(plain accuracy — IDK counts as wrong). The composite rescales MCQ to 0–100 — treatoverallas an indicative aggregate, not a normalized benchmark score. - Judge self-preference: SLDS is judged by DeepSeek-V4-Pro, so DeepSeek models' SLDS scores may carry mild self-preference bias.
- MCQ difficulty scales with option count: every model drops from mcq_4 → mcq_16, as expected.
- Top tier is tight: Nemotron, Kimi, both DeepSeek variants and MiniMax-M3 sit within ~2.5 points overall. Nemotron leads SLDS and LEXam OQ; GLM-5.2 leads translation.
- GLM-5.2 is the best translator (66.24 avg, top on
slt) but Apertus posts the highestsdst(71.68). GLM's overall is dragged down by weak MCQ and the lowest LEXam-OQ of that tier — a strong-translation / weaker-exam profile. - Apertus-v1.5-70B is a near-frontier translator (65.21, third overall; best
sdst) whose MCQ collapses to 14.02, same failure mode as Llama-3.3-70B. - Mistral-Medium-3.5-128B lands mid-pack at 48.11: solid translation (58.69) and MCQ (46.97) but weak SLDS (33.97).
- MiniMax-M3 lands mid-frontier: strong translation (63.98, near the top) but middling MCQ.
- Llama-3.3-70B-Instruct is a translation specialist: it scores 60.77 on translation but only 16.10 on SLDS and 14.92 on MCQ.
- Small open models lag hard on summarization (
slds) and press-release translation (sscprt) — LFM2.5-8B and Olmo are near-floor onslds/sscprt.
Provenance
- Source:
results/results/**/results_*.json(latest run per model). - Regenerate:
./.venv/bin/python -m swiss_legal_evals.aggregatethen... -m swiss_legal_evals.plot. - Per-family bar charts:
plots/*.html. Tidy tables:results/summary_long.csv,results/summary_family_mean.csv,results/summary.csv.
Not included
- qwen3.5-397b — paused (never run to completion).
Xet Storage Details
- Size:
- 7.44 kB
- Xet hash:
- 8df9bc212f39a79fe387feccff9fc099b44e16cef4639def41ea11e4d96c6413
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.