joelniklaus's picture
|
download
raw
7.44 kB

SwissLegalEvals — Results Report

Generated: 2026-07-28 · Models completed: 15 / 15 · Profile: default (full, uncapped — 55,861 samples/model)

Benchmarks: SLDS (landmark-decision summarization), SwiLTra-Bench (legal translation: sdst, slt, sscprt), LEXam (open questions lexam_oq + multiple-choice lexam_mcq_*).


Scoring key

Family What it measures Metric Scale
slds Landmark decision summarization LLM-judge (DeepSeek-V4-Pro) 0–100
sdst Decision-summary translation (text level) LLM-judge (gpt-4o-mini) 0–100
slt Law translation (paragraph level) LLM-judge (gpt-4o-mini) 0–100
sscprt Supreme-court press-release translation LLM-judge (gpt-4o-mini) 0–100
lexam_oq Legal open questions LLM-judge (DeepSeek-R1) 0–100
lexam_mcq_{4,8,16}_idk Legal MCQ with "I don't know" option trad_score (plain accuracy) 0–1 (shown ×100)

trad_score is plain accuracy (1 if correct, else 0; picking IDK scores 0 like any wrong answer). The separate idk_score is the calibration metric (+1 correct, 0 for IDK, −1 for a wrong/unparseable answer) — that is the one that rewards abstaining when uncertain and penalizes confident wrong answers. The LEXam MCQ table reports idk_score (trad_score).

translation_avg = mean of sdst, slt, sscprt. mcq_avg = mean of the three MCQ trad_scores (×100). overall = mean of slds, lexam_oq, translation_avg, mcq_avg.


Overall ranking

Composite over the four task groups. Hy-MT2-30B ran translation only (SwiLTra-Bench), so its overall is translation-only and not directly comparable.

Rank Model Overall Translation SLDS LEXam OQ MCQ avg
Hy-MT2-30B-A3B (translation only) (61.94) 61.94
1 NVIDIA-Nemotron-3-Ultra-550B 59.03 63.42 55.13 66.45 51.14
2 Kimi-K2.6 58.64 65.34 49.81 63.87 55.55
3 DeepSeek-V4-Pro 58.29 64.22 50.49 63.64 54.81
4 DeepSeek-V4-Flash 57.29 61.28 50.97 63.06 53.87
5 MiniMax-M3 56.91 63.98 49.55 63.39 50.71
6 GLM-5.2 55.46 66.24 52.27 58.77 44.55
7 gemma-4-31B-it 54.77 57.18 48.04 57.82 56.03
8 gpt-oss-120b 51.44 58.48 46.27 55.61 45.40
9 Mistral-Medium-3.5-128B 48.11 58.69 33.97 52.81 46.97
10 Qwen3.5-35B-A3B 42.77 42.53 46.73 43.77 38.06
11 Apertus-v1.5-70B 40.22 65.21 37.99 43.67 14.02
12 Llama-3.3-70B-Instruct 33.00 60.77 16.10 40.20 14.92
13 Olmo-3.1-32B-Think 26.04 20.79 18.39 38.52 26.45
14 LFM2.5-8B-A1B 24.75 37.75 11.08 24.51 25.66

Translation (SwiLTra-Bench) — all 15 models

Rank Model Translation avg sdst slt sscprt
1 GLM-5.2 66.24 71.01 67.25 60.47
2 Kimi-K2.6 65.34 70.11 66.34 59.55
3 Apertus-v1.5-70B 65.21 71.68 67.06 56.88
4 DeepSeek-V4-Pro 64.22 69.22 62.23 61.22
5 MiniMax-M3 63.98 66.99 65.83 59.11
6 NVIDIA-Nemotron-3-Ultra-550B 63.42 69.56 63.66 57.02
7 Hy-MT2-30B-A3B 61.94 66.60 63.45 55.76
8 DeepSeek-V4-Flash 61.28 65.70 63.92 54.23
9 Llama-3.3-70B-Instruct 60.77 68.12 62.94 51.26
10 Mistral-Medium-3.5-128B 58.69 65.48 58.47 52.12
11 gpt-oss-120b 58.48 63.80 59.47 52.17
12 gemma-4-31B-it 57.18 63.47 48.93 59.16
13 Qwen3.5-35B-A3B 42.53 45.50 30.61 51.48
14 LFM2.5-8B-A1B 37.75 41.97 40.98 30.30
15 Olmo-3.1-32B-Think 20.79 23.97 28.80 9.61

LEXam multiple-choice (idk_score (trad_score) ×100)

Model mcq_4 mcq_8 mcq_16 avg
DeepSeek-V4-Pro 32.18 (62.42) 12.72 (54.63) -4.33 (47.37) 13.52 (54.81)
gemma-4-31B-it 33.66 (66.40) 11.25 (55.45) -7.54 (46.23) 12.45 (56.03)
Kimi-K2.6 37.85 (67.85) 10.80 (54.84) -11.46 (43.97) 12.40 (55.55)
DeepSeek-V4-Flash 30.49 (62.05) 9.07 (52.95) -6.00 (46.60) 11.19 (53.87)
NVIDIA-Nemotron-3-Ultra-550B 33.59 (63.06) 7.16 (48.67) -15.29 (41.68) 8.49 (51.14)
MiniMax-M3 28.34 (62.42) 7.70 (50.06) -20.21 (39.66) 5.28 (50.71)
Mistral-Medium-3.5-128B 19.16 (57.54) -5.04 (46.02) -24.81 (37.35) -3.56 (46.97)
gpt-oss-120b 18.71 (58.82) -16.08 (41.58) -28.18 (35.79) -8.51 (45.40)
GLM-5.2 2.88 (50.76) -4.90 (44.90) -23.98 (38.01) -8.67 (44.55)
Qwen3.5-35B-A3B 8.49 (52.91) -12.58 (32.00) -41.37 (29.28) -15.15 (38.06)
Olmo-3.1-32B-Think 1.76 (32.32) -27.87 (25.30) -51.35 (21.75) -25.82 (26.45)
LFM2.5-8B-A1B -9.97 (32.89) -35.61 (26.57) -61.34 (17.52) -35.64 (25.66)
Apertus-v1.5-70B -48.21 (24.54) -74.82 (11.62) -88.17 (5.91) -70.40 (14.02)
Llama-3.3-70B-Instruct -48.22 (25.89) -75.51 (12.22) -86.71 (6.65) -70.15 (14.92)

(Hy-MT2-30B did not run LEXam.)


Notes & caveats

  • Different scales: judge families are 0–100; MCQ is a 0–1 trad_score (plain accuracy — IDK counts as wrong). The composite rescales MCQ to 0–100 — treat overall as an indicative aggregate, not a normalized benchmark score.
  • Judge self-preference: SLDS is judged by DeepSeek-V4-Pro, so DeepSeek models' SLDS scores may carry mild self-preference bias.
  • MCQ difficulty scales with option count: every model drops from mcq_4 → mcq_16, as expected.
  • Top tier is tight: Nemotron, Kimi, both DeepSeek variants and MiniMax-M3 sit within ~2.5 points overall. Nemotron leads SLDS and LEXam OQ; GLM-5.2 leads translation.
  • GLM-5.2 is the best translator (66.24 avg, top on slt) but Apertus posts the highest sdst (71.68). GLM's overall is dragged down by weak MCQ and the lowest LEXam-OQ of that tier — a strong-translation / weaker-exam profile.
  • Apertus-v1.5-70B is a near-frontier translator (65.21, third overall; best sdst) whose MCQ collapses to 14.02, same failure mode as Llama-3.3-70B.
  • Mistral-Medium-3.5-128B lands mid-pack at 48.11: solid translation (58.69) and MCQ (46.97) but weak SLDS (33.97).
  • MiniMax-M3 lands mid-frontier: strong translation (63.98, near the top) but middling MCQ.
  • Llama-3.3-70B-Instruct is a translation specialist: it scores 60.77 on translation but only 16.10 on SLDS and 14.92 on MCQ.
  • Small open models lag hard on summarization (slds) and press-release translation (sscprt) — LFM2.5-8B and Olmo are near-floor on slds/sscprt.

Provenance

  • Source: results/results/**/results_*.json (latest run per model).
  • Regenerate: ./.venv/bin/python -m swiss_legal_evals.aggregate then ... -m swiss_legal_evals.plot.
  • Per-family bar charts: plots/*.html. Tidy tables: results/summary_long.csv, results/summary_family_mean.csv, results/summary.csv.

Not included

  • qwen3.5-397b — paused (never run to completion).

Xet Storage Details

Size:
7.44 kB
·
Xet hash:
8df9bc212f39a79fe387feccff9fc099b44e16cef4639def41ea11e4d96c6413

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.