Update LEXam-hard evaluation result
Browse filesAdds this model's score on the [LEXam-hard](https://huggingface.co/datasets/joelniklaus/LEXam-hard) benchmark, the 518 LEXam open questions the strongest open models score lowest on.
The score is the DeepSeek-R1-0528 judge grade (0-100) over those questions, aggregated as SwissLegalEvals aggregates LEXam (mean of the German and English means), recomputed from the per-sample outputs of the [SwissLegalEvals](https://huggingface.co/blog/joelniklaus/swiss-legal-evals) run (lighteval, LEXam paper prompts, one response per question, no tools). The raw outputs are in the public `joelniklaus/SwissLegalEvals` bucket; the recomputation is `reproduction/lexam_hard_results.py` in the dataset repository.
.eval_results/lexam-hard.yaml
CHANGED
|
@@ -1,12 +1,12 @@
|
|
| 1 |
- dataset:
|
| 2 |
id: joelniklaus/LEXam-hard
|
| 3 |
task_id: lexam_hard
|
| 4 |
-
revision:
|
| 5 |
-
value:
|
| 6 |
date: '2026-07-27'
|
| 7 |
source:
|
| 8 |
url: https://huggingface.co/buckets/joelniklaus/SwissLegalEvals
|
| 9 |
name: SwissLegalEvals per-sample details (lighteval)
|
| 10 |
user: joelniklaus
|
| 11 |
notes: lighteval, LEXam paper prompts, one response per question, no tools; DeepSeek-R1-0528 judge;
|
| 12 |
-
|
|
|
|
| 1 |
- dataset:
|
| 2 |
id: joelniklaus/LEXam-hard
|
| 3 |
task_id: lexam_hard
|
| 4 |
+
revision: 1bd50ee3ac80286fed27e5b89f03f4d8b4494244
|
| 5 |
+
value: 31.89
|
| 6 |
date: '2026-07-27'
|
| 7 |
source:
|
| 8 |
url: https://huggingface.co/buckets/joelniklaus/SwissLegalEvals
|
| 9 |
name: SwissLegalEvals per-sample details (lighteval)
|
| 10 |
user: joelniklaus
|
| 11 |
notes: lighteval, LEXam paper prompts, one response per question, no tools; DeepSeek-R1-0528 judge;
|
| 12 |
+
mean of the German and English means over the 518 questions, 0-100
|