SeaWolf-AI commited on
Commit
9eb2e46
·
verified ·
1 Parent(s): a393f62

eval: LEXam-hard 45.72

Browse files
Files changed (1) hide show
  1. .eval_results/lexam_hard.yaml +9 -0
.eval_results/lexam_hard.yaml ADDED
@@ -0,0 +1,9 @@
 
 
 
 
 
 
 
 
 
 
1
+ - dataset:
2
+ id: joelniklaus/LEXam-hard
3
+ task_id: lexam_hard
4
+ value: 45.72
5
+ date: '2026-09-30'
6
+ source:
7
+ url: https://huggingface.co/FINAL-Bench/Darwin-180B-RSI
8
+ name: Model Card
9
+ notes: "LEXam-hard, all 518 open questions; single sample (temperature=1.0, top_p=0.95, top_k=20); thinking budget 32,768 tokens, responses truncated at 32K regenerated with a 120K budget (60 items); judged by DeepSeek-R1-0528 with the LEXam paper judge prompt (eval.yaml); score = mean of German and English mean grades x 100 (de 44.37, en 47.07); 1 item without a parsable grade counted as 0; bf16"