Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -44,6 +44,8 @@ sampling; reasoning = `completion_tokens_details.reasoning_tokens`):
|
|
| 44 |
| GPQA-Diamond (full 198) | **78.3%** | 1,485 / 969 |
|
| 45 |
| CRUXEval-I (full 800, input prediction) | **92.1%** | 197 / 83 |
|
| 46 |
| CRUXEval-O (full 800, output prediction) | **92.9%** | 146 / 96 |
|
|
|
|
|
|
|
| 47 |
|
| 48 |
CRUXEval was run with the official Meta harness (direct prompts, official
|
| 49 |
extraction, exec-based scoring, temp 0.2) — code *understanding*
|
|
|
|
| 44 |
| GPQA-Diamond (full 198) | **78.3%** | 1,485 / 969 |
|
| 45 |
| CRUXEval-I (full 800, input prediction) | **92.1%** | 197 / 83 |
|
| 46 |
| CRUXEval-O (full 800, output prediction) | **92.9%** | 146 / 96 |
|
| 47 |
+
| HumanEval+ (164, official EvalPlus, greedy) | **90.2%** (93.9% base) | 43 / 28 |
|
| 48 |
+
| MBPP+ (378, official EvalPlus, greedy) | **78.6%** (92.9% base) | 91 / 25 |
|
| 49 |
|
| 50 |
CRUXEval was run with the official Meta harness (direct prompts, official
|
| 51 |
extraction, exec-based scoring, temp 0.2) — code *understanding*
|