XReyRobert commited on
Commit
574917e
·
verified ·
1 Parent(s): 7b10e14

Slim Smoke24 publication card

Browse files
Files changed (1) hide show
  1. README.md +6 -10
README.md CHANGED
@@ -131,11 +131,11 @@ Completed artifact checks:
131
  - Index metadata total size matches the local safetensor shards.
132
  - The remote artifact contains the expected five safetensor shards.
133
 
134
- Terminal-Bench 2.0 Smoke24 result:
135
 
136
- | Run | Score | Success rate | Wall-time | Output tokens | Observed decode | LLM API time | Local efficiency |
137
- |---|---:|---:|---:|---:|---:|---:|---:|
138
- | `nex-n2-mini-gptq-pro` | `14/24` | `58.3%` | `314.6m` | `1670.6k` | `140.8 tok/s` | `197.4m` | `0.64x` |
139
 
140
  Smoke24 is a fixed 24-task Terminal-Bench 2.0 comparison corpus, not a full
141
  Terminal-Bench leaderboard run. In this harness, Nex-N2-mini GPTQ-Pro tied the
@@ -143,13 +143,9 @@ Qwen3.6 27B GPTQ reference on solved tasks but used more wall time and far more
143
  output tokens. That makes it a useful candidate for further serving and
144
  generation-control tuning, not an efficiency leader in this specific test.
145
 
146
- Full local report and CSVs are included under:
147
 
148
- - [`benchmarks/terminal-bench-2.0/tb20_smoke24_model_comparison_20260617.md`](benchmarks/terminal-bench-2.0/tb20_smoke24_model_comparison_20260617.md)
149
- - [`benchmarks/terminal-bench-2.0/tb20_smoke24_model_comparison_20260617.csv`](benchmarks/terminal-bench-2.0/tb20_smoke24_model_comparison_20260617.csv)
150
- - [`benchmarks/terminal-bench-2.0/tb20_smoke24_local_efficiency_metrics_20260617.csv`](benchmarks/terminal-bench-2.0/tb20_smoke24_local_efficiency_metrics_20260617.csv)
151
-
152
- ![Terminal-Bench Smoke24 local efficiency](benchmarks/terminal-bench-2.0/tb20_smoke24_score_llm_cost_per_success_bubble.png)
153
 
154
  ## MTP And Vision Status
155
 
 
131
  - Index metadata total size matches the local safetensor shards.
132
  - The remote artifact contains the expected five safetensor shards.
133
 
134
+ Terminal-Bench 2.0 Smoke24 result and associated vLLM serving measurements:
135
 
136
+ | Run | Score | Success rate | Wall-time | Output tokens | Observed decode | LLM API time |
137
+ |---|---:|---:|---:|---:|---:|---:|
138
+ | `nex-n2-mini-gptq-pro` | `14/24` | `58.3%` | `314.6m` | `1670.6k` | `140.8 tok/s` | `197.4m` |
139
 
140
  Smoke24 is a fixed 24-task Terminal-Bench 2.0 comparison corpus, not a full
141
  Terminal-Bench leaderboard run. In this harness, Nex-N2-mini GPTQ-Pro tied the
 
143
  output tokens. That makes it a useful candidate for further serving and
144
  generation-control tuning, not an efficiency leader in this specific test.
145
 
146
+ Task list and harness shape:
147
 
148
+ - [`benchmarks/terminal-bench-2.0/smoke24_task_list_20260616.md`](benchmarks/terminal-bench-2.0/smoke24_task_list_20260616.md)
 
 
 
 
149
 
150
  ## MTP And Vision Status
151