XReyRobert commited on
Commit
076fdac
·
verified ·
1 Parent(s): bbe9e60

Replace remaining informal benchmark labels

Browse files
Files changed (1) hide show
  1. README.md +2 -2
README.md CHANGED
@@ -234,7 +234,7 @@ the observed vLLM GPTQ-Marlin runtime behavior for this deployment shape.
234
  <span style="font-size: 11px; color: #64748b; font-weight: 500;">318 / 350 single-pass run</span>
235
  </div>
236
  <div style="border: 1px solid #e2e8f0; padding: 16px; border-radius: 10px; background: #f8fafc; text-align: center; box-shadow: inset 0 2px 4px rgba(0,0,0,0.02);">
237
- <span style="font-size: 11px; font-weight: 700; color: #7c3aed; text-transform: uppercase; display: block; margin-bottom: 6px; letter-spacing: 0.5px;">🧠 vs Qwopus v2</span>
238
  <span style="font-size: 24px; font-weight: 800; color: #10b981; display: block;">+0.86 pp</span>
239
  <span style="font-size: 11px; color: #64748b; font-weight: 500;">315 / 350 selected-subset run</span>
240
  </div>
@@ -358,7 +358,7 @@ Terminal-Bench 2.0 Smoke24 is a fixed 24-task coding-agent comparison corpus.
358
  It is useful for fast regression and local serving comparison, but it is not a
359
  full Terminal-Bench leaderboard submission.
360
 
361
- This Ornith GPTQ-Pro run used the long-context card-validation shape:
362
  `max_model_len=262144`, `max_input_tokens=220000`, 30 minute task timeout,
363
  32 CPU / 48 GiB sandbox, `thinking_token_budget=32768`, `max_output_tokens=40000`,
364
  temperature `1.0`, top-p `0.95`, top-k `20`, and `preserve_thinking=true`.
 
234
  <span style="font-size: 11px; color: #64748b; font-weight: 500;">318 / 350 single-pass run</span>
235
  </div>
236
  <div style="border: 1px solid #e2e8f0; padding: 16px; border-radius: 10px; background: #f8fafc; text-align: center; box-shadow: inset 0 2px 4px rgba(0,0,0,0.02);">
237
+ <span style="font-size: 11px; font-weight: 700; color: #7c3aed; text-transform: uppercase; display: block; margin-bottom: 6px; letter-spacing: 0.5px;">🧠 vs XReyRobert/Qwopus3.6-27B-v2-GPTQ-Pro-v1</span>
238
  <span style="font-size: 24px; font-weight: 800; color: #10b981; display: block;">+0.86 pp</span>
239
  <span style="font-size: 11px; color: #64748b; font-weight: 500;">315 / 350 selected-subset run</span>
240
  </div>
 
358
  It is useful for fast regression and local serving comparison, but it is not a
359
  full Terminal-Bench leaderboard submission.
360
 
361
+ `XReyRobert/Ornith-1.0-35B-GPTQ-Pro-FOEM-4bit-g128-ns256` used the long-context card-validation shape:
362
  `max_model_len=262144`, `max_input_tokens=220000`, 30 minute task timeout,
363
  32 CPU / 48 GiB sandbox, `thinking_token_budget=32768`, `max_output_tokens=40000`,
364
  temperature `1.0`, top-p `0.95`, top-k `20`, and `preserve_thinking=true`.