# Terminal-Bench 2.0 Smoke24 Task List This file documents the fixed 24-task Smoke24 corpus used for the model-card summary results. It intentionally lists the task names and harness shape only; it does not publish per-task outcomes, model-by-model traces, or aggregate CSV details. ## Corpus - Benchmark: Terminal-Bench 2.0 - Corpus ID: `tb20-coder-smoke24-fast-success-failure-20260616` - Created: 2026-06-16 - Size: 24 tasks - Selection policy: 12 shortest prior successes and 12 shortest prior failures from a recovery-corrected Qwopus3.6-27B-v2-GPTQ-Pro-v1 aggregate. ## Harness Shape - Agent: `terminus-2` - Concurrency: `1` - Sandbox: 32 CPU / 48 GiB RAM - Task timeout: 30 minutes - Max output: 40k tokens - Thinking token budget: 32768 - Sampling: temperature `1.0`, top-p `0.95`, top-k `20` ## Tasks ### Prior Success Group - `git-leak-recovery` - `prove-plus-comm` - `fix-git` - `modernize-scientific-stack` - `kv-store-grpc` - `openssl-selfsigned-cert` - `headless-terminal` - `multi-source-data-merger` - `fix-code-vulnerability` - `nginx-request-logging` - `build-pmars` - `code-from-image` ### Prior Failure Group - `feal-differential-cryptanalysis` - `raman-fitting` - `log-summary-date-ranges` - `video-processing` - `sqlite-with-gcov` - `sanitize-git-repo` - `torch-pipeline-parallelism` - `count-dataset-tokens` - `configure-git-webserver` - `cancel-async-tasks` - `mteb-retrieve` - `hf-model-inference`