# Supported harnesses for the HuggingEnvs data-agent suite.
#
# ONE source of truth. Sweep scripts read this file rather than carrying their own list, so a
# support decision is made once and cannot drift between Stage 1, Stage 2 and Stage 3.
#
# Measured 2026-09-01, Qwen3.5-4B held fixed, indices 0 (scalar-variant) + 21 (json-variant).
# Blank lines and #-comments are ignored. See SUITE_FINDINGS.md for the full table.

# ── stable: safe for cross-model and cross-harness comparison ────────────────────────────────────
mini-swe-agent      # 6/5 turns
opencode            # 6/9
openhands-sdk       # 4/4
qwen-coder          # 4/4
vibe                # 4/4
gemini-cli          # 9/8

# ── works, but budget the turn count ────────────────────────────────────────────────────────────
swe-agent           # 37/45 turns
trae-agent          # 112/200 turns -- ~20x the stable harnesses

# ── works; graded cleanly but scored 0.0 on the json-variant task ───────────────────────────────
claude-code         # 5/5
codex               # 3/8 -- also the only harness observed to produce no answer.txt at all
pi                  # 5/7
terminus-2          # 5/6, both answers wrong: a real 0.0, not a failure

# ── NOT SUPPORTED (excluded 2026-09-01) ─────────────────────────────────────────────────────────
# goose      — erratic cost: 6 turns on one rollout, 347 on another of the SAME task. Unbudgetable,
#              and a sweep cannot be sized when one harness can consume 50x its expected wall-clock.
# openclaw   — `. ~/.nvm/nvm.sh && nvm use 22 && openclaw agent --local ...` exits 1. The agent never
#              makes a model call, so the suite's missing-answer default scores it a countable 0.0.
# kimi-cli   — reached 14 turns then produced no reward at all (`rewards={}`), both tasks.
