a11oy / benchmarks /restraint /README.md
betterwithage's picture
feat(restraint): a11oy Restraint — governed + measured frugality gate
714d76d verified
|
Raw
History Blame
1.73 kB

a11oy Restraint — benchmark (Ponytail methodology, our measurements)

Two arms — no-skill baseline vs a11oy-restraint — over the same five everyday tasks Ponytail uses (email validator, JS debounce, CSV sum, React countdown, FastAPI rate-limit). Code LOC is counted from fenced code blocks; tokens and latency come from the API. Median reported.

Honesty labels (Doctrine v11)

  • MEASURED — printed ONLY for a run we actually executed on our stack.
  • SAMPLE — an illustrative fixture derived from our ladder model (the /api/a11oy/v1/restraint/bench endpoint returns this when no model run is wired on the Space).
  • ROADMAP — the overall bench label until a real run is wired.

We never reprint Ponytail's published numbers as ours.

Reproduce

cp ../../.env.example ../../.env     # add your model key
npx promptfoo@latest eval -c benchmarks/restraint/promptfooconfig.yaml --repeat 10
npx promptfoo@latest view

Once a run completes, paste the measured medians into a results/ file and the UI/endpoint can surface them as MEASURED for that run only.

Citation (Ponytail, MIT)

a11oy Restraint adopts the 6-rung ladder and the lite/full/ultra intensity levels from the open-source Ponytail coding-agent skill (https://github.com/DietrichGebert/ponytail, MIT, © 2026 DietrichGebert) — adopted + governed, not invented here. Ponytail's published results (80–94% less code, 47–77% cheaper, 3–6× faster; median of 10 runs across Haiku/Sonnet/Opus) are cited as Ponytail's, never claimed as ours. Our contribution is the governance (signed DSSE receipts + advisory Λ), the measured-on-our-stack benchmark, and the J/token energy tie-in.