Spaces:
Runtime error
Runtime error
File size: 1,468 Bytes
a6a5d8e | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 | ---
license: other
license_name: proprietary
license_link: https://github.com/szl-holdings/a11oy/blob/main/LICENSE
tags:
- benchmark
- receipts
- governance
- mathcomp
- mirror-not-canonical
pretty_name: A11oy staged test-results schema
---
# A11oy test-results — staged dataset schema
This directory defines the future `SZLHOLDINGS/a11oy-test-results` dataset
layout. It is a **schema and manifest only** in this revision.
GitHub remains canonical. Hugging Face is a generated mirror for review.
## Current claim status
- No live benchmark score is claimed.
- No leaderboard metric is claimed.
- No benchmark corpus is redistributed here.
- No model-index metrics are published.
Competition-math benchmark scoring remains staged until corpus digest, receipts, reproducible
tooling, and judge agreement are present.
## Future dataset layout
```text
README.md
MANIFEST.json
benchmark-map.json
schemas/manifest.schema.json
schemas/result-row.schema.json
schemas/receipt-envelope.schema.json
samples/staged/*.jsonl
results/*.jsonl
receipts/*.jsonl
```
Only schema examples or receipt-backed staged dry-run artifacts may appear
before a sealed run exists. Real results require:
1. immutable corpus digest;
2. raw-score reporting;
3. three-judge panel;
4. append-only receipt chain;
5. unsupported-claim rejection;
6. GitHub CI validation.
Validate the current staged manifest with:
```bash
npm run hf:test-results:audit
npm run benchmark:audit
```
|