rezanehzati's picture
Refine SEGMENT model card, evaluation and loading documentation
411c900 verified
|
Raw History Blame Contribute Delete
6.35 kB
# Evaluation
This page separates local answer-quality measurements, container runtime checks and the team-confirmed finals submission. The underlying records remain in [EVALUATION.json](EVALUATION.json), [PLATFORM_COMPATIBILITY.json](PLATFORM_COMPATIBILITY.json) and [SUBMISSION_STATUS.json](SUBMISSION_STATUS.json).
## Development comparison
Two registered serving implementations were evaluated with the same fixed Qwen3.5-27B judge and revision, hardware envelope and 20-question batching protocol. The board contains 4,000 questions from ten original source videos, with 400 questions per video. Analysis clusters by original video, rather than qID-derived clip filenames.
The reference is a registered implementation derived from the historical 333 framework. Its agreement with the actual 333 Docker image's answers has not been established. The candidate measurement likewise does not mean that all 4,000 questions ran through the final Docker archive.
| Population | Reference correct | Candidate correct | Reference micro accuracy | Candidate micro accuracy | Difference |
| --- | ---: | ---: | ---: | ---: | ---: |
| Full development board | 1,940 / 4,000 | 2,373 / 4,000 | 48.500% | 59.325% | +10.825 pp |
| Strict paired sensitivity | 1,906 / 3,940 | 2,334 / 3,940 | 48.376% | 59.239% | +10.863 pp |
| Excluded diagnostic subset | 34 / 60 | 39 / 60 | 56.667% | 65.000% | +8.333 pp |
On the full board, equal-video weighting gives the same difference as micro accuracy. A paired two-level bootstrap gives a descriptive 95% interval of **+9.049 to +12.501 percentage points** (1,000 draws; seed 42). All ten source-video differences are positive, ranging from +7.5 to +13.25 points. There are 650 paired wins, 217 losses and 3,133 ties. The strict sensitivity's equal-video difference is +10.877 points, with interval +9.209 to +12.588.
The strict analysis excludes pairs `0004`, `0130` and `0196`. Separately recovered pairs `0074`, `0111` and `0123` were accepted as whole pairs under a policy that did not use scores. Historical failed attempts remain failures; recovery does not establish first-attempt reliability. Both accepted result sets have complete score rows and no parsing failures.
The grader runs completed on September 10, 2026 at 10:53:44 UTC for the candidate and 10:55:13 UTC for the reference. The paired report completed at 11:07:14 UTC; its SHA-256 is `a73eedbcbdfeac8f8500d4730980984518209b035f9571d6c10293390ae7035b`. [EVALUATION.json](EVALUATION.json) retains the observation identities and report binding.
## Limits of the comparison
The board was repeatedly used for model selection and has inherited training and data exposure. It is not an independent holdout, and it has not been established as the platform's out-of-distribution population. The confidence intervals are conditional on this board and are unadjusted for repeated selection or multiple capability comparisons.
**Causal-consequence reasoning is a measured weakness.** Capability 5b regressed on 62 cases from six source videos: βˆ’17.742 points in micro accuracy and βˆ’16.881 in the equal-video estimate, with a descriptive interval of βˆ’34.671 to βˆ’1.957. Capability 5a ties in micro accuracy on 34 cases and is slightly negative under equal-video weighting. These fine-capability results are separate from official ID/OOD buckets and finals ranking votes.
The local scorer passed 14 convenience controls: six positive and eight negative. Those controls test response and parser boundaries. They do not measure agreement with the organizer's judge, false-positive rates on real questions, or official score calibration.
No matched comparison in the candidates' actual serving forms establishes that R4 exceeds N3C-v2 or the original 333 image. Historically, the platform recorded 0.5207852833 for 333 and 0.5199292864 for N3C-v2 under different batching. Those scores do not supply an R4 forecast. No official finals score, significance-adjusted bucket ranking or Copeland outcome is inferred from the local results.
## Runtime qualification
The R4 qualification matrix completed eight test cases, nine application invocations and 1,593 answer executions across 533 unique qIDs. It covered ordinary and reversed 20-question inputs, 500 loose clips, the same 500 clips in ZIP archives at 4 GiB and 16 GiB scratch limits, default and repeated canonical calls, and controlled per-question recovery.
The ordinary 20-question and loose 500-question outputs matched their saved references. Reversed inputs matched the same image's ordinary outputs; both ZIP runs matched its loose 500-question outputs. The canonical fixture retained three documented pre-generation deferrals. A 23-question recovery fixture produced exactly its three intended empty answers and preserved the 20 healthy outputs. These counts therefore include repeated questions and reduced neural coverage; they are not semantic accuracy measurements.
The 500-question runs took approximately 1,226.6–1,267.1 seconds by independent runtime measurement. Self-reported answer latencies, sampled resource maxima and independent elapsed time remain distinct in the records. These fixtures do not establish every input size, cold-cache condition, general order invariance, platform scratch allocation or continuous resource peak. See [platform compatibility](PLATFORM_COMPATIBILITY.md) for recovery behavior and the remaining failure boundaries.
Earlier archive tests and failures remain preserved in [EVALUATION.json](EVALUATION.json) and the [immutable pre-refinement report](https://huggingface.co/HeyDonto/SURGFIELD-ORena-2026-SEGMENT-FINAL/blob/24133f8d4a491845db2a19c84c2d182f8b4e4831/EVALUATION.md). Their image identities and test conditions must be retained when interpreting them. In particular, reproducing R3's ZIP extraction failure locally does not prove the cause of the earlier failed platform submission.
## Finals status
HeyDonto Labs confirmed successful submission of the R4 archive, with no errors reported, on September 14, 2026. That is the date of confirmation, rather than a recorded submission timestamp. The platform identifiers, terminal logs, actual resource settings, final score and ranking have not been independently captured here. [SUBMISSION_STATUS.json](SUBMISSION_STATUS.json) records the evidence and exact artifact identity.