Refine SEGMENT model card, evaluation and loading documentation
Browse files- EVALUATION.md +22 -39
- LINEAGE.md +34 -15
- LOAD.md +40 -30
- ORGANIZER_ACCESS.md +4 -8
- PLATFORM_COMPATIBILITY.md +43 -7
- README.md +34 -12
EVALUATION.md
CHANGED
|
@@ -1,62 +1,45 @@
|
|
| 1 |
-
#
|
| 2 |
|
| 3 |
-
|
| 4 |
|
| 5 |
-
|
| 6 |
|
| 7 |
-
|
| 8 |
|
| 9 |
-
|
| 10 |
-
|
| 11 |
-
The full retained ZIP500 was qualified locally at the stated 4 GiB and 16 GiB tmpfs profiles, alongside loose500, healthy20, reversed20, default and persistent canonical calls, and controlled recovery. These fixtures do not establish the platform scratch quota, all possible input sizes, continuous resource peaks, cold-cache requirements, general order invariance, new semantic scores, independent-holdout performance or official platform success. Canonical deferrals and deliberately failed recovery rows are reduced scientific coverage. The prior R3 4 GiB eager-ZIP failure is a conditional local capacity reproduction, not proof of the organizer failure cause. Earlier images and failed attempts remain historical.
|
| 12 |
-
|
| 13 |
-
Measured grader completions: September 10, 2026, 10:53:44 UTC (candidate) and 10:55:13 UTC (reference). Complete readback and actual GPU-return checks passed for both. Paired reporting completed at 11:07:14 UTC.
|
| 14 |
|
| 15 |
| Population | Reference correct | Candidate correct | Reference micro accuracy | Candidate micro accuracy | Difference |
|
| 16 |
| --- | ---: | ---: | ---: | ---: | ---: |
|
| 17 |
-
| Full
|
| 18 |
| Strict paired sensitivity | 1,906 / 3,940 | 2,334 / 3,940 | 48.376% | 59.239% | +10.863 pp |
|
| 19 |
-
|
|
| 20 |
-
|
| 21 |
-
The full board has 400 cases from each of ten original source videos. Its equal-video estimate therefore matches the micro estimate: +10.825 percentage points, with descriptive paired two-level bootstrap 95% interval [+9.049, +12.501] points (1,000 draws, seed 42). All ten source-video differences are positive, ranging from +7.5 to +13.25 points. There are 650 paired wins, 217 losses and 3,133 ties. The strict sensitivity equal-video difference is +10.877 points, interval [+9.209, +12.588]. All score rows and parsing checks are complete; neither mode has a parsing failure.
|
| 22 |
-
|
| 23 |
-
These intervals are conditional on this development board and are unadjusted for repeated model selection or multiple capabilities. The board has inherited training and adaptive-selection exposure. It is not an independent finals holdout.
|
| 24 |
-
|
| 25 |
-
A material fine-capability weakness remains: 5b, causal consequence reasoning, regressed on 62 cases from six source videos. Its micro difference is −17.742 points and equal-video difference −16.881 points, with unadjusted descriptive interval [−34.671, −1.957]. This must remain visible alongside overall gains. The 34-case 5a micro result ties; its equal-video difference is slightly negative. These are local fine-capability results, not official capability-family/ID-OOD buckets or ranking votes.
|
| 26 |
-
|
| 27 |
-
The measured population compares the registered reference and candidate serving implementations with the same fixed Qwen3.5-27B judge, revision, local control check, hardware envelope and 20-case batching protocol. Candidate archive ed85dc914e846f0586c01e71a1427dd64caba0a80844b02e28feedfcb0dcce88 separately passed its exact-container 20-case H100 test. This report does not mean all 4,000 cases ran through that Docker archive, does not establish best across all historical alternatives, and does not predict an official finals score.
|
| 28 |
-
|
| 29 |
-
Strict-excluded pairs 0004, 0130 and 0196 remain distinct from recovered pairs 0074, 0111 and 0123. All three historical failed attempts retain their failed status. The fixed recovery policy selected whole pairs without using scores. No first-attempt reliability or model-only causal claim is made.
|
| 30 |
-
|
| 31 |
-
The paired report SHA256 is `a73eedbcbdfeac8f8500d4730980984518209b035f9571d6c10293390ae7035b`. Reference accepted observation is `c613cfe00d4576f671cdcb587be378e2bf55e476c057f13bb741d9f604363cf1`; candidate observation is `2b5f77c9f4a913cecaae405401ac05c15d045ba896d204ed8ff0d5759450d3d4`. The monorepo custody record `packages/rehearsal-segment/review/preparation/closed3-full4000-comparison-actual-r1/CUSTODY.json` retains both actual returns, complete readbacks, commands, reporting inputs and machine-readable report.
|
| 32 |
-
|
| 33 |
-
|
| 34 |
-
This private model-card follow-up was prepared on September 10, 2026. The historical local runtime evidence in the following paragraphs applies only to archive `ed85dc914e846f0586c01e71a1427dd64caba0a80844b02e28feedfcb0dcce88`, Docker config `d8fcfca448b374aab5d37c23382a02ef24367d7d66f6c5fbe194b65eb07f308b`, and the unchanged immutable asset/metadata revisions in LOAD.md. The exact20 Docker process took 282.298466seconds, with sampled GPU memory maximum 43,589 MiB, no OOM termination and positive first-frame diagnostics for all 20 clips.
|
| 35 |
|
| 36 |
-
|
| 37 |
|
| 38 |
-
The
|
| 39 |
|
| 40 |
-
|
| 41 |
|
| 42 |
-
|
| 43 |
|
| 44 |
-
The
|
| 45 |
|
| 46 |
-
|
| 47 |
|
| 48 |
-
|
| 49 |
|
| 50 |
-
|
| 51 |
|
|
|
|
| 52 |
|
| 53 |
-
|
| 54 |
|
| 55 |
-
|
| 56 |
|
| 57 |
-
The
|
| 58 |
|
| 59 |
-
|
| 60 |
|
|
|
|
| 61 |
|
| 62 |
-
|
|
|
|
| 1 |
+
# Evaluation
|
| 2 |
|
| 3 |
+
This page separates local answer-quality measurements, container runtime checks and the team-confirmed finals submission. The underlying records remain in [EVALUATION.json](EVALUATION.json), [PLATFORM_COMPATIBILITY.json](PLATFORM_COMPATIBILITY.json) and [SUBMISSION_STATUS.json](SUBMISSION_STATUS.json).
|
| 4 |
|
| 5 |
+
## Development comparison
|
| 6 |
|
| 7 |
+
Two registered serving implementations were evaluated with the same fixed Qwen3.5-27B judge and revision, hardware envelope and 20-question batching protocol. The board contains 4,000 questions from ten original source videos, with 400 questions per video. Analysis clusters by original video, rather than qID-derived clip filenames.
|
| 8 |
|
| 9 |
+
The reference is a registered implementation derived from the historical 333 framework. Its agreement with the actual 333 Docker image's answers has not been established. The candidate measurement likewise does not mean that all 4,000 questions ran through the final Docker archive.
|
|
|
|
|
|
|
|
|
|
|
|
|
| 10 |
|
| 11 |
| Population | Reference correct | Candidate correct | Reference micro accuracy | Candidate micro accuracy | Difference |
|
| 12 |
| --- | ---: | ---: | ---: | ---: | ---: |
|
| 13 |
+
| Full development board | 1,940 / 4,000 | 2,373 / 4,000 | 48.500% | 59.325% | +10.825 pp |
|
| 14 |
| Strict paired sensitivity | 1,906 / 3,940 | 2,334 / 3,940 | 48.376% | 59.239% | +10.863 pp |
|
| 15 |
+
| Excluded diagnostic subset | 34 / 60 | 39 / 60 | 56.667% | 65.000% | +8.333 pp |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 16 |
|
| 17 |
+
On the full board, equal-video weighting gives the same difference as micro accuracy. A paired two-level bootstrap gives a descriptive 95% interval of **+9.049 to +12.501 percentage points** (1,000 draws; seed 42). All ten source-video differences are positive, ranging from +7.5 to +13.25 points. There are 650 paired wins, 217 losses and 3,133 ties. The strict sensitivity's equal-video difference is +10.877 points, with interval +9.209 to +12.588.
|
| 18 |
|
| 19 |
+
The strict analysis excludes pairs `0004`, `0130` and `0196`. Separately recovered pairs `0074`, `0111` and `0123` were accepted as whole pairs under a policy that did not use scores. Historical failed attempts remain failures; recovery does not establish first-attempt reliability. Both accepted result sets have complete score rows and no parsing failures.
|
| 20 |
|
| 21 |
+
The grader runs completed on September 10, 2026 at 10:53:44 UTC for the candidate and 10:55:13 UTC for the reference. The paired report completed at 11:07:14 UTC; its SHA-256 is `a73eedbcbdfeac8f8500d4730980984518209b035f9571d6c10293390ae7035b`. [EVALUATION.json](EVALUATION.json) retains the observation identities and report binding.
|
| 22 |
|
| 23 |
+
## Limits of the comparison
|
| 24 |
|
| 25 |
+
The board was repeatedly used for model selection and has inherited training and data exposure. It is not an independent holdout, and it has not been established as the platform's out-of-distribution population. The confidence intervals are conditional on this board and are unadjusted for repeated selection or multiple capability comparisons.
|
| 26 |
|
| 27 |
+
**Causal-consequence reasoning is a measured weakness.** Capability 5b regressed on 62 cases from six source videos: −17.742 points in micro accuracy and −16.881 in the equal-video estimate, with a descriptive interval of −34.671 to −1.957. Capability 5a ties in micro accuracy on 34 cases and is slightly negative under equal-video weighting. These fine-capability results are separate from official ID/OOD buckets and finals ranking votes.
|
| 28 |
|
| 29 |
+
The local scorer passed 14 convenience controls: six positive and eight negative. Those controls test response and parser boundaries. They do not measure agreement with the organizer's judge, false-positive rates on real questions, or official score calibration.
|
| 30 |
|
| 31 |
+
No matched comparison in the candidates' actual serving forms establishes that R4 exceeds N3C-v2 or the original 333 image. Historically, the platform recorded 0.5207852833 for 333 and 0.5199292864 for N3C-v2 under different batching. Those scores do not supply an R4 forecast. No official finals score, significance-adjusted bucket ranking or Copeland outcome is inferred from the local results.
|
| 32 |
|
| 33 |
+
## Runtime qualification
|
| 34 |
|
| 35 |
+
The R4 qualification matrix completed eight test cases, nine application invocations and 1,593 answer executions across 533 unique qIDs. It covered ordinary and reversed 20-question inputs, 500 loose clips, the same 500 clips in ZIP archives at 4 GiB and 16 GiB scratch limits, default and repeated canonical calls, and controlled per-question recovery.
|
| 36 |
|
| 37 |
+
The ordinary 20-question and loose 500-question outputs matched their saved references. Reversed inputs matched the same image's ordinary outputs; both ZIP runs matched its loose 500-question outputs. The canonical fixture retained three documented pre-generation deferrals. A 23-question recovery fixture produced exactly its three intended empty answers and preserved the 20 healthy outputs. These counts therefore include repeated questions and reduced neural coverage; they are not semantic accuracy measurements.
|
| 38 |
|
| 39 |
+
The 500-question runs took approximately 1,226.6–1,267.1 seconds by independent runtime measurement. Self-reported answer latencies, sampled resource maxima and independent elapsed time remain distinct in the records. These fixtures do not establish every input size, cold-cache condition, general order invariance, platform scratch allocation or continuous resource peak. See [platform compatibility](PLATFORM_COMPATIBILITY.md) for recovery behavior and the remaining failure boundaries.
|
| 40 |
|
| 41 |
+
Earlier archive tests and failures remain preserved in [EVALUATION.json](EVALUATION.json) and the [immutable pre-refinement report](https://huggingface.co/HeyDonto/SURGFIELD-ORena-2026-SEGMENT-FINAL/blob/24133f8d4a491845db2a19c84c2d182f8b4e4831/EVALUATION.md). Their image identities and test conditions must be retained when interpreting them. In particular, reproducing R3's ZIP extraction failure locally does not prove the cause of the earlier failed platform submission.
|
| 42 |
|
| 43 |
+
## Finals status
|
| 44 |
|
| 45 |
+
HeyDonto Labs confirmed successful submission of the R4 archive, with no errors reported, on September 14, 2026. That is the date of confirmation, rather than a recorded submission timestamp. The platform identifiers, terminal logs, actual resource settings, final score and ranking have not been independently captured here. [SUBMISSION_STATUS.json](SUBMISSION_STATUS.json) records the evidence and exact artifact identity.
|
LINEAGE.md
CHANGED
|
@@ -1,30 +1,49 @@
|
|
| 1 |
-
#
|
| 2 |
|
| 3 |
-
|
| 4 |
|
| 5 |
-
|
| 6 |
|
| 7 |
-
The
|
| 8 |
|
| 9 |
-
The selected
|
| 10 |
|
| 11 |
-
|
| 12 |
|
| 13 |
-
|
| 14 |
|
| 15 |
-
The
|
| 16 |
|
| 17 |
-
The
|
| 18 |
|
| 19 |
-
The
|
| 20 |
|
|
|
|
| 21 |
|
| 22 |
-
|
| 23 |
|
| 24 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 25 |
|
| 26 |
-
|
| 27 |
|
| 28 |
-
|
| 29 |
|
| 30 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Model and data lineage
|
| 2 |
|
| 3 |
+
The submitted R4 application combines two distinct Qwen-based model paths, an OWLv2 detector and six inherited prior tables. [LINEAGE.json](LINEAGE.json), [MODEL_IDENTITY.json](MODEL_IDENTITY.json) and [MODEL_FILES.json](MODEL_FILES.json) preserve the exact component identities. [IMAGE_IDENTITY.json](IMAGE_IDENTITY.json) identifies the current serving archive.
|
| 4 |
|
| 5 |
+
## Main model
|
| 6 |
|
| 7 |
+
The private N3 backbone is the uniform FP32 mean of checkpoints S1001, S1003 and S1004 at step 17,535, serialized in BF16. Its combined named-shard digest is `a78c72786e43b779ab886d583d9a7da3fb6ed07e026b681cc8a7270b216acf8e`. Its 18 files, including private configuration and processor metadata, are retained byte for byte.
|
| 8 |
|
| 9 |
+
The selected continuation is `original-n3-main-lr2e5-step000900`: learning rate 2e-5, seed 908 and gradient accumulation of eight. The selected training prefix contains 7,200 draws and 4,130 unique rows from an eligible population of 7,207 rows. Later training steps and their exposure are not attributed to this checkpoint.
|
| 10 |
|
| 11 |
+
Its LoRA adapter has rank 16, alpha 32 and 15,335,424 parameter elements. Serving attaches one FP32 default adapter to the BF16 N3 backbone and leaves it unmerged. Adapter SHA-256: `0240d87b54503bd32502f5c77f45a60dbc404f69fa51e04b29ed6f84906697b5`. See [main selected exposure](main-SELECTED_EXPOSURE.json).
|
| 12 |
|
| 13 |
+
## Temporal model
|
| 14 |
|
| 15 |
+
The temporal path uses the public Qwen3-VL-8B-Instruct backbone at revision `0c351dd01ed87e9c1b53cbc748cba10e6187ff3b`. This is separate from the private N3 main backbone.
|
| 16 |
|
| 17 |
+
The selected continuation is `original-temporal-lr5e6-step000600`: learning rate 5e-6, seed 908 and gradient accumulation of eight. Its selected prefix contains 4,800 draws and 2,912 unique rows from 4,680 eligible rows.
|
| 18 |
|
| 19 |
+
The checkpoint contains the complete updated adapter initialized from the inherited temporal LoRA adapter, identified as a47. It is applied once to the public BF16 backbone and merged once during initialization; a47 is not applied as an additional adapter. Adapter SHA-256: `87fc446ac0d8cf6a3d6147f3d959ec21bdee1c806e00dc74f5a1ffe9e6ac7878`. The normalized serving configuration and original producer configuration remain separately identified in [LINEAGE.json](LINEAGE.json). See [temporal selected exposure](temporal-SELECTED_EXPOSURE.json).
|
| 20 |
|
| 21 |
+
## Training data and derived annotations
|
| 22 |
|
| 23 |
+
The inherited N3 corpus contains 46,758 rows:
|
| 24 |
|
| 25 |
+
| Source | Rows |
|
| 26 |
+
| --- | ---: |
|
| 27 |
+
| Organizer data | 33,581 |
|
| 28 |
+
| Paraphrases | 4,621 |
|
| 29 |
+
| Count annotations | 5,217 |
|
| 30 |
+
| Convention annotations | 2,317 |
|
| 31 |
+
| Adopted human-review annotations | 1,022 |
|
| 32 |
|
| 33 |
+
The continuation corpus is separately identified. Restricting the new continuation to organizer data does not remove inherited clinical or development-board exposure. Historical temporal smoke and full-training populations also remain separate. The team's recorded human-labeler and eligibility interpretation is not an organizer eligibility ruling. [TRAINING_RESOURCES.json](TRAINING_RESOURCES.json) retains producer identities and selected-checkpoint resource records.
|
| 34 |
|
| 35 |
+
The separately retained annotation packet contains the 13,177 derived rows across four JSONL files. It intentionally excludes all 33,581 organizer rows and underlying video pixels. It is not the complete ancestor corpus or a complete signature ledger. [ANNOTATION_SCOPE.json](ANNOTATION_SCOPE.json) supplies counts, hashes and provenance limits. A separate private dataset URL, applicable recipient conditions and delivery receipt remain to be recorded; row locators and review statements alone do not establish complete signature custody or redistribution permission.
|
| 36 |
|
| 37 |
+
## Detector and priors
|
| 38 |
+
|
| 39 |
+
OWLv2 uses the pinned `google/owlv2-large-patch14-ensemble` revision `95e26936e865f87db1742128404b3c035d47d89d`.
|
| 40 |
+
|
| 41 |
+
The six prior files retain their original bytes. `cholec_priors_v38.json` was derived from zero-shot OWLv2 detections on 613 frames from ten CholecTrack20 videos. This is prior construction, rather than a new optimizer-training corpus. The remaining duration, quadrant, stem and mode tables retain their organizer-training provenance. The release excludes the v40 prior and obsolete neural checkpoints.
|
| 42 |
+
|
| 43 |
+
## Serving revisions
|
| 44 |
+
|
| 45 |
+
The forward-decoding update changed frame access while preserving the tested sample indices, RGB bytes, source timestamps, frame limits and deadline. R3 corrected fixed SEGMENT routing and recorded pre-generation model deferrals. R4 then changed four wrapper files to limit ZIP staging and contain recoverable per-question errors.
|
| 46 |
+
|
| 47 |
+
R4 retains the 51 inherited serving source files and all 50 model/prior assets, including the selected adapters, prompts, model arithmetic and generation settings. These compatibility changes introduced no gradient updates. Details are in [PLATFORM_COMPATIBILITY.md](PLATFORM_COMPATIBILITY.md).
|
| 48 |
+
|
| 49 |
+
Upstream model notices, dataset restrictions and derived-resource terms are documented in [NOTICES.md](NOTICES.md). This repository does not grant blanket redistribution rights over all training data or derived assets.
|
LOAD.md
CHANGED
|
@@ -1,17 +1,17 @@
|
|
| 1 |
-
#
|
| 2 |
|
| 3 |
-
|
| 4 |
|
| 5 |
-
|
|
|
|
|
|
|
|
|
|
| 6 |
|
| 7 |
-
|
| 8 |
|
| 9 |
-
|
| 10 |
-
- Metadata revision **`bc9e092422beac9ae51005694824cdc70205e0d5`** contains `MODEL_FILES.json`, verifiers, component-loading code, notices and source copies. Its parent is the asset revision, but only the small metadata files are downloaded from it below.
|
| 11 |
|
| 12 |
-
|
| 13 |
-
|
| 14 |
-
Use an authenticated account with read access to the private repository. Download assets and metadata into **different local directories**:
|
| 15 |
|
| 16 |
```sh
|
| 17 |
python3 - <<'PYLOAD'
|
|
@@ -34,37 +34,47 @@ snapshot_download(repo_id=repo, revision=metadata_revision,
|
|
| 34 |
PYLOAD
|
| 35 |
```
|
| 36 |
|
| 37 |
-
The
|
| 38 |
|
| 39 |
-
|
| 40 |
-
|---|---|---|
|
| 41 |
-
| `main_model/` | `/app/artifacts/segment_final/main_model/` | Private N3 BF16 backbone and its saved processor metadata |
|
| 42 |
-
| `main_adapter/` | `/app/artifacts/segment_final/main_adapter/` | Selected step900 FP32 default LoRA; unmerged |
|
| 43 |
-
| `temporal_adapter/` | `/app/artifacts/segment_final/temporal_adapter/` | Complete updated low600 adapter; applied once |
|
| 44 |
-
| `upstream/qwen3-vl-8b-instruct/` | `/app/.cache/huggingface/hub/models--Qwen--Qwen3-VL-8B-Instruct/snapshots/0c351dd01ed87e9c1b53cbc748cba10e6187ff3b/` | Main processor plus temporal public backbone/processor |
|
| 45 |
-
| `upstream/owlv2-large-patch14-ensemble/` | `/app/.cache/huggingface/hub/models--google--owlv2-large-patch14-ensemble/snapshots/95e26936e865f87db1742128404b3c035d47d89d/` | OWLv2 detector and processor |
|
| 46 |
-
| `priors/` | `/app/artifacts/` | Exact six inherited prior files |
|
| 47 |
-
|
| 48 |
-
The public cache's `refs/main` resolves to its pinned revision for OWLv2's unchanged model-ID call. Snapshot files may be materialized as regular files; the model release paths do not require distributing the archive's internal symlinks. Full default-entrypoint reproduction also needs the exact serving source/bootstrap and frozen dependencies, which are already present in the corresponding Docker image. Do not treat a generic HF pipeline as that serving application.
|
| 49 |
-
|
| 50 |
-
The default Docker entrypoint is `/opt/conda/bin/python -I -S /opt/orena/segment/bootstrap.py`, user `app`, working directory `/app`, with offline HF/model settings. Preserve it. See the current [image identity](IMAGE_IDENTITY.json) and [platform compatibility record](PLATFORM_COMPATIBILITY.md) for the exact serving archive and its completed or pending validation; the historical metadata snapshot does not report current platform status.
|
| 51 |
-
|
| 52 |
-
The frozen main `load_main()` verifies the exact N3 and adapter files before and after loading, compares private/public tokenizer vocabularies and chat templates, loads BF16 N3 with SDPA on CUDA, then attaches exactly one selected default adapter. It does not merge the main adapter or cast its FP32 parameters. The frozen temporal loader verifies its exact public base path in the normalized adapter configuration, loads public Qwen BF16/SDPA, applies low600 once, and calls `merge_and_unload()`. The detector uses FP32 OWLv2 and its existing prompts. `model.safetensors` is selected ahead of the inactive `.bin` duplicate by the pinned Transformers resolver.
|
| 53 |
|
| 54 |
-
|
| 55 |
-
|
| 56 |
-
For local asset verification only (standard library, no model execution):
|
| 57 |
|
| 58 |
```sh
|
| 59 |
python3 -I -B ./segment-final-metadata/verify_assets.py --root ./segment-final-assets --manifest ./segment-final-metadata/MODEL_FILES.json
|
| 60 |
```
|
| 61 |
|
| 62 |
-
|
|
|
|
|
|
|
| 63 |
|
| 64 |
-
|
| 65 |
|
| 66 |
```sh
|
| 67 |
python3 -I -B ./segment-final-metadata/load_components.py --root ./segment-final-assets --manifest ./segment-final-metadata/MODEL_FILES.json --component all
|
| 68 |
```
|
| 69 |
|
| 70 |
-
The
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Download and loading
|
| 2 |
|
| 3 |
+
The model components and loading tools are available at two immutable revisions. Access requires an authenticated Hugging Face account with permission to read this private repository.
|
| 4 |
|
| 5 |
+
| Contents | Revision |
|
| 6 |
+
| --- | --- |
|
| 7 |
+
| 50 model and prior assets | `35fc28af1d55a5a9ae8096ff2e2cd2c961aa7119` |
|
| 8 |
+
| Asset manifest, verification and component-loading tools | `bc9e092422beac9ae51005694824cdc70205e0d5` |
|
| 9 |
|
| 10 |
+
These revisions reproduce the component files and loading tools. The current R4 serving archive is selected separately by [IMAGE_IDENTITY.json](IMAGE_IDENTITY.json); its archive SHA-256 is `e477d307ff2e09f469d83569ec1ce8703daeef34edbba633a3f240eab955de00`. See [platform compatibility](PLATFORM_COMPATIBILITY.md) for full-application execution and [submission status](SUBMISSION_STATUS.json) for the team's confirmation.
|
| 11 |
|
| 12 |
+
## 1. Download
|
|
|
|
| 13 |
|
| 14 |
+
Use an environment with `huggingface_hub` installed and authentication already configured. Keep model assets separate from the loader and verification metadata:
|
|
|
|
|
|
|
| 15 |
|
| 16 |
```sh
|
| 17 |
python3 - <<'PYLOAD'
|
|
|
|
| 34 |
PYLOAD
|
| 35 |
```
|
| 36 |
|
| 37 |
+
The metadata allowlist avoids downloading the large weights a second time. It also excludes that historical revision's `README.md`, `LOAD.md` and `IMAGE_IDENTITY.json`. Its retained manifests and lineage documents describe component ancestry and may identify the earlier ed85 image; those fields do not select the current serving archive. Use the current image record linked above.
|
| 38 |
|
| 39 |
+
## 2. Verify all files
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 40 |
|
| 41 |
+
On Linux/POSIX, run the standard-library verifier before loading:
|
|
|
|
|
|
|
| 42 |
|
| 43 |
```sh
|
| 44 |
python3 -I -B ./segment-final-metadata/verify_assets.py --root ./segment-final-assets --manifest ./segment-final-metadata/MODEL_FILES.json
|
| 45 |
```
|
| 46 |
|
| 47 |
+
The verifier reads all 50 assets to EOF and checks their SHA-256 hashes and sizes against the immutable manifest: 36,966,713,444 bytes in total. It uses POSIX filesystem interfaces, downloads no files and invokes no GPU. `MODEL_FILES.json` and `SERVING_ASSETS.json` preserve the same immutable inventory; the tokenizer, processor and configuration files are included among the assets.
|
| 48 |
+
|
| 49 |
+
## 3. Optional component-loading check
|
| 50 |
|
| 51 |
+
In the frozen tensor environment, the following command verifies the assets and loads the main, temporal and OWLv2 components on CUDA:
|
| 52 |
|
| 53 |
```sh
|
| 54 |
python3 -I -B ./segment-final-metadata/load_components.py --root ./segment-final-assets --manifest ./segment-final-metadata/MODEL_FILES.json --component all
|
| 55 |
```
|
| 56 |
|
| 57 |
+
The standalone helper generates no answers. It has been reviewed and checked at source/CPU level; this helper has not been independently qualified as a complete GPU serving application. The container's runtime qualification is separate. Individual component choices are `main`, `temporal` and `owlv2`, in addition to `all`.
|
| 58 |
+
|
| 59 |
+
The main loader checks private N3 and adapter bytes before and after loading, compares private/public tokenizer vocabularies and chat templates, loads BF16 N3 with SDPA and attaches exactly one FP32 default adapter without merging it. The helper relocates only the loader's local asset paths while retaining those checks.
|
| 60 |
+
|
| 61 |
+
The temporal loader uses public Qwen BF16/SDPA, applies the complete step-600 temporal adapter (`low600` in the serving code) once and calls `merge_and_unload()` once, followed by evaluation mode. OWLv2 uses its frozen FP32 model and processor sequence. The source copies are retained under `component_source/`.
|
| 62 |
+
|
| 63 |
+
## Full serving application
|
| 64 |
+
|
| 65 |
+
The submitted Docker image includes the complete bootstrap, router, sampler and frozen dependencies. Preserve its default entrypoint `/opt/conda/bin/python -I -S /opt/orena/segment/bootstrap.py`, user `app`, working directory `/app` and offline model settings. A generic Hugging Face pipeline or the component helper does not reproduce the full application.
|
| 66 |
+
|
| 67 |
+
The source environment pins Torch 2.6.0 with CUDA 12.4, torchvision 0.21.0, Transformers 4.57.6 and PEFT 0.19.1; remaining versions are in [requirements-frozen.txt](requirements-frozen.txt). Use the exact image environment when comparing serving behavior. This dependency record does not establish answer equivalence in an arbitrary host environment.
|
| 68 |
+
|
| 69 |
+
| Repository path | Frozen image path | Purpose |
|
| 70 |
+
|---|---|---|
|
| 71 |
+
| `main_model/` | `/app/artifacts/segment_final/main_model/` | Private N3 BF16 backbone and its saved processor metadata |
|
| 72 |
+
| `main_adapter/` | `/app/artifacts/segment_final/main_adapter/` | Selected step900 FP32 default LoRA; unmerged |
|
| 73 |
+
| `temporal_adapter/` | `/app/artifacts/segment_final/temporal_adapter/` | Complete updated low600 adapter; applied once |
|
| 74 |
+
| `upstream/qwen3-vl-8b-instruct/` | `/app/.cache/huggingface/hub/models--Qwen--Qwen3-VL-8B-Instruct/snapshots/0c351dd01ed87e9c1b53cbc748cba10e6187ff3b/` | Main processor plus temporal public backbone/processor |
|
| 75 |
+
| `upstream/owlv2-large-patch14-ensemble/` | `/app/.cache/huggingface/hub/models--google--owlv2-large-patch14-ensemble/snapshots/95e26936e865f87db1742128404b3c035d47d89d/` | OWLv2 detector and processor |
|
| 76 |
+
| `priors/` | `/app/artifacts/` | Exact six inherited prior files |
|
| 77 |
+
|
| 78 |
+
The public OWLv2 cache's `refs/main` points to its pinned revision for the unchanged model-ID call. Snapshot files can be materialized as regular files; the repository layout does not require distributing the archive's internal symlinks. The pinned resolver selects `model.safetensors` ahead of the inactive `.bin` duplicate, which is excluded from this release.
|
| 79 |
+
|
| 80 |
+
Main rendering uses up to 16 frames; temporal rendering uses up to 128 native frames with a maximum image side of 768 pixels. Generation is greedy with at most 64 new tokens, and the container loads its models eagerly. [LINEAGE.md](LINEAGE.md) and [PLATFORM_COMPATIBILITY.md](PLATFORM_COMPATIBILITY.md) document the distinct model paths and recovery behavior.
|
ORGANIZER_ACCESS.md
CHANGED
|
@@ -1,11 +1,7 @@
|
|
| 1 |
-
#
|
| 2 |
|
| 3 |
-
|
| 4 |
|
| 5 |
-
Attribution: HeyDonto Labs. Responsible contact: Reza Nehzati, Ph.D., via [rezanehzati](https://huggingface.co/rezanehzati).
|
| 6 |
|
| 7 |
-
|
| 8 |
-
|
| 9 |
-
Asset commit: `35fc28af1d55a5a9ae8096ff2e2cd2c961aa7119`. Metadata/tools commit: `bc9e092422beac9ae51005694824cdc70205e0d5`. The upload coordinator supplies the completed revision pair in final instructions or a separate receipt; no commit self-pin is required. No asset-copy, upload or access operation was performed by this draft author. FRAME membership and permissions are unchanged and outside this handoff.
|
| 10 |
-
|
| 11 |
-
The separate 13,177-row derived annotation packet requires its own private dataset URL, recipient-condition acceptance and delivery receipt. It is not a model-file attachment and does not contain the full46,758-row ancestor corpus or video pixels. No public annotation release or broader dataset permission is inferred here.
|
|
|
|
| 1 |
+
# Access and contact
|
| 2 |
|
| 3 |
+
The [SEGMENT model repository](https://huggingface.co/HeyDonto/SURGFIELD-ORena-2026-SEGMENT-FINAL) remains private under the agreed publication hold. Downloads require an authenticated account with repository read access. The team's recorded configuration uses a dedicated SEGMENT resource group with automatic joining disabled. The invitation to `orena-dkfz` remains deferred at the team's request; organizer membership and effective read access have not been confirmed. Successful finals submission does not establish repository access or permission for public release.
|
| 4 |
|
| 5 |
+
Attribution: **HeyDonto Labs**. Responsible contact: **Reza Nehzati, Ph.D.**, via [rezanehzati on Hugging Face](https://huggingface.co/rezanehzati). Use [LOAD.md](LOAD.md) for the immutable download revisions and [SUBMISSION_STATUS.json](SUBMISSION_STATUS.json) for the submission evidence.
|
| 6 |
|
| 7 |
+
The separate 13,177-row derived annotation packet requires its own private dataset URL, applicable recipient-condition acceptance and delivery receipt. It contains neither the full 46,758-row ancestor corpus nor source-video pixels, and is not attached as a model file. See [ANNOTATION_SCOPE.json](ANNOTATION_SCOPE.json) and [NOTICES.md](NOTICES.md) for its scope and restrictions. Public annotation release and broader dataset redistribution rights have not been established.
|
|
|
|
|
|
|
|
|
|
|
|
PLATFORM_COMPATIBILITY.md
CHANGED
|
@@ -1,13 +1,49 @@
|
|
| 1 |
-
#
|
| 2 |
|
| 3 |
-
|
| 4 |
|
| 5 |
-
|
| 6 |
|
| 7 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 8 |
|
| 9 |
-
|
| 10 |
|
| 11 |
-
The
|
| 12 |
|
| 13 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Serving and runtime qualification
|
| 2 |
|
| 3 |
+
The current artifact is the R4 container identified in [IMAGE_IDENTITY.json](IMAGE_IDENTITY.json). HeyDonto Labs confirmed its successful finals submission, with no errors reported, on September 14, 2026. The exact confirmation scope is recorded in [SUBMISSION_STATUS.json](SUBMISSION_STATUS.json).
|
| 4 |
|
| 5 |
+
## Container identity and interface
|
| 6 |
|
| 7 |
+
| Identity | Value |
|
| 8 |
+
| --- | --- |
|
| 9 |
+
| Archive SHA-256 | `e477d307ff2e09f469d83569ec1ce8703daeef34edbba633a3f240eab955de00` |
|
| 10 |
+
| Archive size | 35,613,293,722 bytes |
|
| 11 |
+
| Immutable GCS generation | `1789092734891901` |
|
| 12 |
+
| Docker image/configuration SHA-256 | `3f43fce1081dc5db454447dbe4678947e5cefec014635e611d8d1032a0ac71c4` |
|
| 13 |
+
| Build OCI manifest SHA-256 | `ba9ddb1e78a45040803ae6f0e68c66de7a538f5e3e86397777e5fc2c488106b0` |
|
| 14 |
|
| 15 |
+
Archive, configuration and manifest hashes identify different objects. The same-generation archive passed full gzip/CRC/tar verification, all 26 ordered layer DiffID checks and an actual classic Docker import. That import exposed the configuration identity but no imported manifest digest.
|
| 16 |
|
| 17 |
+
The container runs as user `app` from `/app`, with entrypoint `/opt/conda/bin/python -I -S /opt/orena/segment/bootstrap.py` and offline model settings. It reads the SEGMENT request, FO definitions and per-question clips, supplied loose or in ZIP archives, and writes `/output/answer.json`. The container always uses the SEGMENT task. Question times retain the original source-video clock.
|
| 18 |
|
| 19 |
+
## Input handling and recovery
|
| 20 |
+
|
| 21 |
+
R4 stages one selected ZIP member at a time for a startup probe or answer, verifies it through CRC/EOF and SHA-256, closes decoder handles and removes the staged file before continuing. Loose clips are read directly. Model initialization remains eager.
|
| 22 |
+
|
| 23 |
+
A recoverable error attributed to one question produces an empty answer for that qID, with elapsed latency and failure telemetry, while successful answers are retained and later questions continue. Unsafe or ambiguous archives, source mutation, cleanup-integrity failures, required model initialization failures, environment violations and fixed-track failures remain fatal.
|
| 24 |
+
|
| 25 |
+
Existing pre-generation deferrals to rules remain separately recorded. A valid output file can therefore include rule-based or empty answers; it does not establish healthy neural inference for every question.
|
| 26 |
+
|
| 27 |
+
The four R4 wrapper files are available under [platform_wrapper/](https://huggingface.co/HeyDonto/SURGFIELD-ORena-2026-SEGMENT-FINAL/tree/main/platform_wrapper). The 51 inherited serving source files and all 50 model/prior assets remain unchanged from R3, as do prompts, model arithmetic, frame limits and generation settings.
|
| 28 |
+
|
| 29 |
+
## Measured local coverage
|
| 30 |
+
|
| 31 |
+
The H100 qualification matrix completed eight test cases and nine application invocations, producing 1,593 answer executions across 533 unique qIDs:
|
| 32 |
+
|
| 33 |
+
| Test condition | Observed result |
|
| 34 |
+
| --- | --- |
|
| 35 |
+
| Ordinary and reversed 20-question inputs | Ordinary outputs matched the saved reference; reversed outputs matched the same image's ordinary outputs. |
|
| 36 |
+
| 500 loose clips | All contents matched the saved candidate reference. |
|
| 37 |
+
| The same 500 clips in ZIP archives, with 4 GiB and 16 GiB scratch limits | Both runs matched the same image's loose-clip outputs. |
|
| 38 |
+
| Default and repeated canonical calls | Outputs retained the three documented pre-generation deferrals. |
|
| 39 |
+
| Controlled 23-question recovery fixture | Exactly three intended empty answers; all 20 healthy outputs preserved. |
|
| 40 |
+
|
| 41 |
+
The independently measured 500-question runtimes were approximately 1,226.6–1,267.1 seconds. [PLATFORM_COMPATIBILITY.json](PLATFORM_COMPATIBILITY.json) and [EVALUATION.json](EVALUATION.json) retain the detailed identities and measurements. Content agreement is distinct from semantic accuracy; repeated questions and deliberate failures are included in these counts.
|
| 42 |
+
|
| 43 |
+
## Failure history and remaining limits
|
| 44 |
+
|
| 45 |
+
R3's eager extraction of the retained approximately 21 GB, 500-clip ZIP reproduced an `ENOSPC` failure locally with 4 GiB scratch, before model loading. R4 completed that fixture at the same scratch limit. This establishes a repaired local capacity failure, rather than the cause of the earlier platform submission failure, whose failing-case logs were not captured.
|
| 46 |
+
|
| 47 |
+
Earlier images also exposed input-alias and canonical-interval admission issues. Their tests and failures remain historical in [EVALUATION.json](EVALUATION.json) and the [immutable prior compatibility report](https://huggingface.co/HeyDonto/SURGFIELD-ORena-2026-SEGMENT-FINAL/blob/24133f8d4a491845db2a19c84c2d182f8b4e4831/PLATFORM_COMPATIBILITY.md). Narrow successful tests did not cover all later failure conditions.
|
| 48 |
+
|
| 49 |
+
The local matrix does not establish universal input-size limits, general order invariance, continuous resource peaks, cold-cache requirements or the organizer's scratch allocation. Sampled resource maxima, independent elapsed times and self-reported answer latencies are separate measurements. Actual submitted resource settings and terminal platform logs have not been independently recorded here. Successful submission supplies no official answer-quality score; see [EVALUATION.md](EVALUATION.md).
|
README.md
CHANGED
|
@@ -14,26 +14,48 @@ tags:
|
|
| 14 |
- orena-focus
|
| 15 |
- private-evaluation
|
| 16 |
---
|
| 17 |
-
# SURGFIELD ORena2026 SEGMENT FINAL
|
| 18 |
|
| 19 |
-
|
| 20 |
|
| 21 |
-
**
|
| 22 |
|
| 23 |
-
|
| 24 |
|
| 25 |
-
|
| 26 |
|
| 27 |
-
|
| 28 |
|
| 29 |
-
|
| 30 |
|
| 31 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 32 |
|
| 33 |
-
The
|
| 34 |
|
| 35 |
-
|
| 36 |
|
| 37 |
-
|
| 38 |
|
| 39 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 14 |
- orena-focus
|
| 15 |
- private-evaluation
|
| 16 |
---
|
|
|
|
| 17 |
|
| 18 |
+
# SURGFIELD ORena 2026 SEGMENT FINAL
|
| 19 |
|
| 20 |
+
**HeyDonto Labs** · Responsible contact: **Reza Nehzati, Ph.D.** ([Hugging Face profile](https://huggingface.co/rezanehzati))
|
| 21 |
|
| 22 |
+
SURGFIELD answers questions about surgical video segments for the ORena 2026 SEGMENT challenge. This repository contains the exact model components, configurations, tokenizer and processor assets, loading tools and documentation associated with the submitted R4 container.
|
| 23 |
|
| 24 |
+
HeyDonto Labs confirmed on September 14, 2026 that R4 was successfully submitted for SEGMENT finals, with no errors reported. The submitted archive SHA-256 is `e477d307ff2e09f469d83569ec1ce8703daeef34edbba633a3f240eab955de00`. September 14 is the confirmation date; the submission timestamp, platform identifiers, terminal logs and final score have not been independently recorded here. See [submission status](SUBMISSION_STATUS.json) and [image identity](IMAGE_IDENTITY.json).
|
| 25 |
|
| 26 |
+
## Method
|
| 27 |
|
| 28 |
+
The serving application combines visual models with fixed answer-format rules and six inherited prior tables. It preserves the original source-video clock when interpreting questions and sampling frames.
|
| 29 |
|
| 30 |
+
| Component | Role and configuration |
|
| 31 |
+
| --- | --- |
|
| 32 |
+
| Main visual-language model | Private N3 Qwen3-VL-8B backbone with the selected step-900 continuation adapter. The backbone uses BF16; the FP32 LoRA adapter remains unmerged. |
|
| 33 |
+
| Temporal model | A separate, pinned public Qwen3-VL-8B backbone with the selected step-600 temporal adapter, applied and merged once. |
|
| 34 |
+
| Object detector and priors | Pinned OWLv2 model, six prior tables and fixed routing/formatting rules. |
|
| 35 |
|
| 36 |
+
The main path samples up to 16 frames; the temporal path supports up to 128 native frames with a maximum image side of 768 pixels. Generation is greedy, with at most 64 new tokens. The private N3 and public Qwen backbones are distinct, and both are required. [Training and component lineage](LINEAGE.md) describes their construction and exposure.
|
| 37 |
|
| 38 |
+
## Download and use
|
| 39 |
|
| 40 |
+
Follow [LOAD.md](LOAD.md) to download the 50 model and prior files from immutable revisions, verify their checksums and optionally load the components. [MODEL_FILES.json](MODEL_FILES.json) and [WEIGHTS_SHA256SUMS](WEIGHTS_SHA256SUMS) identify the files.
|
| 41 |
|
| 42 |
+
The component-loading example generates no answers. Reproducing the complete application requires the corresponding Docker image, its frozen environment, routing and preprocessing. See [platform compatibility](PLATFORM_COMPATIBILITY.md) for the entrypoint, recovery behavior and measured runtime coverage.
|
| 43 |
+
|
| 44 |
+
## Evaluation
|
| 45 |
+
|
| 46 |
+
A development comparison used a fixed Qwen3.5-27B judge, 20-question batches and 4,000 questions from ten original source videos. These measurements cover registered implementations; a full 4,000-question evaluation of the submitted R4 Docker archive has not been performed.
|
| 47 |
+
|
| 48 |
+
| Registered serving implementation | Correct | Local micro accuracy |
|
| 49 |
+
| --- | ---: | ---: |
|
| 50 |
+
| Development candidate (registered implementation) | 2,373 / 4,000 | 59.325% |
|
| 51 |
+
| Reference derived from the historical 333 implementation | 1,940 / 4,000 | 48.500% |
|
| 52 |
+
|
| 53 |
+
The evaluation set was reused during development and has inherited data-exposure limitations. The comparison does not reproduce the original 333 Docker image, establish superiority over every historical candidate, or predict the official finals score. Causal-consequence questions regressed by 17.742 percentage points on 62 cases. [EVALUATION.md](EVALUATION.md) reports the paired sensitivity analysis, scorer limitations and runtime evidence separately.
|
| 54 |
+
|
| 55 |
+
## Data, access and limitations
|
| 56 |
+
|
| 57 |
+
Training includes organizer data and inherited derived annotations. The separately retained 13,177-row annotation packet excludes organizer rows and video pixels; its delivery and access are separate from this model repository. See [lineage](LINEAGE.md) and [annotation scope](ANNOTATION_SCOPE.json).
|
| 58 |
+
|
| 59 |
+
The repository remains private under the agreed publication hold. Organizer access is deferred; [ORGANIZER_ACCESS.md](ORGANIZER_ACCESS.md) records the access and delivery requirements. Upstream model licenses and dataset restrictions remain component-specific; [NOTICES.md](NOTICES.md) defines their scope.
|
| 60 |
+
|
| 61 |
+
This is a research and challenge system without clinical deployment qualification. Some inputs can use rule-based answers or explicit empty answers after a recoverable per-question failure. Runtime completion therefore does not establish neural coverage or answer correctness.
|