rezanehzati commited on
Commit
411c900
·
verified ·
1 Parent(s): 24133f8

Refine SEGMENT model card, evaluation and loading documentation

Browse files
Files changed (6) hide show
  1. EVALUATION.md +22 -39
  2. LINEAGE.md +34 -15
  3. LOAD.md +40 -30
  4. ORGANIZER_ACCESS.md +4 -8
  5. PLATFORM_COMPATIBILITY.md +43 -7
  6. README.md +34 -12
EVALUATION.md CHANGED
@@ -1,62 +1,45 @@
1
- # SEGMENT evaluation evidence
2
 
3
- **Finals submission: owner confirmed on September 14, 2026 that repaired R4 was successfully submitted for SEGMENT finals, with no errors reported.** The submitted archive SHA256 is `e477d307ff2e09f469d83569ec1ce8703daeef34edbba633a3f240eab955de00`. September 14 is the confirmation date; the actual submission time, platform UUIDs, terminal logs, score and ranking were not supplied. See [SUBMISSION_STATUS.json](SUBMISSION_STATUS.json) for the exact identity and evidence scope. The model weights, source files and immutable loading revisions are unchanged.
4
 
5
- **September 11 local qualification record: R4 runtime, archive and import checks completed.** The exact archive is `e477d307ff2e09f469d83569ec1ce8703daeef34edbba633a3f240eab955de00` (35,613,293,722 bytes), immutable generation `1789092734891901`. The build manifest is `ba9ddb1e78a45040803ae6f0e68c66de7a538f5e3e86397777e5fc2c488106b0` and configuration `3f43fce1081dc5db454447dbe4678947e5cefec014635e611d8d1032a0ac71c4`. Full compressed/raw archive verification covers all 26 ordered DiffIDs; actual same-generation classic Docker import reports config `3f43fce1081dc5db454447dbe4678947e5cefec014635e611d8d1032a0ac71c4` and no imported manifest digest. The agreed eight cases completed 9 application invocations and 1593 answer executions across 533 unique qIDs. Healthy20 and loose/ZIP500 matched saved content; reverse20 matched the same-image ordinary20, and both ZIP quotas matched the same-image loose500. Canonical repeated calls retained the exact disclosed deferrals. The mixed recovery case blanked only its declared failed questions and preserved later healthy outputs. These are bounded runtime/interface/content checks; no new semantic grade or official final result is claimed.
6
 
7
- The completed development comparison is between two registered serving implementations under the same fixed Qwen3.5-27B judge and 20-case batching protocol: candidate 2373/4000 (59.325%) versus named reference 1940/4000 (48.500%), +10.825 percentage points. Strict sensitivity is 2334/3940 versus 1906/3940. All ten original-source-video differences are positive; 5b causal consequence reasoning regressed 17.742 micro points on 62 cases from six sources. These are adaptively used development results with inherited training/data-exposure limits, not an independent holdout, an official finals score or an all-candidate ranking. The registered local scorer passed14 convenience controls (6positive,8negative). These check response/parser boundaries; they do not estimate organizer-judge agreement, real-question false-positive rates, official calibration or a finals score. All 4,000 were not run through a Docker archive. Source-derived routing equality on the 4,000 intervals 1–299s does not by itself transfer these scores to new image bytes.
8
 
9
- R4 retains the image-owned literal SEGMENT task and the original request clocks. It stages only the selected ZIP member for each startup probe or answer, verifies that member through CRC/EOF and SHA-256, closes decoder handles, and removes the lease before continuing. Loose input remains a direct read. Model initialization is still eager. An attributed ordinary per-question exception produces empty content for that qID with measured latency and failure telemetry; successful answers are preserved and later questions continue. Unsafe or ambiguous archives, source mutation, cleanup integrity, required model initialization, environment and fixed-track failures remain fatal. The inherited precise pre-generation rules deferrals remain separately disclosed as degraded coverage. The four changed image files are the bootstrap, contract, probe entrypoint and new ZIP helper; all 51 unchanged inherited serving source files, including cap_fmt, all 50 model/prior assets, prompts, model arithmetic and generation settings remain unchanged.
10
-
11
- The full retained ZIP500 was qualified locally at the stated 4 GiB and 16 GiB tmpfs profiles, alongside loose500, healthy20, reversed20, default and persistent canonical calls, and controlled recovery. These fixtures do not establish the platform scratch quota, all possible input sizes, continuous resource peaks, cold-cache requirements, general order invariance, new semantic scores, independent-holdout performance or official platform success. Canonical deferrals and deliberately failed recovery rows are reduced scientific coverage. The prior R3 4 GiB eager-ZIP failure is a conditional local capacity reproduction, not proof of the organizer failure cause. Earlier images and failed attempts remain historical.
12
-
13
- Measured grader completions: September 10, 2026, 10:53:44 UTC (candidate) and 10:55:13 UTC (reference). Complete readback and actual GPU-return checks passed for both. Paired reporting completed at 11:07:14 UTC.
14
 
15
  | Population | Reference correct | Candidate correct | Reference micro accuracy | Candidate micro accuracy | Difference |
16
  | --- | ---: | ---: | ---: | ---: | ---: |
17
- | Full operational board | 1,940 / 4,000 | 2,373 / 4,000 | 48.500% | 59.325% | +10.825 pp |
18
  | Strict paired sensitivity | 1,906 / 3,940 | 2,334 / 3,940 | 48.376% | 59.239% | +10.863 pp |
19
- | Strict-excluded diagnostic | 34 / 60 | 39 / 60 | 56.667% | 65.000% | +8.333 pp |
20
-
21
- The full board has 400 cases from each of ten original source videos. Its equal-video estimate therefore matches the micro estimate: +10.825 percentage points, with descriptive paired two-level bootstrap 95% interval [+9.049, +12.501] points (1,000 draws, seed 42). All ten source-video differences are positive, ranging from +7.5 to +13.25 points. There are 650 paired wins, 217 losses and 3,133 ties. The strict sensitivity equal-video difference is +10.877 points, interval [+9.209, +12.588]. All score rows and parsing checks are complete; neither mode has a parsing failure.
22
-
23
- These intervals are conditional on this development board and are unadjusted for repeated model selection or multiple capabilities. The board has inherited training and adaptive-selection exposure. It is not an independent finals holdout.
24
-
25
- A material fine-capability weakness remains: 5b, causal consequence reasoning, regressed on 62 cases from six source videos. Its micro difference is −17.742 points and equal-video difference −16.881 points, with unadjusted descriptive interval [−34.671, −1.957]. This must remain visible alongside overall gains. The 34-case 5a micro result ties; its equal-video difference is slightly negative. These are local fine-capability results, not official capability-family/ID-OOD buckets or ranking votes.
26
-
27
- The measured population compares the registered reference and candidate serving implementations with the same fixed Qwen3.5-27B judge, revision, local control check, hardware envelope and 20-case batching protocol. Candidate archive ed85dc914e846f0586c01e71a1427dd64caba0a80844b02e28feedfcb0dcce88 separately passed its exact-container 20-case H100 test. This report does not mean all 4,000 cases ran through that Docker archive, does not establish best across all historical alternatives, and does not predict an official finals score.
28
-
29
- Strict-excluded pairs 0004, 0130 and 0196 remain distinct from recovered pairs 0074, 0111 and 0123. All three historical failed attempts retain their failed status. The fixed recovery policy selected whole pairs without using scores. No first-attempt reliability or model-only causal claim is made.
30
-
31
- The paired report SHA256 is `a73eedbcbdfeac8f8500d4730980984518209b035f9571d6c10293390ae7035b`. Reference accepted observation is `c613cfe00d4576f671cdcb587be378e2bf55e476c057f13bb741d9f604363cf1`; candidate observation is `2b5f77c9f4a913cecaae405401ac05c15d045ba896d204ed8ff0d5759450d3d4`. The monorepo custody record `packages/rehearsal-segment/review/preparation/closed3-full4000-comparison-actual-r1/CUSTODY.json` retains both actual returns, complete readbacks, commands, reporting inputs and machine-readable report.
32
-
33
-
34
- This private model-card follow-up was prepared on September 10, 2026. The historical local runtime evidence in the following paragraphs applies only to archive `ed85dc914e846f0586c01e71a1427dd64caba0a80844b02e28feedfcb0dcce88`, Docker config `d8fcfca448b374aab5d37c23382a02ef24367d7d66f6c5fbe194b65eb07f308b`, and the unchanged immutable asset/metadata revisions in LOAD.md. The exact20 Docker process took 282.298466seconds, with sampled GPU memory maximum 43,589 MiB, no OOM termination and positive first-frame diagnostics for all 20 clips.
35
 
36
- The untouched exact archive now also passed a captured 500-case, default-entrypoint H100 Docker test. The guard completed all 500 original answer calls with healthy model initialization, no recorded engine failures and zero guard input deferrals. Every answer content matched the **candidate's evaluated serving implementation under B20** for these same 500 cases. This is not 100% semantic accuracy, not a content comparison against the historical pre-evaluation reference, and not evidence isolating a causal batch-size effect or establishing general order invariance. The 500 cases comprise 216 clips from original source0020 and 284 from 0021; prepared request.videoID values are clip aliases and are not independent source videos.
37
 
38
- The actual Docker process ran from 12:49:28.234458 to 13:11:09.278933UTC on September 10, taking 1,301.044475seconds. The independent attach interval was 1,301.398154seconds. This attempt enforced 4 CPU cores and 32 GiB memory on one H100, with a 7,200-second admitted runtime cap; the registered500 profile's 9,144-second value remains separately identified. The answer-loop log reports 1,014.1seconds and the answer-latency sum 1,013.126716seconds; these are self-reported, not independent per-answer measurements. The host observer stopwatch excludes initial registration/input hashing and prior checks.
39
 
40
- Across 584 retained resource samples, GPU memory reached 42,429 MiB and main-process RSS 6,453,096,448bytes. Cgroup memory.peak reached 34,359,738,368bytes (32GiB), including cache/tmpfs. Cgroup accounting is distinct from process RSS and VRAM; sampled maxima are not continuous peaks. Docker reported no OOM kill and cgroup OOM/OOM-kill event counts were zero. Nonzero memory-limit max events alone do not establish OOM.
41
 
42
- Startup first-frame diagnostics were positive for 150 clips. The 60-second diagnostic budget expired, leaving 350 startup probes unmeasured; successful later answer processing does not backfill those first-frame proofs. These skips are distinct from the guard's observed near-still input deferrals, whose literal list was empty. Inputs were 500 loose clips plus request.json, with no serving-input ZIP. The separately accepted ZIP20 test below supplies narrow ZIP extraction evidence; organizer-platform and 500-case ZIP-capacity qualification remain separate.
43
 
44
- The prior500 launcher stopped before START at its scheduling gate and remains a failed attempt, as do the earlier preserved loader/controller failures. The successful separately admitted run does not retrospectively turn them into passes or establish first-attempt reliability. The complete 500 readback passed at 13:11:29UTC. Acceptance SHA256 is `c8410ed845b0875189ae55eb7ae8a1571f81c876e36c53c614b9f44c1b6751b6`; original summary SHA256 is `47692048634f47758654f3b0bf24119ee041429b37955829f96a36b5c5f309d2`; the source-label-corrected summary is `2d2f76b0702028ef52a738c1bbda5b72db0bb185edbc27c5832799df4f4a7fb2`. The correction distinguishes clip aliases from the two original source videos and leaves all health/content/timing/resource results unchanged.
45
 
46
- The same unchanged image subsequently passed a separately admitted ZIP20 test with the prior loose20 clips and identical question order. The serving input contained one ZIP archive and no loose media files; the guard resolved all 20 contained media members and completed all 20 answer calls. Every answer content matched the prior same-image loose20 output. The ZIP_STORED archive was 754,197,783 bytes, SHA256 `9bc6931274c95ab84358fa86b24adf6ffcccc6bb0323e0b513ae94edd5c69436`, with 754,195,577 uncompressed media bytes; this small archive fit the harness's 16 GiB tmpfs. It does not establish capacity for the roughly 21 GB captured500 corpus as a ZIP or identify organizer scratch storage.
47
 
48
- ZIP20's actual Docker process took 148.082458 seconds and its attach interval 149.756455 seconds, within the separate 900-second admitted cap. There were 67 retained runtime/resource observations, zero guard input deferrals, no OOM termination and completed scoped cleanup. These process durations are observations under different run/cache conditions, not a causal speedup measurement. There was no independent post-extraction file-hash census. The complete metadata readback finished at 13:15:25 UTC; receipt SHA256 is `53532b03e69bac7cecd9bb98051e164400c766a7cf3045658fe2e9f9fa0e6058`. The result is local same-image ZIP20 extraction/execution evidence, not organizer-platform qualification.
49
 
50
- The private repository does not imply final submission sign-off. The earlier ancestor evaluation follow-up changed only README.md, EVALUATION.md and EVALUATION.json; all 50 assets, immutable LOAD revision pins, training/exposure disclosures and existing 4,000/3,940-case statistics remain unchanged. Raw videos, gold answers, per-case outputs and judge data are not added by this presentation update.
51
 
 
52
 
53
- ## Historical 43d4 local evidence, followed by failed canonical platform tryout
54
 
55
- **Historical 43d4 local evidence.** The historical 43d4 archive is `43d4fc4fd7a39e4884a5179176a6b6a436956a146494876783d6287855a82825` (35,613,238,247 bytes, generation `1789057108435274`). It has configuration `f7e22f1822d7a152ec91999dbb639ec63378346ef3f5144723c5b453748f97b9`. The resident build manifest is `4eb060d790e39af7b56dcc263b73a0a039ded5cd4dd898c0bda33e6e13c5bf63`. The CPU worker's classic Docker store loaded the exact configuration ID `sha256:f7e22f1822d7a152ec91999dbb639ec63378346ef3f5144723c5b453748f97b9` and all 24 ordered DiffIDs; it did not expose an imported manifest digest. That field remains null. [IMAGE_IDENTITY.json](IMAGE_IDENTITY.json) records their full joins. The original ed85 platform failure remains preserved; it is not a failure of the historically tested derivative, which later failed the canonical interval admission check.
56
 
57
- The complete replacement-image acceptance combines 6 invocations and 600 answer executions across two retained attempts, not 600 unique questions. The first default 20-answer invocation completed with healthy guards and matching contents; its enclosing attempt then failed when a case-sensitive absence parser rejected Docker's lowercase no-such-object message after successful removal. That raw failure and watchdog return code 125 remain failed. A separately claimed continuation used one persistent container for linked loose 20-input, real-root loose 20-input, linked ZIP 20-input/FO, repeated ZIP 20-input/FO, and linked loose 500-input. All accepted guards passed with empty input-deferral lists. Repeated FO contents matched; the FO fixture was reconstructed from the image package and is not an independent capture of organizer socket bytes. Non-FO answers matched the frozen candidate baseline; these are content comparisons, not new semantic grading. The 500-input case's independent host child interval was 1140.333899 seconds; startup first-frame probes were positive for 151 clips and did not cover the other 349. Successful answers do not fill those missing startup proofs. The combined independent host child intervals total 1718.896785 seconds, including the first attempt's 115.555531 seconds. Self-reported answer latencies, attached child intervals and sampled GPU/cgroup measurements remain separate in the machine-readable metadata. Cgroup memory.peak is cumulative within the persistent container, rather than an isolated per-invocation maximum. Library tokenizer/processor/dtype warnings were retained; no behavior change or fallback is inferred from those warnings. Local persistent tests do not attest the deployed shim version, organizer execution or full 500-input ZIP capacity.
58
 
59
- The earlier nine-file image reconciliation changed five existing metadata files and added four files, leaving 77 prior files unchanged at 86 total. All 50 model/prior assets and component loaders remain byte-identical; selected training/annotation facts, original 4,000/3,940 statistics, 5b regression and inherited/adaptive exposure limits are preserved. Original MODEL_FILES.json and SERVING_ASSETS.json retain ancestor image metadata because the immutable loader pins their raw bytes; current-image scope is explicit in the new identity/compatibility records and additive LINEAGE fields. No final form or organizer declaration is submitted.
60
 
 
61
 
62
- The subsequent LOAD.md presentation-only update identifies the fixed serving archive through the current IMAGE_IDENTITY.json and PLATFORM_COMPATIBILITY records, and excludes three historical presentation files from the immutable metadata download. Asset revision A, metadata/tools revision M, all 50 asset bytes, and the verification/component-loading commands remain unchanged.
 
1
+ # Evaluation
2
 
3
+ This page separates local answer-quality measurements, container runtime checks and the team-confirmed finals submission. The underlying records remain in [EVALUATION.json](EVALUATION.json), [PLATFORM_COMPATIBILITY.json](PLATFORM_COMPATIBILITY.json) and [SUBMISSION_STATUS.json](SUBMISSION_STATUS.json).
4
 
5
+ ## Development comparison
6
 
7
+ Two registered serving implementations were evaluated with the same fixed Qwen3.5-27B judge and revision, hardware envelope and 20-question batching protocol. The board contains 4,000 questions from ten original source videos, with 400 questions per video. Analysis clusters by original video, rather than qID-derived clip filenames.
8
 
9
+ The reference is a registered implementation derived from the historical 333 framework. Its agreement with the actual 333 Docker image's answers has not been established. The candidate measurement likewise does not mean that all 4,000 questions ran through the final Docker archive.
 
 
 
 
10
 
11
  | Population | Reference correct | Candidate correct | Reference micro accuracy | Candidate micro accuracy | Difference |
12
  | --- | ---: | ---: | ---: | ---: | ---: |
13
+ | Full development board | 1,940 / 4,000 | 2,373 / 4,000 | 48.500% | 59.325% | +10.825 pp |
14
  | Strict paired sensitivity | 1,906 / 3,940 | 2,334 / 3,940 | 48.376% | 59.239% | +10.863 pp |
15
+ | Excluded diagnostic subset | 34 / 60 | 39 / 60 | 56.667% | 65.000% | +8.333 pp |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
16
 
17
+ On the full board, equal-video weighting gives the same difference as micro accuracy. A paired two-level bootstrap gives a descriptive 95% interval of **+9.049 to +12.501 percentage points** (1,000 draws; seed 42). All ten source-video differences are positive, ranging from +7.5 to +13.25 points. There are 650 paired wins, 217 losses and 3,133 ties. The strict sensitivity's equal-video difference is +10.877 points, with interval +9.209 to +12.588.
18
 
19
+ The strict analysis excludes pairs `0004`, `0130` and `0196`. Separately recovered pairs `0074`, `0111` and `0123` were accepted as whole pairs under a policy that did not use scores. Historical failed attempts remain failures; recovery does not establish first-attempt reliability. Both accepted result sets have complete score rows and no parsing failures.
20
 
21
+ The grader runs completed on September 10, 2026 at 10:53:44 UTC for the candidate and 10:55:13 UTC for the reference. The paired report completed at 11:07:14 UTC; its SHA-256 is `a73eedbcbdfeac8f8500d4730980984518209b035f9571d6c10293390ae7035b`. [EVALUATION.json](EVALUATION.json) retains the observation identities and report binding.
22
 
23
+ ## Limits of the comparison
24
 
25
+ The board was repeatedly used for model selection and has inherited training and data exposure. It is not an independent holdout, and it has not been established as the platform's out-of-distribution population. The confidence intervals are conditional on this board and are unadjusted for repeated selection or multiple capability comparisons.
26
 
27
+ **Causal-consequence reasoning is a measured weakness.** Capability 5b regressed on 62 cases from six source videos: −17.742 points in micro accuracy and −16.881 in the equal-video estimate, with a descriptive interval of −34.671 to −1.957. Capability 5a ties in micro accuracy on 34 cases and is slightly negative under equal-video weighting. These fine-capability results are separate from official ID/OOD buckets and finals ranking votes.
28
 
29
+ The local scorer passed 14 convenience controls: six positive and eight negative. Those controls test response and parser boundaries. They do not measure agreement with the organizer's judge, false-positive rates on real questions, or official score calibration.
30
 
31
+ No matched comparison in the candidates' actual serving forms establishes that R4 exceeds N3C-v2 or the original 333 image. Historically, the platform recorded 0.5207852833 for 333 and 0.5199292864 for N3C-v2 under different batching. Those scores do not supply an R4 forecast. No official finals score, significance-adjusted bucket ranking or Copeland outcome is inferred from the local results.
32
 
33
+ ## Runtime qualification
34
 
35
+ The R4 qualification matrix completed eight test cases, nine application invocations and 1,593 answer executions across 533 unique qIDs. It covered ordinary and reversed 20-question inputs, 500 loose clips, the same 500 clips in ZIP archives at 4 GiB and 16 GiB scratch limits, default and repeated canonical calls, and controlled per-question recovery.
36
 
37
+ The ordinary 20-question and loose 500-question outputs matched their saved references. Reversed inputs matched the same image's ordinary outputs; both ZIP runs matched its loose 500-question outputs. The canonical fixture retained three documented pre-generation deferrals. A 23-question recovery fixture produced exactly its three intended empty answers and preserved the 20 healthy outputs. These counts therefore include repeated questions and reduced neural coverage; they are not semantic accuracy measurements.
38
 
39
+ The 500-question runs took approximately 1,226.6–1,267.1 seconds by independent runtime measurement. Self-reported answer latencies, sampled resource maxima and independent elapsed time remain distinct in the records. These fixtures do not establish every input size, cold-cache condition, general order invariance, platform scratch allocation or continuous resource peak. See [platform compatibility](PLATFORM_COMPATIBILITY.md) for recovery behavior and the remaining failure boundaries.
40
 
41
+ Earlier archive tests and failures remain preserved in [EVALUATION.json](EVALUATION.json) and the [immutable pre-refinement report](https://huggingface.co/HeyDonto/SURGFIELD-ORena-2026-SEGMENT-FINAL/blob/24133f8d4a491845db2a19c84c2d182f8b4e4831/EVALUATION.md). Their image identities and test conditions must be retained when interpreting them. In particular, reproducing R3's ZIP extraction failure locally does not prove the cause of the earlier failed platform submission.
42
 
43
+ ## Finals status
44
 
45
+ HeyDonto Labs confirmed successful submission of the R4 archive, with no errors reported, on September 14, 2026. That is the date of confirmation, rather than a recorded submission timestamp. The platform identifiers, terminal logs, actual resource settings, final score and ranking have not been independently captured here. [SUBMISSION_STATUS.json](SUBMISSION_STATUS.json) records the evidence and exact artifact identity.
LINEAGE.md CHANGED
@@ -1,30 +1,49 @@
1
- # Exact selected lineage
2
 
3
- HeyDonto Labs; responsible contact Reza Nehzati, Ph.D. ([profile](https://huggingface.co/rezanehzati)).
4
 
5
- The private N3 backbone is the uniform FP32 mean of S1001, S1003 and S1004 at step17535, serialized in BF16. Its combined named-shard digest is `a78c72786e43b779ab886d583d9a7da3fb6ed07e026b681cc8a7270b216acf8e`. Its 18 files are retained byte-for-byte, including its private config and processor metadata.
6
 
7
- The selected main continuation is **original-n3-main-lr2e5-step000900**: learning rate 2e-5, seed908, gradient accumulation8. The selected prefix contains 7,200 draws and 4,130 unique rows from the eligible 7,207-row continuation training population. Its adapter is `0240d87b…`, rank16/alpha32, 15,335,424 parameter elements. Serving retains one FP32 default LoRA on the BF16 N3 backbone, unmerged. Later producer steps and exposures are not attributed to this checkpoint.
8
 
9
- The selected temporal continuation is **original-temporal-lr5e6-step000600**: learning rate 5e-6, seed908, gradient accumulation8. The selected prefix contains 4,800 draws and 2,912 unique rows from the eligible 4,680-row population. The `87fc446a…` checkpoint contains the complete updated adapter initialized from a47; a47 must not be applied as a second adapter. The serving configuration is the exact normalized `0ff6dad4…` file, while producer config `d621c666…` remains separately identified. The frozen temporal loader applies it once to public Qwen revision0c351, then merges at initialization.
10
 
11
- The inherited N3 source corpus `4b1ee90b…` contains 46,758 rows: 33,581 organizer rows, 4,621 paraphrases, 5,217 count rows, 2,317 convention rows and 1,022 adopted human-review rows. The continuation corpus `66362830…` is separate. Official-only new continuation does not erase inherited clinical/development exposure. Historical temporal smoke/full populations and the owner-reported human-labeler/eligibility interpretation remain separate disclosures; no new organizer eligibility decision is asserted.
12
 
13
- The private annotation packet `ee69dafa…` retains the known 13,177 derived rows across four JSONL files. It excludes all 33,581 organizer rows and video pixels; row locators and review statements do not establish complete signature custody or redistribution permission. `ANNOTATION_SCOPE.json` lists exact counts and hashes. Its intended private organizer venue is a preserved owner decision; actual access and delivery receipts must be recorded independently.
14
 
15
- The six inherited prior files remain byte-identical. `cholec_priors_v38.json` arose from zero-shot OWLv2 detections on 613 frames from 10 CholecTrack20 videos; it is not a new optimizer-training corpus. The remaining duration, quadrant, stem and mode tables retain their organizer-training provenance. No v40 prior or obsolete neural checkpoint is included.
16
 
17
- The original forward candidate changed the frame sampler (`cd146570…`). The later R3 compatibility derivative additionally fixes image-owned SEGMENT routing in cap_fmt_inferrer.py and records exactly attributed pre-generation deferrals in its bootstrap; the original timestamps, sampler/model behavior and all other 51 serving files remain unchanged. Public Qwen, private N3, both selected adapters, all six priors, prompts, precision, frame limits and generation settings are unchanged. `TRAINING_RESOURCES.json` preserves selected-save timing/peak scopes and producer identities; `MODEL_IDENTITY.json` preserves every selected file identity.
18
 
19
- The known derived packet contains all 13,177 listed extensions: 4,621 paraphrases, 5,217 count remints, 2,317 convention remints and 1,022 adoption rows. Exclusion of the 33,581 organizer rows is intentional, not evidence that those rows must be republished. The unresolved items are the separate private annotation URL/access/delivery receipts and complete signature provenance.
20
 
 
21
 
22
- ## R4 serving compatibility update
23
 
24
- **Finals submission: owner confirmed on September 14, 2026 that repaired R4 was successfully submitted for SEGMENT finals, with no errors reported.** The submitted archive SHA256 is `e477d307ff2e09f469d83569ec1ce8703daeef34edbba633a3f240eab955de00`. September 14 is the confirmation date; the actual submission time, platform UUIDs, terminal logs, score and ranking were not supplied. See [SUBMISSION_STATUS.json](SUBMISSION_STATUS.json) for the exact identity and evidence scope. The model weights, source files and immutable loading revisions are unchanged.
 
 
 
 
 
 
25
 
26
- **September 11 local qualification record: R4 runtime, archive and import checks completed.** The exact archive is `e477d307ff2e09f469d83569ec1ce8703daeef34edbba633a3f240eab955de00` (35,613,293,722 bytes), immutable generation `1789092734891901`. The build manifest is `ba9ddb1e78a45040803ae6f0e68c66de7a538f5e3e86397777e5fc2c488106b0` and configuration `3f43fce1081dc5db454447dbe4678947e5cefec014635e611d8d1032a0ac71c4`. Full compressed/raw archive verification covers all 26 ordered DiffIDs; actual same-generation classic Docker import reports config `3f43fce1081dc5db454447dbe4678947e5cefec014635e611d8d1032a0ac71c4` and no imported manifest digest. The agreed eight cases completed 9 application invocations and 1593 answer executions across 533 unique qIDs. Healthy20 and loose/ZIP500 matched saved content; reverse20 matched the same-image ordinary20, and both ZIP quotas matched the same-image loose500. Canonical repeated calls retained the exact disclosed deferrals. The mixed recovery case blanked only its declared failed questions and preserved later healthy outputs. These are bounded runtime/interface/content checks; no new semantic grade or official final result is claimed.
27
 
28
- R4 retains the image-owned literal SEGMENT task and the original request clocks. It stages only the selected ZIP member for each startup probe or answer, verifies that member through CRC/EOF and SHA-256, closes decoder handles, and removes the lease before continuing. Loose input remains a direct read. Model initialization is still eager. An attributed ordinary per-question exception produces empty content for that qID with measured latency and failure telemetry; successful answers are preserved and later questions continue. Unsafe or ambiguous archives, source mutation, cleanup integrity, required model initialization, environment and fixed-track failures remain fatal. The inherited precise pre-generation rules deferrals remain separately disclosed as degraded coverage. The four changed image files are the bootstrap, contract, probe entrypoint and new ZIP helper; all 51 unchanged inherited serving source files, including cap_fmt, all 50 model/prior assets, prompts, model arithmetic and generation settings remain unchanged.
29
 
30
- The full retained ZIP500 was qualified locally at the stated 4 GiB and 16 GiB tmpfs profiles, alongside loose500, healthy20, reversed20, default and persistent canonical calls, and controlled recovery. These fixtures do not establish the platform scratch quota, all possible input sizes, continuous resource peaks, cold-cache requirements, general order invariance, new semantic scores, independent-holdout performance or official platform success. Canonical deferrals and deliberately failed recovery rows are reduced scientific coverage. The prior R3 4 GiB eager-ZIP failure is a conditional local capacity reproduction, not proof of the organizer failure cause. Earlier images and failed attempts remain historical.
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Model and data lineage
2
 
3
+ The submitted R4 application combines two distinct Qwen-based model paths, an OWLv2 detector and six inherited prior tables. [LINEAGE.json](LINEAGE.json), [MODEL_IDENTITY.json](MODEL_IDENTITY.json) and [MODEL_FILES.json](MODEL_FILES.json) preserve the exact component identities. [IMAGE_IDENTITY.json](IMAGE_IDENTITY.json) identifies the current serving archive.
4
 
5
+ ## Main model
6
 
7
+ The private N3 backbone is the uniform FP32 mean of checkpoints S1001, S1003 and S1004 at step 17,535, serialized in BF16. Its combined named-shard digest is `a78c72786e43b779ab886d583d9a7da3fb6ed07e026b681cc8a7270b216acf8e`. Its 18 files, including private configuration and processor metadata, are retained byte for byte.
8
 
9
+ The selected continuation is `original-n3-main-lr2e5-step000900`: learning rate 2e-5, seed 908 and gradient accumulation of eight. The selected training prefix contains 7,200 draws and 4,130 unique rows from an eligible population of 7,207 rows. Later training steps and their exposure are not attributed to this checkpoint.
10
 
11
+ Its LoRA adapter has rank 16, alpha 32 and 15,335,424 parameter elements. Serving attaches one FP32 default adapter to the BF16 N3 backbone and leaves it unmerged. Adapter SHA-256: `0240d87b54503bd32502f5c77f45a60dbc404f69fa51e04b29ed6f84906697b5`. See [main selected exposure](main-SELECTED_EXPOSURE.json).
12
 
13
+ ## Temporal model
14
 
15
+ The temporal path uses the public Qwen3-VL-8B-Instruct backbone at revision `0c351dd01ed87e9c1b53cbc748cba10e6187ff3b`. This is separate from the private N3 main backbone.
16
 
17
+ The selected continuation is `original-temporal-lr5e6-step000600`: learning rate 5e-6, seed 908 and gradient accumulation of eight. Its selected prefix contains 4,800 draws and 2,912 unique rows from 4,680 eligible rows.
18
 
19
+ The checkpoint contains the complete updated adapter initialized from the inherited temporal LoRA adapter, identified as a47. It is applied once to the public BF16 backbone and merged once during initialization; a47 is not applied as an additional adapter. Adapter SHA-256: `87fc446ac0d8cf6a3d6147f3d959ec21bdee1c806e00dc74f5a1ffe9e6ac7878`. The normalized serving configuration and original producer configuration remain separately identified in [LINEAGE.json](LINEAGE.json). See [temporal selected exposure](temporal-SELECTED_EXPOSURE.json).
20
 
21
+ ## Training data and derived annotations
22
 
23
+ The inherited N3 corpus contains 46,758 rows:
24
 
25
+ | Source | Rows |
26
+ | --- | ---: |
27
+ | Organizer data | 33,581 |
28
+ | Paraphrases | 4,621 |
29
+ | Count annotations | 5,217 |
30
+ | Convention annotations | 2,317 |
31
+ | Adopted human-review annotations | 1,022 |
32
 
33
+ The continuation corpus is separately identified. Restricting the new continuation to organizer data does not remove inherited clinical or development-board exposure. Historical temporal smoke and full-training populations also remain separate. The team's recorded human-labeler and eligibility interpretation is not an organizer eligibility ruling. [TRAINING_RESOURCES.json](TRAINING_RESOURCES.json) retains producer identities and selected-checkpoint resource records.
34
 
35
+ The separately retained annotation packet contains the 13,177 derived rows across four JSONL files. It intentionally excludes all 33,581 organizer rows and underlying video pixels. It is not the complete ancestor corpus or a complete signature ledger. [ANNOTATION_SCOPE.json](ANNOTATION_SCOPE.json) supplies counts, hashes and provenance limits. A separate private dataset URL, applicable recipient conditions and delivery receipt remain to be recorded; row locators and review statements alone do not establish complete signature custody or redistribution permission.
36
 
37
+ ## Detector and priors
38
+
39
+ OWLv2 uses the pinned `google/owlv2-large-patch14-ensemble` revision `95e26936e865f87db1742128404b3c035d47d89d`.
40
+
41
+ The six prior files retain their original bytes. `cholec_priors_v38.json` was derived from zero-shot OWLv2 detections on 613 frames from ten CholecTrack20 videos. This is prior construction, rather than a new optimizer-training corpus. The remaining duration, quadrant, stem and mode tables retain their organizer-training provenance. The release excludes the v40 prior and obsolete neural checkpoints.
42
+
43
+ ## Serving revisions
44
+
45
+ The forward-decoding update changed frame access while preserving the tested sample indices, RGB bytes, source timestamps, frame limits and deadline. R3 corrected fixed SEGMENT routing and recorded pre-generation model deferrals. R4 then changed four wrapper files to limit ZIP staging and contain recoverable per-question errors.
46
+
47
+ R4 retains the 51 inherited serving source files and all 50 model/prior assets, including the selected adapters, prompts, model arithmetic and generation settings. These compatibility changes introduced no gradient updates. Details are in [PLATFORM_COMPATIBILITY.md](PLATFORM_COMPATIBILITY.md).
48
+
49
+ Upstream model notices, dataset restrictions and derived-resource terms are documented in [NOTICES.md](NOTICES.md). This repository does not grant blanket redistribution rights over all training data or derived assets.
LOAD.md CHANGED
@@ -1,17 +1,17 @@
1
- # Local loading from two immutable revisions
2
 
3
- These revisions reproduce the unchanged **model components and loading tools**. For full serving, the locally qualified R4 archive has SHA-256 `e477d307ff2e09f469d83569ec1ce8703daeef34edbba633a3f240eab955de00`. Its exact GCS generation, configuration, four compatibility-source identities and bounded local runtime evidence are recorded in [IMAGE_IDENTITY.json](IMAGE_IDENTITY.json) and [PLATFORM_COMPATIBILITY.md](PLATFORM_COMPATIBILITY.md) at the repository revision containing this page. Prior archives remain historical. Use those current records to select the serving archive.
4
 
5
- The metadata download below remains pinned to revision M and excludes its archived `README.md`, `LOAD.md` and `IMAGE_IDENTITY.json`. The retained `MODEL_FILES.json` and lineage/binding documents describe component ancestry and may contain the ed85 ancestor's image identity; those bindings do not select the current platform artifact. The 50 model/prior files, immutable A/M revisions, manifests and component loaders remain unchanged. The fixed wrapper is already inside the current Docker archive and is identified by the current records linked above.
 
 
 
6
 
7
- The model assets and the loading tools have separate immutable commits:
8
 
9
- - Asset revision **`35fc28af1d55a5a9ae8096ff2e2cd2c961aa7119`** contains the50 exact model/prior files.
10
- - Metadata revision **`bc9e092422beac9ae51005694824cdc70205e0d5`** contains `MODEL_FILES.json`, verifiers, component-loading code, notices and source copies. Its parent is the asset revision, but only the small metadata files are downloaded from it below.
11
 
12
- The upload coordinator fills both IDs in the final handoff instructions after each commit and readback. Neither commit embeds its own ID: a later documentation-only presentation commit or the separately supplied `REVISION_PAIR.json` identifies this pair. Unsubstituted placeholders are not valid revisions; do not substitute a floating branch.
13
-
14
- Use an authenticated account with read access to the private repository. Download assets and metadata into **different local directories**:
15
 
16
  ```sh
17
  python3 - <<'PYLOAD'
@@ -34,37 +34,47 @@ snapshot_download(repo_id=repo, revision=metadata_revision,
34
  PYLOAD
35
  ```
36
 
37
- The first directory intentionally has no metadata files. The verification commands below take the asset root and the metadata manifest explicitly. The second download's allowlist excludes every model/prior asset, so it does not fetch the large weights a second time. Its archived LOAD.md may retain a metadata-revision placeholder because it was written before its own commit; use the completed pair supplied in these final instructions. Full local EOF checks, not a model-card or LFS pointer alone, establish downloaded payload identity.
38
 
39
- | Repository path | Frozen image path | Purpose |
40
- |---|---|---|
41
- | `main_model/` | `/app/artifacts/segment_final/main_model/` | Private N3 BF16 backbone and its saved processor metadata |
42
- | `main_adapter/` | `/app/artifacts/segment_final/main_adapter/` | Selected step900 FP32 default LoRA; unmerged |
43
- | `temporal_adapter/` | `/app/artifacts/segment_final/temporal_adapter/` | Complete updated low600 adapter; applied once |
44
- | `upstream/qwen3-vl-8b-instruct/` | `/app/.cache/huggingface/hub/models--Qwen--Qwen3-VL-8B-Instruct/snapshots/0c351dd01ed87e9c1b53cbc748cba10e6187ff3b/` | Main processor plus temporal public backbone/processor |
45
- | `upstream/owlv2-large-patch14-ensemble/` | `/app/.cache/huggingface/hub/models--google--owlv2-large-patch14-ensemble/snapshots/95e26936e865f87db1742128404b3c035d47d89d/` | OWLv2 detector and processor |
46
- | `priors/` | `/app/artifacts/` | Exact six inherited prior files |
47
-
48
- The public cache's `refs/main` resolves to its pinned revision for OWLv2's unchanged model-ID call. Snapshot files may be materialized as regular files; the model release paths do not require distributing the archive's internal symlinks. Full default-entrypoint reproduction also needs the exact serving source/bootstrap and frozen dependencies, which are already present in the corresponding Docker image. Do not treat a generic HF pipeline as that serving application.
49
-
50
- The default Docker entrypoint is `/opt/conda/bin/python -I -S /opt/orena/segment/bootstrap.py`, user `app`, working directory `/app`, with offline HF/model settings. Preserve it. See the current [image identity](IMAGE_IDENTITY.json) and [platform compatibility record](PLATFORM_COMPATIBILITY.md) for the exact serving archive and its completed or pending validation; the historical metadata snapshot does not report current platform status.
51
-
52
- The frozen main `load_main()` verifies the exact N3 and adapter files before and after loading, compares private/public tokenizer vocabularies and chat templates, loads BF16 N3 with SDPA on CUDA, then attaches exactly one selected default adapter. It does not merge the main adapter or cast its FP32 parameters. The frozen temporal loader verifies its exact public base path in the normalized adapter configuration, loads public Qwen BF16/SDPA, applies low600 once, and calls `merge_and_unload()`. The detector uses FP32 OWLv2 and its existing prompts. `model.safetensors` is selected ahead of the inactive `.bin` duplicate by the pinned Transformers resolver.
53
 
54
- The selected source pins Torch2.6.0/CUDA12.4, torchvision0.21.0, Transformers4.57.6, PEFT0.19.1 and the remaining versions in `requirements-frozen.txt`. Use the exact image environment for serving equivalence; this list is a source declaration, not a new host installation instruction or a claim that an arbitrary environment reproduces answers. Main rendering uses up to16 frames and temporal rendering up to128 native frames, max side768; generation is greedy with max64 new tokens. Runtime model loading is eager in the current candidate. No lazy-init change is part of this release.
55
-
56
- For local asset verification only (standard library, no model execution):
57
 
58
  ```sh
59
  python3 -I -B ./segment-final-metadata/verify_assets.py --root ./segment-final-assets --manifest ./segment-final-metadata/MODEL_FILES.json
60
  ```
61
 
62
- This reads all 50 assets to EOF and validates SHA256 and size. It creates no model, downloads no file and invokes no GPU. HF uploader access, organizer read access and final commit/readback receipts are separate completion gates.
 
 
63
 
64
- For an explicit component-loading check in that frozen tensor environment:
65
 
66
  ```sh
67
  python3 -I -B ./segment-final-metadata/load_components.py --root ./segment-final-assets --manifest ./segment-final-metadata/MODEL_FILES.json --component all
68
  ```
69
 
70
- The example first verifies all 50 model/prior files. It then imports the exact `7a4df645…` main-loader source, relocates only its two local path constants and retains all its original before/after byte, parameter, adapter and processor checks. The temporal operations use the frozen source sequence with local paths: BF16 public Qwen/SDPA, one low600 `PeftModel.from_pretrained`, then one `merge_and_unload`, followed by eval. OWLv2 uses the frozen FP32 model/processor sequence. The exact three component-loading source files are preserved under `component_source/`. This example performs actual GPU model loading when invoked; it has only been source/CPU validated in this preparation. It generates no answers and does not adopt evaluation flags or replace the complete serving bootstrap/router/sampler.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Download and loading
2
 
3
+ The model components and loading tools are available at two immutable revisions. Access requires an authenticated Hugging Face account with permission to read this private repository.
4
 
5
+ | Contents | Revision |
6
+ | --- | --- |
7
+ | 50 model and prior assets | `35fc28af1d55a5a9ae8096ff2e2cd2c961aa7119` |
8
+ | Asset manifest, verification and component-loading tools | `bc9e092422beac9ae51005694824cdc70205e0d5` |
9
 
10
+ These revisions reproduce the component files and loading tools. The current R4 serving archive is selected separately by [IMAGE_IDENTITY.json](IMAGE_IDENTITY.json); its archive SHA-256 is `e477d307ff2e09f469d83569ec1ce8703daeef34edbba633a3f240eab955de00`. See [platform compatibility](PLATFORM_COMPATIBILITY.md) for full-application execution and [submission status](SUBMISSION_STATUS.json) for the team's confirmation.
11
 
12
+ ## 1. Download
 
13
 
14
+ Use an environment with `huggingface_hub` installed and authentication already configured. Keep model assets separate from the loader and verification metadata:
 
 
15
 
16
  ```sh
17
  python3 - <<'PYLOAD'
 
34
  PYLOAD
35
  ```
36
 
37
+ The metadata allowlist avoids downloading the large weights a second time. It also excludes that historical revision's `README.md`, `LOAD.md` and `IMAGE_IDENTITY.json`. Its retained manifests and lineage documents describe component ancestry and may identify the earlier ed85 image; those fields do not select the current serving archive. Use the current image record linked above.
38
 
39
+ ## 2. Verify all files
 
 
 
 
 
 
 
 
 
 
 
 
 
40
 
41
+ On Linux/POSIX, run the standard-library verifier before loading:
 
 
42
 
43
  ```sh
44
  python3 -I -B ./segment-final-metadata/verify_assets.py --root ./segment-final-assets --manifest ./segment-final-metadata/MODEL_FILES.json
45
  ```
46
 
47
+ The verifier reads all 50 assets to EOF and checks their SHA-256 hashes and sizes against the immutable manifest: 36,966,713,444 bytes in total. It uses POSIX filesystem interfaces, downloads no files and invokes no GPU. `MODEL_FILES.json` and `SERVING_ASSETS.json` preserve the same immutable inventory; the tokenizer, processor and configuration files are included among the assets.
48
+
49
+ ## 3. Optional component-loading check
50
 
51
+ In the frozen tensor environment, the following command verifies the assets and loads the main, temporal and OWLv2 components on CUDA:
52
 
53
  ```sh
54
  python3 -I -B ./segment-final-metadata/load_components.py --root ./segment-final-assets --manifest ./segment-final-metadata/MODEL_FILES.json --component all
55
  ```
56
 
57
+ The standalone helper generates no answers. It has been reviewed and checked at source/CPU level; this helper has not been independently qualified as a complete GPU serving application. The container's runtime qualification is separate. Individual component choices are `main`, `temporal` and `owlv2`, in addition to `all`.
58
+
59
+ The main loader checks private N3 and adapter bytes before and after loading, compares private/public tokenizer vocabularies and chat templates, loads BF16 N3 with SDPA and attaches exactly one FP32 default adapter without merging it. The helper relocates only the loader's local asset paths while retaining those checks.
60
+
61
+ The temporal loader uses public Qwen BF16/SDPA, applies the complete step-600 temporal adapter (`low600` in the serving code) once and calls `merge_and_unload()` once, followed by evaluation mode. OWLv2 uses its frozen FP32 model and processor sequence. The source copies are retained under `component_source/`.
62
+
63
+ ## Full serving application
64
+
65
+ The submitted Docker image includes the complete bootstrap, router, sampler and frozen dependencies. Preserve its default entrypoint `/opt/conda/bin/python -I -S /opt/orena/segment/bootstrap.py`, user `app`, working directory `/app` and offline model settings. A generic Hugging Face pipeline or the component helper does not reproduce the full application.
66
+
67
+ The source environment pins Torch 2.6.0 with CUDA 12.4, torchvision 0.21.0, Transformers 4.57.6 and PEFT 0.19.1; remaining versions are in [requirements-frozen.txt](requirements-frozen.txt). Use the exact image environment when comparing serving behavior. This dependency record does not establish answer equivalence in an arbitrary host environment.
68
+
69
+ | Repository path | Frozen image path | Purpose |
70
+ |---|---|---|
71
+ | `main_model/` | `/app/artifacts/segment_final/main_model/` | Private N3 BF16 backbone and its saved processor metadata |
72
+ | `main_adapter/` | `/app/artifacts/segment_final/main_adapter/` | Selected step900 FP32 default LoRA; unmerged |
73
+ | `temporal_adapter/` | `/app/artifacts/segment_final/temporal_adapter/` | Complete updated low600 adapter; applied once |
74
+ | `upstream/qwen3-vl-8b-instruct/` | `/app/.cache/huggingface/hub/models--Qwen--Qwen3-VL-8B-Instruct/snapshots/0c351dd01ed87e9c1b53cbc748cba10e6187ff3b/` | Main processor plus temporal public backbone/processor |
75
+ | `upstream/owlv2-large-patch14-ensemble/` | `/app/.cache/huggingface/hub/models--google--owlv2-large-patch14-ensemble/snapshots/95e26936e865f87db1742128404b3c035d47d89d/` | OWLv2 detector and processor |
76
+ | `priors/` | `/app/artifacts/` | Exact six inherited prior files |
77
+
78
+ The public OWLv2 cache's `refs/main` points to its pinned revision for the unchanged model-ID call. Snapshot files can be materialized as regular files; the repository layout does not require distributing the archive's internal symlinks. The pinned resolver selects `model.safetensors` ahead of the inactive `.bin` duplicate, which is excluded from this release.
79
+
80
+ Main rendering uses up to 16 frames; temporal rendering uses up to 128 native frames with a maximum image side of 768 pixels. Generation is greedy with at most 64 new tokens, and the container loads its models eagerly. [LINEAGE.md](LINEAGE.md) and [PLATFORM_COMPATIBILITY.md](PLATFORM_COMPATIBILITY.md) document the distinct model paths and recovery behavior.
ORGANIZER_ACCESS.md CHANGED
@@ -1,11 +1,7 @@
1
- # Private organizer access
2
 
3
- Repository: [HeyDonto/SURGFIELD-ORena-2026-SEGMENT-FINAL](https://huggingface.co/HeyDonto/SURGFIELD-ORena-2026-SEGMENT-FINAL).
4
 
5
- Attribution: HeyDonto Labs. Responsible contact: Reza Nehzati, Ph.D., via [rezanehzati](https://huggingface.co/rezanehzati).
6
 
7
- ROOT reports the repository is private in the dedicated SEGMENT resource group, with automatic joining disabled. Organizer access status: **deferred_by_owner**. The verified organizer account is `orena-dkfz`; invitation and access are intentionally deferred. No organizer membership or repository-read claim is made. The repository remains on private hold. A private repository's creation does not itself prove organizer access.
8
-
9
- Asset commit: `35fc28af1d55a5a9ae8096ff2e2cd2c961aa7119`. Metadata/tools commit: `bc9e092422beac9ae51005694824cdc70205e0d5`. The upload coordinator supplies the completed revision pair in final instructions or a separate receipt; no commit self-pin is required. No asset-copy, upload or access operation was performed by this draft author. FRAME membership and permissions are unchanged and outside this handoff.
10
-
11
- The separate 13,177-row derived annotation packet requires its own private dataset URL, recipient-condition acceptance and delivery receipt. It is not a model-file attachment and does not contain the full46,758-row ancestor corpus or video pixels. No public annotation release or broader dataset permission is inferred here.
 
1
+ # Access and contact
2
 
3
+ The [SEGMENT model repository](https://huggingface.co/HeyDonto/SURGFIELD-ORena-2026-SEGMENT-FINAL) remains private under the agreed publication hold. Downloads require an authenticated account with repository read access. The team's recorded configuration uses a dedicated SEGMENT resource group with automatic joining disabled. The invitation to `orena-dkfz` remains deferred at the team's request; organizer membership and effective read access have not been confirmed. Successful finals submission does not establish repository access or permission for public release.
4
 
5
+ Attribution: **HeyDonto Labs**. Responsible contact: **Reza Nehzati, Ph.D.**, via [rezanehzati on Hugging Face](https://huggingface.co/rezanehzati). Use [LOAD.md](LOAD.md) for the immutable download revisions and [SUBMISSION_STATUS.json](SUBMISSION_STATUS.json) for the submission evidence.
6
 
7
+ The separate 13,177-row derived annotation packet requires its own private dataset URL, applicable recipient-condition acceptance and delivery receipt. It contains neither the full 46,758-row ancestor corpus nor source-video pixels, and is not attached as a model file. See [ANNOTATION_SCOPE.json](ANNOTATION_SCOPE.json) and [NOTICES.md](NOTICES.md) for its scope and restrictions. Public annotation release and broader dataset redistribution rights have not been established.
 
 
 
 
PLATFORM_COMPATIBILITY.md CHANGED
@@ -1,13 +1,49 @@
1
- # Fixed SEGMENT task and explicit degraded coverage
2
 
3
- **Finals submission: owner confirmed on September 14, 2026 that repaired R4 was successfully submitted for SEGMENT finals, with no errors reported.** The submitted archive SHA256 is `e477d307ff2e09f469d83569ec1ce8703daeef34edbba633a3f240eab955de00`. September 14 is the confirmation date; the actual submission time, platform UUIDs, terminal logs, score and ranking were not supplied. See [SUBMISSION_STATUS.json](SUBMISSION_STATUS.json) for the exact identity and evidence scope. The model weights, source files and immutable loading revisions are unchanged.
4
 
5
- **September 11 local qualification record: R4 runtime, archive and import checks completed.** The exact archive is `e477d307ff2e09f469d83569ec1ce8703daeef34edbba633a3f240eab955de00` (35,613,293,722 bytes), immutable generation `1789092734891901`. The build manifest is `ba9ddb1e78a45040803ae6f0e68c66de7a538f5e3e86397777e5fc2c488106b0` and configuration `3f43fce1081dc5db454447dbe4678947e5cefec014635e611d8d1032a0ac71c4`. Full compressed/raw archive verification covers all 26 ordered DiffIDs; actual same-generation classic Docker import reports config `3f43fce1081dc5db454447dbe4678947e5cefec014635e611d8d1032a0ac71c4` and no imported manifest digest. The agreed eight cases completed 9 application invocations and 1593 answer executions across 533 unique qIDs. Healthy20 and loose/ZIP500 matched saved content; reverse20 matched the same-image ordinary20, and both ZIP quotas matched the same-image loose500. Canonical repeated calls retained the exact disclosed deferrals. The mixed recovery case blanked only its declared failed questions and preserved later healthy outputs. These are bounded runtime/interface/content checks; no new semantic grade or official final result is claimed.
6
 
7
- R4 retains the image-owned literal SEGMENT task and the original request clocks. It stages only the selected ZIP member for each startup probe or answer, verifies that member through CRC/EOF and SHA-256, closes decoder handles, and removes the lease before continuing. Loose input remains a direct read. Model initialization is still eager. An attributed ordinary per-question exception produces empty content for that qID with measured latency and failure telemetry; successful answers are preserved and later questions continue. Unsafe or ambiguous archives, source mutation, cleanup integrity, required model initialization, environment and fixed-track failures remain fatal. The inherited precise pre-generation rules deferrals remain separately disclosed as degraded coverage. The four changed image files are the bootstrap, contract, probe entrypoint and new ZIP helper; all 51 unchanged inherited serving source files, including cap_fmt, all 50 model/prior assets, prompts, model arithmetic and generation settings remain unchanged.
 
 
 
 
 
 
8
 
9
- The full retained ZIP500 was qualified locally at the stated 4 GiB and 16 GiB tmpfs profiles, alongside loose500, healthy20, reversed20, default and persistent canonical calls, and controlled recovery. These fixtures do not establish the platform scratch quota, all possible input sizes, continuous resource peaks, cold-cache requirements, general order invariance, new semantic scores, independent-holdout performance or official platform success. Canonical deferrals and deliberately failed recovery rows are reduced scientific coverage. The prior R3 4 GiB eager-ZIP failure is a conditional local capacity reproduction, not proof of the organizer failure cause. Earlier images and failed attempts remain historical.
10
 
11
- The earlier 43d4 qualification covered six H100 invocations and 600 answer executions across 520 unique questions, including repeated20 cases and a loose500 case. It omitted the owner canonical long intervals and invalid-media control. Its151 positive 500 startup probes left 349 unmeasured; its ZIP coverage was 20, not 500. Those measurements remain historical and do not qualify R4. The completed R2 diagnostic prototype returned 10 answers twice with identical contents but failed its guard on two adaptive_n duration-domain deferrals and one no-frame control. That interface diagnosis is not an all-neural-health pass.
12
 
13
- The exact four R4 image files are available for inspection under platform_wrapper/. The inherited component_source/cap_fmt_inferrer.py and all component loaders remain unchanged. R4 retains the image-owned literal SEGMENT task and the original request clocks. It stages only the selected ZIP member for each startup probe or answer, verifies that member through CRC/EOF and SHA-256, closes decoder handles, and removes the lease before continuing. Loose input remains a direct read. Model initialization is still eager. An attributed ordinary per-question exception produces empty content for that qID with measured latency and failure telemetry; successful answers are preserved and later questions continue. Unsafe or ambiguous archives, source mutation, cleanup integrity, required model initialization, environment and fixed-track failures remain fatal. The inherited precise pre-generation rules deferrals remain separately disclosed as degraded coverage. The four changed image files are the bootstrap, contract, probe entrypoint and new ZIP helper; all 51 unchanged inherited serving source files, including cap_fmt, all 50 model/prior assets, prompts, model arithmetic and generation settings remain unchanged.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Serving and runtime qualification
2
 
3
+ The current artifact is the R4 container identified in [IMAGE_IDENTITY.json](IMAGE_IDENTITY.json). HeyDonto Labs confirmed its successful finals submission, with no errors reported, on September 14, 2026. The exact confirmation scope is recorded in [SUBMISSION_STATUS.json](SUBMISSION_STATUS.json).
4
 
5
+ ## Container identity and interface
6
 
7
+ | Identity | Value |
8
+ | --- | --- |
9
+ | Archive SHA-256 | `e477d307ff2e09f469d83569ec1ce8703daeef34edbba633a3f240eab955de00` |
10
+ | Archive size | 35,613,293,722 bytes |
11
+ | Immutable GCS generation | `1789092734891901` |
12
+ | Docker image/configuration SHA-256 | `3f43fce1081dc5db454447dbe4678947e5cefec014635e611d8d1032a0ac71c4` |
13
+ | Build OCI manifest SHA-256 | `ba9ddb1e78a45040803ae6f0e68c66de7a538f5e3e86397777e5fc2c488106b0` |
14
 
15
+ Archive, configuration and manifest hashes identify different objects. The same-generation archive passed full gzip/CRC/tar verification, all 26 ordered layer DiffID checks and an actual classic Docker import. That import exposed the configuration identity but no imported manifest digest.
16
 
17
+ The container runs as user `app` from `/app`, with entrypoint `/opt/conda/bin/python -I -S /opt/orena/segment/bootstrap.py` and offline model settings. It reads the SEGMENT request, FO definitions and per-question clips, supplied loose or in ZIP archives, and writes `/output/answer.json`. The container always uses the SEGMENT task. Question times retain the original source-video clock.
18
 
19
+ ## Input handling and recovery
20
+
21
+ R4 stages one selected ZIP member at a time for a startup probe or answer, verifies it through CRC/EOF and SHA-256, closes decoder handles and removes the staged file before continuing. Loose clips are read directly. Model initialization remains eager.
22
+
23
+ A recoverable error attributed to one question produces an empty answer for that qID, with elapsed latency and failure telemetry, while successful answers are retained and later questions continue. Unsafe or ambiguous archives, source mutation, cleanup-integrity failures, required model initialization failures, environment violations and fixed-track failures remain fatal.
24
+
25
+ Existing pre-generation deferrals to rules remain separately recorded. A valid output file can therefore include rule-based or empty answers; it does not establish healthy neural inference for every question.
26
+
27
+ The four R4 wrapper files are available under [platform_wrapper/](https://huggingface.co/HeyDonto/SURGFIELD-ORena-2026-SEGMENT-FINAL/tree/main/platform_wrapper). The 51 inherited serving source files and all 50 model/prior assets remain unchanged from R3, as do prompts, model arithmetic, frame limits and generation settings.
28
+
29
+ ## Measured local coverage
30
+
31
+ The H100 qualification matrix completed eight test cases and nine application invocations, producing 1,593 answer executions across 533 unique qIDs:
32
+
33
+ | Test condition | Observed result |
34
+ | --- | --- |
35
+ | Ordinary and reversed 20-question inputs | Ordinary outputs matched the saved reference; reversed outputs matched the same image's ordinary outputs. |
36
+ | 500 loose clips | All contents matched the saved candidate reference. |
37
+ | The same 500 clips in ZIP archives, with 4 GiB and 16 GiB scratch limits | Both runs matched the same image's loose-clip outputs. |
38
+ | Default and repeated canonical calls | Outputs retained the three documented pre-generation deferrals. |
39
+ | Controlled 23-question recovery fixture | Exactly three intended empty answers; all 20 healthy outputs preserved. |
40
+
41
+ The independently measured 500-question runtimes were approximately 1,226.6–1,267.1 seconds. [PLATFORM_COMPATIBILITY.json](PLATFORM_COMPATIBILITY.json) and [EVALUATION.json](EVALUATION.json) retain the detailed identities and measurements. Content agreement is distinct from semantic accuracy; repeated questions and deliberate failures are included in these counts.
42
+
43
+ ## Failure history and remaining limits
44
+
45
+ R3's eager extraction of the retained approximately 21 GB, 500-clip ZIP reproduced an `ENOSPC` failure locally with 4 GiB scratch, before model loading. R4 completed that fixture at the same scratch limit. This establishes a repaired local capacity failure, rather than the cause of the earlier platform submission failure, whose failing-case logs were not captured.
46
+
47
+ Earlier images also exposed input-alias and canonical-interval admission issues. Their tests and failures remain historical in [EVALUATION.json](EVALUATION.json) and the [immutable prior compatibility report](https://huggingface.co/HeyDonto/SURGFIELD-ORena-2026-SEGMENT-FINAL/blob/24133f8d4a491845db2a19c84c2d182f8b4e4831/PLATFORM_COMPATIBILITY.md). Narrow successful tests did not cover all later failure conditions.
48
+
49
+ The local matrix does not establish universal input-size limits, general order invariance, continuous resource peaks, cold-cache requirements or the organizer's scratch allocation. Sampled resource maxima, independent elapsed times and self-reported answer latencies are separate measurements. Actual submitted resource settings and terminal platform logs have not been independently recorded here. Successful submission supplies no official answer-quality score; see [EVALUATION.md](EVALUATION.md).
README.md CHANGED
@@ -14,26 +14,48 @@ tags:
14
  - orena-focus
15
  - private-evaluation
16
  ---
17
- # SURGFIELD ORena2026 SEGMENT FINAL
18
 
19
- **Finals submission: owner confirmed on September 14, 2026 that repaired R4 was successfully submitted for SEGMENT finals, with no errors reported.** The submitted archive SHA256 is `e477d307ff2e09f469d83569ec1ce8703daeef34edbba633a3f240eab955de00`. September 14 is the confirmation date; the actual submission time, platform UUIDs, terminal logs, score and ranking were not supplied. See [SUBMISSION_STATUS.json](SUBMISSION_STATUS.json) for the exact identity and evidence scope. The model weights, source files and immutable loading revisions are unchanged.
20
 
21
- **September 11 local qualification record: R4 runtime, archive and import checks completed.** The exact archive is `e477d307ff2e09f469d83569ec1ce8703daeef34edbba633a3f240eab955de00` (35,613,293,722 bytes), immutable generation `1789092734891901`. The build manifest is `ba9ddb1e78a45040803ae6f0e68c66de7a538f5e3e86397777e5fc2c488106b0` and configuration `3f43fce1081dc5db454447dbe4678947e5cefec014635e611d8d1032a0ac71c4`. Full compressed/raw archive verification covers all 26 ordered DiffIDs; actual same-generation classic Docker import reports config `3f43fce1081dc5db454447dbe4678947e5cefec014635e611d8d1032a0ac71c4` and no imported manifest digest. The agreed eight cases completed 9 application invocations and 1593 answer executions across 533 unique qIDs. Healthy20 and loose/ZIP500 matched saved content; reverse20 matched the same-image ordinary20, and both ZIP quotas matched the same-image loose500. Canonical repeated calls retained the exact disclosed deferrals. The mixed recovery case blanked only its declared failed questions and preserved later healthy outputs. These are bounded runtime/interface/content checks; no new semantic grade or official final result is claimed.
22
 
23
- Private model and resource handoff by **HeyDonto Labs**. Responsible contact: **Reza Nehzati, Ph.D.**, via [rezanehzati](https://huggingface.co/rezanehzati).
24
 
25
- This candidate answers segment-level surgical-video questions through fixed rules, six inherited priors, OWLv2 detections, a private N3 Qwen main model with the selected step900 adapter, and a public Qwen temporal model with the selected step600 adapter. Its forward-decoding sampler changes frame access while preserving the tested sample indices, RGB bytes, timestamps, frame limits and deadline. It introduces no new gradient updates.
26
 
27
- `SERVING_ASSETS.json` lists all **50 required model/prior files**: 22 private assets, 14 pinned public Qwen files, 8 pinned OWLv2 files and 6 priors. `LOAD.md` downloads immutable asset and metadata revisions into separate directories and explains the frozen loading contract. LOAD pins asset revision `35fc28af1d55a5a9ae8096ff2e2cd2c961aa7119` and metadata/tools revision `bc9e092422beac9ae51005694824cdc70205e0d5` separately. The private N3 backbone and public Qwen backbone are distinct; both are required. The inactive OWLv2 `.bin` duplicate and obsolete adapters are excluded.
28
 
29
- R4 retains the image-owned literal SEGMENT task and the original request clocks. It stages only the selected ZIP member for each startup probe or answer, verifies that member through CRC/EOF and SHA-256, closes decoder handles, and removes the lease before continuing. Loose input remains a direct read. Model initialization is still eager. An attributed ordinary per-question exception produces empty content for that qID with measured latency and failure telemetry; successful answers are preserved and later questions continue. Unsafe or ambiguous archives, source mutation, cleanup integrity, required model initialization, environment and fixed-track failures remain fatal. The inherited precise pre-generation rules deferrals remain separately disclosed as degraded coverage. The four changed image files are the bootstrap, contract, probe entrypoint and new ZIP helper; all 51 unchanged inherited serving source files, including cap_fmt, all 50 model/prior assets, prompts, model arithmetic and generation settings remain unchanged.
30
 
31
- The completed development comparison is between two registered serving implementations under the same fixed Qwen3.5-27B judge and 20-case batching protocol: candidate 2373/4000 (59.325%) versus named reference 1940/4000 (48.500%), +10.825 percentage points. Strict sensitivity is 2334/3940 versus 1906/3940. All ten original-source-video differences are positive; 5b causal consequence reasoning regressed 17.742 micro points on 62 cases from six sources. These are adaptively used development results with inherited training/data-exposure limits, not an independent holdout, an official finals score or an all-candidate ranking. The registered local scorer passed14 convenience controls (6positive,8negative). These check response/parser boundaries; they do not estimate organizer-judge agreement, real-question false-positive rates, official calibration or a finals score. All 4,000 were not run through a Docker archive. Source-derived routing equality on the 4,000 intervals 1–299s does not by itself transfer these scores to new image bytes.
 
 
 
 
32
 
33
- The full retained ZIP500 was qualified locally at the stated 4 GiB and 16 GiB tmpfs profiles, alongside loose500, healthy20, reversed20, default and persistent canonical calls, and controlled recovery. These fixtures do not establish the platform scratch quota, all possible input sizes, continuous resource peaks, cold-cache requirements, general order invariance, new semantic scores, independent-holdout performance or official platform success. Canonical deferrals and deliberately failed recovery rows are reduced scientific coverage. The prior R3 4 GiB eager-ZIP failure is a conditional local capacity reproduction, not proof of the organizer failure cause. Earlier images and failed attempts remain historical.
34
 
35
- Training lineage and selected-prefix exposure are disclosed in `LINEAGE.md` and the accompanying JSONs. The separately retained annotation packet has 13,177 derived rows, not the complete 46,758-row ancestor corpus; it contains no underlying video pixels or complete signature ledger.
36
 
37
- The repository remains private for the authorized organizer/team handoff. Upstream model files retain their Apache-2.0 notices. Dataset and derived-resource terms are separate, including the HeiCo publication condition and LapChole access restrictions; see `NOTICES.md`. This card grants no new rights to restricted training data, and does not declare a blanket license over all derived assets. It is research/challenge material, without clinical deployment qualification.
38
 
39
- MODEL_FILES.json and its identical SERVING_ASSETS.json retain the byte-pinned asset inventory used by the unchanged loader. Their embedded image_identity records ed85 ancestry; it is not the current image selection. IMAGE_IDENTITY.json supplies the explicit current-image join and per-file scope, and LINEAGE.json preserves the training/annotation history while adding the same current serving-image identity. LOAD.md continues to pin the unchanged immutable asset A and metadata M revisions.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
14
  - orena-focus
15
  - private-evaluation
16
  ---
 
17
 
18
+ # SURGFIELD ORena 2026 SEGMENT FINAL
19
 
20
+ **HeyDonto Labs** · Responsible contact: **Reza Nehzati, Ph.D.** ([Hugging Face profile](https://huggingface.co/rezanehzati))
21
 
22
+ SURGFIELD answers questions about surgical video segments for the ORena 2026 SEGMENT challenge. This repository contains the exact model components, configurations, tokenizer and processor assets, loading tools and documentation associated with the submitted R4 container.
23
 
24
+ HeyDonto Labs confirmed on September 14, 2026 that R4 was successfully submitted for SEGMENT finals, with no errors reported. The submitted archive SHA-256 is `e477d307ff2e09f469d83569ec1ce8703daeef34edbba633a3f240eab955de00`. September 14 is the confirmation date; the submission timestamp, platform identifiers, terminal logs and final score have not been independently recorded here. See [submission status](SUBMISSION_STATUS.json) and [image identity](IMAGE_IDENTITY.json).
25
 
26
+ ## Method
27
 
28
+ The serving application combines visual models with fixed answer-format rules and six inherited prior tables. It preserves the original source-video clock when interpreting questions and sampling frames.
29
 
30
+ | Component | Role and configuration |
31
+ | --- | --- |
32
+ | Main visual-language model | Private N3 Qwen3-VL-8B backbone with the selected step-900 continuation adapter. The backbone uses BF16; the FP32 LoRA adapter remains unmerged. |
33
+ | Temporal model | A separate, pinned public Qwen3-VL-8B backbone with the selected step-600 temporal adapter, applied and merged once. |
34
+ | Object detector and priors | Pinned OWLv2 model, six prior tables and fixed routing/formatting rules. |
35
 
36
+ The main path samples up to 16 frames; the temporal path supports up to 128 native frames with a maximum image side of 768 pixels. Generation is greedy, with at most 64 new tokens. The private N3 and public Qwen backbones are distinct, and both are required. [Training and component lineage](LINEAGE.md) describes their construction and exposure.
37
 
38
+ ## Download and use
39
 
40
+ Follow [LOAD.md](LOAD.md) to download the 50 model and prior files from immutable revisions, verify their checksums and optionally load the components. [MODEL_FILES.json](MODEL_FILES.json) and [WEIGHTS_SHA256SUMS](WEIGHTS_SHA256SUMS) identify the files.
41
 
42
+ The component-loading example generates no answers. Reproducing the complete application requires the corresponding Docker image, its frozen environment, routing and preprocessing. See [platform compatibility](PLATFORM_COMPATIBILITY.md) for the entrypoint, recovery behavior and measured runtime coverage.
43
+
44
+ ## Evaluation
45
+
46
+ A development comparison used a fixed Qwen3.5-27B judge, 20-question batches and 4,000 questions from ten original source videos. These measurements cover registered implementations; a full 4,000-question evaluation of the submitted R4 Docker archive has not been performed.
47
+
48
+ | Registered serving implementation | Correct | Local micro accuracy |
49
+ | --- | ---: | ---: |
50
+ | Development candidate (registered implementation) | 2,373 / 4,000 | 59.325% |
51
+ | Reference derived from the historical 333 implementation | 1,940 / 4,000 | 48.500% |
52
+
53
+ The evaluation set was reused during development and has inherited data-exposure limitations. The comparison does not reproduce the original 333 Docker image, establish superiority over every historical candidate, or predict the official finals score. Causal-consequence questions regressed by 17.742 percentage points on 62 cases. [EVALUATION.md](EVALUATION.md) reports the paired sensitivity analysis, scorer limitations and runtime evidence separately.
54
+
55
+ ## Data, access and limitations
56
+
57
+ Training includes organizer data and inherited derived annotations. The separately retained 13,177-row annotation packet excludes organizer rows and video pixels; its delivery and access are separate from this model repository. See [lineage](LINEAGE.md) and [annotation scope](ANNOTATION_SCOPE.json).
58
+
59
+ The repository remains private under the agreed publication hold. Organizer access is deferred; [ORGANIZER_ACCESS.md](ORGANIZER_ACCESS.md) records the access and delivery requirements. Upstream model licenses and dataset restrictions remain component-specific; [NOTICES.md](NOTICES.md) defines their scope.
60
+
61
+ This is a research and challenge system without clinical deployment qualification. Some inputs can use rule-based answers or explicit empty answers after a recoverable per-question failure. Runtime completion therefore does not establish neural coverage or answer correctness.