rezanehzati commited on
Commit
409bdf8
·
verified ·
1 Parent(s): 9c213c8

Refine PROCEDURE model card and documentation

Browse files
Files changed (7) hide show
  1. DATA_PROVENANCE.md +34 -34
  2. LICENSE_PROVENANCE.md +13 -13
  3. LOAD.md +25 -19
  4. MODIFICATIONS.md +15 -15
  5. ORGANIZER_ACCESS.md +13 -6
  6. README.md +45 -32
  7. UPSTREAM_NOTICES.md +12 -10
DATA_PROVENANCE.md CHANGED
@@ -1,50 +1,50 @@
1
- # Data provenance and disclosure boundary
2
 
3
- This document covers the retained PROCEDURE candidate and its selected Dense48/Point320 lineage. It does not substitute the diet of another pfull, FRAME, SEGMENT, private-data continuation or later experimental checkpoint. [LINEAGE.json](LINEAGE.json) records exact evidence hashes and distinguishes recovered metadata, intended recipe, observed consumption and selected checkpoint bytes.
4
 
5
- ## Selected adapter training records
6
 
7
- | Record | Dense48 | Point320 specialist |
8
- |---|---|---|
9
- | Selected checkpoint | 12,910; final-save and checkpoint tensors match | Independently accepted terminal 12,910 |
10
- | Initialization | Checkpoint 11,709, tensor SHA-256 `684922901db549df459e586e2d62f1535b3f6820cf6b62e7cf482a409e9ac60e` | Same exact initializer; fresh optimizer |
11
- | Recorded recipe | 51,635 rows, eight ranks, seed 42, learning rate 0.0002, requested 48 frames, maximum side 576 | Same recorded geometry and learning rate; exact point-target-restored diet below |
12
- | Proof limit | Recovery establishes selected tensor/config bytes and saved training metadata, not the exact historical corpus revision or complete per-rank input/update history | Actual terminal, consumed-input and reload evidence exists; it does not eliminate inherited exposure or annotation/decode limitations |
13
 
14
- The selected Point diet is SHA-256 `3f14484182a01b7e5e8eafad437665cc7fb7975cbeb5511296e05147412ad081`: 51,635 rows over 128 source-video identities. Its preparation restored 4,164 target answers from checked organizer sources (2,089 native views and 2,075 middle views), preserving all current windows, question strings and other columns. It did not restore original windows or guarantee that an event appears in the sampled frames.
15
 
16
- The source-lineage audit of the unchanged parent population binds 46,742 rows to organizer HeiCo/LapChole source material. A further 363 transfer rows bind to the paired parent rather than organizer faces. The 4,530 jury rows retain an unresolved original-annotation bridge. This is not a claim that every training row has a fully reconstructed original annotation source.
17
 
18
- The Point runtime recorded 103,280 consumed examples across repeated training draws, not 103,280 unique rows. Of these, 78,788 realized 48 frames and 24,492 had a requested/realized mismatch; diagnostics also record 36 substitutions, ten jittered successes and 174,825 rejected blank frames. Completion is therefore not described as perfect data/render fidelity. Later corpus filtering cannot erase possible exposure inherited through checkpoint 11,709.
19
 
20
- ## Named sources and restrictions
21
 
22
- | Source | Relevant provenance | Distribution scope |
23
- |---|---|---|
24
- | [HeiCo-FOCUS-VQA](https://huggingface.co/datasets/orena-dkfz/heico-focus-vqa) | Organizer colorectal VQA training material and derived views/statistics. The audited parent source covers organizer track-specific training faces; no claim is made that every ancestor diet used the same revision. | The source records CC BY-NC-SA 4.0 plus a challenge-publication restriction. Historical campaign records preserve dataset-specific challenge-use rulings and annotation-delivery conditions; they are not a new redistribution grant or proof that an obligation has been completed. |
25
- | [LapChole-FOCUS-VQA](https://huggingface.co/datasets/orena-dkfz/lapchole-focus-vqa) | Organizer laparoscopic cholecystectomy VQA training material, derived point views and the selected organizer-derived v39 prior. | Access and data handling are governed by the organizer Data Usage Agreement, not a generic open-data license. The published conditions restrict data sharing, public release and reidentification. Consult the original agreement for model/data distinctions. |
26
- | Jury and transfer-derived rows | 4,530 jury rows retain an original-annotation gap; 363 transfer rows are traced to their paired parent. | Raw rows and source media are not distributed here. This unresolved history is not filled with an invented public-dataset attribution or a claim of complete clearance. |
27
- | Private V2 surgery and its derivatives | A separately identified six-predicate diagnostic used for evaluation only; no qualified matched model-accuracy comparison resulted. | Evaluation-only; excluded from adaptation and training-rule selection. No V2 media or reference labels are included. |
28
 
29
- The [PROCEDURE rules](https://procedure.orena-focus-challenge.org/rules/) and captured method-description form require actual data-source disclosure, including descriptions of private sources. This repository does not claim that the method form, architecture figure, restricted annotation handoff or publication obligations have been submitted or acknowledged. The current source pages are attribution/access references; they do not replace pinned historical training evidence.
30
 
31
- ## Included derived runtime priors
32
 
33
- The package includes six small JSON assets under `image_root/app/artifacts/` in addition to learned weights:
34
 
35
- | File | Role and provenance limit |
36
- |---|---|
37
- | `train_modes.json` | Training modes by track, capability and answer format; no embedded dataset revision or complete derivation receipt. |
38
- | `x11_stem_table.json` | PROCEDURE question-stem answer modes; no embedded dataset revision or complete derivation receipt. |
39
- | `stem_table.json` | Track/capability/format/stem modes; source requires at least five training questions and 40% support, but the asset has no complete derivation receipt. |
40
- | `fo_quadrant_priors_colorectal.json` | HeiCo colorectal spatial priors, attributed by the exact loader source; no embedded dataset revision. |
41
- | `cholec_priors_v39_organizer.json` | Selected organizer LapChole prior derived using the shipped detector; embedded record describes 57 videos, 3,477 frames and 14,258 detections. This is the v39 artifact, not the historical CholecTrack20-derived v38 fallback. |
42
- | `duration_priors.json` | HeiCo procedure-training duration constants. The constructor reads the asset, while the retained L16 applicability flag is disabled. |
43
 
44
- ASSET_MANIFEST.json binds their bytes. These assets may contain source-derived answer modes and question stems; they are not described as purely learned parameters. Their embedded historical rulings are preserved as provenance rather than converted into a new legal conclusion.
45
 
46
- ## What is withheld and what the scores mean
47
 
48
- Private surgical media and evaluation banks are not distributed here. The original inference source may retain unit-test fixtures, and the above runtime priors are explicitly included. No raw Parquet/JSONL training dataset, per-question evaluation-output bank or evaluation gold is selected for this handoff.
49
 
50
- The 1,087-question local development board was repeatedly used for model/inference selection. It contains zero challenge OOD rows and only one clinical-flagged row. Its local scores cannot establish a clinical accuracy estimate, untouched generalization, an official platform result or a finals rank. The retained failed rows and the distinction between historical and fresh native protocols are reported in README.md and EVALUATION_SUMMARY.json.
 
1
+ # Training data and provenance
2
 
3
+ This document describes the Dense48 and Point320 adapters used in the submitted PROCEDURE model. Other foundations, tracks and later private-data experiments have separate training histories. [LINEAGE.json](LINEAGE.json) records the evidence for checkpoint identity, saved metadata, intended recipes and observed training inputs.
4
 
5
+ ## Selected adapters
6
 
7
+ | Training record | Dense48 | Point320 specialist |
8
+ | --- | --- | --- |
9
+ | Selected checkpoint | 12,910; final-save and checkpoint tensors match | Terminal checkpoint 12,910, independently checked |
10
+ | Initialization | Checkpoint 11,709, tensor SHA-256 `684922901db549df459e586e2d62f1535b3f6820cf6b62e7cf482a409e9ac60e` | Same initializer; fresh optimizer |
11
+ | Recorded recipe | 51,635 rows, eight ranks, seed 42, learning rate 0.0002, requested 48 frames, maximum side 576 | Same recorded geometry and learning rate; point-target-restored dataset described below |
12
+ | Evidence limits | Recovered tensors, configuration and saved metadata establish the selected checkpoint, but not the exact historical corpus revision or complete per-rank input/update history | Terminal, consumed-input and reload records are available; inherited exposure and annotation/decode limitations remain |
13
 
14
+ The selected Point training dataset has SHA-256 `3f14484182a01b7e5e8eafad437665cc7fb7975cbeb5511296e05147412ad081`: 51,635 rows over 128 source-video identities. Preparation restored 4,164 target answers from checked organizer sources: 2,089 native views and 2,075 middle views. Windows, question strings and other columns were retained. Restoring an answer does not guarantee that the corresponding event is visible in the sampled input.
15
 
16
+ The parent dataset audit traces 46,742 rows to organizer HeiCo/LapChole material. Another 363 transfer rows are traced to their paired parent records. For 4,530 jury rows, the link to the original annotation remains unresolved. Complete original-source provenance is therefore unavailable for part of the training history.
17
 
18
+ Point training recorded 103,280 consumed examples across repeated draws, not unique examples. Of these, 78,788 realized 48 frames and 24,492 had a requested/realized frame-count mismatch. Diagnostics record 36 substitutions, ten jittered successes and 174,825 rejected blank frames. These observations limit claims about training input fidelity. Filtering later datasets does not remove possible exposure inherited through checkpoint 11,709.
19
 
20
+ ## Sources and usage restrictions
21
 
22
+ | Source | Contribution | Applicable scope |
23
+ | --- | --- | --- |
24
+ | [HeiCo-FOCUS-VQA](https://huggingface.co/datasets/orena-dkfz/heico-focus-vqa) | Organizer colorectal VQA training material and derived views/statistics. The parent audit covers track-specific organizer training records; ancestor revisions are not fully established. | The source records CC BY-NC-SA 4.0 and a challenge-publication restriction. Historical records retain dataset-specific challenge-use decisions and annotation-delivery conditions; these do not establish new redistribution rights or completion of those obligations. |
25
+ | [LapChole-FOCUS-VQA](https://huggingface.co/datasets/orena-dkfz/lapchole-focus-vqa) | Organizer laparoscopic cholecystectomy VQA material, derived point views and the selected organizer-derived v39 prior. | Governed by the organizer Data Usage Agreement, including restrictions on sharing, public release and reidentification. Consult that agreement for its treatment of models and data. |
26
+ | Jury and transfer-derived rows | 4,530 jury rows with an unresolved original-annotation link; 363 transfer rows traced to their paired parent records. | Raw rows and media are excluded from this repository. The unresolved history prevents a claim of fully reconstructed provenance or complete clearance. |
27
+ | Private V2 surgery and derivatives | One evaluation-only diagnostic with six predicates; no qualified matched accuracy comparison resulted. | Excluded from training, adaptation and training-rule selection. Media and reference labels are not included. |
28
 
29
+ The [PROCEDURE rules](https://procedure.orena-focus-challenge.org/rules/) and recorded method-description form require disclosure of actual data sources, including private sources. This repository does not establish completion or organizer acknowledgement of the method form, architecture figure, restricted annotation delivery or publication obligations. The linked source pages provide attribution and access terms; the pinned training records document historical usage.
30
 
31
+ ## Included runtime priors
32
 
33
+ Six small JSON files under `image_root/app/artifacts/` supplement the learned parameters. [ASSET_MANIFEST.json](ASSET_MANIFEST.json) records their sizes and hashes.
34
 
35
+ | File | Role and provenance |
36
+ | --- | --- |
37
+ | `train_modes.json` | Training answer modes by track, capability and format. No embedded dataset revision or complete derivation record. |
38
+ | `x11_stem_table.json` | PROCEDURE question-stem answer modes. No embedded dataset revision or complete derivation record. |
39
+ | `stem_table.json` | Answer modes by track, capability, format and stem. The source requires at least five training questions and 40% support; the asset lacks a complete derivation record. |
40
+ | `fo_quadrant_priors_colorectal.json` | HeiCo colorectal spatial priors, attributed by the included loader source. No embedded dataset revision. |
41
+ | `cholec_priors_v39_organizer.json` | Organizer LapChole prior derived with the included detector. Its record describes 57 videos, 3,477 frames and 14,258 detections. This selected v39 asset is distinct from the historical CholecTrack20-derived v38 fallback. |
42
+ | `duration_priors.json` | HeiCo procedure-training duration constants. The constructor reads the asset, but the selected L16 applicability flag is disabled. |
43
 
44
+ These assets contain source-derived answer modes, question stems or statistics. Their recorded historical usage decisions remain provenance records; they do not establish new usage rights.
45
 
46
+ ## Distribution and evaluation limits
47
 
48
+ The repository excludes raw Parquet/JSONL training datasets, private surgical media, annotation datasets, per-question evaluation outputs and reference answers. The exact inference source may retain original unit-test fixtures, and the six runtime priors listed above are included.
49
 
50
+ The local 1,087-question board was repeatedly used for development and selection. It contains no challenge OOD questions and only one clinical-flagged question. Its results do not estimate clinical accuracy, untouched generalization, official platform performance or final rank. [README.md](README.md) and [EVALUATION_SUMMARY.json](EVALUATION_SUMMARY.json) report the historical and fresh native protocols separately, retaining failed questions in their stated denominators.
LICENSE_PROVENANCE.md CHANGED
@@ -1,17 +1,17 @@
1
- # License provenance
2
 
3
- This document reports component terms and their evidence. It does not grant a new blanket license for the aggregate private package or assert that model, source, dataset and annotation terms are interchangeable.
4
 
5
- | Component | Recorded terms or status | Evidence and scope |
6
- |---|---|---|
7
- | Qwen/Qwen3.8-27B upstream metadata and model | Apache License 2.0 | Revision `1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0`; original [LICENSE](image_root/app/artifacts/q38_base/LICENSE) and README match independently fetched pinned upstream bytes. |
8
- | Prequantized Qwen base | Quantized derivative of the recorded Qwen base | Exact selected tensor/config hashes and quantization configuration are in ASSET_MANIFEST.json and LINEAGE.json. Retain the Qwen notices; no new aggregate-package license is declared here. |
9
- | Dense48 and Point320 LoRA adapters | Team-selected adaptations with upstream and data provenance | Exact terminal checkpoint ancestry is in LINEAGE.json. The existence of an upstream Apache license does not establish unrestricted rights to all data-derived material. |
10
- | google/owlv2-large-patch14-ensemble auxiliary cache | Apache License 2.0 as recorded by the model card | Pinned revision `95e26936e865f87db1742128404b3c035d47d89d`; [Google's original card](https://huggingface.co/google/owlv2-large-patch14-ensemble/blob/95e26936e865f87db1742128404b3c035d47d89d/README.md), matching cached SHA-256 `a2e10c3916166f08eaf2ab43ca1eb63c6116df228dea655c56e4c8e1607ecfe9`. |
11
- | Team inference source | Preserved exact source; no additional blanket grant recorded by this document | All 85 repaired source file hashes are in SOURCE_MANIFEST.json. Existing notices remain intact. |
12
- | Third-party Python/CUDA runtime | Per-package terms | Installed versions are observed in EXACT_DEPENDENCIES.json; the complete runtime is retained in the original OCI image rather than repackaged here. |
13
- | Organizer/private data, annotations and six derived runtime priors | Source-specific access and usage terms | DATA_PROVENANCE.md and LINEAGE.json preserve the recorded sources, restrictions and remaining evidence limits. Private media and evaluation-bank payloads are excluded. |
14
 
15
- The repository is private, and configured organizer read access is not a license grant or permission to make restricted data public. No `license: apache-2.0` label is applied to the whole model card merely because the base and auxiliary model report Apache terms. Original attribution and applicable license text must accompany components where redistribution is permitted.
16
 
17
- The exact inventory's LICENSES.json is the machine-readable provenance companion. Dataset terms and any recorded challenge-specific permissions should be read at their original source; this document does not replace those instruments or declare a separate obligation completed.
 
1
+ # Component licenses
2
 
3
+ License terms apply separately to model weights, serving source, runtime software and data-derived assets. No single license is declared for this entire private repository.
4
 
5
+ | Component | Recorded terms or status | Evidence |
6
+ | --- | --- | --- |
7
+ | Qwen/Qwen3.8-27B model and metadata | Apache License 2.0 | Revision `1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0`; the included [LICENSE](image_root/app/artifacts/q38_base/LICENSE) and README match independently retrieved upstream copies. |
8
+ | Prequantized Qwen base | Quantized derivative of the Qwen base | Tensor/configuration hashes and quantization settings are in [ASSET_MANIFEST.json](ASSET_MANIFEST.json) and [LINEAGE.json](LINEAGE.json). Qwen attribution and notices remain applicable. |
9
+ | Dense48 and Point320 LoRA adapters | Team adaptations with upstream and training-data provenance | [LINEAGE.json](LINEAGE.json) records the selected checkpoints. The upstream Apache license does not establish unrestricted rights to all data-derived material. |
10
+ | google/owlv2-large-patch14-ensemble detector | Apache License 2.0 as recorded by its model card | Revision `95e26936e865f87db1742128404b3c035d47d89d`; [Google's card](https://huggingface.co/google/owlv2-large-patch14-ensemble/blob/95e26936e865f87db1742128404b3c035d47d89d/README.md) matches the cached copy, SHA-256 `a2e10c3916166f08eaf2ab43ca1eb63c6116df228dea655c56e4c8e1607ecfe9`. |
11
+ | Team inference source | Existing notices preserved; no additional blanket license granted here | [SOURCE_MANIFEST.json](SOURCE_MANIFEST.json) hashes the 85 repaired serving files. |
12
+ | Third-party Python/CUDA runtime | Individual package terms | [EXACT_DEPENDENCIES.json](EXACT_DEPENDENCIES.json) records installed versions. The complete runtime remains in the reference OCI image. |
13
+ | Organizer/private data, annotations and six derived runtime priors | Source-specific access and usage terms | [DATA_PROVENANCE.md](DATA_PROVENANCE.md) and [LINEAGE.json](LINEAGE.json) document sources, restrictions and remaining gaps. Private media and per-question evaluation records are excluded. |
14
 
15
+ The private repository's organizer read access does not authorize public redistribution of restricted material. Applicable attribution and license text must accompany components wherever redistribution is permitted. The model card therefore has no repository-wide `license: apache-2.0` declaration.
16
 
17
+ [LICENSES.json](LICENSES.json) provides the machine-readable license inventory. Original dataset agreements and recorded challenge-specific permissions remain the authoritative terms; this document does not replace them or establish completion of separate submission or publication obligations.
LOAD.md CHANGED
@@ -1,10 +1,10 @@
1
- # Loading the retained PROCEDURE assets
2
 
3
  Repository: `HeyDonto/SURGFIELD-ORena-2026-PROCEDURE-FINAL`.
4
 
5
- ## Pin the repaired source release
6
 
7
- The original model-asset custody revision remains `6a7fc05f196b98d751e3a14775d60f1161166d2a`. Its prior full download and CPU checks concern the original asset/source snapshot and are preserved as history; they do not qualify this repaired source as an HF GPU load. Download the current repaired source and unchanged assets together at the immutable source release below. Do not combine the new verifier with an old 84-source snapshot.
8
 
9
  ```text
10
  ORIGINAL_ASSET_PROVENANCE_REVISION=6a7fc05f196b98d751e3a14775d60f1161166d2a
@@ -23,11 +23,17 @@ snapshot = snapshot_download(
23
  )
24
  ```
25
 
26
- Use locally configured private-repository authentication; never put a token in recorded source or commands. The publisher binds SOURCE_REVISION in a second small metadata commit after the source commit exists. No model weights are uploaded again.
27
 
28
- ## Preserve the image paths
29
 
30
- Most files are stored under `image_root/` with their original absolute image path made relative: `/app/artifacts/...` becomes `image_root/app/artifacts/...`, and `/app/runtime_zoom_entry.py` becomes `image_root/app/runtime_zoom_entry.py`. The [ASSET_MANIFEST.json](ASSET_MANIFEST.json) lists 46 physical asset files totaling 33,967,875,210 bytes and 11 relative OWLv2 snapshot aliases. [SOURCE_MANIFEST.json](SOURCE_MANIFEST.json) lists the 85 exact repaired source files. [IMAGE_IDENTITY.json](IMAGE_IDENTITY.json) binds the current repaired image and archive, with original 8d3 ancestry. Both OWLv2 weight formats are retained by the asset inventory; this is a conservative complete cache set, not a minimum-file claim. Hugging Face disallows a `.cache` repository path segment, so the OWLv2 cache alone is stored under `image_root/app/hf_cache/` instead of its original `/app/.cache/`. The manifests preserve both `image_path` and `repo_path`, and aliases explicitly record `resolved_repo_path`. These manifests supply every selected path, byte size and SHA-256. Verify those values before using any component; directory names and adapter configuration alone do not establish tensor identity. Download canonical physical files once, then recreate the eleven relative aliases exactly from their repository mappings when reconstructing the cache. To restore the original image filesystem semantics in a separately assembled root, map the contents of `image_root/app/hf_cache/` back to `/app/.cache/`, preserving the relative links. Do not edit the baked `HF_HOME=/app/.cache/huggingface` configuration and call that an unchanged native run; this package does not overwrite a running container automatically. Never turn an alias into an unrelated model download.
 
 
 
 
 
 
31
 
32
  Run the bundled verifier from the trusted documentation checkout against the downloaded snapshot:
33
 
@@ -35,15 +41,15 @@ Run the bundled verifier from the trusted documentation checkout against the dow
35
  python3 -B verify_files.py procedure-final-repaired --restore-aliases
36
  ```
37
 
38
- It authenticates both manifests, streams the SHA-256 of all 131 selected physical asset/source files, and verifies or creates the eleven prescribed relative aliases. It refuses a mismatched existing alias rather than overwriting it. Without `--restore-aliases` it performs read-only verification. The pass result is file verification, not a model-load or GPU-runtime result.
39
 
40
- The Dense48 and Point320 adapters are distinct 247,513,400-byte files even though their 1,135-byte configuration files match. Both use the selected checkpoint 12,910. Do not merge, rename or exchange them, substitute another base revision, regenerate int8 weights, or install a later research adapter when reproducing this candidate.
41
 
42
- ## Load the model components for inspection or custom research
43
 
44
- The exact image reports Python **3.11.11** and the following installed versions: Transformers **5.14.0**, PEFT **0.20.0**, Torch **2.13.0**, bitsandbytes **0.50.1**, Accelerate **1.14.0**, safetensors **0.8.0**, tokenizers **0.22.2**, huggingface-hub **1.29.0**, Pillow **12.3.0**, and PyAV **16.1.0**. These are observed installed versions, not the older `PYTORCH_VERSION` environment label. Matching these numbers alone does not reconstruct the CUDA libraries or full image.
45
 
46
- [load_components.py](load_components.py) is a concrete CUDA component-loading example derived from the shipped `vlm_cholec.py` and `object_specialist.py` calls. It verifies the extracted files, loads the processor from `image_root/app/artifacts/q38_base`, loads the existing prequantized base from `q38_base_int8_prequant`, attaches Dense48 as `default`, and loads Point320 as `point_object` while preserving RNG state for that second attachment. It does not supply a new quantization configuration or merge the adapters.
47
 
48
  ```python
49
  from load_components import load
@@ -59,7 +65,7 @@ finally:
59
  model.eval()
60
  ```
61
 
62
- For a separate auxiliary-detector component experiment, use the relocated, exact OWLv2 snapshot rather than a model name resolved through the network:
63
 
64
  ```python
65
  from pathlib import Path
@@ -78,20 +84,20 @@ owl_model = Owlv2ForObjectDetection.from_pretrained(
78
  owl_model.eval()
79
  ```
80
 
81
- This optional snippet follows the original detector's processor/model API and assumes the manifest verifier has restored the cache aliases. It is not a second whole-pipeline qualification.
82
 
83
- The example checks exact package versions and CUDA availability. Its API sequence was checked with explicit CPU doubles and source-call comparison. Actual downloaded-asset CPU configuration, processor and header checks passed separately; the CUDA component-loading function itself has not been executed as a new GPU load. It intentionally does not recreate the native question router, warmups, per-request controller witnesses, input shell or temporal budget policy. Use the repaired image below for its separately recorded native qualification; the HF component example itself has no new GPU qualification.
84
 
85
- ## Reproduce the demonstrated runtime
86
 
87
- The extracted files are not a standalone operating-system or Python environment. Loading them with an arbitrary current Transformers/PEFT installation is not the qualified native runtime. The current repaired OCI image is:
88
 
89
  ```text
90
  us-central1-docker.pkg.dev/heydonto-425716/surgfield/surgfield-proc-q38-validation@sha256:f0590e6d79097a37f502260cf665624bd4ca87e9cbda66d240d6d4db7d0bb63d
91
  ```
92
 
93
- Its image configuration is `sha256:94d0791cb96f3ac9248e9ec7c918e3bbdc5960646b48d201cb478d488fcda576`. The archive has SHA-256 `b63dcf1bbee6554942a78edc823dfc4b38a89dba54f029acdd28a8c91bfa4a2d` and size 52,775,393,310 bytes, at generation `1789119563478283` of `gs://heydonto-surgfield-research/finals/procedure/candidates/20260911/procedure-diagnostic-repair-20260911-0846-b/procedure-diagnostic-repair-20260911-0846-b.tar.gz`. Original 8d3 remains immutable ancestry at revision `b2176cf6d0f77e58e34a90583041b1a929417fb3`. Registry/GCS access is separate from HF access; the archive is not duplicated here.
94
 
95
- The original default command is `/opt/conda/bin/python -B /app/runtime_zoom_entry.py --submission`, working directory `/app`, UID/GID 1000. The shell reads `/input/request.json`, organizer `/input/FO_definitions.json` and the supplied video layout, and writes `/output/answer.json` containing `{qID, content, latency}` records. Preserve the original input contract and use only data you are authorized to process. Private evaluation fixtures and reference answers are not supplied here.
96
 
97
- Historical8d3 native score checks used H100 with 16 CPUs and 196,608 MiB; current repaired-image execution checks are separately scoped in RELEASE_VERIFICATION.json. These resources describe those experiments; they are not a proven minimum or a claim that every deployment meets the original process budget. Preserve the image's default entrypoint, flags and environment for an exact comparison.
 
1
+ # Loading the PROCEDURE model assets
2
 
3
  Repository: `HeyDonto/SURGFIELD-ORena-2026-PROCEDURE-FINAL`.
4
 
5
+ ## Download a pinned snapshot
6
 
7
+ The immutable source release below contains the 85 repaired serving files and unchanged model assets. The original asset release remains `6a7fc05f196b98d751e3a14775d60f1161166d2a`. Its full-download and CPU validation results apply to that earlier snapshot. The current verifier expects the repaired 85-file source tree, so it is incompatible with the earlier 84-file snapshot.
8
 
9
  ```text
10
  ORIGINAL_ASSET_PROVENANCE_REVISION=6a7fc05f196b98d751e3a14775d60f1161166d2a
 
23
  )
24
  ```
25
 
26
+ Private-repository access requires locally configured Hugging Face authentication.
27
 
28
+ ## File layout and verification
29
 
30
+ Most files retain their original image paths under `image_root/`: `/app/artifacts/...` maps to `image_root/app/artifacts/...`, and `/app/runtime_zoom_entry.py` maps to `image_root/app/runtime_zoom_entry.py`.
31
+
32
+ [ASSET_MANIFEST.json](ASSET_MANIFEST.json) lists 46 physical asset files totaling 33,967,875,210 bytes and 11 relative OWLv2 snapshot aliases. [SOURCE_MANIFEST.json](SOURCE_MANIFEST.json) lists the 85 repaired serving files. [IMAGE_IDENTITY.json](IMAGE_IDENTITY.json) records the repaired image, archive and original 8d3 ancestry. The asset inventory includes both OWLv2 weight formats to preserve the complete selected cache; it is not a minimum download set.
33
+
34
+ Hugging Face disallows a `.cache` repository path segment. The OWLv2 cache is therefore stored at `image_root/app/hf_cache/`, corresponding to `/app/.cache/` in the original image. Manifests record both `image_path` and `repo_path`; aliases also have an explicit `resolved_repo_path`. Each selected file has a recorded size and SHA-256. These hashes identify the payload independently of its directory name or adapter configuration.
35
+
36
+ Download each physical file once and recreate the eleven relative aliases from the manifest mappings. When assembling a filesystem with the original image layout, map the contents of `image_root/app/hf_cache/` back to `/app/.cache/` and preserve the relative links. The native configuration remains `HF_HOME=/app/.cache/huggingface`. The snapshot and examples below do not modify a running container or fetch replacement models for cache aliases.
37
 
38
  Run the bundled verifier from the trusted documentation checkout against the downloaded snapshot:
39
 
 
41
  python3 -B verify_files.py procedure-final-repaired --restore-aliases
42
  ```
43
 
44
+ The verifier authenticates both manifests, streams SHA-256 checks over all 131 selected physical asset/source files, and verifies or creates the eleven specified aliases. An existing alias with a different target causes verification to fail. Without `--restore-aliases`, verification is read-only. A successful result establishes file integrity; model loading and GPU execution require separate validation.
45
 
46
+ Dense48 and Point320 are distinct 247,513,400-byte adapters, although their 1,135-byte configuration files match. Both use checkpoint 12,910. Reproducing the released model requires these separate adapters, the recorded base revision and the existing prequantized int8 weights, without merging or exchanging adapters or regenerating the quantized base.
47
 
48
+ ## Load components for inspection or research
49
 
50
+ The image reports Python **3.11.11** and these installed versions: Transformers **5.14.0**, PEFT **0.20.0**, Torch **2.13.0**, bitsandbytes **0.50.1**, Accelerate **1.14.0**, safetensors **0.8.0**, tokenizers **0.22.2**, huggingface-hub **1.29.0**, Pillow **12.3.0**, and PyAV **16.1.0**. These values come from the installed packages; the older `PYTORCH_VERSION` environment label is not authoritative. Matching package versions alone does not reproduce the CUDA libraries or complete image environment.
51
 
52
+ [load_components.py](load_components.py) provides a CUDA loading example based on the shipped `vlm_cholec.py` and `object_specialist.py` calls. It verifies the files, loads the processor from `image_root/app/artifacts/q38_base` and the prequantized base from `q38_base_int8_prequant`, then attaches Dense48 as `default` and Point320 as `point_object`. RNG state is preserved during the second adapter attachment. The loader uses the saved quantization configuration and keeps the adapters separate.
53
 
54
  ```python
55
  from load_components import load
 
65
  model.eval()
66
  ```
67
 
68
+ The auxiliary OWLv2 detector can also be loaded from its exact local snapshot:
69
 
70
  ```python
71
  from pathlib import Path
 
84
  owl_model.eval()
85
  ```
86
 
87
+ This optional example uses the original detector's processor/model API and requires the verifier to have restored the cache aliases.
88
 
89
+ The component loader checks package versions and CUDA availability. Its API sequence was checked against the serving source using CPU test doubles. Separate tests on the downloaded assets passed configuration, processor and tensor-header checks. The CUDA component-loading function itself has not been executed as a new GPU load. These component examples do not reproduce the native question router, warmups, per-request control checks, input shell or temporal budget policy. Validation of the repaired image's full runtime is recorded separately in [RELEASE_VERIFICATION.json](RELEASE_VERIFICATION.json).
90
 
91
+ ## Run the validated image
92
 
93
+ The extracted files do not include a standalone operating system or Python environment. The repaired OCI image provides the environment used for native execution validation:
94
 
95
  ```text
96
  us-central1-docker.pkg.dev/heydonto-425716/surgfield/surgfield-proc-q38-validation@sha256:f0590e6d79097a37f502260cf665624bd4ca87e9cbda66d240d6d4db7d0bb63d
97
  ```
98
 
99
+ Its image configuration is `sha256:94d0791cb96f3ac9248e9ec7c918e3bbdc5960646b48d201cb478d488fcda576`. The archive has SHA-256 `b63dcf1bbee6554942a78edc823dfc4b38a89dba54f029acdd28a8c91bfa4a2d` and size 52,775,393,310 bytes, at generation `1789119563478283` of `gs://heydonto-surgfield-research/finals/procedure/candidates/20260911/procedure-diagnostic-repair-20260911-0846-b/procedure-diagnostic-repair-20260911-0846-b.tar.gz`. Original 8d3 remains available as immutable ancestry at revision `b2176cf6d0f77e58e34a90583041b1a929417fb3`. Registry/GCS access is separate from Hugging Face access; the archive is not duplicated in this repository.
100
 
101
+ The default command is `/opt/conda/bin/python -B /app/runtime_zoom_entry.py --submission`, with working directory `/app` and UID/GID 1000. The shell reads `/input/request.json`, organizer `/input/FO_definitions.json` and the supplied video layout. It writes `/output/answer.json` with `{qID, content, latency}` records. Inputs must follow this contract and be authorized for use. Private evaluation fixtures and reference answers are not included.
102
 
103
+ Historical 8d3 native score checks used H100 with 16 CPUs and 196,608 MiB of host memory. Current repaired-image execution checks are documented separately in [RELEASE_VERIFICATION.json](RELEASE_VERIFICATION.json). These are test allocations; minimum resource requirements and guaranteed process runtimes have not been established. Reproduction uses the image's default entrypoint, flags and environment.
MODIFICATIONS.md CHANGED
@@ -1,22 +1,22 @@
1
- # Modifications and selected-runtime scope
2
 
3
  Prepared by **HeyDonto Labs** for `HeyDonto/SURGFIELD-ORena-2026-PROCEDURE-FINAL`. Contact: **Reza Nehzati, Ph.D.**
4
 
5
- The current source is copied byte-for-byte from the qualified repaired image `sha256:94d0791cb96f3ac9248e9ec7c918e3bbdc5960646b48d201cb478d488fcda576`. Original 8d3 is retained as ancestry at immutable revision `b2176cf6d0f77e58e34a90583041b1a929417fb3`. No model weights, adapter configuration, router, prompt, frame sampler, temporal window or inference clock is changed by this repair.
6
 
7
- ## Reliability maintenance
8
 
9
- The five net changes are `VEHICLE_INPUTS.json`, `solution/entrypoint/probe_entrypoint.py`, new `solution/entrypoint/lazy_video_staging.py`, `solution/shape_mode_diagnostic.py` and its policy. ZIP metadata is catalogued before serving; requested clips are materialized one at a time, released after request/loader cleanup, and the owned staging directory is removed. Input-root aliases remain allowed; unsafe descendant links and archive paths are refused.
10
 
11
- Diagnostics are bounded to 16,384 events/16 MiB. Recording stops after its event/byte quota or narrow ENOSPC/EDQUOT exhaustion while inference and policy/model/identity checks continue. A normal-exit saturation receipt is best effort and may be absent after forced termination. A cleanup failure may leave one bounded partial record. The separate runtime_media 512-file/128 MiB limit is unchanged and may retain a coarse temporal answer when refinement metadata fills it.
12
 
13
- This maintenance is an execution-reliability change, not a new score experiment.
14
 
15
  ## Model adaptation
16
 
17
- The selected runtime records upstream `Qwen/Qwen3.8-27B` with actual configuration architecture `Qwen3_5ForConditionalGeneration`. It uses a prequantized bitsandbytes int8 vision-language base and two separate LoRA adapters, Dense48 and Point320, selected at checkpoint 12,910. Each selected adapter has rank 32, alpha 64 and dropout 0.05. “Point320” names its 320-tensor adapter payload, not 320 training updates. The selected Point adapter completed 12,910 optimizer updates; it is the point-target-restored quality arm, distinct from later private-data continuation experiments. The base's recorded model identifier and its actual architecture field must both be retained in the inventory; the family name is not a substitute for configuration and tensor hashes.
18
 
19
- Dense48 is the default selected adapter; Point320 is separately loaded for the configured specialist route. The combined selected adapter payload and runtime lock are checked as distinct parts of the release. Exact model and source inventories define the release; files with similar directory or checkpoint names are not equivalent.
20
 
21
  ## Selected tensor files
22
 
@@ -26,16 +26,16 @@ Dense48 is the default selected adapter; Point320 is separately loaded for the c
26
  | `/app/artifacts/proc_q38_point_ck12910/adapter_model.safetensors` | 247,513,400 | `24762d93abdd35704e0087feefb5a9e4d0794d6b7cea39d0df195f93829481a7` |
27
  | `/app/artifacts/q38_base_int8_prequant/model.safetensors` | 29,924,045,158 | `238b0622e3b2446e71daeaef8289f9a81560ed26cc280f47182a0475917b6f04` |
28
 
29
- The two adapter configuration files share SHA-256 `52e9ce4678fc4bcf4e717245952665c921c13a28247c67b679eeafdb25c4f6f8`. Their matching configuration does not make the adapter tensors identical.
30
 
31
- ## Inference composition
32
 
33
- The retained source composes question routing, sampled video frames, model generation, format validation, rule-based answers and output hygiene. Six training/annotation-derived JSON prior and stem-statistic assets are included alongside learned weights, as listed in DATA_PROVENANCE.md. A source-bound 120-second temporal refinement policy uses actual decoded timestamps and retains the verified coarse answer when the original budget guard declines a refinement. Runtime observations distinguish loader return, per-question generation and rules/fallback behavior.
34
 
35
- Private cache preparation, run-as-application setup, input staging and the original default launcher are part of the demonstrated behavior. Merely invoking the base model with a prompt does not reproduce the full PROCEDURE pipeline.
36
 
37
- ## Excluded successors and legacy material
38
 
39
- The shipping-guard `441` image, 240-second policy experiment, private-data training successors, Gate512/LR1920 or recovered1440 branches, V2 observer derivatives and synthetic warmup experiments are not this candidate. Older weights and optimizer state retained in the original image are separately classified by the inventory and excluded from the selected Hugging Face asset set. Their presence in the archive does not make them active selected components.
40
 
41
- The original archive remains the lineage reference. The Hugging Face package is an explicitly selected extraction and is not claimed to reproduce every incidental cache file or the entire container filesystem.
 
1
+ # Model and runtime modifications
2
 
3
  Prepared by **HeyDonto Labs** for `HeyDonto/SURGFIELD-ORena-2026-PROCEDURE-FINAL`. Contact: **Reza Nehzati, Ph.D.**
4
 
5
+ The serving source matches the validated repaired image `sha256:94d0791cb96f3ac9248e9ec7c918e3bbdc5960646b48d201cb478d488fcda576` byte-for-byte. Original 8d3 is preserved at immutable revision `b2176cf6d0f77e58e34a90583041b1a929417fb3`. The repair leaves model weights, adapter configuration, routing, prompts, frame sampling, temporal windows and inference clocks unchanged.
6
 
7
+ ## Reliability repairs
8
 
9
+ Five serving files changed: `VEHICLE_INPUTS.json`, `solution/entrypoint/probe_entrypoint.py`, the new `solution/entrypoint/lazy_video_staging.py`, `solution/shape_mode_diagnostic.py` and its policy. ZIP metadata is indexed before serving. Requested clips are extracted one at a time and released after request processing and video-loader cleanup; the process removes its own staging directory. Input-root aliases remain supported, while unsafe descendant links and archive paths are rejected.
10
 
11
+ Diagnostics are limited to 16,384 events and 16 MiB. Recording stops when either quota is reached or an ENOSPC/EDQUOT storage error occurs. Inference continues, with policy, model and identity checks still enforced. If recording saturates, the logger attempts a final saturation record on normal exit; forced termination may prevent it. A cleanup failure may leave one bounded partial record. The separate `runtime_media` limit remains 512 files/128 MiB; reaching it may leave a temporal answer at the coarse stage because refinement metadata cannot be retained.
12
 
13
+ These changes address execution reliability. No new accuracy measurements accompany the repair.
14
 
15
  ## Model adaptation
16
 
17
+ The recorded upstream model is `Qwen/Qwen3.8-27B`; its shipped configuration specifies `Qwen3_5ForConditionalGeneration`. Both identifiers are retained in the inventory, alongside the exact configuration and tensor hashes. The model uses a prequantized bitsandbytes int8 vision-language base and two separate LoRA adapters, Dense48 and Point320, selected at checkpoint 12,910. Each adapter has rank 32, alpha 64 and dropout 0.05.
18
 
19
+ Point320 refers to the adapter's 320 tensors. The selected Point adapter completed 12,910 optimizer updates using restored point-target labels. It is separate from later private-data training experiments. Dense48 is the default adapter; Point320 is loaded for the configured specialist route. The adapter payloads and runtime configuration lock are verified separately, with exact identities recorded in the model and source inventories.
20
 
21
  ## Selected tensor files
22
 
 
26
  | `/app/artifacts/proc_q38_point_ck12910/adapter_model.safetensors` | 247,513,400 | `24762d93abdd35704e0087feefb5a9e4d0794d6b7cea39d0df195f93829481a7` |
27
  | `/app/artifacts/q38_base_int8_prequant/model.safetensors` | 29,924,045,158 | `238b0622e3b2446e71daeaef8289f9a81560ed26cc280f47182a0475917b6f04` |
28
 
29
+ The adapter configuration files share SHA-256 `52e9ce4678fc4bcf4e717245952665c921c13a28247c67b679eeafdb25c4f6f8`. Their configurations match, but the tensor payloads are distinct.
30
 
31
+ ## Inference pipeline
32
 
33
+ The pipeline combines question routing, frame sampling, model generation, format validation, rule answers and output formatting. Six training/annotation-derived JSON prior and stem-statistic assets accompany the learned weights; their origins are described in [DATA_PROVENANCE.md](DATA_PROVENANCE.md).
34
 
35
+ The retained 120-second temporal refinement policy uses decoded timestamps and keeps the validated coarse answer when the remaining-budget check declines refinement. Runtime records distinguish model loading, per-question generation, rule answers and fallbacks. Cache preparation, execution as the application user, input staging and the default launcher are also part of the validated pipeline; a base-model call alone does not reproduce these steps.
36
 
37
+ ## Other checkpoints and cached files
38
 
39
+ Other experimental checkpoints and runtime variants are excluded from this release. The inventory separately identifies older weights and optimizer state present in the original image; these are inactive and excluded from the selected Hugging Face assets.
40
 
41
+ The original archive remains a lineage reference. The Hugging Face package contains the selected model assets and serving source, rather than every incidental cache file or the full container filesystem.
ORGANIZER_ACCESS.md CHANGED
@@ -1,13 +1,20 @@
1
  # Organizer access
2
 
3
- Repository and clean submission-form URL: [HeyDonto/SURGFIELD-ORena-2026-PROCEDURE-FINAL](https://huggingface.co/HeyDonto/SURGFIELD-ORena-2026-PROCEDURE-FINAL).
4
 
5
- Maintainer: **HeyDonto Labs**. Corresponding contact: **Reza Nehzati, Ph.D.**
 
6
 
7
- The repository is private. The initial access setup records the dedicated resource group **ORena 2026 PROCEDURE organizer review** (id `6aa329480ec8205ac77ce835`) for this repository and read access for `orena-dkfz` through that group. Organization-wide access remains `no_access`, automatic joining is disabled, and unrelated FRAME, SEGMENT and annotation groups were unchanged in the setup comparison.
8
 
9
- This is a configured-access observation. It does not establish that an organizer has accepted an invitation, acknowledged receipt, downloaded the assets, verified their hashes or successfully loaded the model. Organizer acknowledgement has not been received at this preparation stage.
10
 
11
- The private asset commit `6a7fc05f196b98d751e3a14775d60f1161166d2a` has been created and passed a fresh full-file download/hash comparison. The initial empty repository commit is not an asset release. LOAD.md pins the actual asset commit, and later documentation updates must continue to point to it. This does not claim a recipient has downloaded or loaded the model.
12
 
13
- Access is intended for authorized challenge review. This repository excludes private surgical media, annotation banks and evaluation-bank payloads. The preserved source may retain original unit-test fixtures. Any separately required restricted-data or annotation handoff must use its own authorized channel and disclosure; this model repository does not imply that such a handoff has occurred.
 
 
 
 
 
 
 
1
  # Organizer access
2
 
3
+ Repository and submission-form URL: [HeyDonto/SURGFIELD-ORena-2026-PROCEDURE-FINAL](https://huggingface.co/HeyDonto/SURGFIELD-ORena-2026-PROCEDURE-FINAL).
4
 
5
+ **Maintainer:** HeyDonto Labs
6
+ **Responsible contact:** Reza Nehzati, Ph.D.
7
 
8
+ ## Recorded access configuration
9
 
10
+ The repository is private. The recorded setup uses a dedicated resource group, **ORena 2026 PROCEDURE organizer review** (ID `6aa329480ec8205ac77ce835`), containing only this repository. The organizer account `orena-dkfz` has Read access through that group and organization-wide role `no_access`. Automatic joining is disabled. The setup comparison confirmed that FRAME, SEGMENT and annotation permissions were unchanged.
11
 
12
+ Organizer acknowledgement of access has not been received. The configuration alone does not confirm that a recipient has downloaded, verified or loaded the model.
13
 
14
+ ## Release to download
15
+
16
+ Follow [LOAD.md](LOAD.md) for the immutable repaired source and model release at revision `0de0658c69a972ce8b92b9742f40beb62c9134f1`. The original model assets remain traceable to `6a7fc05f196b98d751e3a14775d60f1161166d2a`, which passed a full download and file-hash comparison. Later source and documentation updates preserve those model assets. File verification and model execution results are reported separately.
17
+
18
+ ## Distribution scope
19
+
20
+ Access is intended for authorized challenge review. Private surgical media, annotation datasets and per-question evaluation records are excluded. Original unit-test fixtures may remain in the preserved source. Any separately required restricted-data or annotation delivery must use its authorized channel; this model repository does not establish that such a delivery has occurred.
README.md CHANGED
@@ -11,57 +11,70 @@ tags:
11
  ---
12
  # SURGFIELD ORena 2026 PROCEDURE FINAL
13
 
14
- **HeyDonto Labs** · Corresponding contact: **Reza Nehzati, Ph.D.**
15
 
16
- This private repository contains a reliability repair of the retained PROCEDURE8d3 candidate. The int8 base, Dense48, Point320, auxiliary assets and six runtime-prior JSON files are unchanged. Five net serving files implement bounded per-clip staging and diagnostic saturation handling; [MODIFICATIONS.md](MODIFICATIONS.md) describes them. No private media or evaluation banks are distributed. Original source unit-test fixtures may remain.
17
 
18
- The current source tree has 85 files at `image_root/app/`. Model assets retain immutable custody at `6a7fc05f196b98d751e3a14775d60f1161166d2a`; [LOAD.md](LOAD.md) binds the repaired source release. The historical downloaded-asset CPU checks did not perform GPU generation. This update does not claim an HF GPU component qualification or a new accuracy score.
19
 
20
- ## Current repaired image
21
 
22
- | Identity | Value |
23
- |---|---|
24
- | Configuration | `sha256:94d0791cb96f3ac9248e9ec7c918e3bbdc5960646b48d201cb478d488fcda576` |
25
- | Registry | `us-central1-docker.pkg.dev/heydonto-425716/surgfield/surgfield-proc-q38-validation@sha256:f0590e6d79097a37f502260cf665624bd4ca87e9cbda66d240d6d4db7d0bb63d` |
26
- | Archive SHA-256 | `b63dcf1bbee6554942a78edc823dfc4b38a89dba54f029acdd28a8c91bfa4a2d` |
27
- | Archive bytes | 52,775,393,310 |
28
- | Native entrypoint | `/opt/conda/bin/python -B /app/runtime_zoom_entry.py --submission` |
29
- | Identity | UID/GID 1000; Python 3.11.11 |
30
 
31
- Current native/storage qualification is recorded separately in [RELEASE_VERIFICATION.json](RELEASE_VERIFICATION.json). The current 81-layer image consists of 79 selected 8d3 layers, one staging repair layer and one diagnostic repair layer.
 
32
 
33
- ## Final submission status
34
 
35
- On September 14, 2026, the responsible submitter confirmed that the PROCEDURE archive with SHA-256 `b63dcf1bbee6554942a78edc823dfc4b38a89dba54f029acdd28a8c91bfa4a2d` was successfully submitted for the final submission without errors. This is an owner-reported outcome; the actual submission timestamp and official score or rank were not supplied. [SUBMISSION_STATUS.json](SUBMISSION_STATUS.json) records the exact archive, image and source-revision identities.
36
 
37
- This status update preserves the model assets, serving source and immutable loading revision. The earlier [RELEASE_VERIFICATION.json](RELEASE_VERIFICATION.json) remains a historical record of the September 11 qualification and publication; its then-pending platform status is superseded by the dated owner report above. No new accuracy result or independently retrieved organizer approval is claimed.
38
 
39
- ## Method
40
 
41
- The serving path combines an int8 vision-language base with separate Dense48 and Point320 LoRA adapters at checkpoint 12,910, a question router, answer-format validation and a budget-aware temporal refinement policy. The retained temporal policy uses a 120-second refinement window and 48 sampled frames. A refinement may be declined by the original remaining-budget guard; completion does not mean every question used a second model pass. Rule answers and degraded fallbacks are distinct from observed model generation.
 
 
42
 
43
- The recorded upstream base is `Qwen/Qwen3.8-27B`; the shipped configuration architecture is `Qwen3_5ForConditionalGeneration`. Point320 denotes its 320-tensor adapter payload, not 320 optimizer steps. Exact configuration, adapter hyperparameters, auxiliary assets and upstream revisions are disclosed in the asset inventory and [MODIFICATIONS.md](MODIFICATIONS.md). Older weights retained inside the original image are listed separately from the active selected payload. They are not silently substituted into this handoff.
44
 
45
- ## Historical evaluation of original 8d3
46
 
47
- The following numbers belong to original 8d3 and were not rerun on this repair. These are local research results, not official challenge standings. [EVALUATION_SUMMARY.json](EVALUATION_SUMMARY.json) provides aggregate counts and the independent review hash without releasing per-question data.
 
 
48
 
49
  | Evaluation protocol | Selected candidate | Comparator | Processing failures retained |
50
- |---|---:|---:|---|
51
- | Historical local 1,087-question protocol | 615/1,087 correct; five-bucket mean 0.6191042 | Not a fresh native comparison | Historical protocol recorded 1,087 answered |
52
- | Fresh native default-entrypoint comparison, 1,087 questions per image | 579/1,087 correct; five-bucket mean 0.5593201 | Exact pfull best-platform variant: 444/1,087; mean 0.4653223 | Selected 67; comparator 124 |
53
- | Jointly native- and pool-qualified sensitivity population | 511/910 correct | 421/910 correct | Subset only; primary denominators remain 1,087 each |
 
 
 
 
 
 
 
 
 
 
54
 
55
- The fresh comparison used the unchanged images on H100 with 16 CPUs and 196,608 MiB host memory, and the original external process allowances of `120 + 30 × number_of_questions` seconds. The selected candidate's 67 failures comprise 28 questions in two stream-cap terminations and 39 questions in three groups excluded by the frozen validation-demotion rule. Comparator failures include nine original process-pool timeouts. Failed rows were not removed from the primary scores.
 
 
 
 
 
56
 
57
- The 1,087-question development population contains no challenge OOD rows and only one clinical-flagged row. It was used during development and selection; repeated use, correlated questions within videos and selection bias limit transfer claims. The results establish neither a clinical accuracy estimate nor a finals Copeland rank, and do not establish official platform superiority.
58
 
59
- The original 8d3 archive also completed the organizer's ten-question canonical compatibility fixture in two local input layouts. Owner-supplied platform try-out output matched those ten answer strings. Each layout showed nine question-bound VLM answers and one rule answer for the original invalid clip. This is an execution-compatibility observation, not an accuracy benchmark or independent organizer approval of this repository.
60
 
61
- A separate private, one-surgery, six-predicate diagnostic retained all 12 selected-image observations as processing/observation failures; it yielded no qualified matched accuracy comparison. Its media, labels and derivatives remain evaluation-only and are not included here.
62
 
63
- ## Data and intended use
64
 
65
- This release supports authorized challenge review and surgical-video VQA research. Dataset provenance, annotation sources and restrictions are recorded in [DATA_PROVENANCE.md](DATA_PROVENANCE.md). Private surgical media, annotation banks and evaluation-bank payloads are not distributed with the model. This research system has no established clinical safety or diagnostic performance and is not validated for patient-care decisions.
66
 
67
- Component licenses and notices are documented in [UPSTREAM_NOTICES.md](UPSTREAM_NOTICES.md) and [LICENSE_PROVENANCE.md](LICENSE_PROVENANCE.md). [ORGANIZER_ACCESS.md](ORGANIZER_ACCESS.md) separates configured repository access from organizer acknowledgement.
 
11
  ---
12
  # SURGFIELD ORena 2026 PROCEDURE FINAL
13
 
14
+ **HeyDonto Labs** · Responsible contact: **Reza Nehzati, Ph.D.**
15
 
16
+ SURGFIELD is a surgical-video question-answering system developed for the ORena 2026 PROCEDURE track. It combines a quantized vision-language model, two task-specific LoRA adapters, question routing and temporal refinement. This private repository provides the model assets and serving source corresponding to the submitted container.
17
 
18
+ On September 14, 2026, the responsible submitter confirmed successful final submission of archive `b63dcf1bbee6554942a78edc823dfc4b38a89dba54f029acdd28a8c91bfa4a2d` without errors. This is the date of that report; the actual submission timestamp and official score or rank were not supplied. [SUBMISSION_STATUS.json](SUBMISSION_STATUS.json) records the confirmation and artifact identities.
19
 
20
+ ## Method
21
 
22
+ The system uses a prequantized bitsandbytes int8 vision-language base with two separate LoRA adapters, both selected at checkpoint 12,910:
 
 
 
 
 
 
 
23
 
24
+ - **Dense48** is the default adapter.
25
+ - **Point320** serves the configured object-specialist route. Its name refers to the adapter's 320 tensors, not its number of training updates.
26
 
27
+ The recorded upstream model is `Qwen/Qwen3.8-27B`; its shipped configuration declares the architecture `Qwen3_5ForConditionalGeneration`. Exact base revisions, adapter settings and tensor hashes are recorded in [LINEAGE.json](LINEAGE.json) and [MODIFICATIONS.md](MODIFICATIONS.md).
28
 
29
+ The serving pipeline includes video preprocessing, question routing, answer-format validation and rule-based responses. Six training-derived prior and statistic files supplement the learned parameters. For eligible temporal questions, the policy can sample 48 frames within a 120-second refinement window, subject to source-validation and remaining-budget checks. A declined refinement retains the coarse answer. Runtime records distinguish model generation from rule-based and degraded responses.
30
 
31
+ ## Loading and reproduction
32
 
33
+ Follow [LOAD.md](LOAD.md) to download the private repository, verify its files and restore the required cache aliases. Use this immutable revision for the submitted model assets and repaired serving source:
34
 
35
+ ```text
36
+ 0de0658c69a972ce8b92b9742f40beb62c9134f1
37
+ ```
38
 
39
+ The repository includes 46 physical model asset files, 85 serving source files and mappings for 11 cache aliases. It includes the model configuration, tokenizer, processor and auxiliary detector assets. The inventories and [WEIGHTS_SHA256SUMS](WEIGHTS_SHA256SUMS) provide file identities.
40
 
41
+ The container is the reference runtime. The standalone HF component-loading example has passed source/API checks with CPU substitutes and separate downloaded-asset CPU checks, but has not completed a new GPU model load. Container execution tests are documented separately in [RELEASE_VERIFICATION.json](RELEASE_VERIFICATION.json).
42
 
43
+ ## Evaluation
44
+
45
+ These local development results were measured on the original selected container, identified as **8d3**. They were not rerun as a full accuracy evaluation on the submitted reliability repair. The five-bucket mean is a macro average and differs from the fraction of questions answered correctly.
46
 
47
  | Evaluation protocol | Selected candidate | Comparator | Processing failures retained |
48
+ | --- | ---: | ---: | --- |
49
+ | Historical local protocol, 1,087 questions | 615/1,087 correct; five-bucket mean 0.6191042 | No fresh native comparator in this protocol | Historical protocol recorded 1,087 answered |
50
+ | Fresh default-entrypoint comparison, 1,087 questions per image | 579/1,087 correct; five-bucket mean 0.5593201 | Exact pfull best-platform variant: 444/1,087; mean 0.4653223 | Selected: 67; comparator: 124 |
51
+ | Sensitivity analysis on jointly execution-qualified questions | 511/910 correct | 421/910 correct | Subset analysis; primary denominators remain 1,087 |
52
+
53
+ The fresh comparison used H100 GPUs, 16 CPUs and 196,608 MiB host memory, with external process allowances of `120 + 30 × number_of_questions` seconds. Selected-image failures comprised 28 questions in two stream-cap terminations and 39 questions in three groups excluded under the fixed validation-demotion rule. The comparator had nine process-pool timeouts. Failed questions remained in the primary denominators. [EVALUATION_SUMMARY.json](EVALUATION_SUMMARY.json) records the protocols, aggregate results and review reference.
54
+
55
+ The board contains no challenge OOD questions and only one clinical-flagged question. Repeated development and candidate selection on this board, together with correlated questions from the same videos, limit generalization claims. These results do not establish clinical accuracy, official platform superiority or a finals rank.
56
+
57
+ The original container also completed the organizer's ten-question canonical compatibility fixture in two local input layouts. Owner-supplied platform try-out outputs matched the ten answer strings. Each layout included nine question-bound VLM answers and one rule answer for an invalid clip. This checks execution compatibility, not accuracy. A separate evaluation-only diagnostic covering six predicates from one private surgery retained all 12 selected-image observations as failures and yielded no qualified matched accuracy comparison.
58
+
59
+ ## Submitted release
60
+
61
+ The submitted container adds two reliability repairs to the original selected model: bounded per-clip ZIP staging and diagnostic saturation handling. The base, adapters, auxiliary assets and inference policies are unchanged. [MODIFICATIONS.md](MODIFICATIONS.md) describes the five changed or added serving files and the remaining runtime limits. A full accuracy comparison of the repaired container has not been measured.
62
 
63
+ | Artifact | Identity |
64
+ | --- | --- |
65
+ | Image configuration | `sha256:94d0791cb96f3ac9248e9ec7c918e3bbdc5960646b48d201cb478d488fcda576` |
66
+ | Registry digest | `sha256:f0590e6d79097a37f502260cf665624bd4ca87e9cbda66d240d6d4db7d0bb63d` |
67
+ | Submission archive SHA-256 | `b63dcf1bbee6554942a78edc823dfc4b38a89dba54f029acdd28a8c91bfa4a2d` |
68
+ | Archive size | 52,775,393,310 bytes |
69
 
70
+ [IMAGE_IDENTITY.json](IMAGE_IDENTITY.json) and [LOAD.md](LOAD.md) provide the full registry and GCS locations, entrypoint and environment. The 81-layer container consists of the original 79 layers and two repair layers. Original model assets remain available at revision `6a7fc05f196b98d751e3a14775d60f1161166d2a`; later documentation revisions preserve those assets.
71
 
72
+ [RELEASE_VERIFICATION.json](RELEASE_VERIFICATION.json) preserves the September 11 qualification record. Its then-pending submission status is superseded by the dated submitter confirmation above; organizer results have not been independently retrieved.
73
 
74
+ ## Data, limitations and access
75
 
76
+ Training sources, derived runtime priors and gaps in historical data lineage are described in [DATA_PROVENANCE.md](DATA_PROVENANCE.md). Private surgical media, annotation datasets and per-question evaluation records are not distributed here. The preserved source may contain original unit-test fixtures. Older inactive weights in the container are identified separately and excluded from the selected HF assets.
77
 
78
+ This release is intended for authorized challenge review and surgical-video VQA research. It has no established clinical safety or diagnostic performance and is not validated for patient-care decisions.
79
 
80
+ Component attribution and applicable terms are documented in [UPSTREAM_NOTICES.md](UPSTREAM_NOTICES.md) and [LICENSE_PROVENANCE.md](LICENSE_PROVENANCE.md). The repository remains private, with organizer access described in [ORGANIZER_ACCESS.md](ORGANIZER_ACCESS.md).
UPSTREAM_NOTICES.md CHANGED
@@ -1,23 +1,25 @@
1
  # Upstream notices
2
 
3
- This private handoff is attributed to **HeyDonto Labs**; contact **Reza Nehzati, Ph.D.** Upstream authors retain their respective copyrights and licenses.
4
 
5
- ## Qwen base and selected adaptations
6
 
7
- The recorded upstream is **Qwen/Qwen3.8-27B**, revision `1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0`. Its configuration declares `Qwen3_5ForConditionalGeneration`. The pinned upstream configuration, README and LICENSE were fetched independently and matched the exact retained image's metadata bytes. The original model card identifies Apache License 2.0. See the [pinned upstream model card](https://huggingface.co/Qwen/Qwen3.8-27B/blob/1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0/README.md) and [pinned license](https://huggingface.co/Qwen/Qwen3.8-27B/blob/1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0/LICENSE).
8
 
9
- The original license is included at [image_root/app/artifacts/q38_base/LICENSE](image_root/app/artifacts/q38_base/LICENSE), SHA-256 `bbedc3fda3305820b977265f01b8619d87570a6739de3a5582c3464840f1e57a`. The model card is retained at [image_root/app/artifacts/q38_base/README.md](image_root/app/artifacts/q38_base/README.md). Team quantization and the selected Dense48/Point320 adapters are described in MODIFICATIONS.md; these changes do not replace the upstream attribution.
10
 
11
- ## OWLv2 auxiliary assets
12
 
13
- The auxiliary cache records **google/owlv2-large-patch14-ensemble**, revision `95e26936e865f87db1742128404b3c035d47d89d`. Its [pinned Google model card](https://huggingface.co/google/owlv2-large-patch14-ensemble/blob/95e26936e865f87db1742128404b3c035d47d89d/README.md) identifies Apache License 2.0 and cites the OWL-ViT/OWLv2 work. The independently retrieved card matches the cached card, SHA-256 `a2e10c3916166f08eaf2ab43ca1eb63c6116df228dea655c56e4c8e1607ecfe9`. Both authenticated cached weight formats and their metadata are retained; the alias map avoids duplicate physical uploads of a logical snapshot file.
14
 
15
- ## Runtime software and excluded cache material
16
 
17
- The reference image uses third-party libraries including PyTorch, Transformers, PEFT, bitsandbytes, Accelerate, safetensors, tokenizers, Hugging Face Hub, Pillow and PyAV. EXACT_DEPENDENCIES.json records the installed versions observed inside the image. Their applicable licenses and notices are separate from the model license; this extracted asset handoff does not redistribute the complete Python/CUDA environment or relicense those packages.
18
 
19
- The original image also contains inactive Qwen3-VL-8B, older adapter/checkpoint and optimizer payloads. They are classified in ASSET_MANIFEST.json and excluded from this selected component handoff. They are not additional active bases for the published loader.
 
 
20
 
21
  ## Data-derived assets
22
 
23
- This repository does not distribute private surgical media or evaluation banks. Its exact source may contain original unit-test fixtures, and its selected runtime assets include six disclosed training/annotation-derived prior and statistic JSON files. These carry their source provenance and restrictions; the base-model license is not a grant of rights to private data. DATA_PROVENANCE.md and LICENSE_PROVENANCE.md record that separate scope.
 
1
  # Upstream notices
2
 
3
+ Prepared by **HeyDonto Labs**. Responsible contact: **Reza Nehzati, Ph.D.** Upstream authors retain their respective copyrights and licenses.
4
 
5
+ ## Qwen base and adapters
6
 
7
+ The recorded upstream model is **Qwen/Qwen3.8-27B**, revision `1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0`. Its configuration declares `Qwen3_5ForConditionalGeneration`. The pinned upstream configuration, README and LICENSE were independently retrieved and matched the metadata in the selected image. The original model card identifies Apache License 2.0. See the [pinned model card](https://huggingface.co/Qwen/Qwen3.8-27B/blob/1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0/README.md) and [license](https://huggingface.co/Qwen/Qwen3.8-27B/blob/1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0/LICENSE).
8
 
9
+ The original [LICENSE](image_root/app/artifacts/q38_base/LICENSE) is included with SHA-256 `bbedc3fda3305820b977265f01b8619d87570a6739de3a5582c3464840f1e57a`, together with the [upstream README](image_root/app/artifacts/q38_base/README.md). [MODIFICATIONS.md](MODIFICATIONS.md) describes the team's quantization and Dense48/Point320 adaptations. Upstream attribution remains applicable to those components.
10
 
11
+ ## OWLv2 detector
12
 
13
+ The auxiliary detector is **google/owlv2-large-patch14-ensemble**, revision `95e26936e865f87db1742128404b3c035d47d89d`. Its [pinned Google model card](https://huggingface.co/google/owlv2-large-patch14-ensemble/blob/95e26936e865f87db1742128404b3c035d47d89d/README.md) identifies Apache License 2.0 and cites the OWL-ViT/OWLv2 work. The independently retrieved card matches the cached copy, SHA-256 `a2e10c3916166f08eaf2ab43ca1eb63c6116df228dea655c56e4c8e1607ecfe9`.
14
 
15
+ Both verified cached weight formats and their metadata are included. [MODEL_ALIASES.json](MODEL_ALIASES.json) records the relative cache links so that each physical file is stored once.
16
 
17
+ ## Runtime software
18
 
19
+ The reference image uses PyTorch, Transformers, PEFT, bitsandbytes, Accelerate, safetensors, tokenizers, Hugging Face Hub, Pillow and PyAV. [EXACT_DEPENDENCIES.json](EXACT_DEPENDENCIES.json) records the installed versions. Each package retains its applicable license and notices. This repository provides extracted model assets and serving source, not the complete Python/CUDA environment.
20
+
21
+ The original image also contains inactive Qwen3-VL-8B weights, older adapters/checkpoints and optimizer files. [ASSET_MANIFEST.json](ASSET_MANIFEST.json) classifies these separately; they are excluded from the selected HF assets and are not active bases in the published loader.
22
 
23
  ## Data-derived assets
24
 
25
+ Six training/annotation-derived prior and statistic JSON files are included with the runtime. Their source-specific restrictions are documented in [DATA_PROVENANCE.md](DATA_PROVENANCE.md) and [LICENSE_PROVENANCE.md](LICENSE_PROVENANCE.md). Base-model licenses do not grant rights to private datasets. Private media and per-question evaluation records are excluded; original unit-test fixtures may remain in the preserved source.