Download DATA_PROVENANCE.md from HeyDonto/SURGFIELD-ORena-2026-PROCEDURE-FINAL: direct link, hf CLI and curl.
- Browser
- Download file 6.79 kB
-
https://huggingface.co/HeyDonto/SURGFIELD-ORena-2026-PROCEDURE-FINAL/resolve/main/DATA_PROVENANCE.md
- Command line
-
hf download hf://HeyDonto/SURGFIELD-ORena-2026-PROCEDURE-FINAL/DATA_PROVENANCE.md
-
curl -L -o DATA_PROVENANCE.md https://huggingface.co/HeyDonto/SURGFIELD-ORena-2026-PROCEDURE-FINAL/resolve/main/DATA_PROVENANCE.md
Training data and provenance
This document describes the Dense48 and Point320 adapters used in the submitted PROCEDURE model. Other foundations, tracks and later private-data experiments have separate training histories. LINEAGE.json records the evidence for checkpoint identity, saved metadata, intended recipes and observed training inputs.
Selected adapters
| Training record | Dense48 | Point320 specialist |
|---|---|---|
| Selected checkpoint | 12,910; final-save and checkpoint tensors match | Terminal checkpoint 12,910, independently checked |
| Initialization | Checkpoint 11,709, tensor SHA-256 684922901db549df459e586e2d62f1535b3f6820cf6b62e7cf482a409e9ac60e |
Same initializer; fresh optimizer |
| Recorded recipe | 51,635 rows, eight ranks, seed 42, learning rate 0.0002, requested 48 frames, maximum side 576 | Same recorded geometry and learning rate; point-target-restored dataset described below |
| Evidence limits | Recovered tensors, configuration and saved metadata establish the selected checkpoint, but not the exact historical corpus revision or complete per-rank input/update history | Terminal, consumed-input and reload records are available; inherited exposure and annotation/decode limitations remain |
The selected Point training dataset has SHA-256 3f14484182a01b7e5e8eafad437665cc7fb7975cbeb5511296e05147412ad081: 51,635 rows over 128 source-video identities. Preparation restored 4,164 target answers from checked organizer sources: 2,089 native views and 2,075 middle views. Windows, question strings and other columns were retained. Restoring an answer does not guarantee that the corresponding event is visible in the sampled input.
The parent dataset audit traces 46,742 rows to organizer HeiCo/LapChole material. Another 363 transfer rows are traced to their paired parent records. For 4,530 jury rows, the link to the original annotation remains unresolved. Complete original-source provenance is therefore unavailable for part of the training history.
Point training recorded 103,280 consumed examples across repeated draws, not unique examples. Of these, 78,788 realized 48 frames and 24,492 had a requested/realized frame-count mismatch. Diagnostics record 36 substitutions, ten jittered successes and 174,825 rejected blank frames. These observations limit claims about training input fidelity. Filtering later datasets does not remove possible exposure inherited through checkpoint 11,709.
Sources and usage restrictions
| Source | Contribution | Applicable scope |
|---|---|---|
| HeiCo-FOCUS-VQA | Organizer colorectal VQA training material and derived views/statistics. The parent audit covers track-specific organizer training records; ancestor revisions are not fully established. | The source records CC BY-NC-SA 4.0 and a challenge-publication restriction. Historical records retain dataset-specific challenge-use decisions and annotation-delivery conditions; these do not establish new redistribution rights or completion of those obligations. |
| LapChole-FOCUS-VQA | Organizer laparoscopic cholecystectomy VQA material, derived point views and the selected organizer-derived v39 prior. | Governed by the organizer Data Usage Agreement, including restrictions on sharing, public release and reidentification. Consult that agreement for its treatment of models and data. |
| Jury and transfer-derived rows | 4,530 jury rows with an unresolved original-annotation link; 363 transfer rows traced to their paired parent records. | Raw rows and media are excluded from this repository. The unresolved history prevents a claim of fully reconstructed provenance or complete clearance. |
| Private V2 surgery and derivatives | One evaluation-only diagnostic with six predicates; no qualified matched accuracy comparison resulted. | Excluded from training, adaptation and training-rule selection. Media and reference labels are not included. |
The PROCEDURE rules and recorded method-description form require disclosure of actual data sources, including private sources. This repository does not establish completion or organizer acknowledgement of the method form, architecture figure, restricted annotation delivery or publication obligations. The linked source pages provide attribution and access terms; the pinned training records document historical usage.
Included runtime priors
Six small JSON files under image_root/app/artifacts/ supplement the learned parameters. ASSET_MANIFEST.json records their sizes and hashes.
| File | Role and provenance |
|---|---|
train_modes.json |
Training answer modes by track, capability and format. No embedded dataset revision or complete derivation record. |
x11_stem_table.json |
PROCEDURE question-stem answer modes. No embedded dataset revision or complete derivation record. |
stem_table.json |
Answer modes by track, capability, format and stem. The source requires at least five training questions and 40% support; the asset lacks a complete derivation record. |
fo_quadrant_priors_colorectal.json |
HeiCo colorectal spatial priors, attributed by the included loader source. No embedded dataset revision. |
cholec_priors_v39_organizer.json |
Organizer LapChole prior derived with the included detector. Its record describes 57 videos, 3,477 frames and 14,258 detections. This selected v39 asset is distinct from the historical CholecTrack20-derived v38 fallback. |
duration_priors.json |
HeiCo procedure-training duration constants. The constructor reads the asset, but the selected L16 applicability flag is disabled. |
These assets contain source-derived answer modes, question stems or statistics. Their recorded historical usage decisions remain provenance records; they do not establish new usage rights.
Distribution and evaluation limits
The repository excludes raw Parquet/JSONL training datasets, private surgical media, annotation datasets, per-question evaluation outputs and reference answers. The exact inference source may retain original unit-test fixtures, and the six runtime priors listed above are included.
The local 1,087-question board was repeatedly used for development and selection. It contains no challenge OOD questions and only one clinical-flagged question. Its results do not estimate clinical accuracy, untouched generalization, official platform performance or final rank. README.md and EVALUATION_SUMMARY.json report the historical and fresh native protocols separately, retaining failed questions in their stated denominators.