SURGFIELD ORena 2026 SEGMENT FINAL
HeyDonto Labs · Responsible contact: Reza Nehzati, Ph.D. (Hugging Face profile)
SURGFIELD answers questions about surgical video segments for the ORena 2026 SEGMENT challenge. This repository contains the exact model components, configurations, tokenizer and processor assets, loading tools and documentation associated with the submitted R4 container.
HeyDonto Labs confirmed on September 14, 2026 that R4 was successfully submitted for SEGMENT finals, with no errors reported. The submitted archive SHA-256 is e477d307ff2e09f469d83569ec1ce8703daeef34edbba633a3f240eab955de00. September 14 is the confirmation date; the submission timestamp, platform identifiers, terminal logs and final score have not been independently recorded here. See submission status and image identity.
Method
The serving application combines visual models with fixed answer-format rules and six inherited prior tables. It preserves the original source-video clock when interpreting questions and sampling frames.
| Component | Role and configuration |
|---|---|
| Main visual-language model | Private N3 Qwen3-VL-8B backbone with the selected step-900 continuation adapter. The backbone uses BF16; the FP32 LoRA adapter remains unmerged. |
| Temporal model | A separate, pinned public Qwen3-VL-8B backbone with the selected step-600 temporal adapter, applied and merged once. |
| Object detector and priors | Pinned OWLv2 model, six prior tables and fixed routing/formatting rules. |
The main path samples up to 16 frames; the temporal path supports up to 128 native frames with a maximum image side of 768 pixels. Generation is greedy, with at most 64 new tokens. The private N3 and public Qwen backbones are distinct, and both are required. Training and component lineage describes their construction and exposure.
Download and use
Follow LOAD.md to download the 50 model and prior files from immutable revisions, verify their checksums and optionally load the components. MODEL_FILES.json and WEIGHTS_SHA256SUMS identify the files.
The component-loading example generates no answers. Reproducing the complete application requires the corresponding Docker image, its frozen environment, routing and preprocessing. See platform compatibility for the entrypoint, recovery behavior and measured runtime coverage.
Evaluation
A development comparison used a fixed Qwen3.5-27B judge, 20-question batches and 4,000 questions from ten original source videos. These measurements cover registered implementations; a full 4,000-question evaluation of the submitted R4 Docker archive has not been performed.
| Registered serving implementation | Correct | Local micro accuracy |
|---|---|---|
| Development candidate (registered implementation) | 2,373 / 4,000 | 59.325% |
| Reference derived from the historical 333 implementation | 1,940 / 4,000 | 48.500% |
The evaluation set was reused during development and has inherited data-exposure limitations. The comparison does not reproduce the original 333 Docker image, establish superiority over every historical candidate, or predict the official finals score. Causal-consequence questions regressed by 17.742 percentage points on 62 cases. EVALUATION.md reports the paired sensitivity analysis, scorer limitations and runtime evidence separately.
Data, access and limitations
Training includes organizer data and inherited derived annotations. The separately retained 13,177-row annotation packet excludes organizer rows and video pixels; its delivery and access are separate from this model repository. See lineage and annotation scope.
The repository remains private under the agreed publication hold. Organizer access is deferred; ORGANIZER_ACCESS.md records the access and delivery requirements. Upstream model licenses and dataset restrictions remain component-specific; NOTICES.md defines their scope.
This is a research and challenge system without clinical deployment qualification. Some inputs can use rule-based answers or explicit empty answers after a recoverable per-question failure. Runtime completion therefore does not establish neural coverage or answer correctness.
Model tree for HeyDonto/SURGFIELD-ORena-2026-SEGMENT-FINAL
Base model
Qwen/Qwen3-VL-8B-Instruct