SURGFIELD ORena 2026 PROCEDURE FINAL
HeyDonto Labs · Responsible contact: Reza Nehzati, Ph.D.
SURGFIELD is a surgical-video question-answering system developed for the ORena 2026 PROCEDURE track. It combines a quantized vision-language model, two task-specific LoRA adapters, question routing and temporal refinement. This private repository provides the model assets and serving source corresponding to the submitted container.
On September 14, 2026, the responsible submitter confirmed successful final submission of archive b63dcf1bbee6554942a78edc823dfc4b38a89dba54f029acdd28a8c91bfa4a2d without errors. This is the date of that report; the actual submission timestamp and official score or rank were not supplied. SUBMISSION_STATUS.json records the confirmation and artifact identities.
Method
The system uses a prequantized bitsandbytes int8 vision-language base with two separate LoRA adapters, both selected at checkpoint 12,910:
- Dense48 is the default adapter.
- Point320 serves the configured object-specialist route. Its name refers to the adapter's 320 tensors, not its number of training updates.
The recorded upstream model is Qwen/Qwen3.8-27B; its shipped configuration declares the architecture Qwen3_5ForConditionalGeneration. Exact base revisions, adapter settings and tensor hashes are recorded in LINEAGE.json and MODIFICATIONS.md.
The serving pipeline includes video preprocessing, question routing, answer-format validation and rule-based responses. Six training-derived prior and statistic files supplement the learned parameters. For eligible temporal questions, the policy can sample 48 frames within a 120-second refinement window, subject to source-validation and remaining-budget checks. A declined refinement retains the coarse answer. Runtime records distinguish model generation from rule-based and degraded responses.
Loading and reproduction
Follow LOAD.md to download the private repository, verify its files and restore the required cache aliases. Use this immutable revision for the submitted model assets and repaired serving source:
0de0658c69a972ce8b92b9742f40beb62c9134f1
The repository includes 46 physical model asset files, 85 serving source files and mappings for 11 cache aliases. It includes the model configuration, tokenizer, processor and auxiliary detector assets. The inventories and WEIGHTS_SHA256SUMS provide file identities.
The container is the reference runtime. The standalone HF component-loading example has passed source/API checks with CPU substitutes and separate downloaded-asset CPU checks, but has not completed a new GPU model load. Container execution tests are documented separately in RELEASE_VERIFICATION.json.
Evaluation
These local development results were measured on the original selected container, identified as 8d3. They were not rerun as a full accuracy evaluation on the submitted reliability repair. The five-bucket mean is a macro average and differs from the fraction of questions answered correctly.
| Evaluation protocol | Selected candidate | Comparator | Processing failures retained |
|---|---|---|---|
| Historical local protocol, 1,087 questions | 615/1,087 correct; five-bucket mean 0.6191042 | No fresh native comparator in this protocol | Historical protocol recorded 1,087 answered |
| Fresh default-entrypoint comparison, 1,087 questions per image | 579/1,087 correct; five-bucket mean 0.5593201 | Exact pfull best-platform variant: 444/1,087; mean 0.4653223 | Selected: 67; comparator: 124 |
| Sensitivity analysis on jointly execution-qualified questions | 511/910 correct | 421/910 correct | Subset analysis; primary denominators remain 1,087 |
The fresh comparison used H100 GPUs, 16 CPUs and 196,608 MiB host memory, with external process allowances of 120 + 30 × number_of_questions seconds. Selected-image failures comprised 28 questions in two stream-cap terminations and 39 questions in three groups excluded under the fixed validation-demotion rule. The comparator had nine process-pool timeouts. Failed questions remained in the primary denominators. EVALUATION_SUMMARY.json records the protocols, aggregate results and review reference.
The board contains no challenge OOD questions and only one clinical-flagged question. Repeated development and candidate selection on this board, together with correlated questions from the same videos, limit generalization claims. These results do not establish clinical accuracy, official platform superiority or a finals rank.
The original container also completed the organizer's ten-question canonical compatibility fixture in two local input layouts. Owner-supplied platform try-out outputs matched the ten answer strings. Each layout included nine question-bound VLM answers and one rule answer for an invalid clip. This checks execution compatibility, not accuracy. A separate evaluation-only diagnostic covering six predicates from one private surgery retained all 12 selected-image observations as failures and yielded no qualified matched accuracy comparison.
Submitted release
The submitted container adds two reliability repairs to the original selected model: bounded per-clip ZIP staging and diagnostic saturation handling. The base, adapters, auxiliary assets and inference policies are unchanged. MODIFICATIONS.md describes the five changed or added serving files and the remaining runtime limits. A full accuracy comparison of the repaired container has not been measured.
| Artifact | Identity |
|---|---|
| Image configuration | sha256:94d0791cb96f3ac9248e9ec7c918e3bbdc5960646b48d201cb478d488fcda576 |
| Registry digest | sha256:f0590e6d79097a37f502260cf665624bd4ca87e9cbda66d240d6d4db7d0bb63d |
| Submission archive SHA-256 | b63dcf1bbee6554942a78edc823dfc4b38a89dba54f029acdd28a8c91bfa4a2d |
| Archive size | 52,775,393,310 bytes |
IMAGE_IDENTITY.json and LOAD.md provide the full registry and GCS locations, entrypoint and environment. The 81-layer container consists of the original 79 layers and two repair layers. Original model assets remain available at revision 6a7fc05f196b98d751e3a14775d60f1161166d2a; later documentation revisions preserve those assets.
RELEASE_VERIFICATION.json preserves the September 11 qualification record. Its then-pending submission status is superseded by the dated submitter confirmation above; organizer results have not been independently retrieved.
Data, limitations and access
Training sources, derived runtime priors and gaps in historical data lineage are described in DATA_PROVENANCE.md. Private surgical media, annotation datasets and per-question evaluation records are not distributed here. The preserved source may contain original unit-test fixtures. Older inactive weights in the container are identified separately and excluded from the selected HF assets.
This release is intended for authorized challenge review and surgical-video VQA research. It has no established clinical safety or diagnostic performance and is not validated for patient-care decisions.
Component attribution and applicable terms are documented in UPSTREAM_NOTICES.md and LICENSE_PROVENANCE.md. The repository remains private, with organizer access described in ORGANIZER_ACCESS.md.