# Model and runtime modifications Prepared by **HeyDonto Labs** for `HeyDonto/SURGFIELD-ORena-2026-PROCEDURE-FINAL`. Contact: **Reza Nehzati, Ph.D.** The serving source matches the validated repaired image `sha256:94d0791cb96f3ac9248e9ec7c918e3bbdc5960646b48d201cb478d488fcda576` byte-for-byte. Original 8d3 is preserved at immutable revision `b2176cf6d0f77e58e34a90583041b1a929417fb3`. The repair leaves model weights, adapter configuration, routing, prompts, frame sampling, temporal windows and inference clocks unchanged. ## Reliability repairs Five serving files changed: `VEHICLE_INPUTS.json`, `solution/entrypoint/probe_entrypoint.py`, the new `solution/entrypoint/lazy_video_staging.py`, `solution/shape_mode_diagnostic.py` and its policy. ZIP metadata is indexed before serving. Requested clips are extracted one at a time and released after request processing and video-loader cleanup; the process removes its own staging directory. Input-root aliases remain supported, while unsafe descendant links and archive paths are rejected. Diagnostics are limited to 16,384 events and 16 MiB. Recording stops when either quota is reached or an ENOSPC/EDQUOT storage error occurs. Inference continues, with policy, model and identity checks still enforced. If recording saturates, the logger attempts a final saturation record on normal exit; forced termination may prevent it. A cleanup failure may leave one bounded partial record. The separate `runtime_media` limit remains 512 files/128 MiB; reaching it may leave a temporal answer at the coarse stage because refinement metadata cannot be retained. These changes address execution reliability. No new accuracy measurements accompany the repair. ## Model adaptation The recorded upstream model is `Qwen/Qwen3.8-27B`; its shipped configuration specifies `Qwen3_5ForConditionalGeneration`. Both identifiers are retained in the inventory, alongside the exact configuration and tensor hashes. The model uses a prequantized bitsandbytes int8 vision-language base and two separate LoRA adapters, Dense48 and Point320, selected at checkpoint 12,910. Each adapter has rank 32, alpha 64 and dropout 0.05. Point320 refers to the adapter's 320 tensors. The selected Point adapter completed 12,910 optimizer updates using restored point-target labels. It is separate from later private-data training experiments. Dense48 is the default adapter; Point320 is loaded for the configured specialist route. The adapter payloads and runtime configuration lock are verified separately, with exact identities recorded in the model and source inventories. ## Selected tensor files | Original image path | Bytes | SHA-256 | |---|---:|---| | `/app/artifacts/proc_q38_dense48_ck12910/adapter_model.safetensors` | 247,513,400 | `500126c5a3a7e813483929225bf2321c82f8a5b59c32ed8aac268db34ab7c4fe` | | `/app/artifacts/proc_q38_point_ck12910/adapter_model.safetensors` | 247,513,400 | `24762d93abdd35704e0087feefb5a9e4d0794d6b7cea39d0df195f93829481a7` | | `/app/artifacts/q38_base_int8_prequant/model.safetensors` | 29,924,045,158 | `238b0622e3b2446e71daeaef8289f9a81560ed26cc280f47182a0475917b6f04` | The adapter configuration files share SHA-256 `52e9ce4678fc4bcf4e717245952665c921c13a28247c67b679eeafdb25c4f6f8`. Their configurations match, but the tensor payloads are distinct. ## Inference pipeline The pipeline combines question routing, frame sampling, model generation, format validation, rule answers and output formatting. Six training/annotation-derived JSON prior and stem-statistic assets accompany the learned weights; their origins are described in [DATA_PROVENANCE.md](DATA_PROVENANCE.md). The retained 120-second temporal refinement policy uses decoded timestamps and keeps the validated coarse answer when the remaining-budget check declines refinement. Runtime records distinguish model loading, per-question generation, rule answers and fallbacks. Cache preparation, execution as the application user, input staging and the default launcher are also part of the validated pipeline; a base-model call alone does not reproduce these steps. ## Other checkpoints and cached files Other experimental checkpoints and runtime variants are excluded from this release. The inventory separately identifies older weights and optimizer state present in the original image; these are inactive and excluded from the selected Hugging Face assets. The original archive remains a lineage reference. The Hugging Face package contains the selected model assets and serving source, rather than every incidental cache file or the full container filesystem.