Video-Text-to-Text
PEFT
Safetensors
English
surgical-video
video-question-answering
long-video-understanding
medical
lora
qwen3.5
miccai-2026
Instructions to use wxyi088/orena-SurgScope with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use wxyi088/orena-SurgScope with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.5-9B") model = PeftModel.from_pretrained(base_model, "wxyi088/orena-SurgScope") - Notebooks
- Google Colab
- Kaggle
Download surgscope_recipe.json from wxyi088/orena-SurgScope: direct link, hf CLI and curl.
- Browser
- Download file 1.81 kB
-
https://huggingface.co/wxyi088/orena-SurgScope/resolve/main/surgscope_recipe.json
- Command line
-
hf download hf://wxyi088/orena-SurgScope/surgscope_recipe.json
-
curl -L -o surgscope_recipe.json https://huggingface.co/wxyi088/orena-SurgScope/resolve/main/surgscope_recipe.json
1.81 kB
| { | |
| "base_model": "Qwen/Qwen3.5-9B", | |
| "adapter": {"type": "LoRA", "rank": 64, "alpha": 128, "dropout": 0.05, | |
| "target_modules": "all linear layers of vision encoder, patch merger and language model"}, | |
| "weights": "equal-weight fp32 mean of LoRA A and B of three adapters trained under three frame-sampling regimes (same base, seed and LoRA shape); merged into the base model in bf16 at load time", | |
| "regimes": { | |
| "r768": {"frames": "768 for every question", "tokens_per_frame_pair": "<= 128", | |
| "lr": 1.4e-4, "global_batch": 64, "max_length": 57344, "checkpoint_epoch": 4}, | |
| "dense-mix": {"frames": "1536 time / counting / aggregation, 1152 otherwise", "tokens_per_frame_pair": "60 / 84", | |
| "lr": 1e-4, "global_batch": 32, "max_length": 57344, "checkpoint_epoch": 14}, | |
| "dense-mix-2": {"frames": "2880 time / counting / aggregation, 1512 otherwise", "tokens_per_frame_pair": "32 / 64", | |
| "lr": 1e-4, "global_batch": 32, "max_length": 65536, "checkpoint_epoch": 6} | |
| }, | |
| "optimisation": {"optimizer": "AdamW", "weight_decay": 0.1, "adam_beta2": 0.95, "schedule": "cosine", | |
| "warmup_ratio": 0.03, "epochs": 15, "precision": "bf16", "seed": 42, "framework": "ms-swift 4.3.2"}, | |
| "training_data": "official PROCEDURE train split of HeiCo-FOCUS-VQA and LapChole-FOCUS-VQA (6,873 question-answer pairs), no external data", | |
| "inference": {"windows": "19 rules on the question text", "regime": "dense-mix", | |
| "refinement": "timestamp answers re-asked on +-10 min and +-2 min windows", | |
| "env": {"VIDEO_MAX_TOKEN_NUM": "128", "VIDEO_MIN_TOKEN_NUM": "32", "FORCE_QWENVL_VIDEO_READER": "torchvision"}, | |
| "decoding": "greedy, max 64 new tokens, thinking disabled"} | |
| } | |