Instructions to use Kavin60606/qwen3-vl-2b-affordance-finetuned with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Kavin60606/qwen3-vl-2b-affordance-finetuned with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("unsloth/qwen3-vl-2b-instruct-unsloth-bnb-4bit") model = PeftModel.from_pretrained(base_model, "Kavin60606/qwen3-vl-2b-affordance-finetuned") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Unsloth Desktop
Qwen3-VL-2B Affordance Fine-Tuned (QLoRA)
A fine-tuned Qwen3-VL-2B-Instruct model for visual affordance detection from egocentric (first-person view) images. The model identifies what actions can be performed on visible objects in a scene.
Model Details
| Property | Value |
|---|---|
| Base Model | Qwen/Qwen3-VL-2B-Instruct |
| Method | QLoRA 4-bit (NF4 quantization) |
| LoRA Rank | 16 |
| LoRA Alpha | 32 |
| LoRA Target Modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
| Trainable Parameters | 17.4M / 1.52B (1.15%) |
| Training Framework | Unsloth + TRL SFTTrainer |
| Training Hardware | NVIDIA RTX 3060 Ti (8GB VRAM) |
| Training Time | ~64 minutes |
| PEFT Version | 0.18.1 |
Training Data
Fine-tuned on sanskxr02/fva_affordance_v1 - an egocentric visual affordance dataset containing:
| Property | Value |
|---|---|
| Total Images | 3,177 frames |
| Train Split | 2,859 (90%) |
| Validation Split | 318 (10%) |
| Image Size | 512 x 288 (16:9, resized to 384x216 during training) |
| Unique Affordances | 377 |
| Task Categories | food_preparation (coffee making), tool_use (wine bottling) |
| View | Egocentric / First-person |
Affordance Examples
pour, pull bottle, pick up bottle, press button, pour coffee into cup, grasp coffee grinder, stir, hold bottle, pack bottles, put cup on counter, etc.
Training Results
| Metric | Value |
|---|---|
| Initial Loss | 3.4227 |
| Final Train Loss | 0.0747 |
| Best Eval Loss | 0.0976 |
| Loss Reduction | 97.8% |
| Epochs | 3 |
| Total Steps | 2,145 |
Training Curves
Inference Results
How to Use
Requirements
pip install torch==2.8.0 torchvision==0.23.0 --index-url https://download.pytorch.org/whl/cu126
pip install unsloth transformers>=4.57.1 peft>=0.14.0 bitsandbytes>=0.49.1
pip install accelerate qwen-vl-utils Pillow
Quick Inference
import os
os.environ["TORCHDYNAMO_DISABLE"] = "1" # Required on Windows
from unsloth import FastVisionModel
from qwen_vl_utils import process_vision_info
# Load model with LoRA adapter
model, tokenizer = FastVisionModel.from_pretrained(
"Kavin60606/qwen3-vl-2b-affordance-finetuned",
load_in_4bit=True,
max_seq_length=2048,
)
FastVisionModel.for_inference(model)
# Prepare input
messages = [
{
"role": "system",
"content": [{"type": "text", "text": "You are a visual affordance detection assistant. Given a first-person view image, identify what actions can be performed on the visible objects."}],
},
{
"role": "user",
"content": [
{"type": "image", "image": "path/to/your/image.jpg"},
{"type": "text", "text": "What affordances are present in this scene?"},
],
},
]
# Generate
input_text = tokenizer.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True
)
image_inputs, _ = process_vision_info(messages)
inputs = tokenizer(
text=[input_text],
images=image_inputs,
return_tensors="pt",
padding=True,
).to(model.device)
output_ids = model.generate(**inputs, max_new_tokens=256)
new_tokens = output_ids[0][inputs["input_ids"].shape[1]:]
response = tokenizer.decode(new_tokens, skip_special_tokens=True)
print(response)
Using with Transformers + PEFT (without Unsloth)
import os
os.environ["TORCHDYNAMO_DISABLE"] = "1"
import torch
from transformers import AutoProcessor, Qwen3VLForConditionalGeneration
from peft import PeftModel
from qwen_vl_utils import process_vision_info
# Load base model in 4-bit
from transformers import BitsAndBytesConfig
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16,
)
base_model = Qwen3VLForConditionalGeneration.from_pretrained(
"Qwen/Qwen3-VL-2B-Instruct",
quantization_config=bnb_config,
device_map="auto",
)
# Load LoRA adapter
model = PeftModel.from_pretrained(base_model, "Kavin60606/qwen3-vl-2b-affordance-finetuned")
processor = AutoProcessor.from_pretrained("Qwen/Qwen3-VL-2B-Instruct")
# Prepare and run inference
messages = [
{
"role": "user",
"content": [
{"type": "image", "image": "path/to/image.jpg"},
{"type": "text", "text": "What affordances are present in this scene?"},
],
},
]
text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
image_inputs, _ = process_vision_info(messages)
inputs = processor(text=[text], images=image_inputs, return_tensors="pt", padding=True).to(model.device)
output = model.generate(**inputs, max_new_tokens=256)
print(processor.decode(output[0], skip_special_tokens=True))
Running on Mac (Apple Silicon M1/M2/M3/M4)
Note: Unsloth and bitsandbytes (4-bit quantization) do not support macOS. Use the standard transformers + peft approach with MPS backend instead. The model loads in float16 (~4GB), which fits on any Mac with 16GB+ unified memory.
Requirements (Mac):
pip install torch torchvision
pip install transformers>=4.57.1 peft>=0.14.0 accelerate
pip install qwen-vl-utils Pillow
# Do NOT install bitsandbytes or unsloth on Mac
Inference (Mac):
import os
os.environ["TORCHDYNAMO_DISABLE"] = "1"
import torch
from transformers import AutoProcessor, Qwen3VLForConditionalGeneration
from peft import PeftModel
from qwen_vl_utils import process_vision_info
# Load base model in float16 on MPS (Apple GPU)
base_model = Qwen3VLForConditionalGeneration.from_pretrained(
"Qwen/Qwen3-VL-2B-Instruct",
torch_dtype=torch.float16,
device_map="mps",
)
# Load LoRA adapter
model = PeftModel.from_pretrained(
base_model,
"Kavin60606/qwen3-vl-2b-affordance-finetuned",
)
processor = AutoProcessor.from_pretrained("Qwen/Qwen3-VL-2B-Instruct")
# Prepare input
messages = [
{
"role": "user",
"content": [
{"type": "image", "image": "path/to/your/image.jpg"},
{"type": "text", "text": "What affordances are present in this scene?"},
],
},
]
text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
image_inputs, _ = process_vision_info(messages)
inputs = processor(
text=[text], images=image_inputs, return_tensors="pt", padding=True
).to("mps")
output = model.generate(**inputs, max_new_tokens=256)
print(processor.decode(output[0], skip_special_tokens=True))
Platform Compatibility Summary
| Platform | Inference | Fine-Tuning | Notes |
|---|---|---|---|
| NVIDIA GPU (Linux/Windows) | Full support | Full support (QLoRA 4-bit) | Recommended. 8GB+ VRAM. |
| Apple Silicon Mac (M1-M4) | Works (MPS, float16) | LoRA only (no QLoRA) | 16GB+ unified memory recommended. No bitsandbytes/Unsloth. |
| Intel Mac | CPU only, very slow | Not practical | Not recommended. |
| CPU only (any OS) | Works but slow | Not practical | For testing only. Use device_map="cpu". |
Repository Structure
.
βββ adapter_config.json # LoRA adapter configuration
βββ adapter_model.safetensors # Fine-tuned LoRA weights (~35MB)
βββ tokenizer.json # Tokenizer files
βββ tokenizer_config.json
βββ preprocessor_config.json
βββ chat_template.jinja
βββ trainer_state.json # Full training log history
β
βββ checkpoints/
β βββ checkpoint-2000/ # Checkpoint at step 2000
β βββ checkpoint-2145/ # Final checkpoint (step 2145)
β
βββ dataset/
β βββ annotations.jsonl # Original annotations (3,177 entries)
β βββ train_msswift.jsonl # Training data - MS-SWIFT format (2,859)
β βββ val_msswift.jsonl # Validation data - MS-SWIFT format (318)
β βββ train_llamafactory.json # Training data - LLaMA-Factory format
β βββ val_llamafactory.json # Validation data - LLaMA-Factory format
β βββ dataset_info.json # LLaMA-Factory dataset registry
β
βββ results/
β βββ training_curves.png # Training loss, grad norm, LR curves
β βββ loss_by_epoch.png # Loss over epochs
β βββ inference_comparison.png # Visual GT vs prediction comparison
β βββ inference_results.json # Full inference results (20 samples)
β
βββ train.py # Training script (Unsloth + TRL)
βββ inference.py # Inference script
βββ analyze_and_infer.py # Analysis + inference + visualization
βββ prepare_finetune_data.py # Dataset preparation script
Training Configuration
# Key hyperparameters
MODEL = "unsloth/Qwen3-VL-2B-Instruct"
QUANTIZATION = "4-bit NF4"
LORA_RANK = 16
LORA_ALPHA = 32
BATCH_SIZE = 1
GRADIENT_ACCUMULATION = 4 # effective batch = 4
LEARNING_RATE = 2e-4
LR_SCHEDULER = "cosine"
WARMUP_RATIO = 0.1
EPOCHS = 3
MAX_SEQ_LENGTH = 2048
OPTIMIZER = "adamw_8bit"
PRECISION = "bfloat16"
GRADIENT_CHECKPOINTING = True
VISION_ENCODER = "frozen" # ViT frozen to save VRAM
Reproduce Training
# 1. Clone and setup
git clone https://huggingface.co/Kavin60606/qwen3-vl-2b-affordance-finetuned
cd qwen3-vl-2b-affordance-finetuned
# 2. Create environment
python -m venv finetune_env
# Linux/Mac:
source finetune_env/bin/activate
# Windows:
.\finetune_env\Scripts\Activate.ps1
# 3. Install dependencies
pip install torch==2.8.0 torchvision==0.23.0 --index-url https://download.pytorch.org/whl/cu126
pip install unsloth peft>=0.14.0 bitsandbytes>=0.49.1 accelerate>=1.2.0
pip install trl>=0.17.0 transformers>=4.57.1 datasets qwen-vl-utils Pillow matplotlib
# 4. Download the dataset images
pip install huggingface_hub
hf download sanskxr02/fva_affordance_v1 --repo-type dataset --local-dir ./dataset_images
# 5. Run training
# Windows:
$env:PYTORCH_CUDA_ALLOC_CONF="expandable_segments:True"
# Linux:
export PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
python train.py
Limitations
- Trained on only 2 task scenarios (coffee making + wine bottling) β may not generalize to other domains
- Only 3,177 training samples β limited diversity
- Original annotations generated by
qwen2-vl-2b-int4β may contain noise - Affordance labels are relatively short (avg 2.2 words)
- Vision encoder was frozen during training to fit in 8GB VRAM
Citation
If you use this model, please cite:
@misc{qwen3vl-affordance-2026,
title={Qwen3-VL-2B Fine-Tuned for Visual Affordance Detection},
author={Kavin60606},
year={2026},
publisher={HuggingFace},
url={https://huggingface.co/Kavin60606/qwen3-vl-2b-affordance-finetuned}
}
- Downloads last month
- 10


