Qwen3-VL-2B Affordance Fine-Tuned (QLoRA)

A fine-tuned Qwen3-VL-2B-Instruct model for visual affordance detection from egocentric (first-person view) images. The model identifies what actions can be performed on visible objects in a scene.

Model Details

Property Value
Base Model Qwen/Qwen3-VL-2B-Instruct
Method QLoRA 4-bit (NF4 quantization)
LoRA Rank 16
LoRA Alpha 32
LoRA Target Modules q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
Trainable Parameters 17.4M / 1.52B (1.15%)
Training Framework Unsloth + TRL SFTTrainer
Training Hardware NVIDIA RTX 3060 Ti (8GB VRAM)
Training Time ~64 minutes
PEFT Version 0.18.1

Training Data

Fine-tuned on sanskxr02/fva_affordance_v1 - an egocentric visual affordance dataset containing:

Property Value
Total Images 3,177 frames
Train Split 2,859 (90%)
Validation Split 318 (10%)
Image Size 512 x 288 (16:9, resized to 384x216 during training)
Unique Affordances 377
Task Categories food_preparation (coffee making), tool_use (wine bottling)
View Egocentric / First-person

Affordance Examples

pour, pull bottle, pick up bottle, press button, pour coffee into cup, grasp coffee grinder, stir, hold bottle, pack bottles, put cup on counter, etc.

Training Results

Metric Value
Initial Loss 3.4227
Final Train Loss 0.0747
Best Eval Loss 0.0976
Loss Reduction 97.8%
Epochs 3
Total Steps 2,145

Training Curves

Training Curves

Loss by Epoch

Inference Results

Inference Comparison

How to Use

Requirements

pip install torch==2.8.0 torchvision==0.23.0 --index-url https://download.pytorch.org/whl/cu126
pip install unsloth transformers>=4.57.1 peft>=0.14.0 bitsandbytes>=0.49.1
pip install accelerate qwen-vl-utils Pillow

Quick Inference

import os
os.environ["TORCHDYNAMO_DISABLE"] = "1"  # Required on Windows

from unsloth import FastVisionModel
from qwen_vl_utils import process_vision_info

# Load model with LoRA adapter
model, tokenizer = FastVisionModel.from_pretrained(
    "Kavin60606/qwen3-vl-2b-affordance-finetuned",
    load_in_4bit=True,
    max_seq_length=2048,
)
FastVisionModel.for_inference(model)

# Prepare input
messages = [
    {
        "role": "system",
        "content": [{"type": "text", "text": "You are a visual affordance detection assistant. Given a first-person view image, identify what actions can be performed on the visible objects."}],
    },
    {
        "role": "user",
        "content": [
            {"type": "image", "image": "path/to/your/image.jpg"},
            {"type": "text", "text": "What affordances are present in this scene?"},
        ],
    },
]

# Generate
input_text = tokenizer.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True
)
image_inputs, _ = process_vision_info(messages)
inputs = tokenizer(
    text=[input_text],
    images=image_inputs,
    return_tensors="pt",
    padding=True,
).to(model.device)

output_ids = model.generate(**inputs, max_new_tokens=256)
new_tokens = output_ids[0][inputs["input_ids"].shape[1]:]
response = tokenizer.decode(new_tokens, skip_special_tokens=True)
print(response)

Using with Transformers + PEFT (without Unsloth)

import os
os.environ["TORCHDYNAMO_DISABLE"] = "1"

import torch
from transformers import AutoProcessor, Qwen3VLForConditionalGeneration
from peft import PeftModel
from qwen_vl_utils import process_vision_info

# Load base model in 4-bit
from transformers import BitsAndBytesConfig

bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.bfloat16,
)

base_model = Qwen3VLForConditionalGeneration.from_pretrained(
    "Qwen/Qwen3-VL-2B-Instruct",
    quantization_config=bnb_config,
    device_map="auto",
)

# Load LoRA adapter
model = PeftModel.from_pretrained(base_model, "Kavin60606/qwen3-vl-2b-affordance-finetuned")
processor = AutoProcessor.from_pretrained("Qwen/Qwen3-VL-2B-Instruct")

# Prepare and run inference
messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "image": "path/to/image.jpg"},
            {"type": "text", "text": "What affordances are present in this scene?"},
        ],
    },
]

text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
image_inputs, _ = process_vision_info(messages)
inputs = processor(text=[text], images=image_inputs, return_tensors="pt", padding=True).to(model.device)

output = model.generate(**inputs, max_new_tokens=256)
print(processor.decode(output[0], skip_special_tokens=True))

Running on Mac (Apple Silicon M1/M2/M3/M4)

Note: Unsloth and bitsandbytes (4-bit quantization) do not support macOS. Use the standard transformers + peft approach with MPS backend instead. The model loads in float16 (~4GB), which fits on any Mac with 16GB+ unified memory.

Requirements (Mac):

pip install torch torchvision
pip install transformers>=4.57.1 peft>=0.14.0 accelerate
pip install qwen-vl-utils Pillow
# Do NOT install bitsandbytes or unsloth on Mac

Inference (Mac):

import os
os.environ["TORCHDYNAMO_DISABLE"] = "1"

import torch
from transformers import AutoProcessor, Qwen3VLForConditionalGeneration
from peft import PeftModel
from qwen_vl_utils import process_vision_info

# Load base model in float16 on MPS (Apple GPU)
base_model = Qwen3VLForConditionalGeneration.from_pretrained(
    "Qwen/Qwen3-VL-2B-Instruct",
    torch_dtype=torch.float16,
    device_map="mps",
)

# Load LoRA adapter
model = PeftModel.from_pretrained(
    base_model,
    "Kavin60606/qwen3-vl-2b-affordance-finetuned",
)
processor = AutoProcessor.from_pretrained("Qwen/Qwen3-VL-2B-Instruct")

# Prepare input
messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "image": "path/to/your/image.jpg"},
            {"type": "text", "text": "What affordances are present in this scene?"},
        ],
    },
]

text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
image_inputs, _ = process_vision_info(messages)
inputs = processor(
    text=[text], images=image_inputs, return_tensors="pt", padding=True
).to("mps")

output = model.generate(**inputs, max_new_tokens=256)
print(processor.decode(output[0], skip_special_tokens=True))

Platform Compatibility Summary

Platform Inference Fine-Tuning Notes
NVIDIA GPU (Linux/Windows) Full support Full support (QLoRA 4-bit) Recommended. 8GB+ VRAM.
Apple Silicon Mac (M1-M4) Works (MPS, float16) LoRA only (no QLoRA) 16GB+ unified memory recommended. No bitsandbytes/Unsloth.
Intel Mac CPU only, very slow Not practical Not recommended.
CPU only (any OS) Works but slow Not practical For testing only. Use device_map="cpu".

Repository Structure

.
β”œβ”€β”€ adapter_config.json              # LoRA adapter configuration
β”œβ”€β”€ adapter_model.safetensors        # Fine-tuned LoRA weights (~35MB)
β”œβ”€β”€ tokenizer.json                   # Tokenizer files
β”œβ”€β”€ tokenizer_config.json
β”œβ”€β”€ preprocessor_config.json
β”œβ”€β”€ chat_template.jinja
β”œβ”€β”€ trainer_state.json               # Full training log history
β”‚
β”œβ”€β”€ checkpoints/
β”‚   β”œβ”€β”€ checkpoint-2000/             # Checkpoint at step 2000
β”‚   └── checkpoint-2145/             # Final checkpoint (step 2145)
β”‚
β”œβ”€β”€ dataset/
β”‚   β”œβ”€β”€ annotations.jsonl            # Original annotations (3,177 entries)
β”‚   β”œβ”€β”€ train_msswift.jsonl          # Training data - MS-SWIFT format (2,859)
β”‚   β”œβ”€β”€ val_msswift.jsonl            # Validation data - MS-SWIFT format (318)
β”‚   β”œβ”€β”€ train_llamafactory.json      # Training data - LLaMA-Factory format
β”‚   β”œβ”€β”€ val_llamafactory.json        # Validation data - LLaMA-Factory format
β”‚   └── dataset_info.json            # LLaMA-Factory dataset registry
β”‚
β”œβ”€β”€ results/
β”‚   β”œβ”€β”€ training_curves.png          # Training loss, grad norm, LR curves
β”‚   β”œβ”€β”€ loss_by_epoch.png            # Loss over epochs
β”‚   β”œβ”€β”€ inference_comparison.png     # Visual GT vs prediction comparison
β”‚   └── inference_results.json       # Full inference results (20 samples)
β”‚
β”œβ”€β”€ train.py                         # Training script (Unsloth + TRL)
β”œβ”€β”€ inference.py                     # Inference script
β”œβ”€β”€ analyze_and_infer.py             # Analysis + inference + visualization
└── prepare_finetune_data.py         # Dataset preparation script

Training Configuration

# Key hyperparameters
MODEL = "unsloth/Qwen3-VL-2B-Instruct"
QUANTIZATION = "4-bit NF4"
LORA_RANK = 16
LORA_ALPHA = 32
BATCH_SIZE = 1
GRADIENT_ACCUMULATION = 4  # effective batch = 4
LEARNING_RATE = 2e-4
LR_SCHEDULER = "cosine"
WARMUP_RATIO = 0.1
EPOCHS = 3
MAX_SEQ_LENGTH = 2048
OPTIMIZER = "adamw_8bit"
PRECISION = "bfloat16"
GRADIENT_CHECKPOINTING = True
VISION_ENCODER = "frozen"  # ViT frozen to save VRAM

Reproduce Training

# 1. Clone and setup
git clone https://huggingface.co/Kavin60606/qwen3-vl-2b-affordance-finetuned
cd qwen3-vl-2b-affordance-finetuned

# 2. Create environment
python -m venv finetune_env
# Linux/Mac:
source finetune_env/bin/activate
# Windows:
.\finetune_env\Scripts\Activate.ps1

# 3. Install dependencies
pip install torch==2.8.0 torchvision==0.23.0 --index-url https://download.pytorch.org/whl/cu126
pip install unsloth peft>=0.14.0 bitsandbytes>=0.49.1 accelerate>=1.2.0
pip install trl>=0.17.0 transformers>=4.57.1 datasets qwen-vl-utils Pillow matplotlib

# 4. Download the dataset images
pip install huggingface_hub
hf download sanskxr02/fva_affordance_v1 --repo-type dataset --local-dir ./dataset_images

# 5. Run training
# Windows:
$env:PYTORCH_CUDA_ALLOC_CONF="expandable_segments:True"
# Linux:
export PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True

python train.py

Limitations

  • Trained on only 2 task scenarios (coffee making + wine bottling) β€” may not generalize to other domains
  • Only 3,177 training samples β€” limited diversity
  • Original annotations generated by qwen2-vl-2b-int4 β€” may contain noise
  • Affordance labels are relatively short (avg 2.2 words)
  • Vision encoder was frozen during training to fit in 8GB VRAM

Citation

If you use this model, please cite:

@misc{qwen3vl-affordance-2026,
  title={Qwen3-VL-2B Fine-Tuned for Visual Affordance Detection},
  author={Kavin60606},
  year={2026},
  publisher={HuggingFace},
  url={https://huggingface.co/Kavin60606/qwen3-vl-2b-affordance-finetuned}
}
Downloads last month
10
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support