How to use from
Docker Model Runner
docker model run hf.co/minhduc168/Qwen3-VL-2B-Instruct-unsloth-bnb-4bit-Vietnamese:BF16
Quick Links

Qwen3-VL-2B-Instruct Vietnamese (4-bit)

Mô hình Qwen3-VL-2B-Instruct được fine-tune cho tác vụ trích xuất thông tin hóa đơn, phiếu thu và đơn thuốc tiếng Việt.
Model hỗ trợ hiểu hình ảnh và văn bản, phù hợp cho các bài toán OCR nâng cao, document understanding và information extraction.


🔥 Điểm nổi bật

  • ✅ Tối ưu cho tiếng Việt
  • ✅ Fine-tune cho bill / invoice / prescription extraction
  • ✅ Phiên bản 4-bit (bnb) giúp giảm VRAM khi inference
  • ✅ Có thể chuyển sang GGUF để chạy local CPU
  • ✅ Tương thích với transformers

📂 Cấu trúc Repository

  • /merged_16bit
    Chứa trọng số bnb 4-bit để chạy với thư viện transformers + bitsandbytes.

  • /gguf
    Phiên bản GGUF dành cho llama.cpp hoặc các engine suy luận local.

    Bao gồm:

    • Qwen3-VL-2B-Instruct-Vietnamese.Q4_K_M.gguf — bản nén 4-bit chất lượng cao
    • Qwen3-VL-2B-Instruct-Vietnamese.mmproj.gguf — file projector xử lý hình ảnh

🚀 Hướng dẫn sử dụng

✅ Với Transformers

from transformers import Qwen2VLForConditionalGeneration, AutoProcessor

model = Qwen2VLForConditionalGeneration.from_pretrained(
    "minhduc168/Qwen3-VL-2B-Instruct-Vietnamese",
    device_map="auto"
)

processor = AutoProcessor.from_pretrained(
    "minhduc168/Qwen3-VL-2B-Instruct-Vietnamese"
)

⚠️ Lưu ý quan trọng khi dùng GGUF (Vision Model)

Đối với các model Vision-Language như Qwen3-VL, khi chuyển sang GGUF:

Bắt buộc cần 2 file:

1️⃣ Model chính (.gguf)
2️⃣ Projector (mmproj.gguf)

👉 Thiếu file projector → model không thể xử lý hình ảnh.


📊 Dataset

Model được huấn luyện trên:minhduc168/dataset-qwen-vlm-extract-bill

Bao gồm:

  • Hóa đơn bán lẻ
  • Phiếu thu
  • Đơn thuốc
  • Chứng từ tiếng Việt

Định dạng instruction-following giúp model trích xuất dữ liệu có cấu trúc chính xác hơn.


🎯 Use Cases

  • Trích xuất thông tin hóa đơn tự động
  • Structured OCR
  • Document AI tiếng Việt
  • Medical / pharmacy bill parsing
  • Fintech document processing

📌 Gợi ý phần cứng

Quantization VRAM đề xuất
4-bit bnb ~6–8GB
GGUF Q4 Chạy được trên CPU (khuyến nghị ≥16GB RAM)

License

Apache-2.0

Downloads last month
7
GGUF
Model size
2B params
Architecture
qwen3vl
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for minhduc168/Qwen3-VL-2B-Instruct-unsloth-bnb-4bit-Vietnamese

Quantized
(3)
this model

Dataset used to train minhduc168/Qwen3-VL-2B-Instruct-unsloth-bnb-4bit-Vietnamese