rag-vietnamese / README.md
thaidinhz1's picture
Add RAGAS eval results: 27-question golden set, avg 0.811
24cbd41
|
Raw History Blame
10.2 kB
metadata
title: Vietnamese RAG System
emoji: 🔍
colorFrom: blue
colorTo: green
sdk: docker
pinned: false

Hệ thống RAG cho Tài liệu Tiếng Việt

Hệ thống RAG (Retrieval-Augmented Generation) cho tài liệu tài chính tiếng Việt — hybrid BM25+vector search, cross-encoder reranking, Contextual Retrieval và ColPali visual retrieval.

Kết quả nổi bật: RAGAS eval 27 câu hỏi — Faithfulness 0.737, Answer Relevancy 0.948, Context Recall 0.748, Average 0.811; ColPali vượt trội text RAG trên trang bảng số liệu (0.90 vs 0.80).

Demo trực tuyến

🔗 Live demo: huggingface.co/spaces/thaidinhz1/rag-vietnamese

🔗 Source code: github.com/thaidinh1206/rag-vietnamese

Lưu ý: Lần đầu tải có thể mất ~60s (cold start). Nginx proxy của HF Spaces free tier buffer SSE — streaming hoạt động đúng trên local/VPS nhưng hiển thị theo batch trên HF Spaces.

Tính năng

Tính năng Chi tiết
Phân tích tài liệu 2 tầng pypdf cho PDF văn bản, EasyOCR (GPU) cho ảnh hóa đơn tiếng Việt
Tìm kiếm lai (Hybrid Search) Qdrant (vector) + BM25 (từ khóa) kết hợp bằng Reciprocal Rank Fusion
NLP tiếng Việt Tách từ underthesea cho BM25, tự động viết lại câu truy vấn
Reranking Cross-encoder bge-reranker-v2-m3 với trích dẫn nội tuyến [1][2]...
Contextual Retrieval Sinh context bằng LLM, prepend vào chunk trước khi embedding (theo paper Anthropic)
ColPali Visual Retrieval vidore/colpali-v1.2 multi-vector MaxSim — tìm kiếm trang PDF dưới dạng ảnh, không cần OCR
Đánh giá RAGAS dùng LLM-as-judge trên bộ câu hỏi báo cáo tài chính
REST API FastAPI với giao diện chat SSE streaming
Triển khai Docker trên HuggingFace Spaces với Qdrant Cloud

Kiến trúc

Pipeline A — Text RAG (production)

flowchart TD
    classDef doc     fill:#dbeafe,stroke:#3b82f6,color:#1e40af,font-weight:bold
    classDef process fill:#dcfce7,stroke:#16a34a,color:#14532d
    classDef store   fill:#fef3c7,stroke:#d97706,color:#92400e,font-weight:bold
    classDef ai      fill:#f3e8ff,stroke:#9333ea,color:#581c87,font-weight:bold
    classDef output  fill:#ccfbf1,stroke:#0d9488,color:#134e4a,font-weight:bold

    subgraph INGEST["📥 Ingest (Offline)"]
        A["Tài liệu PDF / Ảnh"] --> B["Parser\npypdf / EasyOCR"]
        B --> C["Chunker"]
        C -->|có contextual| CR["Contextual Retrieval\nLLM prepend context"]
        C -->|không có contextual| EMB
        CR --> EMB["Gemini Embedding\ngemini-embedding-001 · 3072-dim"]
        EMB --> VDB[("Qdrant Cloud\nvector store")]
        C --> BM25[("BM25 Index\nunderthesea tokenizer")]
    end

    subgraph QUERY["🔍 Query (Online)"]
        Q["Câu hỏi"] --> QR["Query Rewriter\nllama-3.3-70b"]
        QR --> QE["Embed Query"]
        QE --> VS["Vector Search"]
        QE --> BS["BM25 Search"]
        VDB --> VS
        BM25 --> BS
        VS --> RRF["RRF Fusion"]
        BS --> RRF
        RRF --> RE["Reranker\nbge-reranker-v2-m3"]
        RE --> LLM["LLM · Groq\nllama-3.3-70b-versatile"]
        LLM --> ANS["Câu trả lời + Trích dẫn"]
    end

    class A doc
    class B,C process
    class CR,EMB,QR,QE,RE,LLM ai
    class VDB,BM25 store
    class VS,BS,RRF process
    class Q doc
    class ANS output

Pipeline B — ColPali Visual Retrieval (thực nghiệm so sánh)

flowchart TD
    classDef doc     fill:#dbeafe,stroke:#3b82f6,color:#1e40af,font-weight:bold
    classDef process fill:#dcfce7,stroke:#16a34a,color:#14532d
    classDef store   fill:#fef3c7,stroke:#d97706,color:#92400e,font-weight:bold
    classDef ai      fill:#f3e8ff,stroke:#9333ea,color:#581c87,font-weight:bold
    classDef output  fill:#ccfbf1,stroke:#0d9488,color:#134e4a,font-weight:bold

    subgraph INGEST["📥 Ingest (Offline)"]
        A["Trang PDF"] --> B["Render ảnh\npypdfium2 · DPI=100"]
        B --> C["ColPali v1.2\nPaliGemma + LoRA · 128-dim patches"]
        C --> D[("Qdrant Local\nmulti-vector MaxSim")]
    end

    subgraph QUERY["🔍 Query (Online)"]
        Q["Câu hỏi"] --> QE["Encode Query\nColPali text encoder"]
        D --> S["MaxSim Search\n∑ max cosine similarity"]
        QE --> S
        S --> R["Top-k trang PDF"]
    end

    class A,Q doc
    class B,S process
    class C,QE ai
    class D store
    class R output

Công nghệ sử dụng

Thành phần Công nghệ
Embedding Google Gemini gemini-embedding-001 (3072-dim)
LLM Groq llama-3.3-70b-versatile
Vector DB Qdrant Cloud (text) + Qdrant local (ColPali)
Visual Retrieval vidore/colpali-v1.2 (PaliGemma + LoRA, multi-vector MaxSim)
Reranker BAAI/bge-reranker-v2-m3 cross-encoder
NLP tiếng Việt underthesea (tách từ)
API FastAPI + Uvicorn + SSE streaming
Triển khai Docker trên HuggingFace Spaces

Dữ liệu

  • PDF: Báo cáo tài chính doanh nghiệp Việt Nam (niêm yết HOSE, HNX, UPCOM)
  • Ảnh: Bộ dữ liệu MC-OCR — hóa đơn bán lẻ tiếng Việt (EasyOCR GPU, ingest offline)

Cài đặt nhanh (Docker)

# 1. Clone repo
git clone https://github.com/thaidinh1206/rag-vietnamese.git
cd rag-vietnamese

# 2. Tạo file .env
cp .env.example .env
# Thêm API keys: GOOGLE_API_KEY, GROQ_API_KEY

# 3. Chạy
docker-compose up

# 4. Mở trình duyệt
# http://localhost:8000

Cài đặt local

python -m venv venv
venv\Scripts\activate
pip install -r requirements.txt

Tạo file .env:

GOOGLE_API_KEY=...
GROQ_API_KEY=...
# Ingest tài liệu
python ingest.py pdf                # PDF (text pipeline)
python ingest.py pdf --contextual   # PDF với Contextual Retrieval
python ingest.py images             # Ảnh (yêu cầu GPU)
python ingest_colpali.py            # PDF pages dạng ảnh (ColPali pipeline)

# Chạy API
uvicorn api:app --reload --port 8000

# Chế độ CLI
python main.py

# Đánh giá
python evaluate.py data/golden_set_pdf.json  # RAGAS eval báo cáo tài chính (27 câu)

# So sánh hai pipeline
python compare_retrieval.py
python evaluate_colpali.py

So sánh Retrieval: ColPali vs Text RAG

Đánh giá trên 10 câu hỏi báo cáo tài chính tiếng Việt, dùng LLM-as-judge chấm điểm mức độ liên quan của nguồn tài liệu:

Chỉ số Text RAG ColPali
Source Relevance (trung bình) 0.790 0.520
Loại câu hỏi Text RAG ColPali
Bảng số liệu (doanh thu, lợi nhuận) 0.80 0.90 ✓
Tên công ty / mã chứng khoán 0.90 ✓ 0.00–0.50
So sánh nhiều kỳ 0.90 ✓ 0.20
Thông tin ngành / lĩnh vực 0.80 0.90 ✓

Nhận xét: Text RAG thắng tổng thể (0.79 vs 0.52) nhờ BM25 xử lý chính xác tên công ty và mã chứng khoán tiếng Việt. ColPali ngang bằng hoặc vượt trội trên các trang chứa bảng số liệu phức tạp — đúng với trường hợp sử dụng thiết kế của nó (nơi pypdf mất cấu trúc bảng). Kết hợp cả hai pipeline sẽ cho kết quả tối ưu.

Đánh giá RAGAS (LLM-as-judge)

Đánh giá trên 27 câu hỏi báo cáo tài chính tiếng Việt — 4 doanh nghiệp (VNM, FPT, HPG, MSN) × 2 năm (2022–2023), judge model: llama-3.3-70b-versatile:

Metric Score
Faithfulness 0.737
Answer Relevancy 0.948
Context Recall 0.748
Trung bình 0.811

Nhận xét: Answer Relevancy cao (0.948) cho thấy pipeline trả lời đúng trọng tâm câu hỏi. Faithfulness 0.737 xác nhận câu trả lời bám sát tài liệu gốc. Context Recall 0.748 phản ánh hybrid search (BM25 + vector + reranking) tìm được đúng đoạn liên quan trong phần lớn trường hợp.

Hạn chế & Hướng phát triển

Hạn chế hiện tại Hướng cải thiện
OCR nhiễu cao với PDF scan (EasyOCR) Thay bằng Gemini Vision / GPT-4o cho OCR chất lượng cao hơn
Bộ đánh giá nhỏ (27 câu) Mở rộng lên 100+ câu, bổ sung adversarial questions
ColPali chỉ chạy local (GPU + ~500KB/trang) Dùng dịch vụ embedding API hỗ trợ multi-vector
Không có conversation memory Thêm chat history, multi-turn Q&A
BM25 rebuild từ Qdrant khi cold start (~5s) Cache BM25 trên persistent volume hoặc Redis
Chưa có hybrid ColPali + Text RAG Kết hợp MaxSim score với RRF để tận dụng cả hai pipeline

Quyết định kiến trúc

Quyết định Lý do
HF Spaces thay vì Render/Railway Free tier duy nhất có 16GB RAM — đủ cho bge-reranker + Gemini embeddings. Render (512MB) và Railway (0.5GB) không thể load model.
Qdrant Cloud cho text vectors Disk HF Spaces reset khi restart — Qdrant Cloud free tier (1GB) giữ vectors qua các lần deploy.
Qdrant local cho ColPali Multi-vector ColPali ~500KB/trang — quá lớn cho Qdrant Cloud free tier. Dùng cho thực nghiệm so sánh local.
Không chạy OCR trong pipeline live EasyOCR trên CPU quá chậm (~30s/ảnh). Kết quả OCR được baked sẵn vào corpus qua offline ingestion.
Groq thay vì Gemini cho LLM Gemini free tier có rate limit nghiêm ngặt. Groq cho inference nhanh với quota thoải mái hơn.
Contextual Retrieval offline Cần 1 LLM call/chunk (~219 calls). Chạy một lần lúc ingest; query time không tốn thêm latency.
Streaming qua SSE Triển khai tại API layer (/stream). Nginx proxy HF Spaces buffer stream — hoạt động đúng trên deploy trực tiếp.