rag-vietnamese / README.md
thaidinhz1's picture
Add RAGAS eval results: 27-question golden set, avg 0.811
24cbd41
|
Raw History Blame
10.2 kB
---
title: Vietnamese RAG System
emoji: 🔍
colorFrom: blue
colorTo: green
sdk: docker
pinned: false
---
# Hệ thống RAG cho Tài liệu Tiếng Việt
Hệ thống RAG (Retrieval-Augmented Generation) cho tài liệu tài chính tiếng Việt — hybrid BM25+vector search, cross-encoder reranking, Contextual Retrieval và ColPali visual retrieval.
**Kết quả nổi bật:** RAGAS eval 27 câu hỏi — Faithfulness **0.737**, Answer Relevancy **0.948**, Context Recall **0.748**, Average **0.811**; ColPali vượt trội text RAG trên trang bảng số liệu (0.90 vs 0.80).
## Demo trực tuyến
🔗 **Live demo:** [huggingface.co/spaces/thaidinhz1/rag-vietnamese](https://huggingface.co/spaces/thaidinhz1/rag-vietnamese)
🔗 **Source code:** [github.com/thaidinh1206/rag-vietnamese](https://github.com/thaidinh1206/rag-vietnamese)
> **Lưu ý:** Lần đầu tải có thể mất ~60s (cold start). Nginx proxy của HF Spaces free tier buffer SSE — streaming hoạt động đúng trên local/VPS nhưng hiển thị theo batch trên HF Spaces.
## Tính năng
| Tính năng | Chi tiết |
|-----------|---------|
| **Phân tích tài liệu 2 tầng** | pypdf cho PDF văn bản, EasyOCR (GPU) cho ảnh hóa đơn tiếng Việt |
| **Tìm kiếm lai (Hybrid Search)** | Qdrant (vector) + BM25 (từ khóa) kết hợp bằng Reciprocal Rank Fusion |
| **NLP tiếng Việt** | Tách từ underthesea cho BM25, tự động viết lại câu truy vấn |
| **Reranking** | Cross-encoder `bge-reranker-v2-m3` với trích dẫn nội tuyến `[1][2]...` |
| **Contextual Retrieval** | Sinh context bằng LLM, prepend vào chunk trước khi embedding (theo paper Anthropic) |
| **ColPali Visual Retrieval** | `vidore/colpali-v1.2` multi-vector MaxSim — tìm kiếm trang PDF dưới dạng ảnh, không cần OCR |
| **Đánh giá** | RAGAS dùng LLM-as-judge trên bộ câu hỏi báo cáo tài chính |
| **REST API** | FastAPI với giao diện chat SSE streaming |
| **Triển khai** | Docker trên HuggingFace Spaces với Qdrant Cloud |
## Kiến trúc
### Pipeline A — Text RAG (production)
```mermaid
flowchart TD
classDef doc fill:#dbeafe,stroke:#3b82f6,color:#1e40af,font-weight:bold
classDef process fill:#dcfce7,stroke:#16a34a,color:#14532d
classDef store fill:#fef3c7,stroke:#d97706,color:#92400e,font-weight:bold
classDef ai fill:#f3e8ff,stroke:#9333ea,color:#581c87,font-weight:bold
classDef output fill:#ccfbf1,stroke:#0d9488,color:#134e4a,font-weight:bold
subgraph INGEST["📥 Ingest (Offline)"]
A["Tài liệu PDF / Ảnh"] --> B["Parser\npypdf / EasyOCR"]
B --> C["Chunker"]
C -->|có contextual| CR["Contextual Retrieval\nLLM prepend context"]
C -->|không có contextual| EMB
CR --> EMB["Gemini Embedding\ngemini-embedding-001 · 3072-dim"]
EMB --> VDB[("Qdrant Cloud\nvector store")]
C --> BM25[("BM25 Index\nunderthesea tokenizer")]
end
subgraph QUERY["🔍 Query (Online)"]
Q["Câu hỏi"] --> QR["Query Rewriter\nllama-3.3-70b"]
QR --> QE["Embed Query"]
QE --> VS["Vector Search"]
QE --> BS["BM25 Search"]
VDB --> VS
BM25 --> BS
VS --> RRF["RRF Fusion"]
BS --> RRF
RRF --> RE["Reranker\nbge-reranker-v2-m3"]
RE --> LLM["LLM · Groq\nllama-3.3-70b-versatile"]
LLM --> ANS["Câu trả lời + Trích dẫn"]
end
class A doc
class B,C process
class CR,EMB,QR,QE,RE,LLM ai
class VDB,BM25 store
class VS,BS,RRF process
class Q doc
class ANS output
```
### Pipeline B — ColPali Visual Retrieval (thực nghiệm so sánh)
```mermaid
flowchart TD
classDef doc fill:#dbeafe,stroke:#3b82f6,color:#1e40af,font-weight:bold
classDef process fill:#dcfce7,stroke:#16a34a,color:#14532d
classDef store fill:#fef3c7,stroke:#d97706,color:#92400e,font-weight:bold
classDef ai fill:#f3e8ff,stroke:#9333ea,color:#581c87,font-weight:bold
classDef output fill:#ccfbf1,stroke:#0d9488,color:#134e4a,font-weight:bold
subgraph INGEST["📥 Ingest (Offline)"]
A["Trang PDF"] --> B["Render ảnh\npypdfium2 · DPI=100"]
B --> C["ColPali v1.2\nPaliGemma + LoRA · 128-dim patches"]
C --> D[("Qdrant Local\nmulti-vector MaxSim")]
end
subgraph QUERY["🔍 Query (Online)"]
Q["Câu hỏi"] --> QE["Encode Query\nColPali text encoder"]
D --> S["MaxSim Search\n∑ max cosine similarity"]
QE --> S
S --> R["Top-k trang PDF"]
end
class A,Q doc
class B,S process
class C,QE ai
class D store
class R output
```
## Công nghệ sử dụng
| Thành phần | Công nghệ |
|-----------|-----------|
| **Embedding** | Google Gemini `gemini-embedding-001` (3072-dim) |
| **LLM** | Groq `llama-3.3-70b-versatile` |
| **Vector DB** | Qdrant Cloud (text) + Qdrant local (ColPali) |
| **Visual Retrieval** | `vidore/colpali-v1.2` (PaliGemma + LoRA, multi-vector MaxSim) |
| **Reranker** | `BAAI/bge-reranker-v2-m3` cross-encoder |
| **NLP tiếng Việt** | underthesea (tách từ) |
| **API** | FastAPI + Uvicorn + SSE streaming |
| **Triển khai** | Docker trên HuggingFace Spaces |
## Dữ liệu
- **PDF**: Báo cáo tài chính doanh nghiệp Việt Nam (niêm yết HOSE, HNX, UPCOM)
- **Ảnh**: Bộ dữ liệu MC-OCR — hóa đơn bán lẻ tiếng Việt (EasyOCR GPU, ingest offline)
## Cài đặt nhanh (Docker)
```bash
# 1. Clone repo
git clone https://github.com/thaidinh1206/rag-vietnamese.git
cd rag-vietnamese
# 2. Tạo file .env
cp .env.example .env
# Thêm API keys: GOOGLE_API_KEY, GROQ_API_KEY
# 3. Chạy
docker-compose up
# 4. Mở trình duyệt
# http://localhost:8000
```
## Cài đặt local
```bash
python -m venv venv
venv\Scripts\activate
pip install -r requirements.txt
```
Tạo file `.env`:
```
GOOGLE_API_KEY=...
GROQ_API_KEY=...
```
```bash
# Ingest tài liệu
python ingest.py pdf # PDF (text pipeline)
python ingest.py pdf --contextual # PDF với Contextual Retrieval
python ingest.py images # Ảnh (yêu cầu GPU)
python ingest_colpali.py # PDF pages dạng ảnh (ColPali pipeline)
# Chạy API
uvicorn api:app --reload --port 8000
# Chế độ CLI
python main.py
# Đánh giá
python evaluate.py data/golden_set_pdf.json # RAGAS eval báo cáo tài chính (27 câu)
# So sánh hai pipeline
python compare_retrieval.py
python evaluate_colpali.py
```
## So sánh Retrieval: ColPali vs Text RAG
Đánh giá trên 10 câu hỏi báo cáo tài chính tiếng Việt, dùng LLM-as-judge chấm điểm mức độ liên quan của nguồn tài liệu:
| Chỉ số | Text RAG | ColPali |
|--------|----------|---------|
| **Source Relevance (trung bình)** | **0.790** | 0.520 |
| Loại câu hỏi | Text RAG | ColPali |
|-------------|----------|---------|
| Bảng số liệu (doanh thu, lợi nhuận) | 0.80 | 0.90 ✓ |
| Tên công ty / mã chứng khoán | 0.90 ✓ | 0.00–0.50 |
| So sánh nhiều kỳ | 0.90 ✓ | 0.20 |
| Thông tin ngành / lĩnh vực | 0.80 | 0.90 ✓ |
**Nhận xét:** Text RAG thắng tổng thể (0.79 vs 0.52) nhờ BM25 xử lý chính xác tên công ty và mã chứng khoán tiếng Việt. ColPali ngang bằng hoặc vượt trội trên các trang chứa bảng số liệu phức tạp — đúng với trường hợp sử dụng thiết kế của nó (nơi pypdf mất cấu trúc bảng). Kết hợp cả hai pipeline sẽ cho kết quả tối ưu.
## Đánh giá RAGAS (LLM-as-judge)
Đánh giá trên **27 câu hỏi** báo cáo tài chính tiếng Việt — 4 doanh nghiệp (VNM, FPT, HPG, MSN) × 2 năm (2022–2023), judge model: `llama-3.3-70b-versatile`:
| Metric | Score |
|--------|-------|
| Faithfulness | **0.737** |
| Answer Relevancy | **0.948** |
| Context Recall | **0.748** |
| **Trung bình** | **0.811** |
**Nhận xét:** Answer Relevancy cao (0.948) cho thấy pipeline trả lời đúng trọng tâm câu hỏi. Faithfulness 0.737 xác nhận câu trả lời bám sát tài liệu gốc. Context Recall 0.748 phản ánh hybrid search (BM25 + vector + reranking) tìm được đúng đoạn liên quan trong phần lớn trường hợp.
## Hạn chế & Hướng phát triển
| Hạn chế hiện tại | Hướng cải thiện |
|-----------------|----------------|
| OCR nhiễu cao với PDF scan (EasyOCR) | Thay bằng Gemini Vision / GPT-4o cho OCR chất lượng cao hơn |
| Bộ đánh giá nhỏ (27 câu) | Mở rộng lên 100+ câu, bổ sung adversarial questions |
| ColPali chỉ chạy local (GPU + ~500KB/trang) | Dùng dịch vụ embedding API hỗ trợ multi-vector |
| Không có conversation memory | Thêm chat history, multi-turn Q&A |
| BM25 rebuild từ Qdrant khi cold start (~5s) | Cache BM25 trên persistent volume hoặc Redis |
| Chưa có hybrid ColPali + Text RAG | Kết hợp MaxSim score với RRF để tận dụng cả hai pipeline |
## Quyết định kiến trúc
| Quyết định | Lý do |
|-----------|-------|
| **HF Spaces thay vì Render/Railway** | Free tier duy nhất có 16GB RAM — đủ cho bge-reranker + Gemini embeddings. Render (512MB) và Railway (0.5GB) không thể load model. |
| **Qdrant Cloud cho text vectors** | Disk HF Spaces reset khi restart — Qdrant Cloud free tier (1GB) giữ vectors qua các lần deploy. |
| **Qdrant local cho ColPali** | Multi-vector ColPali ~500KB/trang — quá lớn cho Qdrant Cloud free tier. Dùng cho thực nghiệm so sánh local. |
| **Không chạy OCR trong pipeline live** | EasyOCR trên CPU quá chậm (~30s/ảnh). Kết quả OCR được baked sẵn vào corpus qua offline ingestion. |
| **Groq thay vì Gemini cho LLM** | Gemini free tier có rate limit nghiêm ngặt. Groq cho inference nhanh với quota thoải mái hơn. |
| **Contextual Retrieval offline** | Cần 1 LLM call/chunk (~219 calls). Chạy một lần lúc ingest; query time không tốn thêm latency. |
| **Streaming qua SSE** | Triển khai tại API layer (`/stream`). Nginx proxy HF Spaces buffer stream — hoạt động đúng trên deploy trực tiếp. |