Update README.md
Browse files
README.md
CHANGED
|
@@ -9,4 +9,41 @@ metrics:
|
|
| 9 |
base_model:
|
| 10 |
- FacebookAI/xlm-roberta-large
|
| 11 |
pipeline_tag: question-answering
|
| 12 |
-
---
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 9 |
base_model:
|
| 10 |
- FacebookAI/xlm-roberta-large
|
| 11 |
pipeline_tag: question-answering
|
| 12 |
+
---
|
| 13 |
+
|
| 14 |
+
# NamSyntax/xlmr-large-viquad
|
| 15 |
+
|
| 16 |
+
## Model Description
|
| 17 |
+
**NamSyntax/xlmr-large-viquad** is a fine-tuned version of the **XLM-RoBERTa Large** model for Vietnamese Question Answering (QA). It was meticulously fine-tuned on the [UIT-ViQuAD 2.0](https://huggingface.co/datasets/taidng/UIT-ViQuAD2.0) dataset to deliver precise, context-aware answers.
|
| 18 |
+
|
| 19 |
+
This model is a part of the [Vietnamese-QA](https://github.com/NamSyntax/Vietnamese-QA) end-to-end Machine Learning pipeline project. (“If you find this repo helpful, feel free to give it a star ⭐—it really means a lot!”)
|
| 20 |
+
|
| 21 |
+
## Intended Uses
|
| 22 |
+
- **Primary Use:** Extractive Question Answering in Vietnamese. Given a context paragraph and a question, the model extracts the substring representing the answer from the context.
|
| 23 |
+
- **Ecosystem:** It can be instantly integrated into web services. The associated project provides out-of-the-box support for a **Gradio Web UI** and a **FastAPI** REST backend.
|
| 24 |
+
|
| 25 |
+
## Training Data
|
| 26 |
+
The model was fine-tuned on **UIT-ViQuAD 2.0**, a high-quality machine reading comprehension dataset curated specifically for the Vietnamese language.
|
| 27 |
+
|
| 28 |
+
## Training Procedure
|
| 29 |
+
This model was trained using PyTorch and the Hugging Face `Trainer` API. The training pipeline implements:
|
| 30 |
+
- Clean modular architecture for HuggingFace Dataset loading.
|
| 31 |
+
- Preprocessing steps emphasizing text tokenization and chunking logic to handle long documents.
|
| 32 |
+
- Hyperparameter management via `config.yaml`.
|
| 33 |
+
|
| 34 |
+
## How to use
|
| 35 |
+
You can use this model directly with the `pipeline` API from Hugging Face:
|
| 36 |
+
|
| 37 |
+
```python
|
| 38 |
+
from transformers import pipeline
|
| 39 |
+
|
| 40 |
+
qa_pipeline = pipeline("question-answering", model="NamSyntax/xlmr-large-viquad")
|
| 41 |
+
|
| 42 |
+
context = "Hà Nội là thủ đô của nước Cộng hòa Xã hội chủ nghĩa Việt Nam. Thành phố nằm ở phía tây bắc trung tâm vùng đồng bằng châu thổ sông Hồng."
|
| 43 |
+
question = "Thủ đô của Việt Nam là gì?"
|
| 44 |
+
|
| 45 |
+
result = qa_pipeline(question=question, context=context)
|
| 46 |
+
print(result)
|
| 47 |
+
```
|
| 48 |
+
|
| 49 |
+
Alternatively, to experience this model with an interactive UI or to deploy it as a production-grade backend service (FastAPI and Docker), please refer to the official repository: [NamSyntax/Vietnamese-QA](https://github.com/NamSyntax/Vietnamese-QA).
|