NamSyntax commited on
Commit
570c024
·
verified ·
1 Parent(s): c78fabf

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +38 -1
README.md CHANGED
@@ -9,4 +9,41 @@ metrics:
9
  base_model:
10
  - FacebookAI/xlm-roberta-large
11
  pipeline_tag: question-answering
12
- ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9
  base_model:
10
  - FacebookAI/xlm-roberta-large
11
  pipeline_tag: question-answering
12
+ ---
13
+
14
+ # NamSyntax/xlmr-large-viquad
15
+
16
+ ## Model Description
17
+ **NamSyntax/xlmr-large-viquad** is a fine-tuned version of the **XLM-RoBERTa Large** model for Vietnamese Question Answering (QA). It was meticulously fine-tuned on the [UIT-ViQuAD 2.0](https://huggingface.co/datasets/taidng/UIT-ViQuAD2.0) dataset to deliver precise, context-aware answers.
18
+
19
+ This model is a part of the [Vietnamese-QA](https://github.com/NamSyntax/Vietnamese-QA) end-to-end Machine Learning pipeline project. (“If you find this repo helpful, feel free to give it a star ⭐—it really means a lot!”)
20
+
21
+ ## Intended Uses
22
+ - **Primary Use:** Extractive Question Answering in Vietnamese. Given a context paragraph and a question, the model extracts the substring representing the answer from the context.
23
+ - **Ecosystem:** It can be instantly integrated into web services. The associated project provides out-of-the-box support for a **Gradio Web UI** and a **FastAPI** REST backend.
24
+
25
+ ## Training Data
26
+ The model was fine-tuned on **UIT-ViQuAD 2.0**, a high-quality machine reading comprehension dataset curated specifically for the Vietnamese language.
27
+
28
+ ## Training Procedure
29
+ This model was trained using PyTorch and the Hugging Face `Trainer` API. The training pipeline implements:
30
+ - Clean modular architecture for HuggingFace Dataset loading.
31
+ - Preprocessing steps emphasizing text tokenization and chunking logic to handle long documents.
32
+ - Hyperparameter management via `config.yaml`.
33
+
34
+ ## How to use
35
+ You can use this model directly with the `pipeline` API from Hugging Face:
36
+
37
+ ```python
38
+ from transformers import pipeline
39
+
40
+ qa_pipeline = pipeline("question-answering", model="NamSyntax/xlmr-large-viquad")
41
+
42
+ context = "Hà Nội là thủ đô của nước Cộng hòa Xã hội chủ nghĩa Việt Nam. Thành phố nằm ở phía tây bắc trung tâm vùng đồng bằng châu thổ sông Hồng."
43
+ question = "Thủ đô của Việt Nam là gì?"
44
+
45
+ result = qa_pipeline(question=question, context=context)
46
+ print(result)
47
+ ```
48
+
49
+ Alternatively, to experience this model with an interactive UI or to deploy it as a production-grade backend service (FastAPI and Docker), please refer to the official repository: [NamSyntax/Vietnamese-QA](https://github.com/NamSyntax/Vietnamese-QA).