camembert-ner-travel
CamemBERT-base fine-tuned for Named Entity Recognition on French travel sentences. The model extracts departure city (LOC_ORIGIN) and destination city (LOC_DEST) from user queries.
Model description
This model is part of the Travel Order Resolver project, a pipeline that:
- Extracts origin/destination from a French sentence (this model)
- Resolves city names against the SNCF station network
- Computes the optimal train route (Dijkstra / A*)
Base model: camembert-base
Task: Token classification (BIO tagging)
Labels
| Label | Description |
|---|---|
O |
Outside any entity |
B-LOC_ORIGIN |
Beginning of a departure location |
I-LOC_ORIGIN |
Inside a departure location |
B-LOC_DEST |
Beginning of a destination location |
I-LOC_DEST |
Inside a destination location |
Training data
Synthetic dataset of French travel sentences generated from SNCF station data. Noise augmentation applied: lowercase, missing accents, typos (≈30% of samples).
Examples:
"Je veux aller de Paris à Lyon"→Paris= LOC_ORIGIN,Lyon= LOC_DEST"un billet depuis marseille pour bordeaux"→ noisy variant
Training procedure
| Hyperparameter | Value |
|---|---|
| Base model | camembert-base |
| Epochs | 3 |
| Batch size | 16 |
| Learning rate | 5e-5 |
| LR scheduler | linear |
| Warmup steps | 500 |
| Weight decay | 0.01 |
| Training runtime | ~164s |
Evaluation results
Test set
| Metric | Value |
|---|---|
| F1 | 0.9930 |
| Precision | 0.9879 |
| Recall | 0.9981 |
Validation set
| Metric | Value |
|---|---|
| F1 | 0.9921 |
| Precision | 0.9890 |
| Recall | 0.9951 |
Comparison with other runs
| Run | LR | Scheduler | Epochs | F1 test |
|---|---|---|---|---|
| exp_lr2e-5_cosine | 2e-5 | cosine | 3 | 0.9918 |
| exp_lr2e-5_linear | 2e-5 | linear | 3 | 0.9930 |
| exp_lr5e-5_linear | 5e-5 | linear | 3 | 0.9930 |
lr=5e-5 converges ~2× faster than lr=2e-5 (train loss 0.057 vs 0.129 at step 1000) with identical final F1.
Usage
from transformers import pipeline
ner = pipeline(
"token-classification",
model="Aldo26/camembert-ner-travel",
aggregation_strategy="simple",
)
result = ner("Je voudrais un billet de Nantes pour Strasbourg")
# [
# {'entity_group': 'LOC_ORIGIN', 'word': 'Nantes', ...},
# {'entity_group': 'LOC_DEST', 'word': 'Strasbourg', ...}
# ]
Limitations
- Trained on synthetic data — may underperform on very unusual phrasings
- Handles one origin and one destination per sentence
- French only
- Downloads last month
- 20
Model tree for Aldo26/camembert-ner-travel
Base model
almanach/camembert-baseEvaluation results
- F1 (test)self-reported0.993
- Precision (test)self-reported0.988
- Recall (test)self-reported0.998