Instructions to use vllm-sr/Vela-1.0-Encoder-307M-PII with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use vllm-sr/Vela-1.0-Encoder-307M-PII with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="vllm-sr/Vela-1.0-Encoder-307M-PII")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("vllm-sr/Vela-1.0-Encoder-307M-PII") model = AutoModelForTokenClassification.from_pretrained("vllm-sr/Vela-1.0-Encoder-307M-PII", device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 2,406 Bytes
f91a604 8ad7379 f91a604 73f4d3e 8ad7379 f91a604 8ad7379 f91a604 ca6210b 8ad7379 f91a604 8d773b8 f91a604 8d773b8 f91a604 6d0034f 73f4d3e 6d0034f 86d35f9 f91a604 8d773b8 f91a604 8ad7379 8d773b8 8ad7379 73f4d3e 8d773b8 8ad7379 8d773b8 86d35f9 f91a604 73f4d3e | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 | ---
library_name: transformers
license: apache-2.0
pipeline_tag: token-classification
base_model: vllm-sr/Vela-1.0-Encoder-307M
base_model_relation: finetune
language:
- en
- zh
- es
- fr
- de
- ja
tags:
- modernbert
- semantic-router
- vela
---
<div align="center">
<img src="https://vllm-sr.ai/img/vllm-sr-logo.social.png" alt="vLLM Semantic Router" width="560" />
<p>
<a href="https://vllm-sr.ai/"><strong>Docs</strong></a> |
<a href="https://vllm-sr.ai/blog/"><strong>Blog</strong></a> |
<a href="https://vllm-dev.slack.com/archives/C09CTGF8KCN"><strong>Slack</strong></a> |
<a href="https://github.com/vllm-project/semantic-router"><strong>GitHub</strong></a>
</p>
</div>
# Vela PII
Vela PII finds sensitive entity spans for privacy-aware routing and redaction.
**307M parameters 路 Input capacity: 32,768 tokens, including special tokens.**
Outputs use 35 BIO labels across 17 entity types. The example returns spans with Unicode character offsets.
## Evaluation
Exact-span micro F1 (脳100) on the same synthetic development sets, compared with [the original mmBERT32K PII model](https://huggingface.co/vllm-sr/mmbert32k-pii-detector-merged). Higher is better.
| Evaluation | Original mmBERT | Vela |
|---|---:|---:|
| Short inputs 路 888 | 23.66 | **90.76** |
| Controlled 4K context 路 30 | 0.44 | **89.03** |
| Controlled 8K context 路 30 | 0.28 | **89.88** |
| Controlled 16K context 路 30 | 0.36 | **89.73** |
| Controlled 32K context 路 30 | 0.23 | **89.24** |
Synthetic examples cover six languages. Long inputs include sparse entities, densely packed repeated entities and negative examples; micro F1 weights each entity equally. Both models process complete inputs in FP32 with the same exact-span scorer. These development sets informed Vela selection; they are not an independent natural-document benchmark.
## Quick start
With PyTorch and Transformers 4.57.6 or 5.17.0:
```python
from transformers import pipeline
model_id = "vllm-sr/Vela-1.0-Encoder-307M-PII"
model = pipeline("token-classification", model=model_id, aggregation_strategy="simple", device=-1)
text = "Contact Mara Wells at mara.wells@example.com."
assert len(model.tokenizer.encode(text)) <= model.model.config.max_position_embeddings
print(model(text))
```
[Explore the Vela model collection](https://huggingface.co/collections/vllm-sr/vela-10-router-models-6aa555ba70cc6997d6d67798)
|