Token Classification
Transformers
Safetensors
modernbert
semantic-router
vela
hallucination-detection
Instructions to use vllm-sr/Vela-1.0-Encoder-307M-Halu with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use vllm-sr/Vela-1.0-Encoder-307M-Halu with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="vllm-sr/Vela-1.0-Encoder-307M-Halu")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("vllm-sr/Vela-1.0-Encoder-307M-Halu") model = AutoModelForTokenClassification.from_pretrained("vllm-sr/Vela-1.0-Encoder-307M-Halu", device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from vllm-sr/Vela-1.0-Encoder-307M-Halu: direct link, hf CLI and curl.
- Browser
- Download file 3.74 kB
-
https://huggingface.co/vllm-sr/Vela-1.0-Encoder-307M-Halu/resolve/469c730769786b05a8061c88bdb64db581d6a394/README.md
- Command line
-
hf download hf://vllm-sr/Vela-1.0-Encoder-307M-Halu@469c730769786b05a8061c88bdb64db581d6a394/README.md
-
curl -L -o README.md https://huggingface.co/vllm-sr/Vela-1.0-Encoder-307M-Halu/resolve/469c730769786b05a8061c88bdb64db581d6a394/README.md
3.74 kB
| library_name: transformers | |
| license: apache-2.0 | |
| pipeline_tag: token-classification | |
| base_model: llm-semantic-router/Vela-1.0-Encoder-307M | |
| base_model_relation: finetune | |
| datasets: | |
| - KRLabsOrg/lettucedetect-code-hallucination | |
| - KRLabsOrg/lettucedetect-prose-hallucination | |
| tags: | |
| - modernbert | |
| - semantic-router | |
| - vela | |
| - hallucination-detection | |
| <div align="center"> | |
| <img src="https://vllm-sr.ai/img/vllm-sr-logo.social.png" alt="vLLM Semantic Router" width="560" /> | |
| <p> | |
| <a href="https://vllm-sr.ai/"><strong>Docs</strong></a> | | |
| <a href="https://vllm-sr.ai/blog/"><strong>Blog</strong></a> | | |
| <a href="https://vllm-dev.slack.com/archives/C09CTGF8KCN"><strong>Slack</strong></a> | | |
| <a href="https://github.com/vllm-project/semantic-router"><strong>GitHub</strong></a> | |
| </p> | |
| </div> | |
| # Vela Halu | |
| Vela Halu finds answer spans unsupported by the supplied evidence, for grounded answers and response verification. | |
| **307M parameters · Supported input length: 8,192 tokens, including special tokens.** | |
| Labels are `supported` (0) and `hallucinated` (1). Spans use Unicode character offsets in the answer. Evidence support is distinct from real-world factual truth; an empty result does not guarantee correctness. | |
| Answers requiring arithmetic or multi-step reasoning beyond explicit context may be incorrectly flagged. | |
| ## Evaluation | |
| Scores (0-100) on the 10,698-example fixed evaluation set. Higher is better. | |
| | Metric | Vela Halu | | |
| |---|---:| | |
| | Overall span F1 | 64.68 | | |
| | Overall example F1 | 87.49 | | |
| | Code span F1 | 53.04 | | |
| | Tool output span F1 | 64.28 | | |
| Span F1 measures character overlap; example F1 measures whether an answer contains any hallucinated span. Evaluation uses bfloat16, an 8,192-token limit, `only_first` pair truncation, and token scores strictly above 0.5. [Full scores](./scores.json) include all source groups and truncation statistics. | |
| ## Quick start | |
| With PyTorch and Transformers 4.57.6: | |
| ```python | |
| import torch | |
| from transformers import AutoConfig, AutoModelForTokenClassification, AutoTokenizer | |
| model_id = "llm-semantic-router/Vela-1.0-Encoder-307M-Halu" | |
| tokenizer = AutoTokenizer.from_pretrained(model_id, use_fast=True) | |
| config = AutoConfig.from_pretrained(model_id) | |
| config.reference_compile = False | |
| model = AutoModelForTokenClassification.from_pretrained( | |
| model_id, config=config, attn_implementation="sdpa" | |
| ).eval() | |
| context = "The museum opens at 10:00 on Tuesday." | |
| question = "When does the museum open on Tuesday?" | |
| answer = "The museum opens at 09:00 on Tuesday." | |
| prompt = f"User request: {question}\n\n{context}" | |
| inputs = tokenizer(prompt, answer, return_offsets_mapping=True, | |
| return_tensors="pt", truncation=False) | |
| assert inputs.input_ids.shape[1] <= 8192 | |
| sequence_ids = inputs.sequence_ids(0) | |
| offsets = inputs.pop("offset_mapping")[0].tolist() | |
| with torch.inference_mode(): | |
| scores = model(**inputs).logits.float().softmax(-1)[0, :, 1].tolist() | |
| spans, current = [], None | |
| for sequence, (start, end), score in zip(sequence_ids, offsets, scores): | |
| if sequence != 1 or end <= start: | |
| continue | |
| if score > 0.5: | |
| if current is None: | |
| current = {"start": start, "end": end} | |
| else: | |
| current["end"] = max(current["end"], end) | |
| elif current is not None: | |
| spans.append(current) | |
| current = None | |
| if current is not None: | |
| spans.append(current) | |
| print([{**span, "text": answer[span["start"]:span["end"]]} for span in spans]) | |
| ``` | |
| Keep the complete evidence, request and answer within the supported input length. The example checks this limit before inference. | |
| [Explore the Vela model collection](https://huggingface.co/collections/llm-semantic-router/vela-10-router-models-6aa555ba70cc6997d6d67798) | |