Token Classification
Transformers
Safetensors
modernbert
semantic-router
vela
hallucination-detection
File size: 3,738 Bytes
11b96af
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
469c730
 
11b96af
 
469c730
11b96af
 
 
469c730
 
 
 
11b96af
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
---
library_name: transformers
license: apache-2.0
pipeline_tag: token-classification
base_model: llm-semantic-router/Vela-1.0-Encoder-307M
base_model_relation: finetune
datasets:
- KRLabsOrg/lettucedetect-code-hallucination
- KRLabsOrg/lettucedetect-prose-hallucination
tags:
- modernbert
- semantic-router
- vela
- hallucination-detection
---

<div align="center">
  <img src="https://vllm-sr.ai/img/vllm-sr-logo.social.png" alt="vLLM Semantic Router" width="560" />
  <p>
    <a href="https://vllm-sr.ai/"><strong>Docs</strong></a> |
    <a href="https://vllm-sr.ai/blog/"><strong>Blog</strong></a> |
    <a href="https://vllm-dev.slack.com/archives/C09CTGF8KCN"><strong>Slack</strong></a> |
    <a href="https://github.com/vllm-project/semantic-router"><strong>GitHub</strong></a>
  </p>
</div>

# Vela Halu

Vela Halu finds answer spans unsupported by the supplied evidence, for grounded answers and response verification.

**307M parameters · Supported input length: 8,192 tokens, including special tokens.**

Labels are `supported` (0) and `hallucinated` (1). Spans use Unicode character offsets in the answer. Evidence support is distinct from real-world factual truth; an empty result does not guarantee correctness.

Answers requiring arithmetic or multi-step reasoning beyond explicit context may be incorrectly flagged.

## Evaluation

Scores (0-100) on the 10,698-example fixed evaluation set. Higher is better.

| Metric | Vela Halu |
|---|---:|
| Overall span F1 | 64.68 |
| Overall example F1 | 87.49 |
| Code span F1 | 53.04 |
| Tool output span F1 | 64.28 |

Span F1 measures character overlap; example F1 measures whether an answer contains any hallucinated span. Evaluation uses bfloat16, an 8,192-token limit, `only_first` pair truncation, and token scores strictly above 0.5. [Full scores](./scores.json) include all source groups and truncation statistics.

## Quick start

With PyTorch and Transformers 4.57.6:

```python
import torch
from transformers import AutoConfig, AutoModelForTokenClassification, AutoTokenizer

model_id = "llm-semantic-router/Vela-1.0-Encoder-307M-Halu"
tokenizer = AutoTokenizer.from_pretrained(model_id, use_fast=True)
config = AutoConfig.from_pretrained(model_id)
config.reference_compile = False
model = AutoModelForTokenClassification.from_pretrained(
    model_id, config=config, attn_implementation="sdpa"
).eval()

context = "The museum opens at 10:00 on Tuesday."
question = "When does the museum open on Tuesday?"
answer = "The museum opens at 09:00 on Tuesday."
prompt = f"User request: {question}\n\n{context}"
inputs = tokenizer(prompt, answer, return_offsets_mapping=True,
                   return_tensors="pt", truncation=False)
assert inputs.input_ids.shape[1] <= 8192
sequence_ids = inputs.sequence_ids(0)
offsets = inputs.pop("offset_mapping")[0].tolist()
with torch.inference_mode():
    scores = model(**inputs).logits.float().softmax(-1)[0, :, 1].tolist()

spans, current = [], None
for sequence, (start, end), score in zip(sequence_ids, offsets, scores):
    if sequence != 1 or end <= start:
        continue
    if score > 0.5:
        if current is None:
            current = {"start": start, "end": end}
        else:
            current["end"] = max(current["end"], end)
    elif current is not None:
        spans.append(current)
        current = None
if current is not None:
    spans.append(current)
print([{**span, "text": answer[span["start"]:span["end"]]} for span in spans])
```

Keep the complete evidence, request and answer within the supported input length. The example checks this limit before inference.

[Explore the Vela model collection](https://huggingface.co/collections/llm-semantic-router/vela-10-router-models-6aa555ba70cc6997d6d67798)