Instructions to use vllm-sr/Vela-1.0-Encoder-307M-Halu with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use vllm-sr/Vela-1.0-Encoder-307M-Halu with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="vllm-sr/Vela-1.0-Encoder-307M-Halu")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("vllm-sr/Vela-1.0-Encoder-307M-Halu") model = AutoModelForTokenClassification.from_pretrained("vllm-sr/Vela-1.0-Encoder-307M-Halu", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Download README.md from vllm-sr/Vela-1.0-Encoder-307M-Halu: direct link, hf CLI and curl.
- Browser
- Download file 3.74 kB
-
https://huggingface.co/vllm-sr/Vela-1.0-Encoder-307M-Halu/resolve/469c730769786b05a8061c88bdb64db581d6a394/README.md
- Command line
-
hf download hf://vllm-sr/Vela-1.0-Encoder-307M-Halu@469c730769786b05a8061c88bdb64db581d6a394/README.md
-
curl -L -o README.md https://huggingface.co/vllm-sr/Vela-1.0-Encoder-307M-Halu/resolve/469c730769786b05a8061c88bdb64db581d6a394/README.md
library_name: transformers
license: apache-2.0
pipeline_tag: token-classification
base_model: llm-semantic-router/Vela-1.0-Encoder-307M
base_model_relation: finetune
datasets:
- KRLabsOrg/lettucedetect-code-hallucination
- KRLabsOrg/lettucedetect-prose-hallucination
tags:
- modernbert
- semantic-router
- vela
- hallucination-detection
Vela Halu
Vela Halu finds answer spans unsupported by the supplied evidence, for grounded answers and response verification.
307M parameters · Supported input length: 8,192 tokens, including special tokens.
Labels are supported (0) and hallucinated (1). Spans use Unicode character offsets in the answer. Evidence support is distinct from real-world factual truth; an empty result does not guarantee correctness.
Answers requiring arithmetic or multi-step reasoning beyond explicit context may be incorrectly flagged.
Evaluation
Scores (0-100) on the 10,698-example fixed evaluation set. Higher is better.
| Metric | Vela Halu |
|---|---|
| Overall span F1 | 64.68 |
| Overall example F1 | 87.49 |
| Code span F1 | 53.04 |
| Tool output span F1 | 64.28 |
Span F1 measures character overlap; example F1 measures whether an answer contains any hallucinated span. Evaluation uses bfloat16, an 8,192-token limit, only_first pair truncation, and token scores strictly above 0.5. Full scores include all source groups and truncation statistics.
Quick start
With PyTorch and Transformers 4.57.6:
import torch
from transformers import AutoConfig, AutoModelForTokenClassification, AutoTokenizer
model_id = "llm-semantic-router/Vela-1.0-Encoder-307M-Halu"
tokenizer = AutoTokenizer.from_pretrained(model_id, use_fast=True)
config = AutoConfig.from_pretrained(model_id)
config.reference_compile = False
model = AutoModelForTokenClassification.from_pretrained(
model_id, config=config, attn_implementation="sdpa"
).eval()
context = "The museum opens at 10:00 on Tuesday."
question = "When does the museum open on Tuesday?"
answer = "The museum opens at 09:00 on Tuesday."
prompt = f"User request: {question}\n\n{context}"
inputs = tokenizer(prompt, answer, return_offsets_mapping=True,
return_tensors="pt", truncation=False)
assert inputs.input_ids.shape[1] <= 8192
sequence_ids = inputs.sequence_ids(0)
offsets = inputs.pop("offset_mapping")[0].tolist()
with torch.inference_mode():
scores = model(**inputs).logits.float().softmax(-1)[0, :, 1].tolist()
spans, current = [], None
for sequence, (start, end), score in zip(sequence_ids, offsets, scores):
if sequence != 1 or end <= start:
continue
if score > 0.5:
if current is None:
current = {"start": start, "end": end}
else:
current["end"] = max(current["end"], end)
elif current is not None:
spans.append(current)
current = None
if current is not None:
spans.append(current)
print([{**span, "text": answer[span["start"]:span["end"]]} for span in spans])
Keep the complete evidence, request and answer within the supported input length. The example checks this limit before inference.