|
Download README.md from vllm-sr/modernbert-base-32k-haldetect: direct link, hf CLI and curl.
- Browser
- Download file 3.33 kB
-
https://huggingface.co/vllm-sr/modernbert-base-32k-haldetect/resolve/b4a19b290467484d3b3d098bc60ae7b4cd587264/README.md
- Command line
-
hf download hf://vllm-sr/modernbert-base-32k-haldetect@b4a19b290467484d3b3d098bc60ae7b4cd587264/README.md
-
curl -L -o README.md https://huggingface.co/vllm-sr/modernbert-base-32k-haldetect/resolve/b4a19b290467484d3b3d098bc60ae7b4cd587264/README.md
3.33 kB
ModernBERT-base-32k Hallucination Detector
A hallucination detection model fine-tuned on RAGTruth dataset using extended 32K context ModernBERT.
Model Description
This model detects hallucinations in LLM-generated text by classifying each token as either Supported (grounded in context) or Hallucinated (not supported by context).
Key Features
- 32K Context Window: Built on
llm-semantic-router/modernbert-base-32kwith YaRN RoPE scaling - Token-Level Classification: Identifies specific spans that are hallucinated
- RAG Optimized: Trained on RAGTruth benchmark for RAG applications
Performance
| Metric | This Model | LettuceDetect BASE | LettuceDetect LARGE |
|---|---|---|---|
| Example-Level F1 | 77.49% | 75.99% | 79.22% |
| Token-Level F1 | 51.47% | 56.27% | - |
Beats LettuceDetect BASE while supporting 4x longer context (32K vs 8K tokens).
Usage
from transformers import AutoModelForTokenClassification, AutoTokenizer
model_name = "llm-semantic-router/modernbert-base-32k-haldetect"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForTokenClassification.from_pretrained(model_name)
# Format: context + question + answer
text = """Context: The Eiffel Tower is located in Paris, France.
Question: Where is the Eiffel Tower?
Answer: The Eiffel Tower is located in London, England."""
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=8192)
outputs = model(**inputs)
predictions = outputs.logits.argmax(dim=-1)
# 0 = Supported, 1 = Hallucinated
With LettuceDetect Library
from lettucedetect.models.inference import HallucinationDetector
detector = HallucinationDetector(
method="transformer",
model_path="llm-semantic-router/modernbert-base-32k-haldetect"
)
context = "The Eiffel Tower is located in Paris, France."
question = "Where is the Eiffel Tower?"
answer = "The Eiffel Tower is located in London, England."
spans = detector.predict(context, question, answer)
# Returns: [{"text": "London, England", "start": 35, "end": 50, "confidence": 0.95}]
Training Details
Dataset
- RAGTruth: ~13,500 samples (QA, Data-to-Text, Summarization)
- Train/Dev/Test split from original RAGTruth
Configuration
base_model: llm-semantic-router/modernbert-base-32k
max_length: 8192
batch_size: 8
learning_rate: 1e-5
epochs: 6
loss: CrossEntropyLoss
scheduler: None (constant LR)
Hardware
- AMD MI300X GPU (196GB VRAM)
- Training time: ~20 minutes
Limitations
- Trained primarily on English text
- Best performance on RAG-style prompts (context + question + answer format)
- Token-level F1 is lower than example-level F1
Citation
@misc{modernbert-32k-haldetect,
title={ModernBERT-base-32k Hallucination Detector},
author={llm-semantic-router},
year={2025},
url={https://huggingface.co/llm-semantic-router/modernbert-base-32k-haldetect}
}
Acknowledgments
- Built on LettuceDetect framework
- Uses ModernBERT architecture
- Trained on RAGTruth dataset