Instructions to use darcar0/quotebound-27b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use darcar0/quotebound-27b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="darcar0/quotebound-27b") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("darcar0/quotebound-27b", device_map="auto") - PEFT
How to use darcar0/quotebound-27b with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use darcar0/quotebound-27b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "darcar0/quotebound-27b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "darcar0/quotebound-27b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/darcar0/quotebound-27b
- SGLang
How to use darcar0/quotebound-27b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "darcar0/quotebound-27b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "darcar0/quotebound-27b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "darcar0/quotebound-27b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "darcar0/quotebound-27b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use darcar0/quotebound-27b with Docker Model Runner:
docker model run hf.co/darcar0/quotebound-27b
Download evidence_faithful_reasoning_release_brief.md from darcar0/quotebound-27b: direct link, hf CLI and curl.
- Browser
- Download file 4.79 kB
-
https://huggingface.co/darcar0/quotebound-27b/resolve/main/evidence_faithful_reasoning_release_brief.md
- Command line
-
hf download hf://darcar0/quotebound-27b/evidence_faithful_reasoning_release_brief.md
-
curl -L -o evidence_faithful_reasoning_release_brief.md https://huggingface.co/darcar0/quotebound-27b/resolve/main/evidence_faithful_reasoning_release_brief.md
Quotebound 27B
Evidence-Faithful Reasoning Release Brief
Released: 2026-04-07 Author: darcar0
Hugging Face model release:
darcar0/quotebound-27b
Companion files:
technical_note_evidence_faithful_reasoning.mdstandalone_holdout_comparison.svgbenchmark_progression.svg
Executive summary
Quotebound 27B is the standalone model release from Evidence-Faithful Reasoning: a research-engineering project on reasoning that has to stay recoverable from the source text rather than asserted on top of it. The project ships a strict benchmark, the hybrid system that clears that benchmark under the full contract, and a standalone model trained to carry the same behavior on its own.
The contract is strict. On every closed packet of source text, the system has to:
- answer correctly,
- cite the right evidence units,
- quote those units verbatim, and
- abstain with
Insufficient evidence.when the packet does not justify a claim.
The project ends in two finished results from one frame. The
benchmark-facing winner is a hybrid stack — bridge checkpoint-2 plus
deterministic_v3 packet-local quote normalization — that clears every
gate on the frozen held-out probe_v0 benchmark. The downloadable artifact
is Quotebound 27B on Hugging Face, which beats the earlier bridge model on
a fresh mixed public holdout and roughly doubles raw quote-faithful
behavior at the model level.
Quotebound 27B
Quotebound 27B is the strongest standalone model the project produced and the artifact most readers will load first. It is the first standalone checkpoint in the project to hold up across multiple evaluation surfaces beyond the held-out probe.
Fresh 36-task mixed public holdout:
| Stack | Task | Strict | Evidence F1 | Quote F1 |
|---|---|---|---|---|
| Bridge raw | 0.8611 | 0.2222 | 0.8815 | 0.3343 |
| Quotebound raw | 0.8889 | 0.4444 | 0.9093 | 0.6815 |
Bridge + deterministic_v3 |
0.8611 | 0.5833 | 0.8815 | 0.8815 |
Quotebound + deterministic_v3 |
0.8889 | 0.5833 | 0.9093 | 0.9093 |
Quotebound 27B beats the prior bridge model on task accuracy, evidence F1,
and quote F1 in both raw and normalized form, ties normalized strict, and
roughly doubles raw quote F1 (0.3343 → 0.6815) at the model level.
Benchmark-facing winner
The benchmark-facing hybrid stack is the strongest full system from the project. Training moved the model past the older frozen baseline; the deterministic packet-local normalizer was the finishing repair, closing the remaining quote-faithful gap without leaving the closed-packet boundary.
| Metric | Frozen probe_v0 |
|---|---|
| Task success | 1.0000 |
| Strict grounded success | 1.0000 |
| Mean evidence F1 | 1.0000 |
| Mean quote F1 | 1.0000 |
| Verify label accuracy | 1.0000 |
| Grounded QA accuracy | 1.0000 |
| Contrastive consistency | 1.0000 |
| Invalid / missing rate | 0.0000 |
Release boundary
The release has two public faces: Quotebound 27B, the standalone model that loads directly from Hugging Face, and a benchmark-facing hybrid stack that closes the last quote-faithfulness gap on the frozen held-out probe. The split is part of the project story, not hidden behind the fine print.
The public release stops at the point where the strongest benchmark-facing system and the strongest standalone model were both clearly in hand. That keeps the package centered on finished artifacts rather than on local variant history.
Intended use and boundaries
This is a release for reasoning over closed packets of source text, not a general-purpose chatbot replacement. It is built for bounded document QA, claim verification, policy and compliance review, contract reading, and other settings where every answer has to be justified from a fixed body of text.
Important boundaries:
- Perfect
probe_v0belongs to the hybrid stack, not to the standalone adapter alone. - The Hugging Face download is the LoRA adapter only; the benchmark-winning
configuration is adapter +
deterministic_v3. - Frozen
probe_v0item-level contents are intentionally not published with the release.
Release surfaces
- Quotebound 27B on Hugging Face:
darcar0/quotebound-27b - Technical note:
technical_note_evidence_faithful_reasoning.md - Fresh public holdout chart:
standalone_holdout_comparison.svg - Frozen benchmark progression chart:
benchmark_progression.svg