Text Generation
Transformers
Safetensors
PEFT
English
reasoning
evidence-grounding
grounded-qa
attribution
fever
hotpotqa
lora
distillation
research
conversational
Instructions to use darcar0/quotebound-27b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use darcar0/quotebound-27b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="darcar0/quotebound-27b") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("darcar0/quotebound-27b", device_map="auto") - PEFT
How to use darcar0/quotebound-27b with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use darcar0/quotebound-27b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "darcar0/quotebound-27b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "darcar0/quotebound-27b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/darcar0/quotebound-27b
- SGLang
How to use darcar0/quotebound-27b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "darcar0/quotebound-27b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "darcar0/quotebound-27b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "darcar0/quotebound-27b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "darcar0/quotebound-27b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use darcar0/quotebound-27b with Docker Model Runner:
docker model run hf.co/darcar0/quotebound-27b
File size: 4,787 Bytes
dab6650 9c1d54d dab6650 9c1d54d dab6650 9c1d54d dab6650 9c1d54d dab6650 0cdc637 9c1d54d dab6650 9c1d54d dab6650 9c1d54d dab6650 9c1d54d dab6650 9c1d54d dab6650 9c1d54d b314f67 9c1d54d dab6650 9c1d54d dab6650 9c1d54d dab6650 9c1d54d 0cdc637 dab6650 0cdc637 9c1d54d b314f67 9c1d54d dab6650 9c1d54d b314f67 9c1d54d dab6650 9c1d54d dab6650 9c1d54d dab6650 9c1d54d dab6650 0cdc637 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 | # Quotebound 27B
## Evidence-Faithful Reasoning Release Brief
Released: 2026-04-07
Author: darcar0
Hugging Face model release:
[`darcar0/quotebound-27b`](https://huggingface.co/darcar0/quotebound-27b)
Companion files:
- [`technical_note_evidence_faithful_reasoning.md`](./technical_note_evidence_faithful_reasoning.md)
- [`standalone_holdout_comparison.svg`](./standalone_holdout_comparison.svg)
- [`benchmark_progression.svg`](./benchmark_progression.svg)
## Executive summary
Quotebound 27B is the standalone model release from Evidence-Faithful
Reasoning: a research-engineering project on reasoning that has to stay
recoverable from the source text rather than asserted on top of it. The
project ships a strict benchmark, the hybrid system that clears that
benchmark under the full contract, and a standalone model trained to carry
the same behavior on its own.
The contract is strict. On every closed packet of source text, the system
has to:
1. answer correctly,
2. cite the right evidence units,
3. quote those units verbatim, and
4. abstain with `Insufficient evidence.` when the packet does not justify a
claim.
The project ends in two finished results from one frame. The
benchmark-facing winner is a hybrid stack — bridge `checkpoint-2` plus
`deterministic_v3` packet-local quote normalization — that clears every
gate on the frozen held-out `probe_v0` benchmark. The downloadable artifact
is Quotebound 27B on Hugging Face, which beats the earlier bridge model on
a fresh mixed public holdout and roughly doubles raw quote-faithful
behavior at the model level.
## Quotebound 27B
Quotebound 27B is the strongest standalone model the project produced and
the artifact most readers will load first. It is the first standalone
checkpoint in the project to hold up across multiple evaluation surfaces
beyond the held-out probe.
Fresh 36-task mixed public holdout:
| Stack | Task | Strict | Evidence F1 | Quote F1 |
|---|---:|---:|---:|---:|
| Bridge raw | 0.8611 | 0.2222 | 0.8815 | 0.3343 |
| Quotebound raw | 0.8889 | 0.4444 | 0.9093 | 0.6815 |
| Bridge + `deterministic_v3` | 0.8611 | 0.5833 | 0.8815 | 0.8815 |
| **Quotebound + `deterministic_v3`** | **0.8889** | **0.5833** | **0.9093** | **0.9093** |
Quotebound 27B beats the prior bridge model on task accuracy, evidence F1,
and quote F1 in both raw and normalized form, ties normalized strict, and
roughly doubles raw quote F1 (`0.3343` → `0.6815`) at the model level.
## Benchmark-facing winner
The benchmark-facing hybrid stack is the strongest full system from the
project. Training moved the model past the older frozen baseline; the
deterministic packet-local normalizer was the finishing repair, closing the
remaining quote-faithful gap without leaving the closed-packet boundary.
| Metric | Frozen probe_v0 |
|---|---:|
| Task success | **1.0000** |
| Strict grounded success | **1.0000** |
| Mean evidence F1 | **1.0000** |
| Mean quote F1 | **1.0000** |
| Verify label accuracy | **1.0000** |
| Grounded QA accuracy | **1.0000** |
| Contrastive consistency | **1.0000** |
| Invalid / missing rate | **0.0000** |
## Release boundary
The release has two public faces: Quotebound 27B, the standalone model that
loads directly from Hugging Face, and a benchmark-facing hybrid stack that
closes the last quote-faithfulness gap on the frozen held-out probe. The
split is part of the project story, not hidden behind the fine print.
The public release stops at the point where the strongest benchmark-facing
system and the strongest standalone model were both clearly in hand. That
keeps the package centered on finished artifacts rather than on local
variant history.
## Intended use and boundaries
This is a release for reasoning over closed packets of source text, not a
general-purpose chatbot replacement. It is built for bounded document QA,
claim verification, policy and compliance review, contract reading, and
other settings where every answer has to be justified from a fixed body of
text.
Important boundaries:
- Perfect `probe_v0` belongs to the hybrid stack, not to the standalone
adapter alone.
- The Hugging Face download is the LoRA adapter only; the benchmark-winning
configuration is adapter + `deterministic_v3`.
- Frozen `probe_v0` item-level contents are intentionally not published with
the release.
## Release surfaces
- Quotebound 27B on Hugging Face:
[`darcar0/quotebound-27b`](https://huggingface.co/darcar0/quotebound-27b)
- Technical note:
[`technical_note_evidence_faithful_reasoning.md`](./technical_note_evidence_faithful_reasoning.md)
- Fresh public holdout chart:
[`standalone_holdout_comparison.svg`](./standalone_holdout_comparison.svg)
- Frozen benchmark progression chart:
[`benchmark_progression.svg`](./benchmark_progression.svg)
|