quotebound-27b / README.md
darcar0's picture
Upload README.md with huggingface_hub
e5e90d1 verified
|
Raw History Blame
8.29 kB
---
language:
- en
license: apache-2.0
base_model:
- Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2
datasets:
- fever/fever
- hotpotqa/hotpot_qa
library_name: transformers
pipeline_tag: text-generation
tags:
- reasoning
- evidence-grounding
- attribution
- fever
- hotpotqa
- lora
- peft
- distillation
- research
---
# Evidence-Faithful Reasoning
*Pilot 3 is the standalone model release.*
This adapter turns its reasoning-distilled 27B base model into an
evidence-first reader for closed packets of text. I built it because I wanted
a release where reasoning had to prove itself: every answer has to land on the
right evidence, quote that evidence verbatim, and stop with
`Insufficient evidence.` when the packet does not justify a claim. The result
is the strongest standalone model from the project, packaged here as a LoRA
adapter you can load directly tonight.
## Resources & Guides
- [Technical brief (PDF)](./evidence_faithful_reasoning_release_brief.pdf)
- [Technical note](./technical_note_evidence_faithful_reasoning.md)
- [Fresh public holdout chart](./standalone_holdout_comparison.svg)
- [Frozen benchmark progression chart](./benchmark_progression.svg)
- [Release architecture chart](./project_release_arc.svg)
![Fresh public holdout: standalone release vs bridge](./standalone_holdout_comparison.svg)
*Fresh 36-task mixed public holdout: the standalone release beats the earlier
bridge model on task accuracy, evidence F1, and quote F1, while the
packet-local normalizer lifts the full stack to `0.9093` quote F1.*
## Why this release exists
I built this project to force reasoning models to show their work in the only
place that counts: the evidence itself. Fluent answers were not enough. I
wanted a model that had to retrieve the right units, quote them exactly, and
fail closed when the packet ran out. This page leads with the standalone
release because it is the artifact you can load immediately, inspect directly,
and use without reconstructing the whole benchmark stack.
## At a glance
- LoRA adapter on top of
[`Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2`](https://huggingface.co/Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2).
- Strongest standalone model from the project and the release I want people to
download first.
- On a fresh 36-task public holdout, raw task improves from `0.8611` to
`0.8889`, raw strict from `0.2222` to `0.4444`, and raw quote F1 from
`0.3343` to `0.6815` over the earlier bridge model.
- Zero invalid outputs on every reported evaluation surface.
- The project also produced a benchmark-winning hybrid stack, but that is a
separate result described under *Release architecture*.
## Quick start
Pilot 3 is a LoRA adapter. Load the base model and attach the adapter:
```python
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base_id = "Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2"
adapter_id = "darcar0/evidence-faithful-reasoning-pilot-3"
tokenizer = AutoTokenizer.from_pretrained(base_id)
base = AutoModelForCausalLM.from_pretrained(base_id, device_map="auto")
model = PeftModel.from_pretrained(base, adapter_id)
```
The base model is 27B parameters, so load it in your usual quantization.
## Prompt format
This release works best with an evidence-first prompt that makes the answer
subordinate to the cited text. A minimal version:
```
You are answering from a bounded evidence packet only.
Work in this order:
1. Identify the smallest set of packet units that matters.
2. Copy exact quote(s) from those units.
3. Only then give the final answer.
Rules:
- No outside facts.
- Return valid JSON only.
- Every quote must be a verbatim substring of the cited unit.
- Do not paraphrase, ellipsize, or stitch quotes.
- If the packet is insufficient, the `answer` field must be exactly
`Insufficient evidence.`
```
The model then writes a JSON object with this shape:
```json
{
"task_id": "<task id>",
"label": "support|contradict|insufficient|null",
"answer": "<one-sentence answer>",
"evidence_ids": ["unit_id_1", "unit_id_2"],
"quotes": [
{"unit_id": "unit_id_1", "quote": "<exact quote>"}
],
"justification": "<one short sentence tied to the cited evidence>"
}
```
## Evaluation
### Fresh 36-task mixed public holdout
A held-out slice of 18 FEVER verify-claim tasks plus 18 HotpotQA grounded-QA
tasks, drawn from public sources and de-duplicated against every training,
dev, and `probe_v0` row.
| Stack | Task | Strict | Evidence F1 | Quote F1 |
|---|---:|---:|---:|---:|
| Bridge raw | 0.8611 | 0.2222 | 0.8815 | 0.3343 |
| Pilot 3 raw | 0.8889 | 0.4444 | 0.9093 | 0.6815 |
| Bridge + `deterministic_v3` | 0.8611 | 0.5833 | 0.8815 | 0.8815 |
| **Pilot 3 + `deterministic_v3`** | **0.8889** | **0.5833** | **0.9093** | **0.9093** |
The standalone release beats the earlier bridge model on task accuracy,
evidence F1, and quote F1 in both raw and normalized form, ties normalized
strict, and roughly doubles raw quote F1 at the model level.
### Fixed dev triage slice (21 tasks)
| Stack | Task | Strict | Evidence F1 | Quote F1 |
|---|---:|---:|---:|---:|
| Pilot 3 + `deterministic_v3` | 1.0000 | 0.6190 | 0.8320 | 0.7095 |
### Untouched 104-task Hotpot shadow slice
Pilot 3 raw improved quote-faithful behavior over the raw bridge model on this
slice, and pilot 3 + `deterministic_v3` matched bridge +
`deterministic_v3` at the system level. That surface remains a narrative
parity result because the report does not publish per-metric cells for it.
## Release architecture
This project ends in two finished artifacts, not one:
1. **Standalone model release** — this page. Pilot 3 is the strongest
version of the project's evidence-faithful behavior that moved into the
model itself, evaluated across multiple non-`probe_v0` surfaces.
2. **Benchmark-facing hybrid stack** — bridge `checkpoint-2` plus the
`deterministic_v3` packet-local normalizer. That stack is the benchmark
winner and the only configuration that clears every gate on frozen held-out
`probe_v0`.
The separation is deliberate. This page is for the standalone release you can
download now. The benchmark winner is documented here because it explains the
project's full result, not because those perfect `probe_v0` numbers belong to
the adapter alone.
## Intended use
Use this release for work that has to stay inside a fixed body of text:
- bounded document QA with explicit evidence requirements,
- claim verification and grounded QA from closed evidence packets,
- policy, compliance, contract, and internal-document workflows where each
answer must be justified from the provided text,
- research on evidence-faithful reasoning and abstention behavior.
## Limitations
- The downloadable artifact is the LoRA adapter only. The base model is
required.
- The `deterministic_v3` packet-local normalizer is not included in this
download. The benchmark-winning configuration is adapter + normalizer, while
the adapter alone reproduces the standalone-model results shown above.
- Perfect `probe_v0` belongs to the benchmark-facing hybrid stack, not to this
adapter alone.
- Specialized for closed-packet reasoning, not open-ended chat or open-domain
QA.
- Frozen `probe_v0` item-level contents are intentionally not published with
the release.
## Citation
References:
- Base model:
[Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2](https://huggingface.co/Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2)
- Datasets:
[fever/fever](https://huggingface.co/datasets/fever/fever),
[hotpotqa/hotpot_qa](https://huggingface.co/datasets/hotpotqa/hotpot_qa)
- Technical brief (PDF):
[evidence_faithful_reasoning_release_brief.pdf](./evidence_faithful_reasoning_release_brief.pdf)
- Technical note:
[technical_note_evidence_faithful_reasoning.md](./technical_note_evidence_faithful_reasoning.md)
```bibtex
@misc{darcar0_evidence_faithful_reasoning_pilot_3_2026,
title = {Evidence-Faithful Reasoning: Pilot 3},
author = {darcar0},
year = {2026},
howpublished = {Hugging Face model release},
url = {https://huggingface.co/darcar0/evidence-faithful-reasoning-pilot-3}
}
```