--- language: - en license: apache-2.0 base_model: - Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2 datasets: - fever/fever - hotpotqa/hotpot_qa library_name: transformers pipeline_tag: text-generation tags: - reasoning - evidence-grounding - attribution - fever - hotpotqa - lora - peft - distillation - research --- # Evidence-Faithful Reasoning *Pilot 3 is the standalone model release.* This adapter turns its reasoning-distilled 27B base model into an evidence-first reader for closed packets of text. I built it because I wanted a release where reasoning had to prove itself: every answer has to land on the right evidence, quote that evidence verbatim, and stop with `Insufficient evidence.` when the packet does not justify a claim. The result is the strongest standalone model from the project, packaged here as a LoRA adapter you can load directly tonight. ## Resources & Guides - [Technical brief (PDF)](./evidence_faithful_reasoning_release_brief.pdf) - [Technical note](./technical_note_evidence_faithful_reasoning.md) - [Fresh public holdout chart](./standalone_holdout_comparison.svg) - [Frozen benchmark progression chart](./benchmark_progression.svg) - [Release architecture chart](./project_release_arc.svg) ![Fresh public holdout: standalone release vs bridge](./standalone_holdout_comparison.svg) *Fresh 36-task mixed public holdout: the standalone release beats the earlier bridge model on task accuracy, evidence F1, and quote F1, while the packet-local normalizer lifts the full stack to `0.9093` quote F1.* ## Why this release exists I built this project to force reasoning models to show their work in the only place that counts: the evidence itself. Fluent answers were not enough. I wanted a model that had to retrieve the right units, quote them exactly, and fail closed when the packet ran out. This page leads with the standalone release because it is the artifact you can load immediately, inspect directly, and use without reconstructing the whole benchmark stack. ## At a glance - LoRA adapter on top of [`Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2`](https://huggingface.co/Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2). - Strongest standalone model from the project and the release I want people to download first. - On a fresh 36-task public holdout, raw task improves from `0.8611` to `0.8889`, raw strict from `0.2222` to `0.4444`, and raw quote F1 from `0.3343` to `0.6815` over the earlier bridge model. - Zero invalid outputs on every reported evaluation surface. - The project also produced a benchmark-winning hybrid stack, but that is a separate result described under *Release architecture*. ## Quick start Pilot 3 is a LoRA adapter. Load the base model and attach the adapter: ```python from peft import PeftModel from transformers import AutoModelForCausalLM, AutoTokenizer base_id = "Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2" adapter_id = "darcar0/evidence-faithful-reasoning-pilot-3" tokenizer = AutoTokenizer.from_pretrained(base_id) base = AutoModelForCausalLM.from_pretrained(base_id, device_map="auto") model = PeftModel.from_pretrained(base, adapter_id) ``` The base model is 27B parameters, so load it in your usual quantization. ## Prompt format This release works best with an evidence-first prompt that makes the answer subordinate to the cited text. A minimal version: ``` You are answering from a bounded evidence packet only. Work in this order: 1. Identify the smallest set of packet units that matters. 2. Copy exact quote(s) from those units. 3. Only then give the final answer. Rules: - No outside facts. - Return valid JSON only. - Every quote must be a verbatim substring of the cited unit. - Do not paraphrase, ellipsize, or stitch quotes. - If the packet is insufficient, the `answer` field must be exactly `Insufficient evidence.` ``` The model then writes a JSON object with this shape: ```json { "task_id": "", "label": "support|contradict|insufficient|null", "answer": "", "evidence_ids": ["unit_id_1", "unit_id_2"], "quotes": [ {"unit_id": "unit_id_1", "quote": ""} ], "justification": "" } ``` ## Evaluation ### Fresh 36-task mixed public holdout A held-out slice of 18 FEVER verify-claim tasks plus 18 HotpotQA grounded-QA tasks, drawn from public sources and de-duplicated against every training, dev, and `probe_v0` row. | Stack | Task | Strict | Evidence F1 | Quote F1 | |---|---:|---:|---:|---:| | Bridge raw | 0.8611 | 0.2222 | 0.8815 | 0.3343 | | Pilot 3 raw | 0.8889 | 0.4444 | 0.9093 | 0.6815 | | Bridge + `deterministic_v3` | 0.8611 | 0.5833 | 0.8815 | 0.8815 | | **Pilot 3 + `deterministic_v3`** | **0.8889** | **0.5833** | **0.9093** | **0.9093** | The standalone release beats the earlier bridge model on task accuracy, evidence F1, and quote F1 in both raw and normalized form, ties normalized strict, and roughly doubles raw quote F1 at the model level. ### Fixed dev triage slice (21 tasks) | Stack | Task | Strict | Evidence F1 | Quote F1 | |---|---:|---:|---:|---:| | Pilot 3 + `deterministic_v3` | 1.0000 | 0.6190 | 0.8320 | 0.7095 | ### Untouched 104-task Hotpot shadow slice Pilot 3 raw improved quote-faithful behavior over the raw bridge model on this slice, and pilot 3 + `deterministic_v3` matched bridge + `deterministic_v3` at the system level. That surface remains a narrative parity result because the report does not publish per-metric cells for it. ## Release architecture This project ends in two finished artifacts, not one: 1. **Standalone model release** — this page. Pilot 3 is the strongest version of the project's evidence-faithful behavior that moved into the model itself, evaluated across multiple non-`probe_v0` surfaces. 2. **Benchmark-facing hybrid stack** — bridge `checkpoint-2` plus the `deterministic_v3` packet-local normalizer. That stack is the benchmark winner and the only configuration that clears every gate on frozen held-out `probe_v0`. The separation is deliberate. This page is for the standalone release you can download now. The benchmark winner is documented here because it explains the project's full result, not because those perfect `probe_v0` numbers belong to the adapter alone. ## Intended use Use this release for work that has to stay inside a fixed body of text: - bounded document QA with explicit evidence requirements, - claim verification and grounded QA from closed evidence packets, - policy, compliance, contract, and internal-document workflows where each answer must be justified from the provided text, - research on evidence-faithful reasoning and abstention behavior. ## Limitations - The downloadable artifact is the LoRA adapter only. The base model is required. - The `deterministic_v3` packet-local normalizer is not included in this download. The benchmark-winning configuration is adapter + normalizer, while the adapter alone reproduces the standalone-model results shown above. - Perfect `probe_v0` belongs to the benchmark-facing hybrid stack, not to this adapter alone. - Specialized for closed-packet reasoning, not open-ended chat or open-domain QA. - Frozen `probe_v0` item-level contents are intentionally not published with the release. ## Citation References: - Base model: [Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2](https://huggingface.co/Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2) - Datasets: [fever/fever](https://huggingface.co/datasets/fever/fever), [hotpotqa/hotpot_qa](https://huggingface.co/datasets/hotpotqa/hotpot_qa) - Technical brief (PDF): [evidence_faithful_reasoning_release_brief.pdf](./evidence_faithful_reasoning_release_brief.pdf) - Technical note: [technical_note_evidence_faithful_reasoning.md](./technical_note_evidence_faithful_reasoning.md) ```bibtex @misc{darcar0_evidence_faithful_reasoning_pilot_3_2026, title = {Evidence-Faithful Reasoning: Pilot 3}, author = {darcar0}, year = {2026}, howpublished = {Hugging Face model release}, url = {https://huggingface.co/darcar0/evidence-faithful-reasoning-pilot-3} } ```