quotebound-27b / README.md
darcar0's picture
Upload README.md with huggingface_hub
e5e90d1 verified
|
Raw History Blame
8.29 kB
metadata
language:
  - en
license: apache-2.0
base_model:
  - Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2
datasets:
  - fever/fever
  - hotpotqa/hotpot_qa
library_name: transformers
pipeline_tag: text-generation
tags:
  - reasoning
  - evidence-grounding
  - attribution
  - fever
  - hotpotqa
  - lora
  - peft
  - distillation
  - research

Evidence-Faithful Reasoning

Pilot 3 is the standalone model release.

This adapter turns its reasoning-distilled 27B base model into an evidence-first reader for closed packets of text. I built it because I wanted a release where reasoning had to prove itself: every answer has to land on the right evidence, quote that evidence verbatim, and stop with Insufficient evidence. when the packet does not justify a claim. The result is the strongest standalone model from the project, packaged here as a LoRA adapter you can load directly tonight.

Resources & Guides

Fresh public holdout: standalone release vs bridge

Fresh 36-task mixed public holdout: the standalone release beats the earlier bridge model on task accuracy, evidence F1, and quote F1, while the packet-local normalizer lifts the full stack to 0.9093 quote F1.

Why this release exists

I built this project to force reasoning models to show their work in the only place that counts: the evidence itself. Fluent answers were not enough. I wanted a model that had to retrieve the right units, quote them exactly, and fail closed when the packet ran out. This page leads with the standalone release because it is the artifact you can load immediately, inspect directly, and use without reconstructing the whole benchmark stack.

At a glance

  • LoRA adapter on top of Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2.
  • Strongest standalone model from the project and the release I want people to download first.
  • On a fresh 36-task public holdout, raw task improves from 0.8611 to 0.8889, raw strict from 0.2222 to 0.4444, and raw quote F1 from 0.3343 to 0.6815 over the earlier bridge model.
  • Zero invalid outputs on every reported evaluation surface.
  • The project also produced a benchmark-winning hybrid stack, but that is a separate result described under Release architecture.

Quick start

Pilot 3 is a LoRA adapter. Load the base model and attach the adapter:

from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

base_id = "Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2"
adapter_id = "darcar0/evidence-faithful-reasoning-pilot-3"

tokenizer = AutoTokenizer.from_pretrained(base_id)
base = AutoModelForCausalLM.from_pretrained(base_id, device_map="auto")
model = PeftModel.from_pretrained(base, adapter_id)

The base model is 27B parameters, so load it in your usual quantization.

Prompt format

This release works best with an evidence-first prompt that makes the answer subordinate to the cited text. A minimal version:

You are answering from a bounded evidence packet only.

Work in this order:
1. Identify the smallest set of packet units that matters.
2. Copy exact quote(s) from those units.
3. Only then give the final answer.

Rules:
- No outside facts.
- Return valid JSON only.
- Every quote must be a verbatim substring of the cited unit.
- Do not paraphrase, ellipsize, or stitch quotes.
- If the packet is insufficient, the `answer` field must be exactly
  `Insufficient evidence.`

The model then writes a JSON object with this shape:

{
  "task_id": "<task id>",
  "label": "support|contradict|insufficient|null",
  "answer": "<one-sentence answer>",
  "evidence_ids": ["unit_id_1", "unit_id_2"],
  "quotes": [
    {"unit_id": "unit_id_1", "quote": "<exact quote>"}
  ],
  "justification": "<one short sentence tied to the cited evidence>"
}

Evaluation

Fresh 36-task mixed public holdout

A held-out slice of 18 FEVER verify-claim tasks plus 18 HotpotQA grounded-QA tasks, drawn from public sources and de-duplicated against every training, dev, and probe_v0 row.

Stack Task Strict Evidence F1 Quote F1
Bridge raw 0.8611 0.2222 0.8815 0.3343
Pilot 3 raw 0.8889 0.4444 0.9093 0.6815
Bridge + deterministic_v3 0.8611 0.5833 0.8815 0.8815
Pilot 3 + deterministic_v3 0.8889 0.5833 0.9093 0.9093

The standalone release beats the earlier bridge model on task accuracy, evidence F1, and quote F1 in both raw and normalized form, ties normalized strict, and roughly doubles raw quote F1 at the model level.

Fixed dev triage slice (21 tasks)

Stack Task Strict Evidence F1 Quote F1
Pilot 3 + deterministic_v3 1.0000 0.6190 0.8320 0.7095

Untouched 104-task Hotpot shadow slice

Pilot 3 raw improved quote-faithful behavior over the raw bridge model on this slice, and pilot 3 + deterministic_v3 matched bridge + deterministic_v3 at the system level. That surface remains a narrative parity result because the report does not publish per-metric cells for it.

Release architecture

This project ends in two finished artifacts, not one:

  1. Standalone model release — this page. Pilot 3 is the strongest version of the project's evidence-faithful behavior that moved into the model itself, evaluated across multiple non-probe_v0 surfaces.
  2. Benchmark-facing hybrid stack — bridge checkpoint-2 plus the deterministic_v3 packet-local normalizer. That stack is the benchmark winner and the only configuration that clears every gate on frozen held-out probe_v0.

The separation is deliberate. This page is for the standalone release you can download now. The benchmark winner is documented here because it explains the project's full result, not because those perfect probe_v0 numbers belong to the adapter alone.

Intended use

Use this release for work that has to stay inside a fixed body of text:

  • bounded document QA with explicit evidence requirements,
  • claim verification and grounded QA from closed evidence packets,
  • policy, compliance, contract, and internal-document workflows where each answer must be justified from the provided text,
  • research on evidence-faithful reasoning and abstention behavior.

Limitations

  • The downloadable artifact is the LoRA adapter only. The base model is required.
  • The deterministic_v3 packet-local normalizer is not included in this download. The benchmark-winning configuration is adapter + normalizer, while the adapter alone reproduces the standalone-model results shown above.
  • Perfect probe_v0 belongs to the benchmark-facing hybrid stack, not to this adapter alone.
  • Specialized for closed-packet reasoning, not open-ended chat or open-domain QA.
  • Frozen probe_v0 item-level contents are intentionally not published with the release.

Citation

References:

@misc{darcar0_evidence_faithful_reasoning_pilot_3_2026,
  title        = {Evidence-Faithful Reasoning: Pilot 3},
  author       = {darcar0},
  year         = {2026},
  howpublished = {Hugging Face model release},
  url          = {https://huggingface.co/darcar0/evidence-faithful-reasoning-pilot-3}
}