dLLM PRM Gap
Collection
Adapters and trajectory artifacts for matched-compute PRM guidance and ORM reranking in discrete diffusion reasoning. • 13 items • Updated • 1
How to use YanZhanPKU/dLLM-PRM-Gap-bidir-dream7b with PEFT:
Task type is invalid.
Public adapter release for dLLM PRM Gap. bidirectional attention with mask-aware mean pooling over intermediate denoising states
This repository contains only trainable adapter parameters and the reward head. It does not include the base model. The release configuration is recorded in
config.json; the arXiv paper defines the release scope and citation.
adapter.safetensors — compact adapter weightsconfig.json — public base-model id, architecture, and provenancefrom huggingface_hub import snapshot_download
from prm.checkpointing import load_diffusion_prm
path = snapshot_download("YanZhanPKU/dLLM-PRM-Gap-bidir-dream7b")
model, report = load_diffusion_prm(
checkpoint=path,
model_path="Dream-org/Dream-v0-Instruct-7B",
local_files_only=False,
)
Role: process-reward model used by PRM Guided, Hybrid, and the snapshot diagnostics.
Paper: https://arxiv.org/abs/2609.35472.
Base model
Dream-org/Dream-v0-Instruct-7B
Task type is invalid.