YanZhanPKU
Publish dLLM PRM Gap release
3c14026
|
Raw History Blame Contribute Delete
1.9 kB
metadata
license: mit
base_model: Dream-org/Dream-v0-Instruct-7B
tags:
  - dllm-prm-gap
  - discrete-diffusion
  - gsm8k
  - process-reward-model
  - dream-7b
pipeline_tag: text-classification
library_name: peft

dLLM PRM Gap · Bidirectional PRM

💻 Code   •   🤗 Collection   •   📄 Paper

News

  • 2026-09: Accepted at NeurIPS 2026.

Public adapter release for dLLM PRM Gap. bidirectional attention with mask-aware mean pooling over intermediate denoising states

This repository contains only trainable adapter parameters and the reward head. It does not include the base model. The release configuration is recorded in config.json; the arXiv paper defines the release scope and citation.

Files

  • adapter.safetensors — compact adapter weights
  • config.json — public base-model id, architecture, and provenance

Load

from huggingface_hub import snapshot_download
from prm.checkpointing import load_diffusion_prm

path = snapshot_download("YanZhanPKU/dLLM-PRM-Gap-bidir-dream7b")
model, report = load_diffusion_prm(
    checkpoint=path,
    model_path="Dream-org/Dream-v0-Instruct-7B",
    local_files_only=False,
)

Role: process-reward model used by PRM Guided, Hybrid, and the snapshot diagnostics.

Paper: https://arxiv.org/abs/2609.35472.