Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection Paper • 2609.35932 • Published 6 days ago • 7
Why Deterministic PRM Guidance Underperforms in Discrete Diffusion Reasoning Paper • 2609.35472 • Published 6 days ago • 2
Byte Authority Collection Byte Authority code and evaluation records for reserved-token prompt injection. • 2 items • Updated 4 days ago • 1
dLLM PRM Gap Collection Adapters and trajectory artifacts for matched-compute PRM guidance and ORM reranking in discrete diffusion reasoning. • 13 items • Updated 5 days ago • 1
SLCA-GRPO Collection Segment-Locked Credit Assignment for tool-calling RL — processed data splits and the reference implementation. • 2 items • Updated 9 days ago • 1
Entropy-Valley Collection Training-free length selection for masked diffusion MT (EMNLP 2026). Code: github.com/Entropy-Valley/Entropy-Valley • 5 items • Updated 9 days ago • 3
SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL Paper • 2609.29050 • Published 10 days ago • 13
RefineSVG: Visual Feedback-Driven Reinforcement Learning for Image-to-SVG Generation Paper • 2607.27699 • Published Jul 30 • 1
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Paper • 2608.08491 • Published Aug 9 • 1
Length-Adaptive Decoding for Masked Diffusion Machine Translation Paper • 2608.22274 • Published Aug 23 • 4
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them Paper • 2509.21117 • Published Sep 25, 2025 • 30
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Paper • 2508.06026 • Published Aug 8, 2025 • 15