DACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language Models

Apple Machine Learning Research Version 1 original current

Imported from official source

Diffusion large language models are a compelling alternative to autoregressive models, yet existing RL methods for diffusion treat all denoising steps as equally important and rely on biased, high-variance likelihood estimates. …

This version

Version
1 of 1
Recorded
September 20, 2026 19:52
Change
Initial
Content hash
deeb64ba047a654f308b8db96a04fca1
All versions
Revision history

Officially records where a publication came from, not whether it is true. Imported records are reproduced from an organization's own official source.