
Delta distillation transfers reasoning gains without reward models
Researchers propose delta distillation, a refinement to on-policy reinforcement learning that sidesteps reward model bottlenecks by extracting reasoning gains directly from teacher models. Rather than copying output distributions, the method captures the delta between a tuned model and its pre-instruction baseline, isolating learned reasoning patterns for transfer. This addresses a real friction point in post-training: reward models often constrain signal quality. The approach matters for teams scaling reasoning-focused LLMs, as it offers a more granular supervision path that could improve efficiency in capability transfer without external reward annotation.58

























