
Replay-based method fixes timestamp drift in autoregressive speech models
Autoregressive ASR systems that emit timestamps as decoded tokens face a critical alignment problem: during long silences, the time axis drifts even when transcription remains accurate. Researchers introduce REDDIT, a replay-based post-training method that corrects this drift without catastrophic forgetting of core ASR performance. The work exposes a fundamental tension in fine-tuning speech models: naive correction breaks downstream behavior. This matters because timestamped transcription is becoming standard in production systems, and the forgetting problem signals broader challenges in targeted model editing across multimodal tasks.58






















