
You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories
Researchers have uncovered that reinforcement learning trajectories in LLMs exhibit extreme low-rank structure, with most performance gains captured by rank-1 approximations that scale linearly with training. This finding enables RELEX, a compute-efficient extrapolation method that predicts future model checkpoints from brief observation windows using linear regression. The discovery has immediate practical implications for RLVR training efficiency and suggests deeper geometric regularities in how LLMs adapt during reasoning-focused fine-tuning, potentially reshaping how labs approach scaling and checkpoint management.62



























