Modelwire
Subscribe

Supervised finetuning outperforms expectations with strategic data sampling

Researchers challenge the conventional hierarchy between supervised finetuning and reinforcement learning in posttraining, demonstrating that SFT can achieve stronger generalization than previously assumed when data distribution is optimized appropriately. Rather than redesigning loss functions, the work focuses on curating training data to capture on-policy learning benefits while retaining access to offline expert trajectories. This finding reshapes assumptions about capability transfer and catastrophic forgetting, suggesting practitioners may be underutilizing SFT's potential and that the frontier between SFT and RL methods is less rigid than industry consensus implies.

Modelwire context

Explainer

The paper's actual contribution is narrower than it appears: it's not that SFT is secretly powerful, but that SFT with deliberately constructed on-policy data can match RL-like generalization without RL's computational cost. The key move is data design, not algorithm redesign.

This sits directly alongside the sampling-focused work from late September. The 'Strategically Diverse Sampling' paper (Sept 25) showed that training data quality matters more than quantity when you curate for solution diversity rather than correctness alone. This new result extends that insight: if you construct SFT data to include the model's own on-policy outputs (not just expert trajectories), you recover benefits practitioners thought required RL. The 'Explore Broadly, Reason Sharply' work (Sept 29) made a similar point about sidestepping RL via smarter sampling. Together, these three papers suggest a pattern: the SFT-vs-RL debate was partly a data curation problem masquerading as an algorithmic one.

If teams report that on-policy SFT data constructed this way degrades performance on held-out benchmarks not seen during curation (GPQA Diamond, MATH-500), that signals the gains are overfitting to the training distribution rather than genuine generalization. Conversely, if the method holds on truly out-of-distribution tasks, it validates the claim that data design, not loss functions, was the bottleneck.

Coverage we drew on

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. arXiv cs.LG originally reported this story as “Finetuning with Sampling: SFT Learns Better Than You Think”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

Related

Self-training via strategic diversity outperforms correctness-only filtering

arXiv cs.CL·

Parallel tempering replaces RL for efficient small model reasoning

arXiv cs.LG·

Language agents improve through self-explanation without reinforcement learning

arXiv cs.CL·
Supervised finetuning outperforms expectations with strategic data sampling · Modelwire