
MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time Scaling
MaxProof demonstrates a shift in how frontier labs approach mathematical reasoning: rather than scaling model size alone, the framework orchestrates test-time computation across proof generation, verification, and refinement using tournament selection over candidate populations. The M3 model's achievement of gold-medal performance on IMO 2025 and USAMO 2026 signals that structured search and ensemble verification can push reasoning capabilities beyond what single-pass inference delivers. This matters because it reframes the scaling frontier from parameter count to inference-time orchestration, a pattern likely to influence how labs tackle other hard reasoning tasks.72


























