
Reasoning without Gold Standards: A Proxy-Judge Theory of Autoformalization
Autoformalization, the task of converting informal reasoning into machine-checkable formal proofs, has hit a scaling wall: expert-validated reference solutions don't exist at meaningful scale, and multiple valid formalizations exist for single arguments. This paper proposes a reference-free evaluation framework using structured property checks across multiple dimensions instead of exact-match scoring. The shift matters because it reframes how AI systems can be trained and evaluated on open-ended reasoning tasks where ground truth is ambiguous or expensive. This pattern extends beyond mathematics to any domain where correctness admits multiple valid solutions, reshaping how researchers approach training data scarcity in formal reasoning.62




























