Back to blog

This Simple Trick Beats Gold-Medal Math

Based on research by Yafu Li, Runzhe Zhan, Haoran Zhang, Shunkai Zhang, Yizhuo Li

Gold-medal performance in the world’s toughest math and physics competitions used to require massive, specialized systems. Now, researchers have shown that a surprisingly simple scaling recipe can turn a standard model into an Olympiad-level solver. This breakthrough challenges the assumption that elite reasoning demands complex, fragmented architectures.

The team developed a unified method to upgrade reasoning models. They started by training the model to rigorously check its own work using a reverse-perplexity curriculum. This instilled a habit of self-correction. Next, they used a two-stage reinforcement learning process. The first stage focused on verifiable rewards, while the second refined proof-level logic. Finally, they boosted performance with test-time scaling, allowing the model to think longer and more deeply during problem-solving.

The result is a 30-billion-parameter model named SU-01. It achieves gold-medal-level scores on the International Mathematical Olympiad and International Physics Olympiad. What makes this striking is its efficiency. The model was trained on just 340,000 short trajectories and 200 reinforcement learning steps. Despite this modest training data, it handles problems with reasoning chains exceeding 100,000 tokens. It also generalizes well to scientific domains beyond math and physics.

The key takeaway is that rigorous reasoning does not need to be complicated. By focusing on self-checking behaviors and scalable training steps, researchers have proven that high-level problem-solving can be achieved with a simple, unified approach. This democratizes access to elite reasoning capabilities, showing that method matters more than sheer scale.

Source: arXiv:2605.13301

This post was generated by staik AI based on the academic publication above.