On-Policy AI Learning Is Overrated
Based on research by Julianna Piskorz, Antonin Berthon, Mihaela van der Schaar
We have long been told that on-policy learning is the gold standard for training AI models, promising better generalization and less catastrophic forgetting. But what if this widely held belief is overstated? A new systematic study challenges the assumption that using a model’s own generated data is inherently superior, revealing a more complex reality behind how these systems actually learn.
Researchers conducted a controlled experiment to isolate the specific impact of rollout policies during the distillation process, where a smaller model learns from a larger one. By independently varying the policy, the direction of the Kullback-Leibler divergence, and the learning rate across the Llama3 and Qwen2.5 model families and reasoning tasks spanning scientific, medical, and arithmetic domains, they sought to separate the effects of policy from other variables. This rigorous approach allowed them to pinpoint exactly which factors drive performance improvements or degradation in the final model.
The results were surprising. Contrary to popular belief, the choice of rollout policy did not play the central role many assumed. Instead, the direction of the KL divergence was the primary driver of task performance and output coverage, while the learning rate governed how much the model forgot or updated its parameters. Forward KL proved remarkably robust to changes in policy, maintaining strong and stable performance. Reverse KL, however, was far more sensitive and favored data generated by the student model itself. While on-policy data did help with harder variants of the Countdown arithmetic task initially, this advantage often disappeared after further reinforcement learning.
The takeaway is clear: on-policy rollouts are not a magic bullet. Their value depends critically on the specific objective, the evaluation setting, and the optimization hyperparameters used. Rather than blindly adopting on-policy methods, practitioners must carefully consider the KL direction and learning rate, as these factors often matter more than the source of the training data. This nuanced understanding could reshape how we approach efficient and effective model training in the future.