New Math Reveals Why AI Models Behave Similarly
Based on research by Dario Picozzi
Large language models are becoming eerily similar, yet we still do not fully understand the hidden structure they share or how to tweak one behavior without breaking another. New research reveals that the mathematical shape of a model’s predictions, known as Fisher-Rao geometry, holds the key to this mystery. By mapping how models assign probabilities to the next word, researchers have found a universal blueprint that connects behavior across vastly different AI architectures.
This geometry acts like a semantic map that remains consistent regardless of the underlying code. Across transformers, state-space models, and recurrent networks, the output geometries agree far more strongly than their internal activation patterns. This shared structure supports the transfer of semantic categories and aligns closely with human word choices. The alignment improves with scale, training, and predictive accuracy, suggesting that as models get smarter, their internal logic converges on a common, human-like framework.
The findings also expose surprising consequences for how models acquire facts. While pretraining corpus statistics can predict which facts a model will learn, randomized experiments show that deeper evidence substantially delays acquisition across every tested architecture. This delay challenges assumptions about how quickly models update their knowledge, revealing a complex interplay between data depth and learning speed that persists regardless of the model type.
Most importantly, this geometric insight enables precise control. It prescribes minimum-disturbance interventions that allow researchers to edit or steer models with minimal collateral damage. Updates learned on one prompt can transfer to unseen ones while better preserving existing behaviors than traditional methods. This reusable control mechanism improves steering, editing, attribution, dictionary learning, and fine-tuning, offering a powerful new tool for managing the evolving capabilities of large language models.