Back to blog

AI Fails Physics Research 33% of Time

Based on research by Yigeng Jiang, Tengchao Yang, Taoyong Cui, Jiaxing Wan, Yuan Wang

Imagine an AI that doesn’t just answer questions but actively conducts scientific experiments, navigating complex physics and chemistry problems with the rigor of a human researcher. This is the promise of deep research agents, systems built on large language models to handle autonomous, multi-step scientific reasoning. Yet, despite their potential to accelerate discovery, we currently have no reliable way to measure if these digital scientists are actually up to the task.

Researchers have now introduced PhySciBench, a comprehensive benchmark designed to fill this critical evaluation gap. It consists of 200 expert-curated questions spanning physics and chemistry, reflecting real-world scientific workflows across six distinct task categories. The results were sobering. Even the most advanced baseline system, Gemini Deep Research, managed only a 33.5 percent accuracy rate. The analysis revealed three persistent flaws in current AI systems: they struggle with long reasoning chains, fail to transfer knowledge effectively between steps, and lack the ability to self-verify results using physical principles.

To address these shortcomings, the team developed DelveAgent, a modular multi-agent framework featuring adaptive planning, dual-granularity memory, and a hierarchical reflection mechanism grounded in physics. This specialized architecture proved significantly more effective than general-purpose models. Across four scientific benchmarks, DelveAgent improved accuracy by up to 7.5 percentage points while cutting inference costs to roughly one-third of the strongest baseline. The findings suggest that architectural specialization is key to making autonomous research reliable. With data and code now publicly available, this work sets a new standard for evaluating AI in the physical sciences, proving that tailored systems can outperform generic giants in complex scientific domains.

Source: arXiv:2606.18648

This post was generated by staik AI based on the academic publication above.