Back to blog

AI Passes Tests But Fails Real Jobs

Based on research by Yiyou Sun, Xinyang Han, Weichen Zhang, Yuanbo Pang, Tianyu Wang

AI systems are crushing standardized tests, yet they remain frustratingly useless in actual professional environments. This disconnect suggests the problem isn’t intelligence, but how we measure it. We have been grading students on multiple-choice questions while ignoring whether they can actually do the job.

Researchers have introduced a new benchmark called Agents' Last Exam, or ALE, to fix this. Unlike previous tests that check for isolated facts, ALE evaluates AI agents on long-horizon, real-world tasks with verifiable outcomes. Developed with over 250 industry experts, it covers 1,000 tasks across 13 industry clusters, focusing on non-physical work defined by standard occupational taxonomies. The goal is to measure sustained performance on economically valuable workflows rather than quick, artificial wins.

The results reveal a stark reality. Even on the hardest tier of tasks, mainstream AI models achieve an average full pass rate of just 2.6 percent. This tiny number proves that current benchmarks are failing to predict real-world utility. The gap between benchmark success and actual economic impact is wide, and most systems are nowhere near saturation.

ALE is designed as a living benchmark, continuously growing as new workflows are onboarded. It aims to shift the focus from leaderboard rankings to GDP-relevant impact. If we want AI to matter in the professional world, we need to stop testing it on puzzles and start testing it on work.

Source: arXiv:2606.05405

This post was generated by staik AI based on the academic publication above.