Why Your AI Code Fails to Create Images
Based on research by Haoyu Zhao, Zihao Zhang, Xudong Wang, Jiaxi Gu, Zuxuan Wu
You ask an AI to write code that draws a picture. The code runs perfectly. The result looks nothing like what you asked for. This frustrating disconnect is known as the Program-to-Visual gap, and it is the central problem a new framework called MaLiang-Harness aims to solve. It transforms visual generation from a one-shot gamble into a persistent, inspectable process of construction, verification, and revision.
The researchers define this gap as the discrepancy between a program that executes correctly and the visual output it produces. To bridge it, MaLiang-Harness organizes multimodal large language models into a loop where the evolving program, its history, and its verification share a common reference. It uses a Persistent Executable Generation state to keep track of context, a Traceable Generation Process to link edits to rendered evidence, and Revision-aware Editing and Verification to check the current state before finalizing. This system coordinates planning, execution, and visual feedback across different rendering backends, ensuring the final image or video matches the intent.
The results reveal a startling mismatch between general AI capability and actual visual generation performance. When tested on MaLiang-IBench and MaLiang-VBench, GPT-6-Astra achieved a perfect 100% generation success rate, with 96.0% of image tasks and 76.9% of video tasks meeting all quality thresholds. However, the study shows that models with similar general benchmark scores can differ substantially in their ability to satisfy specific visual requirements. This suggests that standard metrics are poor predictors of how well an AI can translate executable code into accurate visual outcomes.
MaLiang-Harness provides a systematic basis for understanding how models convert code into visuals, exposing both the potential of programmable generation and the limitations of current evaluation methods. By making the process traceable and revision-aware, it offers a clearer path toward reliable visual creation. The framework is now available for researchers to explore the nuances of this critical translation layer in AI-generated media.