AI coding tools make mistakes 1 in 4 times
New research from the University of Waterloo reveals that top artificial intelligence coding tools make errors in approximately one out of every four attempts when generating structured outputs. Published as "StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs" in Transactions on Machine Learning Research and scheduled for presentation at ICLR 2026, the study highlights persistent challenges in ensuring AI-generated software responses are accurate and consistent within development workflows. As Large Language Models become more integrated into software engineering, the industry has shifted toward using structured outputs. Companies like OpenAI, Google, and Anthropic have introduced features that force AI responses to adhere to predefined formats such as JSON, XML, or Markdown. These formats are designed to make AI data easier for humans to read and for software systems to process automatically. However, the study demonstrates that current technology is not yet reliable enough for this purpose. The researchers evaluated 11 different Large Language Model models across 44 tasks and 18 distinct structured output formats. The results showed that even the most advanced commercial models achieved only about 75% accuracy in following these structural rules. Open-source models performed slightly lower, with accuracy rates closer to 65%. Dongfu Jiang, a Ph.D. student at Waterloo and co-first author of the paper, emphasized that the evaluation measured more than just syntax. The study assessed whether the outputs were not only formatted correctly but also accurate in solving the specific tasks assigned. According to the findings, while the models perform adequately on text-related tasks, they struggle significantly when the work involves generating images, video, or full websites. The research was a collaborative effort led by Jiang, undergraduate student Jialin Yang, and Assistant Professor Wenhu Chen. The project incorporated data annotations from 17 other researchers from Waterloo and international institutions, reflecting a growing trend where students transition from being data annotators to leading their own benchmarking studies. The core implication of the study is that while structured outputs represent a significant advancement for software development, they are not reliable enough to function without human oversight. Jiang noted that developers may deploy AI agents to handle specific tasks, but they must maintain significant supervision to catch errors. This limitation raises important questions about how confidently developers can rely on AI to automate critical parts of the coding process. The study underscores the gap between the theoretical potential of AI to streamline development and the practical reality of its current performance. Even with improved structural constraints, the failure rate remains high enough to prevent fully autonomous operation. As the field moves forward, the research suggests that human review remains an essential component of the development cycle, serving as a necessary safeguard against the errors that persist in even the most sophisticated models.
