The gap between AI evaluations and their real-world performance is widening, and as enterprises charge toward greater autonomy for AI agents, this chasm poses significant risks. Despite mounting evidence that current evaluation methods fail to predict post-deployment performance accurately, companies are pressing forward, often with minimal human oversight.

Key Takeaways
- 50% of enterprises reported AI feature failures post-deployment that passed internal evaluations.
- Only 5% of organizations trust automated evaluations completely.
- Two-thirds are either allowing or planning to allow AI deployments without human intervention.
- The most popular evaluation tools are provider-native options, yet 17% use no dedicated tool.
- A majority plan to revisit or change their evaluation approach in the next year.
The Evaluation Gap: Autonomy Outpaces Assurance
In a survey of 157 enterprises, a critical theme emerges: the **evaluation gap**. This term describes the distance between the autonomy granted to AI agents and the confidence organizations have in the evaluations that are supposed to guarantee agent reliability. Notably, **half** of these enterprises experienced at least one instance where an AI feature, preloved in trials, floundered when faced with customer interaction.
Why Evaluations Fall Short
Automated evaluations are supposed to test potential AI issues before deployment. However, only **5%** of organizations expressed full trust in these evaluations. A prevailing concern is that these tests don’t align well with **real-world outcomes**. Imagine a sports coach relying solely on practice game performance to select their team—only to find the players choke under real match pressure. Similarly, AI evaluations tick the boxes in isolation but miss the mark in actual implementation.
Increasing Autonomy Despite Trust Issues
Paradoxically, even with waning trust in evaluations, a substantial two-thirds of enterprises are either implementing or evolving towards **zero-human deployment** systems. **34%** of these organizations already allow some low-risk deployments without human verification, while **32%** are prepping for such transitions. The trend indicates a burgeoning trust in autonomous operations despite current evaluation frailties.
The Disjointed Evaluation Landscape
Surprisingly, the evaluation landscape is fragmented. Many rely on **provider-native evaluations** or, in some cases, no structured tool at all. This reliance on basic or non-existent tools underscores a troubling **gap in assurance**, as critical evaluations are often left to be performed by in-house scripts or provider-built checker systems, which might not capture all necessary nuances.
Monitoring Focused on Functionality, Not Correctness
While once-deployed, organizations frequently evaluate AI performance based on system **functionality**—is the AI running smoothly, responding quickly, with minimal errors? However, this overlooks the content quality, leading to situations where a proficiently run system delivers incorrect outputs. Monitoring often centers on whether the system functions without addressing whether it’s providing **correct answers**.
Investment in Oversight Rising
Organizations are beginning to invest more in **human review workflows**, indicating an understanding of the necessity for human intervention in a landscape dominated by AI. Yet, while autonomy efforts increase, the appetite for oversight grows, demonstrating an essential balance between automation and human judgment.
Towards a Future of Trustworthy AI
The confluence of autonomy and assurance is a puzzle that enterprises strive to solve. As reliance on AI grows, so too will the need for evaluations that genuinely align with real-world applications. This evolving landscape suggests that while automation offers incredible efficiencies, the role of human oversight remains indispensable to avoid costly mistakes. Moving forward, innovation in evaluation methodologies will be critical in ensuring AI not only matches expectations in theory but consistently meets them in practice.
