AI Agent · Agentic AI · VentureBeat AI
Catching 6, Snapped up on cost, measured on consistency
Compiled by KHAO Editorial — aggregated from 2 sources + 1 reference discovered via search. See llms.txt for citation guidance.
◌ Single Source
Enterprises buy evaluation tooling on economics and trust it on repeatability.
Key facts
- Technology/Software is the largest industry at 23%, followed by Retail/Consumer (15%), Healthcare/Life Sciences (12%), and Manufacturing (10%)
- Two-thirds of organizations (66%) already permit fully automated, zero-human-in-the-loop deployment for low-risk agents (34%) or are actively engineering their pipelines to allow it within twelve
- Responses are filtered to organizations with 100 or more employees (n=157), drawn from a single survey in June 2026; because this is one wave rather than a pooled multi-month sample, the report reads
- Only 36% report no such failure, and the remainder either run no pre-deployment evaluations (8%) or don’t track the root cause closely enough to know (6%)
Summary
Across 157 enterprises, organizations are granting AI agents more autonomy while trusting the evaluations meant to gate that autonomy less. This wave of VentureBeat Pulse Research examines how technical leaders measure agent performance: which reliability and evaluation platforms they use, how they select and trust them, what breaks in production, and how far they are willing to let agents run without a human in the loop. The central finding is an evaluation gap, the distance between the autonomy enterprises are granting their agents and the trust they place in the evaluations meant to govern it. What makes the gap consequential is the direction of travel. VentureBeat fielded this survey as part of its ongoing Pulse Research series, this survey, the Agentic Reliability & Evals tracker, focused on how technical leaders evaluate agent performance and reliability.