AI Agent Reliability Audit
We tested 75 realistic workflows across KPI analysis, SQL generation, permission boundaries, and multi-turn reasoning.
75production-like workflows
68%overall pass rate
12critical/high findings
Critical failures
Unauthorized data exposure3/20 permission testsUser roles with region-limited access received global or restricted-region outputs.
Silent SQL aggregation errors11/35 analytics testsIncorrect joins and grains caused plausible but materially wrong KPI values.
Unsupported causal claims14/25 root-cause testsThe agent attributed business changes to causes not supported by retrieved data.
Most common failure modes
18%wrong metric definition
15%incorrect time window
22%hallucinated root cause
31%failed clarification
9%prompt/model regression
Recommended fixes
- Enforce metric dictionary lookup before SQL generation.
- Add permission-aware retrieval and tool-call filters.
- Require uncertainty statements when evidence is insufficient.
- Add SQL validation for joins, grains, and aggregation risks.
- Introduce the regression suite into CI/CD for prompt and model changes.
Discuss a project
Need AI that is useful, measurable, and safe to run?
Tell us what you are building, evaluating, analysing, or trying to automate. We will help choose the right service path.
Prefer email? drew@agent-reliability.com