Example deliverable

AI Agent Reliability Audit

We tested 75 realistic workflows across KPI analysis, SQL generation, permission boundaries, and multi-turn reasoning.

75production-like workflows
68%overall pass rate
12critical/high findings

Critical failures

Unauthorized data exposure3/20 permission testsUser roles with region-limited access received global or restricted-region outputs.
Silent SQL aggregation errors11/35 analytics testsIncorrect joins and grains caused plausible but materially wrong KPI values.
Unsupported causal claims14/25 root-cause testsThe agent attributed business changes to causes not supported by retrieved data.

Most common failure modes

18%wrong metric definition
15%incorrect time window
22%hallucinated root cause
31%failed clarification
9%prompt/model regression

Recommended fixes

  • Enforce metric dictionary lookup before SQL generation.
  • Add permission-aware retrieval and tool-call filters.
  • Require uncertainty statements when evidence is insufficient.
  • Add SQL validation for joins, grains, and aggregation risks.
  • Introduce the regression suite into CI/CD for prompt and model changes.

Discuss a project

Need AI that is useful, measurable, and safe to run?

Tell us what you are building, evaluating, analysing, or trying to automate. We will help choose the right service path.

Prefer email? drew@agent-reliability.com