Failure taxonomy

Hidden failure modes in production AI agents.

The scariest failures are not obvious hallucinations. They are confident, reasonable-looking outputs that quietly violate data, policy, tool-use, or business assumptions.

KPI ambiguity

The user asks why sales went down, but the agent never clarifies gross vs net sales, region, or time period.

Time-window errors

The agent treats last quarter as calendar quarter when the business uses fiscal quarters.

Wrong grain

Store-level data is aggregated as region-level data, producing misleading conclusions.

Join errors

Sales and product tables are joined incorrectly, duplicating revenue or dropping rows.

Causal overclaiming

The agent says weather caused a decline without evidence in the available data.

Permission boundaries

A user with UK-only access receives US or global data in an answer or tool call.

Sparse data overinterpretation

The agent turns low-volume noise into confident explanations.

Multi-turn context failure

The user changes region or date midway and the agent silently uses stale context.

Tool failure recovery

A SQL query fails and the agent hallucinates an answer instead of repairing the query.

Metric definition mismatch

Revenue should mean net sales excluding VAT, but the agent uses gross sales.

Discuss a project

Need AI that is useful, measurable, and safe to run?

Tell us what you are building, evaluating, analysing, or trying to automate. We will help choose the right service path.

Prefer email? drew@agent-reliability.com