AI Support

How to Measure AI Support Quality (Without Fooling Yourself)

Short answer

Measure AI support with genuine resolution rate as the headline, reopen rate and satisfaction as guardrails, and a weekly sample of twenty answers checked against source documentation. AI errors are systematic, so aggregate metrics miss them until hundreds of customers are affected.

Key takeaways

  • AI errors repeat, so one unnoticed pattern reaches everyone before metrics move.
  • Sample deliberately: escalations, reopens, and low scores, not just random.
  • Check claims against the source document, not for tone.
  • A metric you pay on is the metric that needs the most scrutiny.

Quality processes designed for human agents miss the failure mode that matters most in AI support. Here is a framework that does not.

Why AI quality is a different problem

A human agent gives one incomplete answer. They might catch it, a colleague might, or the customer replies and it gets fixed.

An ungrounded system gives the same incomplete answer to everyone who asks that question, hundreds of times, before any aggregate metric moves enough to notice. The errors are systematic rather than random, which means the detection method has to be different.

That is the whole argument for weekly sampling. It is not diligence theatre, it is the only mechanism that catches a pattern early.

The three numbers

Genuine resolution rate. The share of conversations that ended with a correct answer and no repeat contact within a set window. Define it before you measure it, and exclude abandonment and handoffs explicitly.

Reopen rate. The share of resolved conversations that come back. This is the honesty check on everything else, and it is the metric most teams do not track at all.

Satisfaction, segmented. By topic and by whether the issue was resolved or escalated. Aggregate satisfaction hides the topics where the system is failing.

Report all three together, always. Resolution rate alone can be inflated by loosening the definition. Paired with reopen rate it cannot.

The weekly sample

Twenty answers a week. Not random.

Weight the sample toward:

  • Escalations. Why did it stop? Should it have stopped sooner or later?
  • Reopened conversations. The customer came back. Read the original answer and find what was missing.
  • Low satisfaction scores. Small in number, high in signal.
  • Your highest-volume topic. One systematic error here affects the most people.

Then add five genuinely random ones to catch what the weighting misses.

What to check

Not tone. Fluency and correctness are unrelated in generated text, and a reviewer reading for tone will pass a confident invention every time.

Check each factual claim against the source document it came from. Specifically:

  • Is every number, date, and policy term supported by a source?
  • Is the answer complete, or does it stop before a condition the customer will hit?
  • Did it state a limitation it should have?
  • Would a customer acting on this answer get the outcome they expected?

That last question is the one that matters. An answer can be technically accurate and still lead someone into a wall.

The four failure types

Categorise every failure you find. The category tells you the fix, and it is almost never a model problem.

Content gap. Documentation does not cover it. Write the article.

Content contradiction. Two sources disagree and the system picked one. Fix the content.

Incomplete retrieval. The answer came from half a procedure. Fix the article structure so procedures stay whole.

Missing escalation. The question needed judgement. Tighten the rule.

Tracking the distribution over time is useful in itself. If content gaps dominate, your roadmap is clear. If missing escalations dominate, your rules are too loose.

The metric that should worry you

Rising resolution rate with rising reopen rate.

That combination means the definition of resolution is loosening or the answers are getting shallower. It is the clearest signal of a programme going wrong, and it will not show up in any single metric.

Similarly, rising deflection with flat satisfaction means customers are giving up rather than being helped, which is why deflection is the wrong headline metric.

Audit what you are billed on

If you pay per resolution, the vendor counts the resolutions. That is a reasonable arrangement and it deserves verification.

Ask for the ability to audit a sample of billed resolutions and actually use it, quarterly. Check that each one meets the definition in the contract: the issue closed, the customer did not come back, no human was involved.

A vendor unwilling to allow this is asking for trust the arrangement does not earn. A vendor who allows it and is never asked is being given a pass they did not request.

Reporting to leadership

Three numbers and one sentence:

  • Genuine resolution rate, with the definition stated.
  • Reopen rate.
  • Satisfaction.
  • One sentence on what the weekly sample found.

Do not report deflection, containment, or conversations handled. They are activity metrics that count abandonment as success, and a programme reported on them will eventually be contradicted by churn data in a way that costs you credibility.

The half hour that pays for itself

Twenty answers a week is roughly thirty minutes. It is the highest-return quality activity available in AI support, and it is the first thing dropped when people get busy.

Put it in a calendar, give it an owner, and treat the findings as a standing agenda item. Every team that does this finds something in the first month.

Frequently asked questions

How many AI answers should I review?
Twenty a week, weighted toward escalations, reopened conversations, and low satisfaction scores rather than purely random. That finds systematic issues far faster than a large random sample.
What should I check when reviewing an AI answer?
Whether each factual claim is supported by the source document it cites. Reviewing for tone catches nothing, because fluency and correctness are unrelated in generated text.
Why do AI errors need different quality processes?
Because they are systematic rather than random. A human agent makes one incomplete answer. An ungrounded system makes the same incomplete answer hundreds of times before any aggregate metric moves.
What is the single most important AI support metric?
Reopen rate. It is the honesty check on resolution rate, deflection, and every cost saving claim, and it is the metric most teams do not track at all.
ai support qualitysupport qaresolution ratesupport metrics

Resolve, don't deflect.

See Fidiora resolve a ticket, capture a lead, and keep the bill predictable.

See Pricing