Measuring AI Resolution Rate Honestly (Deflection Is Not Resolution)

Measuring AI Resolution Rate Honestly (Deflection Is Not Resolution)

TL;DR

Most published AI resolution rates count conversations the AI merely ended. Honest measurement separates three numbers — genuine resolution (answered, no re-open within 72 hours, no human requested), clean escalation with full context, and deflection — then uses the escalated topics to fill content gaps and gives leadership numbers it can plan headcount around.

Every AI support vendor publishes a resolution rate. Almost none of them publish the definition behind it. That gap is where trust goes to die: if a customer closes the tab out of frustration, most dashboards count it as "resolved by AI."

We think the honest definition has three parts.

Resolution means the question was answered

A conversation is resolved when the customer's question was actually addressed — not when the conversation ended. In Chattering, a conversation only counts toward AI resolution when all three are true:

  • The AI produced an answer grounded in your knowledge base or connected sources, not a generic reply.
  • The customer did not re-open the conversation or start a new one on the same topic within 72 hours.
  • The customer did not ask for a human at any point after the answer.

Everything else is a deflection, and we label it that way in Reports. Deflections are not worthless — sometimes the customer really did just need a link — but they are a different number, and mixing the two inflates the headline.

Escalation is a feature, not a failure

The fastest way to poison a support experience is an AI that refuses to hand off. We treat a clean escalation — full context, conversation summary, and the customer never repeating themselves — as a successful AI outcome. It shows up in its own column.

When you review your numbers weekly, the question is not "how do we push resolution from 62% to 80%?" It is "which topics escalate most, and what content or workflow would let the AI genuinely resolve them?"

The 72-hour re-open window

The single biggest source of fake resolution is measuring at close time. A customer who got a wrong answer often leaves quietly and comes back the next day angrier. We attribute that return conversation to the original one, which retroactively flips it from resolved to re-opened.

This makes week-one numbers look worse than a competitor demo. It also makes month-three numbers real, which is what you actually run a support team on.

What to do with the honest number

Once deflection, resolution, and escalation are separated, three moves follow:

  1. Fill the content gaps. Every escalated topic with volume is a help-center article waiting to be written. The knowledge base editor shows which articles the AI cites most — and which questions found nothing to cite.
  2. Tune the handoff threshold. If the AI cannot ground an answer in your sources, let it hand off early instead of improvising. A 5% lower AI-resolution number that saves ten angry customers a week is a good trade.
  3. Report all three numbers to leadership. "AI resolves 55%, cleanly escalates 30%, deflects 15%" is a statement a CFO can plan headcount around. "AI handles 85%" is not.

Honest measurement is slower to look impressive and faster to become true. If you want to see the three-column report on your own traffic, start on the free plan — the metrics work the same on day one.

Frequently asked questions

What counts as an AI-resolved conversation?

The AI answered from your knowledge base, the customer didn't re-open the topic within 72 hours, and never asked for a human after the answer. Everything else is logged as a deflection or an escalation.

Why separate deflection from resolution?

A deflection ends a conversation; a resolution answers it. Mixing them inflates the headline number and hides the topics your content doesn't cover yet.

Is a handoff to a human a failure?

No — a clean escalation with full context is a successful AI outcome. It's tracked in its own column so you can see which topics need better content versus more staffing.

Why do week-one numbers look lower than vendor demos?

The 72-hour re-open window retroactively flips bad answers from resolved to re-opened. Demos usually measure at close time, which counts frustrated exits as wins.

Keep reading

Answer it once. Chattering remembers.

Free trial · no credit card required