The line that gets the most nods in a sales call is “our AI resolves 95% of conversations on its own.” It sounds like you could cut your support headcount in half tomorrow. But ask one follow-up question, “how is that calculated?”, and the answer usually gets vague fast. The number is easy to engineer: count every AI reply as a resolution, call one response a closed loop, and hide later customer follow-ups outside the reporting window.
Up front: we don’t promise a fixed resolution rate
YundaDesk will not put “AI resolution rate no lower than X%” in a contract. That’s not evasiveness. It’s because the real number depends on your knowledge base quality, how complex your products are, and who your customers are. Anyone who hands you a universal promise before they’ve seen your data either hasn’t tested it seriously, or is showing you a formula built to look good rather than to be true. The conversation worth having isn’t “what number will you promise,” it’s “how do we measure it, can it keep improving, and can every single case be checked.”
First, separate Handled, Resolved, and Stayed Resolved
Many “automation rate” claims look strong because three different definitions are blended together. At minimum, separate these layers:
| Measure | Plain-English meaning | What it proves | What it does not prove |
|---|---|---|---|
| Handled | AI picked up the conversation and replied | How much front-door volume AI covered | Whether the customer’s issue was solved |
| Resolved | No handoff, no negative feedback, and the conversation closed at that moment | AI may have completed one support interaction | Whether the customer came back later |
| Stayed Resolved | The same issue was not reopened within a 24-72 hour observation window | The interaction probably stayed solved | That every complex case should be automated |
A more honest independent resolution rate should be close to Stayed Resolved, not Handled. If a vendor calls “AI replied” the same thing as “AI resolved,” they are reporting reception volume, not resolution quality. Automation rate is not customer experience by itself. It has to be calibrated by whether customers come back, ask for a human, or leave negative feedback.
From AI-handled volume to truly resolved cases
Formulas you can reuse
Define the observation window first, then define numerator and denominator. This set works well for a monthly dashboard, a pilot readout, or a vendor comparison:
| Metric | Suggested formula | How to read it |
|---|---|---|
| AI handled rate | AI-replied conversations / AI-eligible conversations | Measures AI coverage, not resolution. |
| AI first-contact resolution | AI conversations with no handoff and no later agent supplement / AI-replied conversations | Measures whether the first AI interaction was complete enough. |
| True AI resolution rate | AI conversations with no handoff, no negative feedback, and no same-issue reopen inside the observation window / AI-replied conversations | The closest main metric for “AI resolved it independently.” |
| Deflection rate | Conversations that would likely have entered the agent queue but were closed by AI / conversations that would likely have entered the agent queue | Measures how much agent load AI reduced. |
| Handoff rate | AI conversations later transferred to a human / AI-replied conversations | A high rate points to knowledge, permission, or confidence-boundary gaps. |
| 72h re-contact rate | Same-issue re-contacts within 72 hours / conversations marked resolved by AI | Calibrates how much “apparently resolved” volume leaked later. |
| AI-resolved CSAT | Satisfied ratings among AI-resolved conversations / rated AI-resolved conversations | Measures the customer experience of AI-resolved cases. |
| Knowledge gap rate | Conversations that generated a learning suggestion because knowledge was missing, policy was unclear, or the answer was corrected / AI-replied conversations | Shows where the knowledge base still owes work. |
The important point: do not rely on one headline percentage. Handled rate tells you whether AI covered volume, handoff rate and re-contact rate tell you whether that coverage held, AI-resolved CSAT tells you whether customers accepted the result, and knowledge gap rate tells you what to improve next.
What actually counts as “resolved”
A conversation only counts as genuinely resolved by AI on its own if it meets all of these:
- The customer did not reopen the same issue within a follow-up window (say, 24-72 hours)
- No agent stepped in or added anything to the conversation
- The customer did not leave negative feedback or explicitly ask for a human
- The question itself had a clear boundary (like tracking a shipment or checking a return policy), not something that needed human judgment on a messy edge case
Put differently: if a customer follows up with “still not quite clear, can someone look at my actual order,” that thread shouldn’t count as an independent resolution, no matter how confident the AI’s earlier reply sounded. Resolution rate is supposed to measure where the customer’s experience ended up, not whether the AI said something.
Three common ways the number gets inflated
None of these three tricks looks like outright fraud on its own. Stacked together, they can dress up a system with a real resolution rate of 60% as “95% automated.” So when you’re handed a resolution rate figure, ask what counts as the numerator and denominator, how long the observation window is, and whether repeat questions were excluded.
Four dimensions worth checking instead of one headline number
Rather than arguing over a single percentage, break it into four things you can actually verify:
| Dimension | Question to ask |
|---|---|
| Coverage | Is this rate across all conversations, or just a handful of easy scenarios? |
| Observation window | How long does a customer need to stay away before it counts as “not reopened”? |
| Audit path | Has a human sampled and reviewed conversations marked “resolved” by the AI? |
| Trend, not snapshot | Is this a launch-day number, or a curve over the last quarter? |
The last one is the most overlooked and the most important: a single snapshot means nothing; the direction it’s moving in does. An honest curve that crawls from 55% to 72% over a few months is worth more trust than a 95% figure that appeared out of nowhere, and it actually tells you whether your knowledge base and team’s effort are paying off.
Honest resolution rate should be a trend, not a pitch number
Why controlled learning is what makes the number improve for real
YundaDesk’s signature “gets smarter with use” is the actual mechanism behind a real resolution-rate improvement, not a formula tweak. Whenever the AI can’t answer, an agent fills in the answer, or an agent hits “correct the AI,” that experience becomes a pending learning suggestion that lands on the owner’s review desk. Only once you approve it does it become a skill or a piece of knowledge, and only then does the next similar question have a chance of being handled independently.
That means the improvement in resolution rate is explainable. Last month the AI couldn’t answer “how do returns work for overseas warehouse orders,” an agent answered it, you approved that learning suggestion, and this month the independent resolution rate for that question type should tick up. You can point to the exact learning record behind it instead of a vague “we upgraded the model.” For more on how that loop actually works, see how “gets smarter with use” works.
Traceability: every “resolved” conversation can be pulled back up
The other half of honest measurement is traceability. If a vendor can’t let you trace “resolution rate” back to individual conversation logs, the number can’t really be verified, and it can’t be used to improve anything either.
In YundaDesk’s shared workspace, every AI reply, every handoff to an agent, and every case where an agent stepped in and finished the answer is a fully logged conversation record you can filter by time, channel, or issue type and review. That’s also why high-risk cases, such as refunds, compensation, and price changes, never get counted toward “AI resolved.” Those always go through human approval; AI never executes them on its own, and they shouldn’t be credited to an “automated resolution” scoreboard either. For where that boundary actually sits, see how AI-first, human-backed responsibility is split.
A checklist for checking your own number
If you want to check whether your team’s, or a vendor’s, resolution rate is actually honest, run it through this:
- Does “resolved” require the customer not reopening the same issue within 24-72 hours
- Does the denominator cover all channels and question types, not just the easy scenarios
- Has a human sampled and reviewed conversations marked “resolved”
- Can you see handled rate, handoff rate, 72h re-contact rate, and true AI resolution rate together
- Are refunds, compensation, and price changes excluded from the “AI resolved” count
- Once a learning suggestion is approved, can you verify the resulting change in resolution rate for that question type
Running through this list will tell you more than memorizing any “industry average resolution rate” ever will.
An honest resolution rate should survive you pulling ten random conversations and checking them yourself. If a vendor won’t let you do that, the number probably isn’t worth much.