New AI Agent can now build your knowledge base, connect channels and invite agents — all by chat Try it now
YundaDesk
PricingBlogWhy FreeHelp
Start freeLog in
Questions?Contact sales
Compare

How to Evaluate AI Support Accuracy Claims from Vendors

"95% resolution rate" sounds reassuring until you ask how it was measured. Here's a vendor-evaluation checklist for cross-border support teams to question fixed accuracy claims and check whether learning is actually testable, traceable, and reversible.

YundaDesk Team 2025-09-13Updated 2026-07-15 7 min read

The line that makes buyers relax fastest in a vendor pitch is “our AI resolves 95% of conversations.” It sounds reassuring, right up until you ask how that number was calculated. Whose conversations make up the denominator, and how long was the observation window — if a vendor can’t answer both, that number is marketing copy, not measurement.

Resolution rate depends heavily on your industry, average order value, return policy, and how complete your knowledge base is — the same platform produces a different number for every customer, which is why the only rate we quote carries its conditions: 60%+ of routine questions answered instantly out of the box, 90%+ resolution after a month of human coaching — never a bare, context-free number. What actually deserves verification is how the AI learns, whether it learned correctly, and whether that learning can be tested and rolled back. This is a checklist you can bring straight into a vendor evaluation call.

Why a single fixed resolution number should make you suspicious

Any vendor who quotes an isolated percentage should get three follow-up questions: what’s the denominator, what’s the scenario, and what’s the time window. “Resolution rate” can mean “AI replied and the customer didn’t follow up,” or it can mean “the customer explicitly rated it positive” — and both definitions conveniently exclude conversations that got escalated to a human, since by definition those are conversations the AI didn’t resolve alone.

Cross-border ecommerce makes this worse. A shipping-status question and a return-policy dispute are not remotely the same difficulty, and peak season traffic looks nothing like off-season traffic — one blended percentage cannot capture that variance. What’s worth watching during evaluation is whether answers are grounded, whether unanswerable questions get escalated, and whether every learning event leaves a trail — those are standards you can actually hold a vendor to on a call.

Published research is a good baseline, but verifying accuracy still comes down to your own historical conversations — treat those as the actual test set when you evaluate a vendor.

DATA

Beyond accuracy claims, look for repeatable productivity evidence

+14%Resolutions per agent after adding a generative AI assistant
+34%Resolutions per novice agent after adding a generative AI assistant
Source: Stanford/MIT "Generative AI at Work" study

Checklist one: how was this number actually calculated

Bring this table into your next vendor call:

What to ask A red flag answer A credible answer
Does the denominator include escalated conversations or exclude them “It’s roughly all conversations” Can break out the escalation share on the spot
Is the metric “no follow-up question” or “explicit positive rating” Can’t say Distinguishes the definitions and gives separate numbers for each
Which industry or customer segment produced this number “Our customers are all similar” Acknowledges that knowledge base maturity varies a lot in cross-border retail
Is the observation window a week or six months Only offers a single total Can show how the number trends over time

If a vendor can answer all four with specifics they can break down, the number was probably measured seriously. If every answer is a one-line deflection, the percentage was likely written by marketing as a pitch line.

Checklist two: what happens when the AI can’t answer

More important than any fixed number is what happens the moment the AI doesn’t know the answer or gets it wrong. The AI-first, human-backed boundary determines whether a customer gets left hanging when things go sideways.

  • When there’s no grounding in the knowledge base, does the AI honestly escalate, or does it fabricate a plausible-sounding answer?
  • When a customer asks for a human, is it a one-click handoff, or does the customer have to repeat themselves?
  • Are refunds, compensation, and price changes always routed to human approval, with no exceptions?
  • When a conversation escalates, can the agent see what the AI already said and the customer’s sentiment so far?

If any of these four comes back as “not sure” or “the vendor never mentioned it,” press further — don’t let a polished resolution number carry the conversation.

Checklist three: is “gets smarter with use” actually controlled

“Our AI learns automatically and gets smarter the more you use it” is another common pitch line, and it often means the model is quietly absorbing every historical conversation with no visibility into what it learned, when, or whether a mistake can be undone.

YundaDesk breaks learning into steps you can actually review: when the AI can’t answer, when an agent supplies the answer, or when an agent explicitly corrects the AI, those signals generate a pending learning suggestion. A manager reviews it in an approval queue before it takes effect, and only then does it become a skill, a knowledge entry, or a customer memory. Every suggestion traces back to the conversation that triggered it and the person who approved it, and it can be rolled back with one click if it turns out wrong. The full breakdown is in this piece on how AI keeps getting smarter.

Ask directly: can learning suggestions be exported as a list? Confirm who has permission to approve them, and whether a mistake means calling the vendor’s engineers or rolling it back yourself — both are worth nailing down before you sign.

Checklist four: is the knowledge base the real bottleneck

More often than not, “inaccurate AI” traces back to a knowledge base that’s stale, incomplete, or contradictory — the model itself carries less of the blame. What’s worth pinning down during evaluation is how the knowledge base gets maintained:

Maintenance method Good fit for Risk to watch
Document upload Stable policy and FAQ content Accuracy drops fast if documents aren’t kept current
Website or help-center crawling Content that’s already maintained elsewhere Ask about crawl frequency and coverage scope
Manual Q&A entry High-frequency but scattered support phrasing Needs ongoing input or it goes stale

The relationship between the knowledge base and AI accuracy matters far more than parameter count. If a vendor only talks about the model and never about how the knowledge base is built, updated, or flagged when it’s out of date, that’s a signal worth noting.

Checklist five: test it yourself — don’t just watch a demo

The real way to verify a claim is to run it against your own support scenarios — a vendor’s case study screenshots won’t tell you anything about yours. Ask to:

  1. Test knowledge base coverage against your own historical tickets or common questions.
  2. Deliberately ask a few questions the knowledge base doesn’t cover, and see whether the AI admits it doesn’t know or fabricates an answer.
  3. Have an agent actually edit a few AI drafts, then confirm those edits show up in a review queue and don’t quietly vanish into a black box.
  4. Run through a refund scenario and confirm the decision has to go through human approval — the AI doesn’t get to decide alone.

A vendor willing to run this test with you, versus one that only wants to show you a case study video, gives you your answer before you sign anything.

Checklist six: is billing tied to the accuracy claim

Some vendors bill by resolution outcome, which creates a hidden incentive to loosen the definition of “resolved” so the bill goes up. YundaDesk plans include AI credits with no per-conversation or per-outcome surcharge, so billing is predictable and there’s no incentive to inflate a resolution number to justify a higher invoice. Ask a vendor directly: if your accuracy number went down, would your bill change too? If the answer is yes, the fairness of that number is worth questioning.


A fixed resolution rate doesn’t survive three follow-up questions: what’s the denominator, what happens when the AI can’t answer, and is learning actually under control. What’s worth confirming before you sign is how the knowledge base gets maintained, whether high-risk actions like refunds always require human approval, and whether the vendor will let you test it against your own scenarios — that tells you more than any polished percentage. YundaDesk’s AI agent is built to make all three reviewable and auditable, with the specifics on cost on the pricing page.

Run this playbook in your own workspace

AI answers first, humans back up, every step is revertible — everything in this article can be put into practice in YundaDesk.