New AI Agent can now build your knowledge base, connect channels and invite agents — all by chat Try it now
YundaDesk
PricingBlogChannels
Start freeLog in
Questions?Contact sales
Compare

How to Evaluate AI Support Accuracy Claims from Vendors

\"95% resolution rate\" sounds reassuring until you ask how it was measured. Here's a vendor-evaluation checklist for cross-border support teams to question fixed accuracy claims and check whether learning is actually testable, traceable, and reversible.

YundaDesk Team 2025-09-13Updated 2026-07-10 7 min read

The line that makes buyers relax fastest in a vendor pitch is “our AI resolves 95% of conversations.” It sounds reassuring, right up until you ask how that number was calculated. What’s the denominator? Whose conversations? Over what window? If the answer is vague, that number is marketing copy, not measurement.

We don’t promise a fixed resolution rate, because resolution depends heavily on your industry, average order value, return policy, and how complete your knowledge base is. The same platform produces a different number for every customer. What we do guarantee is that how the AI learns, whether it learned correctly, and whether that learning can be tested and rolled back are all verifiable. This is a checklist you can bring straight into a vendor evaluation call.

Why a single fixed resolution number should make you suspicious

Any vendor who quotes an isolated percentage should get three follow-up questions: what’s the denominator, what’s the scenario, and what’s the time window. “Resolution rate” can mean “AI replied and the customer didn’t follow up,” or it can mean “the customer explicitly rated it positive” — and both definitions conveniently exclude conversations that got escalated to a human, since by definition those are conversations the AI didn’t resolve alone.

Cross-border ecommerce makes this worse. A shipping-status question and a return-policy dispute are not remotely the same difficulty, and peak season traffic looks nothing like off-season traffic. One blended percentage cannot capture that variance. The honest position is: no promise of a fixed resolution rate, but a guarantee that answers are grounded, unanswerable questions get escalated, and every learning event leaves a trail.

Published research is more useful for judging whether AI has productivity potential than for accepting a vendor’s accuracy claim. Treat public numbers as a baseline, then use your own historical conversations as the test set.

DATA

Beyond accuracy claims, look for repeatable productivity evidence

+14%Resolutions per agent after adding a generative AI assistant
+34%Resolutions per novice agent after adding a generative AI assistant
Source: Stanford/MIT "Generative AI at Work" study

Checklist one: how was this number actually calculated

Bring this table into your next vendor call:

What to ask A red flag answer A credible answer
Does the denominator include escalated conversations or exclude them “It’s roughly all conversations” Can break out the escalation share on the spot
Is the metric “no follow-up question” or “explicit positive rating” Can’t say Distinguishes the definitions and gives separate numbers for each
Which industry or customer segment produced this number “Our customers are all similar” Acknowledges that knowledge base maturity varies a lot in cross-border retail
Is the observation window a week or six months Only offers a single total Can show how the number trends over time

If a vendor can answer all four with specifics they can break down, the number was probably measured by a product team. If every answer is a one-line deflection, the percentage was likely written by marketing, not measured by anyone.

Checklist two: what happens when the AI can’t answer

More important than any fixed number is what happens the moment the AI doesn’t know the answer or gets it wrong. The AI-first, human-backed boundary determines whether a customer gets left hanging when things go sideways.

  • When there’s no grounding in the knowledge base, does the AI honestly escalate, or does it fabricate a plausible-sounding answer?
  • When a customer asks for a human, is it a one-click handoff, or does the customer have to repeat themselves?
  • Are refunds, compensation, and price changes always routed to human approval, with no exceptions?
  • When a conversation escalates, can the agent see what the AI already said and the customer’s sentiment so far?

If any of these four comes back as “not sure” or “the vendor never mentioned it,” that’s worth pressing on instead of letting a polished resolution number carry the conversation.

Checklist three: is “gets smarter with use” actually controlled

“Our AI learns automatically and gets smarter the more you use it” is another common pitch line, and it often means the model is quietly absorbing every historical conversation with no visibility into what it learned, when, or whether a mistake can be undone.

YundaDesk breaks learning into steps you can actually review: when the AI can’t answer, when an agent supplies the answer, or when an agent explicitly corrects the AI, those signals generate a pending learning suggestion. A manager reviews it in an approval queue before it takes effect, and only then does it become a skill, a knowledge entry, or a customer memory. Every suggestion traces back to the conversation that triggered it and the person who approved it, and it can be rolled back with one click if it turns out wrong. The full breakdown is in this piece on how AI keeps getting smarter.

Ask directly: can learning suggestions be exported as a list? Who has permission to approve them? If something goes wrong, do you need to call the vendor’s engineers, or can you roll it back yourself?

Checklist four: is the knowledge base the real bottleneck

More often than not, “inaccurate AI” isn’t a model problem — it’s a knowledge base that’s stale, incomplete, or contradictory. Instead of asking “how advanced is your model,” ask how the knowledge base gets maintained:

Maintenance method Good fit for Risk to watch
Document upload Stable policy and FAQ content Accuracy drops fast if documents aren’t kept current
Website or help-center crawling Content that’s already maintained elsewhere Ask about crawl frequency and coverage scope
Manual Q&A entry High-frequency but scattered support phrasing Needs ongoing input or it goes stale

The relationship between the knowledge base and AI accuracy matters far more than parameter count. If a vendor only talks about the model and never about how the knowledge base is built, updated, or flagged when it’s out of date, that’s a signal worth noting.

Checklist five: can you test it yourself, not just watch a demo

The real way to verify a claim is to run it against your own support scenarios, not a vendor’s case study screenshots. Ask to:

  1. Test knowledge base coverage against your own historical tickets or common questions.
  2. Deliberately ask a few questions the knowledge base doesn’t cover, and see whether the AI admits it doesn’t know or fabricates an answer.
  3. Have an agent actually edit a few AI drafts, and confirm those edits genuinely show up in a review queue instead of disappearing into a black box.
  4. Run through a refund scenario and confirm it’s routed to human approval rather than decided by the AI alone.

A vendor willing to run this test with you, versus one that only wants to show you a case study video, gives you your answer before you sign anything.

Checklist six: is billing tied to the accuracy claim

Some vendors bill by resolution outcome, which creates a hidden incentive to loosen the definition of “resolved” so the bill goes up. YundaDesk plans include AI credits with no per-conversation or per-outcome surcharge, so billing is predictable and there’s no incentive to inflate a resolution number to justify a higher invoice. Ask a vendor directly: if your accuracy number went down, would your bill change too? If the answer is yes, the fairness of that number is worth questioning.


A fixed resolution rate sounds reassuring, but it doesn’t survive three follow-up questions: what’s the denominator, what happens when the AI can’t answer, and is learning actually under control. What’s worth confirming before you sign is how the knowledge base gets maintained, whether learning suggestions are traceable and reversible, and whether high-risk actions are always routed to a human. To see how this review loop actually works, check out the product details, or compare it against your own scenario on the pricing page.

Run this playbook in your own workspace

AI answers first, humans back up, every step is revertible — everything in this article can be put into practice in YundaDesk.