Every support lead has sat in a review meeting saying some version of “our satisfaction score looks fine, so why is repeat purchase still dropping?”
The problem usually isn’t the data. It’s that the team is treating three unrelated metrics as one number. NPS dips two points while CSAT climbs. Someone tries to use CES to judge brand loyalty and the conclusion makes no sense. NPS, CSAT, and CES each measure a different thing, and using the wrong one for the wrong moment just adds noise.
NPS measures “would you bring a friend”
Net Promoter Score asks one question: how likely are you to recommend us to a friend or colleague, on a 0-10 scale. 9-10 are promoters, 7-8 are passives, 0-6 are detractors. NPS = % promoters minus % detractors.
NPS vs CSAT vs CES: the industry baseline behind the metric
It’s not measuring whether this particular conversation went well. It’s measuring overall trust and long-term willingness to stick with the brand. That makes NPS a good fit for:
- Quarterly or biannual brand-health check-ins
- Reviewing overall sentiment after peak season wraps up
- Deciding whether a market or customer segment deserves more investment
The trap is that NPS is slow and coarse. A customer might have had a perfectly smooth interaction today but still score a 6 because of a shipping delay three weeks ago. Using NPS to judge a single support interaction is like using an annual physical to judge whether tonight’s dinner tasted good — wrong scale entirely.
CSAT measures “were you happy with this one”
Customer Satisfaction Score asks about a specific moment: how satisfied were you with this interaction, usually on a 1-5 scale or a simple satisfied/not-satisfied split. CSAT = satisfied responses / total responses.
CSAT naturally attaches to a specific touchpoint — a conversation closing, a ticket resolving, a return being processed. It answers “how did we do just now,” which makes it fast-turnaround and fine-grained enough to use as a reference for individual agent quality, or to break down by channel: is CSAT on WhatsApp noticeably lower than on the website widget? Is that a script problem or a response-time problem?
Its limitation is that it only captures whether the outcome felt good, not how much effort it took to get there. A customer might get bounced around three times before finally getting an answer, then still rate it 5/5 satisfied because the problem eventually got solved — the friction in between never shows up in the score.
CES measures “how much work did this take”
Customer Effort Score asks: how much effort did it take to get this resolved, typically on a 1-7 scale where lower means easier. It’s the most overlooked of the three, and arguably the one with the biggest impact on repeat purchase.
A large body of research points to the same pattern: customers don’t usually churn because they weren’t satisfied — they churn because it was exhausting. Transferred three times, asked to repeat the order number twice, stuck waiting after the AI couldn’t answer. CES is the right lens for:
- Complex resolutions (returns, billing disputes, multi-step troubleshooting)
- Whatever happens right after a handoff to a human agent — was the transition smooth?
- The moment AI support couldn’t answer and escalated to a person
A low CES score usually points to a broken process, not a rude agent — which is exactly why it’s better used to audit systems and knowledge bases than to grade individual staff.
The three ways teams blend these metrics and regret it
| Mistake | What happens |
|---|---|
| Using NPS to judge a single conversation | The agent gets blamed for a shipping problem from three months ago |
| Tracking CSAT without CES | Customers say “satisfied” but exhausted; repeat purchase drops and nobody knows why |
| Sending one survey type all through peak season | Patience is already thin during peak; the wrong survey type amplifies the wrong signal |
These metrics aren’t interchangeable — they’re complementary. CES tells you where the process is broken. CSAT tells you how one interaction felt. NPS tells you what all those interactions add up to over time.
A practical read: if CES is bad but CSAT is just okay, the customer probably “got through it, but it hurt.” They won’t churn immediately, but NPS will quietly slide a few months later. Watching CES gives you an earlier warning than waiting on NPS.
The part teams usually miss: surveys need to match the moment, not be one-size-fits-all
Most teams don’t fail because they picked the wrong metric — they fail because sending surveys is manual and scattered. One survey tool on the website widget, another for email follow-ups, data sitting in different places, and nobody has time to stitch it back together.
That’s exactly why the survey should follow the touchpoint: CSAT when a conversation closes, CES right after a handoff to a human, NPS on a periodic cadence — not the same questionnaire dumped into every channel. YundaDesk’s shared inbox keeps AI and human agents in the same conversation thread, so all three survey types can trigger at the right moment and roll up automatically by channel, by agent, and by issue type. Is CES on WhatsApp worse than on email? Is CSAT higher on conversations the AI resolved alone versus ones that got escalated? You don’t need to export and reassemble that comparison by hand.
Which one to use: a quick reference
- Want to know if the brand is earning more trust over time — track NPS quarterly, and don’t try to explain its swings with single-conversation data.
- Want to know if a service moment went well — track CSAT, attached to conversation or ticket closure, broken down by channel and agent.
- Want to know if customers are being put through the wringer — track CES, especially around complex issues and human handoffs. It’s an early warning for process health.
You don’t need all three to be perfect. You just need to be clear on which question each one is answering — that’s worth far more than a headline “we’re at 4.8 satisfaction” number in a deck.
A commonly misread signal: how AI support affects the numbers
A lot of teams worry that letting AI handle conversations will drag satisfaction down. In practice, the more common pattern is the opposite for CES: conversations the AI resolves on its own tend to score lower effort, because the customer isn’t waiting or repeating context. But when the AI can’t answer and doesn’t hand off promptly, CSAT drops noticeably. That’s exactly why the boundary between what AI handles and what a human backs up needs to be clear — the AI should escalate when it should, not keep trying past the point where it’s helping[1]. For more on where that line sits, see AI-First, Human-Backed: Where the Line Actually Sits.
If your knowledge base has gaps, the AI will fail to answer more often, and both CES and CSAT will take the hit. That’s usually a knowledge base coverage problem, not an agent performance problem.
Metrics exist to surface problems early, not to look good in a report. Put NPS, CSAT, and CES in their right places and you’ll spot where repeat purchase starts to slip long before a quarterly review shows the numbers all trending down at once.