CSAT Calculator
Customer satisfaction from a five-point distribution, with the confidence interval and the same data counted the three ways teams count it.
CSAT 80.6%
counting top two boxes, 4 and 5
satisfied 116 of 144
95% interval 73.3% to 86.2%
width 12.9% wide
The same data, counted the other ways
top two boxes, 4 and 5 80.6% ← yours
top box only, 5 44.4%
top three boxes, 3 to 5 93.1%
mean score 4.15 out of 5
The distribution
5 64 44.4% ███████████
4 52 36.1% █████████
3 18 12.5% ███
2 6 4.2% █
1 4 2.8% █
Response rate 27.7% of 520 sent
did not answer 376
To halve that interval about 583 responses
CSAT is 80.6% by the top two boxes, 4 and 5 rule. The same responses
give 44.4% counting only 5s and 93.1% counting 3 upwards. The rule is a
choice, it changes the number by more than most improvements do, and a
score quoted without it is not comparable to anything.
The interval is 73.3% to 86.2%, which is the honest width at 144
responses. Movement inside that range is not a change, and reporting it
as one is how a quarterly review ends up discussing noise.
A 27.7% response rate means the other 376 people are missing from the
score, and they are not a random sample of the rest. People with a
strong opinion answer, which is why a survey after a resolved ticket
scores higher than the same survey sent to everybody.
CSAT, NPS and CES measure different things and none of them substitutes
for another. CSAT asks about one interaction, NPS asks about the
relationship, and CES asks how hard it was. A support team with high
CSAT and high effort scores is being pleasant about a process that
should not exist.
When the survey is sent changes the answer. Immediately after a resolved
ticket is the highest-scoring moment there is; a week later, or after an
unresolved one, is a different number about the same service. Keep the
timing fixed or the trend measures the timing.
A five-point scale and a ten-point scale are not convertible. Neither
are scales with different labels: "satisfied" and "very good" are not
the same question, and translating one into the other loses the
comparison you were trying to make.
The mean of a satisfaction scale is a weak statistic because the
intervals between the points are not equal, and the distance from 1 to 2
is not the distance from 4 to 5. It is printed above because people ask
for it, and the distribution is the thing to read.
Output is valid and updates as you type.
Fix the highlighted fields to update the output.
CSAT is the share of people who said they were satisfied. The arithmetic is a division. The argument is about what counts as satisfied.
The usual convention is the top two boxes of a five-point scale, 4 and 5. It is a convention, not a definition, and it is doing more work than the data: the same 144 responses give 80.6 percent on the top two boxes, 44.4 percent counting only the 5s, and 93.1 percent if you count neutral as satisfied. All three are honest. A score quoted without its rule is not comparable to anything, including last quarter’s score from the same team.
The second thing usually missing is the interval. CSAT is a proportion measured on a sample, and at 144 responses the 95 percent interval is nearly thirteen points wide.
How to use
- Put in how many people gave each score from 1 to 5.
- Pick what counts as satisfied.
- Put in how many surveys you sent, which is the most useful number here.
Example
144 responses out of 520 surveys sent:
CSAT 80.6%
counting top two boxes, 4 and 5
satisfied 116 of 144
95% interval 73.3% to 86.2%
width 12.9% wide
The same data, counted the other ways
top two boxes, 4 and 5 80.6% ← yours
top box only, 5 44.4%
top three boxes, 3 to 5 93.1%
mean score 4.15 out of 5
The distribution
5 64 44.4% ███████████
4 52 36.1% █████████
3 18 12.5% ███
2 6 4.2% █
1 4 2.8% █
Response rate 27.7% of 520 sent
did not answer 376
A 12.9 point interval means a move from 80.6 to 84 next quarter is not an improvement you can demonstrate. It also means the 376 people who did not answer matter more than the rule you chose.
Pitfalls
Say which boxes you counted. Every comparison with an industry benchmark, a competitor or your own past depends on it, and published benchmarks frequently do not state theirs.
The response rate is the real number. People with a strong opinion answer surveys and everyone else does not. A CSAT that rises while the response rate falls usually means the unhappy majority stopped replying, which reads as an improvement on the dashboard.
Timing is part of the measurement. A survey sent immediately after a resolved ticket is the highest-scoring moment available. The same survey a week later, or after an unresolved ticket, is a different number about the same service. Fix the timing or the trend measures the timing.
CSAT, NPS and CES are not substitutes. CSAT asks about one interaction, NPS about the relationship, CES about how hard it was. A team with high CSAT and high effort scores is being pleasant about a process that should not exist.
Scales do not convert. A five-point and a ten-point score are not comparable, and neither are scales with different labels. “Satisfied” and “very good” are different questions, and mapping one onto the other quietly destroys the comparison you were making.
The mean is a weak statistic here. The distance from 1 to 2 is not the distance from 4 to 5, so averaging ordinal labels is an approximation. It is printed because people ask for it; the distribution is what to read.
Watch the 1s, not the average. Twenty 5s and four 1s average well and the four are the ones who will tell other people. A distribution with a tail at the bottom is a different problem from one clustered in the middle, and the same CSAT covers both.
Compatibility
Arithmetic in the browser: nothing is uploaded and nothing is stored.
The interval is Wilson’s score interval, the same one the other rate tools on this site use, chosen
because the textbook p ± z·√(p(1−p)/n) can return a bound outside 0 to 100 percent at exactly the
response counts support teams have. The sample size needed to halve the interval is solved from the
normal approximation, which is the right tool for planning.
The three counting rules are computed from the same distribution every time, so the comparison is arithmetic rather than a claim. The test suite fixes all three against the example distribution.
The interval assumes the respondents are a random sample of your customers. They are not: they are the people who chose to answer. No formula repairs that, which is why the response rate is printed next to the score.