Most loyalty programmes report a retention number that cannot be true. They compare enrolled customers against everyone else, find that members buy more often, and call the difference the programme's effect. It almost never is. The customers who join a loyalty programme are the ones who already intended to come back โ you selected for loyalty rather than creating it, and the report is measuring that selection.
This guide covers how to tell whether a loyalty programme actually changes behaviour: the metrics worth tracking, the holdout group that makes the number real, the traps that inflate it, and the point at which more measurement stops paying for itself.
The number almost everyone reports, and why it misleads
The standard dashboard compares members to non-members on repeat-purchase rate. The gap looks large and the programme looks successful.
The problem is selection. Enrolment is voluntary, so it is chosen disproportionately by frequent buyers, people who already like the brand, and anyone who expects to buy again soon. A programme that did nothing at all would still show a gap, because the two groups were different before the programme existed. Any business that has run this comparison has a number, but not an answer.
The question that matters is incremental: how many purchases happened that would not have happened otherwise. That is a smaller number than the dashboard shows, and it is the only one worth putting in a budget decision.
Three metrics that are worth the trouble
Repeat-purchase rate within a fixed window. The share of customers who buy again within, say, 90 days of their first order. Fix the window and never move it; a rate quoted without a window can be made to say anything, because a longer window always looks better.
Cohort retention curves. Group customers by the month they first bought, then track what share of each cohort is still buying in month 1, 2, 3 and so on. Curves are harder to fool than averages: a programme that works shows later cohorts sitting above earlier ones at the same age. If every cohort has the same shape, whatever you changed did not reach the customer.
Revenue per customer over a fixed horizon. Not lifetime value โ that requires a survival model and assumptions you probably have not validated. Take a concrete horizon, twelve months say, and measure actual revenue per acquired customer. It is unglamorous, it is comparable across periods, and nobody can argue with it.
The holdout group is the whole method
If you want a defensible number, withhold the programme from a random slice of customers and compare.
Pick a percentage โ five or ten per cent is usually enough for a business with meaningful order volume โ assign it randomly, and exclude it from enrolment prompts and reward messaging. Then measure the same metric on both groups. The difference is your programme's actual effect, and unlike the member-versus-non-member comparison it survives scrutiny, because randomisation removes the selection that ruins the naive version.
Two practical notes. Randomise at the customer level, not by store, region or channel, or you are measuring the store rather than the programme. And leave the holdout alone long enough to see a repeat cycle โ if your typical repurchase interval is two months, a three-week test tells you nothing.
The obvious objection is that a holdout means deliberately not rewarding some customers. That is true, and it is the cost of knowing. Weigh it against the alternative, which is spending the reward budget indefinitely without evidence that it changes anything.
Traps that inflate the result
Rewarding the purchase the customer was going to make anyway. Most cashback and points programmes pay on every qualifying order. A large share of that spend goes to purchases that needed no incentive. This is not a reason to stop; it is a reason to know the proportion, because it sets the real cost per incremental sale.
Counting enrolment as engagement. Sign-ups measure the sign-up flow, not the programme. Customers enrol when prompted at checkout and then never think about it again.
Reading a seasonal bump as a programme effect. If the programme launched in the run-up to your strongest quarter, the curve will rise on its own. Compare against the same period last year, or against the holdout.
Letting the reward change the basket rather than the frequency. Sometimes a programme moves customers to buy the same amount in fewer, larger orders. Order count falls, revenue is flat, and the dashboard reads as a failure or a success depending on which metric is on it. Track both.
When the mechanic matters more than the measurement
Measurement tells you whether a programme works; it does not tell you why it did not. When the number comes back flat, the usual cause is not the reward size but the timing โ the reward lands too long after the action for the customer to connect the two, or it requires a claim step that most people never complete.
This is where a pre-sale mechanic behaves differently from a classic rebate: rewarding an action the customer takes before buying, and paying it out quickly, gives you a dated event to measure against rather than a slow accrual nobody notices. The distinction, and when each is appropriate, is covered in our guide to social cashback and in the cashback versus points comparison.
When this is the wrong thing to measure
If you are pre-product-market-fit or running low order volumes, stop here. Retention measurement needs enough customers and enough repeat cycles to separate signal from noise, and below that threshold you will read random variation as strategy. A business doing a few hundred orders a month cannot run a meaningful holdout: the confidence interval will be wider than any effect you could detect.
The same applies to genuinely low-frequency categories. If customers buy from you once every three years, a retention programme is not the lever โ referral and margin are โ and a year of measurement will produce a curve with nothing in it.
And if the programme exists for reasons other than repeat purchase โ collecting first-party data, building a contactable audience, meeting a retailer's requirement โ then measure that instead. Holding it to a retention number it was never designed to move produces a false negative and kills something useful.
What to do on Monday
Fix one window and one metric. Carve out a random holdout. Wait one full repurchase cycle. Compare. Everything else โ tiers, gamification, reward size experiments โ is premature until that loop exists, because without it you cannot tell an improvement from a coincidence.
For how the reward mechanic itself is designed, costed and paid out, see TikJoy's Cashback API.