A vendor tells you their programme produces a 94 per cent improvement in collaboration. You would like to know whether that is true. You have about twenty minutes, no access to the underlying data, and no particular training in research methods. This happens to L and D buyers, HR directors and heads of department constantly, and the usual outcome is that the number is either believed or dismissed on instinct.
There is a middle option. You do not need to evaluate a study properly to work out whether a claim is worth taking seriously. You need to know which four or five questions expose the weak ones, and most weak claims fail on the first two.
Where the numbers usually come from
Most impressive-sounding figures in the skills market come from one of three places, and it is worth being able to tell them apart.
The first is a satisfaction survey administered at the end of a programme. This produces high numbers reliably, because people who have just spent two engaging days with a good facilitator feel positive about it. It measures the experience. It is often reported in language that implies it measured a result.
The second is self-assessed capability, usually before and after. This is more useful, and it has a specific and well documented weakness. People who learn about a topic often revise their view of how competent they were beforehand, so genuine learning can show up as no change or even a decline. The reverse also happens. A confident presentation of a model can raise self-rating without changing anything.
The third is a comparison between people who took part and people who did not. This is the strongest of the three and it introduces the selection problem. If participants volunteered, or were nominated by their manager, they differ from the comparison group in ways that predict the outcome regardless of the programme. Motivated people were going to improve anyway.
None of these are worthless. All three tell you something. The problem is when a number generated by the first is described in the language of the third.
The questions that do most of the work
Ask these roughly in order. You can usually stop early.
What exactly was measured, and who measured it? Ask whether the outcome was self-reported, rated by someone else, or observed in actual work. Self-reported confidence and observed behaviour are different claims and are frequently reported in the same sentence. If nobody can tell you what the measurement instrument was, that is your answer.
Compared with what? A before and after change tells you very little on its own, because people improve over time for many reasons. Ask what the comparison group was and how people ended up in one group rather than the other. If there was no comparison group, the claim is about change, not about the programme.
How long after? Measured on the last day of a programme, almost anything looks effective. Ask for the interval between the intervention and the measurement. Three months is a reasonable minimum for a behaviour claim. If the answer is “immediately after”, you are looking at a reaction, not an effect.
Who dropped out, and are they in the number? If a hundred people started and sixty completed, and the result is based on those sixty, the figure describes the people who stayed. The ones who left are frequently the ones for whom it did not work. This single issue inflates a very large share of published results.
Who paid for it and who ran it? Not disqualifying on its own. Plenty of good evaluation is funded by the organisation that built the thing. But an internally run, internally funded evaluation with no comparison group and no independent measurement is marketing with a methods section.
Reading the effect, not just the significance
Two further points are worth knowing, because they separate people who can read this material from people who cannot.
Statistical significance is not a measure of size. It tells you the result is unlikely to be chance, given the sample. With a large enough group, a trivial difference becomes significant. Ask how big the difference actually was and whether it would matter to you in practice. A significant improvement of two percentage points in a self-rating is not a reason to buy anything.
Percentages without base rates are close to meaningless. A 50 per cent improvement could be a move from four per cent to six per cent. Ask for the raw numbers on both sides. Vendors who report only the percentage change are usually reporting the more flattering of two available figures.
What this looks like in a buying decision
A learning director was comparing two providers of a leadership development programme. The first presented a headline figure of 89 per cent improvement in leadership effectiveness. The second offered a more modest claim about improved decision quality in a group of about two hundred managers.
Five questions took about half an hour across two calls. The 89 per cent turned out to be a post-programme self-rating with no comparison group, no baseline, and no follow-up beyond the final session. The number was real and it described how people felt on the day.
The second provider’s evaluation had a comparison group of managers on a waiting list, measurement at four months, and reported its dropout figures without being asked. The effect was smaller. It was also considerably more likely to be real.
The learning director chose the second, and the more useful outcome was internal. She had asked her own team the same five questions about their existing programmes and found that most of their internal reporting would not have survived them either. That was harder to raise than the vendor conversation and worth more.
Common traps
Treating peer review as a guarantee. It filters out some bad work and lets plenty through. Small samples, short follow-up and selective reporting all appear in published papers. Read the method section rather than the abstract.
Dismissing everything without an experiment. Randomised trials are rare in organisational settings for practical and ethical reasons. Well designed comparison studies, longitudinal tracking and honest case documentation are all useful. The realistic standard is not proof, it is whether the evidence is better than an anecdote and honest about its limits.
Accepting a famous name in place of evidence. A widely used model is not necessarily a validated one. Several of the most popular instruments in corporate learning have weak published support for the claims made about them. Popularity records adoption, not accuracy.
Applying the scrutiny only to vendors. Most internal L and D reporting would not survive the questions above. If you demand rigour externally and report satisfaction scores internally, people will notice, and the standard will not hold.
Reflection questions
Worth putting to your team before your next procurement or programme review.
- For the last capability claim you accepted, do you know whether the outcome was self-reported or observed?
- Which numbers does your function report upward, and would they survive the five questions?
- When you last saw a striking percentage, did you ask what the base rate was?
- Do you know the dropout figures for your own programmes, and are they in the results you report?
- What would you have to stop claiming if you applied your vendor standard to yourself?
None of this requires a background in statistics. It requires being willing to ask what was measured, compared with what, how long afterwards, and who was left out, and then to sit through the pause while somebody works out how to answer. The quality of that pause usually tells you what you need to know.