How to Tell Whether a Skills Claim Is Backed by Anything

Research & Evidence · ·  6 min read ·  

Frequently Asked Questions

How do you evaluate a training provider's impact claims?

Ask five questions: what exactly was measured and by whom, compared with what, how long after the programme, who dropped out, and who funded the evaluation. Most weak claims fail on the first two.

What is wrong with end-of-programme satisfaction scores?

They reliably produce high numbers, because people who have just spent two engaging days with a good facilitator feel positive about it. They measure the experience, and they are often reported in language that implies they measured a result.

Why does a comparison group matter?

Without one, a before and after figure describes change over time rather than the effect of the programme. People improve for many reasons. Ask what the comparison group was and how people ended up in one group rather than the other.

What does statistical significance actually tell you?

That a result is unlikely to be chance, given the sample. It says nothing about how large the difference was. With a big enough group a trivial difference becomes significant, so ask about the size of the effect and whether it would matter in practice.

Why do dropout numbers matter so much?

If a hundred people start, sixty finish, and the result is based on those sixty, the figure describes the people it worked for. Those who left are frequently the ones it did not work for. This single issue inflates a large share of published results.

Is peer-reviewed research automatically reliable?

It filters out some poor work and lets plenty through. Small samples, short follow-up and selective reporting all appear in published papers. Read the method section rather than the abstract, and check the interval between the intervention and the measurement.

A vendor tells you their programme produces a 94 per cent improvement in collaboration. You would like to know whether that is true. You have about twenty minutes, no access to the underlying data, and no particular training in research methods. This happens to L and D buyers, HR directors and heads of department constantly, and the usual outcome is that the number is either believed or dismissed on instinct.

There is a middle option. You do not need to evaluate a study properly to work out whether a claim is worth taking seriously. You need to know which four or five questions expose the weak ones, and most weak claims fail on the first two.

Where the numbers usually come from

Most impressive-sounding figures in the skills market come from one of three places, and it is worth being able to tell them apart.

The first is a satisfaction survey administered at the end of a programme. This produces high numbers reliably, because people who have just spent two engaging days with a good facilitator feel positive about it. It measures the experience. It is often reported in language that implies it measured a result.

The second is self-assessed capability, usually before and after. This is more useful, and it has a specific and well documented weakness. People who learn about a topic often revise their view of how competent they were beforehand, so genuine learning can show up as no change or even a decline. The reverse also happens. A confident presentation of a model can raise self-rating without changing anything.

The third is a comparison between people who took part and people who did not. This is the strongest of the three and it introduces the selection problem. If participants volunteered, or were nominated by their manager, they differ from the comparison group in ways that predict the outcome regardless of the programme. Motivated people were going to improve anyway.

None of these are worthless. All three tell you something. The problem is when a number generated by the first is described in the language of the third.

The questions that do most of the work

Ask these roughly in order. You can usually stop early.

What exactly was measured, and who measured it? Ask whether the outcome was self-reported, rated by someone else, or observed in actual work. Self-reported confidence and observed behaviour are different claims and are frequently reported in the same sentence. If nobody can tell you what the measurement instrument was, that is your answer.

Compared with what? A before and after change tells you very little on its own, because people improve over time for many reasons. Ask what the comparison group was and how people ended up in one group rather than the other. If there was no comparison group, the claim is about change, not about the programme.

How long after? Measured on the last day of a programme, almost anything looks effective. Ask for the interval between the intervention and the measurement. Three months is a reasonable minimum for a behaviour claim. If the answer is “immediately after”, you are looking at a reaction, not an effect.

Who dropped out, and are they in the number? If a hundred people started and sixty completed, and the result is based on those sixty, the figure describes the people who stayed. The ones who left are frequently the ones for whom it did not work. This single issue inflates a very large share of published results.

Who paid for it and who ran it? Not disqualifying on its own. Plenty of good evaluation is funded by the organisation that built the thing. But an internally run, internally funded evaluation with no comparison group and no independent measurement is marketing with a methods section.

Reading the effect, not just the significance

Two further points are worth knowing, because they separate people who can read this material from people who cannot.

Statistical significance is not a measure of size. It tells you the result is unlikely to be chance, given the sample. With a large enough group, a trivial difference becomes significant. Ask how big the difference actually was and whether it would matter to you in practice. A significant improvement of two percentage points in a self-rating is not a reason to buy anything.

Percentages without base rates are close to meaningless. A 50 per cent improvement could be a move from four per cent to six per cent. Ask for the raw numbers on both sides. Vendors who report only the percentage change are usually reporting the more flattering of two available figures.

What this looks like in a buying decision

A learning director was comparing two providers of a leadership development programme. The first presented a headline figure of 89 per cent improvement in leadership effectiveness. The second offered a more modest claim about improved decision quality in a group of about two hundred managers.

Five questions took about half an hour across two calls. The 89 per cent turned out to be a post-programme self-rating with no comparison group, no baseline, and no follow-up beyond the final session. The number was real and it described how people felt on the day.

The second provider’s evaluation had a comparison group of managers on a waiting list, measurement at four months, and reported its dropout figures without being asked. The effect was smaller. It was also considerably more likely to be real.

The learning director chose the second, and the more useful outcome was internal. She had asked her own team the same five questions about their existing programmes and found that most of their internal reporting would not have survived them either. That was harder to raise than the vendor conversation and worth more.

Common traps

Treating peer review as a guarantee. It filters out some bad work and lets plenty through. Small samples, short follow-up and selective reporting all appear in published papers. Read the method section rather than the abstract.

Dismissing everything without an experiment. Randomised trials are rare in organisational settings for practical and ethical reasons. Well designed comparison studies, longitudinal tracking and honest case documentation are all useful. The realistic standard is not proof, it is whether the evidence is better than an anecdote and honest about its limits.

Accepting a famous name in place of evidence. A widely used model is not necessarily a validated one. Several of the most popular instruments in corporate learning have weak published support for the claims made about them. Popularity records adoption, not accuracy.

Applying the scrutiny only to vendors. Most internal L and D reporting would not survive the questions above. If you demand rigour externally and report satisfaction scores internally, people will notice, and the standard will not hold.

Reflection questions

Worth putting to your team before your next procurement or programme review.

  • For the last capability claim you accepted, do you know whether the outcome was self-reported or observed?
  • Which numbers does your function report upward, and would they survive the five questions?
  • When you last saw a striking percentage, did you ask what the base rate was?
  • Do you know the dropout figures for your own programmes, and are they in the results you report?
  • What would you have to stop claiming if you applied your vendor standard to yourself?

None of this requires a background in statistics. It requires being willing to ask what was measured, compared with what, how long afterwards, and who was left out, and then to sit through the pause while somebody works out how to answer. The quality of that pause usually tells you what you need to know.

An internally run, internally funded evaluation with no comparison group and no independent measurement is marketing with a methods section.

Key Takeaways

  • Self-reported confidence and observed behaviour are different claims, and they are routinely reported in the same sentence.
  • With no comparison group, a before and after figure describes change over time, not the effect of the programme.
  • Results based only on people who completed describe the people it worked for. Dropouts inflate a large share of published figures.
  • Statistical significance is not size. With a big enough sample, a trivial difference becomes significant.
  • Most internal L and D reporting would not survive the same questions people apply to vendors.

Your Action Plan

  • Ask your next vendor what the measurement instrument was and who administered it.
  • Ask what the comparison group was and how people ended up in one group rather than the other.
  • Ask for the interval between the programme and the measurement, and treat “immediately after” as a reaction score.
  • Request the dropout numbers and check whether they are included in the headline figure.
  • Put the same five questions to your own internal reporting before your next board update.

Quick Skill Self-Check

Rate yourself honestly on each statement. This is just for you.

I ask what was measured, and by whom, before I accept a number.

Rarely
Always

I check what a result was compared against, not only how large it sounds.

Rarely
Always

I ask for the base rate when someone quotes me a percentage.

Rarely
Always

I notice when I accept weak evidence because it confirms what I already believed.

Rarely
Always

I apply the same standard to our own reporting as I do to a vendor's.

Rarely
Always

I can sit through the silence while someone works out how to answer.

Rarely
Always

Want more practical insights on human skills, leadership, and the future of work?

Get research-backed strategies and actionable ideas delivered to your inbox.

Stay in the loop

i2Skills

Assess & develop the human-centered skills that set great leaders and teams apart in the age of AI and constant change.