Written by the Inclusive Developmental and Therapy Center therapy team · medically reviewed by Dr Muhammad Suffyan, MB BS (GMC 8023727) · Last reviewed July 2026
Parents of a child with additional needs are surrounded by confident claims. A programme promises a breakthrough; a headline announces that a food or a therapy ‘causes’ or ‘cures’ something; a glossy brochure cites ‘studies’. Knowing how to weigh these claims — to tell strong evidence from weak, and a genuine effect from wishful thinking — is one of the most protective skills a family or a clinician can have. This guide is about that skill.
It explains what evidence-based practice actually means (it is broader than most people think), how researchers rank the strength of different kinds of study, and the handful of concepts — reliability, validity, control groups, effect size — that separate a trustworthy finding from a flashy one. It is readable for an interested parent, but the definitions are precise enough for a student or trainee to rely on. The goal is not to make you cynical about all research, but to make you a fair and confident judge of it.
What evidence-based practice really means
‘Evidence-based’ is one of the most misused phrases in the field. Many people assume it means ‘whatever the research says, do that’. It doesn’t. Evidence-based practice rests on three legs, and all three matter: the best available research evidence, the clinician’s own expertise and judgement, and the values and circumstances of the family and child. Remove any one leg and the stool falls over.
This three-part definition has real consequences. Research evidence alone cannot tell you what to do with a particular child, because a study describes averages across a group, not the individual in front of you — that is where clinical expertise comes in. And even the best-supported intervention is the wrong choice if it clashes with a family’s values, culture, resources or the child’s own wishes. A therapy a family cannot sustain, or that distresses the child, is not ‘evidence-based’ for them merely because a journal supports it.
So when you hear that something is ‘evidence-based’, the fair question is not just ‘is there research?’ but ‘is there good research, interpreted by skilled judgement, applied in a way that fits this family?’ Holding all three together is the whole art.
The hierarchy of evidence: not all studies are equal
Studies are not interchangeable, and researchers picture their relative strength as a pyramid — the hierarchy of evidence. The idea is simple: designs higher up the pyramid do more to protect against bias and chance, so they carry more weight. Understanding the ladder lets you judge a claim by the kind of study behind it, before you even read the details.
Near the bottom sits the case study or case series: a detailed description of one child or a handful of children. These are valuable for generating ideas and describing something new, but they cannot tell you whether a treatment works, because there is nothing to compare against and no protection from coincidence. A step up are observational studies — cohort and case-control designs — which follow or compare groups without assigning treatment. They can reveal associations across many people, but because the groups differ in ways beyond the treatment, they cannot firmly establish cause.
Higher still is the randomised controlled trial (RCT), which is the strongest design for testing whether a treatment actually causes an effect (its logic is explained below). At the very top sit the systematic review and the meta-analysis. A systematic review uses an explicit, repeatable method to find and appraise every relevant study on a question; a meta-analysis statistically combines the results of those studies into a single, more precise estimate. Because they pool many studies rather than relying on any one, well-conducted reviews and meta-analyses are treated as the highest tier of evidence. One caveat worth remembering: a review is only as good as the studies it contains — pooling weak studies does not manufacture strong evidence.
Why the randomised controlled trial is so powerful
The randomised controlled trial deserves a section of its own, because its logic is the engine behind most confident claims that a treatment ‘works’. An RCT has two essential ingredients. First, a control group: some participants receive the treatment being tested while others receive a comparison — no treatment, usual care, or a placebo — so the researchers have something to measure the treatment against. Without a control group you cannot know whether children would have improved anyway, simply through maturation, extra attention or the passage of time.
Second, and crucially, participants are assigned to those groups at random. Randomisation is a quietly brilliant idea: by allocating children to groups by chance, it tends to balance out all the other factors — severity, age, family circumstances, motivation — that might otherwise differ between the groups and distort the result. This is how an RCT tackles confounding: a confounder is a hidden third factor linked to both the treatment and the outcome that can create a misleading impression of cause. Randomisation spreads confounders evenly across the groups, known and unknown alike.
Many good trials add blinding, meaning that the people involved don’t know who received the real treatment. In a single-blind study the participants don’t know; in a double-blind study neither the participants nor the people assessing the outcomes know. Blinding matters because expectation is powerful — a parent or an assessor who believes a child received a promising therapy may genuinely perceive more improvement. Blinding keeps that hope from masquerading as a result.
Reliability and validity: two different questions
Whenever a study measures something — a child’s language, attention or behaviour — you should ask two separate questions about that measurement: is it reliable, and is it valid? They sound similar but mean different things, and a measure can have one without the other.
Reliability is about consistency. A reliable measure gives the same answer when the situation hasn’t truly changed: the same result if the same child is tested twice in a short window (test–retest reliability), or if two different examiners score the same performance and agree (inter-rater reliability). An unreliable measure is like a bathroom scale that shows a different weight each time you step on it — you can’t trust any single reading.
Validity is about accuracy: does the measure actually capture what it claims to? A test can be perfectly reliable yet invalid — a scale that always reads six kilograms too heavy is wonderfully consistent and consistently wrong. In child development this matters constantly: a questionnaire might reliably measure something, but is that something really ‘anxiety’, or is it capturing shyness, or language difficulty? The classic way to hold the two together: reliability is hitting the same spot every time; validity is hitting the actual target. You need both, and reliability is necessary but not sufficient for validity.
Sample size, correlation and the traps in between
A few more ideas separate a sound study from a shaky one. The first is sample size — how many participants took part. Small studies are wobbly: with only a handful of children, an impressive-looking result can easily be a fluke of who happened to be included. Larger samples give more stable, trustworthy estimates and make it less likely that chance alone produced the finding. When you see a dramatic claim resting on five or ten children, treat it as a hint worth following up, not a conclusion.
The second, and perhaps the most important idea in this whole guide, is that correlation is not causation. Two things can rise and fall together without one causing the other. A famous everyday example: across a summer, ice-cream sales and drowning rates both climb — but ice cream does not cause drowning; hot weather independently drives both. In child development the same trap is everywhere. If children who receive a certain therapy tend to do better, it may be the therapy — or it may be that families who can access that therapy differ in other ways (income, time, support) that are the real cause. Only a design that controls for those other factors, like an RCT, can turn an association into a confident claim about cause.
Effect size and significance: does it matter, or just exist?
Suppose a study is well designed and finds a real effect. There is still one more question, and it is the one marketing most often hides: how big is the effect? This is what effect size captures — the magnitude of a difference, not merely whether one exists. A therapy might reliably improve a score by a tiny, real-but-trivial amount, or by a large, life-changing amount, and the word ‘works’ covers both.
This is where two easily-confused ideas must be separated: statistical significance and clinical significance. Statistical significance — often reported as a p-value — addresses only one narrow question: how likely is it that a result this large arose by chance if the treatment truly did nothing? A ‘statistically significant’ result (conventionally p less than 0.05) means the finding is probably not a fluke. It says nothing about whether the effect is big enough to matter. With a large enough sample, a difference far too small to make any real difference to a child’s life can still be ‘statistically significant’.
Clinical significance asks the question that actually matters to families: is the effect large enough to make a real, meaningful difference in everyday life? A result can be statistically significant but clinically trivial, or, occasionally, clinically promising but not yet statistically established. The mature reader wants both: confidence that the effect is real (statistical significance and a reasonable sample), and evidence that it is big enough to be worth the time, cost and effort (a meaningful effect size). ‘Significant’ in a headline almost always means the first and quietly hopes you’ll assume the second.
How to read a bold claim critically
Put it all together and you have a simple, calm checklist for any striking claim — including the marketing around a therapy. What kind of study is behind it: a single case, or a controlled trial and review? Was there a control group and randomisation, or just a before-and-after story? How many children took part? Is the effect merely statistically significant, or is it big enough to matter in real life? And is the source independent of the people selling the product?
A few warning signs should raise your eyebrows: claims of a cure or a breakthrough that works for everyone; testimonials and dramatic individual stories offered in place of controlled evidence; the word ‘studies’ with no way to check them; and a treatment sold as effective for a long, unrelated list of conditions at once. None of these prove a therapy is worthless, but each is a reason to slow down and ask for better evidence.
The point of all this is not to reject hope or to dismiss every new idea — genuine progress does come from careful research, and families are right to want the best for their child. The point is balance: to give bold claims exactly the weight the evidence behind them earns, no more and no less. That habit of fair, patient scepticism protects both your child’s time and your own, and it is the single most useful thing a critical reader of research can carry into every conversation.
Key takeaways
- Evidence-based practice stands on three legs together: the best research evidence, clinical expertise, and the family’s values and circumstances — research alone is not enough.
- The hierarchy of evidence runs from case studies (weakest) up through observational studies and RCTs to systematic reviews and meta-analyses (strongest), because higher designs better protect against bias and chance.
- An RCT establishes cause through a control group plus random allocation, which balances out confounders; blinding further stops expectation from masquerading as a result.
- Reliability is consistency (the same answer each time); validity is accuracy (measuring what you claim) — a measure can be reliable yet invalid, and reliability alone is never enough.
- Correlation is not causation: two things moving together may share a hidden common cause, so only a controlled design can turn an association into a claim about cause.
- Statistical significance (a p-value) only says a result is probably not chance; clinical significance and effect size say whether it is big enough to matter — always ask for both.
For students & professionals
A few deeper points worth knowing if you’re studying this area — think of it as a study aid, not a replacement for your course or supervisor.
- Memorise the hierarchy of evidence and why it is ordered as it is: each step up (case study → cohort/case-control → RCT → systematic review/meta-analysis) adds protection against bias and chance. A meta-analysis pools studies for a more precise estimate but inherits the weaknesses of the studies it contains.
- Be able to explain RCT logic precisely: the control group provides a comparison, and randomisation balances known and unknown confounders across groups so that a difference in outcome can be attributed to the treatment.
- Distinguish reliability from validity with the target analogy, and know the subtypes you will be examined on: test–retest and inter-rater reliability; and that reliability is necessary but not sufficient for validity.
- Define a confounder as a variable associated with both the exposure and the outcome that can produce a spurious association, and know that randomisation (not just statistical adjustment) is the strongest defence against unknown confounders.
- Keep effect size and the p-value firmly separate: a p-value addresses only whether an effect is likely non-chance, while effect size quantifies magnitude — and with a large sample a trivial effect can still reach statistical significance without being clinically significant.