Learn · Assessment & Diagnosis

How Assessment & Diagnosis Actually Work: Screening, Standardised Testing and Reaching a Diagnosis

Written by the Inclusive Developmental and Therapy Center therapy team · medically reviewed by Dr Muhammad Suffyan, MB BS (GMC 8023727) · Last reviewed July 2026

When a child is referred for concerns about their development, families often meet a confusing wall of words: screening, assessment, standard scores, percentiles, differential diagnosis. It can feel as though a number on a report is going to decide everything. It isn’t. Assessment is a careful, staged process of gathering information from many sources, and a good clinician treats every test score as one piece of evidence to be weighed, never as the final verdict on a child.

This guide walks through that process from start to finish. It explains the difference between a quick screen and a full assessment, shows you how to actually read the numbers on a standardised report, and describes how a multidisciplinary team moves from raw observations to a diagnosis. It is written so that an interested parent can follow it, but the definitions are precise enough that a student or trainee can rely on them. Throughout, keep one idea in mind: a diagnosis is a doorway to the right support, not a summary of who the child is.

Screening versus full assessment: two different jobs

A screening test and a full assessment answer different questions, and confusing the two causes a great deal of needless worry. A screen asks a single, broad question: ‘Is there enough concern here to look more closely?’ It is deliberately short, low-cost and designed to be given to many children. Its job is to sort children into ‘probably fine’ and ‘worth a closer look’ — nothing more. A screen never diagnoses anything.

Because screens are quick, they accept a trade-off. A good screen is designed to be sensitive, meaning it catches most children who genuinely have a difficulty, even at the cost of also flagging some children who turn out to be developing typically (a ‘false positive’). That is why a ‘fail’ or ‘refer’ result on a screen is an invitation to investigate, not bad news in itself. Common examples in child development include the M-CHAT-R for autism-related concerns and the ASQ as a general developmental screen.

A full assessment is the closer look. It is longer, carried out by a qualified professional, and draws on standardised tests, structured observation, developmental history and information from parents and teachers. Where a screen gives a yes/no signal, an assessment builds a detailed profile of a child’s strengths and difficulties across several areas. Only an assessment — usually across more than one session and often more than one professional — can support a diagnosis.

Reading the numbers: standard scores, percentiles and scaled scores

Standardised tests report results in a few standard currencies, and once you know how to read them the reports stop being intimidating. Start with the raw score: simply the number of items the child got right or the points they earned. A raw score on its own is almost meaningless, because it doesn’t tell you what is typical for the child’s age. To make it meaningful, the raw score is converted by comparing it to a large sample of children of the same age — the ‘norm’ group.

The most common converted currency is the standard score. On most major tests, standard scores are set so that the average (mean) is 100 and the standard deviation (SD) is 15. The standard deviation is just a measure of how spread out scores are. Roughly two-thirds of all children score within one SD of the mean — that is, between 85 and 115 — and about 95% score within two SDs, between 70 and 130. So a standard score of 100 is exactly average; 85 sits one SD below average; 70 sits two SDs below and is a common threshold for a ‘significant’ difficulty.

A percentile rank says the same thing in everyday language: it is the percentage of children in the norm group who scored at or below this child. A percentile of 50 is bang in the middle. A percentile of 16 means the child scored higher than 16% of peers (and corresponds to a standard score of about 85). Crucially, percentiles are not percentages of questions answered correctly, and they are not evenly spaced — the gap in raw ability between the 50th and 60th percentile is much smaller than the gap between the 5th and 15th, because most children cluster near the middle.

Subtests are often reported as scaled scores, which use a different scale: a mean of 10 and a standard deviation of 3, usually running from 1 to 19. A scaled score of 10 is average; 7 is one SD below. These subtest scaled scores are typically combined to produce the standard scores (mean 100) for broader areas. Knowing which scale you are reading — 100/15 or 10/3 — is the single most common thing people get wrong.

Basal, ceiling and the confidence interval

Two more technical terms appear constantly on reports: basal and ceiling. Because a single test covers a wide age range, the examiner does not administer every item to every child. The basal is the point at which the child is assumed to pass all easier items — usually a run of consecutive correct answers — so testing need not start from the very beginning. The ceiling is the point at which the child gets a run of consecutive items wrong, and testing stops because harder items would only add failures. Establishing a clear basal and ceiling is how a test is scored efficiently and fairly; errors in setting them are a common source of scoring mistakes.

No test score is a single exact point, and honest reports say so using a confidence interval. A confidence interval is a band around the score — for example, ‘standard score 82, 95% confidence interval 76–88’ — that reflects the fact that all measurement contains some error. It means that if the child were tested many times, their true score would fall within that band on 95% of occasions. The practical lesson is profound: you should read a score as a range, not a pinpoint. A child who scores 82 has not meaningfully ‘improved’ if they score 85 next time; both sit comfortably within the same confidence band.

Norm-referenced versus criterion-referenced tests

There are two fundamentally different ways to give a test result meaning, and they answer different questions. A norm-referenced test compares a child to other children. The standard scores and percentiles above are all norm-referenced: they tell you how this child performs relative to a representative sample of peers. This is what you want when the question is ‘Is this child developing differently from most children their age?’

A criterion-referenced test, by contrast, compares a child to a fixed standard or skill, ignoring how other children perform. It asks ‘Can this child do this specific thing?’ — for example, ‘Can they produce the /s/ sound correctly in the middle of words?’ or ‘Can they follow a two-step instruction?’ A driving test is an everyday criterion-referenced test: you pass by meeting the standard, regardless of how others did. Criterion-referenced results are especially useful for planning therapy and setting goals, because they describe exactly what a child can and cannot yet do.

Good assessment usually uses both. Norm-referenced tests establish whether there is a difficulty and how significant it is relative to peers; criterion-referenced measures then map the specific skills to target. Neither is ‘better’ — they simply answer different questions, and a report that leans on only one is giving you half the picture.

Informal and dynamic assessment: fairness for every child

Standardised norm-referenced tests have a built-in limitation: they are only fair when the child being tested resembles the children in the norm group. For a bilingual child, a child from a different cultural background, or a child with limited experience of formal testing, a standardised score can badly underestimate ability — not because the child has a disorder, but because the test was not built for them. This is why skilled clinicians lean heavily on informal and dynamic assessment.

Informal assessment gathers evidence outside the fixed procedures of a standardised test: observing play, sampling a child’s natural language, listening to how they talk with family, and using parent and teacher report. It captures what a child does in real life rather than in a testing room.

Dynamic assessment goes a step further and is one of the most valuable tools for telling a genuine disorder apart from a difference of experience. Instead of just measuring what a child already knows, it follows a test–teach–retest pattern: the examiner assesses a skill, then actively teaches it, then re-assesses to see how much the child gained. The key idea is ‘modifiability’. A child who learns quickly and applies the new skill widely after brief teaching is showing a difference (they simply hadn’t been exposed before); a child who makes little progress despite strong, focused teaching is more likely to have an underlying disorder. For bilingual and culturally diverse children, this teach-and-watch approach is far fairer than a one-off standardised score, because it measures learning potential rather than prior opportunity.

From information to diagnosis: how the pieces come together

A diagnosis is a reasoned conclusion, not the output of a single test. It usually begins with a thorough developmental history: the pregnancy and birth, early milestones, medical and family history, and the pattern of concerns over time. History matters enormously, because development is a story that unfolds — when a skill appeared, or disappeared, can be as informative as any test.

Alongside history sits direct observation of the child, ideally in more than one setting, and information from the people who know them best. In most robust services this is a multidisciplinary team effort: a paediatrician, psychologist, speech and language therapist, occupational therapist and others each contribute their piece, so that no single professional’s view decides the outcome. This is particularly important for complex presentations such as autism, where guidance calls for a team rather than one clinician working alone.

The intellectual heart of diagnosis is differential diagnosis: systematically considering all the conditions that could explain a child’s profile and reasoning about which best fits — and, just as importantly, which can be ruled out. Many things overlap. A child who is not talking might have a hearing loss, a developmental language disorder, autism, a global developmental delay, or simply be a late bloomer; a hearing test alone can change everything. Differential diagnosis is the discipline of not leaping to the first plausible label.

Finally, clinicians match the overall picture against formal diagnostic criteria. Two systems dominate. The DSM-5 (the American Psychiatric Association’s Diagnostic and Statistical Manual, fifth edition) is used widely in research and in North American practice. The ICD-11 (the World Health Organization’s International Classification of Diseases, eleventh revision) is the international standard used by health systems in the UK and much of the world. The two are broadly aligned and deliberately harmonised in many areas, but they are not identical in how some conditions are named and grouped, which is why a report may cite one, the other, or both.

A label is a doorway, not the whole child

It is worth ending where families feel it most. A diagnosis can be a relief — a name that unlocks funding, therapy, school support and a community of others who understand — or it can feel like a heavy word attached to someone you love. Both reactions are normal. What matters is holding the label in proportion.

A diagnostic category describes a pattern of difficulties that a child shares with others; it does not describe their personality, their potential, their humour or the particular way they light up. Two children with the same diagnosis can be strikingly different. The right way to use a diagnosis is as a doorway: it opens access to the right kind of help and a shared language for the team around the child. Everything that makes the child themselves still has to be discovered person to person, and that is the part no test will ever score.

Key takeaways

  • A screen only sorts children into ‘probably fine’ and ‘look more closely’; it never diagnoses, and a ‘refer’ result is an invitation to investigate, not a verdict.
  • Most standard scores use a mean of 100 and a standard deviation of 15, so 85 is one SD below average and 70 (two SDs below) is a common significance threshold; subtest scaled scores instead use a mean of 10 and SD of 3.
  • A percentile rank is the percentage of peers scoring at or below the child — not a percentage of questions correct — and percentiles are bunched together near the middle.
  • Every score carries measurement error, so read it as a confidence interval (a band) rather than an exact point.
  • Norm-referenced tests compare a child to peers; criterion-referenced tests compare a child to a fixed skill or standard — good assessment uses both, plus informal and dynamic assessment for fairness.
  • Diagnosis comes from developmental history, observation, a multidisciplinary team and differential diagnosis matched against DSM-5 or ICD-11 criteria — and the resulting label is a doorway to support, not the whole child.

For students & professionals

A few deeper points worth knowing if you’re studying this area — think of it as a study aid, not a replacement for your course or supervisor.

  • Be fluent in converting between scales: a standard score of 85 (mean 100, SD 15) equals roughly the 16th percentile and one SD below the mean; a subtest scaled score of 7 (mean 10, SD 3) is the same relative position. Always state which scale you are reading.
  • Norm-referenced answers ‘how does this child compare to peers?’; criterion-referenced answers ‘can this child do this specific skill?’ Know which question each result is answering before you interpret it.
  • Dynamic assessment (test–teach–retest, measuring modifiability) is your strongest tool for separating a language difference from a disorder in bilingual or culturally diverse children, because it measures learning potential rather than prior exposure.
  • Differential diagnosis means generating the full set of conditions that could explain a profile and reasoning to the best fit while ruling others out — for a non-verbal child, always consider hearing loss, DLD, autism and global developmental delay before concluding.
  • Know the two classification systems: DSM-5 (APA, dominant in research and North America) and ICD-11 (WHO, the international and UK standard). They are broadly harmonised but differ in naming and grouping of some conditions.
Take the first step

Questions about this topic — or your child?

We’re always happy to explain things in plain language, at home or in your studies.

MPS Road, Block A Model Town, Multan (near Bloomfield Hall School, Street No. 2) · Mon–Sat, 10 AM – 7 PM

Call Now WhatsApp
Chat with us