A little perspective for a lot of exam pressure. You're in the right place.

Assessments

K10 validity and reliability: evidence, accuracy, and limitations

The K10 is among the best-studied brief distress screens available, with strong evidence for detecting common mood and anxiety disorders at population level. That evidence concerns screening performance in samples, not certainty about any individual result.

What do validity and reliability mean?

Reliability concerns consistency of measurement. Validity concerns whether the evidence supports a particular interpretation or use. Neither is a permanent guarantee covering every language, population, and setting.

A screening scale can perform very well at identifying groups likely to have a disorder while remaining unable to tell one person what they have. Asking “valid for what purpose, and for whom?” is more informative than treating validation as a quality badge.

How was the K10 developed?

Kessler and colleagues used item response theory to select items that discriminate well across the range of distress, drawing on a large item pool. The aim was maximum precision from the smallest number of questions, for use in the redesigned US National Health Interview Survey.

That method is why the scale is short without being crude, and why the items look like an odd mixture. They were chosen for measurement performance rather than for thematic tidiness.

How well does the K10 detect disorders?

Furukawa and colleagues evaluated the K10 and K6 against diagnostic interviews in the Australian National Survey of Mental Health and Well-Being. Screening for mood and anxiety disorders, the K10 produced an area under the receiver operating characteristic curve of 0.90, with a 95% confidence interval of 0.89 to 0.91.

The K6 was close behind at 0.89, and both outperformed the GHQ-12, then a widely used standard, which produced 0.80. The authors preferred the K6 for general screening on grounds of brevity and consistency, while noting the K10 may do better for severe disorders.

An area under the curve describes how well a scale separates cases from non-cases across all possible thresholds in that sample. It is not the probability that your own result is correct.

Where do the four bands come from?

Andrews and Slade used the Australian National Survey of Mental Health and Well-Being to provide normative data on the K10, relating the range of possible scores to symptoms, disability, service use, and diagnosis. That work underpins the banded interpretations now in circulation.

The specific 10–15, 16–21, 22–29, and 30–50 groupings used in these explainers are the ones the Australian Bureau of Statistics applies in its health surveys, drawn from that normative work together with other population research. A second set of cut-points, commonly 10–19, 20–24, 25–29, and 30–50, is also widely used by services and websites.

Neither grouping is a clinical rule derived for individuals. Both are summaries of what scores in each range looked like across survey populations, which is why the labels are worded as levels of distress rather than as conditions, and why crossing a boundary is not a clinical event.

What do screening statistics not tell you?

Screening inevitably produces missed cases and false positives. A threshold that works well in one population may perform differently in another, and how common a condition is in a population affects how a positive result should be read.

A student population during an exam period is not the general adult population in a national survey. Scores may be elevated across the group for reasons that are situational, which affects what a given total implies.

None of this makes the K10 unreliable. It means a total is a starting point for a conversation rather than a conclusion.

What are the main limitations?

Self-report depends on recollection, interpretation, and the response options offered. It provides the person’s perspective, which is valuable, but cannot supply every part of an assessment.

A total also compresses very different answer patterns. Two people scoring 26 may be describing entirely different four weeks with entirely different needs.

  • It does not identify a condition or establish a cause.
  • It does not assess risk, safety, or urgency.
  • It does not measure functional impact unless the K10+ is used.
  • Two scoring conventions are in circulation, which makes online totals easy to misread.
  • Band thresholds derive largely from one national survey population.
  • It does not predict exam grades or determine academic eligibility.

How were these educational explainers prepared?

These pages explain the questionnaire using published research and official health information, with original student examples. The examples are fictional and illustrate interpretation rather than treatment outcomes.

This is an educational synthesis. ExamStressCheck does not claim to have independently validated the K10 or clinically reviewed these explainers. The linked sources let you check the underlying material, including version-specific scoring guidance.

This page is educational. A questionnaire score cannot establish a diagnosis or explain the cause of your symptoms. If you are struggling, speak with a counselor or healthcare professional, whatever your score.

Read how to ask for support or find a helpline in your country.