GAD-7 validity and reliability: evidence and limitations
What did the original GAD-7 study find?
Spitzer and colleagues studied 2,740 adults in 15 US primary care clinics; 965 received a follow-up diagnostic interview. At a threshold of 10, reported sensitivity for generalized anxiety disorder was 89% and specificity was 82%.
The study reported internal consistency of 0.92 and test-retest reliability of 0.83. These findings describe the research setting and methods. They are not a promise of the same performance in an unsupervised online student population.
What do reliability statistics tell us?
Internal consistency concerns how closely the items relate to one another. Test-retest reliability concerns the similarity of scores when a measure is repeated under the study conditions. Neither statistic is a percentage probability that an individual has a disorder.
Löwe and colleagues studied a general-population sample of 5,030 people and reported internal consistency of 0.89. Their work supported a one-dimensional structure and supplied population evidence beyond the original clinical setting.
When reading a reliability figure, ask which sample produced it and how the measure was administered. A strong coefficient does not, by itself, establish a suitable cutoff or confirm the cause of a person’s symptoms.
What do sensitivity and specificity mean?
Sensitivity is the proportion of people with the target condition who screen positive. Specificity is the proportion without that condition who screen negative. They describe different groups, so neither tells you directly the probability that your own positive result is a diagnosis.
The chance that a positive result represents the target condition also depends on how common that condition is in the population being screened. This is one reason the same questionnaire can have different practical implications in a specialist clinic and a general student survey.
Lowering a threshold can identify more potential cases while also increasing false positives. Raising it can reduce false positives while missing more cases. Choosing a threshold depends on the purpose and consequences of screening.
Does newer evidence match the original accuracy figures?
A 2025 Cochrane review estimated sensitivity of approximately 64% and specificity of 91% for GAD-7 detection of generalized anxiety disorder at the recommended threshold or the closest eligible lower threshold in its core range. The review found substantial variation across studies.
These pooled estimates differ from the original study because the underlying evidence covers other samples and settings. The useful conclusion is not that one percentage is the universal accuracy of the test. It is that screening results need clinical and population context.
What evidence applies to students and adolescents?
Adolescent clinical research and studies in university populations provide additional evidence, but they answer different questions. Correlation with another anxiety measure is not the same as agreement with an independent diagnosis.
For example, the adolescent study linked below involved participants who already had GAD. The college study examines measurement properties in a specific undergraduate sample. Neither establishes that the GAD-7 distinguishes exam anxiety from other sources of anxiety in every student.
Why do websites report different meaningful-change numbers?
They may be using different definitions, statistical methods, or populations. A reliable change index asks whether a difference exceeds an estimate of measurement error. A minimal clinically important difference asks about the size of change considered important under a study’s method.
Those concepts should not be collapsed into a universal pass/fail rule for recovery. A result should be discussed alongside your experience and functioning. Our results page explains one published change estimate and its limits.
What should I check before trusting an online GAD-7 result?
Look for an identified instrument, unchanged scoring, clear instructions, and a source for the interpretation. Research supporting the original measure does not automatically validate a website’s rewritten questions, personalized advice, percentile chart, or recommendation algorithm.
This page summarizes published evidence for educational use. It does not claim that ExamStressCheck has conducted an independent validation study or a clinical review. The research links allow you to examine the evidence behind the explanations.
- Does the form use the correct version and time period?
- Does the explanation distinguish screening from diagnosis?
- Are numerical claims linked to a study with a relevant population?
- Does the website explain its data handling before collecting answers?
- Are recommendations sensitive to your age, language, location, and needs?
This page is educational. A questionnaire score cannot establish a diagnosis or explain the cause of your symptoms. If you are struggling, speak with a counselor or healthcare professional, whatever your score.
Read how to ask for support or find a helpline in your country.