On this page6 sections
The PHQ-9 is nine questions long and is probably the most administered mental health instrument in the world. Its design is unusually direct: each of its nine items corresponds to one of the nine diagnostic criteria for a major depressive episode. It is less a questionnaire than a checklist wearing one.
In short
- Developed by Kroenke, Spitzer and Williams, published in 2001 in the Journal of General Internal Medicine.
- Nine items, each mapping directly onto a diagnostic criterion for major depression.
- Each item scores 0–3 over a two-week window, giving a total from 0 to 27.
- A cutoff of 10 is the conventional screening threshold, with reported sensitivity and specificity both around 88%.
- Item nine asks about self-harm — which makes the instrument safety-critical and changes how it must be deployed.
- Free to use without permission, which is why it is everywhere.
What the PHQ-9 measures#
The PHQ-9 measures the frequency of depressive symptoms over the preceding two weeks. Its nine items cover, in order: loss of interest or pleasure, low mood, sleep disturbance, fatigue, appetite change, feelings of failure or self-blame, concentration difficulty, psychomotor change, and thoughts of self-harm or being better off dead.
The one-to-one mapping onto diagnostic criteria is the design. It is what allows a nine-item questionnaire to function both as a severity measure and as a structured screen, and it is why the instrument reads as clinical rather than conversational.
How it is scored#
Each item is rated for how often the symptom has been present over the last two weeks: not at all (0), several days (1), more than half the days (2), or nearly every day (3). The nine items are summed for a total from 0 to 27.
| Total score | Conventional band |
|---|---|
| 0–4 | Minimal |
| 5–9 | Mild |
| 10–14 | Moderate |
| 15–19 | Moderately severe |
| 20–27 | Severe |
A score of 10 or above is the conventional threshold at which further clinical assessment is indicated. In the validation literature this cutoff reports sensitivity and specificity of roughly 88% each for major depression — good for a nine-item screen, and nowhere near good enough to diagnose anyone.
A change of around 5 points is generally treated as clinically meaningful, which is what makes the PHQ-9 useful for tracking response to treatment rather than only for initial screening.
Item nine, and why it changes everything#
The ninth item asks about thoughts of being better off dead or of hurting yourself. Its presence transforms the instrument from a questionnaire into something with a duty of care attached.
In clinical settings a non-zero response on item nine triggers a defined protocol: direct risk assessment, and escalation where indicated. The item is not scored and filed. It is acted on.
An instrument that can surface risk must be deployed somewhere that can respond to it.
This is the single most important reason a consumer product should not casually administer the full PHQ-9. Asking the question creates an obligation, and an app that collects the answer without a route to help has taken on the obligation and failed it.
Variants#
- PHQ-2 — the first two items only, covering low mood and loss of interest. Used as an ultra-brief first-stage screen; a positive result prompts the full nine.
- PHQ-8 — the nine-item version with the self-harm item removed. Common in large population surveys, where no protocol exists to respond to a risk disclosure. Its existence is itself an acknowledgement of the point above.
- PHQ-15 and the wider PHQ suite — companion instruments covering somatic symptoms, anxiety (GAD-7) and other domains, from the same programme of work.
Limits worth knowing#
- It is a screener, not a diagnosis. A score of 18 does not mean a person has major depression. It means a clinician should look properly.
- Somatic items confound. Fatigue, appetite change and sleep disturbance are produced by plenty of things that are not depression — chronic illness, pregnancy, shift work, medication.
- Two weeks is a fixed window. Repeated fortnightly readings can miss shorter cycles entirely.
- Symptom framing is culture-bound. Distress is not described the same way everywhere, which is part of why wellbeing-framed instruments such as the WHO-5 travel more easily.
- Self-report under low mood is unreliable in a specific direction. Depression affects recall and self-appraisal, which are exactly the faculties the instrument depends on.
Why consumer apps should be careful with it#
The PHQ-9 is free to use, short, and widely recognised. That combination makes it tempting to drop into any product with a wellbeing feature, and a great many products have done exactly that.
It is usually a mistake, for the reason set out above. Item nine can surface an active risk disclosure, and an instrument that can surface risk belongs somewhere that can respond to it. Collecting that answer with no route to help is not a neutral act — it takes on a duty of care and then fails it.
The existence of the PHQ-8 is the field's own answer to this. It exists precisely so that researchers running large surveys, with no clinician attached and no protocol to escalate to, can measure depressive symptom load without asking a question they are not equipped to act on. A consumer product is in the same position, and generally has less excuse.
Common questions#
- What is the PHQ-9?
- The PHQ-9 is a nine-item questionnaire measuring the frequency of depressive symptoms over the previous two weeks. Each item corresponds to one of the nine diagnostic criteria for a major depressive episode. It was published by Kroenke, Spitzer and Williams in 2001 and is widely used as a screening and severity-monitoring tool in primary care.
- How is the PHQ-9 scored?
- Each of the nine items is rated 0 to 3 according to how often the symptom occurred over the past two weeks, giving a total from 0 to 27. Conventional severity bands are 0 to 4 minimal, 5 to 9 mild, 10 to 14 moderate, 15 to 19 moderately severe, and 20 to 27 severe.
- What PHQ-9 score indicates depression?
- A score of 10 or above is the conventional threshold indicating that further clinical assessment is warranted, with reported sensitivity and specificity of about 88% each for major depression. No PHQ-9 score diagnoses depression on its own; diagnosis requires clinical assessment.
- What is the difference between the PHQ-9 and the PHQ-8?
- The PHQ-8 is the PHQ-9 with the ninth item, which asks about thoughts of self-harm, removed. It is used mainly in large population surveys where there is no protocol in place to respond to a disclosure of risk.
immli is a wellbeing and understanding tool, not a medical service. Questions about this page go to hello@immli.me.
Back to top ↑