How Depression Is Measured May Shape How Well Psychotherapy Appears to Work

A new review finds that inconsistent outcome measures in depression psychotherapy trials distort treatment effects and weaken the reliability of the evidence.

3
707

A new systematic historical and meta-analytic review published in the Journal of Affective Disorders found inconsistent and unreliable outcome measurements in depression psychotherapy trials.

Antonia A. Sprenger of the Institute of Social Medicine and Health Systems Research and colleagues set out to investigate which outcome measurement instruments (OMIs) are used in depression psychotherapy trials and how their use has changed over time, with the aim of identifying gaps and potential biases in outcome measurement practices.

The authors found that, despite the increasing use of OMIs in depression research, measurement practices have become increasingly inconsistent, and that different tools can shape how effective psychotherapy appears. The authors argue this undermines the interpretability, generalizability, and reliability of study findings. They suggest developing a standardized, agreed-upon set of outcome measures for depression for more reliable and comparable research findings.

The authors write:

“over the past 50 years, the BDI-1, BDI-2, PHQ-9 and HDRS have been the most frequently used OMIs for depression in psychotherapy studies, with an increasing proportion of studies using the latter two OMIs. Overall, the use of OMIs for depression in adult psychotherapy research has become increasingly heterogeneous over the past 50 years. This is particularly concerning since some OMIs for depression, particularly CES-D and GDS, were found to be associated with systematically different treatment effect estimates in psychotherapy studies. This may hinder meaningful synthesis of results across studies, therefore compromising the evidence available to guideline developers, clinicians, and patients. Developing a core outcome set for future depression psychotherapy trials represents a promising approach to improve the quality of depression measurement and promote comparability across studies.”

Depression is highly prevalent and associated not only with significant personal suffering but also substantial economic burden. Psychotherapy is widely regarded as a first-line treatment for depression, with its efficacy supported by more than 800 randomized controlled trials (RCTs). However, the strength of this evidence depends heavily on how treatment outcomes are measured. Without reliable measurement, it becomes impossible to assess the efficacy of various treatments.

Problems With Measuring Depression

Defining and measuring depression has long been debated in psychology, resulting in the development of a large and diverse set of outcome measurement instruments (OMIs). More than 280 OMIs are currently available for use in depression research, most of which are used to assess symptom severity based on patients’ self-reports. However, these instruments vary widely in what they measure.

According to the authors:

“OMIs can vary considerably in the content they measure (content validity), the representation of the underlying theory (internal structure), and the cognitive processes people engage in when using the OMI (response process).”

Given the lack of an agreed-upon definition of depression, different measurement tools measure different constructs. This creates practical challenges for interpreting research findings. These challenges are particularly pronounced for meta-analyses as they are collections of studies that may all be using different OMIs..

To address these issues, the authors conducted a meta-epidemiological study examining the use of OMIs in depression psychotherapy trials over the past 50 years, assessing trends in measurement consistency and the influence of different instruments on estimated treatment effects. The goal was to inform more deliberate OMI selection and future measurement development.

Methods

The researchers conducted a systematic literature search of four databases up to September 2024. They included studies that were randomized controlled trials comparing psychotherapy with control conditions in adults with depression that assessed depression symptoms through the use of OMIs, only including measurement tools that appear in at least ten treatment comparisons.

They analyzed the change in OMI usage over time and how much variety there was in their use. Additionally, they used statistical analysis methods to investigate whether the type of OMI used influenced how effective psychotherapy appeared to be. 492 trials were included, published between 1977 and 2024, using 17 different OMIs for depression.

OMIs and Psychotherapy

Across the psychotherapy trials reviewed, the authors found that 17 different outcome measurement instruments (OMIs) were used to assess depression. The most frequently used measures were the BDI-1, BDI-2, HDRS, and PHQ-9. Over the past 50 years, use of the BDI-1 and BDI-2 remained relatively stable, while use of the PHQ-9 and HDRS increased in psychotherapy research.

At the same time, the overall variety of depression measurement tools increased, indicating growing heterogeneity in how outcomes are assessed. Importantly, the authors also found that the choice of OMI was associated with different treatment effect estimates. Studies using the CES-D tended to report larger effects, while those using the GDS reported smaller effects, suggesting that how depression is measured can shape how effective psychotherapy appears.

In discussing these findings, the authors note that the rising use of the PHQ-9 is unsurprising given its clinical ease of use, frequent recommendation by national bodies, and close alignment with DSM criteria. However, the PHQ-9 was originally designed as a screening tool rather than a measure of treatment change, and its use may bias estimates by inflating rates of major depressive disorder.

The HDRS, which focuses on observable clinical symptoms, showed the greatest overlap with domains that matter to patients, but its psychometric limitations raise concerns about its continued use as a gold standard. The consistent use of the BDI-1 and BDI-2 likely reflects their long history and strong theoretical grounding, despite well-documented limitations related to cultural generalizability, dimensionality, and measurement invariance.

The authors write:

“Taken together, we found that psychotherapy studies have employed many different OMIs for depression. Although some OMIs, particularly the PHQ-9 and HDRS, were used more frequently over time, previous literature indicates that their psychometric properties are limited, raising concerns about the robustness of the evidence base.”

Overall, the study shows that psychotherapy research over the past five decades has become less consistent in how depression is measured, with a growing number of instruments in use and little agreement on how best to define it. When too many different tools are used, findings across studies become harder to compare, and results from meta-analyses may reflect measurement choice rather than treatment effects.

At the same time, the authors caution against relying on a single “objective” measure, noting that depression is a heterogeneous experience that cannot be fully captured by one instrument alone. Measurement diversity can reflect meaningful differences in theory, context, and lived experience.

As they note:

“The bias OMIs for depression can introduce to treatment effect estimates, coupled with the growing heterogeneity in their administration and the uncertainty about what they actually measure, is concerning. Especially because meta-analyses typically pool results across varying OMIs for depression without distinguishing whether the observed effects reflect the intervention or the choice of OMI.”
Conclusion

To address this problem, the authors argue that developing a core outcome set for depression psychotherapy trials is a promising solution. This would involve agreeing on key outcomes that all trials should measure, defining relevant domains with input from patients and clinicians, and selecting or developing instruments that are valid, reliable, and feasible across contexts.

Taken together, the findings of this study raise deeper questions about whether depression can be fully understood or quantified through standardized outcome measurement instruments. While the authors emphasize the methodological consequences of measurement heterogeneity for psychotherapy research, these results also resonate with a growing body of research that critiques the assumptions underlying many commonly used outcome measures and clinical definitions of depression, which often obscure the complex and context-dependent nature of real lived experience, as well as the diverse and dynamic ways recovery is experienced and understood.

****

Sprenger, A. A., Harrer, M., Miguel, C., Illing, S., Kuper, P., Buntrock, C., Karyotaki, E., Fried, E. I., Cuijpers, P., & Apfelbacher, C. (2026). Inconsistent outcome measurement in depression psychotherapy trials: A systematic historical and meta-analytic review over the past 50 years. Journal of Affective Disorders, 397, 120873. (Link)

Previous articleIn Defense of Instability in Mental Health Recovery
Next articleMad in America’s 10 Most Popular Articles in 2025
Ally Riddle
Ally is pursuing a master's in interdisciplinary studies through New York University's XE: Experimental Humanities & Social Engagement. She uses the relationship between anthropology, public health, and the humanities to guide her research. Her current interests lie at the intersection of literature and psychology as a method to reframe the way we think about different mental states and experiences. Ally earned a bachelor's degree from the University of Minnesota in Biology, Society, & Environment.

3 COMMENTS

  1. ***What happened to the other comments?***

    “Given the lack of an agreed-upon definition of depression, different measurement tools measure different constructs.”

    In other words, following Steve’s observation of lipstick on a pig, here’s more language to cosmetically conceal bullshit as science that’s no more than another circular exercise in confirmation bias (junk in, junk out). For far too long, most of modern history, such pseudoscientific numerology has been quantifying qualitative dimensions of human experience with no better basis than institutionalized practice of superstition. But there’s rationale for the irrationality and method to the madness.

    Whether so-called hard or soft sciences, knowledge is power which has been increasingly monopolized and centralized under capitalist class rule as service industries of, by, and for corporate state control (aka fascism). Industrial science particularly serves ruling class interests in population control, from physics, chemistry, or biology largely amounting to branches of the military industrial complex monopolizing the means of violence for hard control, as with mass murder, to psychology, sociology, or anthropology delivering softer (?) methods of control, as with mass manufactured lies of propaganda and psywar.

    “The welfare of the people has always been the alibi of tyrants.” (Camus) Lower-level ranks and micromanagers in these fields of knowledge may be indoctrinated to believe that professional expertise serves the common good. Yet dutiful role performance, just doing one’s job, in ruling institutions governing research and development and what counts as knowledge are bound to function as secular equivalents of medieval priestly castes, refining religious myth as reasons for the way the world is, much to the advantage of ruling powers, but with science now the opiate of the masses.

    Nearly half a century ago, Stephen Jay Gould definitively debunked IQ’s “mismeasure of man.” Yet today IQ tests remain ritual for classifying us into identities and demographics on data charts recalling bureaucracy of concentration camps, including such useful fictions as race, to recall other targets of eugenics besides ‘mental defectives’ still in practice as well (cf. Hernstein and Murray’s The Bell Curve). The scientific magic of numbers and data (or as Twain said, “lies, damned lies, and statistics”) is that the select ‘disciplines’ which deploy them have been able to misdirect our own intelligent attention and agency to alleged authoritative sources serving authoritarian purposes. This expropriation ensures false consciousness among the general population as to the ways the world works to the gain of ruling powers.

    Fear is a crucial factor in this con game. For example: Are your children at risk for whatever’s made up as mental challenges? Be sure to get them tested right away, so they can be placed on some prefabricated and ever adjustable chart with numbers by their names (which may necessitate ‘special treatment’). Of course by now after generations of medicalizing and pathologizing us every which way, the tests (and function creep of fascism) have become more mandatory, simply taken for granted as the way things are. The final solution, so to speak, is to alienate us from ourselves and our own autonomy until we are entirely dispossessed as biodigital cogs in the machinery of production, profit, and power. Stay tuned to the latest data mining of AI. We’re ‘well’ on our way, at least as long as we keep consenting to sell ourselves like slaves to our masters.

    Report comment

LEAVE A REPLY