Fluoxetine is recommended as a first-line antidepressant to treat pediatric depression in major guidelines, and it is one of the most widely prescribed antidepressants for children and adolescents. However, modern studies find that fluoxetine has lost its efficacy and can now be considered clinically equivalent to placebo. In this piece, I will explain this puzzling phenomenon, which is also crucial for judging the efficacy of other common antidepressants. This issue has not yet been adequately addressed, which can lead to potential harm for young patients.
A more elaborated version of this blog is available here.

Fluoxetine’s loss of efficacy
Two influential reviews published in 2016 and 2020 in prestigious Lancet journals [1,2] report that “For efficacy, only fluoxetine was statistically significantly more effective than placebo (standardised mean difference [SMD] -0.51, 95% credible interval [CrI] -0.99 to -0.03).” [1]
The authors of the 2016 review concluded that “when considering the risk–benefit profile of antidepressants in the acute treatment of major depressive disorder, these drugs do not seem to offer a clear advantage for children and adolescents. Fluoxetine is probably the best option to consider when a pharmacological treatment is indicated.” [1]
Based on these findings, major guidelines recommend fluoxetine as the first-line antidepressant for treating pediatric depression [3]. However, in 2021, a review from the prestigious Cochrane Collaboration [4] reported only half of the efficacy of the 2016 and 2020 reviews, and with much greater precision: SMD = -0.20 (95% CI -0.28 to -0.11). Furthermore, the quality of the evidence in the Cochrane review was rated as “moderate,” in contrast to the Lancet reviews, which reported a “low certainty.”
Fluoxetine can now be considered clinically equivalent to placebo, with a drug-placebo difference of only -2.8 points on the Childhood Depression Rating scale (CDRS-R) which has a range from 17 to 113 (Fig. 1). The Cochrane review concluded that “most newer antidepressants may be associated with small and unimportant reductions in depression symptoms compared with placebo, which raises the question of whether they should be used at all.” Interestingly, neither the Cochrane authors nor most treatment guidelines addressed the loss of efficacy.
Our attempts to explain the loss of fluoxetine’s efficacy
Our initial hypothesis was that there were weaker effects in more recent trials due to a phenomenon known as the “novelty bias.” Indeed, our own meta-analysis provided evidence for this [3] (Fig. 2).

Upon further investigation, we discovered that the Lancet reviews included a small fluoxetine trial with an extremely large effect size (SMD >4) which is highly improbable in medicine. We identified severe problems with this trial and concluded that it is an untrustworthy “zombie” trial. Next, we aimed to reproduce the Lancet reviews and the impact of the zombie trial. When we excluded the zombie, the efficacy of fluoxetine in the Lancet reviews decreased from SMD -0.51 to SMD -0.26 and -0.29. These results were comparable to our reanalysis of the Cochrane review (SMD -0.27). Our findings were rejected by five journals but a preprint is available [5] and, after the five rejections, it will soon be published in a journal.
Why can a treatment lose its efficacy?
Our meta-analysis revealed that estimates of fluoxetine’s efficacy decreased over time, regardless of the zombie trial (Fig. 2). This finding aligns with the novelty bias, a well-known and widespread phenomenon in medicine. Several methodological biases and problematic research practices can explain the novelty bias.
Preregistration, reporting bias, and outcome switching
Earlier trials were conducted before registration and reporting of trials became mandatory. From around 2000 onwards, depending on the field and guidelines, trials have had to be pre-registered. This prevents the study methodology and data analysis from being changed post-hoc to produce more favorable results and ensures that studies with unfavorable results are not left unreported. Figure 3 illustrates the dramatic impact of mandatory registration and reporting. Before mandatory registration, most of the published research findings were positive; afterwards positive trials became the exception.

A 2008 review of antidepressant trials, most of which were conducted before mandatory registration, reported that nearly all failed trials were unpublished [6]. Due to mandatory registration, an increasing number of negative studies have come to light, resulting in a decline in efficacy over time. This likely applies to the database on pediatric antidepressants as well. Indeed, Eli Lilly’s first pediatric fluoxetine trial from 1986 was negative and was never published [7].
Sponsorship Bias
Another important potential mechanism that can explain the decline in efficacy estimates is sponsorship bias. When a new drug is introduced, trials are usually conducted and sponsored by the manufacturer of the drug. For example, Eli Lilly, the manufacturer of fluoxetine, sponsored the two early trials published in 1997 and 2002 [8,9]. Neither trial was pre-registered. A third influential early trial was Treatment of Adolescent Depression Study (TADS), which was publicly funded; however, most of the authors had financial ties to Eli Lilly, so this study can hardly be considered independent. All three trials reported statistically and probably clinically significant results for fluoxetine (see Fig. 2). Later trials were mostly conducted by competing companies, using fluoxetine as an established comparator drug. In these trials, fluoxetine’s efficacy was much smaller and often failed to be statistically significant.
Unblinding and expectancy effects
The decrease in efficacy may be attributed to biases related to expectations and unblinding. Fluoxetine was one of the first SSRIs studied for pediatric depression. It was already a blockbuster drug for adults and was considered safer and better tolerated than previous antidepressants. Consequently, high expectations were placed on fluoxetine when it was initially studied in young patients [10]. Such expectations can bias trial results. A recent reanalysis of the TADS trial revealed that patients, their parents, clinicians, and even independent (blinded) evaluators could guess better than chance who was treated with fluoxetine or placebo [11]. The study also demonstrated the impact of expectancy: there was a difference of 10 CDRS-R points (SMD = 0.7) between those who thought they were on fluoxetine versus those who thought they were on placebo. In contrast, there was only a small difference between those who were actually on fluoxetine or placebo (difference of 2.7 points, SMD = 0.2). Actual treatment allocation could not significantly predict treatment outcome variance above that explained by correct guessing. Importantly, these findings contradict the claim that study participants are unblinded because the drug is effective.
The impact of unblinding may be reduced by decreasing the probability of receiving placebo, such as in multi-arm trials. In more recent trials with multiple arms it has become more difficult to guess exactly which group one is in. Consistent with this assumption, efficacy estimates were substantially higher in the subset of trials with only a 50% chance of receiving placebo (SMD 0.22) than in all trials (SMD 0.12) [12]. Furthermore, positive expectations about the effectiveness of SSRIs have likely decreased due to growing concerns about their efficacy and potential risks, such as increased suicidality.
What we may learn from dropout rates
The aforementioned biases can also impact dropout rates. As a participant in a trial of a promising new drug, you may be disappointed if you end up in the placebo group. You may be able to guess this because of a lack of effects, which may lead you to leave the study. Experiencing adverse events can also lead to dropping out. Therefore, dropout rates can be used to evaluate the acceptability of treatments and their overall harm-benefit ratio [13]. If a treatment is safe and effective, there should be fewer dropouts in the treatment group than in the placebo group.
I conducted meta-analysis/regression analyses that revealed changes in dropout rates in fluoxetine trials over time (Fig. 4). In earlier trials, fewer patients dropped out of the drug arms than the placebo arms, but this trend is not evident in more recent trials.

Furthermore, when fluoxetine was the experimental drug compared only to placebo, there were fewer dropouts in the drug arms than in the placebo arms (Fig. 5). In contrast, when fluoxetine was the comparator drug alongside a more novel drug, the dropout rates of fluoxetine and placebo were similar. Therefore, these results again suggest that fluoxetine is no longer more acceptable than placebo in recent trials.

Implications for other antidepressants
To my knowledge, no one has investigated whether similar biases found for fluoxetine have affected the efficacy estimates of other antidepressants. Now, let us examine the characteristics of placebo-controlled trials of three commonly used antidepressants: sertraline, citalopram, and escitalopram (Table 1). None of the trials included another antidepressant for comparison and all were sponsored by the manufacturer. Only one trial was registered and only one trial was not rated as having a high risk of bias in the Cochrane review. Therefore, these trials are likely affected by the same biases as observed in early fluoxetine trials. Nonetheless, most trials were negative and had small effects. Overall, none of the drugs were statistically significant in the Lancet reviews. In the Cochrane review, only sertraline was statistically significant, but we could not reproduce this finding.
Table 1. Characteristics of clinical trials and meta-analytic results.
| Drug/Study | Registered | Sponsored | Compared with | Outcome | Positive trial? | high ROB? |
| Sertraline | ||||||
| Wagner 2003a | No | Yes | Placebo only | -0.25 (-0.51 to 0.01) | no | yesc,d |
| Wagner 2003a | No | Yes | Placebo only | -0.20 (-0.45 to 0.05) | no | yesc,d |
| NMA Zhou 2020 | -0.11 (-0.71 to 0.49) | |||||
| NMA Hetrick 2021a | -3.51 (-6.99 to -0.04)b -0.23 (-0.48 to 0.02) |
|||||
| Citalopram | ||||||
| Wagner 2004 | No | Yes | Placebo only | -0.34 (-0.61 to -0.09) | yes/no | yes c,e |
| von Knorring 2006 | No | Yes | Placebo only | -0.01 (-0.25 to 0.24) | no | yes c,e |
| NMA Zhou 2020 | -0.18 (-0·89 to 0.55) | |||||
| NMA Hetrick 2021a | -2.90 (-8.23 to 2.43)b
-0.35 (-0.71 to 0.02) |
|||||
| Escitalopram | ||||||
| Wagner 2006 | No | Yes | Placebo only | -0.13 (-0.34 to 0.08) | no | no |
| Emslie 2009 | Yes | Yes | Placebo only | -0.21 (-0.40 to -0.02) | yes | yes c |
| NMA Zhou 2020 | -0.17 (-0.88 to 0.54) | |||||
| NMA Hetrick 2021a | -2.62 (-2.62 to 0.04)b
-0.17 (-0.39 to 0.05) |
aSMDs from the Cochrane review are from our reproduction because they did not report SMDs. The Lancet review (Zhou et al., 2020) included trials with any depression measure and included psychotherapy trials, too. The Cochrane review only included drug trials and only those with the CDRS-R as outcome. bCDRS-R points. cSelective reporting. dOther ROB. eAttrition bias. Abbreviations: NMA: network meta-analysis. ROB: Risk of Bias.
Dropout rates were numerically lower in the placebo arms than in the drug arms in all but one trial (Table 2). This suggests that, when considering overall dropout as a combined measure of efficacy and safety, these three antidepressants were comparable or inferior to treatment with a placebo.
Table 2. Dropout rates in the trials for sertraline, citalopram and escitalopram
| Dropout | ||
| Drug/Study | Drug | Placebo |
| Sertraline | ||
| Wagner 2003a | 32/97
(33.0%) |
14/91
(15.3%) |
| Wagner 2003a | 14/92
(15.2%) |
17/96
(17.7%) |
| Both combined | 46/189 (24.3%) | 31/187 (16.6%) |
| Citalopram | ||
| Wagner 2004 | 22/93 (23.7%) |
18/85 (21.2%) |
| von Knorring 2006 | 45/124 (36.3%) | 46/120
(35.7%) |
| Escitalopram | ||
| Wagner 2006 | 30/132 (22.7%) | 21/136 (15.4%) |
| Emslie 2009 | 32/158 (20.3%) | 25/158 (15.8%) |
Professional ignorance of these findings
First, our review of major treatment guidelines published or updated after the Cochrane review was published revealed that the reduced efficacy of fluoxetine has not been adequately acknowledged and fluoxetine continues to be recommended as a first-line antidepressant for pediatric depression [3].
Second, some key opinion leaders (KOLs) argue that older trials are less biased [14]. For example they argue that earlier trials were conducted in fewer and more experienced study centers, resulting in higher quality. However, we could not find evidence for these arguments in the fluoxetine database [15]. Furthermore, these KOLs ignore the fact that single-center trials can be affected by researcher bias and should therefore be interpreted cautiously [16]. Indeed, older fluoxetine trials from single or few centers were rated as being at high risk of bias, according to the Lancet and Cochrane reviews [1,2,4].
Third, there was a scientifically problematic shift back to relying solely on statistical significance while ignoring the lack of clinically meaningful effects. The Cochrane review concluded that “most newer antidepressants may be associated with small and unimportant reductions in depression symptoms compared with placebo, which raises the question of whether they should be used at all.” However, based on significant (or nearly significant) findings, they also concluded that “if medications are to be used, there is evidence to support a greater range of options for first-line prescribing of antidepressants including sertraline, escitalopram, duloxetine as well as fluoxetine to guideline recommendations that recommend fluoxetine alone (NICE 2019).”
This conclusion was uncritically adopted in the German guidelines from 2025 [17], which recommended the use of fluoxetine, sertraline, and escitalopram for moderate to severe depression. It is argued that there is “strong evidence” for these drugs. Contrast this with the actual finding in the Cochrane review: “there was probably a small unimportant difference between the following NGAs and placebo [sertraline, fluoxetine, escitalopram].” [Emphasis mine.] It seems that the guideline panel simply equated statistical significance with “a strong evidence base,” while disregarding the clinical unimportance of the effects. Furthermore, the findings from the previous Lancet reviews were ignored. These reviews did not find even statistically significant effects for sertraline or escitalopram.
Fourth, the novelty bias has been completely ignored in the interpretation of the findings for antidepressants other than fluoxetine. Commonly used SSRIs, such as escitalopram and sertraline, were investigated a long time ago, without mandatory registration, and almost exclusively by the drug companies. These drugs were never tested by competing companies or independent organizations. Consequently, the efficacy of these drugs is likely overestimated.
Fifth, the concerning results regarding dropout are not acknowledged. If a drug is effective and well tolerated, there should be a lower dropout rate in the drug arms compared to the placebo arms. However, this was only observed in early fluoxetine trials, likely due to methodological biases. For other antidepressants, dropout rates are numerically lower in the placebo arms, despite the methodological biases favoring the drugs.
Finally, the sobering results of independent reanalyses have not been adequately acknowledged. The most prominent example is Study 329, in which paroxetine was deemed “generally well tolerated and effective” [18]. However, independent researchers revealed that paroxetine was not superior to placebo and that there was a significant increase in suicidal events with paroxetine use [19].
A reanalysis of the two early fluoxetine trials (Emslie 1997, 2002) [20] found that “Essential information was missing and there were unexplained numerical inconsistencies. (1) The efficacy outcomes were biased in favour of fluoxetine by differential dropouts and missing data. The efficacy on the Children’s Depression Rating Scale-Revised was 4% of the baseline score, which is not clinically relevant. Patient ratings did not find fluoxetine effective. (2) Suicidal events were missing in the publications and the study reports. Precursors to suicidality or violence occurred more often on fluoxetine than on placebo.”
A recent reanalysis [21] of the influential TADS study concluded that “In contrast to the original TADS Team’s reporting, there was a higher level of harm uncovered in allocation groups taking fluoxetine, including 11 unreported suicide-related adverse events.”
Another example is the reanalysis of the CIT-MD-18 study on citalopram [22]. It was revealed that the publication was ghost-written and that “the published article contained efficacy and safety data inconsistent with the protocol criteria. Procedural deviations went unreported imparting statistical significance to the primary outcome, and an implausible effect size was claimed; positive post hoc measures were introduced and negative secondary outcomes were not reported; and adverse events were misleadingly analysed.”
In their conclusion, the authors of the TADS reanalysis stated that “The findings of this as well as several other reanalyses suggest that antidepressant trial results cannot be taken at face value” [21]. Evidence syntheses and treatment guidelines generally lack this justified scepticism.
Consequences for the harm-benefit assessment
Evidence from clinical trials suggests that all antidepressants investigated for treating pediatric depression can be considered clinically equivalent to placebo. This is especially true when considering the biases revealed in the analysis of fluoxetine trials, which were ignored when interpreting the evidence for other antidepressants. Consequently, even minor adverse events can lead to an unfavorable harm-benefit ratio. Adverse events that occur more frequently during treatment with SSRIs than placebo include sexual dysfunction, gastrointestinal problems, insomnia, and suicidality. Therefore, the harm-benefit ratio is problematic for many patients. Unfortunately, current treatment recommendations and clinical practices are based on an uncorrected, uncritical interpretation of the evidence, which can harm patients. Necessary corrections are overdue.
References
1 Cipriani A, Zhou X, Del Giovane C, et al. Comparative efficacy and tolerability of antidepressants for major depressive disorder in children and adolescents: a network meta-analysis. The Lancet. 2016;388:881–90. doi: 10.1016/S0140-6736(16)30385-3
2 Zhou X, Teng T, Zhang Y, et al. Comparative efficacy and acceptability of antidepressants, psychotherapies, and their combination for acute treatment of children and adolescents with depressive disorder: a systematic review and network meta-analysis. Lancet Psychiatry. 2020;7:581–601. doi: 10.1016/S2215-0366(20)30137-1
3 Plöderl M, Lyus R, Horowitz MA, et al. The loss of efficacy of fluoxetine in pediatric depression: explanations, lack of acknowledgment, and implications for other treatments. J Clin Epidemiol. 2026;189:112016. doi: 10.1016/j.jclinepi.2025.112016
4 Hetrick SE, McKenzie JE, Bailey AP, et al. New generation antidepressants for depression in children and adolescents: a network meta-analysis. Cochrane Database Syst Rev. 2021;2021. doi: 10.1002/14651858.CD013674.pub2
5 Lyus R, Naudet F, van Valkenhoef G, et al. A Re-Appraisal of Three Network Meta-Analyses to Explain the Discrepancy in Findings for the Efficacy of Fluoxetine for the Treatment of Depression in Children and Adolescents. medRxiv. 2025;2025.09.07.25334757. doi: 10.1101/2025.09.07.25334757
6 Turner EH, Matthews AM, Linardatos E, et al. Selective publication of antidepressant trials and its influence on apparent efficacy. N Engl J Med. 2008;358:252–60. doi: 10.1056/NEJMsa065779
7 Eli Lilly. Clinical Study Summary: Study B1Y-MC-HCCJ. 1986.
8 Emslie GJ, Rush AJ, Weinberg WA, et al. A double-blind, randomized, placebo-controlled trial of fluoxetine in children and adolescents with depression. Arch Gen Psychiatry. 1997;54:1031–7. doi: 10.1001/archpsyc.1997.01830230069010
9 Emslie GJ, Heiligenstein JH, Wagner KD, et al. Fluoxetine for Acute Treatment of Depression in Children and Adolescents: A Placebo-Controlled, Randomized Clinical Trial. J Am Acad Child Adolesc Psychiatry. 2002;41:1205–15. doi: 10.1097/00004583-200210000-00010
10 Boussageon R, Gougeon A, Kassaï B. Some additional considerations on the evidence for fluoxetine in pediatric depression. J Clin Epidemiol. 2026;112137. doi: 10.1016/j.jclinepi.2026.112137
11 Jureidini J, Moncrieff J, Klau J, et al. Treatment guesses in the Treatment for Adolescents with Depression Study: Accuracy, unblinding and influence on outcomes. Aust N Z J Psychiatry. 2024;58:355–64. doi: 10.1177/00048674231218623
12 Feeney A, Hock RS, Fava M, et al. Antidepressants in children and adolescents with major depressive disorder and the influence of placebo response: A meta-analysis. J Affect Disord. 2022;305:55–64. doi: 10.1016/j.jad.2022.02.074
13 Sharma T, Guski LS, Freund N, et al. Drop-out rates in placebo-controlled trials of antidepressant drugs: A systematic review and meta-analysis based on clinical study reports. Int J Risk Saf Med. 2019;30:217–32. doi: 10.3233/JRS-195041
14 Walkup JT, Strawn JR. Depressive disorders in children and adolescents. In: Thapar A, Pine DS, Cortese S, et al., eds. Rutter’s Child and Adolescent Psychiatry and Psychology. Wiley 2025:898–914.
15 Plöderl M, Lyus R, Naudet F. Can fluoxetine’s diminished efficacy in pediatric depression be explained by study sites, baseline severity, age, and psychological interventions? A secondary exploratory meta-regression. J Psychiatr Res. 2026;202:160–3. doi: 10.1016/j.jpsychires.2026.08.018
16 Unverzagt S, Prondzinsky R, Peinemann F. Single-center trials tend to provide larger treatment effects than multicenter trials: a systematic review. J Clin Epidemiol. 2013;66:1271–80. doi: 10.1016/j.jclinepi.2013.05.016
17 Kloek M, Zsigo C, Klingele C, et al. S3-Leitlinie Behandlung von depressiven Störungen bei Kindern und Jugendlichen. Deutsche Gesellschaft für Kinder- und Jugendpsychiatrie, Psychosomatik und Psychotherapie e.V. (DGKJP) 2025.
18 Keller MB, Ryan ND, Strober M, et al. Efficacy of Paroxetine in the Treatment of Adolescent Major Depression: A Randomized, Controlled Trial. J Am Acad Child Adolesc Psychiatry. 2001;40:762–72. doi: 10.1097/00004583-200107000-00010
19 Le Noury J, Nardo JM, Healy D, et al. Restoring Study 329: efficacy and harms of paroxetine and imipramine in treatment of major depression in adolescence. BMJ. 2015;h4320. doi: 10.1136/bmj.h4320
20 Gøtzsche PC, Healy D. Restoring the two pivotal fluoxetine trials in children and adolescents with depression. Int J Risk Saf Med. 2022;33:385–408. doi: 10.3233/JRS-210034
21 Aboustate N, Jureidini J, Woodman R, et al. Restoring TADS: RIAT reanalysis of the Treatment for Adolescents with Depression Study. Int J Risk Saf Med. 2025;9246479251337879. doi: 10.1177/09246479251337879
22 Jureidini JN, Amsterdam JD, McHenry LB. The citalopram CIT-MD-18 pediatric depression trial: Deconstruction of medical ghostwriting, data mischaracterisation and academic malfeasance. Int J Risk Saf Med. 2016;28:33–43. doi: 10.3233/JRS-160671













