What a decade of depression trials reveals about results reporting
Registering a trial and actually reporting its results are meant to be two steps in the same obligation. A cross-sectional study of 442 depression trial protocols registered on ClinicalTrials.gov between January 2008 and May 2019 found a substantial gap between the two: median time from study completion to posted results was over two years, and five years from initial registration.
The delay itself distorts the available evidence
That's a long enough gap to matter for reasons beyond simple tidiness. A systematic review or meta-analysis conducted at any given point in time can only draw on results that have actually been posted by then. If posting is delayed by a median of two-plus years after completion, and if the trials that get posted promptly differ systematically from those that don't, whether by outcome, sponsor type, or result direction, then any evidence synthesis conducted before the slower trials catch up is working from a skewed, non-random subset of the true evidence base.
This isn't a hypothetical risk. It's precisely the mechanism by which selective or delayed reporting distorts systematic reviews: not through fabrication, but through timing. A negative or null finding that takes longer to write up and post, whether from lower priority, less enthusiasm, or genuine reluctance, systematically under-represents itself in any evidence synthesis conducted before it eventually appears.
The effect sizes themselves carried a specific warning sign
Among the protocols where the researchers could calculate an effect size, 134 in total, the median effect was small, 0.16. That's unremarkable on its own; small effects are common and not evidence of a problem. What is notable is that for 28% of these protocols, the observed effect ran contrary to the expected direction, meaning the trial's actual result pointed the opposite way from what the study had been designed to detect.
A rate that high is worth sitting with. It doesn't necessarily indicate anything went wrong with any individual trial. It does suggest that a meaningful share of depression trials produce results that don't confirm their starting hypothesis, and that any evidence base drawing primarily on trials with results in the expected direction, because those get reported faster or more completely, would present a systematically rosier picture than the underlying research actually supports.
A specific, avoidable reporting gap compounded the problem
Beyond timing, the study identified a concrete methodological gap in how results were actually reported: between-group effect size calculations frequently had to be based on post-treatment data alone, because pre-treatment data wasn't consistently provided in the posted results. That's not a subtle statistical nuance. Baseline data is foundational to interpreting a between-group comparison correctly, and its inconsistent inclusion in posted trial results is a data completeness problem sitting on top of the timing problem, compounding the difficulty of drawing an accurate picture from what's actually available on the public record.
Why this matters beyond depression trials specifically
The researchers' own conclusion draws the throughline directly: failure to post results in a timely manner, combined with incomplete statistical reporting, can lead to overestimates of treatment effects in systematic literature reviews. Depression trials were this study's specific focus, but the underlying mechanism, delayed and incomplete public reporting distorting downstream evidence synthesis, isn't specific to one condition. It's a structural feature of how clinical trial reporting works wherever the same gaps between completion, reporting, and reporting completeness show up.
What this means for how results reporting actually gets handled
A few practical implications follow for any study team responsible for eventually posting results:
- Treat the results-posting deadline as a genuine operational milestone, not an administrative afterthought that happens whenever capacity allows. A median two-year delay after completion suggests this step is being deprioritised relative to other post-trial work across a large share of studies, not just occasionally slipping.
- Baseline data needs to be captured and retained in a form that makes it straightforward to include in the eventual results posting. The gap this study found, between-group calculations relying on post-treatment data alone because baseline data wasn't consistently available, is precisely the kind of problem prevented by structured, consistently captured data from the study's outset, rather than baseline information scattered across less accessible records by the time results reporting happens.
- A result running contrary to the expected direction is not a reason to deprioritise reporting it. Given that 28% of the protocols in this study showed exactly that pattern, ensuring reporting speed doesn't correlate with whether the result was the one that was hoped for is directly relevant to the integrity of the wider evidence base, not just this individual study's record.
The broader point is that the credibility of published clinical evidence depends on more than any single trial being conducted well. It depends on the reporting pipeline behind it, timeliness and completeness, functioning consistently across trials regardless of what the result turned out to be. This study's findings suggest that pipeline still has real gaps, and that those gaps have a measurable effect on what the published evidence actually shows.