Why more data isn't always better in clinical trials
Picture a nutritional study where participants are asked to log their meals, mood, energy, gut symptoms, supplement intake, sleep quality, and daily activity, every day, for six weeks.
That sounds comprehensive. In practice, it tends to look like this: strong completion rates in week one, patchy data by week three, and a monitoring team spending most of their time chasing missing entries rather than reviewing anything meaningful. More data, less insight.
It's an easy trap to fall into, because every individual field looks reasonable in isolation. A reviewer suggests adding a sleep quality question because it might be relevant. A safety lead wants a supplement log in case of an interaction. Nobody sets out to overload a study. Each addition is small, defensible, and considered on its own merits, which is exactly why the total ends up so much larger than anyone intended.
This isn't a new problem, and it's not unique to any one study. A methods paper on shortening patient-reported outcome measures makes the point plainly: collecting long, cumbersome instruments burdens participants, increases research costs, and can actively reduce the quality of the data collected, not just its volume. The authors' response wasn't to argue for collecting less out of principle. It was to develop a formal statistical method for choosing which items to keep, so that shortening a measure is a deliberate design decision rather than an afterthought once fatigue has already set in.
The problem isn't the intention. It's the assumption that volume and value are the same thing. Here's what actually happens when data collection isn't kept in check:
| What increases | What suffers |
|---|---|
| Number of fields and forms | Participant completion rates |
| Query volume from monitors | Time spent on real safety signals |
| Variables sent to analysis | Proportion of variables ever used |
| Study team burden | Team focus on the research question |
The useful filter isn't "could we collect this?" It's: if this field came back completely blank, would it change the study conclusions? If the answer is no, it probably doesn't belong. If it does belong, the question becomes how to collect it with the least possible friction: reducing frequency from daily to weekly, making a field optional with a quick "nothing to report" tap, or rotating questions across visits rather than asking everything at once.
None of this means being careless about completeness. It means being deliberate about what completeness actually requires, in the same spirit as choosing outcome measure items through a formal selection process rather than including every question a reviewer has ever suggested.
There's a version of "thorough" that produces rich, trustworthy data. And there's a version that produces a dataset so wide and patchy that the analysis team spends weeks cleaning variables that were never going to be used. The difference usually comes down to whether someone asked the hard question during protocol design: not "what could we collect?" but "what do we actually need?"
That question is worth revisiting after the protocol is finalised too, not just before. A field that seemed essential during design can turn out, three months into a study, to be the one nobody has looked at since the first monitoring visit. Cutting it mid-study is harder than never adding it, which is exactly why the filter matters most at the point when adding a field is still cheap and removing one hasn't yet become a protocol amendment.
More data is fine. More data without purpose is just noise wearing a research hat.