Building a minimal dataset that still tells the whole story
There's a temptation in research, especially digital research where adding a new field costs almost nothing, to collect everything and sort it out later. It feels responsible. Thorough. Future-proof. But "while we're at it" thinking has a habit of creating studies that are harder to run, harder to clean, and harder to finish.
This isn't a problem unique to any one field or a matter of individual discipline. An international consensus project building a minimum data set for degenerative cervical myelopathy research went through nine formal phases, from an initial longlist of outcomes down to a final set of just four validated measurement tools, specifically because getting from "everything that could matter" to "what actually matters" turned out to need a structured process, not just good intentions. If a field with decades of prior research still needed a dedicated multi-year consensus exercise to answer that question properly, it's a reasonable bet that most individual studies are answering it too casually.
Here's a practical way to build a dataset that's genuinely minimal without being bare.
01. Start with the non-negotiables
What absolutely must be measured? Not ideally. Not "it would be interesting." What does the study fail to answer without? These fields get priority in everything: app interface placement, reminder scheduling, monitoring attention. If the primary outcome is a weekly gut symptom score, that score gets surfaced first, scheduled first, and cleaned first.
02. Ask what supports those outcomes
Some fields don't answer the primary question, but they help interpret it. Adherence indicators. Contextual data like timing or environment. Participant feedback that flags issues with a specific visit. These belong, but only if they're doing a specific job. "We might need this" is not a job.
03. Challenge everything else
For every remaining field, ask three things: what decision will this data inform, if this field came back completely blank what would we actually miss, and could it be estimated or inferred from something else we're already collecting? Most fields that fail this test don't disappear from the protocol. They get collected less frequently, made optional, or rotated across visits, in much the same way the myelopathy consensus project didn't discard every candidate outcome outright but sorted them into a smaller core set and a wider, optional data element list.
04. Design the participant experience around what's left
Once the essential fields are identified, the interface should reflect that hierarchy. A participant shouldn't have to scroll past ten optional questions to reach the one that matters. And the emotional experience of completion matters too. A short form that a participant finishes in 90 seconds and closes feeling good about is more valuable than a comprehensive one they abandon halfway through.
The practical gains are real: shorter forms, fewer queries, cleaner exports, faster analysis. But the less obvious benefit is focus. When the dataset is minimal by design, the team stays closer to the research question, because the research question is all that's left.
Collect enough to answer the question. Not more.