Understanding partial compliance in remote data collection
Remote studies rarely collect a perfectly complete dataset, and it is a mistake to expect one. A participant misses a diary entry here, skips a check-in there, or starts an adverse event report they never finish. None of this means something has gone wrong: it is simply what happens once data collection moves out of a clinic and into people's daily lives. The harder question is not how to prevent every gap, but what to do with the data once the gaps appear.
Older models treated completeness as the price of admission, dropping anyone who missed enough visits from the analysis entirely. That made more sense when data was collected face to face. In decentralised and real-world studies, it fits poorly: insisting on perfect compliance does not make the data more reliable, it just quietly removes the participants who found the protocol hardest to follow, often exactly the group whose experience matters most.
The shape of the problem also shifts with what is being collected. A nutritional trial logging meals daily produces a different pattern of gaps to an observational study syncing wearable data in the background, which looks different again from a clinical trial built around scheduled ePRO check-ins. One review of an endometriosis symptom diary found that missingness clustered in predictable, explainable ways rather than appearing as random noise. Treating every gap as identical "non-compliance" throws away exactly the pattern a well-designed study should be looking for.
Decide what "usable" means before the study starts
These read as statistical questions, but they are just as much a design brief:
- Is five days out of seven enough for a reliable weekly average?
- Does one missed check-in invalidate the rest of that week's entries?
- What is the minimum threshold for including a participant in the primary analysis?
Once a threshold is agreed, the platform can be built to support it rather than treating a gap as an error. The agreement shouldn't sit with the data team alone. Statisticians, study leads, and where relevant an ethics committee should sign off before the study goes live, with the decision documented in the statistical analysis plan.
Read the shape of the gaps, not just their number
| Pattern | What it usually means |
|---|---|
| A single missed entry | Usually just noise |
| A steady decline over several weeks | A signal worth investigating |
| A quiet gap followed by a sudden batch of back-filled entries | A different signal again, often reminder-related |
A comparison of paper- and web-based outcome collection after thoracic surgery found the two methods differed not just in completeness but in when gaps occurred, a reminder that missingness says as much about the collection method as about the participant. Declining engagement can point to fatigue or a side effect; entries clustered before a deadline suggest reminders aren't landing; gaps confined to one form usually mean that screen has a usability problem.
Dashboards that surface these patterns as they happen, rather than after database lock, make this practical to act on. A coordinator checking one weekly can catch a slow decline before it becomes a withdrawal, echoing what a smartphone-based surveillance study of patient-reported outcomes in advanced cancer found: routine, low-friction monitoring caught changes in adherence early enough to respond, rather than only after the data was already lost.
Design out the gaps that are avoidable
Some missingness is unavoidable. Some is created by the study itself:
- Daily entries required when a weekly average would do just as well
- No recovery path after missing a few days, so participants who fall behind feel there's no point returning
- Copy that frames a missed task as a failure rather than something normal
- No visible sense of progress to sustain motivation over a long study
- Onboarding that never explains why completeness matters
None of this needs a redesign. Often it's one reminder rewritten, or one optional field added so a participant can explain a gap rather than leave it blank.
Respond in proportion to what happened
A missed entry doesn't automatically need a formal query. Often a short, supportive message is enough, and coordinators who treat partial compliance as ordinary rather than alarming tend to get better results from that outreach. Telling a disengaging participant apart from one simply having a difficult week takes judgement: the tone of a message, and how quickly it's sent, can decide whether someone returns or quietly stops responding.
Report how it was handled, not just what was collected
Whatever thresholds were agreed belong in the final write-up. Stating plainly how partial data was handled, what counted as usable, and how much was excluded gives readers a reason to trust the result, and gives the next study team something to build on rather than relearning the same lessons from scratch.
Partial compliance is not a flaw to engineer away. It is a normal feature of collecting data from real people living real lives, and studies designed with that in mind end up with data that is more inclusive, more honest, and ultimately more credible: not despite the gaps, but because of how openly they were handled.
References
- Patterns of missing data in the use of the endometriosis symptom diary
- Data Quality of Longitudinally Collected Patient-Reported Outcomes After Thoracic Surgery: Comparison of Paper- and Web-Based Assessments
- PROutine: a feasibility study assessing surveillance of electronic patient reported outcomes and adherence via smartphone app in advanced cancer