What a systematic review of smartphone EMA found about compliance
Ecological momentary assessment, prompting participants to report on their state in the moment, repeatedly, across a study period, has become a standard tool for capturing data that a single retrospective questionnaire can't reach. A systematic review of smartphone-based EMA studies measuring well-being found a reporting gap worth noting before getting to the results themselves: only 47.2% of the studies reviewed reported a compliance rate at all.
The compliance gap starts with whether anyone measured it
That figure matters on its own terms. If fewer than half of published EMA studies report how many of the requested assessments participants actually completed, then a substantial share of the published literature on well-being fluctuations is working from data of unknown completeness. A finding drawn from a dataset where 90% of prompted assessments were completed is a different kind of evidence than the same finding drawn from a dataset where only 40% were, and readers can't tell which they're looking at when compliance simply isn't reported.
Among the studies that did report it, mean compliance was 71.6%. That's a reasonable, workable figure for an intensive data collection method, not a red flag on its own. But it's also meaningfully short of the near-complete response rate that some study designs implicitly assume when they treat EMA data as a continuous, gapless signal rather than a sample with real, uneven gaps in it.
Study designs varied enormously, with no clear optimal format
The reviewed studies varied widely in structure: average study duration was 12.8 days, with well-being assessed anywhere from 2 to 12 times per day depending on the specific study. That's a wide range for what's nominally the same measurement approach, and it reflects a field where researchers are making genuinely different choices about the frequency-versus-burden trade-off without a clearly established optimal answer to point to.
This heterogeneity has a direct practical consequence: findings from a study prompting participants twice daily over two weeks and a study prompting twelve times daily over the same period aren't measuring quite the same thing, even though both would be described as "EMA of well-being." Comparing or combining results across such different designs is inherently harder than the shared label suggests.
Context turned out to explain more than time of day
One of the review's more interesting substantive findings concerns what actually drives well-being fluctuations. Across the reviewed studies, well-being was found to be higher in evenings and on weekends, a pattern that might read as simply "people feel better outside of work hours." But when studies accounted for a participant's actual location and activity at the time of each assessment, these day-of-week and time-of-day fluctuations disappeared. In other words, it wasn't evening or weekend timing itself driving higher well-being. It was what people were actually doing and where they were, being in nature and engaging in physical activity were both associated with higher well-being, while working was associated with lower well-being, and evenings and weekends simply correlated with more of the former and less of the latter.
That distinction matters for how EMA data gets interpreted. A naive analysis treating "time of assessment" as the explanatory variable would have concluded something about circadian or weekly rhythms in well-being. The more careful analysis, accounting for context, found the real driver was activity and environment, a genuinely different and more actionable finding.
What the review recommends, and why it matters for study design
The review's authors identified a specific gap in the existing literature: most studies focus on group-level comparisons, contrasting average well-being across conditions or populations, while research into individual differences in well-being patterns and fluctuations remains comparatively rare. Given that EMA is uniquely suited to capturing exactly this kind of within-person variation over time, that's a notable underuse of the method's particular strength.
Their broader recommendation was for improved standardisation across measurement instruments, objective data collection methods, and analytical approaches, directly addressing the heterogeneity the review documented.
Practical implications for a study using EMA
A few points follow for anyone designing a study around this kind of intensive, repeated assessment:
- Report compliance as a matter of course, not an optional detail. The finding that over half the reviewed studies didn't is a gap worth actively avoiding, not repeating.
- Plan for real-time compliance monitoring, not just end-of-study calculation. A 71.6% average compliance figure, calculated only after the study concludes, is far less useful than visibility into which participants are falling behind while there's still time to re-engage them.
- Account for context, not just timing, when analysing patterns in the resulting data. The review's finding that location and activity explained fluctuations that looked, at first glance, like simple time-of-day effects is a caution against a purely time-series reading of EMA data.
- Match assessment frequency to what the study can actually sustain, given the wide range of frequencies used across the reviewed studies and the absence of a clearly superior standard to default to.
The underlying message is that EMA's value depends heavily on execution details that are easy to under-specify: how compliance is tracked, how frequently participants are prompted, and how context gets folded into the analysis. The method itself is well established. Doing it well, on the evidence of this review, still isn't.