What a passive EHR extraction pilot actually achieved
Manual transcription between a participant's electronic health record and a study's data capture system is one of the more obvious places for errors to creep into trial data, and one of the more obvious places to remove them entirely if the systems can just talk to each other directly. A multi-centre observational study in diabetic patients tested exactly that: passive extraction of data from eConsent, ePRO, and EHR sources, loaded straight into an EDC system across three sites, rather than staff re-entering it by hand.
The core promise held up
The headline result was straightforward and important: using eSource data eliminated transcription errors. That's not a marginal improvement to an existing process; it removes an entire category of data quality risk at the source, rather than catching and correcting it downstream through query resolution.
The practical integration also worked more smoothly than a more sceptical prediction might have expected. Extracting data from EHR systems across three separate sites required only minimal data transformation and normalisation to feed into the EDC, rather than the substantial custom engineering that connecting genuinely different systems across different sites might imply. The pipeline captured medication changes, healthcare encounters, and lab results as they occurred in standard clinical practice, rather than requiring a separate, parallel data entry process layered on top of routine care.
The actual numbers show where completion drops off
The study enrolled 48 diabetic participants on metformin monotherapy. HbA1c measurements, the core clinical outcome, were captured from 33 of the 48 participants at both baseline and the 12-week follow-up. Electronic patient-reported outcomes, using the SF-12 survey, were collected from all 48 participants at baseline, but only 28 by the 12-week endpoint.
That's a real, specific completion gap worth stating plainly rather than glossing over: a drop from 48 to 28 on the ePRO measure between baseline and 12 weeks is a meaningful decline, even in a study whose central innovation was reducing the operational friction of data collection. Passive extraction from EHR and eConsent addresses the transcription and integration problem cleanly. It doesn't automatically solve participant-side engagement over a follow-up period, since ePRO completion still depends on the participant actually completing the survey, in a way that passively extracted clinical data, sourced from encounters that were going to happen anyway, doesn't.
That distinction is worth being precise about. The HbA1c completion rate (33/48, around 69%) and the ePRO completion rate at 12 weeks (28/48, around 58%) are measuring different kinds of dependency: one on clinical encounters continuing to occur and being captured, the other on a participant independently completing a survey. A study relying on this kind of architecture needs to track both separately, since a strong result on one doesn't guarantee the other.
Why the authors see this as a scalable model, with a caveat
The researchers' own conclusion was optimistic about the approach's potential, describing it as something that could be expanded for larger trials and would significantly reduce staff effort. That's a reasonable extrapolation from a pilot of this size demonstrating the core technical integration works with minimal per-site customisation. The completion gap on the ePRO side is the part of the pilot that a larger-scale application would need to address directly, since scaling up a study without also addressing participant-side follow-through simply reproduces the same completion gap at greater volume.
What this means for a study considering this kind of architecture
A few practical takeaways follow from a pilot this specific:
- Eliminating transcription error is a genuine, immediate win, and one of the clearest cases for this kind of integration. It doesn't depend on participant behaviour and delivers its benefit purely through architecture.
- Minimal per-site transformation across genuinely different EHR systems is achievable, which matters for anyone assuming a multi-site passive extraction approach would require heavy custom engineering for every site individually.
- Participant-reported outcome completion needs its own dedicated attention, separate from the clinical data pipeline. The gap between baseline and 12-week ePRO completion in this pilot is a reminder that removing transcription friction from clinical data doesn't automatically remove engagement friction from participant-facing measures.
- Track completion by data source, not as a single aggregate figure. The different completion trajectories for HbA1c versus ePRO in this pilot would have been invisible behind a single blended completion statistic, and the distinction is exactly what a study team needs to know to act on it.
The broader lesson is that passive data extraction genuinely solves the problem it's designed to solve, transcription accuracy and integration overhead, cleanly and close to completely. It doesn't extend that same benefit to participant-completed measures by default, and a study built around this architecture needs a separate, deliberate plan for the parts of data collection that still depend on the participant actually showing up to do something.