How to know if your eSource is really your source
Not everything entered electronically is eSource. That sentence is worth reading twice, because a lot of study teams assume it is.
The reasoning feels logical: the data goes into a digital system, therefore it's electronic source data, therefore you can call it eSource and proceed accordingly. But whether something counts as eSource depends on where the data was first captured, whether the system maintains a genuine audit trail, and whether everyone involved (monitors, sponsors, investigators) agrees on which record is the authoritative one. Get this wrong and you've built a compliance problem into the foundation of the study.
What "source" actually means
Source data, in regulatory terms, is the original record. The first place a value exists. Whether it's electronic or not is secondary to whether it's primary. So the question isn't "did this data go into an electronic system?" It's "was this system where the data was born, or where it was transferred?"
A site-based eCRF that's locked, audit-trailed, and reviewed by monitors? That's likely eSource. The same form filled in later from a handwritten note found in a desk drawer? That's transcription, not source, regardless of what the platform calls itself. This is the same distinction underpinning the long-standing ALCOA+ data quality standard: for research data to be trustworthy, it needs to be attributable, legible, contemporaneous, original, and accurate, alongside being complete, consistent, enduring, and available. A recent proposal for a formal eSource record system built its entire architecture around exactly this problem, designing a five-step pipeline from initial data collection through to eCRF traceability specifically so that, at any point, a reviewer could trace a value back to where it was actually born rather than trusting a label.
Run this test
For any data point in your study, trace it backwards:
- Where did it first exist?
- Was it entered directly into the system, or transferred from somewhere else?
- If a monitor asked to verify the original entry, what would you show them?
- If the value was ever changed, can you show who changed it, when, and why?
If you can answer all four cleanly, the system is probably functioning as a legitimate source. If any answer involves "we'd have to check the paper backup" or "I think it was entered later from a photo," you have a hybrid situation that needs documentation.
Common grey areas
Participant apps. A participant enters a symptom score on their phone at 7am. That entry is the source, not the eCRF where a coordinator later copies it. If the app is transferring data into a secondary system, the app is the source and the secondary system is a transcription layer. Not the same thing.
Wearables. The device captures heart rate or sleep duration. The investigator types a summary into the eCRF. The wearable is the source. The eCRF is a copy. This matters if the wearable data is ever disputed.
Complex multi-system studies. It's not unusual for a study to use a participant app, a site-based eCRF, and a sponsor reporting environment simultaneously. Each might hold different types of data. Each needs to be traced individually. Declaring "we use eSource" as a blanket statement doesn't hold up when there are six systems involved, and the proposal above found the same problem at scale: two systems rarely use the same underlying data standard by default, so mapping between them has to be handled deliberately rather than assumed to just work.
The documentation question
Whatever you conclude about your source systems, it needs to be written down. The protocol or data management plan should specify, for each data category: where it originates, how it's transferred if at all, who has authority to correct it, and how those corrections are recorded.
This isn't just audit preparation. It's the difference between a team that knows what it's working with and one that discovers the problem at database lock, when there's no longer a straightforward way to reconstruct what actually happened.
eSource isn't a label you apply. It's a status you earn by demonstrating that the data is original, unaltered, and complete in that location. If you can show that clearly, consistently, for every data type, you're in good shape.