The parts of source data verification that still need a human eye
Remote monitoring has come a long way. In digital and decentralised trials, automation can now catch missing fields, flag range violations, and trigger queries without anyone needing to open a binder. But there are still parts of source data verification that depend entirely on human judgement, and it is worth being clear about which parts those are.
What technology does well
Systems supporting eSource or direct data capture bring genuine benefits:
- Highlighting incomplete or out-of-range fields before they reach the monitor
- Tracking timestamps and user actions in the audit trail
- Enabling faster, more frequent review cycles
- Standardising terminology and logic across sites so like-for-like comparison is possible
These functions help monitors identify problems earlier and prioritise where to spend their time. That is real progress, and it goes further than simple range checks. A methods paper describing central statistical monitoring programs applied to real trial data showed the approach could automatically detect fabricated patient records generated to sit suspiciously close to the group average, and could flag entire sites with unusual correlation structures or an implausibly low rate of reported adverse events. When fabricated data was deliberately added to one of the trials studied, the programs caught it.
Where human review is still irreplaceable
The problem with automation is that it checks against rules and statistical assumptions. It does not understand context, and a lot of data quality problems live in context. The same central monitoring research was candid about this limitation: its methods struggled with small trials or sites with fewer than ten patients, and data falsified in a way that didn't match the programs' underlying assumptions could slip through undetected. Every flagged output, the authors noted, still needed a visual, human assessment before anyone acted on it.
Consider some concrete cases:
Readings that are technically in range but clinically implausible. A blood pressure value within acceptable limits that does not match the participant's history, previous readings, or concurrent medications. The system clears it. A trained reviewer spots the inconsistency.
Free-text adverse event descriptions. Automation can flag that a narrative was submitted. It cannot read it for nuance, assess whether it contradicts other entries, or decide whether it warrants escalation.
Cross-system data linkage. When participant app data, eCRF entries, and lab results all need to reconcile, and the integration is imperfect, a human has to make sense of the discrepancies.
Behavioural patterns at sites. A team that completes forms consistently in the hour before a monitoring visit may look clean in the system. A reviewer familiar with the site might recognise this as a pattern worth questioning, in much the same way the statistical monitoring research flagged sites with too few adverse events as worth a second look rather than treating a clean-looking record as automatically reassuring.
The grey areas are where risk concentrates
Most data quality problems are not obviously wrong. They sit in the space between clearly correct and clearly incorrect.
| Grey area | The question a human has to answer |
|---|---|
| A missing value | Genuine skip, or an overlooked field? |
| A form flagged for review | Should it be returned for correction, or accepted with a documented comment? |
| A cluster of similar adverse event reports | Coincidence, or a signal worth escalating? |
| A site flagged by statistical monitoring | Genuine anomaly, or a small-sample artefact the programs weren't built to handle? |
These decisions require someone who understands the protocol, the data model, and how the study is actually running operationally. No amount of statistical sophistication removes the need for that judgement call. It just narrows down which records deserve it.
What good remote systems should do
The goal is not to remove the human review layer. It is to make that layer more effective:
- Surface what needs attention rather than burying it in routine activity
- Reduce time spent on checks that add no value
- Give reviewers fast access to history, context, and linked records
When systems do that well, monitors spend less time checking boxes and more time making the judgements that only they can make, in full awareness that the system's flags are a starting point for investigation rather than a verdict. That is what good oversight looks like in a digital study.