Phone-Based Vitals Screening Results: 5 Field Programs Compared
A benchmark of rPPG field deployment results across five phone-based vitals screening programs, comparing accuracy, sample size, and real-world reach for researchers.

Researchers evaluating camera-based vital signs face a recurring problem: most published accuracy numbers come from controlled laboratory rooms with cooperative volunteers, stable lighting, and a single reference device. Field conditions are different. A market-day screening line, an outdoor compound at midday, a low-end Android phone, and a diverse population across the Fitzpatrick scale all change the picture. To benchmark methods honestly, researchers need rPPG field deployment results reported alongside the sample sizes and conditions that produced them. This comparison pulls together accuracy and reach figures from five published phone-based screening efforts so that method developers and grant reviewers can read them on the same page.
A 2024 meta-analysis of smartphone photoplethysmography found a pooled mean difference of just -0.32 bpm against validated reference methods for heart rate, yet blood pressure estimation in the same body of literature still showed systolic errors above 14 mmHg in several real-world apps.
Reading rPPG field deployment results against a common benchmark
The phrase remote photoplethysmography, or rPPG, covers a family of techniques that recover a pulse signal from subtle color changes in facial or fingertip skin captured by a standard camera. The science is settled enough that heart rate recovery is generally robust. The open questions are about everything downstream: blood pressure, respiratory rate, oxygen saturation, performance across darker skin tones, and whether accuracy holds when the camera leaves the lab. When comparing rPPG field deployment results, three dimensions matter more than a single headline error figure.
- Vital coverage: heart rate alone is far easier than a full panel including blood pressure and SpO2.
- Reference standard: invasive arterial lines, validated cuffs, and ECG produce very different error envelopes.
- Population and setting: sample size, skin tone distribution, and indoor versus outdoor capture decide whether a number generalizes.
A study that reports a 2 bpm heart rate error on 30 indoor volunteers is not comparable to one reporting blood pressure concordance across thousands of patients in four countries. The table below places five reference efforts side by side so the differences are visible rather than buried in abstracts.
| Program / method | Vitals reported | Reach / sample | Reported accuracy | Setting |
|---|---|---|---|---|
| OptiBP multi-country study (Schoettker et al., 2024) | Blood pressure | Multi-country validation, low-resource sites | >90% concordance with invasive BP in 121 anesthesia patients; multi-country accuracy testing | Clinical + field, diverse countries |
| WellFie rPPG application (2023, medRxiv) | HR, RR, SBP, DBP | Validation cohort study | MAPE 2.66% HR, 15.66% RR, 6.06% SBP, 7.05% DBP | Supervised clinical capture |
| ReViSe framework (arXiv, 2023) | HR, SpO2, BP | Multiple public datasets | MAE 1.73-3.95 bpm HR, 1.64% SpO2, 6.7 mmHg SBP, 9.6 mmHg DBP | Benchmark datasets, mixed conditions |
| comestai non-contact PPG app (2024-2025) | HR, SpO2, BP | App validation study | HR MAE 2.96 (99.1%), SpO2 MAE 2.10 (93.4%), SBP MAE 14.24 (61.3%) | Wellness monitoring cohort |
| VitalVideo / camera physiology (Wang et al., arXiv) | HR | 893 subjects, 6 Fitzpatrick skin tones | Large diverse benchmark for cross-skin-tone evaluation | In-the-wild and varied lighting |
What stands out is consistency at the top of the panel and divergence at the bottom. Heart rate errors cluster near or under 3 bpm almost everywhere. Blood pressure tells a different story: systolic mean absolute errors range from about 6 mmHg in a supervised study to over 14 mmHg in a wellness deployment, with diastolic accuracy below 60% in one case. Any benchmark that averages these together hides the signal researchers most need.
Why field numbers diverge from laboratory numbers
The gap between lab and field is not random noise. It traces to a small set of recurring factors that every deployment report should disclose.
- Illumination: end-to-end models tested by researchers studying extreme lighting found traditional rPPG methods degrade sharply outdoors, where uncontrolled sunlight dominates the signal.
- Skin tone: higher melanin absorbs more light and weakens the recovered pulse, a bias documented across smartwatch and camera studies. Diverse training sets such as VitalVideo (893 subjects, six skin tones) exist specifically to measure and reduce this.
- Motion: screening lines involve restless participants, children, and brief capture windows that introduce artifacts absent from seated lab trials.
- Hardware variance: the phones used in community programs are rarely the flagship devices used to validate algorithms.
What this means for benchmarking
A research team cannot fairly compare its own contactless screening accuracy field numbers to a published figure unless both report the same conditions. The most useful deployment papers now break results down by skin tone bin, by lighting category, and by reference device. When those breakdowns are missing, a single MAE figure should be treated as a ceiling, not an expectation.
Industry Applications
Population hypertension screening
Blood pressure is the vital with the highest public health payoff and the weakest contactless accuracy. The OptiBP work tested by Schoettker and colleagues, recognized by the World Health Organization as a promising low-resource technology, reported over 90% concordance against invasive measurement in a controlled cohort, but multi-country testing showed how much site conditions matter. Separately, a digitally enabled hypertension program across Ghana, Kenya, Sierra Leone, and Tanzania analyzed data from more than 63,000 patients between 2019 and 2024, demonstrating that the reach of phone-mediated screening can be enormous even when the measurement device itself is a cuff rather than a camera. The lesson for camera-based methods is that reach is solvable; accuracy at scale is the harder claim.
Community health worker triage
In community health worker programs, the relevant question is rarely whether a phone matches an arterial line. It is whether a 30-second capture can sort a screening line into routine and refer-now groups faster than a single shared cuff. For that triage use, heart rate and respiratory rate accuracy in the WellFie range are often sufficient, while blood pressure remains a flag rather than a diagnosis. This reframing changes how mobile health field trial outcomes should be judged: sensitivity for referral, not millimeter-level agreement, becomes the primary endpoint.
Antenatal and child health
Respiratory rate is a frontline signal for childhood pneumonia and a notoriously hard contactless target, with the WellFie study reporting a 15.66% error for RR against far tighter figures for heart rate. Programs that depend on RR need to validate it independently rather than assume the heart rate accuracy transfers.
Current research and evidence
The strongest recent evidence is the 2024 systematic review and meta-analysis of smartphone photoplethysmography, which found no statistically significant difference from reference methods for resting heart rate, with a pooled mean difference of -0.32 bpm. That settles the heart rate question for most purposes. The ReViSe framework reported heart rate MAEs of 1.73 to 3.95 bpm across datasets along with a 1.64% SpO2 error, while blood pressure errors of 6.7 mmHg systolic and 9.6 mmHg diastolic sat at the optimistic end of the field range. The 2024-2025 comestai validation, by contrast, reported strong heart rate and SpO2 figures but systolic blood pressure accuracy of only 61.3%, a reminder that BP performance is method-specific and not a property of cameras in general.
The fairness literature has matured in parallel. Work on diverse datasets and on filtering techniques such as the elliptic filter has shown measurable error reduction across skin tones, and a UCLA group has published a correction approach for melanin-related bias. For deployments across African populations, these fairness results are not optional context. They are core validity questions, because a method validated mostly on light-skinned cohorts cannot be assumed to perform identically in the field.
The Future of rPPG field deployment results
The next generation of deployment reporting will likely converge on a few shared practices. Stratified accuracy by skin tone and lighting will become an expected table rather than a supplementary note. Reference standards will be reported explicitly so that invasive, cuff, and ECG comparisons are never conflated. And reach metrics, the number of people actually screened and referred, will sit beside accuracy metrics, because a highly accurate method that nobody can deploy at scale serves fewer people than a good-enough method embedded in a working community program. Multimodal approaches that pair camera rPPG with other low-cost signals are also emerging as a route to better blood pressure and fairer cross-skin-tone performance. For researchers, the practical implication is that the most citable future studies will be those that publish their conditions as rigorously as their results.
Frequently asked questions
How accurate is rPPG for heart rate in field conditions?
Across the studies compared here, heart rate is the most reliable contactless vital, with mean absolute errors typically between 1.7 and 3 bpm and a 2024 meta-analysis finding no significant difference from reference methods. Outdoor lighting and motion can degrade this, so field reports should state capture conditions.
Why are blood pressure numbers so much worse than heart rate?
Blood pressure is inferred indirectly from pulse waveform features rather than measured directly, so it is far more sensitive to method, calibration, and population. Reported systolic errors in this comparison ranged from about 6 mmHg in supervised studies to over 14 mmHg in wellness deployments.
Does skin tone affect contactless screening accuracy in the field?
Yes. Higher melanin weakens the recovered pulse signal, a documented bias. Diverse datasets such as VitalVideo and correction methods from groups including a UCLA team aim to close the gap, which makes skin-tone-stratified reporting essential for African deployments.
What sample sizes make a field study credible?
There is no single threshold, but credible studies report enough participants to stratify by skin tone, lighting, and reference device. Reach-focused programs have screened tens of thousands of people, while accuracy validations often involve smaller, carefully referenced cohorts.
Circadify is working on this benchmarking gap directly, gathering real-world screening data from community health worker deployments so that accuracy and reach can be reported together rather than separately. Research teams and public health institutions interested in co-designing a field validation study, or in reading the underlying program data, can find our research write-ups and collaboration contacts at circadify.com/blog.
