Accuracy of Artificial Intelligence in Diagnosing Obstructive Sleep Apnea Using Photoplethysmography: Systematic Review and Meta-Analysis
Journal of Medical Internet Research ·
Background: The traditional method for diagnosing obstructive sleep apnea (OSA) through polysomnography may be expensive and inaccessible. Recent developments in AI propose the use of photoplethysmography (PPG) to aid OSA diagnosis. Objective: This study seeks to evaluate the diagnostic accuracy of an AI-based approach in diagnosing OSA using PPG. Methods: PubMed, Embase, Scopus, Web of Science, and IEEE Xplore were searched from inception to November 3, 2025. The inclusion criteria comprised observational studies that evaluated the accuracy of AI-based methods of OSA diagnosis using PPG compared with conventional sleep testing used in adults. We excluded case reports, case series, reviews, meta-analyses, letters, conference abstracts, pediatric studies, animal studies, foreign language studies, studies with incomplete data, and studies focusing on individual apneic events without patient-level classification. The outcome of interest was the diagnostic accuracy of PPG-based AI models for OSA, as assessed in primary studies using random split tests, cross-validation, or external validation. Independent reviewers extracted data and assessed risk of bias using the QUADAS-2 (Quality Assessment of Diagnostic Accuracy Studies-2) tool. A Bayesian bivariate meta-analysis was used to pool estimates. Further subgroup, sensitivity, and meta-regression analyses were conducted. Overall quality of evidence was assessed using the GRADE (Grading of Recommendations, Assessment, Development and Evaluations) framework. Results: From 12,579 records, we included 13 studies, comprising 9983 participants. All studies were rated as either low or unclear for risk of bias. The overall evidence quality was moderate. AI trained on PPG achieved a pooled sensitivity of 79.6% (95% credible interval [CrI] 55.5%‐93.8%) and specificity of 76.5% (95% CrI 48.2%‐94.0%), compared to conventional diagnosis. The overall specificity increased (apnea-hypopnea index [AHI] ≥5: 63.6%; AHI ≥15: 81.8%; AHI ≥30: 85.1%), but overall sensitivity decreased (AHI ≥5: 87.2%; AHI ≥15: 79.7%; AHI ≥30: 76.7%) with greater categorical AHI severity cutoffs. Additionally, deep learning models achieved a higher specificity (82.9%) than traditional machine learning (63.6%). OSA prevalence and device type were not clearly associated with sensitivity or specificity. Conclusions: AI models trained on PPG have reasonable accuracy and may potentially serve as a low-cost screening tool. However, limitations such as small sample sizes, potential sources of bias, underrepresentation of certain geographic regions, and the influence of potential confounding factors highlight the necessity for further research. Future work should focus on deep learning to improve the feasibility and accessibility of this approach in primary care. Trial Registration: PROSPERO CRD42024534235; https://www.crd.york.ac.uk/PROSPERO/view/CRD42024534235
Background: The traditional method for diagnosing obstructive sleep apnea (OSA) through polysomnography may be expensive and inaccessible. Recent developments in AI propose the use of photoplethysmography (PPG) to aid OSA diagnosis. Objective: This study seeks to evaluate the diagnostic accuracy of an AI-based approach in diagnosing OSA using PPG. Methods: PubMed, Embase, Scopus, Web of Science, and IEEE Xplore were searched from inception to November 3, 2025. The inclusion criteria comprised observational studies that evaluated the accuracy of AI-based methods of OSA diagnosis using PPG compared with conventional sleep testing used in adults. We excluded case reports, case series, reviews, meta-analyses, letters, conference abstracts, pediatric studies, animal studies, foreign language studies, studies with incomplete data, and studies focusing on individual apneic events without patient-level classification. The outcome of interest was the diagnostic accuracy of PPG-based AI models for OSA, as assessed in primary studies using random split tests, cross-validation, or external validation. Independent reviewers extracted data and assessed risk of bias using the QUADAS-2 (Quality Assessment of Diagnostic Accuracy Studies-2) tool. A Bayesian bivariate meta-analysis was used to pool estimates. Further subgroup, sensitivity, and meta-regression analyses were conducted. Overall quality of evidence was assessed using the GRADE (Grading of Recommendations, Assessment, Development and Evaluations) framework. Results: From 12,579 records, we included 13 studies, comprising 9983 participants. All studies were rated as either low or unclear for risk of bias. The overall evidence quality was moderate. AI trained on PPG achieved a pooled sensitivity of 79.6% (95% credible interval [CrI] 55.5%‐93.8%) and specificity of 76.5% (95% CrI 48.2%‐94.0%), compared to conventional diagnosis. The overall specificity increased (apnea-hypopnea index [AHI] ≥5: 63.6%; AHI ≥15: 81.8%; AHI ≥30: 85.1%), but overall sensitivity decreased (AHI ≥5: 87.2%; AHI ≥15: 79.7%; AHI ≥30: 76.7%) with greater categorical AHI severity cutoffs. Additionally, deep learning models achieved a higher specificity (82.9%) than traditional machine learning (63.6%). OSA prevalence and device type were not clearly associated with sensitivity or specificity. Conclusions: AI models trained on PPG have reasonable accuracy and may potentially serve as a low-cost screening tool. However, limitations such as small sample sizes, potential sources of bias, underrepresentation of certain geographic regions, and the influence of potential confounding factors highlight the necessity for further research. Future work should focus on deep learning to improve the feasibility and accessibility of this approach in primary care. Trial Registration: PROSPERO CRD42024534235; https://www.crd.york.ac.uk/PROSPERO/view/CRD42024534235