scieee AI-readable full text Open interactive document viewer

DATA SET OSA

Universidad de La Sabana

Full text

CATEGORIES #NAME OF THE ARTICLE YEAR OF PUBLICATION AUTHORS MAGAZINE COUNTRY OR PLACE TYPE OF STUDY # POPULATION (SEX) # ARTICLES KEY RESULTS AND FINDINGS TYPE OF AI (Deep learning, machine learning, networks neuronal etc.) AI APPLICATION IN THE DIAGNOSIS OF APNEA ADVANTAGES OF APPLIED AI LIMITATIONS OF APPLIED AI FUTURE OPPORTUNITIES AND RECOMMENDATIONS ETHICAL CONSIDERATIONS REPORTED 11 Sleep detection apnea using deep neural networks and single-lead ECG signals 2021 Asghar Zarei, Hossein Beheshti, Babak Mohammadzadeh Asl Biomedical Signal Processing and Control Iran Study experimental Apnea-ECG dataset: 70 subjects (57 men, 13 women), ages 27–63 years, weight 53–135 kg - CNN-LSTM model trained on 1 min ECG segments. Deep Learning (deep neural network combining 2D-CNN + LSTM) Automatic classification of ECG segments as normal vs. apnea; use of the derived AHI to diagnose obstructive sleep apnea dream. It does not require manual feature extraction. Validated only on two public databases, no in prospective clinical population. Explore transfer learning to reduce training time and strengthen models. Use of public databases (Apnea-ECG and UCDDB). - Apnea-ECG dataset: Accuracy 97.21%, sensitivity 94.41%, specificity 98.94%, F1-score 0.96. Extremely high accuracy with single ECG data derivation. Possible confusion with arrhythmias (shared patterns in ECG). Test in the presence of arrhythmias and comorbid conditions.- Classification per patient (per-recording): 100% accuracy and MAE = 2.93 in estimated AHI index. Computationally efficient (processes segments in less than 0.07 s). Need to evaluate transfer learning for improve robustness. Validation in large clinical studies and multicenters.UCDDB dataset: 25 adult subjects suspected of apnea, ages 28–68 years, weight 59.8–128.6 kg - UCDDB dataset: Accuracy 93.70%, sensitivity 90.69%, specificity 95.82%. Potential for portable and home visits. Relatively small population sizes and controlled. Applications in mobile and portable home diagnostic systems.- The model outperformed other traditional methods, including CNN or LSTM networks alone. - Demonstrates that single-lead ECG can be a viable alternative to polysomnography for automated screening.K43 12 AI-Driven Detection of Obstructive Sleep Apnea Using DualBranch Convolutional Neural Networks 2025 Manjur Kolhar, Manahil Muhammad Alfridan, Rayan A Siraj MDPI Saudi Arabia Study of validation/development model roll Deep Learning ECG ataset Apnea-ECG: 70 recordings divided into training (35 records) and test (35 records) Both models, standard CNN and dual-branch CNN, achieved approximately 93-94% accuracy in validation and testing Deep Learning — specifically convolutional neural networks (CNN), including a dual-branch architecture (dualbranch) Binary classification (apnea vs no apnea) a starting from ECG segments, with models trained to detect apnea events Non-invasive Potential for non-invasive clinical use High accuracy / Excellent AUC Improve existing models compared with current advanced methods.The dual-branch CNN model achieved an AUC ROC of ~0.99, indicating excellent ability to distinguish between apnea and non-apnea cases. Potential to be used as a reliable method alternative to PSG They used class balancing techniques (SMOTE) to manage imbalance between apnea and non-apnea events Confusion matrices, ROC curves, and reports were generated. 13 Obstructive sleep apnea detection using ECG-sensor with convolutional neural networks 2020 Xiaowei Wang, Maowei Cheng, Yefu Wang, Shaohui Liu, Zhihong Tian, Feng Jiang, Hongjun Zhang Multimedia Tools and Applications China Study experimental 35 records noted (20 training, 10 test, 5 remaining no specified in detail) The CNN model uses RR-intervals as input features. Deep Learning – Convolutional Neural Network (CNN) with batch normalization, max pooling and softmax Automatic classification of single-lead ECG signals (RR intervals) in apnea vs normal. High accuracy (≈97.8%), 100% sensitivity and specificity 93%. Small sample size (35 records) noted). Extend to larger cohorts and various. Architecture: 3 convolutional layers (with batch normalization and max pooling) + 3 fully connected layers (100, 10, 2 neurons) + final softmax. It does not require multiple PSG signals, only ECG. Limited and public dataset; lacks validation prospective clinical practice. Integration into portable medical devices for home screening.Screening for obstructive sleep apnea using ECG recordings without the need for complete polysomnography Possibility of implementation on devices laptops and wearables. Patients with comorbidities were not evaluated. Validation in real clinical settings. Training with 50 epochs, batch size = 64, Adam optimizer. Reduces diagnostic costs and time Potential overfitting when using data augmentation 14 Automatic detection of obstructive sleep apnea through nonlinear dynamics of singlelead ECG signals 2024 Liangjie Chen, Fenglin Liu, Ying Wang, Qinghui Wang, Chengzhi Yuan, Wei Zeng Applied Intelligence China Study experimental with public database PhysioNet ApneaECG: 70 records (35 training, 35 test); 27–63 years; 30/5 men/women in training, 27/8 on trial; Accuracy 98.27%, Sensitivity 97.68%, Specificity 98.63%, MCC 0.963 (10fold CV). Outperformed SVM and other methods. Detected significant differences in ECG dynamics between normal and OSA. Machine learning:RBF neural networks with deterministic learning theory, using features of TQWT + VMD + PSR (3D phase space) ECG → TQWT (subbands) → VMD (4 main modes) → 3D phase space reconstruction → extraction of Euclidean distances → NN-RBF trained to model dynamics → bank of estimators for classification normal/OSA Extremely high accuracy, using only one ECG channel (plus inexpensive, less invasive), robust against noise, suitable for portable devices, efficient computationally Validation on a single dataset (PhysioNet), need to adjust parameters (Q, J, lag), limited to OSA (does not evaluate central/mixed apnea), limited physiological interpretability Validate in more diverse clinical cohorts, Try central/mixed apnea, adapt parameters in a customized way, combine with additional signals (EEG, SpO₂), explore transfer learning and multimodality Use of public data (PhysioNet Apnea-ECG). No additional ethical approval required. Authors declare no conflicts of interest; funding from the Natural Science Foundation of Fujian Province (China). 15 Detection of obstructive sleep apnea from singleECG channel signals using a CNN-transformer 2023 Hang Liu, Shaowei Cui, Xiaohui Zhao, Fengyu Cong Biomedical Signal Processing and Control China Study experimental with public database PhysioNet ApneaECG (70 records, 401–578 min each one, 100 Hz; notes by minute; ~34,313 With a 3-minute window: ACC 88.2%, SE 78.5%, SP 94.1%, F1 83.4%, AUC 0.947 in the test set; at the record level: 100% accuracy, MAE AHI 4.33. The Transformer outperformed CNN, LSTM, BiLSTM, and GRU. Using a 3-minute window improved performance compared to a 1-minute window. Deep Learning (hybrid 1D-CNN + Transformer encoder-decoder with self-attention) Raw ECG → CNN for representation → Transformer for modeling global temporal context → apnea/normal classification; calculation AHI per minute and per patient It does not require manual feature extraction; it captures global dependencies with Transformer. Improved performance with a 3-minute window, high accuracy per record, potential for home devices. It is an end-to-end model Moderate sensitivity (78.5%), limited base (PhysioNet only), class imbalance (apnea:normal ≈1:5), complex model (may hinder portability) Optimize loss for minorities (apnea), expand datasets with local hospitals, streamline model for wearables, explore multicenter validation Use of public database (PhysioNet Apnea-ECG); authors declare no conflicts of interest; code available on GitHub; no requirement additional ethical approval end, Translated from Spanish to English - www.onlinedoctranslator.com 16 Single-lead ECG based multiscale neural network for obstructive sleep apnea detection 2022 Zhiya Wang, Caijing Peng, Baozhu Li, Thomas Penzel, Ran Liu, Yuan Zhang, Xinge You Internet of Things (Elsevier) Multinational Study experimental development and validation Apnea-ECG dataset: • Accuracy = 90.4% • Sensitivity = 83.3% • Specificity = 94.8% • F1 = 89.6% Deep Learning (multiscale convolutional network URNet = integration) (of ResNet + U-Net) ECG segment classification of 1 and 5 minutes in apnea/non-apnea; indirect estimation of AHI; possibility of monitoring in real-time on portable devices Automatic and multi-scale extraction of relevant ECG characteristics. It does not distinguish between obstructive, central, and mixed. Develop models that extract features directly from ECG raw. Local study approved by the ethics committee of the Chongqing Ninth People's HospitalChina, Germany, Hong Kong Kong 62 patients in the local cohort + 70 recordings of public database Better performance compared to traditional models (CNN, LSTM, HMM). Only use RR intervals, not raw signals from ECG. Expand and diversify clinical bases for improve generalization. Local cohort (10-fold cross-validation): • Average accuracy = 79.5% • Sensitivity = 73.1% • Specificity = 81.5% • F1 = 76.8% High specificity (94.8%) in public database. Lower performance in the local clinical cohort (accuracy < 80%).49 men, 13 women aged 24-77 years Integrate into wearable devices for real-time home monitoring.Compatible with low-end embedded platforms consumption. Comparison with other methods (Table 5, p. 8): URNet outperformed previous traditional and deep learning models (accuracy ≤ 89%). Ablation study (p. 7): It was found that the combination of ResNet and Unet modules improves robustness and avoids saturation problems. Embedded implementation (NVIDIA Jetson Xavier NX): execution time 24.2 s and consumption 17 W, demonstrating viability in portable devices 17 Multi-task feature fusion network for Obstructive Sleep Apnea detection using single-lead ECG signal 2022 Keyan Cao, Xinyang Lv Measurement China Study experimental with databases public (ApneaECG, UCDDB) Apnea-ECG: 70 recordings (≈16,833 segments of training, 16,882 of validation; 6538 OSA, 10,295 normal) / UCDDB: 25 recordings (21 men, 4 In Apnea-ECG: ACC 91.58%, SE 87.29%, SP 94.26%, F1 88.86, AUC 0.9708. With Borderline-SMOTE: ACC 91.13%, SE 90.32%, SP 91.63%. Individual screening: ACC 100%, SE 100%, SP 100%, MAE 2.66, Corr 0.986. Outperformed CNN, CAE, and other reference models. Deep Learning (1D-MTFFNet: CNN with multi-task learning, supervised feature fusion and unsupervised, BorderlineSMOTE for balancing) ECG → RR interval extraction → preprocessing (FIR filter, Rpeaks detection, correction/interpolation) → BorderlineSMOTE for class balancing → 1D-MTFFNet (supervised and unsupervised tasks) → OSA vs normal classification + AHI estimation individual Combination of supervised and unsupervised learning → better feature extraction; Multi-level feature fusion; data balancing improved sensitivity; 100% individual accuracy; Robustness validated with cross-validation. Allows linking data from different channels and to achieve the fusion of multitasking and multilayer features, so that the network can extract more complete information about the network characteristics. Strong cross-validation but only lab data (PhysioNet and UCDDB), lower performance when applied to UCDDB (generalization) limited); relatively complex model → Difficult implementation in wearables. The proposed detection model's structure is still relatively complex, which hinders implementation. its implementation in portable devices. Reduce the complexity of the model for mobile applications, test in cohorts real-world clinical trials, extend to OSA severity, improve device integration laptops. No ethical risks are mentioned; use from public databases (PhysioNet, UCDDB); authors declare no conflicts of interest; funded by National Natural Science Foundation of China 18 Computer-aided obstructive sleep apnea screening from single-lead electrocardiogram using statistical and spectral features and bootstrap aggregating 2016 Ahnaf Rashik Hassan, Dr. Aynal Haque Biocybernetics and Biomedical Engineering Bangladesh Study experimental with database analysis data public and validation of a algorithm automatic Subjects: 35 patients, ages 27–63 years - An apnea detection scheme based on 1-minute ECG segmentation of one lead was proposed. Ensemble learning (Bagging with decision trees) Automatic classification of 1-minute ECG segments as normal or apneic High accuracy with only one ECG lead Validation in a small database (35 subjects) Clinical validation with more samples large and diverse Resistant to overtightening (Bagging) - ANOVA confirmed statistical significance (p < 0.05) in all characteristics. Main results: Accuracy: 85.97% Sensitivity: 84.14% Specificity: 86.83% AUC = 0.81 Aimed at early screening for obstructive sleep apnea using a single channel ECG It was not tested in de novo clinical cohorts with large size Inclusion of elderly patients for generalize resultsLow computational cost → useful in real time and on portable devices No performance report in the presence of comorbidities Potential integration into portable and home monitoring devicesIt does not require complex extraction of multiple signs Need for validation with older adults and different clinical contexts - The performance outperformed several state-of-the-art methods (KNN, SVM, wavelets, ANN). - Bagging was more robust than AdaBoost, Random Forest and other algorithms, showing low variance and resistance to overfitting. - Very low execution time (0.0013 s), making it feasible in real time. - Importance of features: variance, spectral flatness and 19 ApneaNet: A hybrid 1DCNNLSTM architecture for detection of Obstructive Sleep Apnea using digitized ECG 2023 Gaurav Srivastava, Aninditaa Chauhan, Nitigya Kargeti, Nitesh Pradhan, Vijaypal Singh Dhaka Biomedical Signal Processing and Control India Study experimental with public database PhysioNet ApneaECG database:70 records (35 training, 35 test). Includes Group A (apnea), B (borderline), C -Per-segment:Accuracy up to 90.87% (Modified AlexNet+LSTM), Sens 95.48%, Spec 83.43. -ApneaNet:Accuracy 90.13%, Sens 95.14%, Spec 82.06. - With split 28–7: Accuracy improved to 95.69% (AlexNet+LSTM) and 96.37% (ApneaNet). -Per-recording:Accuracy ≈ 97.1% in both models. Hybrid Deep Learning:1D-CNN + LSTM (modified AlexNet and ApneaNet (own). Raw ECG → R peak extraction → intervals RR → preprocessing (median filter + normalization) → 1D-CNN input → LSTM for temporal dependencies → classification apnea/normal. - High sensitivity (≥95%). - Fewer parameters: ApneaNet (0.9M) vs AlexNet original (60M). - Computationally efficient, viable in resource-constrained environments. - Limited dataset (PhysioNet only, 70 records). - External validation absent. - Possible overfitting with few subjects. - Validate in real clinical cohorts and multicentric. - Extend to multimodal signals (SpO₂, respiratory flow). - Optimize for real-time use and portable devices. Use of the public PhysioNet database (Apnea-ECG). Anonymized data, no additional ethical approval required. The authors declare no conflicts of interest. 20 Multimedia Monitoring System of Obstructive Sleep Apnea via a Deep Active Learning Model 2022 Fei Teng, Dian Wang, Yue Yuan, Haibo Zhang, Amit Kumar Singh, Zhihan Lv IEEE Multimedia International: China, New Zealand, India, Sweden Study experimental with public database + prototype in device PhysioNet ApneaECG:70 records (35 training, 35 test), 7–10 h each One. Signs noted by minute like apnea/normal. Per-segment:Accuracy 90.0%, Sens 88.7%, Spec 90.8%. Perrecord:Accuracy 100% (OSA diagnosis per subject). With active learning: Accuracy 92.15% using only 40% of the labeled data. It outperformed previous methods based on CNN and HMM. Deep Learning + Active Learning (lightweight 1D-CNN with 3 conv layers, pooling and softmax; query strategies: uncertainty, density weight, QBC, graph density, random). Raw ECG → R peak detection (Hamilton algorithm) → extraction of RR intervals and R amplitude → cubic interpolation at 900 points → 1D-CNN light → apnea/normal classification. Active learning selects the most valuable segments for labeling. reducing the cost of annotation. Lightweight, fast, suitable for mobile devices; requires less tagged data; high Accuracy per record (100%); applicable to home monitoring with portable sensor + smartphone. Single dataset (PhysioNet), not validated in real clinical cohorts; requires prior extraction of RR intervals (not fully end-to-end); performance depends on the query strategy; lacks validation under noisy conditions. Integrate multimodal data (SpO₂, audio); extend validation in hospitals; optimize data privacy and security; deploy in home applications for mass screening. Use of public data (PhysioNet). No ethical approval required. Funded by the Key Research and Development Program of Sichuan Province. 21 Sleep apnea detection from a single-lead ECG signal with automatic featureextraction through a modified LeNet5 convolutional neural network. 2019 Tao Wang, Changhua Lu, Guohao Shen, Feng Hong PeerJ China Development and validation of deep model learning with public data PhysioNet ApneaECG: 70 recordings of single ECG derivation (35 training, 35 test) Preprocessing: RR intervals and amplitudes were extracted from 5 consecutive 1-minute segments (±2 around the labeled segment). Cubic interpolation and filters were applied. (Figure 1, p. 5). Deep neural network convolutional (CNN) architecture modified LeNet-5 Automatic classification of minute-by-minute ECG segments as apnea or normal, and calculation of AHI per complete recording for diagnosis. It does not require manual feature engineering. Small databases (70 + 25 subjects). Validate with larger and more diverse databases. Improves metrics compared to SVM, KNN, LR and MLP. The notes do not distinguish between apnea and hypopnea, nor centrals. Include differentiation apnea/hypopnea/central. High accuracy per record (97.1%) and sensitivity 100% on PhysioNet.Model: Modified version of LeNet-5 with 1D convolutions, dropout 0.8, and only one fully connected layer to reduce complexity. (Figure 2, p. 6). Low sensitivity in UCD dataset per segment (26.6%). Extend to clinical and home settings with portable devices. UCD dataset: 25 patients (4 women, 21 men) suspects of Apnea, with PSG complete. Based on a single ECG lead → feasible in wearables or portable devices. Possible overfitting in small datasets. Comparison with traditional methods: CNN outperformed SVM, KNN, logistic regression, and MLP. (Table 3, p. 8). Per-segment (PhysioNet, test set): Accuracy 87.6%, Sensitivity 83.1%, Specificity 90.3%, AUC 0.950. Per-recording (PhysioNet): Accuracy 97.1%, Sensitivity 100%, Specificity 91.7%, AUC 0.996, Correlation with actual AHI 0.943. (Table 4, p. 10). Robustness: 10-fold cross-validation showed accuracies between 84.2% and 93.7% for CNN (mean 88.7% ±3.0%), superior to the other methods. (Figure 4, p. 11). Validation on UCD dataset: Per-segment: Accuracy 71.8%, Sensitivity 26.6%, Specificity 86.9%. Per-recording: Accuracy 92.3%, Sensitivity 90.9%, Specificity 100%, Correlation with actual AHI 0.624. (Table 5, p. 12). Comparison with previous literature: Their method (Acc 87.6%) 22 Adoption of Transformer Neural Network to Improve the Diagnostic Performance of 2023 Almarshad MA, AlAhmadi S, Islam MS, BaHammam AS, Soudani A Sensors Saudi Arabia Original study development and validation of deep model learning -AUC up to 0.90 and accuracy 0.82 in independent test set. • It outperforms previous methods based on SpO₂. • Annotations at 1-second resolution, allowing for detailed clinical interpretation. Deep Learning – Transformer Encoder with learnable positional encoding (based on convolutional autoencoder). Automatic classification of segments SpO₂ signal in normal vs. apnea 1-second resolution, It does not require complex filtering or preprocessing. • High temporal granularity (1 s) for better interpretation by doctors. • Outperforms state-of-the-art solutions (SOTA) in accuracy and AUC. • Validation only on the OSASUD dataset; needs evaluation on larger databases and various. • Requires clinical confirmation before use routine. • Validate in more diverse populations and devices spacious. • Combine with other sensors (EEG, ECG) for greater accuracy. • Optimize for device integration laptops and wearables. Publicly accessible data, without the need for ethics committee approval or informed consent; The authors declare that they have no conflicts of interest. 23 Semi-supervised Method with Adaptive Adjustment of Threshold for Detecting Obstructive Sleep Apnea Based on Oxygen Saturation 2024 Linqing Yang, Na Ying, Hongyu Li, Xinyu Lin, Yinfeng Fang, Yong Zhou, Huahua Chen Sensors and Materials China Study experimental with public databases (semi-supervised) either) PhysioNet ApneaECG: 70 recordings, but only 8 with SpO₂ signal (27–63 years, 53–135 kg, men and women). UCDDB (St. Vincent's/UCD): In UCDDB (1s):F1=90.94%, Prec=91.77%, Rec=90.13%, Acc=91.02%.In Apnea-ECG (1min):F1=94.65%, Prec=95.83%, Rec=93.50%, Acc=94.99%. Semi-DynaSeqNet outperformed CNN, LSTM, and CNN+LSTM, improving recall and robustness with pseudolabels. Hybrid Deep Learning (1D-CNN + GRU + Self-Attention + semisupervised thresholding) Raw SpO₂ → 1D-CNN (local features) → GRU (temporal dynamics) → Self-attention (selection of critical events) → semisupervised learning with pseudo-labels and dynamic threshold adjustment → apnea/normal classification UseSpO₂ only,cheap and portable; it combines Local extraction, dynamics and attention; semi-supervised → leverages unlabeled data; improves recall in minority classes High complexity (CNN+GRU+attention), computational demands; limited to 2 public datasets; not clinically validated in real time; requires optimization for implementation in wearables Extend validation to real clinical cohorts; optimize architecture for low power consumption; apply to other biomarkers; Explore integration in oximeters commercials Use of public data (PhysioNet, UCDDB). No additional ethical approval required. Authors declare financing of theZhejiang Provincial Natural Science Foundation;no conflicts of interest. 24 Discriminative methods based on sparse representations of pulse oximetry signals for sleep apnea-hypopnea detection 2017 RE Rolón, LD Larrateguy, LE Di Persia, RD Spies, HL Rufiner Biomedical Signal Processing and Control Argentina Study experimental with database 954 studies of Sleep dataset Heart Health Study (667 training, 287 test); adult population, diagnosis of The MDCS-OD method achieved AUC=0.937, sensitivity 85.65% and specificity 85.92% for AHI≥15. Better performance than previous reference methods. Artificial neural networks (MLP) + sparse coding Automatic detection of apneahypopnea events from pulse oximetry; diagnosisde novo Non-invasive, low cost, reduced size, good diagnostic performance, possibility of being embedded in pulse oximeters for primary care Oximetry signals affected by artifacts (movement, nails, anemia, pigmentation, etc.); higher computational cost in training; unbalanced database in mild cases Possible use in decentralized diagnostic units; support for screening in primary care; integration into oximeters; development of reference networks with specialists It does not report specific ethical considerations. 25 Apnea Detection in Polysomnographic Recordings Using Machine Learning Techniques 2021 Marek Piorecky, Martin Barton, Vlastimil Koudelka, Jitka Buskova, Jana Koprivova, Martin Brunovsky, Vaclava Piorecka Diagnostics Czech Republic Study experimental in laboratory dream 800 PSG initial → 255 records complete (flow nasal+SpO₂), 477 with just flow. Segments of training (351,550), CNN: Acc 84% (AUC 0.903) for apnea, 74% (AUC 0.826) for desaturation. k-NN: 83% and 64% respectively. CNN processed 8h in 36s vs 45min with k-NN. Sensitivity up to 96% with low threshold. Deep Learning (6-layer grid search optimized CNN; ReLU + sigmoid + dropout 0.6) Nasal flow and PSG SpO₂ → filtering/ normalization → 10s segmentation (90% overlap) → CNN (conv, pool, dense) → apnea/ normal classification and desaturation Based on a large PSG database (255 subjects), Rapid detection (seconds vs minutes), clinicianfriendly GUI, configurable with different thresholds (sensitivity vs specificity) Moderate accuracy compared to other studies (84%), overfitting observed → need for early stopping, hardware dependence with AVX, validated only in one center Explore optimization through transfer learning, validate in other cohorts, include more signals (ECG, EEG), improve feature interpretability, expand inter-subject generalization Approved by the Ethics Committee of the National Institute of Mental Health (Czech Republic, code 167/21); informed consent obtained; in accordance with the Declaration of Helsinki. Open source on GitHub. Authors declare no conflicts of interest. 26 OSASformer: A transformerbased model for OSAS screening via multi-source representation fusion 2025 Yuanyuan Hou, Bin Wang, Chengxi Zhang, Qiang Wang, Jiang Li, Pingping Meng, Yongxiang Zhang, Chao Han, Feng Hong, Tong Zhang Knowledge-Based Systems China Study experimental with data from devices portables (PPG) 60 volunteers, 46,081 records of PPG (SpO₂ and frequency cardiac) collected with bracelets intelligent; OSASformer achieved Acc = 0.824, F1 (Normal) = 0.922, F1 (Apnea) = 0.439, F1 (Hypopnea) = 0.477, Diagnosability = 0.878, Discrimination = 0.645. It outperformed CNN, LSTM, LightGBM, and other Transformers (Shapeformer, GTN). It won 2nd place in the JDHealth Global Medical AI Competition. Deep Learning (Transformer dualbranch with shapelet embedding + knowledge-based features + Balanced Batch Sampling) Classification of 60s segments (Normal, Apnea, Hypopnea) using PPG signals (SpO₂ + pulse) of wearables. They transformed the PPG wristband signals into combined representations (shapelets + clinical features) → they processed them with a Two-branch transformer → applied class balancing techniques → evaluated Non-invasive, wearable-based, superior accuracy to baselines, handles class imbalance with Balanced Batch Sampling, incorporates clinical knowledge and robust representation (shapelets + derived features) Low F1 in apnea (0.439), poor discrimination in hypopnea (only 46% correct), risk of confusion between classes; dataset limited to 60 subjects; external validation pending Authors declare no conflicts of interest; approved by the National Natural Science Foundation of China; JDHealth competitor data (2024); low PPG wristband usage consent Test in multicenter studies and with Various devices; compare complete nighttime diagnostics with PSG; explore real-time dynamic detection and application to personalized intervention (e.g. automatic CPAP adjustment) 27 SleepWatcher: Detecting sleep apnea/hypopnea syndrome from wearable devices using deep learning 2025 Hyungbin Kim, Hyubjin Lee, Minsoo Kim, Yon Dohn Chung Biomedical Signal Processing and Control South Korea Study experimental with public database (wearables simulated) UCD Sleep Apnea Database (PhysioNet): 25 subjects (21 men, 4 women, 28–68 years). They were used 20 for training and Stage 1 (normal vs. abnormal): HRV → Acc 89%, Spec 89%, Sen 89; SpO₂ → Acc 86%, Spec 84%, Sen 89; HRV+SpO₂ fusion → Acc 87%, Spec 83%, Sen 91. Stage 2 (apnea vs. hypopnea): Acc 97%, Spec 99% (hypopnea), Sen 76% (apnea). Correct prediction of AHI severity in all 6 test subjects (mild, moderate, severe). Generalization in Apnea-ECG (PhysioNet): HRV Spec 59%, Sen 70; SpO₂ Spec 79%, Sen 97. Deep Learning (2D CNN with images generated by recurrence plotsof HRV and SpO₂, merger at the decision level). Raw HRV and SpO₂ from PSG/wearables → uneven segmentation → transformation to images with recurrence plots → CNNs trained by signal → decision fusion → normal/abnormal classification and apnea/hypopnea + AHI calculation Using easily obtainable signals in wearables (SpO₂, HRV) avoids manual extraction of features. high sensitivity (91% in fusion, 97% in SpO₂ generalization), applicable to home screening Small dataset (20–25 subjects), lower sensitivity for apnea (76%), validation only in public databases, not yet tested on devices real, computational cost of the RP Expand to real clinical cohorts and databases recorded in seconds, optimized for portable hardware (smartwatches), integrated Additional signals (accelerometer, nasal flow), validate in multicenter studies Use of the UCDDB public database and PhysioNet Apnea-ECG. No original data collected. No conflicts of interest declared. Funded by the National Research Council. Foundation of Korea and MSIT. 28 ECG and SpO(2) Signal-Based RealTime Sleep Apnea Detection Using Feed-Forward Artificial Neural Network. 2022 Tanmoy Paul; Omiya Hassan; Khuder Alaboud; Humayara Islam; Md Kamruz Zaman Rana; Syed K. Islam; Abu SM Mosa AMIA Summits on Translational Science USA Study of development and validation of network model neuronal with database public 70 patients Segmentation in 1-minute intervals, using only the first 30 seconds to reduce noise and ambiguity. Artificial neural network feedforward (ANN) Classification of 30-second intervals as apnea or normal based on signals from SpO₂ and ECG It does not require manual feature extraction. Reduced number of records (only 8). Use larger datasets. Used in this Study: 8 records that included ECG and SpO₂ SpO₂ and ECG complement each other to improve consistency. Poor performance with ECG alone, attributed to errors in QRS annotations of the dataset PhysioNet. Design your own QRS detector plus accurate.Signal processing: • SpO₂: artifact removal (<50% or changes >4%), resampling at 1 Hz, 30 samples per segment. • ECG: RR interval extraction using QRS detector, ectopic elimination, 30 samples selected per segment. Simpler and lighter than CNN/RNN, with potential of embedded hardware implementation. Apply pruning and quantization techniques for implementation on low-end hardware consumption. Differences between validation and testing in the combined model. Model capable of working with 30-second segments → suitable for real-time detection. ANN architecture: 5 hidden layers (100, 50, 25, 10, 5 neurons); ReLU activation in hidden layers and Sigmoid activation in output (0 = normal, 1 = apnea). Training: 1000 epochs, 10-fold cross-validation, Adam optimizer, MSE loss function. Processed sequences: • SpO₂: 3,815 (2,293 normal, 1,522 apnea) • ECG: 2,606 (1,534 normal, 1,072 apnea) • Combined: 2,530 (1,512 normal, 1,018 apnea) Cross-validation: • SpO₂: 90.78% ±10.12 • ECG: 80.04% ±7.7 • Combined: 91.83% ±1.51 Independent test: 29 An Explainable Fusion of ECG and SpO₂-Based Models for RealTime Sleep Apnea Detection 2025 Tanmoy Paul, Omiya Hassan, Christina S. McCrae, Syed Kamrul Islam, Abu Saleh Mohammad Mosa MDPI USA Development of model of intelligence artificial explainable to detection in real-time of the apnea obstructive of the dream using signs of electrocardiogram ma (ECG) and saturation of oxygen in blood (SpO₂) Apnea-ECG Database: 32 subjects (25 men, 7 women) This study developed an explainable artificial intelligence model for the real-time detection of obstructive sleep apnea, using ECG and SpO₂ signals. Key findings include: Convolutional neural networks (CNNs) for real-time apnea detection using signals ECG and SpO₂. Development of an explainable AI model for real-time detection of obstructive sleep apnea, using signals from ECG and SpO₂ without the need for extraction features manual. Real-time detection: Ability to identify apnea events during sleep real time. The accuracy of the model may be affected due to the quality of the ECG and SpO₂ signals. Assessment in diverse populations: Conduct studies to evaluate the model's performance in diverse populations and clinical conditions.The model's ability to generalize to populations outside of those used in the training has not been evaluated thoroughly. It does not require manual feature extraction: The model directly processes the signals. raw, facilitating its implementation in portable devices. Development of models without manual feature extraction: ECG and SpO₂ signals were used directly without the need for manual preprocessing, which facilitates their implementation in portable devices. Integration with portable devices: Develop portable devices that integrate the monitoring model continuous sleep apnea. St. Vincent's University Hospital Database: 25 records complete of polysomnography night shift of 21 men and 4 women Explainability: The visual explanations provided by Grad-CAM improve the interpretability of the model. Visual explanations of model decisions: Grad-CAM was implemented to provide visual explanations that identify the signal regions that influence model decisions, improving interpretability. Improving model robustness: Implementing techniques to improve the robustness of the model against signals from low quality or noisy.Performance improvement: Merging individual models improves accuracy and other metrics. evaluation.Fusion of individual models: Combining individual models improved overall performance, achieving 95% accuracy, 94% sensitivity, 97% specificity, and an F1 score of 0.96. Applicability in diverse environments: It can be used in clinics, hospitals, and for monitoring. at home. Comparison with existing models: The proposed models outperformed previous approaches in most evaluation metrics, without requiring preprocessing steps. 30 Apnoea detection using ECG signal machine-based learning classifiers and its performances 2019 H. Nakano; T. Furukawa; T. Tanigawa Journal of Clinical Sleep Medicine Japan Primary school development and validation of algorithm (observational) retrospective) 1852 patients (80% males); average age 52 years; Average BMI 25.7 kg/m². Data from development: 1,548; validation: 304. Deep neural network with 3 convolutional layers and 2 fully connected layers. • High agreement between PSG AHI and the tracheal soundderived index (ICC 0.95). • Sensitivity / specificity for AHI > 5: 0.98/0.76 (AUC 0.99); for AHI > 15: 0.97/0.90 (AUC 0.99); for AHI > 30: 0.92/0.94 (AUC 0.98). • Sleep/wake detection accuracy: 0.88 (kappa 0.63). • Detected event type (central vs obstructive) with good agreement (kappa 0.70). Deep Learning: deep convolutional neural network (DNN with 3 convolutional layers + 2 layers) fully connected). Apnea/hypopnea assessment and sleep/ wake classification from sound recordings tracheal Non-invasive (small neck microphone.) • Calculate total sleep time and avoid underestimation of the IAH. • High concordance with PSG even in severe OSA and in patients with normal BMI. • Able to differentiate between central and obstructive apnea. Training and validation in a single center Japanese. • Population mainly with moderate to severe. • Possible influence of different microphones and ambient noise not extensively evaluated. • It did not include respiratory effortrelated awakenings (RERA). • Validate in other ethnicities and environments clinicians. • Test robustness against ambient noise and different recording devices. • Include RERA and more hypopnea events various. • Explore implementation on portable devices for home use Approved by the Research Ethics Committee of Fukuoka National Hospital (F30-10). Consent Exempt patient with voluntary opt-out option. Funded by the Japanese Society for the Promotion of Science (Grant 17K19928). Authors declared no have conflicts of interest. 31 Obstructive Sleep Apnea Detection Based on Sleep Sounds via Deep Learning 2022 Bochun Wang, Xianwen Tang, Hao Ai, Yanru Li, Wen Xu, Xingjun Wang, Demin Han Nature and Science of Sleep China Study cross 135 participants with snoring usual or sounds respiratory strong during the dream. To propose a novel deep learning method for the automatic detection of apneic events during sleep and to estimate the apnea-hypopnea index (AHI) based solely on sleep sounds obtained by a non-invasive audio recorder. Deep Convolutional Neural Network Neural Network - CNN) called OSAnet. Development of a deep learning method for the automatic detection of Apneic events during sleep and AHI estimation based on sleep sounds obtained by a non-audio recorder invasive. Non-invasive: Uses recorded sleep sounds without physical contact, improving comfort of the patient. Dependence on sound quality: The accuracy of the algorithm can be affected by the quality of the recorded sleep sounds Efficiency: The algorithm provides fast and accurate results for event detection apneic.Study phases: Phase 1: Eligible participants were randomly assigned to training (116 participants) and validation (19 participants) groups for the development of the deep learning algorithm. Phase 2: An independent test group of 59 participants was enrolled to evaluate the algorithm. Potential for home monitoring: The technology allows for the evaluation of sleep apnea sleep in non-clinical settings, facilitating remote monitoring. Results: Algorithm accuracy: 0.81 Sensitivity: 0.78 Specificity: 0.91 Area under the curve (AUC): 0.89 The algorithm showed robust diagnostic performance in identifying severe cases with a sensitivity of 95.6% and a specificity of 91.6%. 32 Screening for obstructive sleep apnea with novel hybrid acoustic smartphone app technology. 2020 Roxana Tiron; Graeme Lyon; Hannah Kilroy; Ahmed Osman; Nicola Kelly; Niall O'Mahony; Cesar Lopes; Sam Coffey; Stephen McMahon; Michael Wren; Kieran Conway; Niall Fox; John Costello; Redmond Shouldice; Katharina Lederer; Ingo Fietze; Thomas Penzel Journal of Thoracic Disease Germany Original study clinical 248 adults (age ≥21 years), is divided into two groups: 128 recordings for training, 120 for testing independent. Synchronized recordings of full PSG were collected along with the Firefly app in the sleep lab. - In the test set (120 recordings), for the clinical threshold AHI ≥15 events/h of sleep: sensitivity 88.3%, specificity 80.0%, AUC ROC of 0.92. - Correlation between the AHI estimated by Firefly and the AHI of the PSG: Pearson r = 0.90 in training, r = 0.81 in the independent test. - In the OSA severity classification (none, mild, moderate, severe), 61.7% of the predictions by Firefly matched the corresponding PSG class. - In cases where there were discrepancies, in 95% of them the difference was only one adjacent severity class; only 6 out of 120 presented an error greater than one class. - Number of false positives: Harm was observed due to periodic limb movements during sleep (PLMS ≥ 15 events/h) — 9 of the false positives (out of 26) were in that category, which is equivalent to 35% of the false positives. - False negatives: identified factors include OSA events (especially hypopneas) not showing typical expected modulation morphologies, physical interference to the equipment or phone (e.g., cover), movements of the mattress or blanket, etc. - The BMI was reasonably balanced between training and testing, and did not appear to be associated with many of the FP or FN errors. - The model can run on Apple and Android smartphones combination of models of machine learning traditional (logistics model, robust linear regression) + digital signal processing + convolutional neural networks (CNN) for detection of snoring and noise It is used for OSA screening: it classifies whether a An individual has AHI ≥15 events/hour (high risk) vs less than that; it also estimates the continuous AHI value for comparison with PSG. Breathing, breathing pauses, snoring, and respiratory effort are detected using passive acoustic signals and active sonar from a smartphone; use of analysis time-frequency envelopes, acoustic characteristics, etc. It does not require any additional hardware other than the user's smartphone. Study conducted in a sleep laboratory, not in completely free domestic environment — it which may affect generalizability. Use multi-night at home to reduce variability and improve accuracy. Informed consent of participants. Compatible with iOS and Android. Evaluations in real domestic environments, with variability in room, noise, telephone position, and covers. Studies conducted in accordance with the Declaration of Helsinki; compliance with privacy and data protection laws from German data (Bundesdatenschutzgesetz). Only one geographical location (Berlin/Germany) for the Collection using PSG; population possibly not representative of all ethnic profiles, environmental factors, habits, room noise, etc It allows potential monitoring in the home for several nights. Potential to integrate data with history clinic or treatments for follow-up (e.g., response to PAP). It can help reduce diagnostic barriers to make screening accessible. The authors declare conflicts of interest Interest: Some have patents related to breathing detection technology/disordered breathing breathing, sensors, etc. Some phone models were not all tested at PSG; hardware variability of microphone/speaker may be affected. It incorporates both active and passive signals for improve robustness. Expand the number of phone models and hardware for robustness against device variations. Good correlation and good balance between sensitivity/specificity for detection of moderate/severe OSA clinical threshold. False negatives are events that do not have the expected morphology, physical interference, or movement of the phone or cover. Clarification: during the performance of In that study, the software/ app was considered a consumer device and not a medical device, so at that time Approval by an independent committee was not considered necessary. as a medical device. There is not enough power to separately assess central apneas (or periodic breathing) given that less than 10% of subjects have central significant apneas. 33 Obstructive Sleep Apnea Screening Using a PiezoElectric Sensor. 2017 Urtnasan Erdenebayar, Jong-Uk Park, Pilsoo Jeong, Kyoung-Joung Read J Korean Med Sci South Korea Study experimental clinical with cohorts 45 patients with OSA (15 mild, 15 moderate, 15 severe; 34 men, 11 women; age average ≈ 55 years) Sensitivity/specificity/accuracy: 72.5/74.2/71.5% (mild), 85.8/80.5/80.0% (moderate), 70.3/77.1/71.9% (severe). Correlation between estimated AHI and PSG R²=0.94. Good stability in detecting snoring and heartbeats. Machine Learning (Support Vector Machine, kernel RBF) Automatic classification of OSA events vs. normal breathing using snoring index and pulse variability extracted from sensor piezoelectric in the neck Low-cost sensor, unaffected by ambient noise, allows recording of snoring and heartbeats simultaneously Discomfort from sensor on the neck, environment of laboratory (not natural sleep), exclusion of central/ mixed apnea and patients with arrhythmias, assumes all OSAs snore, shows small (n=45) Improve algorithm and sensor, expand population, include central/mixed OSA, evaluate in home conditions and in patients with cardiac comorbidities Approved by IRB of Samsung Medical Center (No. 2012-01-063); informed consent obtained; no conflicts of interest 34 Obstructive sleep apnea syndrome detection based on ballistocardiogram m via machine 2019 Weidong Gao, Yibin Xu, Shengshu Li, Yujun Fu, Dongyang Zheng, Yingjia She Mathematical Biosciences and Engineering China Study experimental, algorithm with signs physiological no invasive It does not specify. total size of cohort (HE worked with records of adult patients, With a single model: sensitivity 68%, specificity 73%, accuracy 71%. Withmodel fusion (decision trees + SVM + logistic regression):Sensitivity, specificity, and accuracy increased to 74%, 75% and 75%.Useful method for home screening of OSA. Machine Learning (decision tree + SVM + logistic regression; model fusion) Signal acquisition (BCG (heartbeats) + respiration + noise.) measured how heartbeat and respiration intervals change when there is apnea → with these variables they trained machine learning algorithms. They mainly used decision trees, Non-invasive, electrode-free, low-cost method, applicable at home, improves accuracy with fusion of models Moderate performance (sensitivity and specificity ≈ 74–75%), need to improve robustness to noise, lack of detail on size and diversity sample Explore deep learning, expand datasets, optimize algorithms for greater accuracy, and validate in clinical and home settings. Funded by National Key R&D Program of China; authors declare no conflicts of interest 35 Sleep Apnea Detection Using Multi-ErrorReduction Classification System with Multiple BioSignals 2022 Chung-Ho Su, Ying-Che Kuo, Chun-Wei Tsai, Kuei-Chien Chen Sensors Taiwan Study experimental / validation of deep algorithm learning 70 records nighttime, (35 training, 35 validation) A real-time system using an LSTM network was designed for apnea detection. Deep Learning — Long Short-Term Memory (LSTM) recurrent neural network. Automatic minute-by-minute classification in apnea vs normal using ECG signal. High accuracy (96%) on a standard public database. Validation only on public database (PhysioNet) Apnea-ECG). Validate in diverse clinical populations and big. Capable of processing data in real time. Performance: Overall accuracy: 96.04% Sensitivity: 96.09% Specificity: 96.00% AUC = 0.96 Focused on real-time monitoring and home screening. Limited number of patients (70). Integrate multimodal sensors (oximetry, respiration, accelerometry).Suitable for implementation on devices laptops or telemedicine systems. Not tested in real-world cohorts with comorbidities. Use transfer learning to improve generalization.Managing long-term temporary dependencies. It outperformed classic methods (SVM, Random Forest, CNN) applied to the same database. - The LSTM model was able to process data in real time with low latency, allowing practical use in portable devices. - The ability to manage long time dependencies was highlighted, crucial in apnea where events last ≥10 s. Discussion: - The system can be integrated into wearables with ECG sensors for home monitoring. - The model demonstrated robustness against signal noise. - Comparison with PSG: it does not replace it, but it is useful as a screening and monitoring tool. 36 Sleep Apnea Detection by Tracheal Motion and Sound, and Oximetry via Application of Deep Neural Networks. 2023 Nasim Montazeri Ghahjaverestan; Cristiano Aguiar; Richard Hummel; Xiaoshu Cao; jackson Yu; T. Douglas Bradley Nature and Science of Sleep Canada Original study clinical 233 participants; average age 50 ± 16 years old Correlation with PSG: AHI estimated by BresoDX1 correlated closely with PSG (r = 0.91, p < 0.001). Deep neural networks Automatic analysis of 61-second segments of tracheal-sternal movement, tracheal sound and oximetry for classification respiratory events, calculate AHI and diagnose apnea at different thresholds clinical Use of only two points of contact (tracheasuprasternal + finger with oximeter). It does not measure total sleep time (it uses valid recording time), which may underestimate or overestimate AHI. Validate performance in studies unassisted home care. All participants signed informed consent. Bland-Altman analysis: Mean difference AHI_Breso vs AHI_PSG = −3 events/hour; limits of agreement (95%): −12 to +15 events/ hour → demonstrates good clinical agreement. More stable tracheal sensor against displacement compared to straps. Further optimize the DNN architecture to improve AHI quantification, estimate sleep and wake time, and possibly identify sleep stages. Explicit exclusion of patients with significant comorbidities: neuromuscular, hypoventilation due to obesity, COPD, insufficiency cardiac, etc. Validation performed in a supervised laboratory, not in an unassisted home environment. High sensitivity and specificity at the threshold clinically most relevant (AHI ≥ 15).Diagnostic performance in test set (n = 159): • AHI threshold ≥ 5 (mild or major apnea): Accuracy = 84.3% • AHI threshold ≥ 15 (moderate or greater apnea): Sensitivity = 90.0%, Specificity = 85.9%, Overall Accuracy = 87.4% • AHI threshold ≥ 30 (severe apnea): Accuracy = 92.5% ROC curves: Very high AUC, which confirms robust accuracy of the model (specific values are not reported at all thresholds but are described as high). Importance of signals: The contribution of each type of signal was evaluated: • Using movement + sound + oximetry together produced the best performance. • Removing any of them significantly reduced accuracy. Common errors: Some false positives/negatives were due to differences in valid recording time (BresoDX1) vs total sleep time (PSG), which may alter the AHI estimate. 37 Obstructive Sleep Sleep Apnoea Syndrome Screening Through Wrist-Worn Smartbands: A Machine-Learning Approach 2022 Davide Benedetti; Umberto Olcese; Simone Bruno; Marta Barsotti; Michelangelo Maestri Tassoni; Enrica Bonanni; Gabriele Sicilian; Ugo Faraguna. Nature and Science of Sleep Italy Study observational of validation 78 patients (average age) 57.2 ± 12.9 years). Machine Learning Population screening: binary/graded classification according to AHI cut-offs using Data from wrist-worn smartbands (accelerometer + PPG → HR and parameters of sleep processed by DORMI). Objective data versus self-administered questionnaires reported. Relatively small cohort (n=78) and monocentric. Validation in larger populations and various. Approval by Pisa University Hospital Bioethical Committee (CEAVNO protocol No: 42714).Performance (tables 2 and 3): For AHI ≥5: MCC 0.39, Sensitivity 76.7%, Specificity 66.7%, PPV 88.5%, NPV 46.1%, DOR 6.57. For AHI ≥15: MCC 0.15, Sensitivity 73.3%, Specificity 41.7%, PPV 44%, NPV 71.4%, DOR 1.96. For AHI ≥30: MCC 0.33, Sensitivity 85.7%, Specificity 57.8%, PPV 30.8%, NPV 94.9%, DOR 8.22. Mild vs Moderate-Severe: MCC 0.28, Sensitivity 80%, Specificity 46.7%, PPV 60%, NPV 70%, DOR 3.5. Moderate vs Severe: MCC 0.55, Sensitivity 57.1%, Specificity 93.8%, PPV 88.9%, NPV 71.4%, DOR 20. Less invasive and costly; scalability possible through the use of commercial devices widely distributed. Participants only Caucasian → uncertainty in other skin tones (PPG dependent on tone). Integration of new sensors (SpO₂, blood pressure) when available in consumer devices. 30 women (per (48 men). Written informed consent of the participants. Severity assessment (multi-stage steps) to prioritize who requires study complete diagnosis (CRM/PSG). Higher specificity than STOP-Bang in this study; comparable PPV/NPV in several categories. Dependence on proprietary algorithms (Fitbit + DORMI) for staging and sleep parameters. Combination of physiological data with questionnaires to increase value diagnosis. Compliance with the Declaration of Helsinki. Lower performance compared to methods based on ECG or tracheal accelerometers; hypopneas are main source of disagreement. Classification errors: 14 patients were incorrectly classified in severity, 11 of them with a predominance of hypopneas, suggesting that these are the main source of disagreement. 38 Wearable Sensors and Artificial Intelligence for Sleep Apnea Detection: A Systematic Review 2025 Osa-Sanchez A, RamosMartinez-de-Soria J, Mendez-Zorrilla A, Ruiz IO, Garcia-Zapirain B Journal of Medical Systems Spain Revision systematic 28 studies among 2020 and 2024 - Growing trend in the use of wearables such as patches, watches, rings, etc., combined with convolutional neural networks (CNNs) and transfer learning - The most commonly used data types include ECG, PPG, SpO2, accelerometry, respiratory rate, etc. - Models combining deep neural networks (CNN, RNN/ LSTM) show very good results, provided the data is of good quality and quantity deep neural networks, networks convolutional, recurrent networks, learning classic automatic Detection and classification of apnea/hypopnea events; estimation of the apnea-hypopnea index hypopnea (AHI); differentiation of types of Sleep apnea; early detection using wearables; allows continuous monitoring, possibly at home instead of just in the office laboratory Variability in apnea definitions. Some wearable devices are accurate. limited in certain stages of sleep. Improve standardization in definitions of apnea and assessment protocols.It allows for more accessible diagnosis with portable/wearable devices. - Develop wearables with better sensors, greater comfort, and greater accuracy for different stages of sleep. Possibility of continuous and real-time monitoring real - Dependence on data quality and quantity; problems with artifacts, noise, poor signal. Improved accuracy when using multiple signals (multimodality) - Difficulty in generalizing if the training data does not cover diversity of populations/devices. - more transfer models and deep neural networks that combine multiple signals. Using deep learning models and transferring them improves performance if the data is adequate - more studies with diverse populations, commercial devices, scenarios real domestic use. 39 Comprehensive Evaluation of Machine Learning Techniques for Obstructive Sleep Apnea Detection 2024 Alaa Sheta; Walaa H. Elashmawi; Adel Djellal; Malik Braik; Salim Surani; Sultan Aljahdali; Shyam Subramanian; Parth S. Patel International Journal of Advanced Computer Science and Applications (IJACSA) International (States) United States, Egypt, Algeria) Primary school comparative/exp Experimental of ML models -- Best performance of the Random Forest (RFC) classifier: accuracy in training ~87%, in testing ~65%. ROC ~87%. - Gradient Boosting and SVM also showed good performance. Classic Machine Learning: Random Forest, Gradient Boosting, SVM, Logistic Regression, Gaussian Naive Bayes, K-Nearest Neighbors Prediction of the presence of obstructive sleep apnea using easily obtainable data (demographic, anthropometric) instead of complex physiological signals or PSG; It serves as a screening or pre-selection tool for further diagnostic studies invasive simple/easy to collect data; Low cost Accessibility and practicality Accuracy in the test set relatively low (~65%); risk of overfitting (trainingproof); Robustness is not reported across diverse populations; only demographic/physical variables (not signals) physiological factors) could limit accuracy OSA severity is not stratified Explore inclusion of physiological variables or signal to improve; test models in diverse populations; use more advanced techniques; improve the accuracy of the model under test; possibly combine ML with deep learning; externally validate the models Study funded by Taif University, Saudi Arabia. Data availability statements: the data that Supporting evidence for the findings is available upon request. Averages in training Accuracy: 74.1% Accuracy: 75.8% Recall (Sensitivity): 75.2% F1-score: 75.4% Test averages Accuracy: 65.5% Accuracy: 66.0% Recall (Sensitivity): 67.7% 40 Detecting obstructive sleep apnea by craniofacial imagebased deep learning 2022 Shuai He; Hang Su; Yanru Li; Wen Xu; Xingjun Wang; Demin They have Sleep and Breathing China Primary school development of model of intelligence artificial / observational / validation of algorithm 393 participants AHI ≥ 5 events/h: AUC 0.916, sensitivity 0.95, specificity 0.80; AHI ≥ 15 events/h: AUC 0.812, sensitivity 0.91, specificity 0.73 Averages in AHI ≥5 events/h Sensitivity: 92.8%; Specificity: 84.0%; Precision (PPV): 91.3%; F1-score: 93.3%; Accuracy: 90.0%; AUC: 0.916 Deep Learning — neural networks convolutional applied on craniofacial photographs Automatically detect the presence of OSA, using craniofacial images to estimate the probability of OSA without the need for PSG in all cases Non-invasive; use of relatively few photographs easy to administer; good sensitivity, especially at the low threshold (AHI≥5); potential application in clinical or community settings as a screening tool It can depend heavily on the quality and consistency of the images (angles, lighting, position). - Possible demographic or ethnic bias (Not necessarily validated in external populations with different characteristics) - The lower specificity at higher thresholds (AHI ≥15) suggests limitations for cases moderate/severe External validation in other ethnicities/regions; optimization for realworld image acquisition conditions; integration with other clinical signals or factors to improve accuracy; Explore use with common cameras (phones), automation of image flow; investigate variability due to sex, age, race, etc. institutional affiliations; consecutive recruitment; standard reference PSG; Averages in AHI ≥15 events/h Sensitivity: 90.9%; Specificity: 74.1%; Precision (PPV): 80.6%; F1-score: 85.4%; Accuracy: 82.7%; AUC 0.812 41 Artificial IntelligenceBased Diagnosis of Obstructive Sleep Apnea Syndrome: A Scoping Review 2024 Victor Ravelo; Jorge Fuentes; Marcelo Parra; Gonzalo Muñoz; Sergio Olate International Journal of Morphology Chili Scoping Review 13,293 subjects; age between 18 and 90 years; 9,912 males (≈ 74.56%), 3,381 women (≈ 25.43%) - 14 studies that diagnose OSA using polysomnography. • Average accuracy of around 80% in predicting OSA using variables such as age, gender, anthropometric and imaging measurements. • Many retrospective studies; one prospective study. • Better accuracy when combining image + anthropometric data. Average of metrics (only IAH ≥5, mild) Sensitivity (SE): 0.91 Specificity (SP): 0.94 Accuracy (ACC): 0.92 AUC: 0.95 PPV: 0.99 Predominantly classical machine learning; also some Studies with deep networks. Algorithms mentioned include Support Vector Machines, convolutional neural networks of the type ResNe Predictive assessment of OSA as a prediagnosis or screening, based on craniofacial characteristics and images (2D, 3D, CT, photographs) and measurements anthropometric, • less invasive alternatives lower costs than polysomnography. • accessible variables (facial image, physical measurements) that can facilitate early screening. • Good average accuracy (~80%) • Varied geographical coverage → promise of generalization if they develop well. Many retrospective studies. • Variability in image measurements, image acquisition positions, reference soft tissue vs bone. • Lack of uniformity in models, measurements, and algorithms across studies. • few women • Not all studies use control groups • There are no widely adopted prepolysomnography tests that serve as user guidance. • Incorporate more forward-looking designs robust. • Standardize image acquisition protocols and craniofacial measurements. • Include more diverse populations in terms of gender, ethnicity, and age. • Combine variables from various sources (image + anthropometry + others) biomarkers). • Develop more interpretable and practical models for clinical or community use. Study funded by the University of La Frontera; corresponding author declared; no conflicts of interest reported; ethics institutional and informed consent in the included studies (although as a review, it depends on what was reported in the studies primary) 42 Self-helped detection of obstructive sleep apnea based on automated facial recognition and machine learning 2023 Qi Chen; Zhe Liang; Qing Wang; Chenyao Ma; Yi Lei; John E. Sanderson; Xu Hu; Weihao Lin; Hu Liu; Fei Xie; Hongfeng Jiang; Fang Fang Sleep and Breathing (Springer) China Clinical study transversal with cohorts consecutive of patients subjected to polysomnography, complemented with analysis of photogrammetry facial automated Total: 653 subjects. Two photos of each patient (front and right profile) were captured with a mobile device camera. 68 facial points were automatically marked in each image to extract craniofacial measurements. Three predictive models (logistic regression, XGBoost and CatBoost) were evaluated by combining photographic and clinical data. Machine Learning Classification of subjects with OSA vs. non-OSA based on facial points obtained from photographs and simple clinical data. Non-invasive, economical and fast method. Suboptimal accuracy, especially in OSA mild. Expand studies to more populations diverse and balanced by sex. study approved by the Committee Ethics of Shanghai Jiao Tong University Affiliated Sixth People's Hospital. 55.3% diagnosed with apnea obstructive of the sleep (AHI ≥ 5). It can be applied with photos taken by the user. patient. The sample is predominantly male. (77.2%), which limits generalization. Integrate this methodology into platforms mobile units for community screening. Useful as a mass screening tool or self-diagnosis. All patients signed informed consent.Study conducted on a single population (China) Improve models for detecting mild cases more accuratelyThe CatBoost model was the best: • For moderate to severe OSA (AHI ≥ 15): – Sensitivity = 0.75 – Specificity = 0.66 – Accuracy = 0.71 – AUC = 0.76 • For mild OSA (AHI ≥ 5): lower performance (AUC ≈ 0.69). Sex: 77.2% men. Average age: 45 years. The photogrammetric model alone was less accurate than the combined model with BMI, STOP-Bang, NoSAS, and Epworth Sleepiness Scale. Conclusions: Automated photogrammetry combined with basic clinical practice can be a useful self-diagnostic method for 43 Detection of OSA Through the Application of Deep Learning on Polysomnography Data 2024 Hasan Ulutaş, Recep Sinan Arslan, Muhammet Emin Sahin, Halil Ibrahim Cosar, Cagri Arisoy, Ahmet Sertol Koksal, Mehmet Bakir, Bulent Ciftci Elektronika ir Elektrotechnika Türkiye Study experimental clinical 50 patients attended in the Sleep Laboratory, Dept. of Chest Diseases, Yozgat Bozok Univ.. Data: 19 sensors (7 EEG, 3 EMG, 1 ECG, 2 EOG, 1 SpO₂, 2 thoracic effort, 1 thermistor, 1 pressure, 1 body position). Validation and results They trained and validated with the 50 patients. Overall results: Average accuracy: 96.48%, Precision 97.69%, Recall 95.19%, F1 96.41%, AUC 0.9648. Best patient: Acc 98.35%; worst patient: Acc 93.20%. Demonstrated consistency across different profiles. Deep Learning (DNN feedforward with 4 hidden layers: 64–32–16–8 neurons, sigmoid output). Adam optimizer, learning rate 0.001, batch 50, 100 epochs, early stopping at 32. Raw multichannel PSG → preprocessing and standardization → undersampling for class balancing → DNN feedforward → apnea/hypopnea/normal classification Use ofown clinical dataset,large quantity of signals, high precision, relatively simple model, efficient computation time, applicable as clinical support Relatively small dataset (50 patients), risk of overfitting, not multicenter validated, expensive PSG signals (not yet applicable to wearables), interpretability limited Expand to multicenter cohorts, integrate Explainable AI for clinical trust, exploring channel reduction towards portable configurations, validation longitudinal Ethical approval obtained from the committee by Yozgat Bozok Univ.; informed consent of patients; no conflicts of interest; published under CC BY 4.0 license. In only 50 patients from a single center → the Results do not guarantee generalizability. The model requires full PSG (it is not a simple home screening). The authors themselves acknowledge the need for multicenter validation and validation in larger samples. before considering it a clinical diagnosis definitive. The model was able to detect apnea and hypopnea events from multichannel PSG signals, with performance comparable to human diagnosis, but with limitations, and is a support for diagnosis using PSG, not a replacement for it. 44 Yolo4Apnea: RealTime Detection of Obstructive Sleep Apnea 2020 Sondre Hamnvik; Pierre Bernabé; Sagar Sen Proceedings of the Twenty-nine International Joint Conference on Artificial Intelligence (IJCAI 2020), Demonstrations Track Norway Primary school development of system in real time Data from polysomnograms of 6,441 patients among 1995-1998 Data from polysomnograms 3295 patients between 2001-2003 Total: 9736 YOLO4Apnea can predict apnea events of varying size from abdominal breathing series images and manages to function at a frame rate. higher than 30 fps; with the aim of supplementing or assisting the manual work of sleep technicians - offers levels of confidence to help sleep technician in automatic event certification noted Deep Learning Automatic real-time detection of obstructive apnea events using signal of abdominal breathing, real-time detection Relatively simple sensor (abdominal respiration) capable of operating at variable intervals of apnea; reduces the burden of manual analysis in sleep laboratories can increase throughput/efficiency Accuracy not fully reported depends of the signal quality Extend validations across multiple datasets population Test with other types of signals, optimize for devices with resources limited; investigate explanation of decisions (interpretability); improve robustness against noise and signal artifacts - It might not generalize well to signals from other sources. different sensors or environments;use of YOLO-type convolutional neural network adapted to detect events as if were “objects” in series images temporary need to obtain images of the signal (data transformation) which may introduce latency or processing requirements 45 Real-time apneahypopnea event detection during sleep by convolutional neural networks 2018 Sang Ho Choi, Heenam Yoon, Hyun Seok Kim, Han Byul Kim, Hyun Bin Kwon, Sung Min Oh, Yu Jin Lee, Kwang Suk Park Computers in Biology and Medicine South Korea Study retrospective experimental with databases internal and external (MESA) 179 subjects (129 internal, 50 TABLE); ≥20 years; distributed: 38 no SAHS, 41 mild, 50 moderate, 50 Segment-by-segment: Kappa 0.82, sensitivity 81.1%, specificity 98.5%, accuracy 96.6%. Correlation between estimated AHI and PSG r=0.99, absolute error 3.1 events/h. Severity diagnosis (≥5, 15, 30 events/h): accuracy 94.9%, AUC 0.99 Deep Learning (CNN, 1D convolutional, 3 conv layers + pooling + FC) Automatic detection of AH events in real time from nasal pressure signal, with AHI estimate and classification of severity High accuracy, real-time detection, no requires manual features, generalizable model, reduces analysis time in laboratory It does not distinguish between obstructive, central, and mixed apnea, causes discomfort from the nasal cannula, has low sensitivity in hypopneas, and overestimates AHI, even validated offline Incorporate SpO₂ and EEG to improve and validate In clinical and home settings, extend to central/mixed apnea, use less invasive alternative sensors Approved by IRB of Seoul National University Hospital (No. H-1607-146-778); informed consent; no conflicts of interest 46 Gaussian Mixture Models for Detecting Sleep Apnea Events Using Single Oronasal Airflow Record 2020 Hisham ElMoaqet, Jungyoon Kim, Dawn Tilbury, Satya Krishna Ramachandran, Mutaz Ryalat, Chao-Hsien Chu Applied Sciences USA Study experimental with cohorts (PSG in laboratory dream) 96 patients adults: 10 without apnea, 36 mild, 27 moderate, 23 severe; total ≈ 48,320 events (15,300 OSA, 30,399 CSA, 2,621 MSA) Overall performance: TPR 88.5%, TNR 82.5%, AUC 86.7%. Improved accuracy compared to previous classifiers (neural network and threshold). GMM outperformed the threshold method in PPV (46.6% vs. 28.2%), F1 (61.1% vs. 42.7%), and ACC (83.4% vs. 80.2%). Good performance in OSA and CSA, lower performance in MSA and in mild cases. Overall performance (30 patients in test set): - TPR (sensitivity): 88.5% - TNR (specificity): 82.5% - ACC (accuracy): 83.4% - PPV (positive predictive value): 46.6% (vs 28.2% with thresholds) - F1: 61.1% (vs 42.7% with thresholds) - AUC: 86.7% Probabilistic Machine Learning (Gaussian Mixture Models, generative approach) Automatic identification of apnea/ normal eventsl in oronasal flow windows, calculating class probabilities and subsequent detection 1.Entrance→ They only took the oronasal airflow signal (thermal sensor) from the PSG. 2.Windows→ They divided the signal into two sliding windows: Baseline (600 s) for Define normal breathing. Detection (100 s) to search for events of apnea. 3. Feature Extraction→ they measured relative changes in: Intervals between breaths (Ic). Amplitude of respirations (Ac). 4. Probabilistic modeling→ They trained two Gaussian Mixture Models (GMMs): one for “apnea” and another for “normal”. Non-invasive algorithm, uses a single channel (simpler and cheaper than PSG), greater generalization Between severity levels, PPV is better than thresholds Limitations: lower performance in mild apnea and In MSA, false positives with noise and irregular breathing, sensitivity to class imbalance They recommend combining nasal signal + oronasal flow, adding respiratory effort to classify types of apnea, and improving filtration. noise and validation environments home delivery IRB Approval from the University of Michigan (IRB#HUM00069035); there was no external funding; authors declare no conflicts of interest interest The GMM allowed for the probabilistic modeling of respiratory changes (interval and amplitude) relative to a dynamic baseline, overcoming the limitations of fixed rules and improving generalizability across different types and severities of apnea. It was particularly valuable for mild and moderate patients. 47 Sleep detection apnea from singlechannel electroencephalon gram (EEG) using 2022 Noakes EJ, Moinuddin S, Sanei S, Wasmund T, Harrison SJ PLOS ONE New Zealand Study experimental development and validation of AI model Training: 2,650 patients (Sleep Health Heart Study); Validation: 25 CNN with 69.9% accuracy, MCC 0.375 and AUC 0.804 in cross validation; The model identified delta and beta bands as biomarkers; lower performance in external databases (MCC 0.216 and 0.169). Deep Learning – Convolutional Neural Network (CNN) Single-channel EEG training with 3 convolutional layers, optimized hyperparameters (Hyperband, Adam); validation 10-fold cross-tabulation and testing on datasets external. Low instrumentation (only one EEG channel), potential for home monitoring, model explainability (band analysis) (criticisms). Lower performance on external datasets due to Differences in annotation and segmentation; moderate accuracy; possible stage bias dream. Test on consumer-grade EEG equipment; improve interdataset; explore integration with portable devices. Use of de-identified public databases; ethics committee of the University of Auckland exempted additional approval. 48 Multiple-instance learning for EEG OSA-based event detection 2023 Liu Cheng, Shengqiong Luo, Baozhu Li, Ran Liu, Yuan Zhang, Haibo Zhang Biomedical Signal Processing and Control China Study experimental with public databases and hospitals UCDDB: 25 subjects (EEG C3A2, annotations events apnea). ISRUC: 61 subjects (100 PSG) night shifts, 61 with Mild OSA). Base local hospital: 35 subjects with OSA, apnea central or snoring Primary. Channels EEG C3-A2. In UCDDB: Accuracy 77.3%, Sensitivity 80.3%, Specificity 67.2%. In ISRUC: Accuracy 74.1%, Sensitivity 76.4%, Specificity 62.5%. In Local Hospital: Accuracy 78.6%, Sensitivity 81.3%, Specificity 72.3%. At instance-level (event localization): Accuracy 69–73%, IoU 50–61%. It outperformed DeepSleepNet, AttnSleepNet, and SVM/FCNN methods. Hybrid Deep Learning (Sub-frame Multi-Resolution CNN + Multiple Instance Learning with selfattention + Seq2Seq for event location) raw EEG → division into 60s frames → subframes → multi-resolution extraction (SMRCNN) → MIL mapping with self-attention → frame classification (OSA/normal) and event location (instance-level) → calculation From there It allowsdetect and locateOSA events (not only classify), avoids ambiguity of mixed frames, greater sensitivity and balanced specificity, superior performance to SOTA, applicable to EEG of PSG and ear-EEG Moderate accuracy (<80%), degradation in cross-validation between databases (dataset bias problem), requires specific EEG channel (C3-A2), high complexity Expand to multicenter cohorts, combine with ECG/SpO₂, develop lightweight models for wearables (ear-EEG), Apply domain adaptation and transfer learning Use of public datasets (UCDDB, ISRUC) and local hospital data with ethical approval; supported by the National Natural Science Foundation of China and Tsinghua Precision Medicine Foundation; no conflicts of interest declared 49 Diagnostic accuracy of machine learning and deep learning algorithms for detecting sleep apnea from singlelead ECG: a 2025 Mustafa Eray Kilic; Mehmet Emin Arayici; Oguzhan Ekrem Turan; Yigit Resit Yilancioglu; Emin Evren Ozcan; Mehmet Birhan Yilmaz Sleep Medicine Turkey Revision systematic and meta-analysis of precision diagnostic 84 studies (103 reports) published 20132024; bases of data included PhysioNet ApneaECG, TABLE, Cleveland Family Analysis by segment (1 min): sensitivity 90.3% (95% CI: 88.9–91.6), specificity 91.3% (89.8–92.6), DOR ≈ 98. • Analysis per record (patient): sensitivity 97.7% (95.3– 98.9), specificity 96.9% (93.1–98.7), DOR ≈ 1331. • DL outperformed ML (significantly higher AUC and accuracy, p < 0.05). • Spectrograms/scalograms and raw ECG data performed better than classical HRV metrics. Machine Learning (SVM, FFNN, Adaboost, etc.) and Deep Learning (LSTM, convolutional networks, MTFFNet, etc.). ML/DL algorithms analyze single-lead ECG signals to detect patterns of obstructive sleep apnea, with precision comparable to polysomnography. High sensitivity/specificity (>90%). • Single-lead ECG: inexpensive, portable, potential for home use and type devices smartwatch. • Possibility of massive and continuous screening. Methodological heterogeneity between studies (different AHI thresholds, class handling) unbalanced). • Limited prospective external validation and poor representation of subclinical cases. • ECG only provides indirect evidence of respiratory events and does not distinguish well between subtypes of apnea. Prospective multicenter trials with greater diversity of populations. • Standardize metrics and protocols of validation. • Improve interpretability with explainable AI (SHAP, LIME, Grad-CAM). • Explore integration with wearables and multimodal data fusion. Each original studio had its own own ethical guarantee. Authors report no conflicts of interest relevant interests 50 The Prediction of Obstructive Sleep Apnea Using Data Mining Approaches 2018 Zohreh Manoochehri; Mansour Rezaei; Nader Salari; Habibolah Khazaie; Behnam Khaledi Paveh; Sara Manoochehri Archives of Iranian Medicine Iran Primary school observational cross 333 patients The C5.0 decision tree model had better accuracy (≈ 0.757), sensitivity 0.66, and specificity 0.809. Logistic regression: accuracy ~0.737, sensitivity 0.693, specificity 0.78. Proportion of patients with OSA / without OSA = 208 vs 125 Overall averages of metrics (LRM + C5.0) - - Accuracy: (0.737 + 0.757) / 2 = 0.747 (74.7%) - - Sensitivity: (0.693 + 0.666) / 2 = 0.680 (68.0%) - - Specificity: (0.780 + 0.809) / 2 = 0.795 (79.5%) - - PPV: (0.756 + 0.667) / 2 = 0.712 (71.2%) Classic Machine Learning Predictive model to classify subjects with OSA vs without OSA using clinical data and questionnaires Simpler model, no need for PSG equipment; Use accessible data; can serve as a screening tool; relatively good accuracy and specificity; can prioritize patients for diagnosis complete Relatively small sample; somewhat low sensitivity in the decision model; possible biases if cynical data is incomplete; lack of generalization to other populations Applying the models to larger samples and diverse; improve sensitivity; validate externally; i integrate more predictor variables or sources of data; possible improvement of decision model for equity between accuracy and detection cases Approval of the Ethics Committee of the University of Medical Sciences of Kermanshah; written informed consent from all participants; funding from the same university; no conflict of interest declared interest 51 Efficient Deep Learning Based Hybrid Model to Detect Obstructive Sleep Apnea 2023 Prashant Hemrajani; Vijaypal Singh Dhaka; Geeta Rani; Praveen Shukla; Durga Prasad Bavirisetti Sensors International (India, Norway) Primary school development of Deep model Learning / experiment technical 70 patients (dataset ApneaECG of PhysioNet) MobileNet V1: accuracy ~ 89.5% MobileNet V1 + LSTM: accuracy ~ 90.0% MobileNet V1 + GRU: accuracy ~ 90.29%; specificity ~ 90.72%, sensitivity ~ 90.01% for classifying ECG segments as apneic or normal. A wearable device (“Sleepify”) was built that was capable of capturing ECG, classifying it, and securely sending data to the cloud. Deep Learning: hybrid models that combine a network lightweight convolutional (MobileNet V1) with recurrent networks (LSTM and GRU) Automatic detection/apnea vs non-apnea in short ECG segments (60 s); - High accuracy (~90%) using only ECG signal - Computationally efficient models (modified MobileNet) suitable for devices laptops - Wearable device designed for use domestic - Data transmission security - Improvement over previous methods using only ECG a channel - no severity classification or details of events - Relatively limited dataset - Validation using only the Apnea database - ECG; - no real-world testing with ECG signal collected by wearable under non-real-world conditions controlled - Possible overfitting despite cleaning and partitioning data - Validation in larger populations and various - Test under real-world usage conditions in home - Extend the models for classification - Improve robustness against noise and artifacts in ECG signals - Optimize energy efficiency and wearable comfort The authors declare that there are no conflict of interest There was no external financing. declared Not applicable (Institutional Review Board / Informed Consent) – “Not applicable” in those sections 52 Apnoea detection using ECG signal machine-based learning classifiers and its performances 2023 Rolant Gini J; Dhanalakshmi K Journal of Medical Engineering & Technology India Primary school development / comparative of models of machine learning 70 patients (dataset ApneaECG of PhysioNet) The K-NN classifier obtained the highest accuracy (~92.85%) compared to the other methods evaluated; It was shown that the ECG + ML processing approach can be reliable for apnea detection with a reduced number of features. Classic Machine Learning: SVM, Decision Tree, K-Nearest Neighbors (K-NN), Random Forest ECG-based apnea/non-apnea detection of a channel to facilitate less expensive diagnosis versus use of generalized polysomnography. High accuracy (~92.85%) using few features; use of ECG (very accessible signal); relatively simple ML methods of implement; better cost/convenience compared to more complex approaches. - - -