Auto Machine Learning for Diabetic Retinopathy Screening: A Head-to-Head Multiplatform Comparison Against Human Graders and IDx-DR
American Journal of Ophthalmology, vol.288, pp.267-283, 2026 (SCI-Expanded, Scopus)
- Publication Type: Article / Article
- Volume: 288
- Publication Date: 2026
- Doi Number: 10.1016/j.ajo.2026.04.030
- Journal Name: American Journal of Ophthalmology
- Journal Indexes: Science Citation Index Expanded (SCI-EXPANDED), Scopus, BIOSIS, CINAHL, EMBASE, MEDLINE, Health Research Premium Collection (ProQuest)
- Page Numbers: pp.267-283
- Dokuz Eylül University Affiliated: No
Abstract
Purpose: To benchmark multiple automated machine learning (AutoML) platforms for diabetic retinopathy (DR) screening from fundus photographs using a unified training and evaluationframework, with human consensus grading and an FDA-approved autonomous system (IDx-DR) as reference standards. Design: Retrospective, diagnostic performance and benchmarking study. Methods: Image classifiers were trained on large public datasets labeled according to the International Clinical Diabetic Retinopathy (ICDR) scale (APTOS, n = 5590; DDR, n = 12,524; EyePACS, n = 31,557) after automated image-quality filtering. Performance was evaluated on an independent, institutionally collected patient-level test cohort (n = 726) using the highest DR grade across all images for patient-level classification. The evaluated platforms included Google Vertex AI, Amazon Rekognition, Amazon SageMaker Canvas, AutoGluon, AutoKeras, and Apple CreateML. Models were assessed for 3 screening endpoints—any DR, referable DR (RDR), and sight-threatening DR (STDR)—across probability thresholds of 20%, 50%, and 70%. The primary endpoints were the Area Under the Receiver Operating Characteristic Curve (AUC) and sensitivity for RDR at a 50% decision threshold. Secondary endpoints included specificity, positive predictive value, negative predictive value, accuracy, and F1-score with 95% confidence intervals. Pairwise comparisons were performed using bootstrap testing (n = 1000) for AUC differences and McNemar's test (at a 50% threshold) for binary outcomes, both subjects to Bonferroni correction (P <. 0033). Grad-CAM was applied to locally deployable convolutional neural network–based models. Results: Using human consensus grading as the reference standard, Amazon SageMaker Canvas and AutoGluon demonstrated the strongest overall discrimination, achieving AUC values up to0.96 for STDR and 0.93 to 0.94 for RDR. At the 50% decision threshold, Canvas showed the most balanced performance for RDR (sensitivity 88.3%, specificity 85.5%, accuracy 86.0%), whereasAutoGluon favored sensitivity (any DR sensitivity 95.9%) at the expense of specificity. Vertex AI showed consistently weaker and unstable performance (any DR AUC 0.58; RDR AUC 0.38). Relative to IDx-DR, Amazon Rekognition and Canvas showed the highest agreement, particularly for STDR (AUC up to 0.88-0.90; κ up to ∼0.56). Agreement with human graders was generallylow to moderate (κ ≈ 0.3-0.6) and increased at higher probability thresholds. Conclusions: AutoML platforms can achieve clinically meaningful performance for DR screening. Differences across tools and thresholds reflect their adaptability to diverse clinical settings, underscoring the importance of external validation and threshold calibration.