scieee AI-readable full text Open interactive document viewer

Artificial intelligence for diagnosis and grading of prostate cancer in biopsies: a population-based, diagnostic study

Ström, Peter,Kartasalo, Kimmo,Olsson, Henrik,Ruusuvuori, Pekka,et al

Full text

1 Grading prostate biopsies with artificial 1 intelligence: a diagnostic study 2 Peter Ström*, Kimmo Kartasalo*, Henrik Olsson, Leslie Solorzano, Brett Delahunt, Daniel M Berney, David G 3 Bostwick, Andrew J. Evans , David J Grignon, Peter A Humphrey, Kenneth A Iczkowski, James G Kench, Glen 4 Kristiansen, Theodorus H van der Kwast, Katia RM Leite, Jesse K McKenney, Jon Oxley, Chin-Chen Pan, 5 Hemamali Samaratunga, John R Srigley, Hiroyuki Takahashi, Toyonori Tsuzuki, Murali Varma, Ming Zhou, Johan 6 Lindberg, Cecilia Lindskog, Pekka Ruusuvuori, Carolina Wählby, Henrik Grönberg, Mattias Rantalainen, Lars 7 Egevad, and Martin Eklund 8 9 * Both authors contributed equally to this study. 10 Corresponding author: Dr. Martin Eklund; Department of Medical Epidemiology and Biostatistics, Karolinska 11 Institutet, PO Box 281, SE-171 77 Stockholm, Sweden; [email protected]; +46 737121611 12 13 P Ström (MSc), Department of Medical Epidemiology and Biostatistics, Karolinska Institutet, Stockholm, Sweden. 14 K Kartasalo (MSc), Faculty of Medicine and Health Technology, Tampere University, Tampere, Finland. 15 H Olsson (MSc), Department of Medical Epidemiology and Biostatistics, Karolinska Institutet, Stockholm, 16 Sweden. 17 L Solorzano (MSc), Centre for Image Analysis, Dept. of Information Technology, Uppsala University, Uppsala, 18 Sweden. 19 B Delahunt (MD and Prof), Department of Pathology and Molecular Medicine, Wellington School of Medicine and 20 Health Sciences, University of Otago, Wellington, New Zealand. 21 DM Berney (MD and Prof), Barts Cancer Institute, Queen Mary University of London, London, UK. 22 DG Bostwick (MD and Prof), Bostwick Laboratories, Orlando, FL, USA. 23 AJ Evans (MD), Laboratory Medicine Program, University Health Network, Toronto General Hospital, Toronto, 24 ON, Canada. 25 DJ Grignon (MD and Prof), Department of Pathology and Laboratory Medicine, Indiana University School of 26 Medicine, Indianapolis, IN, USA. 27 PA Humphrey (MD and Prof), Department of Pathology, Yale University School of Medicine, New Haven, CT, 28 USA. 29 KA Iczkowski (MD and Prof), Department of Pathology, Medical College of Wisconsin, Milwaukee, WI, USA. 30 JG Kench (MD and Prof), Department of Tissue Pathology and Diagnostic Oncology, Royal Prince Alfred 31 Hospital and Central Clinical School, University of Sydney, Sydney, NSW, Australia. 32 G Kristiansen (MD and Prof), Institute of Pathology, University Hospital Bonn, Bonn, Germany. 33 TH van der Kwast (MD and Prof), Laboratory Medicine Program, University Health Network, Toronto General 34 Hospital, Toronto, ON, Canada. 35 KRM Leite (MD and Prof), Department of Urology, Laboratory of Medical Research, University of São Paulo 36 Medical School, São Paulo, Brazil. 37 JK McKenney (MD), Pathology and Laboratory Medicine Institute, Cleveland Clinic, Cleveland, OH, USA. 38 J Oxley (MD), Department of Cellular Pathology, Southmead Hospital, Bristol, UK. 39 C Pan (MD), Department of Pathology, Taipei Veterans General Hospital, Taipei, Taiwan. 40 H Samaratunga (MD and Prof), Aquesta Uropathology and University of Queensland, Brisbane, Qld, Australia. 41 JR Srigley (MD and Prof), Department of Laboratory Medicine and Pathobiology, University of Toronto, Toronto, 42 ON, Canada. 43 H Takahashi (MD), Department of Pathology, Jikei University School of Medicine, Tokyo, Japan. 44 T Tsuzuki (MD and Prof), Department of Surgical Pathology, School of Medicine, Aichi Medical University, 45 Nagakute, Japan. 46 M Varma (MD), Department of Cellular Pathology, University Hospital of Wales, Cardiff, UK. 47 M Zhou (MD and Prof), Department of Pathology, UT Southwestern Medical Center, Dallas, TX, USA. 48 J Lindberg (PhD), Department of Medical Epidemiology and Biostatistics, Karolinska Institutet, Stockholm, 49 Sweden. 50 C Lindskog (PhD), Department of Immunology, Genetics and Pathology, Uppsala University, Uppsala, Sweden. 51 P Ruusuvuori (PhD), Faculty of Medicine and Health Technology, Tampere University, Tampere, Finland. 52 C Wählby (PhD and Prof), Centre for Image Analysis, Dept. of Information Technology, Uppsala University, 53 Uppsala, Sweden; BioImage Informatics Facility of SciLifeLab, Uppsala, Sweden. 54 H Grönberg (MD and Prof), Department of Medical Epidemiology and Biostatistics, Karolinska Institutet, 55 Stockholm, Sweden; Department of Oncology, S:t Göran Hospital, Stockholm, Sweden. 56 This is the accepted manuscript of the article, which has been published in The Lancet Oncology. 2020, 21(2), 222-232. https://doi.org/10.1016/S1470-2045(19)30738-7 2 M Rantalainen (PhD), Department of Medical Epidemiology and Biostatistics, Karolinska Institutet, Stockholm, 1 Sweden. 2 L Egevad (MD and Prof), Department of Oncology and Pathology, Karolinska Institutet, Stockholm, Sweden. 3 M Eklund (PhD), Department of Medical Epidemiology and Biostatistics, Karolinska Institutet, Stockholm, 4 Sweden. 5 6 7 3 Abstract 1 Background: An increasing volume of prostate biopsies and a world-wide shortage of 2 urological pathologists puts a strain on pathology departments. Additionally, the high intra3 and inter-observer variability in grading can result in overand undertreatment of prostate 4 cancer. To alleviate these problems, we aimed to develop an artificial intelligence (AI) 5 system with clinically acceptable accuracy for prostate cancer detection, localization, and 6 Gleason grading. 7 8 Methods: We digitized 6,682 needle biopsies from 976 randomly selected participants aged 9 50-69 in the Swedish prospective and population based STHLM3 diagnostic study 10 conducted between May 28, 2012, and Dec 30, 2014 (ISRCTN84445406). The resulting 11 images were used to train deep neural networks for assessing prostate biopsies. The 12 networks were evaluated by predicting the presence, extent, and Gleason grade of 13 malignant tissue for an independent test set comprising 1,631 biopsies from 245 men as well 14 as an external validation set of 330 biopsies from 73 men. We additionally evaluated grading 15 performance on 87 biopsies individually graded by 23 experienced urological pathologists 16 from the International Society of Urological Pathology. We assessed discriminatory 17 performance by receiver operating characteristics (ROC) and tumor extent predictions by 18 correlating predicted millimeter cancer length against measurements by the reporting 19 pathologist. We quantified the concordance between grades assigned by the AI and the 20 expert urological pathologists using Cohen’s kappa. 21 22 Findings: The AI achieved an area under the ROC curve of 0·997 (0·994-0·999) for 23 distinguishing between benign (n=910) and malignant (n=721) biopsy cores on the 24 independent test set and 0·986 (0·972-0·996) on the external validation set (n benign=108; n 25 malignant=222). The correlation between millimeter cancer predicted by the AI and assigned 26 by the reporting pathologist was 0·96 (0·95-0·97) for the independent test set and 0·87 27 (0·84-0·90) for the external validation set. For assigning Gleason grades, the AI achieved an 28 average pairwise kappa of 0·62. This was within the range of the corresponding values for 29 the expert pathologists (0·60 to 0·73). 30 31 Interpretation: An AI can be trained to detect and grade cancer in prostate needle biopsy 32 samples at a level comparable to that of international experts in prostate pathology. Clinical 33 application will reduce pathology workload by culling of benign biopsies and by automating 34 the task of measuring cancer length in positive biopsy cores. An AI with expert level grading 35 4 performance may contribute a second opinion, aid in standardising grading, and provide 1 pathology expertise in parts of the world where it is currently non-existing. 2 3 4 Funding: Swedish Research Council, Swedish Cancer Society, Swedish Research Council 5 for Health, Working Life, and Welfare (FORTE), Swedish eScience Research Center, 6 Academy of Finland [313921], Cancer Society of Finland, Emil Aaltonen Foundation, Finnish 7 Foundation for Technology Promotion, Industrial Research Fund of Tampere University of 8 Technology, KAUTE Foundation, Orion Research Foundation, Svenska Tekniska 9 Vetenskapsakademien i Finland, Tampere University Foundation, Tampere University 10 graduate school, The Finnish Society of Information Technology and Electronics, TUT on 11 World Tour programme and the European Research Council (grant ERC-2015-CoG 12 682810). 13 14 5 Introduction 1 Histopathological evaluation of prostate biopsies is critical to the clinical management of men 2 suspected of having prostate cancer. Despite this importance, the histopathological 3 diagnosis of prostate cancer is associated with several challenges: 4 5 ●More than one million men undergo prostate biopsy in the United States annually.1 6 With the standard biopsy procedure resulting in 10-12 needle cores per patient, more 7 than 10 million tissue samples need to be examined by pathologists. The increasing 8 incidence of prostate cancer in an aging population means that the number of 9 biopsies is likely to further increase. 10 ● It is recognized that there is a shortage of pathologists internationally. In China, there 11 is only one pathologist per 130,000 population, while in many African countries the 12 ratio is of the order of one per million.2,3 Western countries are facing similar 13 problems, with an expected decline in the number of practicing pathologists due to 14 retirement.4 15 ● Gleason grade is the most important prognostic factor for prostate cancer and is 16 crucial for treatment decisions. Gleason grade is based on morphologic examination 17 and is recognized to be notoriously subjective. This is reflected in high intraand 18 inter-pathologist variability in reported grades, as well as both underand over19 diagnosis of prostate cancer.5,6 20 21 A possible solution to these challenges is the application of artificial intelligence (AI) to 22 prostate cancer histopathology. The development of an AI to identify benign biopsies with 23 high accuracy would decrease the workload of pathologists and allow them to focus on 24 difficult cases. Further, an accurate AI could assist the pathologist with the identification, 25 localization and grading of prostate cancer among those biopsies not culled in the initial 26 screening process, thus providing a safety net to protect against potential misclassification of 27 biopsies. AI-assisted pathology assessment could harmonize grading and reduce inter28 observer variability, leading to more consistent and reliable diagnoses and better treatment 29 decisions. 30 31 Using high resolution scanning, tissue samples can be digitized to whole slide images (WSI) 32 and utilized as input for the training of deep neural networks (DNN), an AI technique which 33 has been successful in many fields, including medical imaging.7–10 Despite the many 34 successes of AI, little work has been undertaken in prostate diagnostic histopathology.11–16 35 6 Attempts at grading prostate biopsies by DNNs have been limited to small datasets or 1 subsets of Gleason patterns, and they have lacked analyses of the clinical implications of the 2 introduction of AI-assisted prostate pathology. 3 4 In this study, we aimed to develop an AI with clinically acceptable accuracy for prostate 5 cancer detection, localization, and Gleason grading. To achieve this, we digitized 8,313 6 samples from 1,222 men included in the prospective and population based STHLM3 7 prostate cancer diagnostic study undertaken in 2012-2015.17,18 We evaluated the 8 performance of the model on an independent test set as well as an external validation set 9 (external lab and scanner), and through a comparison with 87 cases of prostate cancer 10 graded by the International Society of Urological Pathology (ISUP) Imagebase panel 11 consisting of 23 experienced urological pathologists.19 12 Methods 13 Study design and participants 14 Between May 28, 2012, and Dec 30, 2014, the prospective and population-based STHLM3 15 screening-by-invitation study (ISRCTN84445406) evaluated a diagnostic model for prostate 16 cancer in men aged between 50 and 69 years residing in Stockholm, Sweden.17,18 STHLM3 17 participants were biopsied if they had PSA ≥ 3 ng/mL or a Stockholm3 test ≥ 10%. Among 18 the 59,159 participants, 7,406 (12·5%) underwent systematic biopsy according to a 19 standardized protocol consisting of 10 or 12 needle cores; with 12 cores being taken from 20 prostates larger than 35 cm3 (Figure 1 and Table 1). Urologists who participated in the study 21 and the study pathologist were blinded to the clinical characteristics of the patients. A single 22 pathologist (L.E.) graded all biopsy cores according to the ISUP grading classification (where 23 Gleason scores 6, 3+4=7, 4+3=7, 8, and 9-10 are reported as ISUP grade 1 to 5, also 24 referred to as Gleason Grade Groups).20 L.E. also delineated cancerous areas using a 25 marker pen and measured the linear cancer extent. 26 27 The biopsy cores were formalin fixed and stained with hematoxylin and eosin. A random 28 selection stratified on ISUP grade of 8,313 biopsies from 1,222 STHLM3 participants was 29 digitized. The cases were chosen to represent the full range of diagnoses, with an over30 representation of high-grade disease. To further enrich the data with high-grade cases, 271 31 slides from 93 men with ISUP 4 and 5 prostate cancers were obtained from outside STHLM3 32 (Figure 1 and Appendix p 3). These slides were re-graded by L.E., digitized and utilized for 33 7 training purposes only. We used 1,631 cores from a random selection of 246 (20%) men to 1 evaluate the performance of the AI (the “independent test set”), while the rest were used for 2 model training. That is, all biopsies from a given man were assigned to either the training or 3 the test dataset.21 4 5 Since slides from different pathology labs differ in appearance and quality due to differences 6 in slide preparation and since WSI characteristics and appearance vary by scanner, it is 7 crucial to assess the performance of DNN models on external labs and scanners (i.e. 8 images of slides from different pathology labs and scanners than the images on which the 9 model was trained) from a real-world clinical setting. We therefore obtained 330 slides (73 10 men) from the Karolinska University Hospital and digitized them on the scanner available at 11 the Karolinska University Hospital pathology lab to replicate their entire workflow of lab 12 processing and slide digitization (the “external validation set”). The selection of slides was 13 enriched for higher ISUP grades to permit evaluation of predictions for these uncommon 14 grades (Table 1). L.E. graded all biopsies in the external test set to avoid confoundment 15 between introducing a different reporting pathologist and a different lab and scanner 16 workflow simultaneously. 17 18 As an additional test set, we digitized 87 cores from the Pathology Imagebase, a reference 19 database launched by ISUP to promote the standardization of reporting of urological 20 pathology.19 These cases were independently reviewed by 23 highly experienced urological 21 pathologists (The ISUP Imagebase panel). Cores from the men in the three test sets were 22 not part of model development and were excluded from any analysis until the final 23 evaluation. 24 25 The study protocol was approved by Stockholm regional ethics committee (permits 26 2012/572-31/1, 2012/438-31/3 and 2018/845–32). For details concerning data collection, 27 see Appendix p 3. 28 Test methods 29 We processed the WSIs with a segmentation algorithm based on Laplacian filtering to 30 identify the regions corresponding to tissue sections and annotations drawn adjacent to the 31 tissue. We then extracted digital pixel-wise annotations, indicating the locations of cancerous 32 tissue of any grade, by identifying the tissue region corresponding to each annotation. To 33 obtain training data representing the morphological characteristics of Gleason patterns 3, 4 34 and 5, we extracted numerous partially overlapping smaller images, or patches, from each 35 8 WSI. We used patch dimensions of 598 x 598 pixels (approx. 540 x 540 µm) at a resolution 1 corresponding to 10X magnification (pixel size approx. 0·90 µm). The process resulted in 2 approximately 5·1 million patches usable for training a DNN (Appendix Figure S1 p 23). 3 4 We used two convolutional DNN ensembles, each consisting of 30 Inception V3 models pre5 trained on ImageNet, with classification layers adapted to our outcome.22,23 The first 6 ensemble performed binary classification of image patches into benign or malignant, while 7 the second ensemble classified patches into Gleason patterns 3 to 5. To reduce label noise 8 in the latter case, we trained the ensemble on patches extracted from cores containing only 9 one Gleason pattern (i.e. cores with Gleason score 3+3, 4+4, or 5+5). Importantly, the test 10 data still contained cores of all grades to provide a real-world scenario for evaluation. Each 11 DNN in the first and the second ensemble thus predicted the probability of each patch being 12 malignant, and whether it represented Gleason pattern 3, 4, or 5, respectively (Appendix 13 Figure S2 p 24). 14 15 Once the probabilities for the Gleason pattern at each location of the biopsy core were 16 obtained from the DNN ensembles, we mapped them to core-specific characteristics (ISUP 17 grade and cancer length) using boosted trees.24 All cores in the training data were used for 18 training the boosted trees. Specifically, aggregated features from the patch-wise probabilities 19 predicted by each DNN for each core were used as input to the boosted trees, and the 20 clinical assessment of ISUP score and cancer length were used as outcomes. The ISUP 21 grade group was assigned based on a Bayesian decision rule of the core-level classifier to 22 obtain ISUP predictions at a clinically relevant operating point (Appendix p 13). 23 Statistical analysis 24 We summarized the operating characteristics of the AI system in a Receiver Operating 25 Characteristic (ROC) curve and the Area Under the ROC Curve (AUC), both on core-level 26 and patient-level. We then specified a range of acceptable sensitivities for potential clinical 27 use and evaluated achieved specificity when compared to the pathology report. The 28 enrichment of high-grade disease in the independent test data and the external validation 29 data may potentially inflate the estimated AUC values since high grades may be easier to 30 discriminate from benign cases compared to ISUP 1 and 2. Therefore, we also estimated the 31 AUC when ISUP 3 to 5 cases were removed from the independent test set and the external 32 validation. 33 34 9 We predicted cancer length in each core and compared it to the cancer length described in 1 the pathology report. The comparison was undertaken on individual cores as well as on 2 aggregated cores (i.e. total cancer length) for each man. Linear correlation was assessed on 3 both all cores and men, as well as restricted to positive cores and men. 4 5 Cohen’s kappa with linear weights was used for evaluating the AI’s performance against the 6 23 experienced urological pathologists on the Imagebase test set. Linear weights emphasize 7 a higher level of disagreement of ratings further away from each other on the ordinal ISUP 8 scale, in accordance with previous publications on the Imagebase study.19 Each of the 87 9 slides in Imagebase was graded by each of the 23 Imagebase panel pathologists, and 10 additionally by the AI. To evaluate how well the AI agreed with the pathologists, we 11 calculated all pair-wise kappas and summarized the average for each of the 23 raters. In 12 addition, we estimated the kappa with a grouping of the Gleason scores in ISUP grades 13 (grade groups) 1, 2-3 and 4-5. We further estimated Cohen’s kappa against the study 14 pathologist’s ISUP grading on the independent test set and the external validation set. For 15 the external validation set, we also estimated Cohen’s kappa after calibrating the 16 probabilities (i.e. scaling the ISUP probabilities before assigning the predicted class). 17 18 We used t-distributed stochastic neighbor embedding (t-SNE) and the deep Taylor 19 decomposition to interpret the representation of the image data learned by the DNN models 20 (Appendix p 17).25 21 22 All confidence intervals (CI) are two-sided with 95% confidence level and calculated from 23 1000 bootstrap samples. DNNs were implemented in Python 3·6·4 using TensorFlow 1·11, 24 and all boosted trees using the Python interface for XGBoost 0·72 (Appendix p 5). 25 26 Role of the funding source 27 The funders had no role in study design, data collection, analysis and interpretation, or 28 writing of the report. The corresponding author had full access to all the data in the study 29 and had final responsibility for the decision to submit for publication. 30 Results 31 We estimated the AUC representing the ability of the AI to distinguish malignant from benign 32 cores to 0·997 (0·994-0·999) for the independent test set and 0·986 (0·972-0·996) for the 33 16 Medical Imaging 2017: Digital Pathology. 2017: 101400S. 1 12 Kallen H, Molin J, Heyden A, Lundstrom C, Astrom K. Towards grading gleason score 2 using generically trained deep convolutional neural networks. Proc. - Int. Symp. 3 Biomed. Imaging. 2016; 2016-June: 1163–7. 4 13 Jiménez del Toro O, Atzori M, Otálora S, et al. Convolutional neural networks for an 5 automatic classification of prostate tissue slides with high-grade Gleason score. Med. 6 Imaging 2017 Digit. Pathol. 2017; 10140: 101400O. 7 14 Arvaniti E, Fricker KS, Moret M, et al. Automated Gleason grading of prostate cancer 8 tissue microarrays via deep learning. Sci Rep 2018; 8: 12054. 9 15 Litjens G, Sánchez CI, Timofeeva N, et al. Deep learning as a tool for increased 10 accuracy and efficiency of histopathological diagnosis. Sci Rep 2016; 6: 26286. 11 16 Campanella G, Hanna MG, Geneslaw L, et al. Clinical-grade computational pathology 12 using weakly supervised deep learning on whole slide images. Nat Med 2019; 13 published online July. DOI:10.1038/s41591-019-0508-1. 14 17 Grönberg H, Adolfsson J, Aly M, et al. Prostate cancer screening in men aged 50-69 15 years (STHLM3): A prospective population-based diagnostic study. Lancet Oncol 16 2015; 16: 1667–76. 17 18 Ström P, Nordström T, Aly M, Egevad L, Grönberg H, Eklund M. The Stockholm-3 18 Model for Prostate Cancer Detection: Algorithm Update, Biomarker Contribution, and 19 Reflex Test Potential. Eur Urol 2018; 74: 204–10. 20 19 Egevad L, Delahunt B, Berney DM, et al. Utility of Pathology Imagebase for 21 standardisation of prostate cancer grading. Histopathology 2018; 73: 8–18. 22 20 Epstein JI, Egevad L, Amin MB, Delahunt B, Srigley JR, Humphrey PA. The 2014 23 international society of urological pathology (ISUP) consensus conference on gleason 24 grading of prostatic carcinoma definition of grading patterns and proposal for a new 25 grading system. Am J Surg Pathol 2016; 40: 244–52. 26 21 Nir G, Karimi D, Goldenberg SL, et al. Comparison of Artificial Intelligence Techniques 27 to Evaluate Performance of a Classifier for Automatic Grading of Prostate Cancer 28 From Digitized Histopathologic Images. JAMA Netw open 2019; 2: e190442. 29 22 Szegedy C, Vanhoucke V, Ioffe S, Shlens J, Wojna Z. Rethinking the Inception 30 Architecture for Computer Vision. Proc. IEEE Comput. Soc. Conf. Comput. Vis. 31 Pattern Recognit. 2016; 2016-Decem: 2818–26. 32 23 Jia Deng, Wei Dong, Socher R, Li-Jia Li, Kai Li, Li Fei-Fei. ImageNet: A large-scale 33 hierarchical image database. 2009 IEEE Conf. Comput. Vis. Pattern Recognit. 2009; : 34 248–55. 35 24 Chen T, Guestrin C. XGBoost. Proc. 22nd ACM SIGKDD Int. Conf. Knowl. Discov. 36 Data Min. - KDD ’16. 2016; : 785–94. 37 17 25 van der Maaten L, Hinton GE. Visualizing High-Dimensional Data Using t-SNE. J 1 Mach Learn Res 2008; 9: 2579–605. 2 26 Nagpal K, Foote D, Liu Y, et al. Development and validation of a deep learning 3 algorithm for improving Gleason scoring of prostate cancer. npj Digit Med 2019. 4 DOI:10.1038/s41746-019-0112-2. 5 27 AI diagnostics need attention. Nature 2018; 555: 285. 6 28 Kweldam CF, Nieboer D, Algaba F, et al. Gleason grade 4 prostate adenocarcinoma 7 patterns: an interobserver agreement study among genitourinary pathologists. 8 Histopathology 2016; 69: 441–9. 9 29 Egevad L, Cheville J, Evans AJ, et al. Pathology Imagebase—a reference image 10 database for standardization of pathology. Histopathology 2017; 71: 677–85. 11 30 Goodfellow IJ, Shlens J, Szegedy C. Explaining and Harnessing Adversarial 12 Examples. CoRR 2014; abs/1412.6. 13 14 15          19 1 Table 1: Subject characteristics among all biopsied men in the STHLM3 study and among men 2 whose biopsies were digitized, tabulated by men (top) and by individual biopsy cores (bottom). No 3 cancer grade information is shown for Imagebase, as the grading of this set of samples was 4 performed independently by multiple observers. Imagebase cancer length was assessed by L.E. 5 20 1 2 Figure 2: ROC curves and AUC for cancer detection by individual cores (solid line) and by men 3 (dashed line) for the independent test set (top) and the external validation set (bottom). 4 5 6 7 Table 2: Sensitivity and specificity at selected points on the ROC curves for cancer detection. The 8 first two columns from left show the number of biopsy cores that could be discarded from further 9 consideration (specificity) and the number of biopsy cores that would need pathological evaluation 10 (sensitivity), respectively. The values in parentheses indicate the corresponding specificity and 11 sensitivity. The next five columns show the number and percentage of missed malignant cores by 12  !+-(G7CF9:CF957<CD9F5H=B;DC=BH,<9F=;<HACGH7C@IAB=B8=75H9GH<9BIA69F5B8D9F79BH5;9C: A=GG9875B79FG5ACB;5@@A9BK=H<75B79F     " ,)+75HH9FD@CHGDF9G9BH=B;H<97CB7CF85B7969HK99B75B79F@9B;H<G9GH=A5H986MH<9!5B8 H<9D5H<C@C;=GH:CF=B89D9B89BHH9GH85H5*9GI@HG5F9G<CKB:CF=B8=J=8I5@7CF9G$+5B85;;F9;5H98 CJ9F 7CF9G :CF 957< A5B )" !+ :CF H<9 =B89D9B89BH H9GH G9H +'( 5B8 9LH9FB5@ J5@=85H=CB G9H '++'% CFF9GDCB8=B; @=B95F 7CFF9@5H=CB 7C9::=7=9BHG 7CADIH98 :CF 5@@ 7CF9G 5B8 A5@=;B5BH 7CF9G CB@M5F9G<CKB=B957<D@CH5H5DC=BHG=BH<9@9:HD@CH5F9>=HH9F985@CB;H<9L5L=G:CF7@5F=HM 22 1 Figure 4: Grading performance on test data. (A) Cohen’s kappa for each pathologist ranked from 2 lowest to the highest. Each kappa value is the average pair-wise kappa for each of the pathologists 3 compared against the others. To account for the natural order of the ISUP scores we used linear 4 weights. The AI is highlighted with a black dot and an arrow. The study pathologist (L.E.) is 5 highlighted with an arrow. Values computed based on all five ISUP scores are plotted in red, while 6 values based on a grouping of ISUP scores commonly used for treatment decision are shown in blue. 7 (B) A confusion matrix on the independent test data of 1631 slides and (C) the external validation 8 data of 330 slides. (D) Results on external validation data are additionally shown following calibration 9 of the slide-level model. This procedure did not involve any model retraining. The pathologist’s (L.E.) 10 grading is shown on the y-axis and the AI’s grading on the x-axis. For the independent test set, 11 Cohen’s kappa with linear weights was 0·83 when considering all cases, and 0·70 when only 12 considering the cases indicated as positive by the pathologist. For the external validation set, the 13 corresponding values were 0·70 and 0·61. Following calibration, the kappa values increased to 0·76 14 and 0·66. The results are presented for an operating point achieving a minimum cancer detection 15 sensitivity of 99%. 16