scieee AI-readable full text Open interactive document viewer

Synthetic MRI improves radiomics‐based glioblastoma survival prediction

Moya Saez, Elisa,Navarro González, Rafael,Cepeda, Santiago,Pérez Núñez, Ángel,Luis García, Rodrigo de,Aja Fernández, Santiago,Alberola López, Carlos

Abstract

Producción Científica

Full text

RESEARCH ARTICLE Synthetic MRI improves radiomics-based glioblastoma survival prediction Elisa Moya-Sáez 1 | Rafael Navarro-González 1 | Santiago Cepeda 2 | ´ Angel Pérez-Núñez 3 | Rodrigo de Luis-García 1 | Santiago Aja-Fernández 1 | Carlos Alberola-L opez 1 1 Laboratorio de Procesado de Imagen, Universidad de Valladolid, Valladolid 2 Departamento de Neurocirugía, Hospital Universitario Río Hortega, Valladolid, Spain 3 Departamento de Neurocirugía, Hospital Universitario 12 de Octubre, Madrid, Spain Correspondence Elisa Moya-Sáez and Rafael Navarro-González, MSc, Laboratorio de Procesado de Imagen, Universidad de Valladolid, Campus Miguel Delibes, Paseo de Belén, 15, 47011 Valladolid, Spain. Email: [email protected]s;[email protected]. uva.es Funding information Fundaci on Científica Asociaci on Española Contra el Cáncer; Ministerio de Ciencia e Innovaci on of Spain, Grant/Award Numbers: RTI 2018-094569-B-I00, PRE2019-089176, PID2020-115339RB-I00 Glioblastoma is an aggressive and fast-growing brain tumor with poor prognosis. Predicting the expected survival of patients with glioblastoma is a key task for efficient treatment and surgery planning. Survival predictions could be enhanced by means of a radiomic system. However, these systems demand high numbers of multicontrast images, the acquisitions of which are time consuming, giving rise to patient discomfort and low healthcare system efficiency. Synthetic MRI could favor deployment of radiomic systems in the clinic by allowing practitioners not only to reduce acquisition time, but also to retrospectively complete databases or to replace artifacted images. In this work we analyze the replacement of an actually acquired MR weighted image by a synthesized version to predict survival of glioblastoma patients with a radiomic system. Each synthesized version was realistically generated from two acquired images with a deep learning synthetic MRI approach based on a convolutional neural network. Specifically, two weighted images were considered for the replacement one at a time, a T2w and a FLAIR, which were synthesized from the pairs T1w and FLAIR, and T1w and T2w, respectively. Furthermore, a radiomic system for survival prediction, which can classify patients into two groups (survival >480 days and ≤480 days), was built. Results show that the radiomic system fed with the synthesized image achieves similar performance compared with using the acquired one, and better performance than a model that does not include this image. Hence, our results confirm that synthetic MRI does add to glioblastoma survival prediction within a radiomics-based approach. KEYWORDS glioblastoma, synthetic MRI, radiomics, survival prediction Abbreviations: AUC, area under the curve; BraTS2020, multimodal brain tumor segmentation; CNN, convolutional neural network; ED, edema; ET, enhancing tumor; ICC, intra-class correlation coefficient; IQR, interquartile range; IR-SE, inversion recovery spin echo; KPS, Karnofsky performance status; LR, logistic regression; MAE, mean absolute error; mRMR, maximum relevance– minimum redundancy; MSE, mean squared error; NET, non-enhancing tumor; PD, proton density; PET, positron emission tomography; PSNR, peak signal to noise ratio; ROI, region of interest; RS, radiomic system; SE, spin echo; SP, survival prediction; SSIM, structural similarity index; SVM, support vector machine; TC, tumor core; TCIA, The Cancer Imaging Archive; T1w, T1weighted; T1w-c, T1w post-contrast; T2w, T2weighted; WT, whole tumor; XGB, extreme gradient boosting. Elisa Moya-Sáez and Rafael Navarro-González contributed equally to this work. Received: 24 February 2022 Revised: 26 April 2022 Accepted: 26 April 2022 DOI: 10.1002/nbm.4754 This is an open access article under the terms of the Creative Commons Attribution License, which permits use, distribution and reproduction in any medium, provided the original work is properly cited. © 2022 The Authors. NMR in Biomedicine published by John Wiley & Sons Ltd. NMR in Biomedicine. 2022;e4754. wileyonlinelibrary.com/journal/nbm 1of14 https://doi.org/10.1002/nbm.4754 1|INTRODUCTION Glioblastoma is an aggressive and fast-growing brain tumor, which comprises the most common kind of malignant brain tumor. 1 Recently, several advances have been achieved in precision oncology and immunotherapy. 2 However, the overall survival remains poor, with approximately 40% survival in the first year after diagnosis and 17% in the second year. 3 Thus, survival prediction (SP) in glioblastoma prognosis is a key task for efficient treatment and surgery planning. Factors commonly used for SPs are age, sex, extent of resection, or Karnofsky performance status (KPS) 4 —a functional status metric for activities of daily living. 5 Although these factors have shown ability to predict survival rates, their results are still poor and nowadays there is no clinical standard based on these metrics. 6 Moreover, some of these factors are subjective (e.g., KPS), so clinicians are forced to rely on their previous experience and intuition. Therefore, a quantitative predictive tool is still an unmet need. Radiomics and medical images are two major actors to achieve such a quantitative predictive tool. Radiomics 7 is a discipline that consists in the extraction of a large number of quantitative features from the images, the selection of a subset of them according to some quality criteria, and the design of an inference engine that will carry out the clinical predictions with the remaining features as inputs. Radiomics has been an useful technique, pioneering in the oncology field, for quantitative decision making in diagnosis, prognosis, and therapeutic response. 8 In particular, radiomic systems (RSs) for SP in glioblastoma could enhance patient management by treatment personalization, giving rise to better outcomes. 9 Regarding the second actor, MRI is a widely used medical imaging technique in different fields—and, particularly, in neuro-oncology—for noninvasive diagnosis and evaluation of disease progression. An MRI scan protocol typically consists of several pulse sequences that provide images with different contrasts, each of which is referred to as a weighted image, and they are collectively referred to as multicontrast images. Each image weighting provides complementary information for diagnosis, since different tissue properties may be more clearly visible in each of them. 10 The use of multicontrast images is also crucial for radiomics. However, multicontrast acquisitions are time consuming, which results in patient discomfort and artifact-prone protocols. For protocol shortening, the emerging field of synthetic MRI lends itself to become a cornerstone. Synthetic MRI comprises methodologies pursued to computationally synthesize realistic MR images from a set of actually acquired images. This discipline has been boosted by deep learning techniques. Synthetic MRI has many potential applications as for image database management, such as retrospective completion of databases with missing weighted images or replacement of artifacted images. Additional applications such as data harmonization 11 or data augmentation 12,13 for segmentation or classification algorithms have recently been proposed. As for multicontrast acquisitions, synthetic MRI may be applied to replace some acquisitions with their synthesized versions, leading to MR protocol shortening, patient wellbeing and higher efficiency. To make this possible, synthesized MR images should be of sufficient quality, and this is usually assessed both visually and with specific image quality assessment metrics. However, it is also essential to measure how these synthesized images will perform with quantitative algorithms. 14 In this work we propose and quantify the application of synthetic MRI to improve a radiomics approach for SP in glioblastoma; both the RS and the image synthesis method are original and the details of their design are fully described. Our purpose is to show that an RS that incorporates an input channel fed by a synthesized image (a) behaves similarly to this system when it is fed with an acquired image and (b) undoubtedly outperforms an RS that does not have this channel. Hence, we validate an MR protocol shortening procedure by means of a glioblastoma SP radiomics-based application. Two weighted images are considered for the synthesis, namely, FLAIR and T2weighted (T2w). We synthesize these images by means of an improved version of our previous deep learning approach for relaxometry maps synthesis. 15 Our results allow us to state that synthetic MRI does add to glioblastoma SP within a radiomics-based approach. 1.1 |Related work The quality assessment of synthesized images generally focuses on their usage in qualitative applications, and it remains to be verified whether quantitative algorithms can reliably work with these synthesized images. In particular, to the best of our knowledge, the performance of synthesized images with a clinical endpoint in radiomics applications has not been thoroughly tested. Recent works make use of synthesized images to improve subsequent quantitative image analysis algorithms. In Reference 16 an MRI synthesis method was presented, and the quality of a tumor segmentation algorithm was used as a benchmark to compare different synthesis procedures. Furthermore, in Reference 17 a generative adversarial network was trained for MRI synthesis and was then used for data augmentation in Parkinson's disease classification. The usage of synthesized data showed an improvement in the classification performance. Moreover, Sikka et al 18 proposed the synthesis of positron emission tomography (PET) images from MRI. The synthesized PET images were then included in an Alzheimer's disease classification task, for which a relevant accuracy improvement was also shown. Regarding the application of synthesized images in glioblastoma, we are aware of two contributions 11,19 that have used synthesized images as input. In Reference 11 image synthesis is used to fill missing contrasts in MRI glioma databases. The overall performance is measured by means of tumor segmentation accuracy, as well as by the merit figures of two RSs, one for tumor grading and the other one for isocitrate 2of14 MOYA-S ´ AEZ ET AL. dehydrogenase-1 (IDH1) status prediction. Completion of the databases with synthesized images produced an improvement in the results of these tasks. However, synthesized images are apparently used for training and testing of the RSs, which implies a coupling between these systems and the corresponding synthesis method. Islam et al 19 proposed a radiogenomics system for overall SP in glioblastoma using an MRI synthesis method to complete the missing images in the dataset. Tumor segmentation and overall SP were tested with the synthesized data. Nevertheless, the classifiers themselves were trained on the synthesized images; hence, this methodology is a data augmentation procedure and benefits in SP may arise not only from the quality of the synthesized images, but also from the fact that more data are used, as pursued in data augmentation. Hence, this methodology does not provide sufficient evidence that acquired images can be replaced by their synthesized versions in a quantitative radiomic application, since two effects are coupled. Furthermore, the authors predominantly use morphological characteristics rather than intensity or texture-based features, although the latter are more directly related to the intensity image values so as to measure the impact of the synthesized images. The aforementioned limitations of Reference 19 are overcome in the present work, since our RS is solely trained with acquired images while synthesized images are exclusively used for testing. In addition, intensity and texture features are also considered in the RS that we propose. 2|MATERIALS AND METHODS In this work we propose the application of synthetic MRI to improve an RS for SP in glioblastoma. Following the flow provided in Figure 1, Section 2.1 describes the datasets used in this work and Section 2.2 the data preprocessing stage. Feature extraction and selection as well as classifier training is described in Section 2.3. Section 2.4 describes the procedure for the synthesis of weighted images, which enter the RS exclusively at test time through the feature calculation block. Finally, Section 2.5 describes the three experiments we have carried out to obtain Results I, II, and III listed in Figure 1. 2.1 |Datasets and MR image acquisition Four different datasets of glioblastoma patients were used in this work. Two of them are publicly available, namely, the BraTS2020 (multimodal brain tumor segmentation) 20 Challenge dataset, and the datasets reachable through TCIA (The Cancer Imaging Archive) 21 —which, in turn, consist of thee sources, namely, the Ivy Glioblastoma Altas Project (Ivy-GAP), the Clinical Proteomic Tumor Analysis Consortium Glioblastoma Multiforme (CPTAC) and The Cancer Genome Atlas Glioblastoma Multiforme (TCGA). The other two (Dataset22 and Dataset24), are private datasets acquired in Hospital Universitario Río Hortega, Valladolid, Spain, and Hospital Universitario 12 de Octubre, Madrid, Spain, respectively. Details of the datasets can be found in Table 1. From all the datasets, we only included those patients (199 patients in total) in whom gross total resection (100% of the enhancing tumor (ET) volume) or near-total resection (>95% of the ET volume) could be performed. The cases selected from the public datasets are referenced in Supporting Information Table S1. For each patient, four MR structural weighted images—T1weighted (T1w), T2w, FLAIR, and T1w post-contrast (T1w-c)—were available. All the acquisitions of the private datasets were performed with IRB approval and informed written consent. See Table 2for details of the acquisition parameters. FIGURE 1 Workflow of the proposed approach. Initially, patients are divided into training (175 patients) and testing (24 patients) sets. The preprocessing pipeline segments tumors and normalizes the contrast intensity. Features are retrieved from the segmented regions of interest (ROIs). After feature selection, relevant features are retained. Five machine learning models for SP were examined. Three different experiment configurations (referred to throughout the text as Experiments/Results I, II, and III) were defined for comparative performance assessment MOYA-S ´ AEZ ET AL.3of14 2.2 |Data preprocessing All weighted images were first co-registered to 1 mm 3 isotropic resolution and skull-stripped 22 following BraTS preprocessing in CaPTk. 23 Then, denoising 24 followed by N4 bias correction 25 were performed in order to obtain white matter 26 and tumor segmentations. nNUnet 27 was utilized to segment the tumor into three distinct regions (ET, non-enhancing tumor (NET), and edema (ED)) for feature extraction. On the other hand, N4 bias correction 25 was applied on the skull-stripped images and these were next normalized by dividing each by the mean intensity of the white matter region contralateral to the tumor. This latter pipeline produces the images input to the RS and to the synthesis procedure. Figure 2depicts the preprocessing pipeline. 2.3 |RS for SP of glioblastoma patients Our RS was trained to classify patients according to the survival criterion (survival > ≤480 days). The threshold of 480 days (i.e., 16 months) was chosen in order to achieve a balance between group sizes in the test dataset. BraTS2020, TCIA, and Dataset22 (175 patients in total, see Table 1) were used to train the RS. Dataset24 (24 patients) was used for testing in coordination with the synthesis method described in Section 2.4. Starting from a total of 117 088 handcrafted features extracted from the weighted images, we trained the RS following a nested crossvalidation scheme (outer, 5-fold; inner, 10-fold). Feature selection methods 28,29 were repeated in each outer split to reduce the possible bias produced if training were done on a single cross-validation split. For each outer split, the model with the lowest Brier loss in the inner split was chosen. Note that five models were selected for the following screening due to the fivefold decision of the outer split. Each of these models was then validated with the validation data corresponding to its outer split. Finally, the model with the best performance, in terms of area under the curve (AUC), was selected. The complete pipeline is outlined in Appendix A. Using the methodology previously outlined, the best model for each of the three scenarios described below was chosen. a. When the four weighted images (i.e., T1w, T2w, FLAIR, and T1w-c) were used as input of the RS, the selected model turned out to be an extreme gradient boosting (XGB) with 17 features, two of which are from FLAIR and two from T2w. TABLE 1 Datasets used in this work. BraTS2020 and TCIA are public datasets, whereas Dataset22 and Dataset24 are private datasets. The number of patients in each dataset is denoted by n. Age is shown as mean standard deviation. Survival is defined as the time in days from diagnosis to death (censored =0) or to the last date the patient was known to be alive (censored =1). The percentages of patients with survival less than 16 months (survival < 16 M) for the different datasets are also displayed (16 months or, equivalently, 480 days) Dataset nAge Survival (IQR) % Censored =1 Survival < 16 M BraTS2020 119 62 12 374 (364) 0% 65.6 % TCIA 34 60 10 521 (482) 5.9 % 58.8 % Dataset22 22 65 10 451 (307) 22.7 % 59.1 % Dataset24 24 57 13 552 (218) 29.2 % 54.2 % Whole datset 199 60 11 447 (346) 7.0 % 62.3 % TABLE 2 All MRI sessions are composed of four structural weighted images, namely, a T1w, a T2w, a FLAIR, and a T1w-c. Details of the scanner and the acquisition parameters are provided if available. The acquisition parameters are echo time (TE), repetition time (TR), and inversion time (TI) Dataset Scanner T1w T2w FLAIR T1w-c BraTS2020 19 institutions NA NA NA NA TCIA 8 institutions TE=2.75–19 ms TE=15–120 ms TE=34.6–155 ms TE=2.1–20 ms TR=352–3379 ms TR=700–6370 ms TR=6000–11 000 ms TR=4.9–3285 ms Dataset22 1.5 T GE TE=6.33–12 ms TE=99–110 ms TE=120–127 ms TE=2.56 ms and TR=360–800 ms TR=2680–8480 ms TR=6000–8000 ms TR=7.96 ms 1.5 T Philips TI=2000 ms Dataset24 1.5 T GE TE=1.83 ms TE=122 ms TE=142 ms TE=2.18 ms TR=5.98 ms TR=4162 ms TR=9350 ms TR=6.85 ms TI=2200 ms 4of14 MOYA-S ´ AEZ ET AL. b. When the channel fed with FLAIR was discarded, the resulting model was a logistic regression (LR) classifier with 16 features. c. When the channel fed with T2w was discarded, a support vector machine (SVM) classifier with 16 features was selected. Hereinafter, these three models are termed XGB17, LR16, and SVM16, respectively. Features selected for each of the previous models are listed in Supporting Information Tables S4, S5, and S6, respectively. 2.4 |Synthesis method using a self-supervised CNN In Reference 15 we proposed a synthetic MRI approach for computation of the T1,T2, and proton density (PD) relaxometry maps and the synthesis of different weighted images from only a pair of inputs. A U-Net convolutional neural network (CNN) trained with synthetic data was employed. However, some synthesized weightings, such as FLAIR, presented relatively low quality presumably due to the exclusively synthetic training. Thus, if only a few image contrasts are of interest, this issue can be overcome by extending the CNN 15 into a self-supervised CNN to be trained with acquired images with the desired contrast. Such an extension has been performed in this work and is graphically represented in Figure 3. The original CNN was configured with two encoders for the two input weighted images, namely, T1w and T2w or T1w and FLAIR (in the case of synthesizing FLAIR or T2w, respectively). Then, the latent representations of each encoder were fused using a pixel-wise max function. Finally, we configured three decoders for generation of the T1,T2, and PD relaxometry maps. In order to extend the CNN into a self-supervised CNN (see Figure 3) we have included a non-trainable lambda layer after the decoder's output. This lambda layer implements ideal equations that describe the MR intensity of the output weighted image as a function of the relaxometry maps and the acquisition parameters. These well known equations can be found in Appendix B. The loss function used to train the self-supervised CNN, named Lsyn, is computed in the weighted image domain as the average of the mean absolute error (MAE) between each acquired image and its synthesized counterpart. Specifically, let mkðxÞ denote the intensity value of the kth acquired image at pixel x, defined in some domain χkℝ2, and let mk synðxÞbe the synthesized image at that pixel location. Then, MAE mk,mk syn  ¼1 χk jj X xχk mkðxÞmk synðxÞ     ð1Þ with mkand mk syn vectors that represent the image intensity values of the acquired and synthesized images respectively in all the pixels belonging to domain χkwith cardinality χk   . Then, the loss function is defined as FIGURE 2 Preprocessing pipeline. Initial images are first co-registered and skull-stripped following the CaPTk pipeline. Afterwards, the different contrast images are denoised and bias-corrected before obtaining the white matter and tumor segmentations. Finally, skull-stripped images are bias-corrected and intensity normalized using the segmentations MOYA-S ´ AEZ ET AL.5of14 Lsyn ¼1 MX M k¼1 MAE mk,mk syn  ,ð2Þ with Mthe overall number of images entering the average (i.e., the batch size). Dataset24 (see Table 1) was used to train the self-supervised CNN following a leave-one-out scheme (i.e., a total of 24 models were trained). To train each model, one patient was used for testing and the remaining 23 patients were randomly split between training (18 patients, approximately 80%) and early-stopping validation (5 patients, approximately 20%). During training, the loss function was optimized using the Adam algorithm with a learning rate αof 1 10 4 . Further, we empirically fixed the batch size to 32 images. We ran the code using the TensorFlow backend 30 on a single NVidia GeForce GTX 1070. The total learning took approximately 1 h of computation time for each model, but execution reduces to a few seconds once the network is fully trained. 2.5 |Experiments We carried out three test experiments to assess the performance of using synthesized images as input of an RS for SP. In all of them the RS was tested with Dataset24. These experiments are detailed next. (I) XGB17 was tested with the acquired T1w, T2w, FLAIR, and T1w-c images as inputs. (II) XGB17 was tested replacing one of the acquired inputs (FLAIR and T2w, one at a time) by its synthesized version. Therefore, in this experiment three of the inputs of the RS were acquired and one was synthesized. As previously stated, these synthesized images were the test images from the leave-one-out scheme of the synthesis method, so no overlap between the training and testing splits occurs in either the synthesis or in the RS. (III) Models LR16 and SVM16 were tested without considering as input FLAIR or T2w, respectively. Note that the RSs used in this third experiment have been built with only three input channels. Performance assessment is twofold. On the one hand, we evaluated the quality of the synthesized images. In addition to visual assessment, we also carried out a quantitative analysis using the well known measures 15 mean squared error (MSE), structural similarity index (SSIM), and peak signal to noise ratio (PSNR). These metrics have been defined within a 3D domain between the synthesized and the acquired images, specifically, within the smallest cube that comprises the foreground of each volume; hence χkin Equation (1) is the intersection between this domain and the kth acquired image. Thereafter, the mean and standard deviation values across patients were calculated. FIGURE 3 Overview of the self-supervised CNN. The inputs of the network are T1w and T2w for the synthesis of FLAIR, and T1w and FLAIR for the synthesis of T2w. Note that all the switches change depending on the weighting we want to synthesize. The lambda layers implement the ideal equations indicated in Appendix B 6of14 MOYA-S ´ AEZ ET AL. On the other hand, in order to compare the performance of the RS with the different experiment configurations, we computed AUC, accuracy, precision, recall, and F1-score as classifier performance metrics. These metrics were reported as the average value over the two classification classes. Additionally, for the sake of completeness, we analyzed the predicted probabilities of survival obtained at the output of the RS for the pairs Experiments I–II and I–III. To this end, we calculated the R2value, customarily used in linear regression, to measure how predicted probabilities of Experiments II and III deviated from the identity function at abscissae equal to the probabilities of Experiment I. The intra-class correlation coefficient (ICC) 31 between these pairs was also measured and boxplots of the probability differences for these pairs were constructed. 3|RESULTS Figure 4shows a representative slice of the synthesized and the corresponding acquired weighted images for several test glioblastoma patients. Overall, the synthesized images are close to the acquired versions regarding structural information and contrast between tissues in both healthy and pathological regions. In particular, the contours and intensities of the different lesion areas are similar in both. Note that in Patient 3 the synthesized FLAIR image does not suffer from motion artifacts, which are present in the acquired image. Additionally, Table 3shows the mean and standard deviation values computed across patients of the synthesis quality metrics between synthesized and acquired images for FLAIR and T2w. SSIM is a value ranging between 0 and 1, and the value 1 is only achievable for two identical images. The high values of PSNR and the low values of MSE show low error between the synthesized and the acquired weighted images for both FLAIR and T2w. All the metrics improve considerably with respect to those obtained by Moya-Sáez et al 15 for FLAIR. Note that the comparison of the values obtained for T2w is not representative since this weighted image was input to the CNN in that work. 15 Figure 5shows the AUC, accuracy, precision, recall, and F1-score achieved for the RS for the three different experiment configurations (Experiments I, II, and III defined in Section 2.5). All the metrics are substantially better in the case of using a synthesized image rather than using a system without such a weighted image for both FLAIR and T2w. The comparison between using an acquired and a synthesized image shows that the performance of the system does not diminish in terms of AUC, and only suffers from a slight degradation in terms of the other performance metrics for FLAIR. Such degradation is not observed with the synthesized T2w. FIGURE 4 A representative axial slice of the images synthesized by the self-supervised CNN for different test patients of Dataset24. A, Synthesized FLAIR images. B, Corresponding actually acquired FLAIR images. C, Synthesized T2w images. D, Corresponding actually acquired T2w images MOYA-S ´ AEZ ET AL.7of14 Figure 6shows scatter plots of the output predicted probabilities of Experiments I–II and I–III. The ground-truth labels are also displayed. Additionally, R2values from the identity linear regressions are provided. A better agreement of points in plots in the upper row (Experiments I–II) compared with the plots in the lower row (Experiments I–III) can be observed. This better agreement is also confirmed with the higher values of R2obtained. Additionally, the ICC values measured are 0.983 and 0.292 for FLAIR and 0.964 and 0.027 for T2w for the pairs Experiments I–II and I–III, respectively. Finally, Figure 7shows boxplots of the probability differences for the pairs Experiments I–II and I–III, for both FLAIR and T2w. As can be seen, the median of the boxplots is closer to zero in Pair I–II than in Pair I–III for both weighted images. A lower interquartile range (IQR) in boxplots of Experiments I–II compared with Experiments I–III can be also observed, together with a median shift from zero in the I–III experiment, an effect which is more prominent for T2w. 4|DISCUSSION In this work, we have thoroughly analyzed the replacement of an actually acquired weighted image with a synthesized version for predicting survival of glioblastoma patients with a completely independent RS. Starting from two acquired weighted images, we synthesized a new weighted image using a CNN-based method. The RS was trained using as input acquired images only. Then, the system was tested using as input acquired images on one side, and replacing one acquired image with a synthesized image on the other. We also compared performance with a system trained from scratch ignoring this additional channel. Results show that multicontrast-demanding quantitative applications, such as radiomics, can be leveraged by synthesized images. Synthesized images may allow widespread usage of these RSs in clinical practice, by retrospectively completing databases with missing modalities and/or replacing artifacted images. Further, these synthesized images have the potential to speed up acquisition protocols by replacing some acquired images with their synthesized versions. In particular, removing FLAIR or T2w from an average brain protocol may reduce the overall scan duration by of the order of 20%, and both of them are artifact-prone sequences due to their sensitivity to motion. Thus, our results allow us to state that synthetic MRI does add to glioblastoma SP within a radiomics-based approach. Indeed, the network described in Reference 15 is prepared to TABLE 3 Synthesis quality metrics used to evaluate the capability to synthesize FLAIR and T2w weighted images. These metrics are the MSE, SSIM, and PSNR. Mean and standard deviation (between parentheses) computed across patients are reported. The metrics were calculated between both the synthesized and the acquired images MSE SSIM PSNR FLAIR 0.0163 0.7595 23.7975 (0.0104) (0.0474) (2.0011) T2w 0.0742 0.7845 25.8934 (0.0319) (0.0616) (1.9032) FIGURE 5 AUC, accuracy, precision, recall, and F1-score of the RS tested with Dataset24 when FLAIR (A) and T2w (B) are acquired, synthesized, or ignored in the whole pipeline 8of14 MOYA-S ´ AEZ ET AL. synthesize more than one contrast, so for protocols with more sequences than those used in this paper more reductions could be potentially achieved. Our work shows that synthesized weighted images are visually similar to the acquired images in both healthy and pathological tissues. Additionally, the values of image quality metrics prove the agreement between the synthesized and the actually acquired images. The performance achieved with the synthesized images in the RS is not only close to the performance achieved using the acquired images, but also substantially better than using a model trained without this weighting. This is confirmed by the classifier performance metrics (i.e., AUC, accuracy, precision, recall, and F1-score) for both FLAIR and T2w. The R2of the identity linear regression and the ICC values also support this finding. Note that an ICC value above 0.9 is considered excellent. 31 It is worth noting that the synthesized images input to the XGB17 model improve the AUC compared with acquired images. This might be caused by the implicit filtering undergone during the synthesis procedure. Moreover, accuracy values FIGURE 6 Scatter plots of the predicted probabilities obtained at the output of the RS for the experiment with the acquired versus the synthesized images (top row) and versus an RS trained from scratch without considering this weighted image as input (bottom row). The plots are shown for FLAIR (A) and T2w (B). Dashed lines represent the threshold fixed in the RS to classify survival (>480 days). Each point represents one glioblastoma patient and its color corresponds to the ground-truth labels. R2values from the identity linear regressions are provided FIGURE 7 Boxplots of the differences between the probabilities obtained at the output of the RS for Experiments I–II and I–III. Each point represents the probability difference for each glioblastoma patient and the dashed line corresponds to zero difference MOYA-S ´ AEZ ET AL.9of14