scieee AI-readable full text Open interactive document viewer

Improving Rectal Tumor Segmentation with Anomaly Fusion Derived from Anatomical Inpainting: A Multicenter Study

CAST

Full text

Improving Rectal Tumor Segmentation with Anomaly Fusion Derived from Anatomical Inpainting: A Multicenter Study Lishan Cai a, b, Mohamed A. Abdelatty a, c, Luyi Han a, d, Doenja M. J. Lambregts a, Joost van Griethuysen a, Eduardo Pooch a, b, Regina G.H. Beets-Tan a, b, Sean Benson a, e, Joren Brunekreef f, h, †, Jonas Teuwen f, g, h, †, * a Department of Radiology, Netherlands Cancer Institute, The Netherlands b GROW School for Oncology and Developmental Biology, Maastricht University Medical Centre, The Netherlands c Department of Diagnostic and Interventional Radiology, Kasr Al-Ainy Hospital, Egypt d Department of Radiology and Nuclear Medicine, Radboud University Medical Centre, The Netherlands e Department of Cardiology, Amsterdam Cardiovascular Sciences, Amsterdam University Medical, Centre, The Netherlands f Department of Radiation Oncology, Netherlands Cancer Institute, The Netherlands g Radboud University Medical Center, Department of Medical Imaging, The Netherlands h University of Amsterdam, Faculty of Science, The Netherlands * Corresponding author † Shared last author Abstract Accurate rectal tumor segmentation using magnetic resonance imaging (MRI) is paramount for effective treatment planning. It allows for volumetric and other quantitative tumor assessments, potentially aiding in prognostication and treatment response evaluation. Manual delineation of rectal tumors and surrounding structures is time-consuming and typically. Over the past few years, deep learning has shown strong results in automated tumor segmentation in MRI. Current studies on automated rectal tumor segmentation, however, focus solely on tumoral regions without considering the rectal anatomical entities and often lack a solid multicenter external validation. In this study, we improved rectal tumor segmentation by incorporating anomaly maps derived from anatomical inpainting. This inpainting was implemented using a U-Net-based model trained to reconstruct a healthy rectum and mesorectum from prostate T2-weighted images (T2WI). The rectal anomaly maps were generated from the difference between the original rectal and reconstructed pseudo-healthy slices during inference. The derived anomaly maps were used in the downstream tumor segmentation tasks by fusing them as an additional input channel (AAnnUNet). Alternative methods for integrating rectal anatomical knowledge were evaluated as baselines, including Multi-Target nnUNet (MTnnUNet), which added rectum and mesorectum segmentation as auxiliary tasks, and MultiAll rights reserved. No reuse allowed without permission. perpetuity. preprint (which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in The copyright holder for thisthis version posted October 16, 2024. ; https://doi.org/10.1101/2024.10.15.24315517doi: medRxiv preprint NOTE: This preprint reports new research that has not been certified by peer review and should not be used to guide clinical practice. Channel nnUNet (MCnnUNet), which utilized rectum and mesorectum masks as an additional input channel. As part of this study, we benchmarked nine models for rectal tumor segmentation on a large multicenter dataset of preoperative T2WI as the baseline and nnUNet outperformed the other eight models on the external dataset. The MTnnUNet demonstrated improvements in both supervised and semi-supervised settings (AI-generated rectum and mesoretum were used) compared to nnUNet, while the MCnnUNet showed benefits only in the semi-supervised setting. Importantly, anomaly maps were strongly associated with tumoral regions, and their integration within AAnnUNet led to the best tumor segmentation results across both settings. The effectiveness of AAnnUNet demonstrated the value of the anomaly maps, indicating a promising direction for improving rectal tumor segmentation and model robustness for multicenter data. 1. Introduction Magnetic Resonance Imaging (MRI) plays a pivotal role in staging rectal cancer and selecting treatment plans, providing valuable information regarding the extent of tumor infiltration within and beyond the bowel wall, and into critical anatomical structures, including perirectal vessels, the mesorectal fascia (MRF), peritoneum, and neighboring pelvic organs (Beets-Tan et al., 2018; MERCURY Study Group, 2007). T2-weighted imaging (T2WI) forms the mainstay of the MRI protocol because of its superior soft tissue contrast to discern the different layers of the rectal wall, mesorectal fat, adjacent vessels, and MRF to allow for detailed local staging (Horvat et al., 2019; Suzuki et al., 2008). Precisely segmenting the tumor is an important task in rectal cancer management. Tumor segmentations are utilized for several purposes including radiation treatment planning, volumetric analysis, and extraction of imaging biomarkers, which may serve as a basis for prognostication and treatment response evaluation. Manual rectal tumor delineation by experienced radiologists is considered the current gold standard. Nevertheless, it is time-consuming and subject to substantial intraand inter-observer variation (Hearn et al., 2020; Irving et al., 2016; Trebeschi et al., 2017). Developing an accurate, generalizable, and robust rectal tumor segmentation model can help reduce this variability and assist in standardizing several steps of diagnostic and therapeutic rectal cancer management. Deep learning (DL) has seen a rapid uptake in several fields, achieving state-of-the-art results in multiple medical image analysis tasks (Razzak et al., 2018). In particular, Convolutional Neural Network-based (CNN) based DL approaches excel in learning image representations from annotated data by using learnable feature extraction filters and sequential convolution, activation, and pooling operations. Among CNN architectures, the U-Net (Ronneberger et al., 2015) and its variants are the most popular architecture in medical image segmentation. Studies (Jian et al., 2018; Knuth et al., 2022; Wang et al., 2018) have explored the ability of U-Net and other 2D CNNs on rectal tumor All rights reserved. No reuse allowed without permission. perpetuity. preprint (which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in The copyright holder for thisthis version posted October 16, 2024. ; https://doi.org/10.1101/2024.10.15.24315517doi: medRxiv preprint segmentation based on T2WI. These CNNs achieved a Dice similarity coefficient score (DSC) ranging from 0.59 to 0.84 for tumor segmentation. Although 2D CNNs are less computationally expensive, 3D models can leverage richer context to improve predictions (Mlynarski et al., 2019). (Hamabe et al., 2022) implemented a 3D U-Net, achieving an averaged DSC of 0.73 (0.60-0.80) over 10-fold cross-validation for rectal tumor segmentation. However, no external validation was done in their study. Besides CNN-based models, transformer-based architectures are also being applied in medical image analysis due to their ability to access long-range semantic information (Xiao et al., 2023). (Li et al., 2023) proposed RTAU-Net, a novel 3D dual path fusion network containing a transformer encoder for extracting global contour information of the tumor. RTAU-Net achieved an averaged DSC of 0.80 and 0.68 in data from two respective medical centers. However, RTAU-Net requires the manual removal of tumor-free slices, which hinders the fully automated implementation of the model. Additionally, RTAU-Net was not compared with the state-of-the-art medical segmentation networks, such as nnUNet (Isensee et al., 2021), a self-configuring implementation of the U-Net architecture, or nnFormer (Zhou et al., 2023), which introduces 3D transformer blocks on top of nnUNet. Besides rectal tumor delineation, some studies also demonstrated that CNNs can accurately delineate anatomical structures such as rectum and mesorectum or perirectal fat with DSC above 0.90 (DeSilvio et al., 2023; Hamabe et al., 2022; Kim et al., 2019). Automated rectum and mesorectum delineation could potentially improve radiological evaluation. The prognosis for rectal cancer depends on how far the tumor has infiltrated the layers of the rectal wall and the mesorectum, and the successful attainment of negative circumferential resection margins (CRMs) through surgical intervention (Nagtegaal et al., 2004). Additionally, Lee et al. (Lee et al., 2019) demonstrated that a 2D model’s variance in tumor regions can be reduced by 90% by incorporating rectal segmentation on the model’s objective. Integrating rectal anatomical knowledge can provide a more comprehensive representation of the T2WI, allowing the model to learn richer and more nuanced patterns, leading to better performance on unseen data. However, the impact of adding rectal anatomical structures including mesorectum for rectal tumor segmentation has not been investigated in a multi-institutional setting. Also, incorporating rectal anatomical structure has so far been limited to adding additional segmentation tasks as presented in (Lee et al., 2019). Unlike for medical challenges with public datasets like brain tumor segmentation (Menze et al., 2015) or clinically significant prostate lesion segmentation (PICAI) (Saha et al., 2023), there is no large multicenter publicly available MRI dataset for rectal cancer studies. This makes it difficult to benchmark different models. An extensive external validation study with multicenter data is highly desirable to compare different deep learning segmentation approaches. In this article, we had the following contributions: All rights reserved. No reuse allowed without permission. perpetuity. preprint (which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in The copyright holder for thisthis version posted October 16, 2024. ; https://doi.org/10.1101/2024.10.15.24315517doi: medRxiv preprint • We developed and evaluated a rectal tumor segmentation model incorporating anomaly maps from anatomical inpainting, showing improved performance. • To generate the anomaly maps, we proposed and evaluated a novel end-to-end rectal anatomical inpainting model, trained exclusively on prostate T2WI, to detect and highlight anomalous areas in rectal T2WI. The model trained on prostate T2WI was applied to rectal T2WI, challenging traditional domain-specific practices and demonstrating the potential for transfer learning. The derived rectal anomaly maps can also be integrated into other clinical downstream tasks. • As part of our study, we benchmarked nine 3D deep learning models for rectal tumor segmentation on a large multicenter dataset of T2WI. • We developed and evaluated a 3D deep learning model specifically to segment rectal anatomical structures, including rectum and mesorectum. • We explored different strategies for incorporating rectal anatomical information into rectal tumor segmentation, including the integration of anomaly maps derived from the inpainting model, adding rectal structures as additional tasks, and utilizing them as prior knowledge. • We released the rectum and mesorectum masks of 100 prostate T2WIs from PICAI dataset, annotated by radiologist. • We uploaded the model weights for rectum and mesorectum segmentation, as well as the weights for MTnnUNet, MCnnUNet, and AAnnUNet. 2 Related Work 2.1 Reconstruction-based Medical Anomaly Detection Supervised learning requires a substantial amount of reliably labeled data, which is often difficult to collect in medical imaging. Therefore, methods requiring partly labeled data (semi-supervised) or no labeling (unsupervised methods) have attracted increased attention. Anomaly detection is a method that can use semi-supervised or unsupervised techniques to address medical imaging tasks such as segmentation. Generative models are frequently used in the field of anomaly detection due to their effectiveness. The mode is often trained to reconstruct images from a specific data distribution (e.g., normal tissue). These models, when confronted with images outside of this distribution (such as those containing tumors), often struggle to accurately reconstruct the anomalous regions, resulting in higher reconstruction errors in those areas Generative adversarial networks (GANs) (Goodfellow et al., 2014), Auto-Encoders (AE), and their variants including Variational AE (VAE) (Kingma and Welling, 2022) and Vector Quantized VAE (VQ-VAE) (van den Oord et al., 2017) were proven to be promising (Baur et al., 2021; Chen et al., 2020; Pinaya et al., 2022) in this field. However, these methods have presented several challenges. They struggle to learn and represent healthy anatomical structures when using full image resolution. Additionally, when reconstructing the entire image, the highlighted regions might not exclusively correspond to diseased areas. As a result, the final anomaly All rights reserved. No reuse allowed without permission. perpetuity. preprint (which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in The copyright holder for thisthis version posted October 16, 2024. ; https://doi.org/10.1101/2024.10.15.24315517doi: medRxiv preprint maps may become more sensitive to other artifacts. Finally, if it is in the self-supervised setting where healthy or normal samples are not available for training, these methods can sometimes can successfully reconstruct anomalies due to a high generalization capacity (Gong et al., 2019). 2.2.1 Partial Image Reconstruction via Inpainting for Medical Anomaly Detection Instead of image-to-image reconstruction methods, inpainting focuses on filling in missing or occluded parts of an image by leveraging the surrounding context. The masks used for inpainting are generally independent of the dataset. (Nguyen et al., 2021) implemented an inpainting-based brain tumor segmentation pipeline for T1-weighted MRI, where anomalous regions were determined by identifying areas of highest reconstruction loss. With prior anatomical knowledge, inpainting can focus solely on high-risk regions, reducing the impact from the background. (Yeganeh et al., 2022) developed an anatomy-aware masking strategy for inpainting to effectively aid the model in learning the shape representation of the organs of interest. (Woo et al., 2024) have proposed a UNet-based model for detecting bone lesions in knee MR images through reconstruction via inpainting and demonstrated that detected anomalies can be further utilized for segmentation. The core idea of anomaly detection through anatomical inpainting involves masking the region of interest (ROI), typically covering the relevant anatomical structures. The inpainting model is trained to fill in the masked area without potential anomalies. The discrepancy between the original and reconstructed images is then utilized to identify anomalies. In this study, the anatomical structures strongly associated with rectal cancer, including the rectum and mesorectum (fatty tissue surrounding the rectum), were masked, and the inpainting model was trained using prostate T2WI images from the PICAI dataset (Saha et al., 2023). Despite the different fields of view between prostate T2WI and rectal T2WI, they overlap in anatomical structures. Most prostate T2WI contains a healthy rectum and mesorectum, ensuring the inpainting model was trained to learn the distribution of the healthy tissues. The inferred reconstructed rectal T2WI slices can then be used to generate anomaly maps, highlighting potentially tumoral areas. 3. Methods 3.1 Rectal Anomaly Detection with Masked T2WI using Anatomical Inpainting The overall pipeline, inspired by (Han et al., 2024, 2023), for detecting rectal anomalies is shown in Figure 2. Prostate T2WI with masked rectum and mesorectum were used to train the model for their reconstruction. The inpainting model was adapted from (Han et al., 2024), an end-to-end MRI sequence generation framework. The framework consists of two stages. In the first stage, only the reconstruction loss is optimized for the encoder and generator. In the second stage, both adversarial loss and cycle-consistent loss are incorporated in addition to the reconstruction loss, with the optimization applied to the encoder, generator, and discriminator. The training of the anatomical All rights reserved. No reuse allowed without permission. perpetuity. preprint (which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in The copyright holder for thisthis version posted October 16, 2024. ; https://doi.org/10.1101/2024.10.15.24315517doi: medRxiv preprint inpainting was based on 2D slices. The inpainting model contains an encoder 𝑬 and a decoder 𝑮. A masked (rectum and mesorectum) 2D T2WI slice 𝑋, can be compressed by 𝑬 into a latent space 𝑧 = 𝑬(𝑋) and 𝑮 can reconstruct the original slice from the latent representation 𝑧. The skip connections were added to recover fine-grained details. To enforce similarity between the generated slices and the actual slices, a supervised reconstruction loss is used: 𝐿𝑟𝑒𝑐 = λ𝑟‖𝑋′− 𝑋‖1+ λ𝑝𝐿𝑝(𝑋′− 𝑋) (1) where 𝑋 is the original slice, and 𝑋′= 𝑮(𝑬(𝑋)) is the restored image. ‖⋅‖1 is the 𝐿1 loss and 𝐿𝑝 is the perceptual loss from pre-trained VGG19, which involves comparing high-level features (not just pixel values) from both the generated and reference images (Johnson et al., 2016). Instead of measuring raw pixel differences, it evaluates how similar the images are in terms of their content and style, based on features extracted from different layers of the VGG19 network. λ𝑟 and λ𝑝 are weight factors, for which the respective values 10 and 0.01 were chosen empirically. For the second stage of the training, the adversarial loss and cycle-consistent loss (Zhu et al., 2017) were added on top of the reconstruction loss to ensure that the inpainted images were both realistic and consistent with the original images. The adversarial loss helps to ensure that the completed regions look realistic and blend seamlessly with the surrounding areas and the cycle-consistent loss focuses on preserving the original structure by ensuring the inpainted image can be accurately reconstructed back to the original. 𝑚𝑖𝑛𝐷𝑚𝑎𝑥𝐺 𝐿𝑎𝑑𝑣 =‖𝑫(𝑋)− 1‖2+‖𝑫(𝑋′)‖2 (2) 𝐿𝑐𝑦𝑐 =‖𝑋′′ − 𝑋‖1 (3) where 𝑋′′ = 𝑮(𝑬(𝑋′)) and ‖⋅‖2 is the 𝐿2 loss and 𝑫 is the discriminator. The anomaly maps are then defined as the absolute differences between the reconstructed slice and the original slice, 𝑀 = |𝑋 − 𝑋′| (4) Let 𝐼 be an image with intensity values. The normalization scheme involved the following steps: 𝑙 = 𝑃𝑒𝑟𝑐𝑒𝑛𝑡𝑖𝑙𝑒0.5(𝐼) ℎ = 𝑃𝑒𝑟𝑐𝑒𝑛𝑡𝑖𝑙𝑒99.5(𝐼) 𝐼𝑛𝑜𝑟𝑚 =𝑚𝑎𝑥(𝐼, 𝑙)− 𝑙 ℎ − 𝑙 First Compute the 0.5th percentile (𝑙) and the 99.5th percentile (ℎ) of the intensity values in the image and normalize the intensity values in the image using the computed percentiles. The rectum was masked with a value of 0, while the mesorectum was masked with a value of 0.5. The model input patch size was (384, 384), with a batch size of 1. We used the AdamW optimizer to train the network (both stages) with β 0.9 and 0.95, an initial learning rate of 0.0001, a weight decay factor of 0.05, and All rights reserved. No reuse allowed without permission. perpetuity. preprint (which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in The copyright holder for thisthis version posted October 16, 2024. ; https://doi.org/10.1101/2024.10.15.24315517doi: medRxiv preprint following a polynomial decay. The model was implemented in PyTorch (Torch version 2.1.2) and the training was conducted on an NVIDIA RTX A6000 GPU. 3.2 Study Design To train the inpainting model, 100 prostate T2WI with manually segmented rectum and mesorectum masks were split into 80 for training and 20 for internal validation. The model was additionally tested on 200 prostate T2WIs and the entire rectal dataset, comprising 705 T2WI. However, because annotating rectal structures is labor-intensive, only 180 rectal randomly selected samples have radiologist-annotated masks for the rectum and mesorectum. To address this, a nnUNet, defined as anatomy nnUNet, was trained specifically to segment the rectum and mesorectum using a dataset of 39 samples (training cohort 1) from a single center, which had the highest number of manually annotated rectum and mesorectum masks among nine centers. Some studies have demonstrated that CNNs can accurately delineate anatomical structures such as rectum and mesorectum or perirectal fat with DSC above 0.90 (DeSilvio et al., 2023; Hamabe et al., 2022; Kim et al., 2019). The model was then evaluated on 141 external samples, see Figure 3. This nnUNet was then used to infer all the rectum and mesorectum masks across the whole rectal dataset. The predicted rectum and mesorectum masks were defined as AI-generated pseudo rectum and mesorectum masks. We incorporated anomaly maps generated by the inpainting model into downstream rectal tumor segmentation tasks by adding them as an additional input for nnUNet, called Anomaly-Aware nnUNet (AAnnUNet). We compared this approach with other strategies for integrating anatomical knowledge into tumor segmentation, including Multi-Target nnUNet (MTnnUNet) and Multi-Channel nnUNet (MCnnUNet). MTnnUNet added rectum and mesorectum segmentation as auxiliary tasks, and MultiChannel nnUNet (MCnnUNet) utilized rectum and mesorectum masks as additional input channels. First, we established a baseline for rectal tumor segmentation by comparing the performance of nine 3D deep learning models, see Figure 1, including UNet (Çiçek et al., 2016), ResUNet (Diakogiannis et al., 2020), UNetR (Hatamizadeh et al., 2022b), SwinUNetR (Hatamizadeh et al., 2022a), AttentionUNet (Atten-UNet) (Oktay et al., 2018), MedFormer (Gao et al., 2023), nnFormer (Zhou et al., 2023), U-Mamba (bot) (Ma et al., 2024) and nnUNet (Isensee et al., 2021). All the models underwent training using training cohort 1 with a 5-fold cross-validation. Subsequently, models were externally tested on the remaining 666 samples from nine centers. MTnnUNet, MCnnUNet, and AAnnUNet were trained on training cohort 1 using 5-fold cross-validation and tested on 666 samples from 9 centers. For the inference of MCnnUNet and AAnnUNet, AI-generated pseudo rectum and mesorectum were used. With AI-generated pseudo anatomical structures, MTnnUNet, MCnnUNet, and AAnnUNet were also trained on a larger cohort comprising 141 samples (training cohort 2) using 5-fold cross-validation All rights reserved. No reuse allowed without permission. perpetuity. preprint (which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in The copyright holder for thisthis version posted October 16, 2024. ; https://doi.org/10.1101/2024.10.15.24315517doi: medRxiv preprint and tested on 564 samples from eight centers. Instead of relying on ground truth rectum and mesorectum masks, AI-generated annotations were employed in training. This method merged AIlabeled anatomical structures with manually labeled tumors, indicating a semi-supervised learning approach. 3.3 Segmentation Models 1. UNet (Çiçek et al., 2016), extends the previous u-net architecture from Ronneberger et al., 2015 (Ronneberger et al., 2015) by replacing all 2D operations with their 3D counterparts. 2. ResUNet (Diakogiannis et al., 2020), is a modified version of UNet. It replaces double convolution layers of UNet with residual blocks from ResNet (He et al., 2016), incorporating shortcut connections for faster convergence. This adaptation works in both 2D and 3D settings, enhancing performance in capturing complex patterns. 3. UNETR (Hatamizadeh et al., 2022b), adopts a ViT-inspired encoder and employs a CNN decoder for 3D image segmentation. The images are initially divided into patches, which are linearly transformed into token embeddings. These tokens undergo processing through a self-attention block, akin to ViT. To manage the quadratic complexity of self-attention, the patch size is set to be relatively large (16) to prevent overly long sequence lengths. 4. SwinUNETR (Hatamizadeh et al., 2022a), reformulates the segmentation task as a sequence-tosequence prediction using a Swin Transformer as the encoder. The encoder is then connected to a Fully Convolutional Neural Network (FCNN)-based decoder through skip connections. 5. Attention UNet (Oktay et al., 2018), introduces an attention-gating module to UNet to enhance its ability to suppress irrelevant regions and highlight salient features crucial for a given task. 6. nnFormer (Zhou et al., 2023), is a 3D transformer for volumetric medical image segmentation nnFormer combines interleaved convolution and self-attention operations. It introduces a special selfattention mechanism to understand both local and global aspects of the image volume. To improve efficiency, it also uses skip attention instead of the usual concatenation or summation operations. 7. MedFormer (Gao et al., 2023), is a transformer-based designed to handle scalable 3D medical image segmentation, including three crucial components: a beneficial inductive bias, hierarchical modeling using linear-complexity attention, and multi-scale feature fusion that combines spatial and semantic information globally. MedFormer can learn from both small and large scale data without pre-training. 8. U-Mamba (Ma et al., 2024), is inspired by the State Space Sequence Models (SSMs) (Gu et al., 2021), which are known for their ability to handle long sequences. The model is designed specifically for biomedical image segmentation with the hybrid CNN-SSM block that integrates the local feature extraction power of convolutional layers with the abilities of SSMs to capture the long-range dependency. All rights reserved. No reuse allowed without permission. perpetuity. preprint (which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in The copyright holder for thisthis version posted October 16, 2024. ; https://doi.org/10.1101/2024.10.15.24315517doi: medRxiv preprint 9. nnUNet (Isensee et al., 2021), is a self-configuring framework for medical image segmentation. It utilizes UNet as its architecture but offers a specialized preprocessing, training technique, and hyperparameter configuration. nnUNet achieves state-of-the-art performance on several medical image segmentation challenges with a relatively simple architectural design. 10. Anatomy nnUNet, is the nnUNet trained to segment rectal-related anatomical structures including the rectum and mesorectum. 11. MTnnUNet, is the nnUNet trained to segment rectum, mesorectum, and rectal tumors. 12. MCnnUNet, is the nnUNet trained to segment rectal tumors with rectum and mesorectum masks as additional input channels. 13. AAnnUNet, is the nnUNet trained to segment rectal tumors with anomaly maps 𝑀 derived from anatomical inpainting. For rectal tumor segmentation, the imaging preprocessing approach was adopted from nnUNet, which included ZScoreNormalization for standardizing intensities, uniform resampling of all images, and a cropping process. All the segmentation models were implemented in PyTorch (Torch version 2.1.2) and trained on an NVIDIA A6000 GPU with randomly initialized weights without transfer learning. The batch size was set to 2 and the models were trained for 1000 epochs with the SGD optimizer. The loss function is the sum of the cross-entropy and Dice loss. During inference, predictions were obtained by averaging the outputs of each model resulting from the 5-fold cross-validation procedure. 3.4 Evaluation Metrics and Statistical Analysis Statistical analysis was conducted in Python (version 3.9) with the SciPy package (version 1.13.1). To measure the performance of image reconstruction, Structural Similarity Index Measurement (SSIM), Peak Signal-to-NoiseRatio (PSNR) were used. To measure the segmentation performance, the Dice Similarity Coefficient score (DSC) and 95% Hausdorff Distance (HD) were utilized on both crossvalidation and external tests. The characteristic differences of cohorts were compared by the KruskalWallis test. The Mann–Whitney U-test was used to compare the difference of indicators among different methods. The model performance differences were calculated using the paired sample t-test. All statistical analyses were two-sided and p-values below 0.05 were regarded as statistically significant. 95% confidence intervals were generated using the bootstrap method with 10,000 replications. 4. Results 4.1 Dataset and Patient Characteristics As a part of a previous institutional review board approved multicenter study project (Bogveradze et al., 2022; Cai et al., 2024; El Khababi et al., 2023; “ESGAR 2020 Book of Abstracts,” 2020; Schurink et al., 2023, 2022) the clinical and imaging data from 1426 patients with biopsy-proven rectal cancer All rights reserved. No reuse allowed without permission. perpetuity. preprint (which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in The copyright holder for thisthis version posted October 16, 2024. ; https://doi.org/10.1101/2024.10.15.24315517doi: medRxiv preprint Acknowledgment The study was supported by the Research High Performance Computing (RHPC) facility of the Netherlands Cancer Institute. Funding sources This study has received funding from the European Union’s Horizon 2020 Research and Innovation Programme under the Marie Skłodowska-Curie grant agreement No 857894. All rights reserved. No reuse allowed without permission. perpetuity. preprint (which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in The copyright holder for thisthis version posted October 16, 2024. ; https://doi.org/10.1101/2024.10.15.24315517doi: medRxiv preprint Reference Baur, C., Denner, S., Wiestler, B., Navab, N., Albarqouni, S., 2021. Autoencoders for unsupervised anomaly segmentation in brain MR images: A comparative study. Med Image Anal 69, 101952. https://doi.org/10.1016/j.media.2020.101952 Beets-Tan, R.G., Lambregts, D.M., Maas, M., Bipat, S., Barbaro, B., Curvo-Semedo, L., Fenlon, H.M., Gollub, M.J., Gourtsoyianni, S., Halligan, S., 2018. Magnetic resonance imaging for clinical management of rectal cancer: updated recommendations from the 2016 European Society of Gastrointestinal and Abdominal Radiology (ESGAR) consensus meeting. European radiology 28, 1465–1475. Bogveradze, N., el Khababi, N., Schurink, N.W., van Griethuysen, J.J.M., de Bie, S., Bosma, G., Cappendijk, V.C., Geenen, R.W.F., Neijenhuis, P., Peterson, G., Veeken, C.J., Vliegen, R.F.A., Maas, M., Lahaye, M.J., Beets, G.L., Beets-Tan, R.G.H., Lambregts, D.M.J., 2022. Evolutions in rectal cancer MRI staging and risk stratification in The Netherlands. Abdom Radiol 47, 38–47. https://doi.org/10.1007/s00261-021-03281-8 Bogveradze, N., Snaebjornsson, P., Grotenhuis, B.A., van Triest, B., Lahaye, M.J., Maas, M., Beets, G.L., Beets-Tan, R.G.H., Lambregts, D.M.J., 2023. MRI anatomy of the rectum: key concepts important for rectal cancer staging and treatment planning. Insights into Imaging 14, 13. https://doi.org/10.1186/s13244-022-01348-8 Cai, L., Lambregts, D.M.J., Beets, G.L., Mass, M., Pooch, E.H.P., Guérendel, C., Beets-Tan, R.G.H., Benson, S., 2024. An automated deep learning pipeline for EMVI classification and response prediction of rectal cancer using baseline MRI: a multi-centre study. npj Precis. Onc. 8, 1–11. https://doi.org/10.1038/s41698-024-00516-x Chen, X., You, S., Tezcan, K.C., Konukoglu, E., 2020. Unsupervised Lesion Detection via Image Restoration with a Normative Prior. Çiçek, Ö., Abdulkadir, A., Lienkamp, S.S., Brox, T., Ronneberger, O., 2016. 3D U-Net: Learning Dense Volumetric Segmentation from Sparse Annotation, in: Ourselin, S., Joskowicz, L., Sabuncu, M.R., Unal, G., Wells, W. (Eds.), Medical Image Computing and ComputerAssisted Intervention – MICCAI 2016. Springer International Publishing, Cham, pp. 424– 432. https://doi.org/10.1007/978-3-319-46723-8_49 Defeudis, A., Mazzetti, S., Panic, J., Micilotta, M., Vassallo, L., Giannetto, G., Gatti, M., Faletti, R., Cirillo, S., Regge, D., 2022. MRI-based radiomics to predict response in locally advanced rectal cancer: Comparison of manual and automatic segmentation on external validation in a multicentre study. European Radiology Experimental 6, 19. DeSilvio, T., Antunes, J.T., Bera, K., Chirra, P., Le, H., Liska, D., Stein, S.L., Marderstein, E., Hall, W., Paspulati, R., Gollamudi, J., Purysko, A.S., Viswanath, S.E., 2023. Region-specific deep learning models for accurate segmentation of rectal structures on post-chemoradiation T2w MRI: a multi-institutional, multi-reader study. Front. Med. 10. https://doi.org/10.3389/fmed.2023.1149056 Diakogiannis, F.I., Waldner, F., Caccetta, P., Wu, C., 2020. ResUNet-a: A deep learning framework for semantic segmentation of remotely sensed data. ISPRS Journal of Photogrammetry and Remote Sensing 162, 94–114. https://doi.org/10.1016/j.isprsjprs.2020.01.013 Dou, M., Chen, Z., Tang, Y., Sheng, L., Zhou, J., Wang, X., Yao, Y., 2023. Segmentation of rectal tumor from multi-parametric MRI images using an attention-based fusion network. Med Biol Eng Comput 61, 2379–2389. https://doi.org/10.1007/s11517-023-02828-9 El Khababi, N., Beets-Tan, R.G.H., Tissier, R., Lahaye, M.J., Maas, M., Curvo-Semedo, L., Dresen, R.C., Nougaret, S., Beets, G.L., Lambregts, D.M.J., Bakers, F.C.H., Barros, P., Bauer, F., de Bie, S.H., Ballantyne, S., Dutra, J.B., Buskov, L., Bogveradze, N., Bosma, G.P.T., Cappendijk, V.C., Castagnoli, F., Charalampos, S., Delli Pizzi, A., Digby, M., Geenen, R.W.F., van Griethuysen, J.J.M., Lafrance, J., Mahajan, V., Malekzadeh, S., Neijenhuis, P.A., Peterson, G.M., Pieters, I., Schurink, N.W., Smit, R., Veeken, C.J., Vliegen, R.F.A., Wray, A., Zeina, A.-R., on behalf of the rectal MRI study group, 2023. Predicting response to chemoradiotherapy in rectal cancer via visual morphologic assessment and staging on baseline MRI: a multicenter and multireader study. Abdom Radiol 48, 3039–3049. https://doi.org/10.1007/s00261-023-03961-7 All rights reserved. No reuse allowed without permission. perpetuity. preprint (which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in The copyright holder for thisthis version posted October 16, 2024. ; https://doi.org/10.1101/2024.10.15.24315517doi: medRxiv preprint ESGAR 2020 Book of Abstracts, 2020. . Insights into Imaging 11, 64. https://doi.org/10.1186/s13244-020-00873-8 FACG, M.A.M., MD, FACR, 2006. Dynamic Radiology of the Abdomen: Normal and Pathologic Anatomy. Springer Science & Business Media. Gao, Y., Zhou, M., Liu, D., Yan, Z., Zhang, S., Metaxas, D.N., 2023. A Data-scalable Transformer for Medical Image Segmentation: Architecture, Model Efficiency, and Benchmark. https://doi.org/10.48550/arXiv.2203.00131 Gong, D., Liu, L., Le, V., Saha, B., Mansour, M.R., Venkatesh, S., Van Den Hengel, A., 2019. Memorizing Normality to Detect Anomaly: Memory-Augmented Deep Autoencoder for Unsupervised Anomaly Detection, in: 2019 IEEE/CVF International Conference on Computer Vision (ICCV). Presented at the 2019 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 1705–1714. https://doi.org/10.1109/ICCV.2019.00179 Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., Bengio, Y., 2014. Generative Adversarial Nets, in: Advances in Neural Information Processing Systems. Curran Associates, Inc. Gu, A., Johnson, I., Goel, K., Saab, K., Dao, T., Rudra, A., Ré, C., 2021. Combining Recurrent, Convolutional, and Continuous-time Models with Linear State Space Layers, in: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 572–585. Hamabe, A., Ishii, M., Kamoda, R., Sasuga, S., Okuya, K., Okita, K., Akizuki, E., Sato, Y., Miura, R., Onodera, K., Hatakenaka, M., Takemasa, I., 2022. Artificial intelligence–based technology for semi-automated segmentation of rectal cancer using high-resolution MRI. PLoS One 17, e0269931. https://doi.org/10.1371/journal.pone.0269931 Han, L., Tan, T., Zhang, T., Huang, Y., Wang, X., Gao, Y., Teuwen, J., Mann, R., 2024. Synthesisbased imaging-differentiation representation learning for multi-sequence 3D/4D MRI. Med Image Anal 92, 103044. https://doi.org/10.1016/j.media.2023.103044 Han, L., Zhang, T., Huang, Y., Dou, H., Wang, X., Gao, Y., Lu, C., Tan, T., Mann, R., 2023. An Explainable Deep Framework: Towards Task-Specific Fusion for Multi-to-One MRI Synthesis, in: Greenspan, H., Madabhushi, A., Mousavi, P., Salcudean, S., Duncan, J., SyedaMahmood, T., Taylor, R. (Eds.), Medical Image Computing and Computer Assisted Intervention – MICCAI 2023. Springer Nature Switzerland, Cham, pp. 45–55. https://doi.org/10.1007/978-3-031-43999-5_5 Hatamizadeh, A., Nath, V., Tang, Y., Yang, D., Roth, H.R., Xu, D., 2022a. Swin UNETR: Swin Transformers for Semantic Segmentation of Brain Tumors in MRI Images, in: Crimi, A., Bakas, S. (Eds.), Brainlesion: Glioma, Multiple Sclerosis, Stroke and Traumatic Brain Injuries. Springer International Publishing, Cham, pp. 272–284. https://doi.org/10.1007/9783-031-08999-2_22 Hatamizadeh, A., Tang, Y., Nath, V., Yang, D., Myronenko, A., Landman, B., Roth, H.R., Xu, D., 2022b. UNETR: Transformers for 3D Medical Image Segmentation. Presented at the Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 574–584. He, K., Zhang, X., Ren, S., Sun, J., 2016. Identity Mappings in Deep Residual Networks, in: Leibe, B., Matas, J., Sebe, N., Welling, M. (Eds.), Computer Vision – ECCV 2016. Springer International Publishing, Cham, pp. 630–645. https://doi.org/10.1007/978-3-319-46493-0_38 Hearn, N., Bugg, W., Chan, A., Vignarajah, D., Cahill, K., Atwell, D., Lagopoulos, J., Min, M., 2020. Manual and semi-automated delineation of locally advanced rectal cancer subvolumes with diffusion-weighted MRI. Br J Radiol 93, 20200543. https://doi.org/10.1259/bjr.20200543 Horvat, N., Carlos Tavares Rocha, C., Clemente Oliveira, B., Petkovska, I., Gollub, M.J., 2019. MRI of rectal cancer: tumor staging, imaging techniques, and management. Radiographics 39, 367–387. Irving, B., Franklin, J.M., Papież, B.W., Anderson, E.M., Sharma, R.A., Gleeson, F.V., Brady, S.M., Schnabel, J.A., 2016. Pieces-of-parts for supervoxel segmentation with global context: Application to DCE-MRI tumour delineation. Medical Image Analysis 32, 69–83. https://doi.org/10.1016/j.media.2016.03.002 All rights reserved. No reuse allowed without permission. perpetuity. preprint (which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in The copyright holder for thisthis version posted October 16, 2024. ; https://doi.org/10.1101/2024.10.15.24315517doi: medRxiv preprint Isensee, F., Jaeger, P.F., Kohl, S.A., Petersen, J., Maier-Hein, K.H., 2021. nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation. Nature methods 18, 203– 211. Isensee, F., Wald, T., Ulrich, C., Baumgartner, M., Roy, S., Maier-Hein, K., Jaeger, P.F., 2024. nnUNet Revisited: A Call for Rigorous Validation in 3D Medical Image Segmentation. Jayaprakasam, V.S., Paroder, V., Gibbs, P., Bajwa, R., Gangai, N., Sosa, R.E., Petkovska, I., Golia Pernicka, J.S., Fuqua, J.L., Bates, D.D., 2022. MRI radiomics features of mesorectal fat can predict response to neoadjuvant chemoradiation therapy and tumor recurrence in patients with locally advanced rectal cancer. European radiology 1–10. Jian, J., Xiong, F., Xia, W., Zhang, R., Gu, J., Wu, X., Meng, X., Gao, X., 2018. Fully convolutional networks (FCNs)-based segmentation method for colorectal tumors on T2-weighted magnetic resonance images. Australas Phys Eng Sci Med 41, 393–401. https://doi.org/10.1007/s13246018-0636-9 Johnson, J., Alahi, A., Fei-Fei, L., 2016. Perceptual Losses for Real-Time Style Transfer and SuperResolution, in: Leibe, B., Matas, J., Sebe, N., Welling, M. (Eds.), Computer Vision – ECCV 2016. Springer International Publishing, Cham, pp. 694–711. https://doi.org/10.1007/978-3319-46475-6_43 Kim, J., Oh, J.E., Lee, J., Kim, M.J., Hur, B.Y., Sohn, D.K., Lee, B., 2019. Rectal cancer: Toward fully automatic discrimination of T2 and T3 rectal cancers using deep convolutional neural network. International Journal of Imaging Systems and Technology 29, 247–259. https://doi.org/10.1002/ima.22311 Kingma, D.P., Welling, M., 2022. Auto-Encoding Variational Bayes. Knuth, F., Adde, I.A., Huynh, B.N., Groendahl, A.R., Winter, R.M., Negård, A., Holmedal, S.H., Meltzer, S., Ree, A.H., Flatmark, K., Dueland, S., Hole, K.H., Seierstad, T., Redalen, K.R., Futsaether, C.M., 2022. MRI-based automatic segmentation of rectal cancer using 2D U-Net on two independent cohorts. Acta Oncologica 61, 255–263. https://doi.org/10.1080/0284186X.2021.2013530 Lambregts, D.M.J., Maas, M., Beets-Tan, R.G.H., 2010. MRI of the Rectum, in: Stoker, J. (Ed.), MRI of the Gastrointestinal Tract. Springer, Berlin, Heidelberg, pp. 205–227. https://doi.org/10.1007/978-3-540-85532-3_13 Lee, J., Oh, J.E., Kim, M.J., Hur, B.Y., Sohn, D.K., 2019. Reducing the Model Variance of a Rectal Cancer Segmentation Network. IEEE Access 7, 182725–182733. https://doi.org/10.1109/ACCESS.2019.2960371 Li, D., Wang, J., Yang, J., Zhao, J., Yang, X., Cui, Y., Zhang, K., 2023. RTAU-Net: A novel 3D rectal tumor segmentation model based on dual path fusion and attentional guidance. Computer Methods and Programs in Biomedicine 242, 107842. https://doi.org/10.1016/j.cmpb.2023.107842 Ma, J., Li, F., Wang, B., 2024. U-Mamba: Enhancing Long-range Dependency for Biomedical Image Segmentation. https://doi.org/10.48550/arXiv.2401.04722 Menze, B.H., Jakab, A., Bauer, S., Kalpathy-Cramer, J., Farahani, K., Kirby, J., Burren, Y., Porz, N., Slotboom, J., Wiest, R., Lanczi, L., Gerstner, E., Weber, M.-A., Arbel, T., Avants, B.B., Ayache, N., Buendia, P., Collins, D.L., Cordier, N., Corso, J.J., Criminisi, A., Das, T., Delingette, H., Demiralp, Ç., Durst, C.R., Dojat, M., Doyle, S., Festa, J., Forbes, F., Geremia, E., Glocker, B., Golland, P., Guo, X., Hamamci, A., Iftekharuddin, K.M., Jena, R., John, N.M., Konukoglu, E., Lashkari, D., Mariz, J.A., Meier, R., Pereira, S., Precup, D., Price, S.J., Raviv, T.R., Reza, S.M.S., Ryan, M., Sarikaya, D., Schwartz, L., Shin, H.-C., Shotton, J., Silva, C.A., Sousa, N., Subbanna, N.K., Szekely, G., Taylor, T.J., Thomas, O.M., Tustison, N.J., Unal, G., Vasseur, F., Wintermark, M., Ye, D.H., Zhao, L., Zhao, B., Zikic, D., Prastawa, M., Reyes, M., Van Leemput, K., 2015. The Multimodal Brain Tumor Image Segmentation Benchmark (BRATS). IEEE Transactions on Medical Imaging 34, 1993–2024. https://doi.org/10.1109/TMI.2014.2377694 MERCURY Study Group, 2007. Extramural depth of tumor invasion at thin-section MR in patients with rectal cancer: results of the MERCURY study. Radiology 243, 132–139. https://doi.org/10.1148/radiol.2431051825 All rights reserved. No reuse allowed without permission. perpetuity. preprint (which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in The copyright holder for thisthis version posted October 16, 2024. ; https://doi.org/10.1101/2024.10.15.24315517doi: medRxiv preprint Mlynarski, P., Delingette, H., Criminisi, A., Ayache, N., 2019. 3D convolutional neural networks for tumor segmentation using long-range 2D context. Computerized Medical Imaging and Graphics 73, 60–72. https://doi.org/10.1016/j.compmedimag.2019.02.001 Nagtegaal, I., Gaspar, C., Marijnen, C., van de Velde, C., Fodde, R., van Krieken, H., 2004. Morphological changes in tumour type after radiotherapy are accompanied by changes in gene expression profile but not in clinical behaviour. The Journal of Pathology 204, 183–192. https://doi.org/10.1002/path.1621 Nguyen, B., Feldman, A., Bethapudi, S., Jennings, A., Willcocks, C.G., 2021. Unsupervised RegionBased Anomaly Detection In Brain MRI With Adversarial Image Inpainting, in: 2021 IEEE 18th International Symposium on Biomedical Imaging (ISBI). Presented at the 2021 IEEE 18th International Symposium on Biomedical Imaging (ISBI), pp. 1127–1131. https://doi.org/10.1109/ISBI48211.2021.9434115 Oktay, O., Schlemper, J., Folgoc, L.L., Lee, M., Heinrich, M., Misawa, K., Mori, K., McDonagh, S., Hammerla, N.Y., Kainz, B., Glocker, B., Rueckert, D., 2018. Attention U-Net: Learning Where to Look for the Pancreas. https://doi.org/10.48550/arXiv.1804.03999 Pinaya, W.H.L., Tudosiu, P.-D., Gray, R., Rees, G., Nachev, P., Ourselin, S., Cardoso, M.J., 2022. Unsupervised brain imaging 3D anomaly detection and segmentation with transformers. Med Image Anal 79, 102475. https://doi.org/10.1016/j.media.2022.102475 Razzak, M.I., Naz, S., Zaib, A., 2018. Deep Learning for Medical Image Processing: Overview, Challenges and the Future, in: Dey, N., Ashour, A.S., Borra, S. (Eds.), Classification in BioApps: Automation of Decision Making. Springer International Publishing, Cham, pp. 323–350. https://doi.org/10.1007/978-3-319-65981-7_12 Ronneberger, O., Fischer, P., Brox, T., 2015. U-Net: Convolutional Networks for Biomedical Image Segmentation, in: Navab, N., Hornegger, J., Wells, W.M., Frangi, A.F. (Eds.), Medical Image Computing and Computer-Assisted Intervention – MICCAI 2015, Lecture Notes in Computer Science. Springer International Publishing, Cham, pp. 234–241. https://doi.org/10.1007/9783-319-24574-4_28 Saha, A., Bosma, J., Twilt, J., Ginneken, B. van, Yakar, D., Elschot, M., Veltman, J., Fütterer, J., Rooij, M. de, Huisman, H., 2023. Artificial Intelligence and Radiologists at Prostate Cancer Detection in MRI — The PI-CAI Challenge. Presented at the Medical Imaging with Deep Learning, short paper track. Schurink, N.W., van Kranen, S.R., Roberti, S., van Griethuysen, J.J.M., Bogveradze, N., Castagnoli, F., el Khababi, N., Bakers, F.C.H., de Bie, S.H., Bosma, G.P.T., Cappendijk, V.C., Geenen, R.W.F., Neijenhuis, P.A., Peterson, G.M., Veeken, C.J., Vliegen, R.F.A., Beets-Tan, R.G.H., Lambregts, D.M.J., 2022. Sources of variation in multicenter rectal MRI data and their effect on radiomics feature reproducibility. Eur Radiol 32, 1506–1516. https://doi.org/10.1007/s00330-021-08251-8 Schurink, N.W., van Kranen, S.R., van Griethuysen, J.J., Roberti, S., Snaebjornsson, P., Bakers, F.C., de Bie, S.H., Bosma, G.P., Cappendijk, V.C., Geenen, R.W., 2023. Development and multicenter validation of a multiparametric imaging model to predict treatment response in rectal cancer. European Radiology 1–10. Suzuki, C., Torkzad, M.R., Tanaka, S., Palmer, G., Lindholm, J., Holm, T., Blomqvist, L., 2008. The importance of rectal cancer MRI protocols on iInterpretation accuracy. World J Surg Onc 6, 89. https://doi.org/10.1186/1477-7819-6-89 Teney, D., Abbasnejad, E., Kafle, K., Shrestha, R., Kanan, C., van den Hengel, A., 2020. On the Value of Out-of-Distribution Testing: An Example of Goodhart’ s Law, in: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 407–417. Trebeschi, S., van Griethuysen, J.J.M., Lambregts, D.M.J., Lahaye, M.J., Parmar, C., Bakers, F.C.H., Peters, N.H.G.M., Beets-Tan, R.G.H., Aerts, H.J.W.L., 2017. Deep Learning for FullyAutomated Localization and Segmentation of Rectal Cancer on Multiparametric MR. Sci Rep 7, 5301. https://doi.org/10.1038/s41598-017-05728-9 van den Oord, A., Vinyals, O., kavukcuoglu, koray, 2017. Neural Discrete Representation Learning, in: Advances in Neural Information Processing Systems. Curran Associates, Inc. All rights reserved. No reuse allowed without permission. perpetuity. preprint (which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in The copyright holder for thisthis version posted October 16, 2024. ; https://doi.org/10.1101/2024.10.15.24315517doi: medRxiv preprint Varoquaux, G., Cheplygina, V., 2022. Machine learning for medical imaging: methodological failures and recommendations for the future. npj Digit. Med. 5, 1–8. https://doi.org/10.1038/s41746022-00592-y Wang, J., Lu, J., Qin, G., Shen, L., Sun, Y., Ying, H., Zhang, Z., Hu, W., 2018. Technical Note: A deep learning-based autosegmentation of rectal tumors in MR images. Medical Physics 45, 2560–2564. https://doi.org/10.1002/mp.12918 Willemink, M.J., Koszek, W.A., Hardell, C., Wu, J., Fleischmann, D., Harvey, H., Folio, L.R., Summers, R.M., Rubin, D.L., Lungren, M.P., 2020. Preparing Medical Imaging Data for Machine Learning. Radiology 295, 4–15. https://doi.org/10.1148/radiol.2020192224 Woo, B., Engstrom, C., Baresic, W., Fripp, J., Crozier, S., Chandra, S.S., 2024. Automated anomalyaware 3D segmentation of bones and cartilages in knee MR images from the Osteoarthritis Initiative. Med Image Anal 93, 103089. https://doi.org/10.1016/j.media.2024.103089 Xiao, H., Li, L., Liu, Q., Zhu, X., Zhang, Q., 2023. Transformers in medical image segmentation: A review. Biomedical Signal Processing and Control 84, 104791. https://doi.org/10.1016/j.bspc.2023.104791 Yeganeh, Y., Farshad, A., Navab, N., 2022. Shape-Aware Masking for Inpainting in Medical Imaging. Zhou, H.-Y., Guo, J., Zhang, Y., Han, X., Yu, L., Wang, L., Yu, Y., 2023. nnFormer: Volumetric Medical Image Segmentation via a 3D Transformer. IEEE Transactions on Image Processing 32, 4036–4045. https://doi.org/10.1109/TIP.2023.3293771 Zhu, J.-Y., Park, T., Isola, P., Efros, A.A., 2017. Unpaired Image-to-Image Translation Using CycleConsistent Adversarial Networks, in: 2017 IEEE International Conference on Computer Vision (ICCV). Presented at the 2017 IEEE International Conference on Computer Vision (ICCV), pp. 2242–2251. https://doi.org/10.1109/ICCV.2017.244 All rights reserved. No reuse allowed without permission. perpetuity. preprint (which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in The copyright holder for thisthis version posted October 16, 2024. ; https://doi.org/10.1101/2024.10.15.24315517doi: medRxiv preprint Table 1 Summary of patient demographic and clinical characteristics of the multicenter dataset Values in age parentheses are the minimum and maximum. Values in parentheses of other items are the percentages. cT, baseline clinical T staging. cN, baseline clinical N staging. Location, tumor location. EMVI+: EMVI positive. EMVI-, EMVI negative. ∗, p values were calculated using the Kruskal-Wallis test between the training cohort, training cohort2 and the external dataset. ∗∗, p values were calculated using the Chi-square test between the training cohort1, training cohort2 and the external dataset. All Training Cohort 1 Training Cohort 2 p-value Age (median, range) 65 (26-88) 65 (39-82) 66 (44-87) 0.75∗, 0.17∗ Sex Female 247 (35%) 12 (31%) 40 (28%) 0.69∗∗, 0.08∗∗ Male 458 (65%) 27 (69%) 101 (72%) 1-2 87 (12%) 3 (8%) 19 (13%) 0.58∗, 0.24∗ cT 3 519 (74%) 34 (87%) 107 (76%) 4 99 (13%) 2 (5%) 15 (11%) 0 268 (38%) 14 (36%) 67 (48%) 0.52∗, 0.001∗ cN 1 264 (37 %) 13 (33%) 54 (38%) 2 173 (25%) 12 (31%) 20 (14%) Low 432 (61%) 25 (64%) 93 (66%) 0.45∗, 0.65∗ Location Middle 236(34%) 14 (36%) 43 (30%) High 37 (5%) 0 (0%) 5 (4%) EMVI EMVI + 274 (39%) 18 (46%) 38 (27%) 0.43∗∗, 0.002∗∗ EMVI - 431 (61%) 21 (54%) 103 (73%) Total 705 39 141 All rights reserved. No reuse allowed without permission. perpetuity. preprint (which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in The copyright holder for thisthis version posted October 16, 2024. ; https://doi.org/10.1101/2024.10.15.24315517doi: medRxiv preprint Table 2 Rectum and Mesorectum inpainting performance in the test data aSSIM ± STD aPSNR ± STD Prostate T2WI (Num = 200) 86.72 ± 4.45 25.87 ± 1.51 Rectal T2WI (Num = 705) 83.38 ± 5.33 23.87 ± 2.56 Rectal T2WI Tumor Regions 39.32 ± 13.98 17.50 ± 3.80 Rectal T2WI Tumor-Free Regions 83.62 ± 4.22 24.76 ± 2.41 aSSIM, averaged SSIM. aPSNR, averaged PSNR. STD: Standard Deviation All rights reserved. No reuse allowed without permission. perpetuity. preprint (which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in The copyright holder for thisthis version posted October 16, 2024. ; https://doi.org/10.1101/2024.10.15.24315517doi: medRxiv preprint Table 3 Comparison of various models on rectal tumor segmentation in the external test (Num = 666, 9 centers) Network UNet ResUNet UNetR SwinUNetR Atten-UNet MedFormer nnFormer U-Mamba nnUNet aDSC (%) 55.5 (53.7, 57.2) 55.8 (54.0, 0.57.7) 42.2 (40.4, 44.1) 41.3 (39.4, 43.3) 55.0 (53.0, 57.0) 57.9 (56.0, 59.8) 44.8 (42.4, 47.3) 57.6 (55.4, 60.0) 62.8 (60.7, 64.8) mDSC (%) 63.2 (61.8, 64.7) 64.8 (63.3, 66.3) 47.8 (44.8, 50.7) 46.2 (42.3, 50.1) 65.2 (63.1, 67.2) 67.0 (65.3, 68.7) 56.7 (51.6, 61.8) 70.9 (69.2, 72.6) 73.2 (71.8, 74.7) aHD (mm) 26.04 (22.85, 29.23) 24.31 (21.46, 27.16) 43.8 (39.5, 48.2) 37.65 (34.23, 41.07) 24.73 (21.51, 27.95) 22.74 (19.77, 25.70) 26.90 (23.39, 30.42) 21.52 (17.60, 25.45) 17.28 (14.63,19.94) mHD (mm) 10.40 (9.32,11.49) 9.39 (8.39, 10.40) 22.0 (18.4, 25.7) 20.66 (18.04, 23.27) 8.69 (7.48, 9.91) 8.15 (6.93, 9.32) 16.13 (13.41, 18.84) 7.45 (6.54, 8.36) 6.36 (5.49, 7.22) aDSC: average Dice Coefficient Similarity. mDSC: median Dice Coefficient Similarity. aHD: average of 95% Hausdorff Distance. mHD: median of 95% Hausdorff Distance. The 95% confidential intervals are presented. All rights reserved. No reuse allowed without permission. perpetuity. preprint (which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in The copyright holder for thisthis version posted October 16, 2024. ; https://doi.org/10.1101/2024.10.15.24315517doi: medRxiv preprint Table 4 Comparison of nnUNet, MTnnUNet, MCnnUNet, AannUNet and Ensemble on rectal tumor segmentation in the external test, fully supervised setting (Num = 666, 9 centers) Network nnUNet MTnnUNet MCnnUNet AA-nnUNet Ensemble aDSC (%) 62.8 (60.7, 64.8) 65.4 (63.7, 67.2) 60.5 (58.6,62.4) 65.6 (63.9, 67.3) 69.1 (67.6, 70.6) mDSC (%) 73.2 (71.8, 74.7) 74.4 (73.4, 75.5) 70.2 (69.0, 71.5) 72.8 (72.0, 73.7) 75.5 (74.5, 76.5) aHD (mm) 17.28 (14.63,19.94) 13.63 (11.87, 15.38) 21.06 (19.51,22.62) 16.69 (14.57,18.82) 14.74 (12.72, 16.75) mHD (mm) 6.36 (5.45, 7.22) 5.61 (5.04, 6.18) 7.99 (7.00, 9.02) 5.61 (4.96, 6.25) 5.09 (4.64, 5.55) aDSC: average Dice Coefficient Similarity. mDSC: median Dice Coefficient Similarity. aHD: average of 95% Hausdorff Distance. mHD: median of 95% Hausdorff Distance. The 95% confidential intervals are presented. All rights reserved. No reuse allowed without permission. perpetuity. preprint (which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in The copyright holder for thisthis version posted October 16, 2024. ; https://doi.org/10.1101/2024.10.15.24315517doi: medRxiv preprint (a) (b) Figure 7: (a) The DSC and 95% HD boxplots of nnUNet, MTnnUNet, and MCnnUNet in a fully supervised manner. (b) The DSC 95% HD boxplots of nnUNet, MTnnUNet, and MCnnUNet in a semi-supervised manner. All rights reserved. No reuse allowed without permission. perpetuity. preprint (which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in The copyright holder for thisthis version posted October 16, 2024. ; https://doi.org/10.1101/2024.10.15.24315517doi: medRxiv preprint Figure 8: The visualization of the segmentation performance of nnUNet, MTnnUNet, MCnnUNet, AAnnUNet, and Ensemble using T2WI, supervised setting. Each row is a different sample from the external test set. The columns from left to right are original T2WI, ground truth, tumor prediction masks from nnUNet, MTnnUNet, MCnnUNet, AAnnUNet, and Ensemble. Figure 9: The example where all algorithms failed to locate and segment the rectal tumor, but the anomaly map highlighted the tumoral region. All rights reserved. No reuse allowed without permission. perpetuity. preprint (which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in The copyright holder for thisthis version posted October 16, 2024. ; https://doi.org/10.1101/2024.10.15.24315517doi: medRxiv preprint