Full text
Received 5 June 2025, accepted 26 June 2025, date of publication 21 July 2025, date of current version 29 July 2025. Digital Object Identifier 10.1109/ACCESS.2025.3591272 Comparative Analysis of Deep Learning-Based Feature Extractors for Change Detection in Automotive Radar Maps HARIHARA BHARATHY SWAMINATHAN 1, ARON SOMMER2, URI IURGEL2, ANDREAS BECKER 3, (Member, IEEE), AND MARTIN ATZMUELLER 1,4 1Semantic Information Systems Group, Osnabrück University, 49074 Osnabrück, Germany 2Aptiv Services Deutschland GmbH, 42119 Wuppertal, Germany 3Faculty of Information Technology, Fachhochschule Dortmund, 44139 Dortmund, Germany 4German Research Center for Artificial Intelligence (DFKI), 49084 Osnabrück, Germany Corresponding author: Harihara Bharathy Swaminathan (hsw[email protected]) ABSTRACT The Siamese network architecture has been applied by deep learning practitioners to find similarities between images. In the domain of autonomous driving, this network configuration has recently gained attention for solving the change detection task, which involves identifying changes in a previously known map of a vehicle’s environment. This is vital, as such deviations may compromise the accuracy and reliability of the map, which is essential for the vehicle’s ability to localize itself and navigate effectively. In this paper, we present a set of experiments involving state-of-the-art deep learning architectures based on both convolution (CNN) and attention mechanisms such as AlexNet, GoogLeNet, VGG, ResNet, Vision Transformer, and Shifted Windows Transformer as possible candidates for the feature extractor backbone module in the Siamese architecture to detect changes caused by the disappearance and appearance of construction zones. Also, we evaluate the performance of these architectures using finetuning, i. e., initializing the convolutional layers with pre-trained weights. In our experimentation, the best results were obtained using VGG16 (CNN), especially when it was initialized using pre-trained weights from the ImageNet-1K dataset. In particular, VGG16 with an average F1 score of 92% on highway datasets outperformed the baseline residual network composed of ResNet18 convolutions by about 13.5%. INDEX TERMS Change detection, automotive radar, occupancy maps, siamese networks. I. INTRODUCTION The performance of state-of-the-art deep neural networks has been evaluated in the past on the basis of their scores in the classification of benchmark datasets such as the ImageNet. In this paper, we focus on feature extractors in a specific domain, i. e., automotive radar. In particular, we evaluate the performance of feature extractors such as AlexNet, GoogLeNet, VGG, ResNet, Vision Transformer (ViT) and Shifted Windows (SWIN) Transformer for the purpose of change detection on automotive radarbased maps. In this work, the task is framed as a binary classification problem, in which the model predicts The associate editor coordinating the review of this manuscript and approving it for publication was Fabrizio Santi . whether the reference map has diverged from the current environment. These predictions help assess the accuracy and relevance of the reference map and indicate when remapping is necessary due to detected changes in the vehicle’s environment. As our main contributions, we present results on how to detect changes caused by the disappearance and appearance of construction zones. Furthermore, we evaluate the performance of these feature extractors after fine-tuning, i. e., initializing the convolutional layers with pre-trained weights. For the purpose of change detection, we use a Siamese network configuration, which consists of two deep neural network (DNN) sub-networks as backbones for feature extraction from an image pair, followed by a decision head for image similarity learning. VOLUME 13, 2025 2025 The Authors. This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ 130629
H. B. Swaminathan et al.: Comparative Analysis of Deep Learning-Based Feature Extractors In our experiments, we focus on the domain of automotive radar, for which realistic open benchmark datasets are rather scarce and need to be tailored to the specific application domain. Therefore, as a dataset, we use real-life automotive radar map image pairs generated over a one-year period around the German city of Wuppertal, particularly on highways with construction zones. Our main findings are summarized as follows: 1) Siamese networks configured with VGG-family of feature extractors outperformed the baseline configuration with a residual network i. e., ResNet18. Specifically, the VGG16 backbone achieved an F1 score increase of about 13.5% on average on test datasets depicting disappearing and appearing highway construction zones. 2) By initializing the convolution layers in VGG using pre-trained weights obtained from classification of ImageNet-1K dataset, we overall achieved faster convergence through the process of fine-tuning. The rest of the paper is structured as follows: Section II discusses related work. After that, Section III presents the change detection dataset along with the applied evaluation metrics. Next, Section IV and Section Vdiscusses our experimentation and results in detail, respectively. Finally, Section VI concludes with a summary and interesting directions for future work. II. RELATED WORK In the past, Siamese networks [1] have been extensively studied for various use cases, ranging from signature verification [2], face verification [3] to change detection in remote sensing applications [4]. Recently, Swaminathan et al. have proposed a Siamese network-based classifier for detecting changes due to the appearance and disappearance of construction zones on radar maps of the environment [5]. As backbones, they used the convolution layers of ResNet-18 as a feature extractor. As an input to the multi-layer perceptron (MLP) network of the decision module, feature vectors from the CNN backbones were concatenated. The intermediate layers of the MLP had a size of 256 and 1, subsequently followed by a sigmoid activation at the end, trained using a binary crossentropy (BCE) loss function. This configuration will serve as the baseline for our work and is shown in Figure 1. FIGURE 1. The baseline siamese network architecture composed of a ResNet-18 feature extractor and MLP network as a decision head. Overall, to the best of the authors’ knowledge, a comparative analysis of existing state-of-the-art deep neural network architectures as candidates for a feature extraction backbone in a Siamese network configuration has not been carried out in the literature before. III. DATASET AND EVALUATION METRICS A. CHANGE DETECTION DATASET For the purpose of training and testing, we use the change detection dataset [5] presented previously by Swaminathan et al. The dataset includes radar maps showing changes such as the appearance and disappearance of construction zones. These maps were collected over a year around Wuppertal, Germany, with a focus on a highway connecting Wuppertal and Düsseldorf, as shown in Figure 2. A negative sample (Figure 2a) shows only irrelevant or minor changes (marked with ovals), while positive samples show construction zones appearing or disappearing, along with some unrelated changes, cf. Figures 2b and 2c. During data collection, certain lanes were closed due to construction, and the test vehicle was driven through a nearby lane on the highway. These maps were generated using measurements of corner radar sensors fitted to the test vehicle. Furthermore, a high precision DGPS measurement unit i.e. the POS LV-220 positioning system from Applanix corporation was used to obtain precise positioning of the vehicle. Each sample of this dataset consists of perfectly aligned bird’s eye-view (BEV) radar map image pairs obtained from different measurement drives of the same environment. Within a specific data sample, terms such as a ‘‘reference map’’ and a ‘‘current map’’ are used to denote images that were captured at different points in time. We divide the training and testing samples, ensuring that a construction zone shown during the training process is not part of the testing one. For our experiments, the image pairs comprising the appearance and disappearance of construction zones in the route towards Düsseldorf are considered as part of the training dataset and the image pairs comprising of appearing (A) and disappearing construction zones (D) in the route towards Wuppertal as part of the testing dataset. In terms of size and distribution, the training set contained a total of 38,353 samples with 55.56% negatives and 44.43% positives and the testing set contained a total of 8967 samples with 74.18% negatives and 25.82% positives. B. EVALUATION METRICS In real-world situations, when considering the change detection dataset, the reference map typically matches the current sensor measurements, and deviations from a reference map rarely occur. Thus, a corresponding change detection dataset features a rather large amount of non-changes (negatives) vs. instances of changes (positives). Since accuracy can be a misleading performance metric for imbalanced datasets, we apply the F1 scores for evaluating the predictions 130630 VOLUME 13, 2025
H. B. Swaminathan et al.: Comparative Analysis of Deep Learning-Based Feature Extractors FIGURE 2. Instances of a non change, an appearing and a disappearing construction zone from the change detection dataset. Relevant changes i.e. construction zones are annotated with white boxes and apparent changes are annotated in ovals. Note: Only relevant changes were taken into account for labeling. of the classifier. Table 1presents the components of the confusion matrix for evaluating classification performance. In order to evaluate model performance, we report the average F1 score across two test datasets featuring disappearing and appearing construction zones on highways. Additionally, we consider the number of Floating Point Operations (FLOPs), a commonly used metric for assessing computational complexity. VOLUME 13, 2025 130631
H. B. Swaminathan et al.: Comparative Analysis of Deep Learning-Based Feature Extractors TABLE 1. Confusion matrix / change detection problem. IV. EXPERIMENTS Below, we present the experimental setups involving state-ofthe-art deep networks such as AlexNet, GoogLeNet, VGG, ResNet, ViT and SWIN transformer as feature extractors. A. IMPLEMENTATION DETAILS We trained all the models using an adam optimizer [6] for 15 epochs. For training, we use mini-batches of 16 image pairs, the Binary Cross-Entropy (BCE) loss function as formalized in Equation (1). LBCE = − 1 N N X i=1 yi·log(p(yi)) +(1 −yi)·log(1 −p(yi)) (1) For vanilla model implementations and pre-trained weights, we apply the torchvision subpackage from pytorch. Usually, the input images contained a single channel and spatial dimensions of 256 ×256. Any other changes to the input are defined in the description part of the corresponding backbone. Due to a change on the overall network configuration, the search for an optimal learning rate was performed individually for the different backbones. For experiments incorporating pre-trained ImageNet weights, the learning rate was set to 0.0001, which was lower than the learning rate used for training models from scratch. B. AlexNet AlexNet [7] is a deep network composed of five convolution layers and three max-pooling layers. For the first two convolutional layers, kernels of size 11 ×11 and 5 ×5 are used alongside kernels of size 3 ×3 for deeper layers. Each branch of the siamese network processes an input image to give an output volume containing 256 output channels with spatial dimensions of 6 ×6. In Figure 3, we have shown the convolution part of this feature extractor, a part of the vanilla AlexNet. Subsequently, we flatten this 3D volume and feed it as input to a fully connected (FC) layer with 512 output nodes to complete the process of feature extraction. C. GoogLeNet The vanilla GoogLeNet [8], shown in Figure 5, contains a total of nine inception blocks followed by a global average pooling layer at the end. As shown in Figure 4, each inception block contains convolutions with kernels such as 1×1, 3 ×3, 5 ×5 and 3 ×3 max pooling to let the network decide the best kernel configuration for the problem. Inside an inception block, these convolutions are carried out FIGURE 3. AlexNet architecture, cf. [7]. in a parallel fashion and the feature maps generated are concatenated together at the end. In every parallel branch of the inception block, 1 ×1 convolution reduces the number of input channels to larger convolutions such as 3 ×3 and 5 ×5. Global average pooling operation instead of the conventional fully-connected layers results in reduction of the number of trainable parameters and computation cost in GoogLeNet. As a result of global average pooling at the end, we extract an one-dimensional vector of 1024 features from each single channel image input. FIGURE 4. Inception module of GoogLeNet, cf. [8]. FIGURE 5. The GoogLeNet architecture, cf. [8]. D. VGG The VGG [9] network is composed of multiple successive smaller 3 ×3 kernels (stride 1) than the use of 11 ×11, 7 ×7 and 5 ×5 kernels. For instance, a single 5×5 convolution layer containing 25 learnable parameters could be replaced by two 3 ×3 convolution layers with 18 trainable parameters. The primary building block of the VGG consists of a sequence of convolutions with 3×3 kernels (stride 1, padding 1) followed by a 2 ×2 max-pooling layer (stride 2). Through the arrangement of VGG blocks in cascades, four different variants of the VGG architecture are possible namely VGG11, VGG13, VGG16 and VGG19, whose convolution stems and bodies were considered for our experiments. Finally, a global average pooling layer 130632 VOLUME 13, 2025
H. B. Swaminathan et al.: Comparative Analysis of Deep Learning-Based Feature Extractors was attached to the last convolution layer to extract a total of 512 features from each single channel image. The experimental setup containing these variants along with their corresponding convolution layers are depicted in Figure 6. FIGURE 6. Multiple variants of VGG architecture, cf. [9]. E. ResNet Residual networks or ResNets [10] use skip connection between layers, acting as a highway that connects earlier layers to far more deeper layers of a CNN due to the exploding and vanishing gradient problem encountered during training of deep networks. A ResNet is made up of a number of residual blocks containing such skip connections, arranged one after the other. A comparison between traditional CNN architecture and a ResNet block with a skip connection is shown in Figure 7. For the experiments, convolution layers of the variants ResNet-18, ResNet-34 and ResNet-50 with a global average pooling layer at the end, as described in Figure 8were considered for feature extraction. While variants such as ResNet-18, ResNet-34 extracted a total of 512 features from a single channel image, ResNet-50 extracted 2048 features from each single channel input image. F. VISION TRANSFORMER (ViT) In ViT [11], an input image is divided into series of square patches, transformed into tokens and encoded using blocks that implement a dot-product based attention mechanism [12] to capture relationship between the patches. For our experiments, we use ViT-B/16 to process three channel images with spatial dimensions 224 ×224, divided into multiple square patches of size 16 ×16 ×3 and transformed into tokens of dimension 196 ×768, where 196 represents sequence length and 768 is the hidden dimension. The transformer encoder consists of twelve blocks in cascade. Each encoder FIGURE 7. Comparison of a traditional deep CNN block and the Residual block in ResNet with skip connections, cf. [10]. FIGURE 8. Layer-wise comparison of ResNet-18, ResNet-34 and ResNet-50 architectures, cf. [10]. block consists of a multi-head attention module followed by a Multi-layer perceptron (MLP) network. As shown in Figure 9, the Vision Transformer processes image patches as tokens, similar to words in NLP transformers. For feature extraction, we calculate an average for all patches instead of taking the features from the class token. Therefore, we obtain a total of 768 features from each input image. G. SWIN TRANSFORMER (SWIN) For our experiments, we use the Swin-V2-S [13] variant to process three channel images of spatial dimensions 256 ×256. Similar to ViT, an image passed through a Swin Transformer is divided into non-overlapping patches. However, a scaled cosine attention mechanism replaces the dot-product attention used in ViT. Instead of applying selfattention over all patches at once, the Swin Transformer computes attention within small windows of the image. These windows are shifted between layers to allow information to flow across different regions. VOLUME 13, 2025 130633
H. B. Swaminathan et al.: Comparative Analysis of Deep Learning-Based Feature Extractors FIGURE 9. Vision Transformer architecture, cf. [11]. The architecture of SWIN follows a hierarchical design with multiple stages, where each stage reduces the image resolution and increases the feature dimension. This structure helps the model capture both local and global information efficiently, as shown in Figure 10. Apart from removing the final classification head, no other changes were made to the standard Swin Transformer for our experiments. As the encoder output, we obtain a total of 768 features from each input image. FIGURE 10. Swin Transformer architecture, cf. [13]. H. VISION TRANSFORMER (ViT) WITH EARLY CONVOLUTIONS We explored the impact of early convolutional layers on Vision Transformers (ViTs), inspired by the paper ‘‘Early Convolutions Help Transformers See Better’’ by Xiao [14]. Instead of the standard ViT approach of simply splitting an image into 16×16 patches as proposed by [11], we integrated a ‘‘convolutional stem’’ at the beginning. This stem uses smaller, successive 3×3 kernels (like those found in VGG networks) to process the 224×224 input images. At the stem’s end, a 1×1 kernel ensures the output correctly matches the 768-dimensional input expected by the Transformer encoder. We then reshape these feature maps into a sequence before feeding them into the first block of a ViT-B/16 transformer. FIGURE 11. Image feature extraction using a hybrid architecture consisting of a convolutional stem and a ViT-B/16 transformer, cf. [14]. Figure 11 illustrates the forward pass for feature extraction from an image input. We tested two distinct convolutional stem configurations: 1) VGG16-Inspired Convolutional Stem: This first setup uses a deep convolutional stem designed to mimic the architecture of VGG16. It consists of thirteen 3×3 convolutional layers, configured with increasing output channels (like 64, 128, 256, 512, etc.), similar to the original VGG16 network shown in Figure 12. This makes it a relatively heavyweight option, packed with many learnable parameters. When used within a Siamese network (where two identical branches process separate inputs), the features extracted from both branches are concatenated before being fed into the decision head’s Multi-Layer Perceptron (MLP). FIGURE 12. VGG16-Inspired Convolutional Stem, cf. [9]. 2) Lightweight Convolutional Stem: In contrast, our second configuration, shown in Figure 13, features a much lighter convolutional stem. This version has only four 3×3 convolutional layers. Each of these layers is followed by a batch normalization (BN) layer, a rectified linear unit (ReLU) activation, and a maxpooling layer. Similar to the VGG16-inspired stem, the output channels for these four layers are sequentially set to 64, 128, 256, and 512. This design has fewer learnable parameters compared to the VGG16-inspired stem. Instead of concatenating the features from both branches, it takes the squared difference between them. This specific operation results in the first layer of the decision head’s MLP having only 768 nodes, as its processing the difference directly rather than a combined input. In Table 2, we have listed all combinations of feature extractors and the corresponding decision heads composed of a MLP network, considered for the experiments. V. RESULTS In this section, we present the average F1 scores obtained by the siamese network configurations listed in Table 2 130634 VOLUME 13, 2025
H. B. Swaminathan et al.: Comparative Analysis of Deep Learning-Based Feature Extractors FIGURE 13. Lightweight Convolutional Stem, cf. [9]. TABLE 2. Feature extractors and decision heads configured in a Siamese network. on the dataset presented previously in Section III. Additionally, we also discuss notable conclusions from experiments where the entire network was trained from scratch as well as fine-tuning the network using weights for backbone layers from pre-trained ImageNet1k [15] dataset. For the purpose of finding the best model, we present the average F1 scores obtained on highway test datasets in Figure 14 and amount of Floating Point Operations (FLOPs) in Figure 15. FIGURE 14. Average F1 score obtained by siamese networks configured with various state-of-the-art feature extractors on highway change detection datasets. The experimental results where fine-tuning was performed with pre-trained weights of ImageNet1k dataset are represented using (p) next to the backbone. On the basis of the F1 scores presented in Figure 14, we observe that convolution based feature extractors such FIGURE 15. Amount of Floating Point Operations (FLOPs) in siamese networks configured with various state-of-the-art backbone feature extractors. as the VGG and ResNet18 outperform the purely attentionbased ViT-B/16 and SWIN. Even though this trend is counterintuitive, we can attribute it to size of the training dataset. Additionally, Vision transformers such as ViT or SWIN take a very long time to converge and are highly-sensitive to the choice of the learning rate, exhibiting substandard optimizability. In contrast, the use of early 3×3 convolutions in configurations such as VGG16 convolutional Stem + ViT-B/16 and Light-weight convolutional Stem +ViT-B/16 has enabled for quicker convergence, robustness to the respective learning rate choice (0.0001) as well as the choice of the applied optimizer (Adam). Overall, VGG feature extractors like VGG11(p), VGG13(p), VGG16(p) and VGG19(p) have outperformed the baseline ResNet18 due to the use of multiple successive smaller 3 ×3 kernels with stride 1. Even though we run into difficulties training a deep, memory intensive network from scratch without any skip connections like VGG, initializing its convolution layers using pre-trained weights enables for faster convergence. This process of fine-tuning a feature extractor that has already been trained for extraction of features from a larger ImageNet1K dataset is more efficient than training it from scratch. VGG16(p) with close to 92% F1 score and ≈79.8 GFLOPs is considered as the best backbone choice for change detection. An increase in model performance by ≈13.5% over the baseline configuration with ResNet18 (81% F1 score) as feature extractor is observed from their corresponding average F1 scores on the test datasets. For the purpose of understanding the contribution from individual layers of VGG16(p) towards model performance, we conducted experiments by freezing them and training the network for a total of 15 epochs. The average F1 scores obtained by the siamese network with their corresponding fine-tuning settings are given in Table 3. Table 3indicates that freezing deeper layers such as 18 to 25 of the VGG16(p) resulted in a significant drop in model performance as compared to freezing earlier layers such as layers 0 to 6. When the complete backbone is frozen VOLUME 13, 2025 130635
H. B. Swaminathan et al.: Comparative Analysis of Deep Learning-Based Feature Extractors TABLE 3. Average F1 score obtained by the siamese network with VGG16 feature extractor on highway change detection test datasets during fine-tuning. and only the weights of FC layers are updated during training the network for transfer learning, we obtain an average F1 score of 68% only. Hence, the process of fine-tuning the model is proven to be more effective than transfer learning. Additionally, the computational feasibility of real-time deployment was assessed by performing post-training static quantization using the PyTorch quantization API, resulting in a model size of 14.4 MB and an inference time of approximately 188 milliseconds on the target hardware, which featured an AMD Ryzen 7 2700X Eight-Core CPU with an x86-based architecture—demonstrating the model’s suitability for real-time applications. A. VISUALIZING MODEL PREDICTIONS WITH GradCAM Since DNNs such as CNNs are often treated as black boxes, GradCAM [16] was introduced for providing visual explanations that help for understanding why a model makes a certain prediction. This technique helps by generating a coarse activation map highlighting which parts of the input image contributed most to a model’s prediction. FIGURE 16. Input maps and their corresponding GradCAM heatmaps of predictions obtained from the siamese network architecture with VGG-16 backbone feature extractor. Note: The highlighted regions of the heatmaps denote the regions on the inputs which contributed the most towards a prediction. According to the GradCAM activation maps of correctly predicted positive (TP) shown in Figure 16a, the network is trained to ignore irrelevant changes and focus directly on those portions of the input map pairs where relevant changes due to appearance of construction zones are located. It can also be observed that the model is also unable to focus on construction zones located away from the vehicle on Figure 16b. For the true negative (TN) prediction shown in Figure 16c, the network focuses on those portions of the images where there are some amount of visible changes, which are not part of the construction zones and thus considered as irrelevant. The occurrence of false positives (FP) is attributed mainly to incorrect labeling as shown by the GradCAM activation maps of incorrectly predicted positive in Figure 16c. Here, the disappearance of a construction zone is not labeled correctly as a Change. Finally, for an example of an incorrectly classified negative (FN) shown in Figure 16d, the feature extractor fails to focus on those portions of the input map pairs where there are relevant changes due to the disappearance of a construction zone. VI. CONCLUSION In this work, we presented a comparative analysis on stateof-the-art deep neural networks using both convolution and attention mechanisms as candidates for backbone feature extraction in a siamese network configuration for the purpose of change detection. A VGG16 backbone managed to outperform the baseline siamese network configuration which used ResNet18 by about 13.5%. In order to obtain best results, fine-tuning using pre-trained weights for convolution layers of VGG was performed. According to these results, we highly recommend the use of VGG-blocks containing multiple successive smaller 3 ×3 kernels with stride 1 as a backbone feature extractor for change detection on automotive radar based environment maps. Furthermore, the performance of attention-based backbones such as ViT and SWIN were also evaluated and found to be inferior than those of their convolution mechanism counterparts. However, the overall performance of an attention-based feature extraction mechanism such as ViT was improved via the help of early convolutions with multiple successive smaller 3 ×3 kernels. A significant challenge in addressing the change detection task is the limited availability of datasets that capture a diverse range of changes, including those arising from seasonal variations and dynamic traffic conditions. For future work, it would be beneficial to perform a domain shift analysis to assess the model’s performance on scenes from previously unobserved locations or environmental conditions. Furthermore, the lack of semantic annotations poses an additional hurdle, as their presence could enable models to be trained for more fine-grained, pixel-level change detection. In future work, we aim to extend our modeling approach and analysis, also targeting out-of-distribution cases, with methodological and architectural refinements. Here, also few-shot learning methods provide interesting future directions, e.g., [17],[18]. REFERENCES [1] Y. Li, C. L. P. Chen, and T. Zhang, ‘‘A survey on Siamese network: Methodologies, applications, and opportunities,’’ IEEE Trans. Artif. Intell., vol. 3, no. 6, pp. 994–1014, Dec. 2022. [2] J. Bromley, I. Guyon, Y. LeCun, E. Säckinger, and R. Shah, ‘‘Signature verification using a ‘Siamese’ time delay neural network,’’ in Proc. Adv. Neural Inf. Process. Syst., vol. 6, 1993, pp. 737–744. 130636 VOLUME 13, 2025
H. B. Swaminathan et al.: Comparative Analysis of Deep Learning-Based Feature Extractors [3] S. Chopra, R. Hadsell, and Y. LeCun, ‘‘Learning a similarity metric discriminatively, with application to face verification,’’ in Proc. IEEE Comput. Soc. Conf. Comput. Vis. Pattern Recognit. (CVPR), vol. 1, Sep. 2005, pp. 539–546. [4] L. Mou, M. Schmitt, Y. Wang, and X. Xiang Zhu, ‘‘A CNN for the identification of corresponding patches in SAR and optical imagery of urban scenes,’’ in Proc. Joint Urban Remote Sens. Event (JURSE), Mar. 2017, pp. 1–4. [5] H. B. Swaminathan, A. Sommer, U. Iurgel, A. Becker, and M. Atzmueller, ‘‘Change detection in automotive radar based occupancy maps using Siamese networks,’’ in Proc. Int. Radar Symp., Jul. 2024, pp. 56–61. [6] P. Qi, W. Zhou, and J. Han, ‘‘A method for stochastic L-BFGS optimization,’’ in Proc. IEEE 2nd Int. Conf. Cloud Comput. Big Data Anal. (ICCCBDA), Apr. 2017, pp. 156–160. [7] A. Krizhevsky, I. Sutskever, and G. E. Hinton, ‘‘ImageNet classification with deep convolutional neural networks,’’ in Proc. Adv. Neural Inf. Process. Syst., vol. 60, 2017, pp. 84–90. [8] C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, ‘‘Going deeper with convolutions,’’ in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), Jun. 2015, pp. 1–9. [9] K. Simonyan and A. Zisserman, ‘‘Very deep convolutional networks for large-scale image recognition,’’ 2014, arXiv:1409.1556. [10] K. He, X. Zhang, S. Ren, and J. Sun, ‘‘Deep residual learning for image recognition,’’ in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), Jun. 2016, pp. 770–778. [11] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, ‘‘An image is worth 16x16 words: Transformers for image recognition at scale,’’ 2020, arXiv:2010.11929. [12] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, ‘‘Attention is all you need,’’ in Proc. Adv. Neural Inf. Process. Syst., vol. 30, 2017, pp. 5998–6008. [13] Z. Liu, H. Hu, Y. Lin, Z. Yao, Z. Xie, Y. Wei, J. Ning, Y. Cao, Z. Zhang, L. Dong, F. Wei, and B. Guo, ‘‘Swin transformer v2: Scaling up capacity and resolution,’’ in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Jun. 2022, pp. 12009–12019. [14] T. Xiao, M. Singh, E. Mintun, T. Darrell, P. Dollár, and R. Girshick, ‘‘Early convolutions help transformers see better,’’ in Proc. Adv. Neural Inf. Process. Syst., 2021, pp. 30392–30400. [15] J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, ‘‘ImageNet: A large-scale hierarchical image database,’’ in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., Jun. 2009, pp. 248–255. [16] R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, ‘‘Grad-CAM: Visual explanations from deep networks via gradient-based localization,’’ in Proc. IEEE Int. Conf. Comput. Vis. (ICCV), Oct. 2017, pp. 618–626. [17] G. Cheng, L. Cai, C. Lang, X. Yao, J. Chen, L. Guo, and J. Han, ‘‘SPNet: Siamese-prototype network for few-shot remote sensing image scene classification,’’ IEEE Trans. Geosci. Remote Sens., vol. 60, 2022. [18] J. J. Valero-Mas, A. J. Gallego, and J. R. Rico-Juan, ‘‘An overview of ensemble and feature learning in few-shot image classification using Siamese networks,’’ Multimedia Tools Appl., vol. 83, no. 7, pp. 19929–19952, Jul. 2023. HARIHARA BHARATHY SWAMINATHAN was born in Chennai, India, in 1993. He received the B.Tech. degree in electronics and communication engineering from the B. S. Abdur Rahman Crescent Institute of Science and Technology, Tamil Nadu, India, in 2014, and the M.Eng. degree in embedded systems for mechatronics from the University of Applied Sciences (FH), Dortmund, Germany, in 2020. He is currently pursuing the Ph.D. degree with the Semantic Information Systems (SIS) Group, Osnabrück University, Germany. His research interests include application of machine learning and deep learning for environment perception of self-driving cars, particularly using data from automotive radars. ARON SOMMER was born in Berlin, Germany, in 1986. He received the Dipl.-Math. Techn. degree in Technomathematics with a technical background in communications engineering from Karlsruhe Institute of Technology (KIT), Karlsruhe, Germany, in 2013, and the Ph.D. degree from the Institute of Information Processing, Leibniz University Hannover, in 2019, where he was involved in a project on synthetic aperture radar. Since 2020, he has been the Radar Algorithm Expert for automotive applications with Aptiv. URI IURGEL received the Dipl.-Ing. degree in electrical engineering (with a focus on information technology) from the University of Duisburg, Germany, in 2000, and the Dr.-Ing. degree from the Technical University of Munich, in 2006. Since 2005, he has been with Delphi/Aptiv, currently holding the position of the Manager Radar Perception and a Senior Expert Sensor Algorithms. His research interests include artificial intelligence, information retrieval, natural language processing, algorithms for environment perception with camera and radar, and radar signal processing, and specifically relates to applications for autonomous driving and advanced driver assistance systems. ANDREAS BECKER (Member, IEEE) was born in Wuppertal, Germany, in 1975. He received the Dipl.-Ing. degree in electrical engineering and the Dr.-Ing. degree from the University of Wuppertal, in 2000 and 2006, respectively. During the doctoral degree, his research was focused on the numerical simulation of electromagnetic problems, including the analysis of borehole radars. From 2007 to 2016, he was with Hella GmbH & Co. KGaA, and Delphi Deutschland GmbH (now Aptiv), eventually achieving the position of the Radar Technical Manager. Since 2017, he has been a Professor of information technology with the University of Applied Sciences Dortmund, Dortmund, Germany. His research interest includes perception and control systems of mobile robots. MARTIN ATZMUELLER is currently a Full Professor with the Institute of Computer Science, Osnabrück University, Germany, where he holds the ROSEN-Group-Endowed Chair of semantic information systems. He is the Founding Spokesperson with the Joint Laboratory on Artificial Intelligence and Data Science, a Founding Member with the Research Unit Data Science, Osnabrück University, and the Scientific Director with the Research Department Cooperative and Autonomous Systems, German Research Center for Artificial Intelligence (DFKI). His research interests include artificial intelligence (AI), data science, and integrative AI systems, where his particular research interests include modeling complex data, explainable AI, interpretable learning, machine perception, and semantic interpretation, and relates to applications in complex integrative AI systems, especially robot control and sensor-based AI systems. VOLUME 13, 2025 130637