Training and Validation of Visual Perception Functions for Autonomous Driving with Synthetic Data
Abstract
Diese Arbeit befasst sich mit der Nutzbarkeit synthetisch erzeugter Daten für das Training und die Validierung visueller Wahrnehmungsfunktionen…
Full text
Training and Validation of Visual Perception Functions for Autonomous Driving with Synthetic Data M.Sc. Korbinian Hagn geb. in Bad Tölz Dissertation zur Erlangung des akademischen Grades Doktor der Ingenieurwissenschaften (Dr.-Ing.) der Technischen Fakultät der Christian-Albrechts-Universität zu Kiel eingereicht im Jahr 2023
Kiel Computer Science Series (KCSS) 2024/2 dated 2024-01-18 ISSN 2193-6781 (print version) ISSN 2194-6639 (electronic version) Electronic version, updates, errata available via https://www.informatik.uni-kiel.de/kcss The author can be contacted via [email protected] Published by the Department of Computer Science, Kiel University Multimedia Information Processing Group Please cite as: Ź Hagn, K. Training and Validation of Visual Perception Functions for Autonomous Driving with Synthetic Data Number 2024/2 in Kiel Computer Science Series. Department of Computer Science, 2024. Dissertation, Faculty of Engineering, Kiel University. @book{Hagn2024, author = {Korbinian Hagn}, title = {Training and Validation of Visual Perception Functions for Autonomous Driving with Synthetic Data}, publisher = {Department of Computer Science, Kiel University}, year = {2024}, number = {2024/2}, doi = {10.21941/kcss/2024/2}, series = {Kiel Computer Science Series}, note = {Dissertation, Faculty of Engineering, Kiel University.} } © 2024 by Korbinian Hagn ii
About this Series The Kiel Computer Science Series (KCSS) covers dissertations, habilitation theses, lecture notes, textbooks, surveys, collections, handbooks, etc. written at the Department of Computer Science at Kiel University. It was initiated in 2011 to support authors in the dissemination of their work in electronic and printed form, without restricting their rights to their work. The series provides a unified appearance and aims at high-quality typography. The KCSS is an open access series; all series titles are electronically available free of charge at the department’s website. In addition, authors are encouraged to make printed copies available at a reasonable price, typically with a print-on-demand service. Please visit http://www.informatik.uni-kiel.de/kcss for more information, for instructions how to publish in the KCSS, and for access to all existing publications. iii
1. Gutachter: Prof. Dr.-Ing. Reinhard Koch Christian-Albrechts-Universität zu Kiel Kiel, Germany 2. Gutachter: Prof. Dr. Hanno Gottschalk Technische Universität Berlin Berlin, Germany Datum der mündlichen Prüfung: 11.01.2024 iv
Zusammenfassung Diese Arbeit befasst sich mit der Nutzbarkeit synthetisch erzeugter Daten für das Training und die Validierung visueller Wahrnehmungsfunktionen beim autonomen Fahren. Synthetisch erzeugte Bilder ermöglichen die Erstellung sicherheitskritischer Szenarien, die in der realen Welt potenziell gefährlich zu erfassen sind, und liefern zusätzlich pixelgenaue Ground-Truth Annotationen. Bei der Anwendung synthetischer Bilder auf Wahrnehmungsfunktionen, die anhand von realen Daten trainiert wurden, stellt sich jedoch das Problem der Überbrückung der Domänenlücke. Dies gilt sowohl für das Training als auch für die Validierung mit synthetischen Bildern. Daher muss die Domänenlücke hinreichend verstanden werden, um synthetische Bilder zu erzeugen, die für das Training und die Validierung verwendet werden können. Es wird eine neue Diskrepanzmetrik eingeführt und angewendet, um die Parameter einer realistischen Sensorsimulation zu optimieren und die Domänenlücke effektiv zu reduzieren. Mehrere Einflussfaktoren auf die Domänenlücke werden hierzu untersucht. Die Faktoren, welche die visuelle Erkennung beeinträchtigen, werden vorgestellt und es wird gezeigt, dass sie einen großen Einfluss auf die Erkennbarkeit von Fußgängern haben. Aus den Erkenntnissen dieser Einflussfaktorenanalyse konnte ein Kalibrierungsverfahren einer gewichteten Verlustfunktion entwickelt werden, um die Wahrnehmungsleistung bei realen Fußgänger zu erhöhen. Neue Methoden zur Validierung werden vorgestellt. Die tiefe Variationsdatensynthese und die Klassifizierung von Faktoren, welche die visuelle Erkennung beeinträchtigen. Während erstere Methode durch parametrisierte probabilistische Bilderzeugung nach Wahrnehmungsfehlern sucht, erkennt die letztere Methode Wahrnehmungsfehler durch die Nichtübereinstimmung eines Klassifikators und der tatsächlichen Erkennung. Die Erkenntnisse aus Training und Validierung flossen ein in die Erstellung von zwei synthetischen Validierungsdatensätzen, VALERIE und SynPeDS. v
Abstract This work deals with the usability of synthetically generated data for training and validation of visual perception functions applied in autonomous driving. Synthetically generated images allow the creation of safety critical scenarios which are potentially dangerous to capture in the real-world and additionally deliver pixel perfect ground truth annotations. However, applying synthetic images to perception functions trained on real-world data poses the problem of bridging the domain gap. This is true for both training and validating with synthetic images. Therefore, the domain gap has to be sufficiently understood to generate synthetically images viable to be used for training and validation. A new domain discrepancy metric is introduced and applied to optimize the parameters of a realistic sensor simulation effectively reducing the domain gap. Several influence factors on the domain gap are disentangled. The visual detection impairing factors are introduced and shown to have a high influence on the detectability of pedestrians. Additionally, these factors are used to calibrate a weighting loss function to increase the perception performance on real-world pedestrians. New methods for perception validation are introduced. The deep variational data synthesis and the classification of visual detection impairing factors. While the former method searches for perception faults by parameterized probabilistic image creation, the latter method detects perception faults by the disagreement of a detectability classifier and the actual detection result. The findings of both training and validating were influencing the creation of two synthetic validation datasets, VALERIE and SynPeDS. vii
Acknowledgements First and foremost I want to express my sincere gratitude to my supervisor and mentor Dr.-Ing. Oliver Grau with whom I worked and researched at Intel Labs. I could learn many things and was able to grow as a researcher under his guidance and support. I want to thank Prof. Dr.-Ing. Reinhard Koch for his academic supervision of this thesis and Prof. Dr. Hanno Gottschalk for reviewing this thesis. I want to thank my colleagues Qutub Syed Sha and Peter Nöst for their help and support throughout this work. I want to thank Intel Labs for giving me the opportunity to work and research on this thesis. I want to thank my brother Josef Hagn and my friend Max Büttner for proofreading this thesis and giving valuable feedback. Additionally, I want to thank Roman Raschke, Eduard Engel, Max Bierling, Benjamin Schramm and Rudolf Lehner for their support and distraction when I was in need for one. I want to sincerely thank my mother Silvia Hagn for giving me the chance to always pursue what I set my mind to and always believing in my success. I want to thank Sabina Materak for her relentless support through all ups and downs this thesis has brought. Last, I want to thank Samuel for giving me the greatest joy with his bright smile. ix
List of Tables 3.1 Number of Cityscapes visual detection impairing factor PCA data points not included in the alpha-shape by sequences of the VALERIE dataset for different values of α........ 36 5.1 Characteristics of each sequence generated at 48.18° N, 11.58° E in the VALERIE dataset. ................ 61 5.2 Features added in SynPeDS dataset per data tranche and data pipeline (physical-based rendering (PBR), real-time engine (RT)). (Source: Publication 7 in Chapter 7.7 [SBF+22]) 63 xvii
List of Acronyms Definition of acronyms used in the thesis: CNN convolutional neural network CDF cumulative distribution function DIW detection impairment weighting DNN deep neural network EMD earth movers distance FID Frechét Inception distance FN false negative FP false positive FPPI false positive per image FPN feature pyramid network GAN generative adversarial network IoU intersection over union IS Inception score KID kernel Inception distance laMR log-average miss rate mIoU mean Intersection over Union MR miss rate PCA principal component analysis RPN region proposal network SSD single shot multibox detector SWD sliced Wasserstein distance TP true positive TPR true positive rate xix
Chapter 1 Introduction Autonomous driving is on the verge of becoming an integral part of our every day life. Powered by the ongoing development leaps of artificial intelligence and machine learning methods the dream of fully automated driving comes within reach. These advancements are highly driven by powerful sensor based perception functions. These functions utilize a range of sensors, such as monocular or stereo cameras, LiDAR and RADAR sensors to perceive the real-world and plan their driving accordingly. As part of these achievements the question of safety emerges more and more urgent. Safety is paramount as part of the development in the automotive industry but was not yet the center of attention in the development process of perception functions. But the recent shift towards safe AI brought up a whole field of research with many nuances, laying more focus on the perceptive part of the automated driving stack. While the datasets used for training the perception functions often stem from real-world sensor recorded data, also the trend to utilize synthetically generated sensor data is fueled by the shift of focus on safe AI. 1.1 Motivation Validation of a fully autonomous driving stack in the real-world is a tedious task involving many hundreds of thousands of kilometers to be driven on the street to assure that every aspect, including safety relevant scenarios, is captured [KP16]. Re-evaluation of autonomous driving functions for every new software iteration is therefore not only time-consuming but also expensive. Simulation on the other hand has been used to validate software or hardware in the loop and is an inexpensive and fast alternative. 1
1. Introduction Focusing on the perception task with monocular camera sensors the simulation part is then mainly the generation, i.e., rendering, of synthetic street scenes. One of the main advantages of synthetically rendered imagery is the ability to provide rich meta annotations such as ground truth information and scene descriptions. Additionally, the long-tailed distribution of automotive street scenes can be sampled and validated without any risk to human life. This long-tail distribution samples are essentially the rare accidents or near-accidents that cannot easily and safely be captured in real-world test scenarios. Using this synthetic data for re-training the perception function then helps to reduce these gaps in the perception function. However, there is one major factor hindering the sole usage of synthetic data: The domain gap. This gap is described as the difference of training and validation datasets distribution. This difference can be as subtle as changes in the color distribution of the images, or as prominent as changes of the buildings or persons due to different geolocations. Understanding and overcoming the domain gap is the key to successfully apply synthetic data for training perception functions but also to improve its applicability for validation. 1.2 Research Question The usage of synthetic data is steadily increasing due to the ease of data generation, i.e., image rendering and annotation, and the capabilities of risk-free synthesis of spurious or dangerous events which is especially useful in testing autonomous driving algorithms. The main research question this thesis tries to answer is twofold: First, how can we utilize solely synthetic data for training visual perception functions which are thereafter applied to real-world datasets. This directly leads to the consequence of understanding the differences between synthetic source and real-world target domain, also named as the domain gap. Measuring this domain gap with the right distance or discrepancy measures has a significant influence in understanding the factors defining the domain gap. With the right measures in place one can find and disentangle the factors influencing the domain gap and continuously improve 2
1.3. Publications the synthetic image generation process to bridge the remaining domain gap. Second, how can this improved synthetic data be used for validation of real-world perception functions. With the improved synthetic data the domain gap does not significantly influence the validation of perception functions anymore. Understanding the shortcomings of current state-ofthe-art semantic segmentation and 2D bounding box pedestrian detectors is essential to improve the safety of such algorithms in the real-world. 1.3 Publications This thesis main contributions are presented in academic publications which are attached in Chapter 7. All publications are listed in their chronological appearance with a short summary of their contents in the following: Publication 1: DNN Analysis through Synthetic Data Variation Qutub Syed Sha, Oliver Grau and Korbinian Hagn, Published in 2020 Proceedings ACM Computer Science in Cars Symposium. [SGH20], Chapter 7.1. This contribution introduces the concept of variational data synthesis to analyze and validate visual perception functions. Through parameterization of a generative content rendering system this approach shows how to create validation data given a previously defined validation goal. Results of this paper explain the influence of pedestrian object occlusion rates and visible pixels towards the detection by two state-of-the-art semantic segmentation models. Publication 2: Improved Sensor Model for Realistic Synthetic Data Generation Korbinian Hagn and Oliver Grau, Published in 2021 Proceedings ACM Computer Science in Cars Symposium. [HG21], Chapter 7.2. This paper proposes a method to improve data synthesis methods for automotive datasets by introducing a realistic sensor simulation model. Additionally, the earth movers distance ( EMD ) based domain divergence measure is introduced and utilized as a parameter optimization criteria to adapt the sensor 3
1. Introduction simulation from synthetic to real-world images overall increasing the cross-domain performance of a semantic segmentation model by over 7%. Publication 3: Optimized Data Synthesis for DNN Training and Validation by Sensor Artifact Simulation Korbinian Hagn and Oliver Grau, Published in: Fingscheidt, T., Gottschalk, H., Houben, S. (eds) Deep Neural Networks and Data for Automated Driving. Springer, Cham. [HG22b], Chapter 7.3. This book chapter is a more in-depth continuation of the previous publication about improving the quality of synthetic data generation methods by realistic sensor simulation. Here, additional focus is laid on the extraction of sensor parameters from realworld datasets and application of the extracted parameters on the sensor simulation. Moreover, the EMD domain divergence criteria is thoroughly compared to the well-established Frechét Inception distance ( FID ) domain distance. It was found that the EMD better projects the generalization performance on the target dataset than the FID. Publication 4: A Variational Deep Synthesis Approach for Perception Validation Oliver Grau, Korbinian Hagn and Qutub Syed Sha, Published in: Fingscheidt, T., Gottschalk, H., Houben, S. (eds) Deep Neural Networks and Data for Automated Driving. Springer, Cham. [GHS22], Chapter 7.4. This book chapter first applies the concept of variational deep data synthesis. It introduces the module of probabilistic scene generation, variation of scene parameters and application of the realistic sensor simulation introduced in previous publications. It is shown by the creation of synthetically generated images with high numbers of diverse objects at various illumination settings that we can effectively validate pedestrian detection algorithms on the influence of factors such as, different occlusion objects, additive Gaussian noise, and pedestrian training data distributions. 4
1.3. Publications Publication 5: Validation of pedestrian detectors by classification of visual detection impairing factors Korbinian Hagn and Oliver Grau, Proceedings of the 17th European Conference on Computer Vision Workshops (ECCVW 2022), 2022, Chapter 7.5. In this publication the concept of visual detection impairment factors are introduced. These factors severely influence the detectability of objects, here pedestrians, for object detectors. We showed how these factors can actually influence the detectability. Additionally, we introduced a classification method of pedestrian objects into detectable and non-detectable according to these factors and applied this classification to validate a real-world pedestrian detector finding training data biases, such as ethnicity or age biases. Publication 6: Increasing pedestrian detection performance through weighting of detection impairing factors Korbinian Hagn and Oliver Grau, 2022 Proceedings ACM Computer Science in Cars Symposium., Chapter 7.6. This paper applies the visual detection impairment factors to calibrate an empirical weighting loss of a real-world pedestrian detector. Training pedestrian samples are here weighted according to their extracted impairment factors. We show that this empirical detection impairment weighting loss (DIW loss) improves the state-of-theart on a real-world pedestrian detection benchmark. Publication 7: SynPeDS – A Synthetic Dataset for Pedestrian Detection in Urban Traffic Scenes Thomas Stauner, Frédérik Blank, Michael Fürst, Johannes Günther, Korbinian Hagn, Philipp Heidenreich, Markus Huber, Bastian Knerr, Thomas Schulik, Karl Leiss, 2022 Proceedings ACM Computer Science in Cars Symposium., Chapter 7.7. This publication is the official paper alongside the publication of the KI-Absicherung project’s synthetic dataset, SynPeDS. It contains ground truth for a plethora of visual perception tasks in the automotive domain, such as semantic segmentation, instance segmentation, 2D and 3D bounding boxes, and pose information. We demonstrate the quality of the dataset by semantic segmentation cross-domain generalization 5
2. Background One influence factor is the additive noise in the training data which is a common technique for image augmentation in training to prevent overfitting [Bis95]. Work by [CCN+16; NCC+18] applied different sensor effects which are not adapted to the target dataset to the training set and reported a degradation of the cross-domain performance. However, modeling camera effects to improve the learning with synthetic data for 2D bounding box detection has been proposed by [CSV+18; LLF+20a]. Learning the camera sensor parameters as a style-loss from a real-world dataset extracted from a VGG-16 [SZ15] model’s feature vector and applying these parameters for training with synthetic data for 2D bounding box detection was shown by [CSV+19] to improve the cross-domain generalization performance. However, using a VGG-16 [SZ15] model trained on the ImageNet [DDS+09] dataset poses the problem of optimizing features unrelated with the actual parameters of the target dataset. We propose in Chapter 3 a method to directly optimize sensor artifact simulation parameters on the target data distribution. 2.3.3 Visual Detection Impairing Factors Visual detection impairing factors are defined to be influential on the capability of a detector to reliably detect an object. Some of these factors, such as occlusion rate or contrast, have been previously studied. For example, the authors of [ZBO+16] conclude in their work that the contrast measure of an object ought to have no influence on the object detection capability, however, we are able to show in Chapters 3 and 4 that this factor has indeed a significant influence. The occlusion rate, i.e., the ratio of visible to occluded and visible pixels, as well as the distance of an object to the observer are well acknowledged to have a significant influence on the detection capability [DWS+11; DWS+09]. In Chapter 3 we introduce several additional influence factors and show in section 4.2.2 the actual influence on the task of pedestrian detection. 2.3.4 Training and Sampling Methods In Chapter 3 a novel sample weighting loss for pedestrian detection based on the notion of visual detection impairing factors is introduced. Sampling 12
2.3. Training with Synthetic Data methods are a well-researched topic with the most popular approaches being importance sampling [KM53] and hard example mining [MGE11; SGG16]. These methods do not assign higher weights but specifically oversample harder training samples. Harder samples in the sense of hard example mining [MGE11; SGG16] are determined by the gradient loss of a sample. Boosting algorithms such as AdaBoost [FS97] iteratively weight miss-classified or harder samples higher. For single-stage detectors the focal loss [LGG+17] gained high popularity by putting higher weight to harder samples steered by the cross-entropy loss of each sample. The focal loss, as well as Re-sampling [CBH+02; DGZ17] or cost-sensitive weighting [Tin00; KHB+17] methods all tackle the class-imbalance problem of single-stage detectors. The class-imbalance occurs due to the relevant objects, e.g., pedestrians, only making up for a small part in the image whereas most of the objects are part of the background. Training a detection head with uniform weights for all objects would actually exaggerate the loss of background objects. Two-stage detectors eliminate this problem by applying a RPN , extracting and balancing only relevant objects from the image before sending these proposals to train the detection head with. Specifically targeting the problem of pedestrian detection are the repulsion loss [WXJ+18] and the aggregation loss [ZWB+18] which enforce the bounding box proposal to be very tight on the pedestrian because in automotive scenarios pedestrians are often in crowds and partially occluded by one another. A somewhat different approach is self-paced learning [KPK10]. Here, the learning of easier examples first is encouraged. The training of easier examples first should prevent the model from being stuck in a bad local optimum. Actually, reaching local minima or even the global minimum is discouraged as this would result in overfitting to the training data which leads to bad cross-domain performance on the validation data. Therefore, regularization techniques have been proposed tackling this problem [KPK10; MMX+17; JMZ+15], improving the robustness against the training set bias [RZY+18]. Chapter 3.2 focuses on offline pre-weighting training samples according to visual detection impairment factors ajar to the human visual system 13
2. Background without additional hyperparameter tuning. This approach is in contrast to most of the mentioned online weighting methods including meta-learning approaches [RZY+18; TP12; ADG+16; LUT+17]. 2.4 Validation of Visual Perception Functions The validation of visual perception functions is of high relevance to guarantee a safe autonomous driving functions. Therefore, current automotive safety standards as the ISO 26262 [ISO18] and ISO DIS 21448 [ISO21b] have been developed, defining so-called safety cases for safety-related functions, which form an argument to achieve only a residual risk by the collection of evidences supporting this claim. A plethora of work focusing on the connection of AI safety and these mentioned standards has already been conducted [BGH17; SQC17; GHP+20; SS20; ACP21; BKS+21]. Findings of these methods are already developed into a new safety standard ISO/AWI PAS 8800 [ISO21a]. Other approaches try to further tie these safety arguments to AI methods by measuring the relevance of established metrics in relation to the safety [CNH+18; HSR+20; SKR+21; CKL21]. Another example of such measure is shown in [LGH+21a] by tying the visual detection impairment factor of distance to the object with the safety relevance of detecting this object. These developments brought up the notion of functional insufficiencies [GMB18] in AI models. Our approaches in Chapter 4 specifically target the lack of generalization [SSH20]. A lack of generalization describes the fault that the model could not perform as expected on the target domain. Validating of a machine learning model for such functional insufficiencies is done by the creation and testing of corner cases [AGG+21; BSS+21]. Corner cases are defined by [BBL+19] as "non-predictable relevant object in relevant location". In other words these corner cases are rare events in the tail of the distribution function of automotive street scenes, such as collisions or near-collisions. In our validation data generation approach we find corner cases by probabilistic sampling of a scene parameter space, i.e., corner case samples are a portion of the data samples generated by this method. 14
2.5. Visual Perception Datasets 2.5 Visual Perception Datasets Several datasets for automotive applications, specifically targeting visual perception, have been proposed. Real-world semantic segmentation datasets such as the Audi Autonomous Driving Dataset (A2D2) [GKM+20], Berkeley Deep Driving Dataset (BDD100K) [YCW+20], Cityscapes [COR+16], India Driving Dataset (IDD) [VSN+19], and Mapillary Vistas [NOR+17] are used throughout this work in cross-domain performance experiments. These datasets form a cross-section of many geolocations including Northand South-America, Europe, and Asia. We constrained our work on Europe and Germany and therefore the Cityscapes dataset with the primary German geolocation is of special interest in this thesis. Another Asian automotive dataset to mention is the ApolloScape [HCG+18] dataset. But due to restrictive licensing this dataset is not used. For the task of pedestrian detection additional bounding box annotations for Cityscapes are provided. This dataset is referred to as CityPersons [ZBS17]. Additionally, the EuroCity Persons (ECP) [BKF+19] is considered for our validation application in Chapter 4. However, several more pedestrian detection dataset are available, such as the WiderPerson dataset [ZXW+19], the Caltech [DWS+09] dataset or the ETH [ELV07] dataset. These datasets are very valuable for training and as adaptation target for domain adaptation techniques, e.g., our realistic sensor simulation. But, even though real-world datasets adhere to existing standard workflows of crowdsourcing annotations [LWZ+16; SDF12; KRF+16], human annotators inherently introduce inaccuracies or errors into these ground truth annotations. These inaccuracies or errors are defined as label noise and have been studied [NDR+13] with mitigation strategies already being proposed [RLA+14; LYS+17; JZL+18; Vah17; HMW+18]. With regard to the validation of visual perception functions this label noise poses a great challenge. Differentiating insufficiencies of perception models due to actual functional insufficiencies or due to erroneous labels in the validation data is a tedious process and can hardly be automated. Synthetic data provides a solution to this problem, as ground truth annotations are always accurate. Provided that the implementation does not contain any errors. Several synthetic datasets for automotive perception task have been proposed, for example the Synthia [RSM+16b] dataset, 15
2. Background the Synscapes [WU18] dataset and the GTAV [RVR+16; RHK17] dataset. These datasets are valuable for benchmarking cross-domain adaptation methods but do not allow for additional creation of corner-case data. Here, autonomous driving simulators such as Carla [DRC+17] and the LGSVL simulator [RST+20] prove to be beneficial [LGH+21b; GHA21]. Procedural methods for road generation [PJX+20] can enhance the capabilities of these methods. Some methods try to reduce the remaining domain gap by synthetization of test images through generative approaches [RBK+21]. But similar to label noise the image inconsistencies introduced by these generative models with regard to the corresponding annotation data makes it unfeasible for validation. In our work we utilize synthetic data from the VALERIE and SynPeDS datasets whose scenes are created by variational methods without the need for a full-fledged autonomous driving simulation. 16
Chapter 3 Training with Synthetic Data In this chapter we present our research on training visual perception functions with synthetic data. If we want to achieve a high performing perception function on the target domain it is necessary to overcome and close the domain gap between synthetic source data and real-world target data. Beginning with Section 3.1, we define and validate an appropriate metric to measure the domain gap from synthetic to real-world datasets. We investigate the influence of a realistic sensor simulation on the domain gap and additionally design an optimization method to adopt these sensor simulation parameters for a real-world automotive dataset. Hereby a novel cross-domain generalization metric based on the EMD or Wasserstein-2 distance is used. This measure is based on a redefinition of the crossdomain per-image performance as a domain gap proxy. Following the investigations of the sensor simulation influence on the domain gap, we disentangle several additional factors on their respective influence on the domain gap. These factors are the number of graphical assets used to create a synthetic dataset, the number of frames to train a perception function and the structural similarities of real-world and synthetic data based on the segmentation heatmaps. In Section 3.2 the detection impairing factors are introduced. These factors represent influential factors that are responsible for impairing the visual detection of pedestrians. Visual detection impairing factors are defined to be influential on the detection abilities by the human vision system. In this section we show how to utilize these impairing factors to detect missing training data in the synthetic source domain by comparing the distributions of visual detection impairing factors of a synthetic and real-world dataset. Furthermore, we show how one can 17
3. Training with Synthetic Data boost the detection performance of a pedestrian detector by weighting training samples according to a visual detection impairing factor calibrated loss function. In Chapter 4 we utilize these factors to validate pedestrian detectors and detect training data biases. 3.1 Overcoming the Domain Gap When training deep neural network ( DNN )- or CNN -based visual perception functions with a synthetic source dataset it is advantageous that the training samples have the same or at least a similar distribution as the target domain distribution. The difference in distributions of source and target domain is known as domain gap. A domain gap from source and target domain will lead to detection performance degradation. Therefore, when using synthetic training data, the sample distribution of the target domain should be modeled as close as possible. However, the target data distribution is not easily available and moreover a complex combination of individual distributions, such as lighting, textures, and even more subtle factors like the number of unique persons in the dataset. Without the possibility of direct comparison of these sample distributions one can instead measure the domain gap by indirect measures. One of these measures is the domain generalization distance, i.e., the distance of a performance measure when trained on the synthetic domain compared to when trained on the target real-world domain. Such measures and metrics are essential if we want to be able to model the synthetic data distribution to increasingly resemble the real-world distribution. Additional, these metrics allow disentangling of certain influence factors on the domain gap, i.e., understanding their influence on the model performance, as we can show in Subsection 3.1.3. 3.1.1 Measuring the domain gap Measuring the domain gap for visual perception learning is based to a large extend on the advances of domain adaptation techniques [THS+17; LCW+15; GL15; THS+18]. Some predominantly used measures in this area 18
3.1. Overcoming the Domain Gap of research are the IS [SGZ+16], the FID [HRU+17], and the KID [BSA+18]. All of these measures are based on either extracting feature vectors ( FID , KID ) or softmax scores from the InceptionV3 [SVI+16] network trained on the ImageNet [DDS+09] dataset. Hereby, the relevant features or scores are extracted from both, the source dataset and the target dataset, and subsequently a distance metric between those is calculated. However, these measures are, while being useful to train generative networks such as generative adversarial network ( GAN ), non-predictive of the actual performance [HG21]. In this work we showed for the task of semantic segmentation that the target domain mIoU performance of a CNN trained with a dataset of lower FID value, i.e. potentially smaller domain gap, is worse than a model trained on a dataset with a higher FID value. From these observations we constructed a new performance based domain discrepancy measure with a close link to the actual cross-domain performance. This new measure is based on the EMD of the target to target domain performance compared to the source to target performance. The performance itself is measured as mIoU on a per-image basis. While this measure is related to the overall cross-domain performance measurement of the mIoU , the EMD -based metric can capture the performance of a perception function at harder and easier to classify images as these would be averaged out in a measure on the average of whole target dataset. This measure is explicitly designed to be a discrepancy measure and not a distance measure of the domain gap. A distance is inherently symmetric and would not capture the difference of training on the source domain and evaluating on the target domain compared to training on the target domain and evaluating on the source domain. A discrepancy is therefore necessary if we want to have our measure strongly tied to the actual domain generalization performance of source to target domain. We prove the hypotheses of the viability that our EMD -based domain discrepancy measure on the per-image performance measure can be used as a proxy for the domain gap. Therefore, we conducted an experiment by training an ensemble of DeeplabV3+ semantic segmentation models with a ResNet101 backbone on the Cityscapes dataset. The weights of each model were initialized randomly, and each model was evaluated after the training on the Cityscapes validation dataset on a per-image basis. The results are mIoU histograms over the Cityscapes validation dataset. Next, 19
3. Training with Synthetic Data we compare the histograms by calculating the cumulative distribution function ( CDF )s of each histogram. On the resulting CDF s we apply the 2-sample Kolmogorov-Smirnov test to each possible pair of the ensemble, and we get a minimum p´value ą 0.95. Resulting, we cannot reject the null-hypotheses that the per-image based performance histogram is not a good proxy of the training dataset. 0 20 40 60 80 100 mIoU [%] 0.0 0.2 0.4 0.6 0.8 1.0 Cumulative count Model 6 Model 5 Model 4 Model 3 Model 2 Model 1 Figure 3.1. CDF of ensembles of DeeplabV3+ models trained on Cityscapes and evaluated on the corresponding validation set. (Source: Publication 2 in Chapter 7.2 [HG21]) In Figure 3.1 one can see the CDF s of the resulting per-image performance histograms. This experiment proves that only the training dataset is responsible for the shape of the per-image based mIoU performance. 3.1.2 Realistic Sensor Simulation The domain gap is the main cause of reduced performance on the target dataset when training with synthetic data, therefore it has proven beneficial to adapt the source synthetic dataset to the target dataset reducing this exact gap. The research area which deals with this task specifically is domain adaptation. While there is a plethora of domain adaptation strategies based on generative models showing impressive results on the visual 20
3.1. Overcoming the Domain Gap adaptation of images on a target domain, there is a major drawback if one chooses to use these adapted images for validating a model. This drawback is the non-deterministic behavior of these generative domain adaptation models on the input images which leads to inconsistencies between the visually adapted image and the original ground truth. An often seen example is the visible Mercedes-Star of the ego car in the Cityscapes dataset that is hallucinated into the synthetic image when adopted to this dataset [RAK22; HTP+18; PEZ+20]. If we use these adopted images to validate a perception model for detection faults, it is hard to impossible to track an error back to the real source of the problem due to the non-matching input and ground truth. However, for validation with synthetic data it is still important to minimize the domain gap of the training dataset to the target validation dataset, as with a high domain gap between these datasets one cannot safely determine if a perception fault was caused due to the domain difference, e.g. different geolocation, or due to actual errors in the detection model. Therefore, understanding the influential factors of the domain gap is a key component to improve validation with synthetic images as well as improving the cross-domain performance when training with synthetic datasets. If these influential factors are well enough understood we can then model, i.e., simulate, these factors and apply them on our synthetic dataset deterministically. One of these influential factors we have identified are the sensors used to capture the real-world imagery data. More specifically the sensor lens artifacts of cameras which are inherently observable in real-world datasets as can be seen for example in an image of the A2D2 dataset in Figure 3.2. Synthetically generated images on the other hand often simulate a pinhole camera model [Stu14] that does not realistically simulate any sensor lens artifacts. In real-world datasets we could observe several sensor lens artifacts such as blur, chromatic aberration, and additive sensor noise. Furthermore, in datasets such as the Cityscapes dataset, the images were captured in a high dynamic range format and then subsequently tonemapped and gamma corrected to an 8-bit integer RGB low dynamic range to be displayable on most modern screens. Tone-mapping and gamma correction have a major influence on the style and even the usability of the image for perception as under exposed areas in an image such as persons in the shadow on the sidewalk can be highlighted. Unfortunately, detailed 21
3. Training with Synthetic Data the Cityscapes dataset with DeeplabV3+ models trained on subsets of the VALERIE and SynPeDS datasets. Figure 3.7 shows the generalization results with the respective cumulative frame counts that were used to train each segmentation model. 0 20000 40000 60000 80000 100000 120000 140000 cumulative frame count 0 10 20 30 40 50 60 70 80 mIoU [%] Tranche 1 Tranche 2 Tranche 3 Tranche 4 Tranche 5 SynPeDS Sequence_0050 Sequence_0060 cityscapes baseline Figure 3.7. Number of training frames per SynPeDS (blue) tranche or VALERIE (red) sequence and overall generalization performance on the Cityscapes dataset. While no model comes close to reaching the baseline performance of 82.34%, the cross-domain performance with Sequences of the VALERIE reach higher mIoU values with far fewer image frames than the SynPeDS dataset. The diversity in the VALERIE dataset continuously improved which is evident by the increasing cross-domain performance, whereas the performance of the VALERIE model even deceased for tranche 4. In tranche 4 a significant pedestrian object distribution bias was introduced into the dataset which we detected with our validation methods as can be seen in section 4.1. Overall it is clearly visible in this result that only increasing the frame count by reiterating the same assets in the scenes is no viable strategy to increase the cross-domain generalization performance. Structural Similarities from Segmentation Heatmaps We investigated the influence of the person asset diversity of a synthetic dataset and found it to be influential on the cross-domain generalization. 28
3.1. Overcoming the Domain Gap (a) VALERIE Sequence 0050 (b) VALERIE Sequence 0058 (c) VALERIE Sequence 0060 (d) VALERIE Sequence 0050-0060 (e) Cityscapes Figure 3.8. Cumulative semantic segmentation heatmaps derived from different VALERIE sequences and from the Cityscapes dataset. Understanding the influence of the placement of objects in the scene is the 29
3. Training with Synthetic Data next logical step to explain the remaining domain gap. This difference of object placements between source and target dataset can be understood as the structural similarities of the respective scenes in the datasets. One method to measure this structural similarities is to compare the semantic ground truth per class of the synthetic and real-world domain. To do this, the number of occurrences for every class individually and per pixel in the ground truth are accumulated across a dataset. Next, the accumulated results are sub-sampled and binned to a 100x100 grid. Thereafter, these results are normalized to the range [ 0,1 ] . The overall result is a 2D histogram of normalized class occurrences for every class in the dataset. Figure 3.8 exemplary depicts these histograms on the progression of VALERIE Sequences ((a) to (d)) compared to the target Cityscapes (e) heatmap for the person class. By visual analysis of the histograms one can determine first differences in the person distribution between early VALERIE sequences and the Cityscapes dataset. The sequences 0050 and 0060 show clear person silhouettes around the main horizontal distribution. This occurs when the number of frames is quite low when calculating these histograms. Comparing the result of combined sequences 0050 to 0060 with the Cityscapes heatmap it is evident that the latter distribution is vertically more spread out around the main horizontal distribution and there is a focus or center on the right side of the histogram. This distribution can be explained by the way the images were recorded. The Cityscapes dataset was recorded from a car in right-hand driving countries where most persons are visible on the nearer right side of the car camera and persons on the opposite side of the road are more often occluded by oncoming cars. To actually measure a difference between the histograms or heatmaps we are utilizing the sliced Wasserstein distance ( SWD ) [BRP+15]. The SWD is a sample based form of the EMD for 2and higher-dimensional data. To compute this distance the source and target 2D histograms are projected by a random sampled vector to a 1-dimensional vector each. Next, the EMD is calculated on the projected vector. This process is repeated for 1000 iterations. The final distance result is the average of all projected distance results. We apply these calculations to the VALERIE and SynPeDS dataset with the target dataset Cityscapes. The SWD results with the cross-domain generalization performance on the person class can be seen in Figure 3.9. 30
3.1. Overcoming the Domain Gap 6.0 6.5 7.0 7.5 8.0 8.5 9.0 9.5 sliced wasserstein distance 0 10 20 30 40 50 60 70 80 person mIoU [%] Tranche 2 Tranche 4 Tranche 5 SynPeDS Sequence_0050 Sequence_0054 Sequence_0058 Sequence_0060 Figure 3.9. Sliced Wasserstein distance per SynPeDS (blue) tranche or VALERIE (red) sequence and overall generalization performance on the Cityscapes dataset. Lower SWD results indicate a greater similarity between source and target person heatmaps, i.e., person placements. The sequences 0058 and 0060 of the VALERIE dataset achieve the highest cross-domain generalization performance and have the lowest SWD to the target dataset. Another interesting observation can be made for these sequences, as the 0058 sequence has a lower SWD as the 0060 sequence. Comparing these sequences again visually in Figure 3.8 one can see that there is a higher person distribution from the middle to right on the main horizontal distribution compared to the distribution of the 0060 sequence. Making the 0058 distribution more similar to the target Cityscapes distribution. For the SynPeDS dataset the results are differing from earlier tranches to later, more mature tranches. The tranche 2 shows the lowest mIoU and highest SWD values, but while tranche 5 has a higher SWD than tranche 4 the cross-domain performance is higher than the latter. With the result from the number of assets we know that tranche 5 has more pedestrian assets than tranche 4 and achieves therefore a better generalization even though the person distribution is dissimilar to the target dataset. The overall dataset of SynPeDS reaches the lowest SWD for this dataset but with higher performance than the VALERIE sequence 0054, again due to a higher number of assets in the former dataset. Concluding, 31
3. Training with Synthetic Data the SWD of the class distribution can be used to measure the similarities of source and target datasets and helps to better understand differences in the datasets. 3.2 Visual Detection Impairing Factors We identified factors influencing the domain gap when training with synthetic data and applying the trained model to real data. Our work also investigates the reasons and factors that impede a successful detection of an object. We show that by understanding these factors we can leverage these to further measure domain distances, increase the detection performance and implement validation strategies to identify missing training data. The factors are termed visual detection impairing factors as they hinder understanding, i.e., impede the detection of an object. We previously focused on the task of semantic segmentation which gave us the capabilities to investigate on the overall scene structure differences. Following, we focus on the task of object detection, and due to the autonomous driving setting on the task of pedestrian detection. Results from the investigations on visual detection impairing factors are published in Publications 5 in Chapter 7.5 and 6 in Chapter 7.6. The visual detection impairing factors we are considering in our work are visualized in Figure 3.10. The factors can be grouped in four major categories. The first category is the location of an object in the image. This category is covered by the bounding box coordinates center locations cx and cy . The second category is the size of an object in the image. For this category we use the width and height dimensions of the bounding box. Additionally, we use the distance to the camera measured in [m] which strongly correlates to the last factor, the actual number of visible pixel of an object in the image. The third category is the occlusion of an object with the single factor occlusion rate. The occlusion rate is the quotient of visible to visible plus occluded pixels of an object. The fourth and last category is the contrast of an object. The contrast measures the visual difference of an object to its background. Low contrast values, e.g., dark clothed person standing in the shadow, makes it harder to detect a person. In our work we use three different formulations of contrast measures. 32
3.2. Visual Detection Impairing Factors cx cy w h+ (a) Bounding Box Coordinates 5m 5·104pxl 10m 1·104pxl (b) Distance and Visible Pixels (c) Occlusion (d) Contrast Measures Figure 3.10. The potential detection performance impairing factors we consider in this work: (a) bounding box coordinates ( ocx , ocy , oh , ow ), (b) distance and number of visible pixels of a pedestrian ( od , ovp ), (c) rate of occlusion ( oocl ), (d) contrast of a pedestrian (red) to its background (blue) calculated by the full pedestrian silhouette ( oc f ull ), segment wise ( ocmean ) and edge wise ( ocedge ). (Source: Publication 5 in Chapter 7.5 [HG23]) The first is calculated by averaging the RGB values of a person instance and calculating the Euclidean distance to the average of the surrounding background pixels. The second contrast measure segments the person into 12 segments and calculates the Euclidean distance to the adjacent background pixels with the results being averaged. The third and last contrast measure considers only a small pixel border of the person instance and calculates the Euclidean distance to the surrounding background. A more in depth explanation and calculation of the individual factors can be found in the Publications 5 in Chapter 7.5 and 6 in Chapter 7.6. 33
3. Training with Synthetic Data In the following subsection we show how we can utilize these factors to find differences in source and target datasets and reason on missing data that has to be added to the source training dataset to improve crossdomain generalization. Subsequently, we show how these factors allow us to improve a pedestrian detector by steering the training loss to focus on harder to detect samples according to these visual impairing factors. 3.2.1 Missing Training Data Detection The previously defined impairment factors can be used to describe a dataset and then use this dataset description to compare it with the description of another dataset. In our experiment we begin by extracting each individual factor for each person in the synthetic VALERIE dataset and the real-world CityPersons dataset. The CityPersons dataset is an extension of the CityScapes dataset and delivers additional bounding box annotations for persons. With the resulting visual impairing factors per person for both datasets one could already calculate differences by just comparing these factors. But due to the factors being closely correlated, e.g., distance and number of visible pixels, it is useful to reduce the redundancy by application of a principal component analysis. After reducing the dimensionality of the impairing factors with the PCA , we can visualize both datasets on a 2D plot by extraction of the first and second major PCA component and plot every pedestrian onto a scatter plot. For subsequent detection of dissimilarities between source and target dataset we can calculate an area around our source synthetic dataset points and find target real-world dataset points that are not enclosed by this area. These non-enclosed data points are persons with specific visual detection impairing factors that are not included in our source dataset. To calculate the surrounding area we use alpha shapes [EM94]. Alpha shapes are a special form of Delaunay triangulation [Del+34] where the α parameter denotes the maximum radius of a circle in the triangulation around each data point. In other words, the distance between points if there will be an edge or not. For low values of α the shape is more rugged and follows the points on the borders more closely, whereas for high values of α the shape will be similar to the complex hull of the points with the actual complex hull being α=8 . We visualized the PCA results for α values α= 0.7, α= 1.0 and α=8 with 34
3.2. Visual Detection Impairing Factors impairing factors extracted for VALERIE sequences 0057, 0057-0058 and 0057-0060 and CityPersons in Figure 3.11. (a) Seq. 0057 α=0.7 (b) Seq. 0057 α=1.0 (c) Seq. 0057 α=8 (d) Seq. 0057-0058 α=0.7 (e) Seq. 0057-0058 α=1.0 (f) Seq. 0057-0058 α=8 (g) Seq. 0057-0060 α=0.7 (h) Seq. 0057-0060 α=1.0 (i) Seq. 0057-0060 α=8 Figure 3.11. Alpha-shapes generated by the PCA of visual detection impairing factors from three different VALERIE sequence combinations and three different α values. Orange indicate VALERIE data points, blue indicate Cityscapes data point inside the alpha-shape, and red indicate Cityscapes data points outside the shape. As mentioned for lower values of α the shape is more rugged and 35
3. Training with Synthetic Data more missing data points, visualized as red points, are detected. For the complex hull shape only few outliers are detected and there is a considerable amount of blue CityPersons data points being in the shape but not covered by orange VALERIE data points. Further, for later sequences, e.g., 0057-0060, the number of missing data points is fewer even with lower values of α . To highlight the progress from VALERIE sequences and better understand the right choice of the α parameter we calculated the number of missing data points for different values of α for each of the three sequences. The results are shown in Table 3.1. Table 3.1. Number of Cityscapes visual detection impairing factor PCA data points not included in the alpha-shape by sequences of the VALERIE dataset for different values of α. α Sequence 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1.0 8 0057 576 278 172 137 106 92 70 51 44 40 6 0057-0058 365 179 121 100 72 52 38 30 28 28 3 0057-0060 302 143 97 73 49 40 35 30 23 19 3 With values of αă 0.7 the alpha shape will omit too many points further off from the main distribution, creating a very dense shape around the points and therefore many outliers will be detected. Whereas for an α=8 , i.e. the complex hull, the shape does not capture the actual shape spanned by the points and too few outliers will be detected. We found reasonable α values are in the interval 0.7 ďαď 1.0, but a visual inspection of the actual produced shapes will nonetheless prove useful. In our work, especially in the creation of the VALERIE dataset, this tool was used to detect missing data points which were subsequently added for each iteration of the sequences. This is visible in the number of outliers in Table 3.1 which continuously decreases for sequences 0057-0058 and 0057-0060 for every αvalue being used. 3.2.2 Detection Impairment Weighting Loss As previously shown, the visual detection impairment factors are useful to detect differences in datasets. In this subsection we introduce another 36
3.2. Visual Detection Impairing Factors approach on how to utilize these factors and boost the pedestrian detection performance of an object detector by steering the training loss towards harder to detect samples according to these factors. The detailed approach is defined in Publication 6 in Chapter 7.6. The general idea behind our approach is to calibrate the training loss of a pedestrian detector to put more attention towards objects which are harder to detected according to the visual detection impairing factor of these samples. The workflow of this approach is depicted in Figure 3.12. Dval Synth Synthetic image data Detector Detection Filter f(.) Extract Detection Impairing Factors Ωm Sample weighting Dtrain CP Real-world image data Dw CP filtered objects o∈Ωm oocl · · · od o10.8 · · · 30 . . .. . ..... . . oO0.6 · · · 34 training.json filename class (u) bbox (v) weight aachen 0.png person {0.5,0.3,0.2,0.3}0.8 aachen 1.png person {0.2,0.1,0.15,0.25}0.1 . . .. . .. . .. . . Figure 3.12. Generation of the training weights for pedestrian objects of the realworld CityPersons dataset. (Source: Publication 6 in Chapter 7.6 [HG23]) First, we have to define which samples are harder to detect. Therefore, we begin by inferencing on a synthetic dataset with a pedestrian detector which was pre-trained on this synthetic domain. For the synthetic data we use the sequences 00057 and 0058 from the VALERIE dataset. In parallel to the detection the impairing factors per pedestrian object are extracted from each inferenced image. The detection results and the corresponding impairing factors are forwarded to the detection filter f( . ) . This filter evaluates the detections with the ground truth annotations and dismisses all true positive detected pedestrians so that only the missed detections 37
4. Validation of Visual Perception Functions Figure 4.2. Example of scene parameter variation, in this case the time of the day is varied, causing dramatic changes in the scene illumination and according contrast variations. (Source: Publication 4 in Chapter 7.4 [GHS22]) sent to the validation flow control. The flow control sets the next scene and parameter variations to be generated in the data synthesis block. These variations can now either follow the ranges given by the scenario description or the previous scene is reused and recreated with a different set of parameters. This choice depends on the result of the evaluation metric. For example if a pedestrian detector evaluates for FN s, i.e. miss detections, the flow control chooses to recreate the scene under different lighting conditions to check if the miss detections were caused by difficult lighting conditions. Some examples of such time of day parameter variations are given in Figure 4.2. The process ends if all parameter combinations are exhausted. The evaluation metric results and the corresponding parameter variations per image are then carefully examined by the validation engineer to pinpoint reasons for the perception flaws as is demonstrated in Chapter 4.1.2. 44
4.1. Variational Deep Data Synthesis for Perception Validation 4.1.1 Generation of Synthetic Validation Data The generation of synthetic data in our variational deep data synthesis approach consists of three major components. The first component is the probabilistic scene generator. It first generates the ground layout of the automotive scene, i.e., the layout for streets, sidewalk and buildings. Next, the generator places three-dimensional objects or assets from the database on the previously defined layout. For example, cars are placed on the street or on parking spots, whereas persons are placed on the sidewalk and on the street. As the name of the component suggests the whole scene, including the object placement, is probabilistically generated, i.e., sampled, from a given range of parameters. These parameters are set by the validation flow control and define for example the min and max width of the street and sidewalks or the probability of a person being placed on the street. The second component is the parameter variation generation. These parameters include, among others, the time of day, and geolocation settings. By changing the time of day parameter the scene lighting can change drastically. The variation generation is therefore the main instrument of the validation flow control to search for parameter combinations that can cause perception flaws. The last component is the sensor & environment simulation. By utilizing Blender 6, which allows the importing, editing, and scripting of 3D content, the predefined scenes are then rendered with the physically based renderer Cycles. Subsequently, the realistic sensor simulation model derived in Chapter 3.1.2 is applied to the rendered image to recreate the sensor impression of a real-world recording. Additionally, metadata from the rendering process is extracted to generate the necessary data for an accurate ground truth annotation for a range of perception tasks. These task include semantic segmentation, instance segmentation, 2D & 3D bounding box detection and depth estimation from monocular images. The continuous development of this synthetization method eventually resulted in the generation of the VALERIE dataset. This dataset is described in more detail in Chapter 5.1. 45
4. Validation of Visual Perception Functions 4.1.2 Validation Results As the focus of this validation approach is on visual perception functions, we validated perception function models for the task of semantic segmentation and 2D bounding box pedestrian detection. The validated segmentation model is the DeeplabV3+ model with a ResNet101 backbone pre-trained on the ImageNet dataset. This model is fine-tuned on a range of real-world and synthetic datasets and validated utilizing our variational deep data synthesis method. For the 2D bounding box pedestrian detection task, the SSD model fine-tuned on tranche 3 of the synthetic SynPeDS dataset is validated. The feature extractor of the SSD model is a ResNet50 backbone pre-trained on the ImageNet dataset. To measure the model performance on the semantic segmentation task, the mIoU is calculated. The performance of the pedestrian detection is measured by calculating the true positive rate ( TPR ). The TPR , also known as the sensitivity, is calculated as the quotient of TP s over the sum of TP s and FN s. This measure allows filtering out all predictions on images with missed pedestrian objects, i.e., with TPRă1. In our work several influence factors on the perception function are validated. Beginning with the influence of the pedestrian distribution in the training data on the perception generalization. Therefore, we use the sequence 0058 of the VALERIE dataset, generated by our variational synthesis method, for training the segmentation model and compare the dataset’s pedestrian distributions to other synthetic datasets and to the target dataset Cityscapes. Next, the influence on the segmentation performance on images with additive Gaussian noise is evaluated. Here, the realistic sensor simulation as part of the data synthesis component allows us to simulate images with increasing additive Gaussian noise. Last, we validate the pedestrian detector and the semantic segmentation model on the influence of different person occluder objects. The parameterization of the scene layout allows exchanging objects in front of the pedestrian in a scene and enables us to evaluate the respective influence of each object on the detection performance. 46
4.1. Variational Deep Data Synthesis for Perception Validation (a) (b) (c) Figure 4.3. Pedestrian distribution over horizontal angle and distance. (a): Cityscapes. (b): SynPeDS Tranche 3. (c): VALERIE synthetic data. (Source: Publication 4 in Chapter 7.4 [GHS22]) Influence of Object Distributions The object, i.e., person distribution of a dataset, was found in Chapter 3.1.3 to have a considerable influence on the cross-domain generalization quality when training with synthetic data. In Figure 4.3 the spatial distribution of pedestrians in the Cityscapes (a), SynPeDS tranche 3 (b), and VALERIE sequence 0058 (c) are shown. These visualizations represent histograms of persons placed in the datasets images at a distance to the camera in 47
4. Validation of Visual Perception Functions [m] and at a horizontal viewing angle of ´ 30 ˝ to 30 ˝ , corresponding to the left and right side of the image. As previously discussed it is of advantage to recreate or match the spatial distribution of objects in the source synthetic domain to the one of the target real-world domain. The Cityscapes dataset is considered again as the target dataset. In this dataset the distributions of pedestrians show a uniform distribution with slight tendency to a positive horizontal angle. Looking into the Cityscapes dataset one notices the right-handed driving in all of these images with oncoming traffic occasionally occluding pedestrians on the left side, explaining the slight tendency of more visible pedestrians on the right side or positive horizontal angles. The SynPeDS tranche 3 distribution shows two sharp parallel lines where most of the datasets pedestrians are located on. This indicates a strong potential bias in the pedestrian distribution. We can strengthen this hypothesis by visual evaluation of perception results as shown in the last subsection of this chapter. The VALERIE Sequence 0058 distribution on the other hand shows a more uniform person distribution which better resembles the Cityscapes distribution. For both synthetic datasets one further interesting observation can be made. The number of pedestrians at further away distances of around 50 m is significantly higher than the one of the Cityscapes dataset. This can be explained by the manual annotations of the real-world dataset. While human annotators have a hard time drawing bounding boxes for persons above such distances, due to the small size of the person, synthetically generated images deliver pixel perfect ground truth information at any distance and size. Influence of Noise on Detection Performance The influence of noise on the detection performance is evaluated by applying Gaussian noise with increasing variance on the input image, followed by subsequent segmentation inference and performance evaluation. In this experiment the DeeplabV3+ segmentation model is trained on the realworld datasets A2D2,Cityscapes, and on the sequence 0058 of the synthetic VALERIE dataset. The variance σ of the Gaussian noise is continuously increased in steps of 1 in the range σ2P[ 0,20 ] . At every step the segmentation performance, mIoU , is measured for each of the three trained models. The resulting graph for this experiment is shown in Figure 4.4. The 48
4.1. Variational Deep Data Synthesis for Perception Validation 0 2 4 6 8 10 12 14 16 18 20 noise variance 0 10 20 30 40 50 60 mIoU [%] DeeplabV3+ trained on A2D2 DeeplabV3+ trained on Cityscapes DeeplabV3+ trained on Synthetic data Figure 4.4. Top: mIoU performance decreases with increasing noise variance. Bottom (left to right): segmentation maps with increasing noise variance σ2P { 0,10,20 } , image pixels xiP[ 0,255 ] . (Source: Publication 4 in Chapter 7.4 [GHS22]) images below the plot are segmentation results of the Cityscapes trained model at σ2 values of { 0,10, 20 } . While the VALERIE and Cityscapes trained models increase in performance for small values of σ2 the model trained on A2D2 continuously declines in performance. The initial increase of the performance can be explained that both Cityscapes and the VALERIE exhibit noise in their respective training images similar to these values of Gaussian noise. For the A2D2 dataset these values seem to mismatch with the noise in the training images as no initial increase of performance can be observed. The source of decline in performance is clearly visible in the prediction image with σ2= 20. Here, object boundaries begin to fray and the segmentation prediction starts to smear out accross classes. 49
4. Validation of Visual Perception Functions Figure 4.5. Scene with variation of occluding objects. Left: 2D bounding box detection. Right: semantic segmentation (Source: Publication 4 in Chapter 7.4 [GHS22]) Influence of Occluder Objects To understand the influence of the occluder object on the pedestrian detection performance we utilized our scene generator to create images of a pedestrian on the street with different occluders in front of it. Next, the SSD and the DeeplabV3+ semantic segmentation model are used to compute the inference on these images. The resulting predictions are shown in Figure 4.5. For all three variations of the occluder object, the SSD predicts two 50
4.2. Classification of Visual Detection Impairing Factors partial bounding boxes. One of those predictions always covers the upper half of the person with low confidence, whereas the other prediction covers a smaller part of the upper half with higher confidence in two of the three images. Even though the occlusions of the person differ significantly, the visibility of the upper half is enough to correctly determine the presence of a pedestrian. Another observation is that the boxes do not cover the spread out arms of the person, which hints towards missing pedestrian poses in the training data or even missing bounding box anchor scales in the SSD model. In the segmentation prediction the conclusions differ understandably. While the person is correctly determined to be present in the images and correctly classified as person in all images, the occluder in front of the person has only little influence on the prediction. Again the spread out arms of the person are not correctly detected, emphasizing the hypothesis of missing training data because both detection models were trained on the same synthetic dataset with no such poses included. Additionally, another detection fault is visible in the segmentation prediction of the person. The road below the pedestrian is in all images classified as sidewalk even though the remaining road is correctly classified. This suggests another training data bias. Here, the training data did not include enough pedestrians standing on the road, but instead the pedestrians were placed on the sidewalk in most of the training images. The segmentation model was trained on the tranche 3 of the SynPeDS dataset. Re-examining the pedestrian distributions in this dataset in Figure 4.3 b), strengthens this hypothesis. The sharp person distribution stems from pedestrians only placed on the sidewalk and none placed on the streets. 4.2 Classification of Visual Detection Impairing Factors In this section, we describe a validation approach for 2D bounding box detectors based on the previously in Chapter 3.2 introduced visual detection impairing factors. This method allows detecting training data biases of pedestrian detectors, such as missing ethnicity or poses. Furthermore, in this result we present additional evidence for the actual object detection performance influence of the visual detection impairing factors. This 51
4. Validation of Visual Perception Functions method is described in full detail in Publication 5 in Chapter 7.5. 4.2.1 Data Bias Detection By definition, the visual detection impairing factors allow separating the set of objects in the dataset, i.e. pedestrians, into detectable and non-detectable subsets. Because the impairment factors define the border between detectable and non-detectable pedestrians. We utilize this observation and train a binary classifier to learn this exact border function. Applying this classifier on the impairing factors extracted from a validation dataset and comparing the classification result with the predictions from a pedestrian detector under test enable us to validate the training data bias of the detector. Especially, if the classifier is carefully trained and predicts the pedestrian to be detectable, but the detector does not detect this pedestrian, then this is a strong evidence on an existing data bias of the training data. Starting by training of the binary classifier. Figure 4.6 depicts the training workflow. To create a training dataset for the classifier we first distinguish the persons of a synthetic training dataset into the detectable and the non-detectable subsets. This is done by predicting the pedestrian bounding boxes in this synthetic training dataset with a pedestrian detector trained on the same synthetic domain. For the synthetic training dataset Dtrain Synth , we use a subset of the VALERIE sequence 0060. The pedestrian detector is trained on the VALERIE sequence 0058. The pedestrian detector is a Cascade R-CNN model with a HRNet backbone. To generate the necessary data to train the binary classifier we extract the visual detection impairment factors from the training dataset. The visual impairment factors used in this method are introduced in Chapter 3.2. In the next step, the prediction results and the extracted impairment factors are forwarded to the Accumulate Results component. Here, the split between detectable and non-detectable class is carried out. Every pedestrian in the ground truth without a successful prediction, i.e. IoUă 0.5, is allocated to the non-detectable class ( s= 0) or to the detectable class ( s= 1) otherwise. The results are stored as ground truth data S and the extracted impairment factors on the other hand are stored into the training dataset Ω as input to train the classifier. With the input Ω and ground truth S , the binary classifier is subse52
4.2. Classification of Visual Detection Impairing Factors Dtrain Synth Pedestrian Detector Accumulate Results Extract Detection Impairing Factors Ω,S Classifier g(.) Train Pedestrian objects o∈Ω detectability s∈ S #oocl · · · ods o10.8 · · · 30 1 . . .. . ..... . .. . . oO0.6 · · · 34 0 Figure 4.6. Training a classifier to distinguish between detectable and nondetectable pedestrian objects. (Source: Publication 5 in Chapter 7.5 [GHS22]) quently trained. The binary classifier is a multi layer perceptron with 4 hidden layers and 20 nodes per layer. As loss function the binary-cross entropy is used. The training data is split into 80% training and 20% validation data. After training, the resulting classifier reaches a F1 score of 0.93 on the validation data. With the successful trained classifier, it is then used to validate a realworld pedestrian detector. The validation workflow is shown in Figure 4.7. Again, we start with a synthetic dataset Dval Synth . This dataset is the remaining subset of the VALERIE sequence 0060 which is now used for validation. The pedestrian detector under test predicts on this dataset while in parallel the visual detection impairing factors are extracted. The detectors under test are two Cascade R-CNN models with HRNet backbones each. The first model is trained on the CityPersons and the second one on the EuroCity Persons dataset. In the next step, the inference and extraction results are accumulated and forwarded to the previously trained binary classifier. The classification results, the pedestrian prediction results and additional metadata Dval Meta from the input dataset Dval Synth are then stored in the result dataset Dvalcl Synth . The additional metadata originates from the asset database used to synthesize the validation dataset and includes additional information per pedestrian object, such as the geolocation, the pose, the 53
5. Validation Datasets cant amount of image and ground truth data was generated. This data was gathered and stored into the respective VALERIE sequences which are used throughout this work. With increased understanding and disentanglement of the relevant influence factors on the synthetic to real-world domain gap the gained knowledge was used to improve the quality of the data generation process. Each sequence therefore corresponds to a stage in the development of the data synthesis pipeline and exhibits distinctive features differentiating them from one another. The most distinctive features of each sequence are described in Table 5.1. With increasing development efforts and increased understanding of the intricacies of influence factors on the domain gap, the structural complexity from sequence to sequence was increased. Structural complexity can hereby be understood as the amalgamation of scene layout, object placements, asset diversity, environment, and sensor simulation. As was shown in Chapter 3.1.3, the best domain generalization performance on the Cityscapes dataset was achieved with the combined dataset of sequences 0054 to 0060. We found that due to the much higher number of frames and lower diversity of this sequence the cross-domain performance after training is reduced. The lower diversity is owed to the fact that this sequence reuses scenes and re-renders them with different lighting conditions. While these scenes are important to be incorporated in a validation dataset they would add a heavy training bias if they are directly included into the training dataset. A viable option for the usage of the 0062 sequence in training would be to sub-sample the data. Exemplary images from the sequences 0058 and 0060 of the VALERIE dataset are shown in Figure 5.1. As described in the Chapter 4.1.1, the dataset includes ground truth annotation data for several perception tasks from monocular images. These tasks include semantic segmentation, instance segmentation, 2D bounding box detection, 3D bounding box detection, and depth estimation. Additionally, for every sequence and every image a scene description file includes all meta information from the data synthesis process. These meta information include the time of day, the definition, direction and position of assets, the definition of the camera and its position in the scene, the visual detection impairing factors per pedestrian, and the overall geolocation position. Especially the meta information, as shown with the visual 60
5.1. VALERIE Table 5.1. Characteristics of each sequence generated at 48.18° N, 11.58° E in the VALERIE dataset. Seq. Characteristics Frames Cameras Scenes per Scene 0050 Fixed street layout 1000 1 2002 crossings Night scenes 0052 Similar to 0050 1800 1 300 Time 5:30 to 21:00 (GMT+1) 0054 2 crossings 480 1 480Few traffic signs Time 10:05 (GMT+1) 0057 2 crossings 1000 1 1000 7 facades Varying street width Time 6:00-20:00 (GMT+1) 0058 Similar to 0057 1395 1 1395 Time 6:30-20:00 (GMT+1) 0059 2 crossings 1306 2 653Varying street width Time 10:30 (GMT +1) 0060 T-junction 1430 2 715Varying street width Random ego-vehicle looking direction 0062 Similar to 0058 10855 2 700Time 7:00-10:12 (GMT+1) South looking direction perception impairing factors, are important for a validation method to better understand and draw conclusions from found perception faults. The VALERIE dataset is planned to be released to the research community to increase the efforts for validating visual perception functions and hopefully increase the safety in autonomous driving as a whole. 61
5. Validation Datasets Figure 5.1. Our fully parameterizable generation pipeline allows rendering pedestrians at any size, occlusion, time of day, and distance to the camera. (Source: Publication 5 in Chapter 7.6 [HG23]) 5.2 SynPeDS The SynPeDS dataset is one of the key results of the KI-Absicherung project. This project and the resulting dataset were a collaborative effort of 28 partners in technology and academia in Germany to set the foundation of research on safeguarding autonomous driving perception functions. The dataset rendering was done by three different partners: Mackevision, 62
5.2. SynPeDS BIT-TS, and Intel. Whereas the first partner implemented the image rendering process in the real-time game engine Unreal [Gam], the other partners implemented the rendering as physically based rendering pipeline. BIT-TS implemented their physically based rendering with the Blender [Fou] render engine whereas Intel used the OSPray [Int] render engine. Because these pipelines where developed independently by each data producing partner, some implementations of ground truth annotations and meta information differ. The dataset is split into 9 different tranches with distinctive features added for each tranche similar to the sequences in the VALERIE dataset. The added features per tranche and for each rendering pipeline are listed in Table 5.2. Table 5.2. Features added in SynPeDS dataset per data tranche and data pipeline (physical-based rendering (PBR), real-time engine (RT)). (Source: Publication 7 in Chapter 7.7 [SBF+22]) Tranche New Features PBR RT 1, 2, 3 Preparation for large-scale data production ˆ ˆ 4 Frame-to-frame variations ˆ ˆ Meta information on AssetIDs ˆ ˆ Bodypart segmentation ˆ Procedural sun model ˆ 5 Sensor noise as post-processing ˆ Procedural clouds model ˆ Ground truth for pose estimation ˆ Meta information on occlusion ˆ 6 Environmental effects: wetness and sun glare ˆ Out-of-distribution assets ˆ Variations of camera sensor parameters ˆ 7 Camera and LiDAR sensor models using PBR with OSPRay and ˆ different LiDAR sensor parameters Meta information on AnimationID ˆ Environmental effects: fog, vignetting ˆ 8 Night scenes with artificial light ˆ 9 Specific user requests for contrast or material ˆ 63
5. Validation Datasets As with the VALERIE dataset, the SynPeDS dataset could take advantage from the findings and disentanglement of domain gap factors. Improvements from tranche 1 to the subsequent tranches are for example the addition of pedestrian assets due to findings of the cross-domain generalization in Chapter 3.1.3 or the application of our sensor simulation on the image data due to findings from Chapter 3.1.2. The analysis of the domain generalization performance and the influence of pedestrian assets on the person class domain generalization capability have been published in conjunction to this dataset in Publication 7 in Chapter 7.7. The SynPeDS dataset includes ground truth annotation data for a range of visual perception tasks. These tasks include instance segmentation, semantic segmentation, bodypart segmentation, 2d bounding box detection, 3d bounding box detection, depth estimation, and pose estimation. Unlike the VALERIE dataset these tasks are not limited to monocular camera image perception. Starting at tranche 7 a realistic LiDAR simulation was added and can be used as input for these perception tasks. The meta information per image is similar to the meta information of the VALERIE dataset but varies strongly throughout the tranches and is very limited for earlier tranches, i.e. tranche 1 to 3. The dataset is available through an industry friendly licensing model. This allows not only academic researchers but also the industry to use the dataset for research and especially development of new perception validation methods. This dataset is the first synthetic dataset with such an open licensing model. 64
Chapter 6 Conclusions This work focuses on the usage of synthetically generated realistic sensor images for training and validation of visual perception functions for autonomous driving applications. The main contributions are divided into three topics: Training with synthetic images, validation with synthetic images, and the characterization of synthetic datasets for training and validation. Training a visual perception function solely on synthetically generated imagery and applying this function on real-world data poses the problem of how to overcome the domain gap. We investigated several domain distance and discrepancy measures and found that a major shortcoming is that these measures do not predict the actual target domain, i.e., generalization, performance. Therefore, we introduced a new earth movers distance ( EMD ) performance based domain discrepancy measure which directly correlates to the expected performance on the target dataset. Furthermore, we showed that we can utilize this EMD measure as optimization loss to optimize the parameters of a realistic sensor simulation on the synthetic data and achieve a reduced domain discrepancy to the target Cityscapes dataset. With help of the metadata that comes along the synthetically generated imagery we were able to entangle several influence factors on the remaining domain gap. These factors include, the number of training assets in the synthetic data, the number of training frames, and the scene structure measured through the sliced Wasserstein distance ( SWD ) of segmentation heatmaps. Additionally, we introduced the visual detection impairment factors. We showed through several experiments that these factors have a significant influence on the detection performance of a pedestrian detector. Extraction of the impairment factors from source and target datasets allows us for example to detect missing training data in 65
6. Conclusions the synthetic images. Last, we utilize the impairment factors to calibrate the training loss of a real-world pedestrian detector on harder to detect samples. With this approach we improve state-of-the-art performance on pedestrian detection. The main tasks for validation of a perception are semantic segmentation and 2D bounding box detection, especially with a focus on pedestrian detection. In our work we introduce a method named variational deep data synthesis for perception validation. This method introduces a fully parameterizable data generation and validation approach for a perception function under test. Synthetic images are hereby generated with regard to the findings of the domain gap factors. Hereby we can effectively reduce the domain discrepancy and guarantee that found perception flaws stem from the perception model or from the training data but not from mismatched domains. We showed how the pedestrian placement during training, the image noise, and the type of the occluding object during inference influence the performance of segmentation and 2D detection functions. Furthermore, a new method to detect training data biases of pedestrian detectors is derived. This method utilizes the visual detection impairment factors and trains a classifier to distinguish between detectable and non-detectable pedestrians by their respective impairment factors. If classification and the actual detection result differ then a training data bias is present. Further, we show that the visual detection impairment factors have significant influence on the detection performance except for the pedestrian placement in the image. The preceding findings of training and validating visual perception functions for autonomous driving influenced the creation process of two synthetic validation datasets. First, the VALERIE dataset which is a product of the data generation process of the deep variational data synthesis method. This dataset consists of physically based rendered images in 8 sequences which were continuously improved by findings from the domain distance entanglement results. Second, the SynPeDS dataset which resulted from the collaborative KI-Absicherung project. This dataset consists of 9 tranches with images both rendered in a physically based renderer and a real-time engine. Similar to the first dataset, this dataset benefited from the findings of the domain distance entanglement experiments and was continuously improved in quality. 66
Overall, this thesis improves the understanding in using synthetic data for training as well as for validation of visual perception functions. Yet, there are several domain gap influence factors still to be found and analyzed. For validation methods this work is part of the initial ignition of a branch of research with high relevancy for safe autonomous perception functions and safe AI in general. 67
Chapter 7 Publications 7.1 Publication 1 DNN Analysis through Synthetic Data Variation Qutub Syed Sha, Oliver Grau and Korbinian Hagn Published in 2020 Proceedings ACM Computer Science in Cars Symposium. [SGH20] Reprinted with permission from Qutub Sayed Sha DOI: 10.1145/3385958.3430479 69
DNN Analysis through Synthetic Data Variation CSCS ’20, December 2, 2020, Feldkirchen, Germany Table 2: Deeplab and Detectrons results on both the pedestrians Models Dark cloth pedestrian Bright cloth pedestrian Deeplab (pACC%) 89.27 (min: 0.071; max: 100) 78.82 (min: 0.065; max: 100) Detectron (pACC%) 87.31 (min: 0.01; max: 100) 92.04 (min: 0.024; max: 100) Figure 4: Performance of Deeplab (top) and dectectron (bottom) over pixel area. adversarial attack on the image but in a controlled, precise and parametric environment. •Dark clothed pedestrian : From various plots shown above, a few of the outlier samples are shown below in the Table 3 and Table 5. This table helps understand the massive variation of DNN performance with little or no noticeable tweaks; it causes the network to fluctuate its output. The pedestrians in these images are 10% occluded. All images have the same parameter variation except the sun time difference and material of the pedestrian. We could see the contrast variation of pACC, as shown in Table 4. Sample 2 shows a tremendous drop from 97.85% to 56.97%. Figure 5: Pixel area based on metadata - occlusion rates, Deeplab (top) and dectectron (bottom). •Bright clothed pedestrian : The images shown in Table 5, and the corresponding Table 6 with inference values explain the stability factor of detectron’s output. They show fewer variations, especially if the pictures contain bright clothed pedestrians. Detectron as opposed to deeplab is invariant towards the shiny and non-shiny surface of the pedestrian’s cloth material. The above examples are considered to explain the arbitrary features and parameters influencing the neural networks with no clear distinctionacknow. Its is hard to see patterns with our rendered images and strong inferences can only be confirmed by iterating over all 7. Publications 76
CSCS ’20, December 2, 2020, Feldkirchen, Germany Qutub Syed Sha, Oliver Grau, and Korbinian Hagn Figure 6: distance evaluation based on pedestrian detection rate, Deeplab (top) and dectectron (bottom). major color variations. This investigation requires a huge dataset to study these parameters and its dependency over the stability of the network to categorize it as a safe algorithm at critical situations. It is merely to motivate the reader that these parameters do play a important role on the detection rates. These parameters are need to be studied in detailed fashion to gain confidence in neural network based systems. 6.4 Discussion All these parameters (pixel area, occlusion rate and distance) considered for the evaluation show case a non linear impact on the detection rate. The models behave quite contrast to each other. The detection rate in detectron model (fig.4) show a downward trend as the pixel area increases, which is not the case when the humans tries to detect the pedestrian. The pedestrian fitted with dark cloths are detected with higher rate by deeplab model. Similarly, pedestrian fitted with bright cloths are detected with higher rate by detectron model. The occlusion plots in fig. 6 follows a clear downward trend as the occlusion rate increases, which is intuitive and as expected. We also see a major dependency on the texture of Table 3: Variation of dark clothed pedestrian: shiny and nonshiny materials sample 1 sample 2 Shiny surface material Non shiny surface material Table 4: Variation of detection rate over changes in parameters: Deeplab and Detectrons results on dark clothed pedestrians for images shown in Table 3 Models sample 1 sample 2 Shiny surface material Deeplab (pACC%) 78.82 67.27 Detectron (pACC%) 98.36 97.85 Non shiny surface material Deeplab (pACC%) 77.15 65.2 Detectron (pACC%) 95.68 56.97 Table 5: Variation of bright clothed pedestrian: shiny and non-shiny materials sample1 sample2 Shiny surface material non-shiny surface material 7.1. Publication 1 77
DNN Analysis through Synthetic Data Variation CSCS ’20, December 2, 2020, Feldkirchen, Germany Figure 7: Pixel area based on metadata - occlusion rates, Deeplab (top) and dectectron (bottom). Table 6: Variation of detection rate over changes in parameters: Deeplab and Detectrons results on dark clothed pedestrians for images shown in Table 5 Models sample 1 sample 2 Shiny surface material Deeplab (pACC%) 76.2 70.3 Detectron (pACC%) 98.52 98.54 Non shiny surface material Deeplab (pACC%) 75.72 70.69 Detectron (pACC%) 97.64 98.51 the material dawned by pedestrians. These factors might be dependent to other internal parameters like surface material, irradiance, and shadows. Hence we see these non linear and contrasting behavior between the models. DNNs are powerful in generalizing the objects in a scene. It can also be invariable to the position of the object. From our experiments we see that they are quite sensitive to the parameters (occlusion, pixel area, surface material, irradiance, and shadows) that are not generally taken into consideration. With these plots, we can conclude that the trained models are not stable and not invariant to scene parameters. The analysis of the cause of these effect would require a deeper investigation and more parameter variation runs. 7 CONCLUSIONS AND OUTLOOK We have presented a new validation approach using computational validation based on synthetic generative data. The approach allows a flexible description of parameters to be varied and to systematically test parameters in our unified validation parameter space. We consider our system as a valuable tool for the analysis of insufficiencies in perception functions. Although the synthetic content presented in this paper is currently not photorealistic, the evaluation based on pACC is covering a good range and seemed to have good distinguishing power for analysis of the chosen network algorithms (Deeplab+Detectron). The fact that the pACC values are on average slightly lower than on real test data (Cityscape) indicate a domain-gap between the real and synthetic data. We plan to investigate if this is due to rendering fidelity or to other effects (like scene complexity and variation). We get a more detailed analysis on the DNNs by closing in the domain gap of real and virtual images. The evaluated experiments show the potential to identify insufficiencies and to relate it to scene or other properties. Moving forward we are implementing more sophisticated analysis methods based on more complex meta-data relations. Further, we will improve the computational efficiency of our validation engine. The latter aspect also includes concepts to enable incremental rendering and separation of computational stages, like computation of sensor errors independent from the rendering, and similar. Both planned extensions will enhance the tool for validation, which allows to explore a high number of parameters and, therefore, provide wider coverage and higher degree of automation for the validation of DNNs for perception functions. ACKNOWLEDGMENTS The research leading to these results was partially funded by the German Federal Ministry for Economic Affairs and Energy within the project “KI Absicherung – Safe AI for Automated Driving”. REFERENCES Wilhelm Burger and Matthew J. Barth. 1995. Virtual Reality for Enhanced Computer Vision. Springer US, Boston, MA, 247–257. https://doi.org/10.1007/978-0-38734904-6_19 Holger Caesar, Jasper Uijlings, and Vittorio Ferrari. 2016. COCO-Stuff: Thing and Stuff Classes in Context. arXiv:1612.03716 [cs.CV] Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. 2017. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE transactions on pattern analysis and machine intelligence 40, 4 (2017), 834–848. Pegasus Consortium. 2020. Pegasus project home page. https://www.pegasusprojekt.de/en/home. Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. 2016. The Cityscapes Dataset for Semantic Urban Scene Understanding. arXiv:1604.01685 [cs.CV] 7. Publications 78
CSCS ’20, December 2, 2020, Feldkirchen, Germany Qutub Syed Sha, Oliver Grau, and Korbinian Hagn Alexey Dosovitskiy, Germán Ros, Felipe Codevilla, Antonio López, and Vladlen Koltun. 2017. CARLA: An Open Urban Driving Simulator. CoRR abs/1711.03938 (2017). arXiv:1711.03938 http://arxiv.org/abs/1711.03938 Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2015. Deep Residual Learning for Image Recognition. arXiv:1512.03385 [cs.CV] Nidhi Kalra and Susan M. Paddock. 2016. Driving to Safety: How Many Miles of Driving Would It Take to Demonstrate Autonomous Vehicle Reliability? https: //www.rand.org/pubs/research_reports/RR1478.html Santa Monica, CA; RAND Corporation. Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. 2016. Feature Pyramid Networks for Object Detection. arXiv:1612.03144 [cs.CV] M. Magnor, O. Grau, O. Sorkine-Hornung, and C. Theobalt (Eds.). 2015. Digital Representations of the Real World: How to Capture, Model, and Render Visual Reality. A K Peters CRC Press. T. Menzel, G. Bagschik, and M. Maurer. 2018. Scenarios for Development, Test and Validation of Automated Vehicles. In 2018 IEEE Intelligent Vehicles Symposium (IV). 1821–1827. https://doi.org/10.1109/IVS.2018.8500406 Shai Shalev-Shwartz, Shaked Shammah, and Amnon Shashua. 2017. On a Formal Model of Safe and Scalable Self-driving Cars. Arxiv (2017). https://arxiv.org/abs/1708.06374 Evan Shelhamer, Jonathan Long, and Trevor Darrell. 2016. Fully Convolutional Networks for Semantic Segmentation. arXiv:1605.06211 [cs.CV] W. Wachenfeld and H. Winner. 2015. Die Freigabe des autonomen Fahrens. In Autonomes Fahren. Technische, rechtliche und gesellschaftliche Aspekte. Springer Vieweg. Josie Wernecke. 1994. The Inventor Mentor: Programming Object-Oriented 3D Graphics with Open Inventor. Addison-Wesley. Yuxin Wu, Alexander Kirillov, Francisco Massa, Wan-Yen Lo, and Ross Girshick. 2019. Detectron2. https://github.com/facebookresearch/detectron2. 7.1. Publication 1 79
7. Publications 7.2 Publication 2 Improved Sensor Model for Realistic Synthetic Data Generation Korbinian Hagn and Oliver Grau Published in 2021 Proceedings ACM Computer Science in Cars Symposium. [HG21] DOI: 10.1145/3488904.3493383 80
Improved Sensor Model for Realistic Synthetic Data Generation Korbinian Hagn Oliver Grau [email protected] oliver[email protected] Intel Deutschland GmbH Neubiberg, Bayern, Germany ABSTRACT Synthetic, i.e., computer generated-imagery (CGI) data is a key component for training and validating deep-learning-based perceptive functions due to its ability to simulate rare cases, avoidance of privacy issues and easy generation of huge datasets with pixel accurate ground-truth data. Recent simulation and rendering engines simulate already a wealth of realistic optical effects, but are mainly focused on the human perception system. But, perceptive functions require realistic images modeled with sensor artifacts as close as possible towards the sensor the training data has been recorded with. In this paper we propose a method to improve the data synthesis by introducing a more realistic sensor model that implements a number of sensor and lens artifacts. We further propose a Wasserstein distance (earth mover’s distance, EMD) based domain divergence measure and use it as minimization criterion to adapt the parameters of our sensor artifact simulation from synthetic to real images. With the optimized sensor parameters applied to the synthetic images for training, the mIoU of a semantic segmentation network ( DeeplabV3+ ) solely trained on synthetic images is increased from 40.36% to 47.63%. CCS CONCEPTS •Computing methodologies →Image manipulation ;Image segmentation;Cross-validation. KEYWORDS datasets, neural networks, sensor simulation, image synthesis, domain adaptation ACM Reference Format: Korbinian Hagn and Oliver Grau. 2021. Improved Sensor Model for Realistic Synthetic Data Generation. In Computer Science in Cars Symposium (CSCS ’21), November 30, 2021, Ingolstadt, Germany. ACM, New York, NY, USA, 9 pages. https://doi.org/10.1145/3488904.3493383 Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from [email protected]. CSCS ’21, November 30, 2021, Ingolstadt, Germany ©2021 Copyright held by the owner/author(s). Publication rights licensed to ACM. ACM ISBN 978-1-4503-9139-9/21/11...$15.00 https://doi.org/10.1145/3488904.3493383 1 INTRODUCTION Training of deep neural networks (DNNs) is increasingly resorting towards computer generated imagery ( CGI ) due to its mitigation of certain issues. First, synthetic data can avoid privacy issues found with recordings of members of the public and on the other hand can automatically produce ground truth data at higher quality and reliability than costly manually labeled data. Moreover, simulations allow synthesis of rare cases and the systematic variation and explanation of critical constellations [Syed Sha et al . 2020] –– a requirement for validation of products targeting safety-critical applications, such as automated driving. Here, the creation of corner cases and scenarios which otherwise could not be recorded in a real-world scenario without endangering other traffic participants is the key argument for validation of perceptive AI with synthetic images. Despite the advantages in CGI methods, training and validation with synthetic images still has challenges: While training with these images does not guarantee a similar performance on real-world images and validation is only valid if one can make sure that the found weaknesses do not stem from the distribution shift from the real to the synthetic image domain. To measure and mitigate this domain shift, metrics have been introduced with various applications in the field of domain adaptation or transfer learning. In domain adaptation these metrics are applied to train generative adversarial network ( GAN ) to adapt on a target feature space [Pan et al . 2009] or to recreate the visual properties of a dataset [Salimans et al . 2016]. On the other hand, the problem of training and validation with synthetic imagery as the source domain is directly related to the predictive performance of a perception algorithm on the target data, which these kinds of metrics struggle to capture [Ravuri and Vinyals 2019]. Additionally, applications of domain adaptation methods often resort to specifically trained DNNs which adapt one domain to the other and therefore add an extra layer of complexity and uncontrollability, whereas the creation of images via a synthesis process allows to understand domain distance influence factors more directly. Camera-recorded images inherently show visual imperfections or artifacts, such as sensor noise, blur, chromatic aberration or image saturation, as can be seen in an image example from the A2D2 [Geyer et al . 2020] dataset in Figure 1 . CGI methods, on the other hand, are usually based on idealized models, i.e., pinhole camera model free of sensor artifacts. In this paper we present an approach to decrease the domain discrepancy of synthetic to real-world imagery for perceptive DNNs by realistically modelling sensor lens artifacts to increase the viability of CGI for training and cross domain validation. 7.2. Publication 2 81
CSCS ’21, November 30, 2021, Ingolstadt, Germany Hagn and Grau Figure 1: Real-world images (here A2D2) exhibit sensor lens artifacts which have to be closely modelled by an image synthetization process to decrease the domain distance of synthetic to real-world datasets to make them viable for training and validation. Therefore, a new interpretation of the domain discrepancy by generalization of the distance of two datasets by the per image performance comparison over a dataset utilizing the Wasserstein or earth mover’s distance (EMD) is presented. This domain discrepancy measure is then used to optimize the proposed sensor lens artifact parametrization and effectively minimize the domain discrepancy between a synthetic and a real-world dataset. We demonstrate how this model is able to decrease the EMD domain discrepancy in an optimization of the parameters as depicted in Figure 5. Further, we compare the domain discrepancies of random, of extracted parameters from the target real-world dataset and of optimized parameters. Additionally, we show that the model trained with optimized sensor artifacts decreases the domain divergence on a wealth of real-world and synthetic datasets compared to a model trained without sensor artifacts. 2 RELATED WORKS This work is related to two areas: synthetic data generation for training and validation and domain distance measures, as used in the field of domain adaptation. Sensor Simulation for Synthetic Image training. The use of synthesized data for development and validation is an accepted technique and has been also suggested for computer vision applications (e.g., [Burger and Barth 1995]). Recently, specifically for the domain of driving scenarios, games engines have been adapted [Dosovitskiy et al. 2017; Richter et al. 2017]. Although games engines provide a good starting point to simulate environments, they usually only offer a closed rendering setup with many trade-offs balancing between real-time constraints and a subjectively good visual appearance for human observers. Specifically the lighting computation in this rendering pipelines are in-transparent. Therefore it does not produce a physically correct imagery, instead only a fixed rendering quality (as a function of lighting computation and tone mapping), resulting in low dynamic range (LDR) output images (typically 8bit per RGB color channel). Recently, physical-based rendering techniques have been applied to the generation of data for training and validation, like Synscapes [Wrenninge and Unger 2018]. For our work we use a dataset in high-dynamic range (HDR) resolution created with the physicalbased Blender Cycles renderer 1 . We implemented a customized tone mapping to 8bit per color channel and sensor simulation, as described in the next section. While there is great interest in understanding the domain distance in the area of domain adaptation via generative strategies, i.e., GAN s, there has been little research regarding sensor artifact influence on training and validation with synthetic images. Other works [da Costa et al . 2016; Nazaré et al . 2018] add different kinds of sensor noise to their training set and reported a degradation of performance, compared to a model trained with no noise in the training set, due to higher task complexity. Adding noise in training is a common technique for image augmentation and can be seen as regularization technique [Bishop 1995] to prevent overfitting. Our task of modeling sensor artifacts for synthetic images extracted from camera images is not aiming to improve generalization through random noise, but to tune the parameters of our sensor model to closely replicate the real-world images and improve generalization on the target data. First results of modeling camera effects to improve synthetic data learning on the task of bounding box detection have been proposed by [Carlson et al . 2018; Liu et al . 2020]. Lin et al. [Liu et al . 2020] additionally state that generalization is an asymmetric measure which should be considered when comparing with symmetric dataset distance measures from literature. Furthermore, Carlson et al. [Carlson et al . 2019] learned sensor artifact parameters from a real-world dataset and applied the learned parameters of their noise sources as image augmentation during training with synthetic data on the task of bounding box detection. However, contrasting our approach, they apply their optimization as style loss on a latent feature vector extracted from a VGG-16 network trained on ImageNet and evaluate the performance on the task of 2D object detection. Domain Distance Measures. A key challenge in domain adaptation approaches is the expression of a difference measure between datasets, also called domain shift. A number of methods were developed to mitigate this shift, for example via unsupervised domain adaptation (e.g., see [Ganin and Lempitsky 2015; Long et al . 2015; Tsai et al . 2018; Tzeng et al . 2017]). Similar, [Hoffman et al . 2018] utilize GAN s to stylize synthetic source domain images to real-world target domain images. However, these approaches of domain adaptation require another DNN to adapt from source to target domain with little to no tuning capability after the adaptation function has been learned, whereas we want to learn parameters of a sensor simulation, eliminating the need for a complex network. To measure the domain shift or domain distance, measures based on the classification output of a discriminator network have been proposed [Salimans et al . 2016] based on the InceptionV3 topology [Szegedy et al . 2016] trained on the ImageNet dataset. The work of [Bińkowski et al . 2018; Heusel et al . 2017] relies on features extracted from the InceptionV3 network to tune domain adaptation approaches. However, these metrics are not predictive of the classification performance when the data is applied as augmentation 1Provided by a project partner (https://www.bit-ts.com/) 7. Publications 82
Improved Sensor Model for Realistic Synthetic Data Generation CSCS ’21, November 30, 2021, Ingolstadt, Germany or replacement when training a discriminator [Ravuri and Vinyals 2019]. Therefore, to measure performance directly, it is essential to train with the adapted or synthetic data and validate on the target data, i.e., cross-evaluation as done by [Ros et al . 2016; Saleh et al . 2018; Wrenninge and Unger 2018]. In our work we build on this baseline and introduce a measure that mitigates weaknesses of a metric over the dataset as explained in Section 3.2. Performance Metrics. The mean intersection over union ( mIoU ) is a widely used performance metric for benchmarking semantic segmentation [Cordts et al . 2016; Varma et al . 2019]. Adaptations and improvements of the mIoU have been proposed which set more weight on the segmentation contour as in [Fernandez-Moral et al . 2018; Rezatofighi et al . 2019]. Performance metrics as the mIoU are computed over the whole validation dataset, i.e., the whole confusion matrix, but there are propositions to apply the mIoU calculation on a per-image basis [Csurka et al. 2013]. A per image comparison mitigates several shortcomings of a single evaluation metric on the whole dataset when used for comparison of classifiers on the same dataset. First, one can distinguish multimodal and unimodal distributions, i.e., strong classification on one half and weak classification on the other half of a set can lead to the same mean as an average classification on all samples. Second, unimodal distributions with the same mean but different shape are also indiscernible under a single dataset averaged metric. This justification led to our choice of a per-image-based mIoU metric as it allows for deeper investigations which are especially helpful when one wants to understand the characteristics that increase or decrease a domain discrepancy. 3 METHODS Given a synthetic ( CGI ) dataset of urban street scenes, our goal is to decrease the domain gap to a real-world dataset for semantic segmentation by realistic sensor artifact simulation. Therefore, we systematically analyze the image sensor noise of the target realworld dataset and use these extracted parametrization for our sensor artifact simulation. To compare our source synthetic dataset with the real-world dataset we contrive a novel per image performancebased metric to measure the generalization distance between the two of them. We utilize a DeeplabV3+ [Chen et al . 2018] semantic segmentation model with a ResNet101 [He et al . 2016] backbone to train and evaluate on the different datasets throughout this paper. Utilizing this discrepancy measure as optimization criteria, we adapt the parameters with random and extracted parameter starting points of our sensor artifact simulation and therefore further decrease the domain distance between synthetic and real-world images. 3.1 Sensor Simulation We implemented a simple sensor model with the principle blocks depicted in Figure 2: The module expects images in linear RGB space. Rendering engines like Blender Cycles 2 can provide these images as results in OpenEXR format3. 2https://www.blender.org/ 3https://www.openexr.com/ We simulate a simple model by applying sensor noise, as added Gaussian Noise (zero mean, variance is a free parameter),chromatic aberration, and blur followed by a simple exposure control (linear tone mapping), finished by non-linear Gamma correction. First, we apply blur by a simple box filter with filter size FxF and a chromatic aberration (CA). The CA is approximated using radial distortions (k1, second order), as defined in OpenCV. The CA is implemented as a per channel (red, green, blue) variation of the k1 radial distortion, i.e., we introduce an incremental parameter ca that affects the radial distortions: k 1 (blue)=−ca ; k 1 (дreen)= 0; k 1 (red)= +ca . As next step we apply Gaussian noise to the input image. Applying a linear function, the pixel values are then mapped and rounded to the target output byte range [ 0 .. 255 ] by applying a linear function. The two parameters of the linear mapping are determined by a histogram evaluation of the input RGB values of image, imitating an auto exposure of a real camera. In our experiments we have set it to saturate 2% (initially) of the brightest pixel values, as these are usually values of very high brightness, like sky or even the sun. Values below the minimum or above the set maximum are mapped to 0or 255 respectively. In the last step we apply gamma correction to achieve the final processed synthetic image: x=(˜ x)γ(1) The parameter γ is an approximation of the sensor non-linear mapping function. For media applications this is usually γ= 2 . 2for the sRGB color space. However, for industrial cameras this is not yet standardized and some vendors do not reveal it 4 . We therefore estimate the parameter as an approximation. Figure 3 depicts the difference of an image with and without simulated sensor artifacts. 3.2 Dataset Discrepancy Measure Our proposed discrepancy measures quantifies per image performance between same topologies trained on different datasets but evaluated on the same dataset. Considering the task of semantic segmentation we chose the mIoU as our base performance metric. We then modify the original mIoU calculation to be calculated on a per image basis instead of the whole evaluated dataset. Following, we introduce the Wasserstein-1 or EMD metric as our domain discrepancy measure. The measure is calculated on the per-image mIoU distribution of two classifiers, one trained on the source domain, the other trained on the target domain, evaluating the test set of the target domain. The mIoU is defined as follows: mIoU =1 SÕ s∈S TPs TPs+FPs+FNs ×100%,(2) With TPs , FNs and FPs being the true-positives, false-negatives, and false-positives of the sth class. Here, S={ 0 , 1 , ..., S− 1 } with S= 11, as we use the 11 classes as can be seen in Table 1. These classes are common in the real and synthetic datasets we considered for evaluation of the sensor artifact parameter optimization outcome in 4.3. 4The providers of the Cityscapes dataset don’t document the exact mapping. 7.2. Publication 2 83
CSCS ’21, November 30, 2021, Ingolstadt, Germany Hagn and Grau Linear Mapping, Rounding + Noise Generator Measurement of Histogram Range Gamma Correction Blur, Chromatic Aberration synthetic image x0 ˜x ∈ [0, ..., 255] ”real” image x Figure 2: Sensor artifact simulation. (a) (b) Figure 3: Left (a): Synthetic images without lens artifacts. Right (b): Applied sensor lens artifacts, including exposure control. Modifying the mIoU to calculate a distribution over the perimage IoU, it takes the following form: IoUn=1 SÕ s∈S TPs,n TPs,n+FPs,n+FNs,n ×100%,(3) where n denotes the n−th image in the evaluated dataset. As typically done with the mIoU , IoUn is measured in %. As we want to compare the distributions of the per image IoU values originating from the evaluation of the target test set by two DNNs, one trained on the source domain and one trained on the target domain, therefore, we apply the Wasserstein distance. The Wasserstein distance as an optimal mass transport metric from [Kolouri et al . 2017] is originally defined for density distributions p and q where inf denotes the infinum, i.e., the lowest transportation cost, Γ(p,q) denotes all joint distributions π , i.e., transportation maps, for (X,Y) which have the marginals p and q as follows: Wr(p,q)=(inf π∈Γ(p,q)∫R×R |X−Y|rdπ)1/r.(4) This distance formulation can be reformulated an was shown by [Ramdas et al. 2017] to be equivalent to the following: Wr(p,q)=(∫∞ −∞ |P(t) − Q(t)|r)1/rdt.(5) Here P and Q denote the respective cumulative distribution function (CDF) of p and q. In our application we only calculate the empirical distributions of p and q, further simplifying the formulation in this case to the function of the order statistics: Wr(ˆ p,ˆ q)=( n Õ i=1 ∥ˆ pi−ˆ qi∥r)1/r,(6) where ˆ p and ˆ q are the empirical distributions of the marginals p and q sorted in ascending order. With r= 1and equal weight distributions we get the EMD which, in other words, measures the area between the respective CDFs with L1as ground distance. We consider a sample size of at least 100 to be sufficient for the EMD calculation to be valid. In our application the Wasserstein-1 distance is calculated as the distance of the distribution ˆ p , i.e., a model trained on the source domain and evaluated on the target domain, to the distribution ˆ q , i.e., a model trained on the target domain evaluated on the target domain. When compared to distance measures such as the Fréchet inception distance ( FID ), which is a special case of the Wasserstein-2 distance when ˆ p and ˆ q in 6 are normally distributed and r= 2, that is by definition a symmetric distance, our measure is a divergence, i.e., the distance from source dataset A to target dataset B can be different to the distance from target dataset B to source dataset A. A divergence also reflects the characteristic of a classifier having different generalization distance when trained on dataset A and evaluated on dataset B or the other way around. Because the ground measure of the signatures, i.e., the IoU per image, is bound to 0 ≤IoUn≤ 100, the EMD measure is then bounded 0 ≤EMD ≤ 100 r with r being defined as the Wasserstein norm. For r=1, the measure is bound with 0≤EMD ≤100. As we want to ensure that we can be certain in using the performance criterion of a dataset as proxy for its domain distribution, we must be certain that we obtain the same distribution when training 7. Publications 84
Improved Sensor Model for Realistic Synthetic Data Generation CSCS ’21, November 30, 2021, Ingolstadt, Germany 0 20 40 60 80 100 mIoU [%] 0 100 200 300 400 500 Cumulative image count Ensemble 6 Ensemble 5 Ensemble 4 Ensemble 3 Ensemble 2 Ensemble 1 Figure 4: CDF of ensembles of DeeplabV3+ models trained on Cityscapes and evaluated on the same. Applying the twosample Kolmogorov-eSmirnov test to each possible pair of the ensemble we get a minimum p−value >0.95. with the same DNN from different, i.e., random, starting conditions. Therefore, we trained six models of the DeeplabV3+ network with the same hyperparameters but by different random weight initialization on the Cityscapes dataset and evaluated them on the validation set calculating the mIoU per image distributions. The resulting distributions of each model in the ensemble are then converted into a CDF as is shown in Figure 4. When comparing CDF s, and to reinforce the claim that the mIoU per image performance distribution is constant when training a model on a dataset, we apply the two-sample Kolmogorov-Smirnov test [Massey Jr 1951] on each pair of distributions in the ensemble. The resulting p-values of the Kolmogorov-Smirnov tests are at least > 0 . 95, hence supporting our hypothesis. 4 RESULTS AND DISCUSSION 4.1 Sensor Parameter Extraction As a baseline for our sensor artifact simulation we analyzed images from the Cityscapes training data and measured the parameters. Sensor noise was extracted from images with uniformly coloured areas ranging from dark to light colors. Chromatic aberration was extracted from images with traffic signs on the outmost edges of the image in horizontal and vertical direction, due to their favorable high contrast of black and white of the traffic signs. Saturation was measured as relative number of pixels in percent that are clipped at the value 255 in all color channels in the RGB images. The resulting measured parameters are then as follows: saturation = 2 . 0%, noise ∼ N( 0 , 3 ) , γ= 0 . 8, F= 4, and ca = 0 . 08. These parameters are later also used as a starting point of our parameter optimization approach. 4.2 Sensor Artifact Optimization Experiment As we want to show that by applying sensor lens artifacts we can effectively decrease the domain discrepancy between synthetic and real-world domain, we further utilize the EMD as dataset discrepancy measure and the extracted sensor parameters from camera images of Cityscapes, we then apply an optimization strategy to iteratively decrease the gap between the Cityscapes and the synthetic dataset [Consortium 2021] receiving the optimal parameters for the sensor simulation. For optimization, we chose to use the trust region method. Specifically, we use the trust region reflective (trf) method as implemented in SciPy [Virtanen et al . 2020]. The trf is a least squares minimization method to find the local minimum of a cost function given certain input variables. As cost function we use the EMD from synthetic model and real-world model predictions on the target test set. The variables as input to the cost function are the parameters of the sensor artifact simulation. The trf method has the capability of bounding the variables to meaningful ranges as this will prevent the parameters to increase into unreasonable values, e.g., increase the additive gaussian noise to values > 10. The stop criterion is met when the increase of parameter step size or decrease of the cost function is below 10−6. The overall description of our optimization method is depicted in Figure 5. Step 1: Initial parameters from the optimization method are applied in the sensor artifact simulation to the synthetic images. Step 2: The DeeplabV3+ model with ResNet101 backbone is pretrained on 15 epochs on the original synthetic dataset is loaded and trained for one epoch on the synthetic dataset with a learning rate of 0 . 1. Step 3: The model parameters are frozen and set to evaluate. Step 4: The model predicts on the validation set of the Cityscapes dataset. Step 5: The remaining domain discrepancy is measured by evaluation of the mIoU per image and calculation of the EMD to the evaluations of a model trained on Cityscapes. Step 6: The resulting EMD is fed as cost to the optimization method. Step 7: New parameters are set for the sensor artifact simulation, or the optimization ends if the stop criteria are met. After iterating the parameter optimization with the trf method we compare our optimized lens artifacts trained model with the unmodified synthetic trained model by their EMD s with a model trained on the Cityscapes dataset. Figure 6 depicts the distributions resulting from this evaluation. The DeeplabV3+ model trained with the optimized sensor artifact simulation applied on the synthetic dataset outperforms the baseline and achieves an EMD score of 26 . 48, while decreasing the domain gap by 6 . 19. The resulting parameters are saturation = 2 . 11%, noise ∼ N( 0 , 3 . 0000005 ) , γ= 0 . 800001, F= 4and ca = 0 . 008000005. The parameters changed only slightly from the starting point, indicating the extracted parameters as good first choice as we can confirm when we optimize from a random initialized starting point. An exemplary visual inspection of the results in Figure 7 helps to understand the distribution shift and therefore the decreased EMD . While the best prediction performance image (TOP) increased only slightly from the synthetic trained model (c) to the sensor artifact optimized model (d), the worst prediction case (Bottom) shows improved segmentation performance for the sensor artifact optimized model (d), in this case even better than the Cityscapes trained model (b). We compare the overall mIoU performance on the Cityscapes datasets between models trained with the initial unmodified synthetic dataset, the synthetic dataset with random initialized lens artifact parameters and the synthetic dataset with extracted parameters from Cityscapes with the baseline of a model trained on 7.2. Publication 2 85
128 K. Hagn and O. Grau 1 Introduction Validation of deep neural networks (DNNs) is increasingly resorting toward computer-generated imagery (CGI) due to its mitigation of certain issues. First, synthetic data can avoid privacy issues found with recordings of members of the public and, on the other hand, can automatically produce vast amounts of data at high quality with pixel-accurate ground truth data and reliability than costly manually labeled data. Moreover, simulations allow synthesis of rare cases and the systematic variation and explanation of critical constellations [SGH20]—a requirement for validation of products targeting safety-critical applications, such as automated driving. Here, the creation of corner cases and scenarios which otherwise could not be recorded in a real-world scenario without endangering other traffic participants is the key argument for the validation of perceptive AI with synthetic images. Despite the advantages of CGI methods, training and validation with synthetic images still have challenges: Training with these images does not guarantee a similar performance on real-world images and validation is only valid if one can verify that the found weaknesses in the validation do not stem from the synthetic-to-real distribution shift seen in the input. To measure and mitigate this domain shift, metrics have been introduced with various applications in the field of domain adaptation or transfer learning. In domain adaptation, the metrics such as FID, kernel inception distance (KID), and maximum mean discrepancy (MMD) are applied to train generative adversarial networks (GANs) to adapt on a target feature space [PTKY09] or to re-create the visual properties of a dataset [SGZ+16]. However, the problem of training and validation with synthetic imagery is directly related to the predictive performance of a perception algorithm on the target data, and these kinds of metrics struggle to correlate with the predictive performance [RV19]. Additionally, applications of domain adaptation methods often resort to specifically trained DNNs, e.g., GANs, which adapt one domain to the other and therefore add an extra layer of complexity and uncontrollability. This is especially unwanted if a validation goal is tested, e.g., to detect all pedestrians, and the domain adaption by a GAN would add additional objects into the scene (e.g., see [HTP+18]) making it even harder to attribute detected faults of the model to certain specifics of the tested scene. Here, the creation of images via a synthesis process allows to understand domain distance influence factors more directly as all parameters are under direct control. Camera-recorded images inherently show visual imperfections or artifacts, such as sensor noise, blur, chromatic aberration, or image saturation, as can be seen in an image example from the A2D2 [GKM+20] dataset in Fig.1. CGI methods, on the other hand, are usually based on idealized models; for example, the pinhole camera model [Stu14] which is free of sensor artifacts. In this chapter, we present an approach to decrease the domain divergence of synthetic to real-world imagery for perceptive DNNs by realistically modeling sensor lens artifacts to increase the viability of CGI for training and validation. To achieve this, we first introduce a model of sensor artifacts whose parameters are extracted 7. Publications 92
Optimized Data Synthesis for DNN Training and Validation … 129 Fig. 1 Real-world images (here A2D2) exhibit sensor lens artifacts which have to be closely modeled by an image synthesization process to decrease the domain distance of synthetic to realworld datasets to make them viable for training and validation from a real-world dataset and then apply it on a synthetic dataset for training and measuring the remaining domain divergence via validation. Therefore, a new interpretation of the domain divergence by generalization of the distance of two datasets by the per-image performance comparison over a dataset utilizing the Wasserstein or earth mover’s distance (EMD) is presented. Next, we demonstrate how this model is able to decrease the domain divergence further by optimization of the initial extracted sensor camera simulation parameters as depicted in Fig.6. Additionally, we compare our results with randomly chosen parameters as well as with randomly chosen and optimized parameters. Last, we strengthen the case for the usability of our EMD domain divergence measure by comparison with the well-known Fréchet inception distance (FID) on a set of real-world and synthetic datasets and highlight the advantage of our asymmetric domain divergence against the symmetric distance. 2 Related Works This chapter is related to two areas: domain distance measures, as used in the field of domain adaptation and synthetic data generation for training and validation. Domain distance measures: A key challenge in domain adaptation approaches is the expression of a distance measure between datasets, also called domain shift. A number of methods were developed to mitigate this shift (e.g., see [LCWJ15,GL15, THSD17,THS+18]). To measure the domain shift or domain distance, the inception score (IS) has been proposed [SGZ+16], where the classification output of an InceptionV37.3. Publication 3 93
130 K. Hagn and O. Grau based [SVI+16] discriminator network trained on the ImageNet dataset [DDS+09] is used. The works of [HRU+17,BSAG21] rely on features extracted from the InceptionV3 network to tune domain adaptation approaches, i.e., the FID and KID. However, these metrics cannot predict if the classification performance increases when adapted data is applied as training data for a discriminator [RV19]. Therefore, to measure performance directly, it is essential to train with the adapted or synthetic data and validate on the target data, i.e., cross-evaluation as done by [RSM+16,WU18,SAS+18]. Performance metrics: The mean intersection-over-union (mIoU) is a widely used performance metric for benchmarking semantic segmentation [COR+16,VSN+18]. Adaptations and improvements of the mIoU have been proposed which put more weight on the segmentation contour as in [FMWR18,RTG+19]. Performance metrics as the mIoU are computed over the whole validation dataset, i.e., the whole confusion matrix, but there are propositions to apply the mIoU calculation on a per-image basis and compare the resulting empirical distributions [CLP13]. A per-image comparison mitigates several shortcomings of a single evaluation metric on the whole dataset when used for comparison of classifiers on the same dataset. First, one can distinguish multimodal and unimodal distributions, i.e., strong classification on one half and weak classification on the other half of a set can lead to the same mean as an average classification on all samples. Second, unimodal distributions with the same mean but different shape are also indiscernible under a single dataset averaged metric. This justification led to our choice of a per-imagebased mIoU metric as it allows for deeper investigations which are especially helpful when one wants to understand the characteristics that increase or decrease a domain divergence. Sensor simulation for synthetic image training: The use of synthesized data for development and validation is an accepted technique and has been also suggested for computer vision applications (e.g., [BB95]). Recently, specifically for the domain of driving scenarios, games engines have been adapted [RHK17,DRC+17]. Although game engines provide a good starting point to simulate environments, they usually only offer a closed rendering set-up with many trade-offs balancing between real-time constraints and a subjectively good visual appearance for human observers. Specifically the lighting computation in the rendering pipelines is intransparent. Therefore, it does not produce a physically correct imagery; instead only a fixed rendering quality (as a function of lighting computation and tone mapping), resulting in output of images having a low dynamic range (LDR) (typically 8-bit per RGB color channel). Recently, physical-based rendering techniques have been applied to the generation of data for training and validation, like Synscapes [WU18]. For our chapter we use a dataset in high dynamic range (HDR) created with the physical-based Blender Cycles renderer.1We implemented a customized tone mapping to 8-bit per color channel and sensor simulation, as described in the next section. 1Provided by the KI Absicherung project [KI 20]. 7. Publications 94
Optimized Data Synthesis for DNN Training and Validation … 131 While there is great interest in understanding the domain distance in the area of domain adaptation via generative strategies, i.e., GANs, there has been little research regarding sensor artifact influence on training and validation with synthetic images. Other works [dCCN+16,NdCCP18] add different kinds of sensor noise to their training set and report a degradation of performance, compared to a model trained with no noise in the training set, due to training of a harder, i.e., noisier, visual task. Adding noise in training is a common technique for image augmentation and can be seen as a regularization technique [Bis95] to prevent overfitting. Our task of modeling sensor artifacts for synthetic images extracted from camera images is not aimed at improving the generalization through random noise, but to tune the parameters of our sensor model to closely replicate the real-world images and improve generalization on the target data. First results of modeling camera effects to improve synthetic data learning on the perceptive task of bounding box detection have been proposed by [CSVJR18, LLFW20]. Lin et al. [LLFW20] additionally state that generalization is an asymmetric measure which should be considered when comparing with symmetric dataset distance measures from literature. Furthermore, Carlson et al. [CSVJR19] learned sensor artifact parameters from a real-world dataset and applied the learned parameters of their noise sources as image augmentation during training with synthetic data on the task of bounding box detection. However, contrasting our approach, they apply their optimization as style loss on a latent feature vector extracted from a VGG-16 network trained on ImageNet and evaluate the performance on the task of 2D object detection. 3 Methods Given a synthetic (CGI) dataset of urban street scenes, our goal is to decrease the domain gap to a real-world dataset for semantic segmentation by realistic sensor artifact simulation. Therefore, we systematically analyze the image sensor artifacts of the real-world dataset and use this extracted parametrization for our sensor artifact simulation. To compare our synthetic dataset with the real-world dataset we contrive a novel per-image performance-based metric to measure the generalization distance between the datasets. We utilize a DeeplabV3+ [CZP+18] semantic segmentation model with a ResNet101 [HZRS16] backbone to train and evaluate on the different datasets throughout this paper. To show the valuable properties of our measure we compare it with the established domain distance, i.e., Fréchet inception distance (FID). Lastly, we use our measure as optimization criteria for adapting the parameters of our sensor artifact simulation with the extracted parameters as starting point and show that we can further decrease the domain distance from synthetic images to real-world images. 7.3. Publication 3 95
132 K. Hagn and O. Grau 3.1 Sensor Simulation We implemented a simple sensor model with the principle blocks depicted in Fig.2: The module expects images in linear RGB space. Rendering engines like Blender Cycles2can provide these images as results in OpenEXR format.3 We simulate a simple model by applying chromatic aberration,blur, and sensor noise, as additive Gaussian noise (zero mean, variance is a free parameter), followed by a simple exposure control (linear tone mapping), finished by non-linear gamma correction. First, we apply blur by a simple box filter with filter size F×Fand a chromatic aberration (CA). The CA is approximated using radial distortions (k1, second order), e.g., [CV14], as defined in OpenCV. The CA is implemented as a per channel (red, green, blue) variation of the k1 radial distortion, i.e., we introduce an incremental parameter ca that affects the radial distortions: k1(blue)= −ca;k1(green)= 0;k1(red)= +ca. As the next step, we apply Gaussian noise to the input image. Applying a linear function, the pixel values are then mapped and rounded to the target output byte range [0, ..., 255]. The two parameters of the linear mapping are determined by a histogram evaluation of the input RGB values of the respective image, imitating an auto exposure of a real camera. In our experiments we have set it to saturate 2% (initially) of the brightest pixel values, as these are usually values of very high brightness, induced by sky or even the sun. Values below the minimum or above the set maximum are mapped to 0 or 255, respectively. In the last step we apply gamma correction to achieve the final processed synthetic image: x=(˜ x)γ(1) The parameter γis an approximation of the sensor non-linear mapping function. For media applications this is usually γ=2.2 for the sRGB color space [RD14]. However, for industrial cameras, this is not yet standardized and some vendors do Linear Mapping, Rounding + Noise Generator Measurement of Histogram Range Gamma Correction Blur, Chromatic Aberration synthetic image x ˜x ∈[0, ..., 255] "real" image x Fig. 2 Sensor artifact simulation 2https://www.blender.org/. 3https://www.openexr.com/. 7. Publications 96
Optimized Data Synthesis for DNN Training and Validation … 133 (a) (b) Fig. 3 Left a: Synthetic images without lens artifacts. Right b: Applied sensor lens artifacts, including exposure control not reveal it.4We therefore estimate the parameter as an approximation. Figure 3 depicts the difference of an image with and without simulated sensor artifacts. 3.2 Dataset Divergence Measure Our proposed distance quantifies per image performance between models trained on different datasets but evaluated on the same dataset. Considering the task of semantic segmentation we chose the mIoU as our base metric. We then modify the mIoU to be calculated per image instead of the confusion matrix on the whole evaluated dataset. Next, we introduce the Wasserstein-1 or earth mover’s distance (EMD) metric as our divergence measure between the per-image mIoU distribution of two classifiers trained on distinct datasets, i.e., synthetic and real-world datasets, but evaluated on the same real-world dataset the second classifier has been trained with. The mIoU is defined as follows: mIoU =1 S s∈S T Ps T Ps+F Ps+F Ns ×100%,(2) 4The providers of the Cityscapes dataset don’t document the exact mapping. 7.3. Publication 3 97
134 K. Hagn and O. Grau with T Ps,F Ns, and F Psbeing the amount of true-positives, false-negatives, and false-positives of the sth class over all images of the evaluated dataset. Here, S= {0,1, ..., S−1}, with S=11, as we use the 11 classes defined in Table1. These classes are the maximal overlap of common classes in the real and synthetic datasets considered for cross-evaluation and comparison of our measure with the Fréchet inception distance (FID), as can be seen later in Sect.4.3, Tables 3 and 4. A distribution over the per-image IoU takes the following form: IoUn=1 S s∈S T Ps,n T Ps,n+F Ps,n+F Ns,n ×100%,(3) where ndenotes the nth image in the validation dataset. Here, IoUnis measured in %. We want to compare the distributions of per-image IoU values from two different models; therefore, we apply the Wasserstein distance. The Wasserstein distance as an optimal mass transport metric from [KPT+17] is defined for density distributions p and qwhere inf denotes the infinium, i.e., lowest transportation cost, (p,q)denotes the set containing all joint distributions π, i.e., transportation maps, for (X,Y) which have the marginals pand qas follows: Wr(p,q)=inf π∈(p,q)R×R |X−Y|rdπ1/r .(4) This distance formulation is equivalent to the following [RTC17]: Wr(p,q)=∞ −∞ |P(t)−Q(t)|r1/r dt.(5) Here Pand Qdenote the respective cumulative distribution functions (CDFs) of p and q. In our application we calculate the empirical distributions of pand q, which simplifies in this case to the function of the order statistics: Wr(ˆp,ˆq)=n i=1 | ˆpi− ˆqi|r1/r ,(6) where ˆpand ˆqare the empirical distributions of the marginals pand qsorted in ascending order. With r=1 and equal weight distributions we get the earth mover’s distance (EMD) which, in other words, measures the area between the respective CDFs with L1as ground distance. We assume a sample size of at least 100 to be enough for the EMD calculation to be valid, as fewer samples might not guarantee a sufficient sampling of the domains. In our experiments we use sample sizes ≥500. 7. Publications 98
Optimized Data Synthesis for DNN Training and Validation … 135 Table 1 Due to differences in label definition of real-world datasets, the class mapping for training and evaluation is decreased to 11 classes that are common in all considered datasets: A2D2 [GKM+20], Cityscapes [COR+16], Berkeley Deep Drive (BDD100K) [YCW+20], Mapillary Vistas (MV) [NOBK17], India Driving Dataset (IDD) [VSN+18], GTAV [RVRK16], our synthetic dataset [KI 20], and Synscapes [WU18] s0 1 2 3 4 5 6 7 8 9 10 Label road sidewalk building pole traffic light traffic sign vegetation sky human car truck 7.3. Publication 3 99
136 K. Hagn and O. Grau 0 20 40 60 80 100 mIoU [%] 0.0 0.2 0.4 0.6 0.8 1.0 Cumulative count Model 6 Model 5 Model 4 Model 3 Model 2 Model 1 Fig. 4 CDF of ensembles of DeeplabV3+ models trained on Cityscapes and evaluated on its validation set. Applying the 2-sample Kolmogorov-Smirnov test to each possible pair of the ensemble, we get a minimum p−value >0.95 The FID is a special case of the Wasserstein-2 distance derived from (6) with p=2 and ˆpand ˆqbeing normally distributed, leading to the following definition: FID = ||µ−µw||2+tr( +w−2(w)1/2), (7) where µand µware the means, and and ware the covariance matrices of the multivariate Gaussian-distributed feature vectors of synthetic and real-world datasets, respectively. Compared to distance metrics such as the FID which by definition is symmetric, our measure is a divergence, i.e., the distance from dataset A to dataset B can be different to the distance from dataset B to dataset A. Being a divergence also reflects the characteristic of a classifier having different generalization distance when trained on dataset A and evaluated on dataset B or the other way around. Because the ground measure of the signatures, i.e., the IoU per image, is bounded to 0 ≤I oUn≤100, the EMD measure is then bounded to 0 ≤E M D ≤100rwith r being the Wasserstein norm. For r=1, the measure is bound with 0 ≤E M D ≤100. To verify whether the per-image IoU of a dataset is a good proxy of a dataset’s domain distribution, we need to verify that the distribution stays (nearly) constant when training from different starting conditions. Therefore, we trained six models of the DeeplabV3+ network with the same hyperparameters but different random initialization on the Cityscapes dataset and evaluated them on the validation set calculating the mIoU per image. The resulting distributions of each model in the ensemble are converted into a CDF as is shown in Fig. 4. To have a stronger empirical evidence of the per-image mIoU performance distribution being constant for a dataset, we apply the two-sample Kolmogorov-Smirnov test on each pair of distri7. Publications 100
Optimized Data Synthesis for DNN Training and Validation … 137 bution in the ensemble. The resulting p-values are at least >0.95, hence supporting our hypothesis. 3.3 Datasets For our sensor parameter optimization experiments we consider two datasets. First, the real-world Cityscapes dataset, which consists of 2,975 annotated images for training and 500 annotated images for validation. All images were captured in urban street scenes in German cities. Second, the synthetic dataset provided by the KIA project [KI 20]. This dataset consists of 21,802 annotated training images and 5,164 validation images. The KI-A synthetic dataset comprises urban street scenes, similar to Cityscapes, and suburban to rural street scenes which are characterized by less traffic and less dense house placements, therefore more vegetation and terrain objects. 4 Results and Discussion 4.1 Sensor Parameter Extraction As a baseline for our sensor simulation, we analyzed images from the Cityscapes training data and measured the parameters. Sensor noise was extracted from about 10 images with uniformly colored areas ranging from dark to light colors. Chromatic aberration was extracted from 10 images with traffic signs on the outmost edges of the image, as can be seen in Fig.5. The extracted values have been averaged over the count of images. The starting parameters of our optimization approach are then as follows: saturation =2.0%, noise ∼N(0,3),γ=0.8, F=4, and ca =0.08. 4.2 Sensor Artifact Optimization Experiment Utilizing the EMD as dataset divergence measure and the extracted sensor parameters from camera images of Cityscapes, we apply an optimization strategy to iteratively decrease the gap between the Cityscapes and the synthetic dataset [KI 20]. For optimization, we chose to use the trust region reflective (trf) method [SLA+15] as implemented in SciPy [VGO+20]. The trf is a least-squares minimization method to find the local minimum of a cost function given certain input variables. The cost function is the EMD from synthetic model and real-world model predictions on the same real-world validation dataset. The variables as input to the cost function are the parameters of the sensor artifact simulation. The trf method has the capability 7.3. Publication 3 101
144 K. Hagn and O. Grau hyperparameters leads to the same per-image performance distributions when these ensemble models are evaluated on the validation set of the training dataset. When utilizing synthetic imagery for validation, the domain gap, due to visual differences between real and computer-generated images, is hindering the applicability of these datasets. As a step toward decreasing the visual differences, we apply the proposed divergence measure as a cost function to an optimization which varies the parameters of the sensor artifact simulation, while trying to re-create the sensor artifacts that the real-world dataset exhibits. As starting point of the sensor artifact parameters, we extracted empirically the values from chosen images of the real-world dataset. The optimization improved the visual difference between the real-world and the optimized synthetic dataset measurably by the EMD and we could show that even when starting with random initialized parameters we can decrease the EMD and increase the mIoU on the target datasets. When measuring the divergence after parameter optimization to other real-world and synthetic datasets, we could show that the EMD decreases for all considered datasets but when measured by the FID only four of the datasets are closer. As the EMD is derived from the mIoU per image, it is an indicator of performance on the target dataset, whereas the FID fails to relate with performance. Effective minimization of the visual difference between synthetic and real-world datasets with the EMD domain divergence measure is one step further toward fully utilizing CGI for validation of perceptive AI functions. Acknowledgements The research leading to these results is funded by the German Federal Ministry for Economic Affairs and Energy within the project “Methoden und Maßnahmen zur Absicherung von KI-basierten Wahrnehmungsfunktionen für das automatisierte Fahren (KI Absicherung)”. The authors would like to thank the consortium for the successful cooperation. References [BB95] W. Burger, M.J. Barth, Virtual reality for enhanced computer vision, in J. Rix, S. Haas, J. Teixeira (eds.), Virtual Prototyping: Virtual Environments and the Product Design Process (Springer, 1995), pp. 247–257 [Bis95] M. Christopher Bishop, Training with noise is equivalent to Tikhonov regularization. Neural Comput. 7(1), 108–116 (1995) [BSAG21] M. Binkowski, D.J. Sutherland, M. Arbel, A. Gretton, Demystifying MMD GANs, Jan. 2021, pp. 1–36. arxiv:1801.01401 [CLP13] G. Csurka, D. Larlus, F. Perronnin, What is a good evaluation measure for semantic segmentation? in Proceedings of the British Machine Vision Conference (BMVC), Bristol, UK, Sept. 2013, pp. 1–11 [COR+16] M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, B. Schiele, The Cityscapes dataset for semantic urban scene understanding, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, June 2016, pp. 3213–3223 [CSVJR18] A. Carlson, K.A. Skinner, R. Vasudevan, M. Johnson-Roberson, Modeling camera effects to improve visual learning from synthetic data, in Proceedings of the European Conference on Computer Vision (ECCV) Workshops, Munich, Germany, Aug. 2018, pp. 505–520 7. Publications 108
Optimized Data Synthesis for DNN Training and Validation … 145 [CSVJR19] A. Carlson, K.A. Skinner, R. Vasudevan, M. Johnson-Roberson, Sensor Transfer: Learning Optimal Sensor Effect Image Augmentation for Sim-to-Real Domain Adaptation, Jan. 2019, pp. 1–8. arxiv:1809.06256 [CV14] V. Chari, A. Veeraraghavan, L. Distortion, Radial distortion, in Computer Vision: A Reference Guide. ed. by K. Ikeuchi (Springer, Boston, MA, 2014), pp. 443–445 [CZP+18] L.-C. Chen, Y. Zhu, G. Papandreou, F. Schroff, H. Adam, Encoder-Decoder with atrous separable convolution for semantic image segmentation, in Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, Sept. 2018, pp. 833–851 [dCCN+16] G.B.P. da Costa, W.A. Contato, T.S. Nazaré, J.E.S. do Batista Neto, M. Ponti, An Empirical Study on the Effects of Different Types of Noise in Image Classification Tasks, Sept. 2016, pp. 1–6. arxiv:1609.02781 [DDS+09] J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, F.-F. Li, ImageNet: a large-scale hierarchical image database, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Miami, FL, USA, June 2009, pp. 248–255 [DRC+17] A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, V. Koltun, CARLA: an open urban driving simulator, in Proceedings of the Conference on Robot Learning CORL, Mountain View, CA, USA, Nov. 2017, pp. 1–16 [FMWR18] E. Fernandez-Moral, R. Martins, D. Wolf, P. Rives, A new metric for evaluating semantic segmentation: leveraging global and contour accuracy, in Proceedings of the IEEE Intelligent Vehicles Symposium (IV), Changshu, China, June 2018, pp. 1051– 1056 [GKM+20] J. Geyer, Y. Kassahun, M. Mahmudi, X. Ricou, R. Durgesh, A.S. Chung, L. Hauswald, V.H. Pham, M. Mühlegg, S. Dorn, T. Fernandez, M. Jänicke, S. Mirashi, C. Savani, M. Sturm, O. Vorobiov, M. Oelker, S. Garreis, P. Schuberth, A2D2: Audi Autonomous Driving Dataset, April 2020, pp. 1–10. arxiv:2004.06320 [GL15] Y. Ganin, V. Lempitsky, Unsupervised domain adaptation by backpropagation, in Proceedings of the International Conference on Machine Learning (ICML), Lille, France, July 2015, pp. 1180–1189 [HRU+17] M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, S. Hochreiter, GANs trained by a two time-scale update rule converge to a local Nash equilibrium, in Proceedings of the Conference on Neural Information Processing Systems (NIPS/NeurIPS), Long Beach, CA, USA, Dec. 2017, pp. 6626–6637 [HTP+18] J. Hoffman, E. Tzeng, T. Park, J.-Y. Zhu, P. Isola, K. Saenko, A. Efros, T. Darrell, CyCADA: cycle-consistent adversarial domain adaptation, in Proceedings of the International Conference on Machine Learning (ICML), Stockholm, Sweden, July 2018, pp. 1989–1998 [HZRS16] K. He, X. Zhang, S. Ren, J. Sun, Deep residual learning for image recognition, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, June 2016, pp. 770–778 [KI 20] KI Absicherung Consortium. KI Absicherung: Safe AI for Automated Driving (2020). Assessed 18 Nov. 2021 [KPT+17] S. Kolouri, S.R. Park, M. Thorpe, D. Slepcev, G.K. Rohde, Optimal mass transport: signal processing and machine-learning applications. IEEE Signal Process. Mag. 34(4), 43–59 (2017) [LCWJ15] M. Long, Y. Cao, J. Wang, M.I. Jordan, Learning transferable features with deep adaptation networks, in Proceedings of the International Conference on Machine Learning (ICML), July 2015, pp. 97–105 [LLFW20] Z. Liu, T. Lian, J. Farrell, B. Wandell, Neural network generalization: the impact of camera parameters. IEEE Access 8, 10443–10454 (2020) [NdCCP18] T.S. Nazaré, G.B.P. da Costa, W.A. Contato, M. Ponti, Deep convolutional neural networks and noisy images, in Proceedings of the Iberoamerican Congress on Pattern Recognition (CIARP), Madrid, Spain, Nov. 2018, pp. 416–424 7.3. Publication 3 109
146 K. Hagn and O. Grau [NOBK17] G. Neuhold, T. Ollmann, S. Rota Bulò, P. Kontschieder, The mapillary vistas dataset for semantic understanding of street scenes, in Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy, Oct. 2017, pp. 4990–4999 [PTKY09] S.J. Pan, I.W. Tsang, J.T. Kwok, Q. Yang, Domain adaptation via transfer component analysis, in Proceedings of the International Joint Conferences on Artificial Intelligence (IJCAI), Pasadena, CA, USA, July 2009, pp. 1187–1192 [RD14] R. Ramanath, M.S. Drew, Color Spaces, in Computer Vision: A Reference Guide. ed. by K. Ikeuchi (Springer, Boston, MA, 2014), pp. 123–132 [RHK17] S.R. Richter, Z. Hayder, V. Koltun, Playing for benchmarks, in Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy, Oct. 2017, pp. 2232–2241 [RSM+16] G. Ros, L. Sellart, J. Materzynska, D. Vazquez, A.M. Lopez, The SYNTHIA dataset: a large collection of synthetic images for semantic segmentation of urban scenes, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, June 2016, pp. 3234–3243 [RTC17] A. Ramdas, N. García Trillos, M. Cuturi, On Wasserstein two sample testing and related families of nonparametric tests. Entropy 19(2), 47 (2017) [RTG+19] H. Rezatofighi, N. Tsoi, J.Y. Gwak, A. Sadeghian, I. Reid, S. Savarese, Generalized intersection over union: a metric and a loss for bounding box regression, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, June 2019, pp. 658–666 [RV19] S.V. Ravuri, O. Vinyals, Seeing is not necessarily believing: limitations of BigGANs for data augmentation, in: Proceedings of the International Conference on Learning Representations (ICLR) Workshops, New Orleans, LA, USA, June 2019, pp. 1–5 [RVRK16] S.R. Richter, V. Vineet, S. Roth, V. Koltun, Playing for data: ground truth from computer games, in Proceedings of the European Conference on Computer Vision (ECCV), Amsterdam, The Netherlands, Oct. 2016, pp. 102–118 [SAS+18] F.S. Saleh, M.S. Aliakbarian, M. Salzmann, L. Petersson, J.M. Alvarez, Effective use of synthetic data for urban scene semantic segmentation, in Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, Sept. 2018, pp. 84–100 [SGH20] Q.S. Sha, O. Grau, K. Hagn, DNN analysis through synthetic data variation, in Proceedings of the ACM Computer Science in Cars Symposium (CSCS), virtual conference, Dec. 2020, pp. 1–10 [SGZ+16] T. Salimans, I. Goodfellow, W. Zaremba, V. Cheung, A. Radford, Xi Chen, Improved Techniques for Training GANs, June 2016, pp. 1–10. arxiv:1606.03498 [SLA+15] J. Schulman, S. Levine, P. Abbeel, M. Jordan, P. Moritz, Trust region policy optimization, in Proceedings of the International Conference on Machine Learning (ICML), Lille, France, July 2015, pp. 1889–1897 [Stu14] P. Sturm, Pinhole Camera Model, in Computer Vision: A Reference Guide. ed. by K. Ikeuchi (Springer, Boston, MA, 2014), pp. 610–613 [SVI+16] C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, Z. Wojna, Rethinking the inception architecture for computer vision, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, June 2016, pp. 2818–2826 [THS+18] Y.-H. Tsai, W.-C. Hung, S. Schulter, K. Sohn, M.-H. Yang, M. Chandraker, Learning to adapt structured output space for semantic segmentation, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, June 2018, pp. 7472–7481 [THSD17] E. Tzeng, J. Hoffman, K. Saenko, T. Darrell, Adversarial discriminative domain adaptation, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, July 2017, pp. 2962–2971 [VGO+20] P. Virtanen, R. Gommers, T.E. Oliphant, M. Haberland, T. Reddy, D. Cournapeau, E. Burovski, P. Peterson, W. Weckesser, J. Bright, S.J. van der Walt, M. Brett, J. Wilson, 7. Publications 110
Optimized Data Synthesis for DNN Training and Validation … 147 K.J. Millman, N. Mayorov, A.R.J. Nelson, E. Jones, R. Kern, E. Larson, C.J. Carey, l. Polat, Y. Feng, E.W. Moore, J. Vand erPlas, D. Laxalde, J. Perktold, R. Cimrman, I. Henriksen, E.A. Quintero, C.R. Harris, A.M. Archibald, A.H. Ribeiro, F. Pedregosa, P. van Mulbregt, and SciPy 1.0 Contributers. Fundamental algorithms for scientific computing in python. SciPy 1.0. Nat. Methods 17, 261–272 (2020) [VSN+18] G. Varma, A. Subramanian, A. Namboodiri, M. Chandraker, C.V. Jawahar, IDD: A Dataset for Exploring Problems of Autonomous Navigation in Unconstrained Environments, Nov. 2018, pp. 1–9. arxiv:1811.10200 [WU18] M. Wrenninge, J. Unger, Synscapes: A Photorealistic Synthetic Dataset for Street Scene Parsing, Oct. 2018, pp. 1–13. arxiv:1810.08705 [YCW+20] F. Yu, H. Chen, X. Wang, W. Xian, Y. Chen, F. Liu, V. Madhavan, T. Darrell, BDD100K: a diverse driving dataset for heterogeneous multitask learning, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), virtual conference, June 2020, pp. 2636–2645 Open Access This chapter is licensed under the terms of the Creative Commons Attribution 4.0 International License (http://creativecommons.org/licenses/by/4.0/), which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons license and indicate if changes were made. The images or other third party material in this chapter are included in the chapter’s Creative Commons license, unless indicated otherwise in a credit line to the material. If material is not included in the chapter’s Creative Commons license and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. 7.3. Publication 3 111
7. Publications 7.4 Publication 4 A Variational Deep Synthesis Approach for Perception Validation Oliver Grau, Korbinian Hagn and Qutub Syed Sha Published in: Fingscheidt, T., Gottschalk, H., Houben, S. (eds) Deep Neural Networks and Data for Automated Driving. Springer, Cham. [GHS22] Reprinted with permission from Oliver Grau DOI: 10.1007/978-3-031-01233-4_13 112
A Variational Deep Synthesis Approach for Perception Validation Oliver Grau, Korbinian Hagn, and Qutub Syed Sha Abstract This chapter introduces a novel data synthesis framework for validation of perception functions based on machine learning to ensure the safety and functionality of these systems, specifically in the context of automated driving. The main contributions are the introduction of a generative, parametric description of threedimensional scenarios in a validation parameter space, and layered scene generation process to reduce the computational effort. Specifically, we combine a module for probabilistic scene generation, a variation engine for scene parameters, and a more realistic sensor artifacts simulation. The work demonstrates the effectiveness of the framework for the perception of pedestrians in urban environments based on various deep neural networks (DNNs) for semantic segmentation and object detection. Our approach allows a systematic evaluation of a high number of different objects and combined with our variational approach we can effectively simulate and test a wide range of additional conditions as, e.g., various illuminations. We can demonstrate that our generative approach produces a better approximation of the spatial object distribution to real datasets, compared to hand-crafted 3D scenes. 1 Introduction This chapter introduces an automated data synthesis approach for the validation of perception functions based on a generative and parameterized synthetic data generation. We introduce a multi-stage strategy to sample the input domain of the possible generative scenario and sensor space and discuss techniques to reduce the required vast amount of computational effort. This concept is an extension and generalizaO. Grau (B)·K. Hagn ·Q. Syed Sha Intel Deutschland GmbH, Lilienthalstraße 15, 85579 Neubiberg, Germany e-mail: oliver[email protected] K. Hagn e-mail: [email protected] Q. Syed Sha e-mail: [email protected] © The Author(s) 2022 T. Fingscheidt et al. (eds.), Deep Neural Networks and Data for Automated Driving, https://doi.org/10.1007/978-3-031-01233-4_13 359 7.4. Publication 4 113
360 O. Grau et al. tion of our previous work on parameterization of the scene parameters of concrete scenarios, called validation parameter space (VPS) [SGH20]. We extend this parameterization by a probabilistic scene generator to widen the coverage of the generated scenarios and a more realistic sensor simulation, which also allows to variate and simulate different sensor characteristics. This ‘deep’ synthesis concept overcomes currently available systems (as discussed in the next section) or manually, i.e., by human-operator-generated synthetic data. We describe, how our synthetic data validation engine makes use of the parameterized, generative content to implement a tool supporting complex and effective validation strategies. Perception is one of the hardest problems to solve in any automated system. Recently, great progress has been made in applying machine learning techniques to deep neural networks to solve perceptional problems. Automated vehicles (AVs) are a recent focus as an important application of perception from cameras and other sensors, such as LiDAR and RaDAR [YLCT20]. Although the current main effort is on developing the hardware and software to implement the functionality of AVs, it will be equally important to demonstrate that this technology is safe. Universally accepted methodologies for validating safety of machine learning-based systems are still an open research topic. Techniques to capture and render models of the real world have matured significantly over the last decades and are now able to synthesize virtual scenes in a visual quality that is hard to distinguish from real photographs for human observers. Computer-generated imagery (CGI) is increasingly popular for training and validation of deep neural networks (DNNs) (see, e.g., [RHK17,Nik19]). Synthetic data can avoid privacy issues found with recordings of members of the public and can automatically produce ground truth data at higher quality and reliability than costly manually labeled data. Moreover, simulations allow synthesis of rare scene constellations helping validation of products targeting safety-critical applications, specifically automated driving. Due to the progress in visual and multi-sensor synthesis, building systems for validation of these complex systems in the data center becomes feasible now and offers more possibilities for the integration of intelligent techniques in the engineering process of complex applications. We compare our approach with methods and strategies targeting testing of automated driving [JWKW18]. The remainder of this chapter is structured as follows: The next section will give an outline of related work in the field. In Sect.3we give an overview of our approach. Section4describes an outline of our synthetic data validation engine, our parameterization, including a realistic sensor simulation, and the effective computation of the required variations. In Sect.5we present evaluation results, followed by Sect.6 with some concluding remarks. 7. Publications 114
A Variational Deep Synthesis Approach for Perception Validation 361 2 Related Work The use of synthesized data for development and validation is an accepted technique and has been also suggested for computer vision applications (e.g., [BB95]). Several methodologies for verification and validation of AVs have been developed [KP16,JWKW18,DG18] and commercial options exist.1These tools were originally designed for virtual testing of automotive functions, such as braking systems, and then extended to provide simulation and management tools for virtual test drives in virtual environments. They provide real-time-capable models for vehicles, roads, drivers, and traffic which are then being used to generate test (sensor) data as well as APIs for users to integrate the virtual simulation into their own validation system. What is getting presented in this chapter is focusing on the validation of perception functions, which is an essential module of automated systems. However, by separating the perception as a component, the validation problem can also be decoupled from the validation of the full driving stack. Moreover, this separation allows, on the one hand, the implementation of various more specialized validation strategies and, on the other hand, there is no need to simulate dynamic actors and the connected problem of interrelations between them and the ego-vehicle. The full interaction of objects is targeted by upcoming standards like OpenScenario.2 Recently, specifically in the domain of driving scenarios, game engines have been adopted for synthetic data generation by extraction of in-game images and labels from the rendering pipeline [WEG+00,RVRK16]. Another virtual simulator system, which gained popularity in the research community, is CARLA [DRC+17], also based on a commercial game engine (Unreal4 [Epi04]). Although game engines provide a good starting point to simulate environments, they usually only offer a closed rendering setup with many trade-offs balancing between real-time constraints and a subjectively good visual appearance to human observers. Specifically, the lighting computation in this rendering pipelines is limited and does not produce physically correct imagery. Instead, game engines only deliver fixed rendering quality typically with 8 bit per RGB color channel and only basic shadow computation. In contrast, physical-based rendering techniques have been applied to the generation of data for training and validation, as in the Synscapes dataset [WU18]. For our experimental deep synthesis work, we use the physical-based open-source Blender Cycles renderer3in high dynamic range (HDR) resolution, which allows realistic simulation of illumination and sensor characteristics increasing the coverage of our synthetic data in terms of scene situations and optical phenomena occurring in real-world scenarios. The effect of sensor and lens effects on perception performance has not been studied a lot. In [CSVJR18,LLFW20], the authors are modeling camera effects to improve synthetic data for the task of bounding box detection. Metrics and parameter estimation of the effects from real camera images are suggested by [LLFW20] and 1For example, Carmaker from IPG or PreScan from TASS International. 2https://www.asam.net/standards/detail/openscenario/. 3https://www.blender.org/. 7.4. Publication 4 115
362 O. Grau et al. [CSVJR19]. A sensor model including sensor noise, lens blur, and chromatic aberration was developed based on real datasets [HG21] and integrated into our validation framework. Looking at virtual scene content, the most recent simulation systems for validation of a complete AD system include simulation and testing of the ego-motion of a virtual vehicle and its behavior. The used test content or scenarios are therefore aimed to simulate environments spanning a huge virtual space and are then virtually driving a high number of test miles (or km) in the virtual world provided [MBM18, WPC20,DG18]. Although this might be a good strategy to validate full AD stacks, one remaining problem for validation of perception systems is the limited coverage of data testing critical scene constellations (sometimes called ‘corner cases’) and parameters that lead to drop in performance of the DNN perception. A more suitable approach is to use probabilistic grammar systems [DKF20, WU18] to generate 3D scenarios which include a catalog of different object classes, and places them relative to each other to cover the complexity of the input domain. In this chapter we demonstrate the effectiveness of a simple probabilistic grammar system together with our previous scene parameter variation [SGH20] with a novel multi-stage strategy. This approach allows to systematically test conditions and relevant parameters for validation of perceptional function in a structured way. 3 Concept and Overview The novelty of the framework introduced in this chapter is the combination of modules for parameterized generation and testing of a wide range of scenarios and scene parameters as well as sensor parameters. It is tailored towards exploration of factors that (hypothetically) define and limit the performance of perception modules. A core design feature of the framework is the consequent parameterization of the scene composition, scene, and sensor parameters into a validation parameter space (VPS) as outlined in Sect.4.2. This parameterization only considers the near proximity of the ego-car or sensor; in other words, only the objects visible to the sensor are generated. This allows a much more well-defined test of constellations involving a specific number of object types, environment topology (e.g., types and dimensions of streets), and relation of objects, usually as an implicit function of where objects are positioned relative in the scene. This leads to a different data production and simulator pipeline than for conventional AV validation which typically provides a virtual world with a large extent to simulate and test the driving functions down to a physical level, inspired by real-world validation and test procedures [KP16,MBM18,DG18,JWKW18,WPC20]. Figure1shows the building blocks of our VALERIE system. The system runs an expansion of the VPS specified in the ‘validation task’ description. Our current implementation is based on a probabilistic description of how to generate the scene 7. Publications 116
A Variational Deep Synthesis Approach for Perception Validation 363 Validation Engineer Validation Automation Data Synthesis Asset Database Scenario Preparation VA LERIE Validation Flow Control Probabilistic Scene Generator Parameter Variation Generator Sensor & Environment Simulation Perception Function Evaluation Metric Ground Truth Fig. 1 Block diagram of the proposed validation approach and defines the parameter variations in the parameter space. In the future, the validation task should also include a more abstract target description of the evaluation metrics. The data synthesis block consists of three sub-components: The probabilistic scene generator generates a scene constellation, including a street layout, and places three-dimensional objects from the asset database according to probabilistic placement rules laid out in the scenario preparation. The parameter variation generator produces variations of that scene, including sun and light settings and variations of the placement of objects (see Fig.2for some examples). The sensor & environment simulation is using a rendering engine to compute a realistic simulation of the sensor impressions. Further, ground truth data is provided through the rendering process, which can be used for a pixel-accurate depth map (distance from camera to scene object) or meta data, like pixel-wise label identifiers of classes or object instances. Depending on the perception task, this information is specifically used for training and evaluation of semantic segmentation (see Sect.4.5). The output of the sensor simulation is passed to the perception function under test and the response to that data is computed. An evaluation metric specific to the validation task is based on the perception response. The ground truth data, as generated by the rendering process is usually required here, e.g., to compute the similarity to the known appearance of objects. In the experiments presented in this chapter we used known performance metrics for DNNs, such as the mean intersection-over-union (mIoU) metric, as introduced by [EVGW+15]. The parameterization along with the computation flow are described in detail in the next section. 7.4. Publication 4 117
370 O. Grau et al. provided by BIT-TS, a project partner of the KI-Absicherung project,7consisting of urban street scenes inspired by the preceding two real-world datasets. All of these datasets are labeled on a subset of 11 classes which are alike in these datasets to provide comparability between the results of the different trained and evaluated models. For the second task, the 2D-bounding box detection, we utilize the single-shot multibox detector (SSD) by [LAE+16], a 2D-bounding box detector trained on the synthetic data for pedestrian detection. This bounding box detector is applied on our variational data in Sect.5. To measure the performance of the task of semantic segmentation, the mean intersection-over-union (mIoU) from the COCO semantic segmentation benchmark task is used [LSD15]. The mIoU is denoted as the intersections between predicted semantic label classes and their corresponding ground truth divided by the union of the same, averaged over all classes. Another performance measure utilized is the pixel accuracy (p Acc) which is defined as follows: pAcc =T P +T N T P +F P +F N +T N .(1) The number of true positives (TP), true negatives (TN), false positives (FP), and true negatives (TN) are used to calculate pAcc, which can also be seen as a measure for correctly predicted pixels over all pixels considered for evaluation. For the 2D-bounding box detection we are interested in cases where, according to our definition, the performance-limiting factors are within bounds where the network should still be able to correctly predict a reasonable bounding box for each object to detect. For each synthesized and inferred image, the true positive rate (TPR) is calculated. The TPR is defined as the number of correctly detected objects (TP) over the sum of correctly detected and undetected objects (TP+FN). As we are interested in prediction failure cases we can then filter out all images with a true positive rate (TPR) of 1 and are left with images where the detection has omitted objects to detect. 4.6 Controller The VALERIE controller (as depicted in Fig.1, validation flow control) executes the validation run. This run can be configured in multiple ways depending on how much synthetic data is generated and evaluated. Two aspects have a major influence on this: First, the specification of parameters to be varied, and second, the used sampling strategy, which also depends on the validation goal. Both aspects are briefly described in the following. Specification of variable validation parameters: As outlined in Sect. 4.2, the approach depends on the provision of a generative scene model. This consists of a parameterized 3D scene model and includes 3D assets in the form of static and 7https://www.ki-absicherung-projekt.de/. 7. Publications 124
A Variational Deep Synthesis Approach for Perception Validation 371 dynamic objects. On top of this, we define variable parameters in this scene as an explicit list, as explained in Sect.4.2. For the specification of a validation run, all or a subset of these parameters are selected and a range and sampling distribution for that specific parameter is added. For example, to vary the x-position of a person in the scene along a line with the uniform or homogeneous distribution and a step size of 1 m, we define {p1, UNIFORM, 1.5, 5.5, 1.0} The parameters refer to the following: Parameter p1 refers to parameter declarations of xposition of person-1 in the example of Sect.4.2. The field UNIFORM refers to a uniform sampling distribution. Other modes include GAUSSIAN (Gaussian distribution). The parameters 1.5, 5.5, 1.0 refer to the parameter range [1.5...5.5] and the initial step size of 1m. Sampling of variable validation parameters: The actual expansion or sampling of the validation parameter space can be further configured and influenced in the VALERIE controller by selecting a sampler and validation strategy or goal. The sampler object provides an interface to the controller to the validation parameter space, considering the parameter ranges and optionally the expected parameter distribution. We support uniform and Gaussian distributions. In our current implementation, the controller can be configured to either sample the validation parameter space by a full grid search, or by a Monte-Carlo random sampling. However, the step size can be iteratively adapted depending on the validation goal. One option here is to automatically refine the search for edge cases (or corner cases) in the parameter space: As an edge case, we consider here a parameter instance, where the evaluation function is changing between an ‘acceptable’ state to a ‘failed’ state (using a continuous performance metric). For our use case of person detection, that means a drop in the performance metric below a threshold. Other validation goals we are planning to implement could be the automated determination of sensitive parameters or (ultimately) more intelligent search through high-dimensional validation parameter spaces. 4.7 Computational Aspects and System Scalability Our approach is designed for execution in data centers. The implementation of the components described above is modular and makes use of containerized modules using docker.8For the actual execution of the modules we use the Slurm9scheduling tool, which allows running our validation engine with a high number of variants in parallel, allowing the exploration of many states in the validation parameter space. 8www.docker.com. 9https://slurm.schedmd.com. 7.4. Publication 4 125
372 O. Grau et al. The results presented here are produced on an experimental setup using six dual Xeon server nodes, each equipped with 380 GB RAM. The runtime of the rendering process as outlined above is mainly determined by the rendering and in the order of 10...15min per frame, using the high-quality physically based rendering (PBR) Cycles render engine. 5 Evaluation Results and Discussion To evaluate the effectiveness of our data synthesis approach, we conducted experiments in generating scenes, variation of a few important parameters, and then we evaluated the perception performance including an analysis of performance-limiting factors, such as occlusions and distance to objects. We used our scene generator to generate variations of street crossings, as depicted in Fig.2. For these examples a base ground is generated first, with flexible topology (crossings, t-junction) and dimensions of streets, sidewalks, etc. In the next step, buildings, persons, and objects, including cars, traffic signs, etc. , are selected from a database and randomly placed by the scene generator, taking into account the probabilistic description and rules. The approach can handle any number of object assets. The current experimental setup includes a total of about 500 assets, with about 60 different buildings, 180 different person models, and other objects, including vegetation, vehicles, and so on. Scene parameter variation: Within the generated scenes, we vary the position and orientations of persons and some occluding objects. Further, we change the illumination by changing the time of the day. This has two main effects: First, it is changing the illumination intensity and color (dominant at sunset and sunrise), and second, it is generating a variation of shadows casted into the scene. In particular, from our experience, the latter creates challenging situations for the perception. Comparison of object distribution: Fig. 5shows the spatial distribution of persons in a) the Cityscapes dataset, b) KI-A tranche 3 dataset, and c) a dataset using our generative scene generator, as depicted in Figs.2and 3. The diagrams present a topview of the respective sensor (viewing cone) and the color encodes the frequency of persons within the sensor viewing cone, i.e., they give a representation of distance and direction of persons in all considered frames of the dataset. The real-world Cityscapes dataset has a distribution that corresponds with most persons located left and right of the center, i.e., the street. There are slightly more persons on the right side, which can be explained by the fact that often sidewalks on the left hand are occluded by vehicles from the other road side. The distribution of our dataset resembles as expected this distribution, with slightly less occupation in the distance. In contrast, the distribution of the KI-A tranche 3 dataset shows a very sharp cumulation of the distribution on what corresponds to a narrow band on the sidewalks of their 3D simulation. 7. Publications 126
A Variational Deep Synthesis Approach for Perception Validation 373 (a) (b) (c) Fig. 5 Pedestrian distribution over horizontal angle and distance. a: Cityscapes. b: KI-A tranche 3. c: Our synthetic data Influence of different occluding objects on detection performance: A number of object and attribute variations are depicted in Fig.6. On the left side, the SSD bounding box detector [LAE+16] is applied to the three images with different occluding objects in front of a pedestrian. In all three images, two bounding boxes are predicted for the same pedestrian. While one bounding box includes the whole body, the second bounding box only covers the non-occluded upper part of the pedestrian. On the right side, the DeeplabV3+ model trained on the KI-A tranche 3 is used to create a semantic map of the same three images. Besides the arms, the pedestrian is detected, even partially through the occluding fence. However, another interesting observation can be made: The ground the pedestrian stands on is always labeled as sidewalk. We interpret this as an indication to a bias in the training data, as the training data does not include enough images of pedestrians on the road, just on the sidewalk. This hypothesis can be further strengthened when we inspect the pedestrian distributions in Fig. 5b, where the pedestrians are distributed narrowly left and right off the street in the middle. Additionally, both bounding box prediction and the 7.4. Publication 4 127
374 O. Grau et al. Fig. 6 Scene with variation of occluding objects. Left: 2D bounding box detection. Right: semantic segmentation semantic segmentation do not include the pedestrian’s arms in their predictions. This can also be attributed to a bias in the training data. Influence of noise on detection performance: An experiment demonstrating our sensor simulation determines the influence of sensor noise on the predictive performance. In Fig.7, Gaussian noise with increasing variance is applied to an image, and three DeeplabV3+ models trained on A2D2, Cityscapes, and a synthetic dataset, respectively, are used to predict on the data. While image color pixels are represented in the range xi∈ [0,255], the noise variance is in the range of σ2∈ [0,20] with a step size of 1. For each noise variance step, the mIoU performance metric on the image prediction per model is calculated. While initially the models trained on Cityscapes and the synthetic dataset increase in performance, all models’ predictive 7. Publications 128
A Variational Deep Synthesis Approach for Perception Validation 375 0 2 4 6 8 10 12 14 16 18 20 noise variance 0 10 20 30 40 50 60 mIoU [%] DeeplabV3+ trained on A2D2 DeeplabV3+ trained on Cityscapes DeeplabV3+ trained on Synthetic data Fig. 7 Top: mIoU performance decreases with increasing noise variance. Bottom (left to right): segmentation maps with increasing noise variance σ2∈ {0,10,20}, image pixels xi∈ [0,255] performance ultimately decreases with an increasing level of sensor noise. The initial increase can be explained to stem from the domain shift of training to validation data, where in the training data a small noise variance can be observed. Analysis of performance-limiting factors: Some scene parameters have a major influence on the perception performance. This includes the occlusion rate of objects, with totally occluded objects that are obviously not detectable or the object size (in pixels) in the images, also with a natural boundary where detection breaks up if the object size is too small. Other performance-limiting factors include contrast and other physically observable parameters. To measure the influence or sensitivity of perception functions against performance-limiting factors we designed an experiment using about 30,000 frames containing one person each. The person is moved and rotated on the sidewalk and on the street. The occlusions are determined by rendering a mask of the person and comparison with the actual instance mask considering occluding objects. A degree of 100% represents a fully occluded object. Figure8shows results of this experiment, each gray dot representing one frame and the colored curves showing regression plots with differently clothed persons. The figure shows a pAcc downwards trend with increasing occlusion rates. The Detectron2 model (trained on Cityscapes) is comparatively more robust than 7.4. Publication 4 129
376 O. Grau et al. Fig. 8 Polynomial regression curves (of order 3) on pedestrian detection rate pAcc of DeeplabV3+ (left) and Detectron2 (right) for various occlusion rates of a pedestrian wearing dark or bright clothes DeeplabV3. The plot shows that Detectron2 offers stable detection with occlusion rates <35% and then the performance drops. DeeplabV3’s (also trained on Cityscapes) performance drops after 15% occlusion rate. The curves are not linearly following a trend due to the fact that there are other scene parameters (sunlight, shadow, direction of pedestrian) which are not constant across the rendered images. What can also be seen in the figures is that, despite the trend of the regression curves, there is a great variation in the data—visible by the widely scattered grey points. That means that the performance depends also on other factors besides the occlusion rate. Figure8is showing one example of analysis possible with the metadata provided by our framework. More parameters are considered in our previous work [SGH20]. Data bias analysis: Another experiment we conducted considers failure cases, i.e., false negatives (FN) of the SSD 2D-bounding box detector regarding pedestrian detection. To accomplish this, we rendered 2640 images with our variational data synthesis engine. These images are then inferred by the SSD model and evaluated. Only pedestrians with a bounding box width greater than 0.1 ×image width and a height of 0.1 ×image height are considered valid for evaluation. Additionally, only objects with an occlusion rate below 25% are considered valid. These restrictions guarantee that pedestrians in the validation are of sufficient size, i.e., close to the camera, and clearly visible due to little occlusion and would therefore be easy to detect. 7. Publications 130
A Variational Deep Synthesis Approach for Perception Validation 377 0 10 20 30 40 Count Detected pedestrian Non-detected pedestrian ID: 1 2 345 6 Clothing: Veiled Casual Veiled Paramedic Business Physician Ethnicity: Arabian Caucasian Arabian Caucasian Caucasian Caucasian Gender: Woman Woman Man Man Man Man Fig. 9 Count of detected and non-detected pedestrians for different pedestrian assets, i.e., different clothing, ethnicity, and gender With these restrictions in place we found that from all the pedestrian assets the synthesis engine placed in the scene, there were six assets that were omitted by the SSD model as can be seen in Fig.9. The asset ID 1 is an Arabian woman wearing traditional clothes effectively veiling the person. Asset ID 2 is a Caucasian woman clothed in summer casual, i.e., short pants and short sleeves, revealing parts of her skin. The second Arabian ethnicity asset with the ID 3 is similar to asset 1 clothed in traditional veiling clothes but of male gender. The remaining assets 4, 5, and 6 are of male gender and Caucasian ethnicity wearing different work clothes, i.e., a blue paramedical outfit for ID 4, business casual jeans and jacket for ID 5, and white physician clothes for ID 6. The asset ID 2 with the summer casual clothed woman is only miss-detected a few times, in most cases the detection worked well, indicating no data bias for this asset. In contrast, the pedestrian asset ID 6 of a physician dressed in white hospital clothing has not been detected at all. Additionally, two of the assets that were relatively most often overlooked by the network are the Arabian clothed woman with asset ID 1, as well as an Arabian clothed man with the ID 3. This result would suggest that these kind of pedestrian assets, i.e., IDs 1, 3, and 6, were not present in the data for training the model and adding them to it will lead to a mitigation of this exact failure case. 7.4. Publication 4 131
378 O. Grau et al. 6 Outlook and Conclusions This chapter has introduced a new generative data synthesis framework for the validation of machine learning-based perception functions. The approach allows a very flexible description of scenes and parameters to be varied and systematical tests of parameter variations in our unified validation parameter space. The conducted experiments demonstrate the benefits of splitting the validation process into scene variation that looks into randomized placement of objects and a variation of scene parameters and sensor simulation. Our simple probabilistic scene generator is scalable and able to produce scenes with a high number of different objects—as provided by an asset database. The spatial distribution of the positioned objects, as demonstrated for persons in Fig.5, is more realistic compared to manually crafted 3D scenes. Along with our sensor simulation (results discussed in the chapter ‘Optimized Data Synthesis for DNN Training and Validation by Sensor Artifact Simulation’ [HG22]), we present a step to close the domain-gap between synthetic and real data. Future work will continue to analyze the influence of other factors, such as rendering fidelity, scene complexity, and composition, to further improve the capabilities of the framework and make it even more applicable for the validation of real-world AI functions. Our experiments with performance-limiting factors, as shown for occlusion rates and object size (as a function of distance to the camera) in the previous section gives clear evidence that the performance of perception functions cannot be characterized by only a few factors. It is, however, a complex function of many parameters and aspects, including scene complexity, scene lighting and weather conditions, and the sensor characteristics. The deep validation approach described in this chapter is addressing this multi-dimensional complexity problem and we designed a system and methodology for flexible validation strategies to span all these parameters at once. Our validation parameterization, as demonstrated in the results section, is an effective way to detect performance problems in perception functions. Moreover, it allows in its flexible design the sampling and a practical computation at scale allowing for deep exploration of the multi-variate validation parameter space. Therefore, we see our system as a valuable tool for the validation of perception functions. Moving forward we are looking into using the deep synthesis approach to implement sophisticated algorithms to support more complex validation strategies. As another key direction we target improvements in the computational efficiency of our validation approach, allowing coverage of more complexity and parameter dimensions. Acknowledgements The research leading to these results is funded by the German Federal Ministry for Economic Affairs and Energy within the project ‘Methoden und Maßnahmen zur Absicherung von KI-basierten Wahrnehmungsfunktionen für das automatisierte Fahren (KI Absicherung)’. The authors would like to thank the consortium for the successful cooperation. 7. Publications 132
A Variational Deep Synthesis Approach for Perception Validation 379 References [BB95] W. Burger, M.J. Barth, Virtual reality for enhanced computer vision, in Virtual Prototyping: Virtual Environments and the Product Design Process, ed. by J. Rix, S. Haas, J. Teixeira (Springer, 1995), pp. 247–257 [COR+16] M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, B. Schiele, The Cityscapes dataset for semantic urban scene understanding, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (Las Vegas, NV, USA, 2016), pp. 3213–3223 [CPK+18] L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, A.L. Yuille, DeepLab: semantic image segmentation with deep convolutional nets, Atrous convolution, and fully connected CRFs. IEEE Trans. Pattern Anal. Mach. Intell. (TPAMI) 40(4), 834–848 (2018) [CSVJR18] A. Carlson, K.A. Skinner, R. Vasudevan, M. Johnson-Roberson, Modeling camera effects to improve visual learning from synthetic data, in Proceedings of the European Conference on Computer Vision (ECCV) Workshops (Munich, Germany, 2018), pp. 505–520 [CSVJR19] A. Carlson, K.A. Skinner, R. Vasudevan, M. Johnson-Roberson, Sensor transfer: learning optimal sensor effect image augmentation for sim-to-real domain adaptation, pp. 1–8 (2019). arXiv:1809.06256 [DG18] W. Damm, R. Galbas, Exploiting learning and scenario-based specification languages for the verification and validation of highly automated driving, in Proceedings of the IEEE/ACM International Workshop on Software Engineering for AI in Autonomous Systems (SEFAIAS) (Gothenburg, Sweden, 2018), pp. 39–46 [DKF20] J. Devaranjan, A. Kar, S. Fidler, Meta-Sim2: learning to generate synthetic datasets, in Proceedings of the European Conference on Computer Vision (ECCV) (Virtual conference, 2020), pp. 715–733 [DRC+17] A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, V. Koltun, CARLA: an open urban driving simulator, in Proceedings of the Conference on Robot Learning CORL (Mountain View, CA, USA, 2017), pp. 1–16 [Epi04] Epic Games, Inc. Unreal Engine Homepage (2004). [Online; accessed 2021-11-18] [EVGW+15] M. Everingham, L.V. Gool, C.K.I. Williams, J. Winn, A. Zisserman, The pascal visual object classes challenge: a retrospective. Int. J. Comput. Vis. (IJCV) 111(1), 98–136 (2015) [GKM+20] J. Geyer, Y. Kassahun, M. Mahmudi, X. Ricou, R. Durgesh, A.S. Chung, L. Hauswald, V.H. Pham, M. Mühlegg, S. Dorn, T. Fernandez, M. Jänicke, S. Mirashi, C. Savani, M. Sturm, O. Vorobiov, M. Oelker, S. Garreis, P. Schuberth, A2D2: Audi autonomous driving dataset (2020), (pp. 1–10). arXiv:2004.06320 [HG21] K. Hagn, O. Grau, Improved sensor model for realistic synthetic data generation, in Proceedings of the ACM Computer Science in Cars Symposium (CSCS) (Virtual Conference, 2021), pp. 1–9 [HG22] K. Hagn, O. Grau, Optimized data synthesis for DNN training and validation by sensor artifact simulation, in Deep Neural Networks and Data for Automated Driving— Robustness, Uncertainty Quantification, and Insights Towards Safety ed. by T. Fingscheidt, H. Gottschalk, S. Houben (Springer, 2022), pp. 149–170 [HZRS16] K. He, X. Zhang, S. Ren, J. Sun, Deep residual learning for image recognition, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (Las Vegas, NV, USA, 2016), pp. 770–778 [JWKW18] P. Junietz, W. Wachenfeld, K. Klonecki, H. Winner, Evaluation of different approaches to address safety validation of automated driving, in Proceedings of the IEEE Intelligent Transportation Systems Conference (ITSC) (Maui, HI, USA, 2018), pp. 491–496 7.4. Publication 4 133