scieee AI-readable full text Open interactive document viewer

Interpretability of deep learning models

Domingo Gregorio, Pablo

Abstract

In recent years we have seen growth on interest for Deep Learning (DL) algorithms on a variety of problems, due to their outstanding performance. This is more palpable on a multitude of fields, where self-learning algorithms are becoming indispensable tools to help professionals solve complex problems. However as these models are getting better, they also tend to be more complex and are sometimes referred to as "Black Boxes". The lack of explanations for the resulting predictions and the inability of humans to understand those decisions seems problematic. In this project, different methods to increase the interpretability of Deep Neural Networks (DNN) such as Convolutional Neural Network (CNN) are studied. Additionally, how these interpretability methods or techniques can be implemented, evaluated and applied to real-world problems, by creating a python ToolBox.

Full text

Interpretability of Deep Learning Models A Degree Thesis submitted to the faculty of Escola T`ecnica Superior d’Enginyeria de Telecomunicaci´o de Barcelona Universitat Polit`ecnica de Catalunya by Pablo Domingo Gregorio In partial fulfilment of the requirements for the Degree of Telecommunications Technologies and Services Major in Audiovisual Systems Advisor: Ver´onica Vilaplana UPC ·ETSETB Barcelona, June 2019 Acknowledgements I would like to express my sincere gratitude to my advisor Ver´onica Vilaplana for her guidance, patience and motivation during the development of this thesis. But also for giving me the opportunity to work in a field that I am passionate about. I also would like to appreciate the continuous encouragement and understanding of my friends and family, especially my grandfather for encouraging my curiosity and my father to instil the critical spirit. I Abstract In recent years we have seen growth on interest for Deep Learning (DL) algorithms on a variety of problems, due to their outstanding performance. This is more palpable on a multitude of fields, where self learning algorithms are becoming indispensable tools to help professionals solve complex problems. However as these models are getting better, they also tend to be more complex and are sometimes referred to as ”Black Boxes”. The lack of explanations for the resulting predictions and the inability of humans to understand those decisions seems problematic. In this project, different methods to increase the interpretability of Deep Neural Networks (DNN) such as Convolutional Neural Network (CNN) are studied. Additionally, we evaluate how these interpretability methods or techniques can be implemented, measured and applied to real-world problems, by creating a python ToolBox. II Resum En els darrers anys hem vist un creixement en l’inter`es pels algoritmes d’aprenentatge profund (DL) a una varietat de problemes, a causa del seu excel·lent rendiment. Aix`o ´es m´es palpable en una multitud de camps, on els algoritmes d’autoaprenentatge s’estan convertint en eines per ajudar als professionals resoldre problemes complexes. Tanmateix, a mesura que aquests models milloren, tamb´e aumenta la seva complexitat, i avegades es denominen ”Caixes Negres”. La manca d’explicacions per a les prediccions resultants i la incapacitat dels humans per entendre aquestes decisions sembla problem`atica. En aquest projecte, s’estudien diferents m`etodes per augmentar la interpretabilitat de Xarxes Neuronals Profundes (DNN) com ara les Xarxes Neuronals Convolucionals (CNN). A m´es, avaluem com es poden implementar i aplicar aquests m`etodes o t`ecniques d’interpretaci´o a problemes reals, creant una Toolbox de python. III Resumen En los ´ultimos a˜nos, hemos visto un crecimiento de inter´es por los algoritmos de aprendizaje profundo (DL) en una variedad de problemas, debido a su excelente rendimiento. Esto es m´as palpable en una multitud de campos, donde los algoritmos de autoaprendizaje se estan convirtiendo en una herramienta indispensable para ayudar a profesionales resolver problemas complejos. Sin embargo, a medida que estos modelos mejoran, tambi´en aumenta su complejidad, y en ocasiones se los denomina ”Cajas Negras”. La falta de explicaciones para las predicciones resultantes y la incapacidad de los humanos para entender esas decisiones parece problem´atica. En este trabajo, se estudian diferentes m´etodos para aumentar la interpretabilidad de Redes Neuronales Profundas (DNN) como las Redes Neuronales Convolucionales (CNN). Adem´as, evaluamos c´omo estos m´etodos o t´ecnicas de interpretabilidad pueden implementarse y aplicarse a problemas reales, mediante la creaci´on de una ToolBox de python. IV Contents 1 Introduction 1 1.1 Context .................................. 1 1.2 ProblemStatement............................ 1 1.3 ProjectOverview............................. 2 2 Theoric Background 3 2.1 Building Blocks of Neural Networks . . . . . . . . . . . . . . . . . . . 3 2.2 ActivationFunctions ........................... 3 2.3 LearningMechanism ........................... 4 2.4 Neural Networks Architectures . . . . . . . . . . . . . . . . . . . . . . 5 3 Interpretability Methods 6 3.1 SurrogateMethods ............................ 6 3.2 FunctionMethods............................. 7 3.2.1 Activation Visualization . . . . . . . . . . . . . . . . . . . . . 7 3.2.2 Gradient.............................. 7 3.2.3 SmoothGrad ........................... 7 3.2.4 Integrated Gradients . . . . . . . . . . . . . . . . . . . . . . . 8 3.3 SignalMethods .............................. 8 3.3.1 Activation Maximization (AM) . . . . . . . . . . . . . . . . . 8 3.3.2 Deconvolution........................... 9 3.3.3 Guided BackPropagation . . . . . . . . . . . . . . . . . . . . . 9 3.4 Attribution Methods . . . . . . . . . . . . . . . . . . . . . . . . . . . 10 3.4.1 Layer-wise Relevance Propagation . . . . . . . . . . . . . . . . 10 3.4.2 Other Attribution methods . . . . . . . . . . . . . . . . . . . . 10 3.5 LocalizationMethods........................... 11 3.5.1 GradCAM and Guided GradCAM . . . . . . . . . . . . . . . . 11 3.5.2 OcclusionMap .......................... 12 3.6 DistanceMethods............................. 12 3.6.1 Distance Robustness . . . . . . . . . . . . . . . . . . . . . . . 12 3.6.2 Similar Distance Examples . . . . . . . . . . . . . . . . . . . . 12 4 Experiments and Results 13 4.1 Experiment Considerations . . . . . . . . . . . . . . . . . . . . . . . . 13 4.2 FunctionMethods............................. 14 4.2.1 Activation Visualization . . . . . . . . . . . . . . . . . . . . . 14 4.2.2 Gradient.............................. 15 4.2.3 SmoothGrad............................ 16 V 4.2.4 Integrated Gradients . . . . . . . . . . . . . . . . . . . . . . . 16 4.3 SignalMethods .............................. 17 4.3.1 Activation Maximization . . . . . . . . . . . . . . . . . . . . . 17 4.3.2 Deconvolution........................... 19 4.3.3 Guided BackPropagation . . . . . . . . . . . . . . . . . . . . . 20 4.4 Attribution Methods . . . . . . . . . . . . . . . . . . . . . . . . . . . 21 4.4.1 LayerWise Relevance Propagation (LRP) . . . . . . . . . . . . 21 4.5 LocalizationMethods........................... 23 4.5.1 GradCAM and Guided GradCAM . . . . . . . . . . . . . . . . 23 4.5.2 OcclusionMaps.......................... 25 4.6 DistanceMethods............................. 26 4.6.1 Distance Robustness . . . . . . . . . . . . . . . . . . . . . . . 26 4.7 Discussion................................. 26 5 Budget 27 6 Conclusions 28 7 Future Work 29 Bibliography 30 Appendix A Open source Interpretability ToolBox 32 Appendix B Additional Details and Results 33 B.1 Activation Visualization Results . . . . . . . . . . . . . . . . . . . . . 33 B.2 GradientResults ............................. 34 B.3 SmoothGradResults ........................... 35 B.4 Integrated Gradients Results . . . . . . . . . . . . . . . . . . . . . . . 36 B.5 Nguyen Solution for Activation Maximization (AM) . . . . . . . . . . 36 B.6 Activation Maximization Results . . . . . . . . . . . . . . . . . . . . 37 B.7 Deconvolution Results . . . . . . . . . . . . . . . . . . . . . . . . . . 38 B.8 Guided BackPropagation Results . . . . . . . . . . . . . . . . . . . . 38 B.9 Layer-Wise Relevance Propagation Results . . . . . . . . . . . . . . . 39 B.10 GradCAM and GuidedGradCAM Results . . . . . . . . . . . . . . . . 41 B.11 Occlusion Maps Results . . . . . . . . . . . . . . . . . . . . . . . . . 42 Appendix C Work Plan 43 C.1 TasksandMilestones........................... 43 C.2 GanttDiagrams.............................. 46 List of Figures 2.1 Neural Networks Structures . . . . . . . . . . . . . . . . . . . . . . . 3 2.2 Forward and Backward Propagation . . . . . . . . . . . . . . . . . . . 4 2.3 Deep Neural Network Architecture . . . . . . . . . . . . . . . . . . . 5 2.4 CNNlayeroperations........................... 5 3.1 DeConvNetStructure........................... 9 3.2 Suppression of Negative Influences . . . . . . . . . . . . . . . . . . . . 9 3.3 Layer-wise Relevance Propagation (LRP) Process . . . . . . . . . . . 10 3.4 GradCAMScheme ............................ 11 3.5 ClustersofClasses ............................ 12 4.1 VGG16Architecture ........................... 13 4.2 Visualization of different Feature Maps . . . . . . . . . . . . . . . . . 14 4.3 Gradient for last Convolution Layers . . . . . . . . . . . . . . . . . . 15 4.4 SmoothGrad on Warplane . . . . . . . . . . . . . . . . . . . . . . . . 16 4.5 Score for different Luminance Intesity values . . . . . . . . . . . . . . 16 4.6 Integrated Gradients on Wolf . . . . . . . . . . . . . . . . . . . . . . 17 4.7 Activation Maximization for Feature Maps . . . . . . . . . . . . . . . 18 4.8 Activation Maximization for Classes . . . . . . . . . . . . . . . . . . . 19 4.9 Deconvolution for Wolf Image . . . . . . . . . . . . . . . . . . . . . . 20 4.10 Guided BackPropagation for Wolf Image . . . . . . . . . . . . . . . . 20 4.11 LRP on Lion & Tiger for α=1 ..................... 22 4.12 Failed LRP for α= 2, β =1 ....................... 22 4.13 Alternative GradCAM System . . . . . . . . . . . . . . . . . . . . . . 23 4.14 Explanation of Tiger . . . . . . . . . . . . . . . . . . . . . . . . . . . 24 4.15 Negative Explanation of Tiger . . . . . . . . . . . . . . . . . . . . . . 24 4.16 Explanation of Warplane . . . . . . . . . . . . . . . . . . . . . . . . . 24 4.17 Negative Explanation of Warplane . . . . . . . . . . . . . . . . . . . . 25 4.18 Occlusion Map for the Class Cheetah . . . . . . . . . . . . . . . . . . 25 B.1 Feature Maps of Second Convolution Layers . . . . . . . . . . . . . . 33 B.2 Feature Maps from the Last three Convolution Layers . . . . . . . . . 33 B.3 GradientonWarplane .......................... 34 B.4 GradientonFly.............................. 34 B.5 GradientonLion ............................. 34 B.6 SmoothGradonGoat........................... 35 B.7 SmoothGradonWolf........................... 35 B.8 SmoothGrad on Basketball . . . . . . . . . . . . . . . . . . . . . . . . 35 VII B.9 Integrated Gradients on Lion . . . . . . . . . . . . . . . . . . . . . . 36 B.10 Integrated Gradients on Hammerhead . . . . . . . . . . . . . . . . . . 36 B.11NguyenScheme.............................. 36 B.12AMforClasses .............................. 37 B.13 Deconvolution Examples . . . . . . . . . . . . . . . . . . . . . . . . . 38 B.14 Guided on Hammerhead . . . . . . . . . . . . . . . . . . . . . . . . . 38 B.15GuidedonWarplane ........................... 39 B.16 Guided on Lion & Tiger . . . . . . . . . . . . . . . . . . . . . . . . . 39 B.17LRPonHammerhead........................... 39 B.18LRPonWolf ............................... 40 B.19LRPonBasketball ............................ 40 B.20 Explanation of Basketball . . . . . . . . . . . . . . . . . . . . . . . . 41 B.21 Explanation of Baseball . . . . . . . . . . . . . . . . . . . . . . . . . 41 B.22 Explanation of Shark . . . . . . . . . . . . . . . . . . . . . . . . . . . 41 B.23 Explanation of Hammerhead . . . . . . . . . . . . . . . . . . . . . . . 42 B.24 Occlusion Map for the Class Warplane . . . . . . . . . . . . . . . . . 42 B.25 Occlusion Map for the Class Tiger . . . . . . . . . . . . . . . . . . . . 42 C.1 Gantt Diagram First Part . . . . . . . . . . . . . . . . . . . . . . . . 46 C.2 Gantt Diagram Second Part . . . . . . . . . . . . . . . . . . . . . . . 46 Chapter 2 Theoric Background Then using that gradient the weights are updated using the equation 2.5, where λis a pre-defined value called learning rate. ωi=ωi−λ·∂Γ(θ) ∂ω (2.5) Finally using the chain rule the gradients of the next layer are performed. This process is done iteratively backwards until the first layer is reached. 2.4 Neural Networks Architectures Depending on the problem to be solved, some architectures will work better than others. In particular, for tabular data DNN are frequently used. These networks are formed by significant amounts of layers known as fully connected layers. These layers, also denominated Dense layers due to the vectorized structure of their neurons, use dot product to obtain the activations. Figure 2.3: Deep Neural Network Architecture For other problems, such as image classification and segmentation, CNN architectures are preferred. They are very similar to NNs. However, unlike ordinary NN, CNN try to reduce the amount of parameters by using local connectivity and sharing parameters between neurons by using Convolution and Pooling layers, where matrix multiplication and down-sampling are used respectively. (a) Convolution Operation (b) MaxPooling Operation Figure 2.4: CNN layer operations Convolution layers contain weight matrices (referred to as filters). The activations of these layers (known as feature maps) apart from extracting spatial information are the results of applying convolutions between the input and these filters. These feature maps are passed to the Pooling layers, in order to reduce the dimensionality and achieve some invariance to translation on the images. Finally, to obtain the predictions of the inputs Dense layers are employed. 5 3. Interpretability Methods To better clarify how interpretability methods work, this project has divided them into six categories inspired from a classification made by Kindermans [5] which divide them into three types. Here three more categories are included with the same objective, separate them based on what these methods try to explain. Surrogate: Attempt to estimate the NN model by using interpretable models, such as linear regression or decision trees, to describe locally or globally the function of the NN. Function: Highlight parts which (when changed) could increase or decrease the output of a certain layer. In other words, try to show how the output changes as the input changes. Signal: Show what input patterns originally caused a given activation on a layer. Can also be seen as a mapping of layer outputs back to the input space. Attribution: Describe how much an individual component influences or contributes to the output, also known as relevance. Localization: Map where the NN is focusing in order to decide between classes. Generate heat maps indicating zones of attention. Distance: Based on the principle that similar inputs should produce similar outcomes, tries to build clusters of classes and present the distance between them. 3.1 Surrogate Methods These methods try to approximate locally or globally the model under study by training other models more interpretable, such as Linear Models or Decisions Trees. Global Surrogate and Local Interpretable Model Explanations (LIME)[6] are the most known. The approximation is achieved by following a series of steps. Acquire the predictions of the complex model from a selected dataset. Select and train an interpretable model (such as linear regression or decision trees). Measure how well the surrogate model is approximating the complex model. Explain how the interpretable model functions. 6 Chapter 3 Interpretability Methods Even though their implementation is very straightforward, they tend to generalize excessively leaving out important distinctions that the model could have done. In addition, this interpretable models have to be trained and possible training errors could appear. 3.2 Function Methods 3.2.1 Activation Visualization This technique does exactly what it says, takes all the activations from all the layers and present them. It is only applied to models where their activations can be visualized for example Decision Trees or CNN. Once the activations are obtained, an evolution of the activations across all layers can be mapped and determine how the activations shaped the final decision. Depending on the model architectures more precise visualization techniques can be used. 3.2.2 Gradient Uses forward and backward propagation to determine which parts of the input would cause an update on the weights of a selected layer, and consequently the activations of that layer. To measure it, the process defined on 2.2 is calculated, where the loss function Γ(θ) chosen is the mean of the activations Afrom the selected layer. Gradient =∂Γ(θ) = ∂A ∂X (3.1) The result obtained by gradients is also known as saliency map. The disadvantage of working with gradients is that the deeper the network is, the noisier and unstable the results become. 3.2.3 Smooth Grad This method attempts to reduce the noise from the results of Gradient by adding noise [7]. This is achieved by applying several times the Gradient method, where each time a Gaussian noise with a certain deviation is added to the input Xand then those results are averaged to obtain a smoother result. SmoothGrad =1 MX m ∂Γ(θ) = ∂A ∂(X+N(0, σ2)) (3.2) 7 Chapter 3 Interpretability Methods 3.2.4 Integrated Gradients As their fellow companion SmoothGrad, this method tries to improve the results of Gradient by computing the integral of the gradients between a baseline value and a limit point [8]. Actually, calculating the integral of the gradients is intractable, so instead, the method uses the Riemann sum, where samples are constructed around an interval between a baseline and a limit, using linear interpolation. Then gradients are calculated for every generated sample and an average sum is used to get the saliency map. As greater the number of samples in between that interval, the more accurate the resulting saliency map will be. IntGrad =1 N· N X n=1 ∂Γ(θ) = ∂A ∂Xn (3.3) For image-oriented NN it is common to generate samples where the luminance goes from zero to the maximum luminance on the input image, in other words the original image. 3.3 Signal Methods 3.3.1 Activation Maximization (AM) Proposed by Simonyan [9] this technique tries to unravel which concepts the DNN has built by identifying which input patterns maximize the activations of a certain layer. To do so, noise is introduced to the model where forward and backward propagation are used, but instead of updating the weights, the updates are applied to the noise input through gradient descent. For CNN this method can be used to visualize which input patterns causes certain feature maps or classes to be activated. In the case of the feature maps, it is possible to appreciate what shapes are captured by the network. In the case of classes, it is useful to guarantee that the characteristics shown match with the class maximized. The results obtained sometimes can be hard to figure out and different modifications have been proposed to enhance the quality of these results. Test different gradient descent techniques, such as Adam [10], Momentum or Nesterov [11]. Apply gaussian or median filters to blur the input during the iteration. Using Generative Adversarial Networks (GAN) to generate better input priors, instead of using Gaussian noise, as Nguyen suggests [12]. More details on the topic can be found on B.5 8 Chapter 3 Interpretability Methods 3.3.2 Deconvolution Mainly used in CNN, the idea is to attach a DeConvNet [13] to the CNN to reconstruct an input back from the activations of a selected layer. This DeConvNet is responsible for reversing all the processes of the CNN by defining two phases. Up Phase: similar to the forward propagation, except that recollects important information called switches at each layer. Down Phase: similar to the backward propagation, except that uses the switches from the up phase to inverse the process done at each layer. Figure 3.1: DeConvNet Structure 3.3.3 Guided BackPropagation Attempts to map the activations of a certain layer back to the input space by suppressing all the negative influences produced on the forward and backward propagation [14]. Then a trace of just the active neurons and positive influences can be obtained. Positive and negative influences are variations that indicate what a neuron detects (positive) or neglects (negative). Figure 3.2: Suppression of Negative Influences 9 Chapter 3 Interpretability Methods 3.4 Attribution Methods 3.4.1 Layer-wise Relevance Propagation LRP [15] is a layer-wise backward propagation technique, where each neuron of the network receives a share of the model output and gets redistributed to its predecessors until the input variables are reached. Two phases are executed. Forward computation: where all the activations caused by a certain input are recollected. Relevance propagation: where conservative mechanisms are used to compute the share of relevance between layers. Figure 3.3: LRP Process The relevance propagation is done through rules, which take the relevance Rkfrom a neuron kand computes the shares across all the neurons from the lower layer to obtain the relevance Rj. The conservation property imposes the equation 3.4 X j Rj=Rk(3.4) The most common rule is the αβRule which allows the appearance of positive and negative influences by giving αmore or less weight, as long as the conditions α−β= 1 and β≥0 are satisfied. The αβRule is defined in 3.5, where ajare the activations of the lower layer, and ω+ jk, ω− jk are the maximum and minimum of the weights from the upper layer. Rj=X k αajω+ jk Pjajω+ jk −βajω− jk Pjajω− jk !·Rk(3.5) 3.4.2 Other Attribution methods Apart from LRP, other procedures have been contemplated to capture the relevance of individual components. For example, methods like DeepLIFT [16] and SHAP 10 Chapter 3 Interpretability Methods values [17] measure the impact that altering an individual feature from a baseline value has on the prediction outcome. However, The most recent PatterNet and PatternAttribution [5] focus on estimating the value α+, which indicates the direction of the significant information from the data. On equation 3.6 xdenote the layer inputs, ythe layer outputs and ωTthe weights of the current layer. α+=E[x·y]−E[x]·E[y] ωT·E[x·y]−ωT·E[x]·E[y](3.6) Both compute the gradients using a layer-wise back propagation as LRP does, but PatterNet replace the weights ωby the directions α+and PatternAttribution replace them by ωα+where is the element-wise multiplication. 3.5 Localization Methods 3.5.1 GradCAM and Guided GradCAM It is known that as the depth of CNN increases, higher-level visual concepts are captured [18], but this spatial information is lost in fully connected layers. Grad Class Activation Map (CAM) [19] exploits this fact by replacing all dense layers with Global Average Pooling (GAP) layers, leaving just one with a softmax activation. Such modification requires new training for the GAP weights ωk. Once the weights have been trained, forward propagation is employed to acquire the feature maps Fkof the last convolutional layer. Then a weighted sum is performed between the associated weights ωkof a pre-selected neuron from the prediction dense layer, and the feature maps Fk. Figure 3.4: GradCAM Scheme As GradCAM lacks the ability to capture fine-grained details, it can be combined with Guided BackPropagation by multiplying both results to obtain what is called as GuidedGradCAM. This is possible because GradCAM heat maps vary between values from 0 to 1. 11 Chapter 3 Interpretability Methods 3.5.2 Occlusion Map Mostly used in image-oriented NN, Occlusion Map is an iterative method that recollects how much the score of a certain class drops in percentage by sliding a window which masks out parts of the input image. A heat map is created showing where the meaningful zones of the input image the model needs to predict correctly. 3.6 Distance Methods 3.6.1 Distance Robustness Constructs class clusters from input data points or predictions made by the model and identifies how close or distant these groupings are [20]. The most efficient is t-SNE which first create probability distributions in high-dimensional space to then use a t-Student distribution to recreate the probability distribution in lowdimensional space. Finally optimizes the embeddings using gradient descent. It is useful to check if similar classes are distant enough to ensure that the model can really distinguish between them. Figure 3.5: Clusters of Classes 3.6.2 Similar Distance Examples Its intention is to show training examples that are similar to new data. It builds a database from predictions obtained through forward propagation on training examples, to then compare them with the predictions of test data points. Finally, with distance metrics it is possible to present the closer candidates from the database. 12 4. Experiments and Results 4.1 Experiment Considerations Before delving into the results obtained in this project, different considerations and decisions have been taken to carry out the experiments. The interpretability methods previously explained assume that a pre-trained model is being analyzed. Both the model and the database chosen are commonly used for DL research to increase the comprehension and reproducibility of the studied methods. The model selected is a CNN known as VGG16, pre-trained on the ImageNet [21] dataset. VGG16 has been obtained from Keras framework [2] who offers a wide range of pre-trained models. The VGG16 architecture is made out of 5 blocks of Convolutional layers, where at the end of each block there is a MaxPooling layer. This causes the image dimensions to shrink at half at the end of each block, while the number of filters duplicates. Figure 4.1: VGG16 Architecture ImageNet is a large visual database aimed to be applied on visual object recognition software research. More than 14 million images are hand-annotated indicating what objects are pictured. In total ImageNet contains more than 20 thousand categories with a typical category, such as ”lion” or ”tiger”, consisting of several hundred images. The database offers image URLs that are freely available directly from their 13 Chapter 4 Experiments and Results webpage, though the actual images are not owned by ImageNet. A set of 117 images have been collected to discuss the results of the methods. This images were extracted from the ImageNet webpage [21]. One of the objectives of this project was to showcase as many methods as possible. However, some methods need restructuring the model architecture or train an entire new model. Therefore, that kind of methods have been discarded for the experiments, due to the amount of time that the implementation and the obtainment of the results would require. Methods such as Surrogate 3.1 or the Attribution methods explained on 3.4.2 have been excluded. Others like Distance methods 3.6 have only been considered for exclusive classes because otherwise these methods need vast amounts of data from different classes to produce conclusive results. 4.2 Function Methods 4.2.1 Activation Visualization Being VGG16 the model under study, the feature maps from some Conv layers have been obtained and presented [22]. The activations from other layers don’t offer as much information as Conv layers provide, since Conv layers are specialized on object detection an image segmentation. (a) Block:1 - Conv:2 (b) Block:3 - Conv:2 (c) Block:5 - Conv:2 Figure 4.2: Visualization of different Feature Maps Given that each Conv layer has multitude filters, feature maps can be visualized from different blocks. With these images, one could determine what patterns are filters extracting from the input image, and in consequence, how the feature maps look. The pixels in black represent zeros and grey areas are values different from zero. More results can be found on B.1. This is useful to see that on the first blocks Conv layers are using simple morphological operations to identify contours and basic shapes 4.2a whereas on deep Conv layers more complex operations like segmentation and object detection are performed 4.2b. However, for deeper feature maps the image definition is low which makes it harder to extract conclusions and interpret the results, due to Pool layers. This effect can be seen on 4.2c. 14 Chapter 4 Experiments and Results Besides the great visualization that Guided BackPropagation offers, it has difficulties distinguishing different classes when two or more appear on the same image as seen on Figure B.16. Other examples can be found on B.8. 4.4 Attribution Methods 4.4.1 LayerWise Relevance Propagation (LRP) LRP focus on assigning a relevance value to each pixel to determine how much a single pixel has contributed to the final decision. On 3.4.1 two processes have been described, although the relevance propagation is the trickiest to implement. The way this phase works can be visualized on the following equation 4.2, where Sc are the scores obtained from the forward computation and R’s are matrices storing all the relevance values at each layer. Sc→Rk→αβRule(ωj,k, aj)→Rj⇒ ⇒Rj→αβRule(ωi,j, ai)→Ri→... →Rinput (4.2) The function αβRule described on 3.5 need to be expressed in a vectorized form to be applied on our VGG16 model. The following formulas 4.3 represent the operations needed to propagate the relevance values from Rkto Rj[26]. ω+ j,k =max(ωj,k), ω− j,k =min(ωj,k) Z+ j=ForwardP ass(ω+ j,k, aj), Z− j=ForwardP ass(ω− j,k, aj) S+ j=Rk Z+ j , S− j=Rk Z− j C+ j=BackwardP ass(S+ j, ω+ j,k), C− j=BackwardP ass(S− j, ω− j,k) Rj=aj·[α·C+ j−β·C− j] (4.3) Note that satisfying the detailed conditions on 3.4.1 for αand β, when α= 1 the negative influences, also known as negative explanations, disappear. Different results for this casuistic have been obtained. It is observable that VGG16 especially focuses on the contours of the targets and in the case of animals their faces. In 4.11 the relevant pixels seem to be the ones that form the lion’s face and the tiger stripes. Even though more relevant pixels are found around the face of the lion, the most relevant pixel is from the tiger. This match the probability score that VGG16 detects (54.46% for the tiger and 35.36% for the lion). 21 Chapter 4 Experiments and Results Figure 4.11: LRP on Lion & Tiger for α= 1 For higher values of αthe intention is to visualize positive and negative influences, but problems have been encountered. The main problem is that for this case relevance values tend to infinity seemingly caused by Conv layers. Clipping was tested, without realizing that then, the values would be just approximations of the real relevance values. Additionally, multiple pixels are clipped no matter the maximum value chosen, generating an inconclusive result 4.12. Figure 4.12: Failed LRP for α= 2, β = 1 Even though the results are not the expected, here the difference in relevance values between the tiger versus the lion is more apparent, as higher relevance (yellow pixels) on the front legs and face of the tiger are attributed compared to the lion. Due to the problems had with values of α > 1 only results for α= 1 have been computed and can be found on B.9 with other insights about the results. 22 Chapter 4 Experiments and Results 4.5 Localization Methods 4.5.1 GradCAM and Guided GradCAM When studying Guided Backpropagation the difficulty to differentiate classes have been discussed. GradCAM uses a system to create heat maps where heat zones are mapped highlighting where the class information is. Yet, the suggested procedure needs the training of the commented GAP layers 3.5.1, which substitute Dense layers. An alternative scheme can be applied [27] to achieve these heatmaps without the necessity of modifying the actual VGG16 architecture. Figure 4.13: Alternative GradCAM System The way this scheme works is to first use the CNN to obtain the class scores, while saving the feature maps from the last Conv layer in the process (in the case of VGG16 block5-conv3). Then gradients are applied and averaged following the equation 4.4, where ycis a selected class, obtaining the weights ωc kthat are markers pointing out how important each feature map Ais for the final decision. ωc k=1 MX iX j ∂yc ∂Ak i,j (4.4) This is very useful because spatial information is preserved as we work with feature maps from Conv layers. Finally, a weighted sum between the ωc kand Ais computed and a ReLU function is applied to ensure only positive values, getting one dimension image, that when normalized results in a heat map. As mentioned on 3.5.1 to solve the lack of pixel explanation the heat maps generated can be used as masks for the results of Guided BackPropagation obtaining then pixel based explanations. 23 Chapter 4 Experiments and Results Figure 4.14: Explanation of Tiger Another thing to note is that negative explanations can be achieved by changing the sign on equation 4.4. That way we can see for one class which zones explain the selected class and also which ones don’t. For example negating the gradients and using the same image as Figure 4.14 we obtain the Figure 4.15. Figure 4.15: Negative Explanation of Tiger Taking into account that lion and tiger classes are pretty close to each other (neurons 291 and 292 respectively on VGG16), seeing this results we can ensure that the model has no problem distinguishing between the two classes. Consequently the intuition tell us that Distance methods would tell that clusters of these two classes are well separated from each other. The particular case at Figure 4.4 caught my attention to verify if VGG16 is focussing on letters when warplanes are classified. To test this casuistic this method can give us insights about what is really happening. Figure 4.16: Explanation of Warplane 24 Chapter 4 Experiments and Results Surprisingly, the model has learned to detect the class Warplane from the sky, not using the letters as it first seem 4.4. Strangely enough when the negative explanation is computed for the same input the result is the following Figure 4.17. Figure 4.17: Negative Explanation of Warplane What is definitely a proof that VGG16 correctly avoids the letters to capture patterns. But at the same time it also dodge pixels from the plane. More examples and discussions are made on B.10. 4.5.2 Occlusion Maps The implementation of this method is easily the simplest, other than Activation Visualization. A mask of a certain size is slided horizontally across the input image to then track the score of a certain class. The parameters of Occlusion Maps are the size of the mask and the stride which tells us how many pixels need to be ignored during the sliding. Notice that if the size of the mask is too little, then the accuracy drops will not be perceived and at the same time very detailed maps will be obtained. In contrast, if the size is too large, huge drops of the class score will be tracked but pixel precision will be lost. Apart from this trade-off, the stride is added to reduce the computation cost of the method, and often low values are used in order to affect as less as possible the resultant heat map. A size of 15 for the mask 3 and a stride of 3 pixels have been used as parameters to obtain the results. Examples with other images can be found on B.11. Figure 4.18: Occlusion Map for the Class Cheetah 25 Chapter 4 Experiments and Results 4.6 Distance Methods 4.6.1 Distance Robustness These type of methods require huge amounts of data from each class to create good cluster estimations and determine how they are distributed across multidimensional space. This amount of samples were not available for this project, so instead we have decided to discuss the separation between classes where VGG16 seem to have problems. To do so, cosine distance 4.5 have been used because it is a normalized distance metric. Similarity(A, B) = A·B kAk·kBk(4.5) The classes that have been studied are between Lion and Tiger and also Warplanes versus Aircraft. Multiple samples from each classes where recollected and their respective scores. Then the average of those class scores have been computed to finally determine the distance between the means of both classes. For the case of Lion and Tiger the average distance is 0.1273 far away from 1 that indicates equality. For the other case, Warplanes distance themselves from Aircraft by 0.0182 on average. Besides that Aircrafts often have warplanes on them, with this result we ensure that VGG16 is capable to distinguish between these two classes. 4.7 Discussion Across this chapter different strategies have been studied and discussed to increase the level of interpretability in this case of the VGG16 model. Being this model a CNN, it has been proven that probably Attribution,AM and Localization methods have provided more insights and interpretability after all. This is caused, in part, because visualizing what the model has learned on each layer plus knowing where the attention is being emphasized, give enough evidence to ensure that VGG16 has no training errors and even in some cases surprise us raising hypothesis of how it is really learning. Considering the previous point, we can verify that visualization is very important when it comes to CNN models. Although in the case of VGG16, Function and some Signal methods like Deconvolution have not provided appealing results it may well not be the case for other CNN with the same objectives, since each model at the end learn differently even with the same training examples. Overall, with this study it can be ensured that VGG16 is well trained and no apparent biases or errors have been found. Apart from that, the implementation of the interpretability methods, have been successful. 26 5. Budget This project has been developed thanks to the resources of the Image Processing Group at UPC. Therefore, the computational cost of the GPU to test the interpretability methods did not have any cost. Besides the expenses due to the computational power required, the main costs of this project come from the salary of the researchers and the time spent in it. The team who carried out the project is formed by me, the author of the thesis, and my advisor. I have considered myself as a junior engineer with a salary of 10e/hour and my advisor as a senior engineer with a salary of 20e/hour. The total length of the project was 23 weeks, starting in mid-January and finishing at the end of June, as illustrated in the Gantt diagram shown on Figure C.1 and Figure C.2. Therefore, the budget can be calculated as shown on the following Table 5.1 Wage/Hour Dedication Weeks Total Junior Engineer 10.00e/h 30h/week 23 6900.00e Senior Engineer 20.00e/h 3h/week 20 1200.00e TOTAL 8100.00e Table 5.1: Budget of the Project 27 6. Conclusions The main goal of this project was to exhibit and demonstrate how multiple interpretability methods work while improving the comprehension of how a trained model behaves. A bunch of interpretability techniques have been applied to a CNN (VGG16), some of them capable of quantitatively increase the intelligibility of the model, which shows that it is possible to remove the Black Box label from this type of models. On the case of study we have also seen that for CNN, visualization and pattern localization are very important when it comes to understanding how CNN build concepts and make decisions. This may not be the case for other model architectures, but that is why a variety of other techniques have been discussed. Moreover a useful classification of these methods have been proposed, inspired on a previous partial classification [5], with the intend of avoiding the possible confusion that one may encounter when deciphering DNN. Even though the best interpretation will always be through a specialized technique for an individual model, these classification can help others have a prior and solid idea of where to start investigating their models. During the analysis of the VGG16 we can also conclude that in order to achieve a global interpretation of any model, a combination of methods from different categories and approaches would help contrast the results obtained. If more than one method points out a particular event, there will be enough evidence to confirm that the observed event happens. Just with one interpretability technique is very difficult to proof possible anomalies on the model. Last but not least, we have managed to create an open source, functional and scalable python ToolBox called iNNterpret [1] that allow users implement all tested methods on this project. To facilitate the use of this ToolBox, methods are found into submodules and the methods are implemented on different classes with the same structure. 28 7. Future Work To end this project, here it is presented a list of possible projects and tasks that have not been tackled or discussed and would be interesting to invest time in them, with the aim of deepen in the impact and possibilities that interpretability can reach in AI field. Try to implement the methods that were not implemented on this project, such as Surrogate methods, Distance methods and the detailed Attribution methods on 3.4.2. A comparison between models and the interpretability methods can be made, to see if techniques applied on one model are consistent in another, or further tweaks need to be applied. The intuition tell us that models would not share the same insights as they are trained differently even with the same objective, and different learned patterns would be observed. Find out, how these techniques can be implemented into new NN architectures, for example GAN’s or other more complex architectures, and understand their behaviours. Keep maintaining the ToolBox by solving possible bugs and adding new methods. 29 Bibliography [1] iNNterpret. Software. https://github.com/paudom/iNNterpret. [2] Keras Framework. Software. https://keras.io. [3] TensorFlow Framework. Software. https : / / www . tensorflow . org / api _ docs/python/tf. [4] Maximilian Alber et al. “iNNvestigate neural networks!” In: (2018). http: //arxiv.org/abs/1808.04260. [5] Pieter-Jan Kindermans et al. “Learning how to explain neural networks: PatternNet and PatternAttribution”. In: (2017). http://arxiv.org/abs/1705. 05598, pp. 1–12. [6] Marco Tulioa Ribeiro, Sameer Singh, and Carlos Guestrin. “”Why Should I Trust You?”: Explaining the Predictions of Any Classifier”. In: (2016). http: //arxiv.org/abs/1602.04938. [7] Daniel Smilkov et al. “SmoothGrad: removing noise by adding noise”. In: (2017). http://arxiv.org/abs/1706.03825, pp. 1–12. [8] Mukund Sundararajan, Ankur Taly, and Qiqi Yan. “Axiomatic Attribution for Deep Networks”. In: (2017). http://arxiv.org/abs/1703.01365. [9] Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. “Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps”. In: (2013). http://arxiv.org/abs/1312.6034, pp. 1–8. [10] Diederik P. Kingma and Jimmy Ba. “Adam: A Method for Stochastic Optimization”. In: (2014). http://arxiv.org/abs/1412.6980, pp. 1–15. [11] Aleksandar Botev, Guy Lever, and David Barber. “Nesterov’s accelerated gradient and momentum as approximations to regularised update descent”. In: Proceedings of the International Joint Conference on Neural Networks (2017). http://arxiv.org/abs/1607.01981, pp. 1899–1903. [12] Anh Nguyen et al. “Synthesizing the preferred inputs for neurons in neural networks via deep generator networks”. In: (2016). http://arxiv.org/abs/ 1605.09304, pp. 1–29. [13] Matthew D Zeiler and Rob Fergus. “Visualizing and Understanding Convolutional Networks”. In: European Conference on Computer Vision (ECCV) 8689 (2013). http://arxiv.org/abs/1311.2901, pp. 818–833. [14] Jost Tobias Springenberg et al. “Striving for Simplicity: The All Convolutional Net”. In: (2014). http://arxiv.org/abs/1412.6806, pp. 1–14. [15] Gr´egoire Montavon, Wojciech Samek, and Klaus-Robert M¨uller. “Methods for Interpreting and Understanding Deep Neural Networks”. In: Digital Signal Processing: A Review Journal (2017). http://arxiv.org/abs/1706.07979, pp. 1–15. 30 Chapter B Additional Details and Results B.6 Activation Maximization Results (a) Gold Fish Class (b) Fly Class (c) Hammerhead Class (d) Basketball Class Figure B.12: AM for Classes On some maximized inputs, characteristics are arduous to find. For example (B.12d) strange patterns are formed, but none is familiar with the elements that appear on a basketball game. For marine animals (B.12a & B.12c) it can also be hard to see fins and elements that remind us of this kind of animals. However, for insects, the network is capable of capture recognizable characteristics as the wings and eyes of a fly (B.12b). 37 Chapter B Additional Details and Results B.7 Deconvolution Results Depending on the feature map selected, different pixels zones or characteristics are back-traced. (a) Original Image (b) Block3-Conv3-127 (c) Original Image (d) Block3-Conv3-127 Figure B.13: Deconvolution Examples B.8 Guided BackPropagation Results All the results presented are from Block 5 and Conv layer 3. Figure B.14: Guided on Hammerhead 38 Chapter B Additional Details and Results Figure B.15: Guided on Warplane Figure B.16: Guided on Lion & Tiger B.9 Layer-Wise Relevance Propagation Results Figure B.17: LRP on Hammerhead High relevance is attributed to the contours of the Hammerhead as well as his eye. Notice that some importance have been assigned to the fish around. Interestingly 39 Chapter B Additional Details and Results the four corners of the image are even more important for the VGG16 than the background water, as if the network wanted to see what is beyond those corners. Figure B.18: LRP on Wolf Here the previous points are again satisfied. Faces are hugely relevant when it comes to predict animals. Figure B.19: LRP on Basketball In the case of not finite concepts like basketball, strangely enough, VGG16 stops focusing on faces and starts paying more attention to contours of objects like the ball or the clothing. Some pavement lines are detected as significant too. 40 Chapter B Additional Details and Results B.10 GradCAM and GuidedGradCAM Results Figure B.20: Explanation of Basketball Instead of the ball, more attention is used on the jerseys. Computing the explanations for the class Baseball it seems that VGG16 focuses more on the legs and hands. Figure B.21: Explanation of Baseball Other learning patterns and distinctions can be can be found. For example the difference between a Shark and Hammerhead classes, which for humans is obvious. Figure B.22: Explanation of Shark 41 Chapter B Additional Details and Results Figure B.23: Explanation of Hammerhead Analyzing the results, VGG16 appears to detect sharks by the top fin, and hammerheads by the lower fins. Furthermore for hammerhead some attention is used for the other fish around the predator, raising the hypothesis that maybe hammerheads are more commonly surrounded by fish than sharks are. B.11 Occlusion Maps Results Figure B.24: Occlusion Map for the Class Warplane Figure B.25: Occlusion Map for the Class Tiger 42 C. Work Plan C.1 Tasks and Milestones Background learning for DL Tasks Take neural networks and DL course Take DNN course Take CNN course Milestone Learning Table C.1: Work Package 1 Learn the basics about Interpretability Tasks Search related papers Prepare summaries. Milestone Learning Table C.2: Work Package 2 Plan the tasks and objectives of the project Tasks Plan the project Specify the objectives, requirements and specifications Prepare the Project Plan document Milestone Project Plan document Table C.3: Work Package 3 43 Chapter C Work Plan Research and read about Interpretability methods Tasks Research what types of methods exist Select a bunch of methods from each category. Get an idea of how the methods work and their objective Milestone Documentation & Research Table C.4: Work Package 4 Implement the methods selected and extract results Tasks Choose a pretrained model to test the methods on Implement, Test and modify available code to obtain results Try different parameters and find the best ones. Milestone Software Table C.5: Work Package 5 Critical Review Tasks Develop the Critical Review document Evaluate the state of the project Milestone Critical Review Document Table C.6: Work Package 6 Design and Build the Toolbox Tasks Define the structure of the Toolbox Adapt the current code to the Toolbox Milestone Software Table C.7: Work Package 7 44 Chapter C Work Plan Test the Toolbox Tasks While building the Toolbox, testing it Milestone Software Table C.8: Work Package 8 Extract Conclusions Tasks Collect the results Define how these could be helpful for the medical industry Milestone Conclusions Table C.9: Work Package 9 Final Deliver Tasks Write the Final Memory document Prepare the Final Presentation Milestone Final Memory Document & Presentation Table C.10: Work Package 10 45 Chapter C Work Plan C.2 Gantt Diagrams Here Gantt diagrams are shown to describe visually the time plan with all the work packages and their time limits. Figure C.1: Gantt Diagram First Part Figure C.2: Gantt Diagram Second Part 46