Full text
Automatic classification of tissue malignancy for breast carcinoma diagnosis Irene Fond´ona, Auxiliadora Sarmientoa,∗,AnaIsabelGarc´ıaa,Mar´ıa Silvestrea, Catarina Eloy b,c,Ant´onio Pol´oniab, Paulo Aguiar d,e aSignal Processing and Communication Department, Engineering School, University of Seville, Seville, Spain bPathology Department, Institute of Molecular Pathology and Immunology (Ipatimup), University of Porto, Porto, Portugal cMedical Faculty, University of Porto, Porto, Portugal dInstitute of Biomedical Engineering (INEB), University of Porto, Porto, Portugal eInstitute for Research and Innovation in Health Sciences (i3S), Porto, Portugal Abstract Breast cancer is the second leading cause of cancer death among women. Its early diagnosis is extremely important to prevent avoidable deaths. However, malignancy assessment of tissue biopsies is complex and dependent on observer subjectivity. Moreover, hematoxylin and eosin (H&E)-stained histological images exhibit a highly variable appearance, even within the same malignancy level. In this paper, we propose a computer-aided diagnosis (CAD) tool for automated malignancy assessment of breast tissue samples based on the processing of histological images. We provide four malignancy levels as the output of the system: normal, benign, in situ and invasive. The method is based on the calculation of three sets of features related to nuclei, colour regions and textures considering local characteristics and global image properties. By taking advantage of well-established image processing techniques, we build a feature vector for each image that serves as an input to an SVM (Support Vector Machine) classifier with a quadratic kernel. The method has been rigorously evaluated, first with a 5-fold cross-validation within an initial set of 120 images, second with an external set of 30 different images and third with images with artefacts included. Accuracy levels range !Fully documented templates are available in the elsarticle package on CTAN. ∗Corresponding author. Tel.: +34 95 448 2176; Fax: +34 95 448 7291 Email address: [email protected] (Auxiliadora Sarmiento) Preprint submitted to Journal of L A T EXTemplates March5,2018 *Revised manuscript (clean) Click here to download Revised manuscript (clean): Computer_in_biology_and_medicine_mar_2018_EDITADO_AJE_revision4.pdf This is an Accepted Manuscript of an article published by Elsevier in Computers in Biology and Medicine Volume 96, 1 May 2018, available at: https://doi.org/10.1016/j.compbiomed.2018.03.003 Copyright 2018 Elsevier. En idUS Licencia Creative Commons CC BY-NC-ND
from 75.8% when the 5-fold cross-validation was performed to 75% with the external set of new images and 61.11% when the extremely difficult images were added to the classification experiment. The experimental results indicate that the proposed method is capable of distinguishing between four malignancy levels with high accuracy. Our results are close to those obtained with recent deep learning-based methods. Moreover, it performs better than other state-of-the-art methods based on feature extraction, and it can help improve the CAD of breast cancer. Keywords: Breast cancer, Computer-aided diagnosis, Digital pathology, Histopathological images, Pattern recognition and classification, Tissue malignancy 1. Introduction Breast cancer is the leading cancer killer among women aged 20–59 years globally [1]. Its survival rate exceeds 80% if detected early but unfortunately decreases to 10–40% when diagnosed in its late stages [2]. Mammography has undergone greater scrutiny than almost any other medical intervention, mainly because when mammography screening is introduced into a population, deaths from breast cancer decline [3]. If a clinically suspicious mass is found in the mammography, a biopsy should also be performed. The obtained tissue samples are stained with hematoxylin and eosin (H&E) and observed by a pathologist with a microscope. This procedure, called histopathological analysis, is considered the foundation of modern oncology, playing a major role in the treatment of other types of diseases [4]. Most breast lesions detected by pathologists are benign: that is, they are not cancerous. Common benign lesions are developmental abnormalities, epithelial and stromal proliferations, inflammatory lesions, and neoplasms. The majority of these benign lesions are not related with an increased risk for subsequent breast cancer. Breast cancer starts when cells in the breast begin to divide and grow out of control in an abnormal way. According to the type of cell they start in, breast cancers can be grouped into two categories: carcinomas and sarcomas [5]. Carcinomas are cancers that begin in the lining layer (epithelial cells) ofthebreast,whereassarcomasbegin in the connective tissue (stromal).Sarcomasaremuchlesscommonthan carcinomas, accounting for less than 1 % of primary breast cancers. Carcinomas can be divided into two broad categories: noninvasive or in 2
situ and invasive. In situ carcinoma is characterized by cancer cells that are confined to the pre-existing normal lobules or ducts. Is there is any evidence that cancer cells invade the surrounding stroma, however limited, the lesion is classified as invasive carcinoma. In situ carcinoma has the potential to progress to invasive cancer if not treated adequately. By visually inspecting histological images of tissue samples, pathologists identify some distinctive patterns thatcharacterizetissue malignancy. In particular, they consider two main aspects: general architecture and cytological details. These experts categorize the images regarding the type of breast cancer and the malignancy within each type. However, subjectivity and other human factors could lead into errors that could critically affect patient care [6]. Moreover, the process is time consuming and expensive, making the development of a computer-aided diagnosis tool (CAD) critical. The significant importance of these tools is intimately associated with the recent developments in the promising field of digital pathology [7, 8]. CAD tools have been widely used in other diseases such as diabetic retinopathy or melanoma. Analogous to the role of CAD algorithms in these applications, the desired tool should aim to complement the specialized opinion of the pathologist by using an objective judgement, making use of quantitative measures [2]. For that purpose, the tool should give information about the malignancy level of the sample under study. Focusing on carcinoma, it would be desirable for the system to divide the images into four categories: normal breast tissue, benign lesion, in situ carcinoma, and invasive carcinoma, as is performed in clinical practice [5]. An example of each image type is depicted in Fig. 1. However, creating a computer vision application for automatic classification of tissue malignancy is not a trivial challenge [6, 9]. Breast cancer is a heterogeneous disease comprising numerous distinct entities [9], and therefore, tissue samples exhibit a highly variable appearance. In this paper, we present a CAD tool for automatic classification of tissue malignancy based on the processing of histological images. The primary contribution of this technique is the possibility of distinguishing among four malignancy levels based on the use of region and nuclei colour and texture features. 3
(a) (b) (c) (d) Figure 1: Representative histological H&E stained images of normal (a), benign (b), in situ (c), and invasive (d) cases. 2. Related Work Automated histopathological image analysis has recently become a significant research problem in medical imaging. There is an increasing need to develop quantitative image analysis methods as a complement to the efforts of pathologists [8, 10]. Regarding research on cancer diagnosis, over the past two decades, a tremendous amount of work has been conducted [11]. 2.1. Generally adopted workflows Currently, there are two families of methods that address the problem of tissue malignancy classification. Thefirstfamilyiscomposedofmethods that use machine learning algorithms based on feature descriptors. These methods follow a similar workflow that comprises three steps: preprocessing, feature extraction, and diagnosis [12, 13, 14, 15, 16, 17, 18, 19, 20, 21]. 4
The preprocessing step has three main objectives. The first one is detecting an area of the image with a significant amount of information called the region of interest (ROI); here, thresholding techniques are widely employed [11]. The second one is noise reduction with some kind of filter. This noise is usually added involuntarily during the staining process and could hinder the segmentation and feature extraction stages. The third objective is colour normalization, which tries to somehow eliminate, or at least reduce, colour variations or uneven illumination [12]. For feature extraction, current methods usually segment certain areas of the image before computing some distinctive characteristics [13]. Images can be processed by looking for tissue areas or complete structures. Another approach is searching for nuclei or individual cells. In both cases, once the elements are segmented, some numerical values are obtained. The numbers of cells, shapes, and areas are just some of the measures usually extracted [13, 14]. State-of-the-art algorithms are designed based on well-established image processing techniques relying on concepts such as texture, intensity level or morphology [11]. Finally, for the diagnosis step, the majority of the proposed techniques use learning algorithms, mainly those based on classifiers such as neural networks, support vector machines (SVMs), and decision trees [10, 11]. The other family of methods, which has recently been developed, is based on deep learning techniques, in which the relevant features are learned directly by the network. In particular, convolutional neural networks (CNNs) have been used to handle two-class breast cancer classification in [22, 23, 24] and the classification into four classes in [25]. With large datasets, deep learning techniques have been applied with great success to the detection, segmentation and recognition of objects and regions in images [26, 27]. However, histopathological datasets are often small, and in this scenario, CNNbased methods might result in overfitting and poor generalization. When overfitting happens, the model learnedfromthetrainingdatasetdoesnot generalize well, and therefore, the performance of the classification decreases when new data is presented. To overcome this serious problem, CNN-based methods for cancer classification usually expand their dataset by taking image patches and by artificially generating samples via affine transformations. The augmented dataset is then used to train the network. Finally, they perform a patch-wise classification and decide the label of the whole image by simple decision fusion methods, such as max-pooling and voting. 5
2.2. System evaluation These previously proposed systems are evaluated according to their performance. However, this is a difficult issue due to the general lack of annotated images and ground truth (GT) [11]. The most realistic and accurate evaluation method is splitting the image database into two groups: one for training and one for testing. As the number of images is not enough, the authors try to find a compromise solution. One of the adopted options is to evaluate the system according to the success obtained in the training stage. However, the use of the same images for training and testing could lead to over-optimistic algorithms. Another option is k-fold cross-validation. This approach randomly partitions the dataset into k groups and uses the k-1 groups left to train the system. This procedure is repeated k times such that each group is used once to test the system. The problem is that the test set must be independent from the training one because it is an example of an unknown future case. k-fold cross-validation may result in dependent samples, and therefore, the result would not be as realistic as needed [14]. Regarding the state-of-the-art methods proposed for breast cancer study, a large number of them are intended to segment specific structures within the images [6, 13, 15, 16, 17, 18, 19]. These methods do not try to provide a level of malignancy but instead aim to serve as an initial step for future CAD tools. Only a few approaches give a malignancy assessment of the tissue samples under study. The majority of them consider only two possible output classes: malignant or benign. In [20], the authors extracted nuclei with a watershed algorithm. Afterwards, they computed a number of features based on these segmented nuclei, such as their density, proximity, and number. They also added other image characteristics computed in windows of various fixed sizes all over the grey level image. Some of these new characteristics were mean intensity level, contrast, correlation, and homogeneity. Gabor decomposition coefficients were also added to the feature vector where the decomposition was performed in various fixed scales separately on each of the colour planes of the original image. The results obtained with the four colour spaces, i.e., RGB, HSV, CIE L*a*b*, and stain decomposition mode or H&E colour space, were presented. The dimensions of the obtained feature vector were too high (3000 or 2000 in terms of the number of colour planes in the colour space under test), making the subsequent classification procedure too slow. Therefore, its dimensionality was reduced to two with the spectral clustering technique. The obtained results showed that the H&E colour space has a higher feature 6
discriminative power than the rest of the colour spaces under test. Regarding the accuracy of the classification procedure, as generally adopted for image processing methods for histological images, a kind of cross-validation with k=10wasusedonasetof58images. Theaccuracyobtainedwas80.2%, which is based on the theory previously explained. The approach presented in [4], as in [20], only classifies samples into two possibilities: malignant and benign. Themethodreliesontheaccuratesegmentation of nuclei, performed with random walker technique in the H plane of the H&E colour space. From these segmented regions, the algorithm computes the mean of the nuclei sizes for each image. To represent the anatomical structure of the image, the Hessian matrix on each pixel was computed, and the deviation from a blob structure, RB, was obtained. As nuclei in benign images are rounded structures, RB must be low, and therefore, the mean RB value was then added to the feature vector. The method also extracts the Fourier coefficients because, as the authors claim in the article, irregular nuclei are characterized by a precise distribution of Fourier frequency coefficients . Searching for this specific distribution, the number of irregular nuclei is obtained per image. As a texture descriptor, the histogram of textons computed in the H colour plane with 11 bins was obtained. Three network related parameters were used as well: the mean cycle-weighted Euclidean length, which takes advantage of the cycle structure present within the cell networks; the average shortest path between nuclei; and the number of connected components in the graph. The resulting feature vector had a length of 27. An SVM classifier was used in the classification stage. The data base under test consisted of an unknown number of images that should not be high, as the authors took patches of equal sizes from the complete images to evaluate the performance of the system, considering them as if they were whole images for their experiment. Although it may have been unavoidable due to the lack of images, this selection could lead to over-optimistic results, as the correlation among patches of a certain image is supposed to be high. The dataset could not represent an unknown real experiment. For computing the accuracy, 10-splits Monte Carlo cross-validation was adopted. For each split, the training and validation sets were composed of 30 and 70 images, respectively. The result was computed with the arithmetic mean of the predictive accuracy scores obtained in each split. Unlike k-fold cross-validation, Monte Carlo method allows to select the percentage of data for training and testing regardless of the number of iterations. But, the selection of the images for each learning-testing split is random, and therefore, one image may 7
be selected in all learning sets, but in none of test set. Also, the results will vary if the analysis is repeated with different random splits, which is known as Monte Carlo variation. The accuracy levels claimed were between 84.86% and 87.14%, agreeing with cross-validation methods in general. The method presented in [28], aims to classify images of carcinoma into three levels of severity. This method supposes that the input image is a carcinoma and therefore does not distinguish between normal and benign cases. The technique processes the grey-level image and extracts 30 classical textural features based on first-order image statistics, the second-order covariance matrix and the third-order run-length matrix. Forthe classification step, three classifiers were compared: k-nearest neighbour (kNN), probabilistic neural network (PNN) and SVM.Theimagedatabaseconsistedof 65 patches, which were obtained from 13 original images. For quality measurement, k-fold cross-validation was used again, obtaining a mean accuracy of 89.5%. However, the database under test is highly redundant, and this numerical value may not represent the overall performance of the method in arealisticcase. 2.3. Contribution The proposed methods main contributions are as follows: 2.3.1. Severity assessment dividing into four classes To the best of our knowledge, there are no feature-based methods for normal, benign, in situ, and invasive categorization of original tissue samples. The methods proposed usually give only two levels: malignant and benign, and only one work [28] classifies images of carcinomas into three levels of severity. Only one of the methods based on deep learning techniques presents classification results in four classes [25]. 2.3.2. Rigorous validation Image database. Even state-of-the-art methods are tested on image datasets with a low number of images. To overcome this problem, researchers use samples of the same image instead of complete original ones. This fact could lead to over-optimistic results due to the high correlation of one sample with another. Moreover, appearance variability, a key characteristic of H&E images, is limited in this kind of testing. To address this, the database used is composed of three sets of images. The first one with 120 images (30 for each class / malignancy level under study) is intended to be as representative as 8
possible of the wide variety of appearances that one could encounter in H&E images. From milk ducts to fat, these images content is complex and variable. Another set of 20 images of the same complexity was reserved to externally validate the method. Finally a new set of 16 images was also tested. This set of images is the most difficult one because it comprises images with artefacts such as text and poor staining. External validation. Prediction models tend to perform better on data on which the model was constructed than on new data. This difference in performance is an indication of the optimism in the apparent performance on the derivation set. Results are often accepted without sufficient consideration of the importance of external validation [29]. To validate our CAD tool, three tests were performed: 5-fold cross-validation within the first set of 120 images, external validation using the second set and the already trained model and external validation using thethirdsetofdiscardedimages. 3. Method The proposed tool has been developed with MATLAB R !R2015a software (The MathWorks Inc., Natick, MA). MATLAB implementations of the proposed algorithm are freely available from [30]. A block diagram of the system is shown in Fig. 2. Conceptually, the approach is simple: after an initial colour normalization step, we build a feature vector for each image and perform a classification procedure based on an SVM classifier. Two different types of characteristics compose the feature vector: colour based and texture based. 3.1. Dataset The image database consists of a total of 156 high-resolution anonymous annotated H&E staining images from breast tumour biopsies, which were made available through the Bioimaging 2015 Grand Challenge [31]. To the best of our knowledge, it is the only public database that provides examples of the four degrees of cancer. The images were all digitized with the same acquisition conditions and with a magnification of 200×.Allofthemhavea fixed size of 2048 ×1536 and were stored in tagged image file format (TIFF). Unlike the majority of the state-of-the-art methods, [4, 10, 14, 25, 28], we do not subsample the images or extract patches from them. On the contrary, we have used the whole images, thereby preserving all the information that 9
serves as a threshold for the others belonging to the circular neighbourhood. Higher pixel values relative to this central one correspond to a value of 1 in the corresponding LBP code. Otherwise, a 0 value would be assigned. Computing the corresponding LBP code for each pixel is similar to identifying different grey-level local variations and therefore identifying local textures. The extracted characteristics are the occurrences of each of the LBP patterns on each image, obtaining 36 features per colour plane under study. To take into account the interaction between colour and texture, we perform the same analysis in the R, G and B colour planes from the RGB colour space, which constitutes a total number of 144 features that are concatenated in order to build the LBP-based feature vector fLBP . 3.4.3. Sparse texture vector This vector is built by radially sorting the values of the neighbours of the pixel under study in the C colour plane. The neighbourhood is a squared window with a fixed size of 5×5 pixels. As in [40], we further process this textural representation to increase the variance between the elements of the texture descriptor and to improve the efficiency. Therefore, we apply the principal component analysis (PCA) technique to retain only 5 of the strongest components. Once the texture of each pixel neighbourhood is calculated for each pixel, we employ a similar technique to that presented in [43]. The obtained sparse texture set is assumed to be distributed as a 5-dimensional Gaussian distribution, the parameters of which are estimated with the expectationmaximization algorithm. Finally, we retain the mean and the diagonal of the covariance matrix of the distribution as the texture features to be added into the local texture feature vector, fst, with a final length of 10. 3.5. Classification With the colour and texture features extracted, we build the final feature vector as in Eq.(1). Its size is 260, and although PCA analysis was performed to reduce this length, the resultant decrease in the quality of the subsequent classification procedure indicated to us that all the components should be preserved. ffinal ={fnc,f rc,f ct,f LBP ,f st}(1) This feature vector serves as the input for the classifier. As presented in results section, we have tested a total number of 9 classifiers, obtaining 16
the best model using the SVM classifier [44] with a quadratic kernel. SVMs classify data by choosing the optimal hyperplanes, which separate data within their respective classes, optimizing classification accuracy for the training set. The best hyperplanes have the largest separation in relation to the given classes [45]. 4. Results In this section, we present the results obtained when the method was evaluated with three different histopathological image datasets. 4.1. Classification results 4.1.1. Experiment with training dataset; dataset number 1 For the first image dataset, which is composed 120 images, we have trained a total number of 9 classifiers: •Simple tree. Decision tree classifiers are conceptually simple yet powerful and widely used classifiers [46]. Their structure is similar to a tree with its root at the top; internal nodes representing a condition based on which the tree splits into edges; and leafs that are nodes that do not get split. The root node represents the entire population, whereas each of the leafs or terminal nodes contains a different class. The maximum number of splits utilized in the tree is 4. •Complex tree. In order to make fine distinctions between classes, complex trees increase the maximum number of split up to 100 [46]. •Bagged tree. The idea of bagging is to obtain the best classifier by combining the results of multiple weak classifiers into a single and strong one [47]. In a bagged tree, the basic classifier is a decision tree. The workflow of the algorithm is simple. For a fixed number of times, T, the algorithm selects at random n training samples with replacement, obtaining T training sets of size n. Some of the samples could be repeated from one set to another due to the replacement option. Then, T decision trees are trained with these sets, with each one trying to fit the model. Their decisions are finally combined with the majority voting rule. Bagging leads to improvements in unstable procedures [48]. 17
•AdaBoost tree. The concept behind an AdaBoost tree is similar to a Bagged tree. However, AdaBoost trees use the information related to the classification difficulty of each training sample to add a weight to it. This weighted sample is input into the next tree for training, and the process is repeated. Due to this procedure, the final trees are focused on classifying hard samples [49]. The other difference from a bagged tree is the final decision. Instead of using a majority vote, the weighted vote technique is adopted. •kNN Cosine. kNN is a classifier that does not use a model to fit the training data and subsequently classify the new samples. Conversely, the training phase for kNN consists of simply storing all known instances and their class labels [50]. In the test phase, the algorithm considers k neighbours of the test sample, assigning it to the class of the majority of them. To obtain these k neighbours, the distances from all of the training samples to the test one are computed. Only the k closest samples in the feature space in terms of the selected distance are kept. The distance metric used is a key factor in the accuracy of the classification procedure. We refer to kNN cosine as a kNN classifier where the distance metric is the cosine distance metric. Medium distinctions between classes are made, and the number of neighbours, k, is set to 10. •kNN Cubic. This is a kNN classifier where the distance metric used is cubic [50]. Again, the distinctions between classes are medium, and the number of neighbours is 10. •SVM linear kernel. SVM classifiers consider a sample as a point in the feature space. However, instead of considering the neighbours of the test sample as in kNN, SVM tries to divide the feature space into regions with each corresponding to one category. When a new sample is tested, the algorithm only checks where it is located. The limits among regions in the space are called hyperplanes. From the number of hyperplanes that might classify the data, SVM defines the best hyperplane as the one that represents the largest separation, or margin, between the two classes [51]. In the practical implementation of SVM, the dot product of feature vectors is performed. However, this dot product operation is replaced by a kernel, that is, a similarity function that is easier to compute [44]. If the kernels used are polynomial regarding 18
Table 1: Obtained TPR and FNR with the classifiers under test TPR (%) FNR (%) Class Class Classifier 12341234 Simple tree 53.3 66.7 63.3 76.7 46.7 33.3 36.7 23.3 Complex tree 46.7 66.7 60.0 76.7 53.5 33.3 40.0 23.3 Bagged tree 50.0 66.7 76.7 83.3 50.0 33.3 23.3 16.7 AdaBoost tree 6.7 70.0 76.7 90.0 93.3 30.0 23.3 10.0 KNN cosine 30.0 53.3 83.3 83.3 70.0 46.7 16.7 16.7 KNN cubic 30.0 56.7 76.7 53.3 70.0 43.3 23.3 46.7 Lineal SVM 66.7 73.3 70.0 76.7 33.3 26.7 30.0 23.3 Quadratic SVM 73.3 80.0 70.0 80.0 26.7 20.0 30.0 20.0 Cubic SVM 73.3 80.0 70.0 76.7 26.7 20.0 30.0 23.3 their order, then they can be linear, quadratic, cubic, etc. We have first tested a linear SVM. •SVM quadratic kernel. In this case, we used a quadratic polynomial as the kernel for the SVM classifier. •SVM cubic kernel. Finally, we have selected the third-order polynomial, cubic, to be the kernel of our SVM classifier. The ability of these 9 classifiers to model the problem has been expressed in terms of their confusion matrices based on 5-fold cross-validation as a test method. Each row of the confusion matrix refers of the actual class identity of test images, and each column indicates the classifier output [10]. As we are identifying four malignancy levels, we obtain 4 ×4 confusion matrices. We also compute the true positive rate (TPR) and false negative rate (FNR) for each of the classes. The TPR, or sensitivity, measured per class is the proportion of images correctly classified in that class in relation to the total number of images belonging to it. The FNR per class is the miss ratio, that is, the proportion of images not classified in that class in relation of the number images belonging to the class. Figure 5 shows the confusion matrices, whereas the TPR and FNR for each of the classifiers under test are depicted in Table 1. The mean accuracy for all the classifiers under test is presented in Table 2. The best classifier, SVM quadratic, obtains an accuracy of 75.8%, while the 19
Figure 5: Confusion matrices obtained using different classifiers. worst, cubic SVM, obtains an accuracy of 54.2%. Considering the obtained results, we have selected the quadratic SVM classifier for the rest of the experiments. 20
Table 2: Obtained accuracy with the classifiers under test Classifier Accuracy (%) Simple tree 65.0% Complex tree 62.5% Bagged tree 69.2% AdaBoost tree 60.8% KNN cosine 62.5% KNN cubic 75.0 % Lineal SVM 71.7 % Quadratic SVM 75.8 % Cubic SVM 54.2% 4.1.2. Experiment with test dataset; dataset number 2 We externally validated the algorithm with a second set of 20 images. The accuracy obtained with the quadratic SVM classifier was 75%. The results are summarized on Table 3. A total number of 15 images were correctly classified by the proposed algorithm. Misclassification occurred mostly when benign images were processed, while all the ”invasive” class images were correctly classified. As benign images show some structures that, although not malignant, are abnormal in appearance, they are sometimes similar to the in situ images where cancer grows within the milk duct. An example is shown in Fig. 6. (a) (b) Figure 6: Benign images (a) appear similar to in situ ones (b), and therefore, the algorithm tends to misclassify them. 21
Table 3: Classification labels for test image dataset 2 along with its ground truth (GT) GT Algoritm result Normal Normal Normal Normal Normal Normal Normal Normal Normal In situ Benign Benign Benign Benign Benign Benign Benign In situ Benign In situ In situ In situ In situ Benign In situ In situ In situ In situ In situ Benign Invasive Invasive Invasive Invasive Invasive Invasive Invasive Invasive Invasive Invasive 22
Table 4: The classification accuracies (%) comparing our method (quadratic SVM) with the results presented in [25] on datasets 2 and 3 (extended dataset). Accuracy (%) Accuracy (%) Classifier Dataset 2 Dataset 3 CNN +Majority 80 75 CNN + Maximum probability 80 62.5 CNN+SVM +Majority 85 68.8 CNN + SVM + Maximum probability 80 62.5 Quadratic SVM 75 61.11 4.1.3. Experiment with the extended test dataset, dataset number 3 To further validate the method, we performed experiments by adding 16 extremely difficult images to dataset number 2. In this extended dataset, malignancy level classification is more ambiguous but still received the same annotation from two independent pathologists (although with a lower degree of confidence). The accuracy achieved in this case is 61.11%. The decrease in accuracy was expected, but the score is still competitive relative to the difficulty in the classification. 4.2. Comparison with the state-of-the-art As mentioned previously, the dataset we have employed in our experiments was first used in the Bioimaging 2015 Grand Challenge [31]. Before our study, this dataset has been used only in the CNN-based approach presented in [25]. This method divides the image into twelve non-overlapping patches. Then, it computes the patch class probability with a trained CNN or a CNN+SVM, and finally, it decides the image class using a majority voting approach. An augmented dataset composed of overlapping patches and rotated and mirrored versions of these patches is built to train the network. This methodology has been applied in order to overcome the small size of the training set and to prevent overfitting. In Table 4, we present a comparison of the results of our method and those of the methods presented in [25] with datasets 2 and 3. 5. Discussion The performance of our feature-based method appears to be close to that of the CNN-based approach of [25]. However, there are various aspects related to CNN-based techniques that we would like to discuss. 23
In the literature, it is well established that when there are not enough training examples, CNNs are vulnerable to overfitting [27]. That is, diagnostic results are only limited to a few specific data instead of general data. One of the most widely used methods to overcome this serious problem is data augmentation [24, 52], but the augmented dataset used for training purposes may not be enough since the augmented samples are still highly correlated. In fact, the choice of augmentation strategy can be more relevant than the type of network architecture used [53]. Additionally, some authors have reported that data augmentation approaches do not improve the accuracy of the CNN when addressing brain tumour segmentation [54]. The method presented in [25] generated an augmented dataset by mirroring and rotating overlapping patches. Even with this data augmentation strategy, the size of the training dataset is still small compared to other public databases, such as the BreakHis dataset used in [22], which contains approximately 1000 images in each class, benign or malignant, for the designated magnification. Therefore, considering the large number of parameters of a CNN, one might think that even with this data augmentation strategy, overfitting might have influenced the results in [25] and that the obtained accuracies of the classifier could be optimistically biased. In contrast, our featured-based method does not artificially increase the size of the dataset by either dividing it into overlapping patches or applying affine transformations. Another related problem with many CNN-based methods is that in general, the prediction of the class of an image based on patch-level predictions is not robust [55]. In fact, as shown in Table 4, the performance of a CNNbased method depends strongly on the strategy used to obtain the image label based on the patch-wise classification. This can be explained because the GT labels of individual patches are not provided in the dataset, and considering that the patch preserves the label of the complete image, they may be inconsistent. To overcome this problem, it is necessary to perform patch classification while taking into account the local dependencies of labels [54], but the method of [25] does not include this step in the design of the network. Finally, we should also mention that our method follows the methodology of pathologists for the classification of the images, as they take into consideration the mixture of structures and textures that can only be appreciated in the whole image, as opposed to the patch-wise classification approach that the CNN-based method uses. 24
6. Conclusions In this paper we have addressed the problem of breast carcinoma malignancy assessment from the point of view of histopathological image processing. This is a key point in early cancer detection that remains unsolved, however, mainly due to the intrinsic complexity of breast cancer images. The development of a system able to automatically clasify histological images could be a great contribution to the medical community, saving money, time, and most importantly lives. State-of-the-art techniques usually provide only two labels for the input images: malignant or benign. This is an incomplete classification that neglects cases that, although not malignant, exhibit suspicious structures. Moreover, malignant cancer could be more or less severe depending on its evolution stage. Our main contribution is a new and efficient system that provides promising results in histopathological image classification and that completely and automatically gives four labels of malignancy: normal, benign, in situ and invasive cancer. For label assignment, our tool automatically computes several features from the images that are used by an SVM quadratic classifier. These features are colour and texture based according to local characteristics and global image properties. We have extensively validated the tool with three experiments performed on a public dataset using a 5-fold cross-validation and two external validations. We have employed the only public dataset that provides four clases of tissue malignancy for the breast cancer diagnosis. The proposed technique has achieved high levels of accuracy for datasets 1-3: 75.8%, 75% and 61.11% respectively. These numerical values are similar to the ones provided by state-of-the-art feature-based techniques, despite the fact that they only classify samples into two or at most three classes and do not externally validate their methods. Unfortunately, due to its small size and difficulty, there is only one method based on a CNN that uses the same database. Although it is not clear whether overfitting affects the results of the CNN approach, our results are also close to those obtained with the CNN-based method. In contrast to the CNN-based approach, which perform patch-by-patch classification, we compute the features in whole images. More importantly, our method simulates the decision process used by pathologists. The achieved accuracy levels suggest that we are in the right direction in terms of definitely solving the problem of breast cancer malignancy assessment. However, the method must be enhanced in order to avoid errors 25
[53] I. Goodfellow, Y. Bengio, A. Courville, Deep Learning, MIT Press, 2016. [54] M. Havaei, A. Davy, D. Warde-Farley, A. Biard, A. Courville, Y. Bengio, C. Pal, P.-M. Jodoin, H. Larochelle, Brain tumor segmentation with deep neural networks, Medical Image Analysis 35 (2017) 18 – 31. [55] L. Hou, D. Samaras, T. M. Kurc, Y. Gao, J. E. Davis, J. H. Saltz, Patch-based convolutional neural network for whole slide tissue image classification, in: Proc. IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2016, pp. 1–10. doi:10.1109/CVPR.2016.266. 32