scieee AI-readable full text Open interactive document viewer

Segmentation and classification of burn images by color and texture information

Acha Piñero, Begoña; Serrano Gotarredona, María del Carmen; Acha Catalina, José Ignacio; Roa Romero, Laura María

Abstract

In this paper, a burn color image segmentation and classification system is proposed. The aim of the system is to separate burn wounds from healthy skin, and to distinguish among the different types of burns (burn depths). Digital color photographs are used as inputs to the system. The system is based on color and texture information, since these are the characteristics observed by physicians in order to form a diagnosis. A perceptually uniform color space (L *u*v *) was used, since Euclidean distances calculated in this space correspond to perceptual color differences. After the burn is segmented, a set of color and texture features is calculated that serves as the input to a Fuzzy-ARTMAP neural network. The neural network classifies burns into three types of burn depths: superficial dermal, deep dermal, and full thickness. Clinical effectiveness of the method was demonstrated on 62 clinical burn wound images, yielding an average classification success rate of 82%

Full text

Segmentation and classification of burn images by color and texture information Begon ˜a Acha Carmen Serrano Jose ´I. Acha A ´rea de Teorı ´adelaSen ˜al y Comunicaciones Escuela Te ´cnica Superior de Ingenieros University of Seville Camino de los Descubrimientos s/n 41092 Sevilla, Spain Laura M. Roa Grupo de Ingenierı ´a Biome ´dica Escuela Te ´cnica Superior de Ingenieros University of Seville Camino de los Descubrimientos s/n 41092 Sevilla, Spain Abstract. In this paper, a burn color image segmentation and classification system is proposed. The aim of the system is to separate burn wounds from healthy skin, and to distinguish among the different types of burns (burn depths). Digital color photographs are used as inputs to the system. The system is based on color and texture information, since these are the characteristics observed by physicians in order to form a diagnosis. A perceptually uniform color space ( L * u * v *)was used, since Euclidean distances calculated in this space correspond to perceptual color differences. After the burn is segmented, a set of color and texture features is calculated that serves as the input to a Fuzzy-ARTMAP neural network. The neural network classifies burns into three types of burn depths: superficial dermal, deep dermal, and full thickness. Clinical effectiveness of the method was demonstrated on 62 clinical burn wound images, yielding an average classification success rate of 82%. © 2005 Society of Photo-Optical Instrumentation Engineers. [DOI: 10.1117/1.1921227] Keywords: color images; burn; image segmentation; burn classification. Paper 04076 received May 12, 2004; revised manuscript received Jan. 11, 2005; accepted for publication Jan. 11, 2005; published online May 11, 2005. 1 Introduction For a successful evolution of a burn injury it is essential to initiate the correct first treatment.1To choose an adequate one, it is necessary to know the depth of the burn, and a correct visual assessment of burn depth highly relies on specialized dermatological expertise. As the cost of maintaining a burn unit is very high, it would be desirable to have an automatic system to give a first assessment in all the local medical centers, where there is a lack of specialists.2,3 The World Health Organization demands that, at least, there must be one bed in a burn unit for each 500000 inhabitants. So, normally, one burn unit covers a large geographic extension. If a burn patient appears in a medical center without burn unit, a telephone communication is established between the local medical center and the closest hospital with burn unit, where the nonexpert doctor describes subjectively the color, shape, and other aspects considered important for burn characterization. The result in many cases is the application of an incorrect first treatment 共very important for a correct evolution of the wound兲, or unnecessary displacements of the patient, involving high sanitary cost and psychological trauma for the patient and family. With the fast advances in technology, computer aided diagnosis 共CAD兲systems are gaining widespread acceptance. However, nowadays, the research in the field of skin color images is developing slowly due to the difficulty of translating human color perception into objective rules, analyzable by a computer. Generally speaking, one can find two main applications about skin color image processing in the literature:4 the assessment of the healing of skin wounds or ulcers,5–9 and the diagnosis of pigmented skin lesions such as melanomas.10–15 The analysis of lesions involves more traditional image processing techniques such as edge detection and object identification, as well as an analysis of the color, irregularity, and shape of the segmented lesion. In wound analysis, the analysis of the colors within the wound site is often more important than the detection of the wound border or the calculation of its area. Particularly, in the case of burn depth determination, focusing on the shape of the burn is irrelevant for predicting its depth. The main characteristics for this purpose are color and texture information, as they are the features observed by physicians in order to give a diagnosis. Automatic burn wound diagnosis is still a largely unexplored field. In the related bibliography, one can find that there is a tendency to investigate objective methods for determining the depth of the burn in order to reduce the subjectivity and the high experience requirement that visual inspection demands. Some research into the relationship between depth and superficial temperature16 has been developed. There are also other works trying to evaluate burn depth by using thermographic images,17 infrared and ultraviolet images,18 radioactive isotopes19 and laser Doppler flux measurements20 On the other hand, there is hardly bibliography about burn depth determination by visual image analysis and processing. Although some research groups apply segmentation algorithms to burn images,5,7,8,21,22 they try to give an assessment of the healing of the burn, so they focused on calculating differences among several aspects such as area, shape, and appearance in order to give a prediction of the healing evoluAddress all correspondence to Begon ˜a Acha. Tel: +34-954487333; E-mail: [email protected] 1083-3668/2005/$22.00 © 2005 SPIE Journal of Biomedical Optics 10(3), 034014 (May/June 2005) 034014-1Journal of Biomedical Optics May/June 2005 䊉Vol. 10(3) Downloaded From: http://spiedigitallibrary.org/pdfaccess.ashx?url=/data/journals/biomedo/22403/ on 03/29/2017 Terms of Use: http://spiedigitallibrary.org/ss/termsofuse.aspx Fig. 1 Different appearances that could present a burn: (a) superficial dermal (blisters), (b) superficial dermal (red), (c) deep dermal, (d) full thickness (beige), (e) full thickness (brown). Fig. 4 Examples of the different 49⫻49 burn images used to train the classifier: (a) superficial dermal (blisters), (b) superficial dermal (red), (c) deep dermal, (d) full thickness (beige), (e) full thickness (brown). Acha et al.: Segmentation and classification... 034014-2Journal of Biomedical Optics May/June 2005 䊉Vol. 10(3) Downloaded From: http://spiedigitallibrary.org/pdfaccess.ashx?url=/data/journals/biomedo/22403/ on 03/29/2017 Terms of Use: http://spiedigitallibrary.org/ss/termsofuse.aspx tion of the wound. To our knowledge, only the group of Afromowitz et al.21,22 tries to give a diagnosis of the burn depth. From this assessment, they estimate the number of days that the wound will take to heal. They measure the optic reflectivity in the red, green, and infrared bands, hypothesizing that it is highly correlated with burn healing time, and they form a false color image that indicates the time of healing, or equivalently, the depth of the burn. The main disadvantage of the method is the complexity and cost of the image acquisition system 共video camera, filter wheel, motor driver, etc.兲. The main contribution of this work is the design of a clinically feasible system for automatic burn wound classification based on visual digital images. First, a protocol for the standardization of the burn image acquisition was designed. This first step was required due to the novelty of the application. Second, a new segmentation algorithm is proposed, which has been proven effective in segmenting burn wound images. Third, once the burnt part is segmented, representative color and texture descriptors are extracted from it. Finally, a neural network classifier processes these descriptors to give an estimation of the burn depth. 2 Materials and Methods 2.1 Burn Characterization There are three main types of burn wounds.1共1兲Superficial dermal burn: when the epidermis and part of the dermis are destroyed. The presence of blisters 共usually brown color兲 and/or a bright red color characterize it. It is painful. 共2兲Deep dermal burn: it is characterized by its pink-whitish color. 共3兲 Full-thickness burn: all the skin thickness is destroyed and skin grafts are needed. A beige-yellow or a dark brown color characterizes it. It is not painful. Although a burn wound is classified in three classes, it can present five different appearances. 共A兲Blisters: they are superficial dermal burns with a bright texture and a rose-brown color. 共B兲Bright red: they are superficial dermal burns with bright red colors and wet appearance. 共C兲Pink-white: they are deep dermal burns with a dotted appearance. 共D兲Yellowbeige: first appearance of full-thickness burns. 共E兲Brown: second appearance of full-thickness burns. Examples of each appearance are shown in Fig. 1. 2.2 Image Acquisition and Calibration The image acquisition was carried out by means of a digital photographic camera, the Canon EOS 300D 共Canon Inc., Tokyo, Japan兲. Any nonspecialized person should be able to acquire data from the patient, because it is not possible to have an expert in each center. A digital photographic camera is easy to utilize and people are used to them. The problems we found that had to be solved when using a digital photographic camera for this application are explained in the following subsections. 2.2.1 Illumination influence The most important source of information for our system in order to classify burn depths is color, which is extremely influenced by the illumination. In hospitals the lighting conditions can change depending on the room where the patient is. Then, measured pixel values depend on the illuminants and with multiple illuminants the measured values cannot be accurately converted to a known color space without some additional information. Therefore, a study about the influence of the different sources of illumination is needed. To perform this study, we photographed the Macbeth ColorChecker DC chart 共Gretag-Macbeth GmbH, Martinsried, Germany兲under three different illuminations: in a darkroom with the built-in flash 共guide number⫽13 m at ISO 100兲, in a darkroom with fluorescent light, and in a room under diffused sunlight. Under these three different situations, we fixed the ISO speed to 100, the f stop (Av)to 20 and we varied the exposure time (Tv). We define that the exposure time is optimum under a particular illuminant when it is the maximum time without saturating any channel. The ratio between the exposure times will give us the influences of the different sources of light. The optimum exposure times were 1/200, 0.6, and 1.6 s for the flash, sunlight, and fluorescent light, respectively. That means that the flash is 320 times stronger than the fluorescent and 120 times stronger than the sunlight. In other words, if we choose Tv⫽1/200 and 8 bits per color component, the fluorescent light will not influence even the least significant bit and the sunlight will influence the two least significant bits. In fact, we took a photograph under both fluorescent and sunlight illuminations with this parameter (Tv⫽1/200)and only these two least significant bits had values different to 0. We can conclude that the xenon flash illumination is sufficiently strong to dominate illumination. That is an important result because in this way we only have to calibrate the images once for each camera, and not for each room where patients are treated. 2.2.2 Calibration An additional problem we encountered is that manufacturers normally do not publish either the red 共R兲, green 共G兲, blue 共B兲 primaries of the camera or the color temperature of the flash. Therefore we need to determine in some way a transformation matrix to convert from measured RGB coordinates to a device-independent color representation system. For this purpose, we find the matrix transformation between RGB and CIE 共Commission Internationale de l’Eclairage XYZ 共device-independent color space兲. In the literature there are many transformation matrices from RGB to XYZ color space, but they are defined for specific illuminants 共D65, D50, etc.兲and specific RGB primaries 共CCIR Rec. 709, FCC-NTSC, etc兲.23 We have developed a calibration method based on the Macbeth ColorChecker DC chart, which is specifically designed for calibration of digital cameras. The Macbeth ColorChecker DC chart has 240 color chips and it is supplied with data giving the CIE XYZ chromaticity coordinates of each chip under D50 illuminant. The 240 chips occupy an area of 12 cm⫻20 cm. Our method finds the transformation matrix from RGB under unknown illuminant to XYZ under D50, and corrects the nonuniformity of the illumination as well as the spatial nonuniformity of the camera sensitivity. This algorithm iteratively performs the following steps: 1. Without correcting the illumination profile and using only three color patches, we calculate the initial matrix M1that converts from RGB under an unknown illuminant to XYZ under D50. 2. In the i’th step, using the 240 color patches in the chart Acha et al.: Segmentation and classification... 034014-3Journal of Biomedical Optics May/June 2005 䊉Vol. 10(3) Downloaded From: http://spiedigitallibrary.org/pdfaccess.ashx?url=/data/journals/biomedo/22403/ on 03/29/2017 Terms of Use: http://spiedigitallibrary.org/ss/termsofuse.aspx and the matrix Mi⫺1,we calculate the profiles, PR,i(x,y),PG,i(x,y),and PB,i(x,y),so that, for each patch, the R,G,Bcorrected with the profiles and multiplied by Mi⫺1are the X,Y,Zvalues specified by the manufacturer of the color chart. That is, for each patch kin the position (xk,yk)the following equation is performed: 冋 PR,i PG,i PB,i 册 ⫽ 冋 1/R共xk,yk兲 1/G共xk,yk兲 1/B共xk,yk兲 册 共Mi⫺1兲⫺1 冋 Xk Yk Zk 册 .共1兲 3. We calculate the three fourth order surfaces, PR,i ⬘(x,y), PG,i ⬘(x,y),and PB,i ⬘(x,y),that match best the profiles PR,i(x,y),PG,i(x,y),and PB,i(x,y)calculated in step 2. Previously, we have experimentally determined that a fourth order surface adequately approximates the sensitivity of the camera and the nonuniformity of the flash illumination altogether. 4. Using this profile, we calculate the matrix Mithat best maps the R,G,Bvalues into the X,Y,Zvalues specified for all the patches in the color chart. To determine this optimum Mithe following mean square error is minimized: ␧2⫽1 240 兺 k⫽1 240 共Xtk⫺Xk兲2⫹共Ytk⫺Yk兲2⫹共Ztk⫺Zk兲2, 共2兲 where Xtk,Ytk,and Ztkare the X,Y, and Zvalues of the k’th color patch, in the position (xk,yk),specified by the manufacturer. 5. Repeat from step 2 until the mean square error ␧begins to grow. It must be emphasized that the matrix Mis the product of two matrices: the transformation from RGB to XYZ under an unknown illuminant and the linear transformation to perform the chromatic adaptation from an unknown illuminant to D50. The matrix obtained with the proposed method is M⫽ 冋 45 60 ⫺19 24 93 ⫺23 33739 册 when the R,G,Bvalues are normalized to one. It should be noted that this matrix Mis specific for each camera, so calibration should be performed for every camera used. 2.2.3 Acquisition protocol The third problem consists of fixing the acquisition protocol so that the photographs are useful for diagnosis. After fixing it we have validated its suitability. The acquisition protocol was developed by an interdisciplinary group formed by burn specialized physicians and technicians.24 The main points of the acquisition protocol were the following: distance between camera and patient should be about 40–50 cm 共to fix this parameter, physicians carried out a careful analysis of photographs taken of different burn wounds from different distances; in the end, they chose 40–50 cm because they could distinguish texture from this distance and, at the same time, they usually had a global vision of the burn兲, healthy skin should appear in the image when possible, the background should be a green/blue sheet 共the ones used in hospitals, because as the blue/green color is so different from the skin colors, the background can be easily rejected by the segmentation algorithm兲, the flash must be on and the camera should be placed parallel to the burn. The parameters of the camera were set to: ISO speed 100, exposure time 1/200 s and aperture 共f stop兲20. In order to validate the acquisition protocol, a survey was done.24,25 For this survey, 38 photographs of all etiologies, locations, and characteristics of the most frequent lesions were taken following the specified protocol. They were presented to a panel of 12 experts in burn diagnosis. The experts had to answer about the certainty in diagnosis 共1–5兲: 1⫽minimal, 3⫽moderate, 5⫽maximum, certainty. A mean of 4.26 in sureness in diagnosis and 84.6% of diagnostic accuracy was answered, whereas diagnostic accuracy of a trained plastic surgeon when looking live at the same 38 burn wounds was 84.3%. 2.3 Burn Wound Segmentation The segmentation approach used here is a supervised pixelbased algorithm based on measures in the CIE L*u*v*color coordinate space. L*u*v*and L*a*b*color representation systems are called uniform systems because Euclidean distances between colors measured in these spaces are very much correlated with color differences according to human perception. They are particularly useful in color image segmentation of natural scenes using histogram-based techniques, in which our method is included. They are slightly different because of the different approaches to their formulation. Nevertheless, both spaces are equally good in perceptual uniformity and provide very good estimates of color difference 共distance兲between two color vectors.23 Therefore, we could have chosen any of these two spaces, but we preferred the L*u*v*one, because the color components a*and b* do not depend on the luminance, and it is known that color perception is strongly influenced by the luminance.26 The following steps show the scheme proposed: 2.3.1 Selection of a small region in the burn wound by the user and preprocessing of the image For a nonexpert physician 共in fact, for most of the people兲it is easy to differentiate burnt skin from normal one. Therefore, the burn wound will be segmented using the color information ofa5⫻5 pixel area around the point that the user selects with the mouse. Before segmenting the image, it is convenient to preprocess it in order to get more homogeneous regions eliminating noise and small structures. To perform this task, an anisotropic diffusion is applied to the color image.27,28 The aim of the diffusion is to make the regions more homogeneous but preserving the edge information. In order to perform the anisotropic diffusion, the approach of separating the diffusion of the chromatic and achromatic information was followed28 as is shown in Fig. 2. First, the image is converted into L*u*v* color coordinate system according to23 Acha et al.: Segmentation and classification... 034014-4Journal of Biomedical Optics May/June 2005 䊉Vol. 10(3) Downloaded From: http://spiedigitallibrary.org/pdfaccess.ashx?url=/data/journals/biomedo/22403/ on 03/29/2017 Terms of Use: http://spiedigitallibrary.org/ss/termsofuse.aspx L*⫽ 再 116 冉 Y Y0 冊 1/3 ⫺16, if Y Y0 ⬎0.008856 903.3 冉 Y Y0 冊 ,otherwise .共3兲 Computation of u*and v*involves intermediate u⬘,v⬘, u0 ⬘,and v0 ⬘quantities defined as u⬘⫽4X X⫹15Y⫹3Z, 共4兲 v⬘⫽9Y X⫹15Y⫹3Z. Finally, u*⫽13L*共u⬘⫺u0 ⬘兲, 共5兲 v*⫽13L*共v⬘⫺v0 ⬘兲. Y0,u0,and v0correspond to the white reference point, which depends on the illuminant 共D50 after the calibration兲. From these coordinates, the hue and chroma components are calculated as H⫽arctan(v*/u*)and C⫽ 冑 (u*)2⫹(v*)2, respectively. A complex quantity is calculated that relates the hue and the chroma as P⫽Cexp(jH). The achromatic anisotropic diffusion, applied to L*,is carried out by means of the discrete formulation27 of the partial differential equation ⳵ ⳵ tL*共x,y,t兲⫽div关 ␣ 共x,y,t兲ⵜL*共x,y,t兲兴,共6兲 where div and ⵜdenote the divergence and the gradient operators, respectively, and ␣ (x,y,t)is a monotonically decreasing function of the image gradient magnitude called the conductance coefficient and is given by ␣ 共x,y,t兲⫽1 1⫹ 冉 兩 ⵜL*共x,y,t兲 兩 ␥ p 冊 2.共7兲 The diffusion constant ␥ pwas selected as the 5% of the maximum value of 兩 ⵜL*(x,y,t) 兩 at each t, an artificial time parameter that denotes the number of diffusion iterations, which was fixed to 20. The chromatic anisotropic diffusion is performed by applying Eq. 共6兲to the complex quantity P ⳵ ⳵ tP共x,y,t兲⫽div关 ␣ 共x,y,t兲ⵜP共x,y,t兲兴,共8兲 where ⵜP(x,y,t)is28 ⵜP共x,y,t兲⫽关ⵜC共x,y,t兲⫹jCⵜH共x,y,t兲兴exp关jH共x,y,t兲兴 共9兲 and separating real and imaginary parts of Eq. 共8兲it follows that ⳵ ⳵ tC⫽div共 ␣ ⵜC兲⫺ ␣ C 兩 ⵜH 兩 2, 共10兲 ⳵ ⳵ tH⫽div共 ␣ ⵜH兲⫹2 ␣ CⵜC•ⵜH, where the spatial and temporal dependencies have been omitted for convenience. To obtain the coefficient ␣ for the complex quantity Pwe need to calculate 兩 ⵜP(x,y,t) 兩 ,which is 兩 ⵜP共x,y,t兲 兩 ⫽ 冑 兩 ⵜC共x,y,t兲 兩 2⫹C2共x,y,t兲 兩 ⵜH共x,y,t兲 兩 2. 共11兲 2.3.2 Conversion to single channel image In this step a gray scale image is obtained from the diffused color image. In this gray scale image, differences between the burnt skin selected by the user and other parts of the image are emphasized. Based on the observation that doctors segment burn wounds by measuring differences among colors, the selection box selected by the user is slid as a mask of size 5⫻5 pixels along the image and, for each pixel in the image under the center of the sliding mask, the following operation is performed:29 f共n,m兲⫽1 MAX 兺 i⫽n⫺⌬ n⫹⌬ 兺 j⫽m⫺⌬ m⫹⌬ dE关p共i,j兲,w共i,j兲兴, 共12兲 where MAX is max n,m(兺i⫽n⫺⌬ n⫹⌬兺j⫽m⫺⌬ m⫹⌬dE„p(i,j),w(i,j)…),⌬ ⫽(L⫺1)/2 with L⫽5, p(i,j)represents a pixel in the diffused image to be segmented in L*u*v*color space, w(i,j) is a pixel of the mask selected by the user, and dE(•),the Euclidean distance between pixels p(i,j)and w(i,j),is defined as Fig. 2 Diffusion filtering separating chromatic and achromatic information. Acha et al.: Segmentation and classification... 034014-5Journal of Biomedical Optics May/June 2005 䊉Vol. 10(3) Downloaded From: http://spiedigitallibrary.org/pdfaccess.ashx?url=/data/journals/biomedo/22403/ on 03/29/2017 Terms of Use: http://spiedigitallibrary.org/ss/termsofuse.aspx dE共p共i,j兲,w共i,j兲兲⫽ 兵 关Lp *共i,j兲⫺Lw *共i,j兲兴2⫹关up *共i,j兲 ⫺uw *共i,j兲兴2⫹关vp *共i,j兲 ⫺vw *共i,j兲兴2 其 1/2.共13兲 2.3.3 Thresholding operation and postprocessing The result of the above step is a gray-scale image where pixels with lowest values are those in the region to be segmented. This image has been carefully designed to emphasize the burnt regions, and a thresholding operation should suffice to get a good segmentation. The histogram of this distance image is multimodal so a method to find a threshold to select the mode in the left of the histogram should be found. This task is carried out in two steps: 共1兲the peaks 共maximum values兲of the different modes present in the histogram are found, and 共2兲the threshold which separates the two modes closest to the left of the histogram is calculated applying Otsu’s thesholding method.30 To perform the first step, the following algorithm is applied to the histogram of the gray-scale image: 共1兲find all peaks in the histogram, that is, all the values in the histogram which are higher than their two neighbors; 共2兲form a new curve with the peaks found in the previous step and then select again the peaks in the new curve; 共3兲remove nonsignificant peaks, i.e., those peaks whose values are less than 1% of the maximum peak value are rejected; 共4兲remove nonsignificant valleys, that is, if two peaks have not a significant valley between them we maintain only the highest of the two peaks. To check if a valley is significant or not, the minimum value between two peaks is found. If this minimum value is greater than 75% of the lowest peak out of the two peaks, then the valley is considered nonsignificant. These four steps are illustrated in Fig. 3. Once we have localized the main modes in the histogram, we have to find the threshold which separates the two modes closest to the left part of the histogram. This task is carried out by applying Otsu’s method,30 which is an adaptive thresholding technique to split a histogram into two classes, c1with gray levels 关1,...,k兴,and c2with gray levels 关k⫹1,...,K兴.Let mi(k)and mTbe the mean intensities for the class ciand for the whole image, respectively. The between-class variance was defined by Otsu as ␴ b 2共k兲⫽ ␻ 1共k兲共m1共k兲⫺mT兲2⫹ ␻ 2共k兲共m2共k兲⫺mT兲2, 共14兲 where ␻ 1(k)and ␻ 2(k)are cumulative sums of the probabilities in each class, that is, ␻ 1(k)⫽兺j⫽1 kpj, ␻ 2(k) ⫽兺j⫽k⫹1 Kpj,and pj⫽xj/Npixels ,where xjis the number of pixels with gray level jin an image and Npixels is the number of pixels with gray levels from 1 to Kin the whole image, that is, the total number of pixels in the image. The optimal threshold k ˆis chosen so that the between-class variance ␴ b 2is maximized. The election of Otsu’s method, among many existing thresholding methods, is due to its simplicity in computation.31 In fact, many modern segmentation algorithms are based in Otsu’s method or use it for comparison.32–34 Finally, by the application of a 3⫻3 median filter, the segmentation result is improved by removing spurious points 共1–4 pixel sized兲, that is, points that have been segmented and do not actually belong to the burn. 2.4 Classification Once the burn is segmented, its depth must be estimated for classification purposes. It has been proven that physicians determine the depth of a burn based on color perception, as well as on some texture aspects. As it has been previously said, L*u*v*space is a perceptually uniform color representation system. Also, the hue and the chroma coordinates are intimately related to the way human beings perceive chromaticity. That is why, in this study, a set of descriptors formed by statistical moments of the histograms obtained for each coordinate of the L*u*v*color space, as well as for the hue and chroma image planes derived from them, have been used. More specifically, the descriptors chosen are: mean of lightness (L*),mean of hue 共H兲, mean of chroma 共C兲, standard deviation of lightness ( ␴ L),standard deviation of hue ( ␴ H), standard deviation of chroma ( ␴ C),mean of u*,mean of v*, standard deviation of u*( ␴ u),standard deviation of v*( ␴ v), skewness of lightness (sL),kurtosis of lightness (kL),skewness of u*(su),kurtosis of u*(ku),skewness of v*(sv)and kurtosis of v*(kv). Afterwards it has been necessary to apply a descriptor selection method to obtain the optimum set for the subsequent classification. 2.4.1 Feature selection The discrimination power of these 16 features is analyzed using the sequential forward selection 共SFS兲method and the Fig. 3 Process of detecting the main peaks in the histogram. (a) Detection of the peaks in the histograms: peaks are marked with circles. (b) Finding the peaks in the histogram of the peaks: peaks from the original histogram are marked with dots and new peaks with circles. (c) Rejection of nonsignificant peaks: peaks from Fig. (b) are marked with dots and peaks selected in this step are marked with circles. (d) Final peaks in the original histogram after the rejection of peaks without a significant valley between them. In this case the three peaks in the former step are accepted. Acha et al.: Segmentation and classification... 034014-6Journal of Biomedical Optics May/June 2005 䊉Vol. 10(3) Downloaded From: http://spiedigitallibrary.org/pdfaccess.ashx?url=/data/journals/biomedo/22403/ on 03/29/2017 Terms of Use: http://spiedigitallibrary.org/ss/termsofuse.aspx sequential backward selection 共SBS兲method35,36 via the Fuzzy-ARTMAP neural network which is detailed in the following subsection. SFS is a bottom-up search procedure where one feature at a time is added to the current feature set. At each stage, the feature to be included in the feature set is selected among the remaining available features which have not been added to the feature set. So the new enlarged feature set yields a minimum classification error comparing to adding any single feature. The algorithm stops when adding a new feature yields an increase of the classification error. The SBS is the top-down counterpart of the SFS method. It starts from the complete set of features and, at each stage, the feature which shows the least discriminatory power is discarded. The algorithm stops when removing another feature implies an increase of the classification error. To apply these two methods, 50 49⫻49 pixel images for each burn appearance have been used 共see Fig. 4兲. As there are five appearances, in all we have 250 49⫻49 pixel images.*One photograph has been taken per burn wound. In general, we selected only one 49⫻49 pixel image per photograph, unless there were different appearances in the same wound. In this case, one 49⫻49 image per appearance was selected. The selection performance is evaluated by fivefold cross validation 共XVAL兲.15 In this sense, the disadvantage of sensitivity to the order of presentation of the training set, that the SBS and SFS methods present,35 is diminished. To perform the XVAL method the 50 images per burn appearance are split into five disjoint subsets. Four of these subsets 共that is, 40 images per appearance兲serve as a training set for the neural network, while the other one 共ten images兲is used as validation set. Then, the procedure is repeated interchanging the validation subset with one of the training subsets, and so on till the five subsets have been used as validation sets. The final classification error is calculated as the mean of the errors for each XVAL run. In Fig. 5 the evolution of the classification error is presented for both selection methods. It can be observed that both curves coincide at the beginning and at the end, but then they separate obtaining a minimum classification error with seven or eight descriptors 共2% error兲for the SFS method, and six descriptors 共1.6% error兲for the SBS method. In fact, this minimum error is again reached with 12 descriptors, although it is reasonable to choose the set of six, because it will imply less complexity in the neural network and shorter processing time. The six descriptors provided by SBS method were chosen as the best feature set: lightness, hue, standard deviation of the hue component, u*chrominance component, standard deviation of the v*component, and skewness of lightness. 2.4.2 FuzzyARTMAP neural network The classifier used is a Fuzzy-ARTMAP neural network. This type of network is based on the Adaptive Resonance Theory developed by Grossberg and Carpenter. Fuzzy-ARTMAP is a supervised learning classification architecture for analogvalue input pairs of patterns.37 The reasons for this choice are that Fuzzy-ARTMAP offers the advantages of well-understood theoretical properties, an efficient implementation, clustering properties that are consistent with human perception, and a very fast convergence. It has also a track record of successful use in industrial and medical applications.38 Other strongpoints of this type of neural network are the small number of design parameters 共the vigilance parameter, ␳ a苸关0,1兴,and the selection parameter, ␣ ⬎0兲, and that the architecture and initial values are always the same, independent of the application. When the input parameters are the features selected by the SBS method above, the network classifies the burn depth of the segmented region into five types: the first and the second belonging to superficial dermal depth, the third to deep dermal, and the fourth and fifth to full thickness. So, the network has six neurons in the input layer and five neurons in the output layer. In the Fuzzy-ARTMAP neural network the architecture is dynamic, so the number of neurons in the hidden layer is fixed during the training and according with the vigilance parameter. 3 Experimental Results The images used to test the burn CAD tool were 62 digital photographs taken by physicians following the acquisition protocol. All the images were diagnosed by a group of plastic surgeons, affiliated with the burn unit of the Virgen del Rocı ´o Hospital, from Seville 共Spain兲. The assessments were validated one week later, as is the common practice when handling burnt patients. The images were 1536⫻1024 pixels and they were stored as JPEG 共high quality兲files. The computer used was a Pentium IV, 1.7 GHz and 256 MB of random access memory. The average run time was 4 min for an image and the programming tool was MATLAB 6.1 共The Mathworks Inc., Natick, Massachusetts兲. *The 250 49⫻49 pixel images are small images showing each one only one burn appearance 共no healthy skin or background兲. Each 49⫻49 pixel image has been validated by two physicians as belonging to a particular depth. Therefore, these 250 images form a database used only for the feature selection step. Fig. 5 Evolution of the classification error for SFS method (䊉) and SBS method (䊊). Acha et al.: Segmentation and classification... 034014-7Journal of Biomedical Optics May/June 2005 䊉Vol. 10(3) Downloaded From: http://spiedigitallibrary.org/pdfaccess.ashx?url=/data/journals/biomedo/22403/ on 03/29/2017 Terms of Use: http://spiedigitallibrary.org/ss/termsofuse.aspx 3.1 Segmentation Results The segmentation algorithm proposed in this paper was tested with 35 out of the 62 images of the database. These 35 images were manually segmented by five physicians. The reason of using 35 photographs instead of 62 is that, although the protocol says that it should appear as healthy and burnt skin, very often the extension of the burn wound is so large that there is only burnt skin in the image. Therefore, in these cases it is not meaningful to compare the segmentation results performed by the physicians and by the algorithm. The segmentation gold standard was obtained by applying the voting method to the regions segmented by the five specialists. In other words, one pixel was considered to belong to the segmented region in the gold standard if most of the physicians had considered it in this way. Once a gold standard was obtained, two parameters were calculated to measure the performances of the segmentation algorithm. The first parameter was the positive predictive value 共PPV兲, which measures the ratio between the number of pixels segmented by the algorithm which fit the segmentation gold standard and the total amount of pixels segmented. The second parameter is called sensitivity 共S兲, and it is the ratio between the number of pixels segmented by the algorithm which fit the segmentation gold standard and the total amount of pixels in the segmentation gold standard. Intuitively it can be seen that the first parameter measures the over segmentation, which would be null if PPV were 1. Likewise, Smeasures the under segmentation. In Table 1 the results for the 35 images are presented. As is shown in this table, almost all the photographs are properly segmented. It must be emphasized that, although the sensitivity tends to be only around 0.8, this is because doctors tend to over segment the burnt region. Therefore, this should not be interpreted as a poor performance of the algorithm. Figures 6–8 show the segmentation results for some images of the three types of depth. Figures 共a兲represent original images and Figs. 共b兲represent the segmented ones. In the segmented images we have marked with yellow color the segmented region. In all the cases, the burn wound was segmented correctly from the normal skin. 3.2 Classification Results To test the classification part we employed the 62 images of the database used for validation 共different from the one used for training兲. The neural network was trained with the 250 49⫻49 pixel images previously cited. The training was performed with ␳ a⫽1and ␣ ⫽0.001. At the end of the training the weights were fixed for the subsequent classification test. For this test the six features were extracted from the segmented part of the 62 images. Classification results are summarized in Table 2. We have used 22 images with superficial dermal burns, 18 with deep dermal burns, and 22 with fullthickness burns. The average success percentage was 82.26%. All superficial dermal burns misclassified were classified by the network as deep dermal ones. All deep dermal burns were misclassified as superficial dermal ones. And, in the case of misclassified full-thickness burns, 80% of them were classified as superficial dermal and 20% as deep dermal. 4 Discussion and Conclusions The classification of burn depths based on visual inspection is a difficult task, which needs a lot of training. That is why in burn related literature there is a constant search for objective methods to determine the depth of a burn. A prototype of one invasive technique is the acquisition of biopsies and their histological study for the burn depth diagnosis.39 This technique, although it can be considered as ‘‘gold standard,’’ is not exempt from problems related to loss of dermis in the burn, to the existence of considerable variability depending on where the biopsy was acquired, and to the fact that this technique is a snapshot view of the lesion, apart from the residual scars provoked by the biopsy acquisition. These inconveniences have directed efforts towards the design of noninvasive procedures. Some noninvasive techniques analyze the perfusion of the burn wound based on the fact that tissue damage is inversely proportional to the vascularization after the lesion.40–42 Nevertheless, in these procedures it is necessary to supply a vital colorant to the patient by intravenous method and it is essential to have an emergency system. Other experimental techniques analyze the changes in optical properties of the skin related to the changes of its vascularization,43 although their application environment is, for the moment, exclusively experimental. In another type of approximation to the problem being studied, the remission-optical measurement exploits the different spectral backscattering effects of burned Table 1 Quantification of segmentation results (PPV: positive predictive value; S: sensitivity). Image PPV S Image PPV S Image 1 0,9309 0,8093 Image 19 0,8303 0,9280 Image 2 0,9314 0,6969 Image 20 0,9627 0,9005 Image 3 0,9391 0,8684 Image 21 0,9196 0,7418 Image 4 0,9302 0,8324 Image 22 0,8752 0,7789 Image 5 0,9614 0,9015 Image 23 0,9559 0,9725 Image 6 0,9741 0,8853 Image 24 0,8622 0,9107 Image 7 0,8807 0,7297 Image 25 0,9082 0,9069 Image 8 0,8984 0,8108 Image 26 0,9646 0,7989 Image 9 0,9618 0,7772 Image 27 0,9320 0,9364 Image 10 0,9737 0,8206 Image 28 0,8711 0,8457 Image 11 0,7928 0,8190 Image 29 0,9569 0,9482 Image 12 0,9624 0,7452 Image 30 0,9571 0,8814 Image 13 0,9806 0,7248 Image 31 0,9134 0,8318 Image 14 0,9424 0,7820 Image 32 0,9327 0,8698 Image 15 0,9384 0,8457 Image 33 0,6990 0,7588 Image 16 0,8327 0,8066 Image 34 0,9192 0,5174 Image 17 0,6420 0,8539 Image 35 0,7701 0,8530 Image 18 0,8788 0,9646 Average 0,9023 0,8301 Acha et al.: Segmentation and classification... 034014-8Journal of Biomedical Optics May/June 2005 䊉Vol. 10(3) Downloaded From: http://spiedigitallibrary.org/pdfaccess.ashx?url=/data/journals/biomedo/22403/ on 03/29/2017 Terms of Use: http://spiedigitallibrary.org/ss/termsofuse.aspx Fig. 6 Segmentation result for a superficial dermal burn. (a) Original image where the selection made by the user is shown with an arrow. (b) Segmented image. Fig. 7 Segmentation result for a deep dermal burn. (a) Original image where the selection made by the user is shown with an arrow. (b) Segmented image. Fig. 8 Segmentation result for a full thickness burn. (a) Original image, which has both superficial dermal burn (the red part) and full-thickness burn (the creamy part). (b) Segmented image. In this case the user has made the selection in the creamy part in order that the algorithm segments all the full-thickness part of the burn. It segments correctly all the full-thickness parts of the image regarding what physicians said. Acha et al.: Segmentation and classification... 034014-9Journal of Biomedical Optics May/June 2005 䊉Vol. 10(3) Downloaded From: http://spiedigitallibrary.org/pdfaccess.ashx?url=/data/journals/biomedo/22403/ on 03/29/2017 Terms of Use: http://spiedigitallibrary.org/ss/termsofuse.aspx