Automatic lane segmentation in TLC images using the continuous wavelet transform
Abstract
This paper describes a new methodology for lane detection in Thin-Layer Chromatography images. An approach based on the continuous wavelet transform is used to enhance the relevant lane information contained in the intensity profile obtained from image data projection. Lane detection proceeds in three phases: the first obtains a set of candidate lanes, which are validated or removed in the second phase; in the third phase, lane limits are calculated, and subtle lanes are recovered. The superior performance of the new solution was confirmed by a comparison with three other methodologies previously described in the literature.
Full text
Hindawi Publishing Corporation Computational and Mathematical Methods in Medicine Volume 2013, Article ID 218415, 19 pages http://dx.doi.org/10.1155/2013/218415 Research Article Automatic Lane Segmentation in TLC Images Using the Continuous Wavelet Transform Bruno Moreira,1,2 António Sousa,1,3 Ana Maria Mendonça,1,2 and Aurélio Campilho1,2 1INEB-Instituto de Engenharia Biom´ edica, Campus da FEUP, Universidade do Porto, Rua Dr. Roberto Frias, s/n, 4200-465 Porto, Portugal 2Faculdade de Engenharia da Universidade do Porto (FEUP), 4200-465 Porto, Portugal 3Instituto Superior de Engenharia do Porto (ISEP), Instituto Polit´ ecnico do Porto, 4200-072 Porto, Portugal Correspondence should be addressed to Bruno Moreira; mor[email protected]p.pt Received 13 May 2013; Revised 22 July 2013; Accepted 1 August 2013 Academic Editor: Carlo Cattani Copyright © 2013 Bruno Moreira et al. This is an open access article distributed under the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited. This paper describes a new methodology for lane detection in Thin-Layer Chromatography images. An approach based on the continuous wavelet transform is used to enhance the relevant lane information contained in the intensity profile obtained from image data projection. Lane detection proceeds in three phases: the first obtains a set of candidate lanes, which are validated or removed in the second phase; in the third phase, lane limits are calculated, and subtle lanes are recovered. The superior performance of the new solution was confirmed by a comparison with three other methodologies previously described in the literature. 1. Introduction This paper focuses on one of the initial components of a screening tool for Fabry disease (FD), based on the automatic analysis of Thin-Layer Chromatography (TLC) images: the lane segmentation. FD is a rare X-linked hereditary metabolic storage disorder caused by genetic abnormalities, which leads to an enzymatic deficiency [1] resulting in the accumulation of excessive quantities of one class of lipids, the sphingolipids, globotriaosylceramide (Gb3) being the most prevalent in FD patients [2]. Although the usual onset of the first symptoms is in childhood, by middle-age, life-threatening complications are often developed in untreated patients [3]. The recent availability of enzymatic replacement therapy, in conjunction with the progressive nature of the disease, has renewed the interest in this disorder and revealed the need for early diagnosis, which can only be achieved with generalized screening programs [4]. The complete diagnosis of FD is very complex, but the first phase is simply based on the detection of an abnormal quantity of Gb3 in urine or blood plasma of the patient. The direct measurement of these compounds can be carried out by using a microtandem mass spectrometer (MS/MS), but its use is very expensive. Another approach, less expensive, is the analysis of patient urine or blood plasma samples using TLC [5]. TLC is a type of liquid chromatography that allows the separation, identification, and visual quantification of a wide variety of components in a mixture [6]. The components to be separated by the chromatographic process are distributed between two phases, a stationary phase and a mobile phase. The solutes, distributed preferentially in the mobile phase, will move more rapidly through the system than those distributed preferentially in the stationary phase. Thus, the solutes will elute in order of their increasing distribution with respect to the stationary phase [7]. At the end of the chromatographic process, the components are spread along a lane and distributed by different bands based on their physical properties (size, molecular weight, etc.) [6]. TLC has the highest sample throughput amongst the chromatographic techniques. Up to 30 different samples and standards can be appliedtoasingleplateinindividuallanesandbeanalyzed at the same time, which explains the spreading use of TLC, as both a screening and confirmation tool, worldwide [8]. For the implementation of a screening tool for Fabry disease, the identification of normal and suspicious individuals will be based on the presence or absence of the disease biomarkers, previously separated by the TLC development
2 Computational and Mathematical Methods in Medicine of urine samples. The screening system is supported by a set of image analysis and classification methods, so as to try to automate the interpretation of TLC chromatograms. One initial and crucial phase, usually called lane segmentation, involves the automatic separation of the individual samples contained in the digital image of a TLC plate. Several software packages that include solutions for lane segmentation can be found in the literature, although most of them were developed for gel electrophoresis images. KODAK 1D [9]softwarehasan“automaticlanefinder”optionthat uses a multiple pass algorithm to determine the lanes on the image, as an alternative to the interaction with the operator. GelBuddy [10] is a user-friendly Java-based software that requires the number of lanes as input. The lane tracks are located by detecting local maxima of intensity profiles, obtained through the integration of pixel values over a set of horizontal sectors. Getlanes [11]operatesonfour-color, fluorescence-based, and electrophoretic gel images. Each gel filecontainsfourfilterimages,whicharesummedtoform a brightness image. The “brightness” profile is obtained by a vertical integration of the brightness image, and a firstdifference approximation to the gradient is computed to identify maxima, which are marked as lane locations. It also uses models of expected lane and interlane spacing and lateral lane behavior to improve tracking on imperfect gels. PyElph [12] is a software tool for gel image analysis and phylogenetics. For lane detection, it computes the maximum value of each column and creates an intensity profile. A threshold set to 70% of the maximum intensity value of the domain is used forlanedetection.Thelanesselectedatthisthresholdlevel areusedforcalculatingthemeanvalueforthelanes’width. The lanes narrower than the mean width are removed from the set, and the mean width of the remaining lanes is again computed. Finally, all the lanes between two thresholds (70% and 15% of the maximum of the domain), and presenting awidthdeviationoflessthan25%fromthemeanwidth, areincludedinthefinalset.LaneRuler[13] approaches the problem by dividing the data region into “zones” such that each zone has the same integrated intensity, thus placing more “nodes” in data-rich portions of the gel compared to its datapoor portions. Generic lane width is then determined based on a Fourier analysis in each zone. Methods developed for lane segmentation include algorithms supported by the application of spatial-domain filters [14] and semiautomatic detection based on the assumption that the lanes have constant width and are equispaced [15]. Bajla et al.[16] proposed a semiautomatic method for lane separation in electrophoretic gel images. This approach is based on a one-dimensional cumulative indicator of lane edges, coupled with a shifted regular initial grid calculated from the aprioriinformation on the number of lanes. In [17], the bands contained in each lane are enhanced and afterwards replaced by their skeletons. Then, lane segmentation is performedbasedonthesebandskeletons.In[18], two methods are proposed, one based on an iterative moving average filter (IMA) and the other using the continuous wavelet transform (CWT). Sousa et al.[19] presented an automatic procedure for lane detection in TLC images based on the detection of local extreme points in the image projection profile, which had been previously smoothed using a nonweighted moving average filter. In the approach described in [20,21], lane detection is accomplished by locating the lane boundaries. The derivative of image intensities in the horizontal direction is calculated, and its values are summed across the vertical direction. The resulting one-dimensional curve has local extremes at the boundaries of the lanes. In [22], the image is horizontally divided into equal parts, and for each one a profile is obtained from its vertical projection. Each profile issmoothedandusedtoestimatethelanecenters(profile maxima) in each image part. The methodology connects local maxima along a lane, by visiting each partitioned image from the bottom to the top, and checking whether two local maxima are within a range of horizontal coordinates. Other methods related to lane detection include algorithms to geometrically correct images with the help of distance maps [23]andactiveshapemodels[24]. Most of the packages and methods that were developed for gel image analysis assume some regularity in lane distribution or require additional inputs, such as the number of lanes in the image. So, this kind of solution does not work well when the images do not present the expected regularity inlanedistribution,asoccursinsomeimagesofourdataset. Therefore, an approach that is able to deal with the specific characteristics of TLC images, without lowering performance when applied to more standard ones, is required in our system. This paper presents a new methodology for automating the detection of lanes in digital images of TLC plates, which is an improved version of the algorithm presented in [25]. The main difference between the two methods is the initial smoothing step, which is applied to the intensity profile resulting from the integration of the image area containing thelanes.Theinitialsmoothingstepisessentialforthesuccess of the following lane segmentation steps, as it allows the removal of noise and irrelevant data. In [25], a Savitzky-Golay (SG) filter [26] was applied to the intensity profile, while in the solution herein proposed, the Continuous Wavelet Transform (CWT) [27] is used for removing both noise and other high frequency components. When compared to other algorithms in the literature, our approach also relies on the detection of extreme points of the intensity profile that results from the projection of the image data. However, both in the method herein described and in [25], after locating the most obvious lanes associated with the extreme points of the smoothed profile, two further steps are performed, one for validating previously included lanes and removing false detections and another for adding very subtle lanes that could not be clearly distinguished. Nevertheless, a new implementation of these two supplementary steps is now proposed using image adaptive parameters instead of the constant settings used in [25]. Theoutlineofthepaperisasfollows.Section 2 describes the proposed method, including the preprocessing sequence appliedtotheimages,theCWT-basedtechniqueforsmoothing the intensity profile, and the lane segmentation process. The dataset description, lane segmentation examples, and some performance measures are presented in Section 3.
Computational and Mathematical Methods in Medicine 3 Input image Figure 1(b) ROI detection and preprocessing Enhanced image ROI Figure 1(c) Profile generation 1-D profile Figure 1(d) Profile smoothing Smoothed profile Figure 1(e) Detection of an Detection of subtle lanes initial set of lanes Initial set of lanes Figure 1(f) Removal of false lanes Lane segmentation Figure 1(g) (a) Column (pixel) 500 1000 1500 2000 200 400 600 800 1000 Line (pixel) (b) Column (pixel) 500 1000 1500 200 400 600 Line (pixel) (c) 0 500 1000 1500 0 5 10 15 20 25 30 Intensity P(x) Column (x) (d) 0 500 1000 1500 0 5 10 15 20 25 30 Intensity P(x) Column (x) (e) Figure 1: Continued.
4 Computational and Mathematical Methods in Medicine Column (pixel) 500 1000 15000 0 200 400 600 Line (pixel) (e) Column (pixel) 500 1000 15000 0 200 400 600 Line (pixel) (f) Figure 1: Main steps of the methodology: (a) block diagram; (b) original TLC image; (c) enhanced image ROI; (d) ROI intensity profile; (e) smoothed profile; (f) initial set of lanes; (g) final lane segmentation. 200 400 600 800 100 200 300 400 500 600 Column (pixel) Line (pixel) (a) 200 400 600 800 5 10 15 20 Intensity P(x) Column (x) (b) Scale 200 400 600 800 1 10 100 Noise range Lane range Baseline range Column (x) 01 Mean coefficientstheof amplitude 200 400 600 800 3 4 5 6 7 8 Intensity P(x) Column (x) (c) (d) (e) (c) Figure 2: (a) ROI of a TLC image. (b) Profile intensity obtained for the same image. (c) Color map representation of the CWT coefficients for the intensity profile. (d) The lane range limits are represented by the red lines. (e) Smoothed profile resulting from the reconstruction of the scales within the lane range. Furthermore, the proposed methodology is compared with three previous solutions, described in [18,19,25]. Results are discussed in Section 4. Some conclusions and further research are included in Section 5. 2. Methodology for Lane Segmentation 2.1. Overview of the Methodology. The processing flow of the method for lane segmentation herein proposed is illustrated in the block diagram of Figure 1(a).TheoriginalRGBimage is acquired, and the region of interest (ROI) is obtained, as presented by the dashed rectangle in Figure 1(b).The ROI is preprocessed, and the enhanced image is shown in Figure 1(c).Theresultingimagecolumnsareintegratedin order to obtain an intensity profile, which is then smoothed using the technique described in detail in Section 2.3.Figures 1(d) and 1(e) show the initial and smoothed profiles, respectively, for the image data of Figure 1(c).Thesmoothed profile will allow the selection of an initial set of potential lanessuchasthoseinFigure 1(e),whichisafterwardsupdated with the removal of false detections. The last phase focuses on the detection of missed lanes and on the localization of lane centers and boundaries, as can be observed in Figure 1(f). 2.2. Image Acquisition and Preprocessing. At the end of the chromatographic process, TLC plates quickly deteriorate,
Computational and Mathematical Methods in Medicine 5 Scale (pixel) 200 400 600 800 1 10 100 Column (x) 0 0.5 1 1.5 Mean amplitude of the coefficients 200 400 600 800 3 4 5 6 7 Intensity P(x) Column (x) (a) (b) (c) (a) Scale (pixel) 200 400 600 800 1 10 100 Column (x) 0 0.5 1 1.5 Mean amplitude of the coefficients 200 400 600 800 3 4 5 6 7 Intensity P(x) Column (x) (d) (e) (f) (b) Scale (pixel) 200 400 600 800 1 10 100 Column (x) 0 0.5 1 1.5 Mean amplitude of the coefficients 200 400 600 800 3 4 5 6 7 Intensity P(x) Column (x) (g) (h) (i) (c) Figure 3: (a) Representation of the scales used to smooth the intensity profile. (b) Cutoff-min set to 50. (c) The effect of the exclusion of lower scales is reflected in the reconstructed profile. (d) Representation of the scales used to smooth the intensity profile. (e) Cutoff-max set to 150. (f) The exclusion of scales containing low frequency information may lead to the appearance of false lanes. (g) Representation of the scales used to smooth the intensity profile. (h) Cutoff values chosen using the proposed methodology (scales 80 and 250). (i) Reconstructed profile. and therefore they need to be scanned as soon as possible. The TLC digital image is acquired in true color, since the specialist will need this information for inspecting the images later, as it helps in the identification of specific biomarkers. A typical TLC image, such as the one shown in Figure 1(b), usually contains two distinct regions: a border and a region of interest (ROI). The border is the external region of the image, and it comes from the adhesive tape used for protecting the plate right after the chromatographic process. This adhesive tape usually has handwritten text with the identification/composition of the samples. It has no relevant information for the image analysis process, but it interferes with the correct detection of the lanes, so this region of the image is discarded. The ROI is the internal region of the image, formed by lanes containing the result of chromatographic development and empty spaces. This is the
6 Computational and Mathematical Methods in Medicine Column (pixel) Line (pixel) 200 400 600 800 1000 100 200 300 400 500 600 700 (a) 1000 5 10 15 20 Intensity P(x) 200 400 600 800 Column (x) (b) 0 0 1 2 0 1 2 0 1 2 Maxima Minima Potential lanes 1000200 400 600 800 Column (x) 0 1000200 400 600 800 Column (x) 0 1000200 400 600 800 Column (x) (c) Column (pixel) Line (pixel) 200 400 600 800 1000 100 200 300 400 500 600 700 (d) Figure 4: (a) ROI of a TLC image. (b) Smoothed intensity profile. (c) Regional maxima (top), regional minima (middle), and set of potential lanes resulting from their combination (bottom). (d) Initial set of lanes represented by their central lines. 0 0 2 4 6 8 200 400 600 800 Intensity P(x) Column (x) (a) Column (pixel) Line (pixel) 200 400 600 800 100 200 300 400 500 (b) Figure 5: (a) Intensity profile (continuous line) and results of the first phase of lane detection (initial set represented by the dashed line); (b) the dashed line represents the removed lane. relevant region for our methodology, which is automatically delineated using a classification-based algorithm described in [28]. The ROI is converted from RGB to grayscale (GS) by retaining the luminance information. The resulting image (GS ROI) is represented using 256 levels of gray and is computed by a weighted sum of the red (R), green (G), and blue (B) components given by GS ROI =0.30R+0.59G+0.11B.(1) The GS ROI is then processed in order to remove the background, which is estimated using a closing morphological
Computational and Mathematical Methods in Medicine 7 Column (pixel) Line (pixel) 200 400 600 800 1000 100 200 300 400 500 600 700 800 900 (a) 0 2 4 6 8 10 12 14 16 0 1000200 400 600 800 Column (x) Intensity P(x) (b) Column (pixel) Line (pixel) 200 400 600 800 1000 100 200 300 400 500 600 700 800 900 (c) Column (pixel) Line (pixel) 200 400 600 800 1000 100 200 300 400 500 600 700 800 900 (d) Figure 6: (a) ROI of a TLC image; (b) Intensity profile (continuous line) and results of the first phase of lane detection (initial set represented by the dashed line); (c) Image ROI with its profile derivative overlapped; (d) Results after the three phases of the lane detection process, with the boundaries of the detected lanes represented by the vertical lines. operator with a square structuring element [29]. The size of the morphological operator is derived from the size of the image and is equal to 10% of the number of image rows, thus ensuring that the structuring element length is higher than the size of the bands. The estimated background is subtracted from the GS ROI, and the image data is projected onto the direction perpendicular to lane development (vertical projection) to integrate the information into an intensity profile that will be used for lane detection. The intensity profile 𝑃(𝑥)of an image 𝐼(𝑥,𝑦)is obtained by averaging the grey levels on each column in the image as defined by 𝑃(𝑥)=1 𝑁 𝑁 ∑ 𝑦=1𝐼(𝑥,𝑦) 𝑥=1,...,𝑀, (2) where 𝑀isthenumberofcolumnsand𝑁is the number of rows in the image. Since during background removal the image is inverted (as can be observed in Figure 1(c)), the intensity profile presents maximal regions where the original image is darker (lane zones) and minimal regions where this image is lighter (empty zones). It is worth mentioning that TLC images are often corrupted by noise on the top rows due to a significant accumulation of compounds that are present in biological materials. This noise makes the image analysis process more difficult, as it influences the profile, and smoothes the transition between regions with and without lanes. Moreover, the top region of a TLC image is not important for the remaining phases of the method, as it does not contain any relevant compounds used as biomarkers. To overcome this problem, we decided to exclude the top 25% of rows from the averaging into the intensity profile. 2.3. Profile Smoothing. The intensity profile obtained by the previously mentioned projection usually presents local variations that can lead to a high number of false lanes. Thus, a smoothed version of the profile, containing just the main
8 Computational and Mathematical Methods in Medicine 500 1000 1500 200 400 600 800 1000 500 1000 1500 2000 Column (pixel) Column (pixel) Line (pixel) 200 400 600 800 1000 Line (pixel) (a) Column (pixel) 200 400 600 800 1000 Column (pixel) 200 400 600 800 1000 200 400 600 800 1000 Line (pixel) 200 400 600 800 1000 Line (pixel) (b) Figure 7: (a) Images from DB1. (b) Images from DB2. intensity variations corresponding to the transitions between lanes zones and empty zones, is required. Moving average filters have been the most common solution applied in other methods for lane segmentation, but this kind of smoothing filter can destroy important signal information. For instance, the peaks of the intensity profile corresponding to the center of the lanes lose height when submitted to a moving average filter. The ideal filter would produce smoothed data without flattening the peaks. To overcome these problems, we propose a smoothing approach based on the CWT. The Wavelet transform offers simultaneous interpretation of the signal in both time and frequency, which allows local, transient, or intermittent components to be elucidated [30]. The wavelet transform provides a series expansion of a signal, using a set of orthonormal-based functions, which are generated by scaling and translation of two functions: the mother wavelet and the scaling function (daughter wavelet). As a result of wavelet analysis, a family of hierarchically organized decompositions is produced, where each level of the hierarchy is associated with a specific scale [27]. Although the discrete wavelet transform (DWT) is a common choice in many applications, we selected the CWT for this particular application because the highest scale resolution is provided by the continuous transform. In the DWT, scales are chosen so that the wavelets are orthogonal, which implies that the scale range will be the smallest one that will not produce loss of information [31]. The scales in the CWT are not constrained, and the wavelets are nonorthogonal. This property, while making the CWT redundant, provides a finely detailed description of a signal in terms of both time and frequency [32]. These characteristics allow an accurate selection of scales that will be important in the smoothing of the intensity profile. The CWT uses a set of wavelets, where each element is constructed from the same function, the original wavelet 𝜓(𝑡)(mother wavelet). Each daughter wavelet is a scaled and shifted version of the mother wavelet [33], according to 𝜓(𝑠,𝜏) (𝑥)=1 √𝑠𝜓(𝑥−𝜏 𝑠),(3) where 𝑠and 𝜏are the scale and translation parameters, respectively [33]. The daughter wavelets include an energy normalization term, 1/√𝑠,thatkeepstheenergyofthese wavelets equal to the energy of the original mother wavelet. For this application, the Morlet wavelet was used as the mother wavelet function, because it is known for its excellent time-frequency localization [34]. It consists of a plane wave modulated by a Gaussian function and is defined by 𝜓0(𝜂)=𝜋−1/4𝑒𝑖𝜔0𝜂𝑒−𝜂2/2,(4)
Computational and Mathematical Methods in Medicine 9 Column (pixel) Line (pixel) 200 400 600 800 100 200 300 400 500 600 700 (a) 0 5 10 15 200 400 600 800 Column (x) Intensity P(x) (b) 200 400 600 800 Column (x) 0 1 2 3 4 5 Intensity P(x) −2 −1 (c) 0 1 2 3 4 5 200 400 600 800 Column (x) Intensity P(x) −2 −1 (d) Figure 8: (a) ROI of a TLC image. (b) Intensity profile with the representation of the regions selected by the ℎ-maxima and ℎ-minima transformations. (c). Profile derivative. The width and amplitude of each lane are represented by the width and height of each rectangle, respectively. (d) The mean values for the lane width and amplitude are represented by the outside rectangle, while the values obtained after the search for subtle lanes are represented by the inner rectangle. where 𝜔0is the nondimensional frequency and 𝜂is a nondimensional “time” parameter. To satisfy the wavelet’s admissibility condition, this function must have a zero mean and be localized in both time and frequency spaces [35]. In our application, for the analysis of the profile resulting from the projection of the ROI data, the scales of interest in the CWT are related to the width of profile peaks, being the most significant ones determined by the presence of lanes. Hence, the most interesting coefficients for reconstructing a reliable smoothed version of the original profile are those whose scales correspond to the expected range of lane widths. After the analysis of the CWT decomposition of several intensity profiles, it was possible to identify a common pattern consisting of three ranges in the scale domain. At low scales, several significant coefficients are generated by high frequency noise (noise range). On the other hand, at very high scales, only the coefficients associated with the profile baseline are found (baseline range). So, the coefficients that will contain lane information are normally located in the middle range and usually present the highest amplitude values (lane range). In order to enhance lane information and at the same time achieve an adequate smoothing, the intensity profile should be reconstructed using only the scales belonging to the lane range, defined by the two cut-off values for scales: the first, herein called cutoff-min, should separate the lane range from thenoiserange,whilethesecondone,cutoff-max,setsapart thelaneandbaselineranges.Bothcut-offvaluesshouldbe situated within the lane range, which is limited by scales 30 and 250. Figure 2 showsanoverviewofthesmoothingprocess using the CWT. Figure 2(a) presents the ROI of a TLC image, whose integration into a 1D profile leads to the result depicted in Figure 2(b). The CWT coefficients for this intensity profile
16 Computational and Mathematical Methods in Medicine (1) Compute profile derivative, 𝐷 (2) Locate lane boundaries (𝑛= number of lanes in the current set) (a) Right borders,𝑅𝑖; 𝑖 = 0,⋅⋅⋅,𝑛; 𝑅0=1 (b) Left borders, 𝐿𝑖; 𝑖 = 1,⋅⋅⋅,𝑛+1; 𝐿𝑛+1 =image width (3) Estimate mlw =1 𝑛∑𝑛 𝑖=1 𝑅𝑖−𝐿𝑖 (4) Estimate mla =1 𝑛∑𝑛 𝑖=1 𝐷(𝑅𝑖)−𝐷(𝐿𝑖) (5) For 𝑖 = 1,⋅⋅⋅,𝑛+1 (a) If 𝐿𝑖−𝑅 𝑖−1 >mlw (i) Compute local maxima positions, 𝑀𝑗 (ii) Compute local minima positions, 𝑚𝑘 (iii) Create the pairs (𝑀𝑗,𝑚𝑘),𝑚𝑘>𝑀 𝑗 (iv) For each pair If 0.6 mlw <𝑚 𝑘−𝑀𝑗<1.4 mlw and 𝐷(𝑚𝑘)−𝐷(𝑀 𝑗)>0.3 𝑚𝑙𝑎 add the new lane to the set. Algorithm1:Subtlelanerecovery. Table3:Comparisonoffinalresultsobtainedusingthemethods[18]and[19] with the intermediate values after the first phase of both the proposed methodology and that described in [25] (DB1). True lanes detected True lanes missed False lanes detected Recall (𝑅) Precision (𝑃) F𝛽-measure (𝛽=1) Real class Akbari et al. [18] 480 171 140 73.7% 77.4% 75.5% Sousa et al. [19] 622 29 58 95.6% 91.4% 93.5% Moreira et al. [25] 644 7 31 98.9% 95.4% 97.1% Proposed method 647 4 28 99.3% 95.8% 97.5% lane in the initial set are represented by a rectangle whose sides are equal to those lane features. Figure 8(d) shows a lane recoveredinthelastphaseofthealgorithm,wheretheoutside rectangle has dimensions 𝑚𝑙𝑤 and 𝑚𝑙𝑎;themeanvalues obtained after averaging for the lanes present in Figure 8(c), and the inside rectangle represents the corresponding values for the recovered lane. 3.3. Comparison of Lane Segmentation Methods. In this subsection, some results of application of the proposed methodology are presented and compared with those obtained using the methods described in [18,19,25]. Figure 9 illustrates the different phases of the herein proposed method for a DB2 image, whose ROI is shown in Figure 9(a). The intensity profile of Figure 9(b) was obtained after smoothing using the CWT. The regions marked by the dashed line are the result of the first phase of lane detection and include seven true lanes correctly located, one true lane incorrectly detected because two central positions (one correct and one false) were assigned for that single lane, and one true lane missed. These results can be observed in Figure 9(c) where the vertical lines correspond to the center of the lanes detected by the algorithm. The false central line wasremovedinthesecondphase,andtheupdatedsetoflanes is shown in Figure 9(d). The third phase allows the recovery ofthemissedlane(Figure 9(e)), along with the determination of lane boundaries as depicted in Figure 9(f). The image shown in Figure 10(a),selectedfromDB1, presents an example where the methodology fails to correctly detect all the existent true lanes. At the final phase, there is still one undetected lane, on the right of the image. This lane has a band located on the top rows that are not included in the intensity profile. However, it is worth mentioning that lanes presenting bands only in the image top rows are not important for our specific application, because FD biomarkers are located in the bottom half of the image. For assessing the importance of the initial smoothing step, a comparison between the approach herein presented and the one described in [25]isshowninFigure 11.Themain difference between these two methods is the initial smoothing step. The ROI of a DB1 image and its original intensity profile are shown in Figures 11(a) and 11(b),respectively. Figures 11(c) to 11(h) present intermediate results of both the methods proposed in [25] (left) and the herein described methodology(right).TheprofilesmoothedusingtheSGfilter stillpresentshighfrequencycomponentsthatareenhanced in the profile derivative. As a consequence, the subtle lane included in the image is only segmented when the CWTbased smoothing is applied (Figure 11(h)). In Figure 12, the proposed methodology is compared with the ones described in [18,19]. In [19], the intensity information is projected onto the horizontal direction, and afterwards, a nonweighted moving average filter is iteratively applied to the profile until the number of potential lanes, which is estimated based on the number of local maxima, remains unchanged. Each lane is delimited by two consecutive maxima of the profile, while the points where the projection is minimal are associated with lane central
Computational and Mathematical Methods in Medicine 17 Table4:Comparisonoffinalresultsobtainedusingthemethods[18]and[19] with the intermediate values after the first phase of both the proposed methodology and that described in [25](DB2). True lanes detected True lanes missed False lanes detected Recall (𝑅) Precision (𝑃) F𝛽-measure (𝛽=1) Real class Akbari et al. [18] 1359 62 49 95.6% 96.5% 96.0% Sousa et al. [19] 1302 120 11 91.2% 99.2% 95.0% Moreira et al. [25] 1316 106 75 92.5% 94.6% 93.5% Proposed method 1350 72 46 94.9% 96.7% 95.8% Table 5: Results obtained for the three methodologies on DB1. True lanes detected True lanes missed False lanes detected Recall (𝑅) Precision (𝑃) F𝛽-measure (𝛽=1) Real class Akbari et al. [18] 480 171 140 73.7% 77.4% 75.5% Sousa et al. [19] 622 29 58 95.6% 91.4% 93.5% Moreira et al. [25] 647 4 20 99.4% 97.0% 98.2% Proposed method 649 2 12 99.7% 98.2% 98.9% positions. We have implemented the method in [18]following the description found in the literature and using as an input parameter the total number of occupied lanes in the chromatographic plate. In order to compare the performance of the methods, the results of the proposed methodology and the ones obtained with the methods described in [18,19] are presented in Figure 12.Figure 12(a) shows the original TLC image (which includes two empty lanes), with the respective intensity profile depicted in Figure 12(b). The set of lanes detected when using the methodology described in [18]isshownin Figure 12(c). This set includes two false detections resulting from local maxima of the CWT-based smoothed intensity profile presented in Figure 12(d).Thenumberoffalselane detections was reduced to one when the method described in [19]wasapplied,butatruelanewasmissed(Figure 12(e)). All lanes were correctly detected by the methodology proposed in this paper, as illustrated by Figure 12(f). 3.4. Comparative Analysis of Methods’ Performance. This section is devoted to the analysis of the performance of methods mentioned in the previous subsections. The four approaches were applied to all the images of the two datasets described in Section 3.1. The parameters for our method were set as detailed in Section 3.2.Forthemethodin[19], the code was provided by the author. Tables 1and 2show the results obtained after each of the three phases of the described methodology, when applied to DB1 and DB2, respectively. In order to get a more reliable evaluation of the influence of the smoothing phase, we have compared the results from the methods in [18,19] with those obtained after the first phase for both the proposed methodology and the one described in [25]. Thus, the refinement steps (the second and third phases of lane segmentation) are excluded from the results, as the improvements of these two last lane segmentation phases could give an unfair advantage to the methods that use them. These results are presented in Tables 3and 4. The final results obtained with the four methodologies are resumed in Tables 5(DB1) and 6(DB2). 4. Discussion A robust smoothing technique combined with the last two phases of the lane segmentation process achieved the best overall performance for both datasets. From the results in Tables 1and 2,itispossibletoconcludethateachofthe refinement steps addresses a different problem. In DB1, because most of the true lanes were detected during the first phase of the lane segmentation process, the improvement introduced by the lane recovery step is residual, but several false lanes were removed during the second phase. The images in DB2 contain a significant number of subtle lanes that,althoughnotincludedintheinitialsetoflanes,were recovered during the third phase. The methodology described in [18] relies on the selection of a specific scale of the CWT, followed by a detection of local extremes on the profile, reconstructed using the coefficients associated with that scale. This characteristic of the method makes it more effective when the lanes to be segmented are uniformly distributed all over the image, creating a periodicity that matches the reconstructed profile, thus justifying the good performance on the detection of true lanes even when they are very subtle. Nonetheless, this same characteristic is responsible for the high number of false detections when the images contain large empty regions corresponding to empty lanes, as happens in some images of DB1. The main drawback of the method in [19]istheattenuation of the small peaks caused by the iterative nonweighted average filtering, making the detection of subtle lanes a very hard task. Indeed, the majority of the 27 missed lanes that result from the application of this method to DB1 are associated with lanes represented by very low intensity peaks in the profile. In DB2, the detection of lanes in image extremes is also a problem, since most are not represented by a significant maximum in the intensity profile after filtering.
18 Computational and Mathematical Methods in Medicine Table 6: Results obtained for the three methodologies on DB2. True lanes detected True lanes missed False lanes detected Recall (𝑅) Precision (𝑃) F𝛽-measure (𝛽=1) Real class Akbari et al. [18] 1359 62 49 95.6% 96.5% 96.0% Sousa et al. [19] 1302 120 11 91.2% 99.2% 95.0% Moreira et al. [25] 1360 62 54 95.6% 96.2% 95.9% Proposed method 1395 27 38 98.1% 97.3% 97.7% The methodology described in [25] achieved good results with the DB1 images through the elimination of false lanes and the recovery of subtle ones, in spite of the use of identical parameter values for all the images in the dataset. However, when the methodology was applied to DB2 images, a lower performance was obtained, in the numbers of both false detections and missed lanes (Table 4). Finally, the importance of the new smoothing solution basedontheCWTisclearlydemonstratedbytheresults obtained after the first phase of the lane segmentation process (Table 5), which outperform all the values achieved by the other three methods. 5. Conclusions We have described a new methodology for lane detection in chromatography images using an innovative technique based on the CWT for decomposing the original signal, followed by the reconstruction of a new smoothed profile, using a set of coefficients in a selected range of scales adapted to each image. This technique has proven to be able to deal with the noise present in the intensity profile, while preserving the main features that are required for the subsequent three phasesofthelanesegmentationprocess.Althoughthefirst one relies on the detection of profile maxima as proposed in other solutions, the novelty of the methodology presented in this paper is the inclusion of two refinement stages, using parameters adapted to image features, to overcome some of the limitations of other approaches. The proposed methodology is fully automatic and does not require the number of lanes as an input, neither a constant spacing between lanes, nor the absence of empty lanes. The smoothing solution based on the CWT also proved to perform better than those found in [19,25], even without the improvement introduced by the refinement phases. Lane segmentation is the initial phase of the development of a screening tool for genetic disorders, and in particular Fabrydisease.However,thisisacrucialpartofthesystem,as an erroneous detection of lanes will prevent its use. As future work,weintendtodevelopadvancedmethodsfortheanalysis of lane patterns aiming at automating the identification of major disease’s biomarkers. References [1] H. Fabry, “Angiokeratoma corporis diffusum-Fabry disease: historical review from the original description to the introduction of enzyme replacement therapy,” Acta Paediatrica,vol.91, no.439,pp.3–5,2002. [2] Y. A. Zarate and R. J. Hopkin, “Fabry’s disease,” The Lancet,vol. 372,no.9647,pp.1427–1435,2008. [3] C. M. Eng, D. P. Germain, M. Banikazemi et al., “Fabry disease: guidelines for the evaluation and management of multi-organ system involvement,” Genetics in Medicine,vol.8,no.9,pp.539– 548, 2006. [4] R. Schiffmann, J. B. Kopp, H. A. Austin III et al., “Enzyme replacement therapy in fabry disease a randomized controlled trial,” JournaloftheAmericanMedicalAssociation,vol.285,no. 21, pp. 2743–2749, 2001. [5] J.M.Miller,Chromatography: Concepts and Contrasts,chapter 11, John Wiley & Sons, Hoboken, NJ, USA, 2005. [6] B. Fried and J. Sherma, Thin-Layer Chromatography, Part I, Chapter 1, Marcel Dekker, New York, NY, USA, 2005. [7] P. A. Sewell, “Liquid chromatography: theory of liquid chromatography,” in Handbook of Methods and Instrumentation in Separation Science,I.D.WilsonandC.Poole,Eds.,pp.591–605, Elsevier, London, UK, 2004. [8] J. Sherma, “Basic TLC techniques, materials, and apparatusin,” in Handbook of Thin-Layer Chromatography, J. Sherma and B. Fried, Eds., pp. 1–62, Marcel Dekker, New York, NY, USA, 2003. [9] J. Pizzonia, “Electrophoresis Gel Image processing and analysis usingtheKODAK1Dsoftware,”BioTechniques,vol.30,no.6, pp.1316–1320,2001. [10]T.ZerrandS.Henikoff,“Automatedbandmappinginelectrophoretic gel images using background information,” Nucleic Acids Research,vol.33,no.9,pp.2806–2812,2005. [11] M. L. Cooper, D. R. Maffitt, J. D. Parsons, L. Hillier, and D. J. States, “Lane tracking software for four-color fluorescencebased electrophoretic gel images,” Genome Research,vol.6,no. 11, pp. 1110–1117, 1996. [12] A. B. Pavel and C. I. Vasile, “PyElph-a software tool for gel images analysis and phylogenetics,” BMC Bioinformatics,vol.13, no. 91, p. 9, 2012. [13]R.T.F.Wong,S.Flibotte,R.Corbettetal.,“LaneRuler: automated lane tracking for dna electrophoresis gel images,” IEEE Transactions on Automation Science and Engineering,vol. 7,no.3,pp.706–708,2010. [14] A. M. C. Machado, M. F. M. Campos, A. M. Siqueira, and O. S. F. De Carvalho, “Iterative algorithm for segmenting lanes in gel electrophoresis images,” in Proceedings of the 1997 10th Brazilian Symposium of Computer Graphic and Image Processing, SIBGRAPI 97, pp. 140–146, October 1997. [15]J.K.ElderandE.M.Southern,“Computer-aidedanalysisof one dimensional restriction fragment gels,” in Nucleic Acid and Protein Sequence Analysis-A Practical Approach,M.J.Bishop and C. J. Rawlings, Eds., pp. 164–172, IRL Press, Oxford, UK, 1987. [16] I. Bajla, I. Holl¨ ander, S. Fluch, K. Burg, and M. Koll´ ar, “An alternative method for electrophoretic gel image analysis in the GelMaster software,” Computer Methods and Programs in Biomedicine,vol.77,no.3,pp.209–231,2005.
Computational and Mathematical Methods in Medicine 19 [17] C. Lin, Y. Ching, and Y. Yang, “Automatic method to compare the lanes in gel electrophoresis images,” IEEE Transactions on Information Technology in Biomedicine,vol.11,no.2,pp.179– 189, 2007. [18] A. Akbari, F. Albregtsen, and K. S. Jakobsen, “Automatic lane detection and separation in one dimensional gel images using continuous wavelet transform,” Analytical Methods,vol.2,no. 9, pp. 1360–1371, 2010. [19] A. V. Sousa, R. Aguiar, A. M. Mendonc¸a, and A. C. Campilho, “Automatic lane and band detection in images of thin layer chromatography,” in Image Analysis and Recognition, LNCS, 3121, A. C. Campilho and M. S. Kamel, Eds., pp. 158–165, Springer, Berlin, Germany, 2004. [20] C. Maramis and A. Delopoulos, “Efficient quantitative information extraction from PCR-RFLP gel electrophoresis images,” in 2010 20th International Conference on Pattern Recognition, ICPR 2010, pp. 2560–2563, tur, August 2010. [21] C. Maramis, D. Karagiannis, and A. Delopoulos, “HPVTyper: a software application for automatic HPV typing via PCR-RFLP gel electrophoresis,” in Human Papillomavirus and Related Diseases-from Bench to Bedside-Research Aspects,D.V.Broeck, Ed., InTech, New York, NY, USA, 2012. [22] S. C. Park, I. S. Na, S. H. Kim et al., “Lanes detection in PCR gel electrophoresis images,” in 11th IEEE International Conference on Computer and Information Technology, CIT 2011 and 11th IEEE International Conference on Scalable Computing and Communications, SCALCOM 2011, pp. 306–313, cyp, September 2011. [23] M. Sotaquir´ a, “On the use of distance maps in the analysis of 1D DNA gel images,” in 2009 International Conference on Digital Image Processing, pp. 172–176, tha, March 2009. [24] P. Barrantes and P. Alvarado, “Lane detection on gel electrophoresis images using active shape models,” in Proceedings of the Conference on Technologies for Sustainable Development (TSD ’11), pp. 43–46, 2011. [25] B. M. Moreira, A. V. Sousa, A. M. Mendonc¸a, and A. C. Campilho, “Automatic lane detection in chromatography images,” in Image Analysis and Recognition, LNCS 7325,A.C. Campilho and M. S. Kamel, Eds., pp. 180–187, Springer, Berlin, Germany, 2012. [26] A. Savitzky and M. J. E. Golay, “Smoothing and differentiation of data by simplified least squares procedures,” Analytical Chemistry,vol.36,no.8,pp.1627–1639,1964. [27]F.Chau,Y.Liang,J.Gao,andX.Shao,Chemometrics-from Basics to Wavelet Transform, Chapter 4,JohnWiley&Sons,New Jersey, NY, USA, 2004. [28] A. V. Sousa, M. C. S´ a-Miranda, A. M. Mendonc¸a, and A. C. Campilho, “Classification-based segmentation of the region of interest in chromatographic images,” in Image Analysis and Recognition, LNCS 6754,A.C.CampilhoandM.S.Kamel,Eds., pp. 68–78, Springer, Berlin, Germany, 2011. [29] M. H. Sedaaghi and Q. H. WU, “The power of morphological filters alone and when combined with linear filtering,” in Mathematical Morphology and Its Applications to Image and Signal Processing,H.J.A.M.HeijmansandJ.B.T.M.Roerdink,Eds., pp. 375–382, Kluwer Academic, Dordrecht, The Netherlands, 1998. [30] P. S. Addison, “Wavelet transforms and the ECG: a review,” Physiological Measurement, vol. 26, no. 5, pp. R155–R199, 2005. [31] J. C. van den Berg, Wavelets in Physics, Chapter 10, Cambridge University Press, Cambridge, UK, 2004. [32] R. K. Young, Wavelet Theory and Its Applications, Chapter 1, Kluwer Academic Publishers, Norwell, Mass, USA, 1993. [33] S. Mallat, A Wavelet Tour of Signal Processing, Chapter 1, Elsevier, Burlington, Mass, USA, 2009. [34] I. Daubechies, “Wavelet transform, time-frequency localization and signal analysis,” IEEE Transactions on Information Theory, vol.36,no.5,pp.961–1005,1990. [35] A.Grinsted,J.C.Moore,andS.Jevrejeva,“Applicationofthe cross wavelet transform and wavelet coherence to geophysical times series,” Nonlinear Processes in Geophysics,vol.11,no.5-6, pp.561–566,2004. [36] P. Soille, Morphological Image Analysis-Principles and Applications, Chapter 6, Springer, New York, NY, USA, 2004. [37] A. V. Sousa, A. Mendonc¸a, and A. Campilho, “Chromatographic pattern classification,” IEEE Transactions on Biomedical Engineering,vol.55,no.6,pp.1687–1696,2008.