Automatic Detection of Facial Midline And Its Contributions To Facial Feature Extraction
Abstract
We propose a novel approach for detection of the facial midline from a frontal face image. Using midline as a guide reduces computational cost required for facial feature extraction (FFE) because the midline is capable of restricting multi-dimensional searching process into one-dimensional search. The proposed method detects the facial midline from an edge image as the symmetry axis using the generalized Hough transformation. Experimental results on the FERET database indicate that the proposed algorithm can accurately detect facial midlines over many different scales and rotation. The total computational time for facial feature extraction has been reduced by a factor of 280 using the midline detected by this method.
Full text
Electronic Letters on Computer Vision and Image Analysis 6(3):55-66, 2008 Automatic Detection of Facial Midline And Its Contributions To Facial Feature Extraction Nozomi NAKAO, Wataru OHYAMA, Tetsushi WAKABAYASHI and Fumitaka KIMURA Graduate School of Engineering, Mie University, 1577 Kurimamachiya–cho, Tsu–shi, Mie 514–8507, Japan Received 17 April 2007; revised 17 June 2007; accepted 17 September 2007 Abstract We propose a novel approach for detection of the facial midline from a frontal face image. Using midline as a guide reduces computational cost required for facial feature extraction (FFE) because the midline is capable of restricting multi-dimensional searching process into one-dimensional search. The proposed method detects the facial midline from an edge image as the symmetry axis using the generalized Hough transformation. Experimental results on the FERET database indicate that the proposed algorithm can accurately detect facial midlines over many different scales and rotation. The total computational time for facial feature extraction has been reduced by a factor of 280 using the midline detected by this method. Keyword: Facial midline, Facial feature detection, Generalized Hough transformation, Biometrics 1 Introduction Biometrics employing a fully automatic face recognition or authentication technologies requires both face detection and recognition[1]. In the face detection problem, we are given an input image that may contain one or more human faces (or it may contain no face at all). The scale of the face is not known in advance. For example, in a 512×768 input image, the face may appear in a small region 64×64 size, or it might occupy the entire range 512 ×768 pixels. The problem is to segment the input image and isolate the face(s). Particularly, it is necessary to determine a tight bounding box around each face that contains just the face (forehead to chin), excluding as much of the hair as possible. Of course, the results of the recognition task[2] depend heavily on how well the detection task has been done. For example, when the bounding boxes are not tight enough, Chen et.al[3] showed that non-face artifacts tend to dominate and hence corrupt the feature extraction process needed for recognition. For a human face, there are important features or landmarks that one can exploit for detection purposes. If the position of these facial features is known, then face detection and localization can be done easily and more accurately. The detection of facial features, though, is computationally expensive; hence it makes sense to apply the detection only in the vicinity of a face and not the entire image (which may contain many non-face artifacts). Even for frontal face images can be observed as the most simple situation in face recognition, there are many parameters to estimate, for instance location of each feature, scale and rotation of faces. If we get any guides that can be utilized for facial feature extraction by a method that is easier than that for facial features, it is possible to reduce total computational costs. Correspondence to: Wataru OHYAMA <[email protected]> Recommended for acceptance by Umapada Pal and P. Nagabhushan ELCVIA ISSN:1577-5097 Published by Computer Vision Center / Universitat Aut`onoma de Barcelona, Barcelona, Spain
The facial midline, in other words the facial symmetry axis, is one of promising candidates for such guides to reduce the computational cost. The extraction of facial midline is equivalent to the detection of facial slant angle and localization of the center point between each eye; hence, the extracted midline can be utilized to normalize the slant and location of the face. This reduces the complexity of facial feature extraction. Also the facial midline has additional contributions for face recognition. For example, Quintiliano et al. [4] reported that symmetrization of face image could improve the performance of face recognition. Symmetrization, which means reconstructing the dark side of the face from the clear side in this case, utilized the facial midline. In this paper, we propose a facial midline detector based on generalized Hough transformation (GHT). This method detects the facial midline from a grayscale image where one frontal face is. Since faces are often slanted in image, the detection method must be robust for these varieties. We present an automatic detection technique of the facial midline and evaluate the performance of the proposed method by experiments with facial images from the FERET database[5]. The proposed method detects the facial midline based on matching a binary edge image of input face and its mirror image. GHT is used for the matching. For binary images, GHT behaves an equivalent algorithm as the correlation method. However it has advantages on computational cost and noise tolerance. In this paper, we also proposed a fast algorithm of GHT for symmetry detection of faces. In contrary to our method X.Chen et al.[6] have proposed an automatic methodology for the facial midline detection. In their method, axes of facial symmetry are detected as those which maximize the Y value that is based on the gray level differences (GLD) between the both sides of the axis. Their approach has the following twofold drawbacks. (1) the Y value is quite sensitive to change of lighting conditions: if faces are illuminated from left or right sides, GLD is easily influenced. (2) It is computationally expensive because the maximization problem for the Y value is solved by a sweeping algorithm: in other words, to find a axis which maximizes the Y value, we have to evaluate all combinations of rotation and position of candidates. Other method has been proposed by Hiremath and Danti[7]. In this method, a face is explained by the Lines–of–Separability (LS) face model which includes the facial symmetry axis. To obtain this LS model, we have to extract both eyes from frontal face image before detect the symmetry axis. From this point of view, this method is observed as a bottom–up approach which is opposite to our method. Another group of methods employ models of facial shape and appearance like Active Appearance Model [8]. These methods utilize pre-trained model or template and fit them to input face image. Our proposed method does not require such preliminary training. The remainder of this paper is organized as follows. In Section 2, we present the proposed methodology used for face midline detection. Section 3 gives experimental results and contributions by this method for facial feature extraction is given in Section 4. Section 5 gives discussions. 2 Proposed Methodology In this section, we present the proposed methodology for facial midline detection. Our method is based on bilateral symmetry of human face and extracts the symmetry axis as the facial midline. To extract the axis reliably, we employ the generalized Hough transformation (GHT)[9][10] that is able to extract non-analytical curves from an image. 2.1 What is the facial midline? We define the facial midline as the perpendicular bisector of the interocular line segment (connecting each eyes). As exampled in Figure 1, when the face in an input image is slanting with the angle θ, the midline should be detected having the same slanting angle. Detecting a line on an image is an equivalent problem of detecting one point at which the line passes and the angle of the line. In Figure 1, the line passing through the point c= (cx, cy)and the angle θis expressed as x−cx sin θ=y−cy cos θ.(1) 56
Figure 1: Example of facial midline as the symmetry axis We can determine these two parameters, cand θ, from a pair of points between which symmetry axis line passes. When two points, p= (px, py)and q= (qx, qy), are symmetrical to each other such that a point con the axis can be expressed as c=p+q 2. And the angle θis obtained as that is orthogonal to the angle of (q−p). Consequently, we can rewrite expression (1) using pand qas follows. x−px+qx 2 sin θ= y−py+qy 2 cos θ,(2) θ= tan−1qy−py qx−px.(3) The problem to solve is to extract this pair of symmetrical points, which are given as examples by pand qin Figure 1. We employ the assumption where a frontal face is globally symmetrical. However, the symmetry of faces is easily corrupted when faces are illuminated form left or right sides. In this case, to reduce the influences by illuminations, we have to combine preprocesses in our method. 2.2 Overview of the methodology The proposed method consists of three main stages, as shown in Figure 2. In the first stage, we apply preprocessing that consists of edge detection, thresholding and noise removal. The input of the proposed method, which is demonstrated by (a), is a grayscale image containing one human face in unoccluded frontal view. The size of image is 512×768 pixels. And the face is nonrigid and has a high degree of variability in scale, location, and slant. The resultant image after the preprocessing contains strong edge components of which lengths are sufficient for the GHT. An example of resultant preprocessed image is shown in (b). The second stage of this method is the GHT. The GHT requires a proper reference point for reasonable execution. The reference point is illustrated by pin (b). The GHT extracts the point that is symmetric to the reference point. The resultant point is called the relevant point in this research, which is denoted by qin (c). In the third stage, using the detected coordinates of two symmetric points pand q, we obtain the facial midline by (2). Brief descriptions of the each process are presented in the following subsections. 2.3 Preprocessing The preprocessing in the proposed method generates a binary edge image from input images. Since the GHT algorithm we employ in the second stage is applicable only to a binary image, it is important to obtain proper 57
Figure 2: Three main stages of the proposed facial midline detection binary images for sufficient results. For instance, the binary image that includes too much noise increases the computational cost of the GHT and easily corrupts the results mean while the image with too little edge components makes the reliability of the GHT significantly weak. At first, edge magnitude of an input image is calculated by using the Sobel operator. Edge image is binarized by p-tile thresholding. In this method, a threshold Tis selected as such that p% of the image area has gray values (i.e. edge magnitude) less than Tand the rest has gray values larger than T. Because the Sobel operator enhances noise in theoriginal image, the resultant binary image might contain some noise if wecould determine the best threshold. To remove the noise, we eliminate edge elements whose length is smaller than Llpixels or grater than Lupixels. Here, Lland Luare two threshold values. These values are estimated form the experiment and it is discussed in Section 2.6. The length of edge components can be obtained by 8–connective boundary following algorithm. After the boundary following, each edge component is represented by the contour code. 2.4 Generalized Hough Transformation The generalized Hough transformation (GHT)is an algorithm to detect objects, which have the same (or similar) shape as a given template, from given binary images. It is empirically known that GHT is robust to both noise and lack of objects in images. For binary images, GHT behaves as a fast algorithm of template matching. To adapt a template to objects in an input image involving variety of poses: scale, position and rotation, we have to transform and apply the template on the input image by repetition; hence the computation cost becomes high. To reduce this computational time, GHT employs voting strategy in a parameter space whose dimensionality is equivalent to the variety of poses. The GHT in this research is aimed at finding the relevant point that is symmetry to the reference point. The assumption of facial bilateral symmetry suggests that the edge image might also be symmetric. So we employ the mirror image of the binary edge image obtained by preprocess as a template. This means that GHT detects the most similar shape object to the mirror image from the binary edge image. When GHT detects the object, we can easily detect the pair of symmetry points. The tasks of GHT in the proposed method are as follows: (1) Selection of the reference point: For GHT, we should select a reference point in an image. The selection of the reference point is arbitrary but very important for reasonable execution because it influences the 58
Figure 3: The generalized Hough transformation in the proposed method performance of the following GHT steps. Sato and Ogawa [11] have observed that using the center of gravity (CG) of edge pixels as the reference point contributes to the most reliable results by GHT. To verify this contribution of CG to the result of GHT, we performed a pilot experiment using 400 facial images same as in Section 2.6. In this pilot experiment, the reference point were switched among several candidates including CG and tested using the performance of midline detection (described in Section 3). The results of the experiments suggest that CG provides the most accurate midline detection. Therefore, we use CG of all black pixels (edge pixels) in the binary edge image as the reference point. CG pis obtained by, p= N P j=0 ej N= N P j=0 ejx N, N P j=0 ejy N!,(4) where, ej= (ejx, ejy)and Ndenote an edge pixel and the total number of edge pixels in the binary edge image, respectively. An example of the reference point is shown as pin Figure 3 (a). (2) Generation of the template image: As described above, we use the mirror image of the binary edge image corresponding to the vertical axis as a template. When the edge image is symmetric corresponding to the vertical axis, the original image and the template might be overlap considerably at the relevant point (Figure 3(b)). (3) Voting in the parameter space: The GHT’s parameter space in this method becomes three dimensional, i.e. qx,qyand rotation θ. They correspond to the object’s variety of poses. Figure 3(c) illustrates the voting process in this method. The sweeping template, which is point symmetric image of the template (b), scans each of all edge pixels in the binary edge image. During sweeping, the corresponding point in the parameter space accumulates the vote from the template image. (4) Detection of the relevant point: The location and the angle of rotation of the template are detected from the point in the parameter space, where the maximum voting value is obtained (Figure 3(d)). 59
Figure 4: Basic idea of fast midline detection 2.5 Fast Algorithm of GHT The above tasks provide the proper information to adapt the template to the binary edge image, though, computational time for these tasks, especially for voting in the three-dimensional parameter space, is not negligible. To reduce this cost, we introduce the following restriction for the parameter space. When a human face displays symmetric corresponding to the vertical axis, in other words the face is straight in image; the vertical position qyof the template (mirror image) is exactly same as that of the original binary edge image. If both the reference and edge images were rotated with the same angle to the opposite direction each other, the change of qybetween the template and the original image is eliminated. This means that the dimensionalities of the parameter space are restricted to two, qxand θ. Figure 3 and Figure 4 illustrates the basic concepts of this method. In this method, the range of facial slanting angle is assumed as θ∈[−15◦,15◦]. Computational time for GHT is significantly reduced by this fast algorithm. In our pilot study, the time for one GHT operation is reduced from 10[s] to 0.15[s] on 2.6 GHz Intel Core2 processor. 2.6 Parameter Settings The proposed method requires some preliminary defined parameters: pfor the p-tile thresholding and Lland Lufor the noise reduction. To determine these parameters, we performed the following preliminary experiment with a data set consisting 400 frontal face images selected randomly from the FERET database. The GHT– based facial midline detection described above is applied on all 400 facial images with each combination of the following parameter settings, p∈ {5,6,7,8,··· ,20}[%],(5) Ll∈ {5,10,15,··· ,30}[pixels],(6) Lu∈ {500,600,700,··· ,1500}[pixels].(7) Each combination of parameters is evaluated for the shift and angle errors (described in the next section). The combination that yields the highest performance: (p, Ll, Lu) = (11,20,1400) is selected. 3 Experiments To verify the effectiveness of the proposed method, we apply the proposed method to the images from FERET database. Some examples of detected midline are shown in Figure 5. The white line in each picture is the detected midline. The face midline over many different scales and rotation has detected correctly. Figure 6 shows examples where the conventional GLD–based method did not yield accurate midline due to the lighting asymmetry on the face; and in contrast, the proposed GHT–based method extracted midline 60
Figure 5: Result of the midline detection Figure 6: visual comparison of extracted midlines accurately. White and yellow lines denote the extracted midlines by the GLD–based method and the GHT– based one, respectively. In these examples, lighting condition on each side of a face is different. The GLD– based method is too sensitive for such difference of lighting, and extracted an axis of local symmetry instead of that of global symmetry. In contrast to this, the proposed GHT–based method could extract ideal midline as an axis of global bilateral symmetry even if edge components rise in the dark side of a face. Next, we quantify the performance of the proposed method by evaluation experiment with 2409 frontal face images from the fa and fb probes in FERET database. For this test, we compare the detected midline with the reference midline obtained from ground-truth eye locations. As used in [6], two measurements, angle error ∆θand distance error s, are used to evaluate the performance of midline detection. The angle error ∆θis the difference between detected and reference midlines. The distance error sis the distance between theses two midlines on the interocular line segment. Figure 7 illustrates these measures. To demonstrate the advantage of our proposed method, we compare the proposed method and the conventional one, which has been proposed in [6], for the angle and the distance errors. Figure 8 shows cumulative histograms of the angle error and distance (shift) error of the detected midlines for the 2409 images in FERET by Chen’s conventional GLD–based method and our proposed method. The conventional GLD–based method 61
Figure 7: angle error and distance error for evaluation Figure 8: Performance evaluation by cumulative histograms was implemented to work on the same condition as our proposed GHT–based method. Both methods were tested on the same data set. 93.48% of the detected midlines are within 5 degrees angle error; this means that the rotations of face in 93.48% of input images are correctly estimated by the proposed method. And 84.52% detected distance error are within 10 pixels; this means that the positions of midline in 84.52% of input images are detected correctly. These results suggest that the performance of our method is superior than that in [6]. This result suggests that the proposed method provides acceptable performance for the midline extraction. The computational time of the proposed method for the 2409 facial images is 369.12[s] by a 2.66 GHz Intel Core2 CPU. The frame rate is 6.53 [frames/s]. We disucussed here some detailed investigation about the angle and distance errors. An input frontal face is expected to slant on the image plane. The midline detection algorithms are required to detect these slanting angle of face; hence, it is important to invest the tendency of the errors with the variety of facial slanting angle. The FERET database has variety in facial slanting angle on its consisting images. We grouped the images that have same slanting angles and calculated the mean of each the angle and distance errors on each group. Figure9 shows the results of evaluation the conventional GLD and the proposed GHT for the relationship between the mean errors and facial slanting angle calculated from the given ground-truths. From the results, the followings are observed: 1. For the most of slanting angles, the proposed GHT yielded smaller mean error than that by the conventional GLD. 2. The distance error of the conventional GLD increased with large slanting angle. In contract, the distance 62
Figure 9: The tendency of mean angle error and mean distance error with the facial slanting Figure 10: Performance comparison between the proposed and the conventional methods. error of the proposed GHT is nearly-constant. Figure 10 (a) to (c) illustrate the comparison of these errors. (a) and (b) illustrate the number of images where the angle and the distance errors are less than 5 degrees and 10 pixels respectively; in other words, the midline is detected correctly. (c) illustrates the number of images where both the conditions of angle and shift are satisfied. These results indicate that the accuracy of midline detection is significantly improved by the proposed method. We also compare these methods for the computational time. Figure 10 (d) indicates that the computational time is reduced from 18.1[s] to 0.15[s] for one input image. 4 Contributions of midline detection to facial feature extraction The contribution of the proposed method to the facial feature extraction (FFE) is considerably significant. Here, we discuss the advantage of the detected midline in FFE. The use of a midline as a guide for feature extraction reduces the computational time required for FFE. In FFE, an algorithm must estimate many parameters, which describe the face, i.e. scale, rotation and position. Midlines which are estimated properly eliminate these estimation tasks for rotation and reduces the range of position variety. 63