scieee AI-readable full text Open interactive document viewer

3D Model Based Pose Invariant Face Recognition from a Single Frontal View

Chen, Qinran; Cham, Wai-kuen

Abstract

This paper proposes a 3D model based pose invariant face recognition method that can recognize a face of a large rotation angle from its single nearly frontal view. The proposed method achieves the goal by using an analytic-to-holistic approach and a novel algorithm for estimation of ear points. Firstly, the proposed method achieves facial feature detection, in which an edge map based algorithm is developed to detect the ear points. Based on the detected facial feature points 3D face models are computed and used to achieve pose estimation. Then we reconstruct the facial feature points' locations and synthesize facial feature templates in frontal view using computed face models and estimated poses. Finally, the proposed method achieves face recognition by corresponding template matching and corresponding geometric feature matching. Experimental results show that the proposed face recognition method is robust for pose variations including both seesaw rotations and sidespin rotations.

Full text

Electronic Letters on Computer Vision and Image Analysis 6(1):13-26, 2007 3D Model Based Pose Invariant Face Recognition from a Single Frontal View Qinran Chen and Wai-kuen Cham Department of Electronic Engineering, The Chinese University of Hong Kong Shatin, New Territories, Hong Kong Received 2 August 2005 ; accepted 16 March 2007 Abstract This paper proposes a 3D model based pose invariant face recognition method that can recognize a face of a large rotation angle from its single nearly frontal view. The proposed method achieves the goal by using an analytic-to-holistic approach and a novel algorithm for estimation of ear points. Firstly, the proposed method achieves facial feature detection, in which an edge map based algorithm is developed to detect the ear points. Based on the detected facial feature points 3D face models are computed and used to achieve pose estimation. Then we reconstruct the facial feature points’ locations and synthesize facial feature templates in frontal view using computed face models and estimated poses. Finally, the proposed method achieves face recognition by corresponding template matching and corresponding geometric feature matching. Experimental results show that the proposed face recognition method is robust for pose variations including both seesaw rotations and sidespin rotations. Key Words: Face Recognition, Pose Estimation, 3D face Model, Single View. 1 Introduction During the last few decades, research on the automatic face recognition (AFR) has received increasing attention, and different face recognition algorithms have been developed. [1-3] give some good reviews in this field. Most AFR algorithms are for face recognition under controlled conditions. For example, satisfactory recognition rates on face images which are uncovered, in frontal view, with neutral expression and controlled lighting have been reported in [7-10]. While some other recognition algorithms such as [1315] have been developed to tackle the variations on different lighting, small occlusions, and facial expressions for frontal view face images. The results are encouraging. The problem related to variations in poses received much attention and many algorithms have been developed to tackle this problem. An early attempt is the 2D appearance based approach which describes faces under varying pose with a set of 2D features and achieves pose analysis and face recognition by comparing these features. [18] presents a method for pose invariant face recognition in the entire eigenspace. Huang et. al [16] achieved pose invariant face recognition in the view-space which is a subspace of the eigen-space. Demir’s method [19] is similar to that of [16], but employing a sub-LDA space as the viewspace. In [11] and [12], this problem was tackled in the discriminant waveletface space and the kernel LDA space respectively. [20] describes a line-based algorithm for pose invariant face recognition. These Correspondence to: [email protected] Recommended for acceptance by E. Martí ELCVIA ISSN: 1577-5097 Published by Computer Vision Center / Universitat Autonoma de Barcelona, Barcelona, Spain 2 Qinran Chen et al. / Electronic Letters on Computer Vision and Image Analysis 6(1):13-26, 2007 appearance based methods can provide good recognition results based on dense sampling of the continuous pose in gallery. However this requirement not only increases the gallery size but also makes the recognition process more time consuming. In addition, when the gallery consists of only one front view image (such as a passport photo) per candidate, these methods cannot work. Therefore the 3D model based approach was proposed. It has stronger generalization to pose variation and is available to achieve pose invariant face recognition from a single frontal view, though its implementation is more complex. In the 3D model based approach, a 3D face model is built to represent the 3D geometry of human faces in 2D images. This approach removes the effect of pose variations on face recognition by estimating and aligning poses with a 3D face model and then extracting features under a uniform pose for classification. Generally, pose estimation is the most critical and challenging operation in the 3D model based approach. In [21-22], fixed generic 3D face models were proposed to be used for all candidates. These methods can achieve pose estimation from a single face image based on affine transforms. However pose of a particular face cannot be estimated accurately by using a fixed 3D model. In [6, 17], simple adaptive 3D face models that can adapted to fit a particular person were proposed. The pose estimation and the model adaptation were achieved synchronously by using geometrical measurements. These methods can obtain effective pose estimation from a single face image. However, [6, 17] can only estimate the sidespin rotations of the face in an image with the assumption that the face has no seesaw rotation. Recently, Blanz et al [23-24] built a 3D morphable face model from a large set of real 3D face data for pose invariant face recognition. Based on this model, the pose estimation and model adaptation were achieved by hybrid geometric information and texture information based optimization. The reported performance of pose estimation in this system is good, but the optimization procedure is very complex and requires large computing time. This paper proposes a model based pose invariant face recognition method to recognize a face from its single nearly frontal view. As a generalization of Lam and Yan’s method [6], our method obtain more robust performance to pose variation and gives following contributions: 1) proposed a edge map based ear point detection algorithm, 2) presents a more general and powerful pose estimation algorithm 3) achieve classification by corresponding template matching and corresponding geometric feature matching. 2 Overview of the Proposed Method In this paper we propose a 3D model based pose invariant face recognition method that can recognize a face from its single nearly frontal view, which assumes that the face has no seesaw rotation and may have small sidespin rotation. The proposed method is composed of four operations: (1) facial feature detection, (2) pose estimation and 3D model adaptation, (3) pose invariant feature extraction, and (4) classification. The block diagram of the proposed method is given in Fig. 1. In the operation of feature detection, beside eye corners, mouth corners, nose tip, eyebrow points, face contour, new facial features in the form of two lower joint points of the ears and the face boundary (called ear points in the following) are detected by an edge map based algorithm. A simple adaptive 3D face model is used to represent the 3D geometry of the face in an image. With the assumption that the gallery faces have no seesaw rotation, we achieve pose estimation and Fig. 1. Overall method architecture of the face recognition. Test image Poses of Faces in all Images M 3D face models Gallery Feature Database Pose Estimation and Model Adaptation Recognition Result M Potential Pose and M Potential Models Gallery Images of M persons Facial Feature Detectio n Pose Estimation and Model Adaptation Pose Invariant Feature Extraction Reconstructed Facial Feature Locations and Templates of Frontal Views Facial Feature Detection Pose Invariant Feature Extraction Classification by Template Matching and Geometric Features Matching Reconstructed Facial Feature Locations and Templates of M Potential Frontal Views Facial Features on all images Facial Features on The Test Image Qinran Chen et al. / Electronic Letters on Computer Vision and Image Analysis 6(1):13-26, 2007 3 model adaptation directly for these images using Lam and Yan’s method and obtain a pose and a face model for each gallery face. Based on each gallery face image and the corresponding estimated face model, an efficient algorithm is proposed to estimate a pose and a face model for a test face. Thus the test face totally has M potential poses and potential face models, where M is the number of the candidates in the gallery. In the pose invariant feature extraction operation, we compute the facial feature points’ locations and the facial feature templates of frontal views and potential ones from the gallery face images and the test image respectively based on the estimated poses and models. Finally, the proposed method achieves classification by comparing the obtained template and geometric features from the gallery images and the test image. 3 Facial Feature Extraction Locating facial features is an important step in face recognition. In the proposed method, the rough face contour, two outside eye corners and , two inside eye corners -, two mouth corners -, a nose tip and tow eyebrow points - (see Fig. 1(a)) are located by using Lam and Yan’s method [6]. The ear points as shown in Fig. 6 are not used in most face recognition algorithms. They are important features for estimating the seesaw rotation in our algorithm. An edge map based algorithm is proposed to detect the ear points in this paper. In the following, we illustrate how the algorithm detects left ear point. (0) p(3) p(1) p(2) p(4) p(5) p (8) p(6) p(7) p First of all, a 2D rotation transform is performed on the input face image I to produce an upright face image Iu, in which the line that holds least square distances to four eye corners is parallel to horizontal axis (see Fig. 2(b)). The facial feature points p in I () ,0,..., j uj=8 8 u correspond to p in I, as shown in Fig. 2(b). Then the modified canny edge detector introduced in [4] is employed to obtain the edge map E of (for example, Fig. 3(a)). From the detected facial feature points and the rough face contour in I () ,0,..., jj= u, a searching region for ear points is determined in E as shown in Fig. 3 (b). The trivial edges in the searching region are eliminated. Canny edge detection may produce disconnected edges which correspond to the continuous contours in Iu. We develop a new edge connection operation to recover such continuity in the edge map E. Thus the connected edges are obtained as shown in Fig. 3 (c). (b) Fig. 2 Facial features in: (a) the input image I, (b) the upright image Iu. (6) p(7) p (5) p (4) p (8) p (0) p(1) p(2) p(3) p (6) p(7) p (5) p (4) p (8) p (0) p(1) p(2) p(3) p (a) (a) (b) (d) Fig. 3 (a) an upright face image. (b) The edge map and the searching window. (b) Connected edge map. (c) The potential face and ear boundaries. y x O a c b d e Γ (c) By filtering out edges that have a nonnegative slope and retaining only the largest connected edges, we obtain the potential face and ear boundaries Γ as shown in Fig. 3 (d). It is assumed that the outermost curve in Γ, says l (i.e. ad in Fig. 3 (d)), should include both ear boundary and face boundary. We estimate the salient points on by R/J curvature based curve partition algorithm [5]. The salient points, which are inward l 4 Qinran Chen et al. / Electronic Letters on Computer Vision and Image Analysis 6(1):13-26, 2007 bending points (illustrated in Fig. 4(a)) satisfying the condition that R/J curvature at these points are larger than a predefined threshold value, will be chosen as possible ear points. In order to exclude the false candidates such as the neck point (the joint point of the face boundary and the neck boundary), anthropocentric constraints are used to verify each possible ear point. The anthropocentric constraints are formed based on some prior knowledge and statistics obtained from some 200 face images. Let denote the vertical distance between l’s top end point ts d ut p and a possible ear point us p , and dlem denote the vertical distance between and , see Fig. 4(b). If d (0) u p(4) u pts > dlem or us p is on the right side of , (0) u pus p will be rejected as the left ear point. If no candidate is viable to pass the verification, it means that no left ear point is detected; otherwise the viable candidate, which holds the largest R/J curvature will be regarded as the detected left ear point. us p ut p (0) u p ts d l (4) u p lem d ∗ ∗ ∗ ∗ Outward bending point Outward bending point Inward bending point Inward bending point l (b) (a) Fig. 4. (a)Bending points on the left face and ear boundary l(b) The elements related to the anthropocentric constraints. Similarly, we can detect the right ear point. In Iu, the left ear point and the right ear point are labelled as and respectively (see Fig. 2(b)). Some examples of the detected ear points are shown in Fig. 5. (9) u p(10) u p Fig. 5 The examples of ear points detected by our method. 4 Pose Estimation and 3D Model Adaptation In this paper, a face image is regarded as a 2D orthogonal projection of a 3D face. While a 3D face is originally posed in the world coordinate system as shown in Fig. 6(a), its projection on the image plane will be a face image in front view. The image plane is always perpendicular to the z-axis of the word coordinate system. When the 3D face has a certain rotation around the origin of the world coordinate, its 2D projection on image plane is a face image with corresponding pose. Any rotation can be uniquely decomposed into three orderly rotations--seesaw rotation, sidespin rotation, and in image plane rotation, which are around xaxis, y-axis and z-axis by x θ , y θ and z θ respectively. In this paper, all gallery images and test images are adjusted to upright face images by 2D rotation operation mentioned in section 2. In addition, it is assumed that the face in each gallery image has no seesaw rotation. Thus for faces in upright gallery images, only Qinran Chen et al. / Electronic Letters on Computer Vision and Image Analysis 6(1):13-26, 2007 5 small sidespin ration angles , for m=1,…M need to be estimated, where M is the number of candidates in the gallery. For the face in an upright test image, we need to the estimate seesaw ration angle () yg m θ − xt θ − and the sidespin ration angle y t θ −. In our method, a 3D face model is used for pose estimation. The face model will be adapted to fit a particular person in the process of pose estimation. 4.1 The adaptive face model A 3D adaptive model similar to that used in Lam and Yan’s method (cylindrical volum with a less convex surface part as face) used to represent the 3D geometry of a head. As shown in Fig. 6(a), facial feature points on the 3D model are labeled as . The 3D model is originally located in world coordinates system under two conditions: (1) the four eye corners are coplanar on x-z plane; (2) the y-z plane is the symmetrical plane of the face model. Therefore we can obtain its frontal view projection on the image plane. () , 0,...,10 jj=P x z y 0 e ε = 1 e ε = (9) P(10) P (4) P(5) P (8) P m ε (0) P(1) P(2) P(3) P (6) P(7) P ( a ) ( b ) 0.15 e ε = y x o ∗ ∗ ∗ ∗ e α e r (0) P(1) P(2) P (3) P Fig. 6 (a) The 3D face model in original pose. (b) The horizontal cross-section of the face model through the eye corners. In [6], the convexity of the less convex surface of the face model is specified by ε . When 0 ε =, the less convex surface area becomes a flat plane, while 1, ε = the model is a cylinder. In our method, considering that the convexity of the surface around mouth is evidently larger than that around eyes, we use m ε and e ε to specify the local convexities of the less convex surface around the mouth and the eyes respectively. In this paper, we set , while . With fixed 0.15 e ε =0.85 m ε =e ε , the structure of the horizontal cross-section of the face model passing through four eye corners is specified by parameters and e re α (illustrated in Fig. 6(b)). The arc passing through points , , and is also a part of a circle. The origin of the world coordinates system is on this cross-section and is marked by O in Fig. 6(b). Similarly, the structure of the horizontal cross-section of the face model passing through two mouth corners is specified by parameters and (0) P(1) P(2) P(3) P m r m α . In addition, the information of ear points is appended in our face model with the assumption that the depth distance (along z-axis) between an ear point and an outside eye corner is . Thus we adapt the simple 3D face model to fit a particular person by estimating parameters , e r e re α , and m rm α from a 2D face image. 4.2 Pose estimation and model adaptation Suppose there are M face images in nearly front view in the gallery, one image per candidate. With the assumption that the face in each gallery image has no seesaw rotation, we can estimate , , , , and , from each gallery face image using Lam and Yan’s method. These 5 parameters specify the 3D model which fits the face in the m-th gallery image. The facial feature points in the m-th gallery face image are represented by corresponding to () yg m θ −() eg rm − () eg m α −() mg rm −() mg m α −1,...,m=M (,) , mj ug− p() j u pon Fig. 2(b), j = 1,…,10. 6 Qinran Chen et al. / Electronic Letters on Computer Vision and Image Analysis 6(1):13-26, 2007 Fig.7. The projection of the left eye corner, the left ear point and the left mouth corner on y-z p lan (when 0 yt θ −=o and () 0 ygi θ −=o), for (a) the test face (b) the i-th gallery face. W O () xti ψ θ − + U V Q () xti θ −z y lem t d− lee t d− (b) W′ O′ () lee g id− z y U′ V′ Q′ ψ ′ () eg ri − ()cos( ()) eg eg ri i α −− () lem g id− ()cos( ()) mg mg ri i α −− (a) For a test image, we propose a new algorithm to estimate xt θ − and y t θ − , and to achieve 3D face model adaptation by estimating parameters et r−, et α − , mt r − , and mt α − . Let j = 1,…,10 represent the facial feature points in the test face image. Assuming that the test image is a face image of the i-th ( ) candidate in the gallery, we first estimate the corresponding potential seesaw rotation angle . As the test face is not surly same as the i-th gallery face, is a potential () , j ut− p[1, ]iM∈ () xt i θ − () xt θ −ixt θ − . Suppose that both the test face and the i-th gallery face have no sidespin rotation 0 yt θ − = o and () 0 yg i θ − = o. The left ear point, the left outside eye point and the left mouth corner on the test face are projected on y-z plane, and the projections of these facial feature points are denoted by W, U and V respectively as shown in Fig. 7(a). The corresponding facial feature points on the i-th gallery face are also also projected on y-z plane, and the projections of these three facial feature points are denoted by W′, U′ and V′ respectively as shown in Fig. 7(b). Line segments WQ and W′Q′ are perpendicular to UV and U′ V′, with intersection Q and Q′. Thus we have following equations: (| | ( ( )) | |)/ | | / x t lee t lem t tg i d d ψθ −− ++ =WQ QU UV − , (1) (| | ( ) | |)/ | | ( )/ ( ) lee g lem g tg d i d i ψ −− ′′ ′ ′′ ′′ +=WQ QU UV (2) where denotes the vertical distance between the left out side eye corner lee t d− (0) ut − p and the left ear point (9) ut − p on the test image, d denotes the vertical distance between lem t− (0) ut − p and the left mouth corner (4) ut − p, denotes the vertical distance between the () lee g di − (,0)i ug − p and (,9)i ug − p on the i-th gallery image, and d denotes the vertical distance between the p and . As it is assumed that the test face and the i-th gallery face belong to the same person, we have lem g− (,0)i ug− (,4)i ug− p k ′ ′ =WQ W Q ,k ′ ′ =QU Q U , ψ ψ ′ =, and k ′ ′ =UV U V . The scaling factor k is used to remove the size difference of the test face and the i-th gallery face. Thus the corresponding potential seesaw rotation angle of the test face can be formulated as follows: 1() () ( ) tan (( (| | ( ) | |) | |)/ | |) lee g lem g xt lee t em t didi itg dd θ ψ −− − − −− ′′ ′ ′′ ′′ ′′ ′ =+−WQ QU QU WQ ψ − . (3) Based on the i-th adapted gallery face model and facial feature points on the i-th gallery image, ′ ′ WQ , ′′ and UV ψ ′can be computed by following equations: 1 tan (( ()sin( ()) ()sin( ())/ ()) mg mg eg eg lemg ri i ri id i ψαα − −−−−− ′=− , (4) 221 ( ( )) ( ( )) cos(tan ( ( )/ ( )) ) lee g e g lee g lem g di ri didi ψ − −− −− ′′ ′ =+⋅WQ − , (5) 221 ( ( )) ( ( )) sin(tan ( ( )/ ( )) ) lee g e g lee g lem g di ri didi ψ − −− −− ′′ ′ =+⋅QU − i θ (6) From (3)-(6), the corresponding seesaw rotation angle xt− can be computed. As () () () lee g g di −− , lem di lee t lem t dd −− as well as the parameters of the i-th gallery model are independent to y t θ −() yg i θ − and , the Qinran Chen et al. / Electronic Letters on Computer Vision and Image Analysis 6(1):13-26, 2007 7 proposed algorithm is available to compute for any () xt i θ − y t θ − and (without the requirement of and ). In the same way, another corresponding potential seesaw rotation angle for the test face can be computed using the information provided in the right feature points. We average these two corresponding potential seesaw rotation angles as the final . In this paper, we assume that the seesaw rotation angle of the test face is in range of -25°to +25°. Thus if or , it will be set at 25°or -25°respectively. If we do not detect any ear point on the i-th gallery face image or on the test face image, then is assigned to be zero. () yg i θ − 0 yt θ −=o() 0 yg i θ −=o () xt i θ − () 25 xt i θ −>o() 25 xt i θ −<− o () xt i θ − Based on the estimated , we can further estimate corresponding, and , and using the geometric information about two outside eye corners and the face contour in the test face image. Let ei () xt i θ −() et ri −() et i α −() yt i θ − Π represent a circle that passes through the outside eye corners on the corresponding potential model of the test face and centers at the origin of the word coordinate system. Let ei Θ be the projection of ei Π on the image plane. If , ei should be a line passing the outside eye corners () 0 xt i θ −=oΘ(0) ut − p and in the test image (Lam and Yan’s algorithm just consider this situation), otherwise ei (3) ut− p Θ should be an ellipse passing the outside eye corners in the test image as shown in Fig. 8. (11) ut − p(11) (11) (, ) ut ut pp −− represents the middle point between xy (0) ut − p and . In this paper, we assume that and (3) ut− p(0) ut− p(3) ut − p are not occluded in an image. ai p (, ) ai ai p x yp and bi bi bi p (, ) p x yp represent the two end points of the long axis of the ellipse ei Θ respectively. (, ) ci ci ci p p x yei p is the center of Θ . Fig.8 ei Θ, the projection of ei Πin test image when () 0 xt i θ −>o and (a) () 0 yt i θ − = o, (b) () 0 yt i θ −>o. y x bi p∗ ∗ ∗ ∗∗∗ ai p ci p (0) ut− p(3) ut − p (11) ut− p (a) ∗ ∗ ∗ ∗ ∗ ∗ (0) ut − p(3) ut− p (11) ut − p bi p ai p ci p (b) The ellipse (,) ei x yΘ in the test image can be formulated as follows: ()cos( ())cos ()sin( ())cos( ())sin ci et yt et yt xt p x ri i ri i i x θλ θθλ −− −− − =+ + , (7) ()sin( ())sin ci et xt p y ri i y θλ −− =− + (8) where (0,2 ) λ π ⊂ is the independent variable and ( ci p x , ci p y) denotes the location of the ellipse’s center. , () et ri −ci p x and ci p y are given by: (0) (3) () (2cos( ())cos ()) et ut ut yt et ri i i θα −−− − − =−pp , (9) (11) (0) (3) tan( ( ))tan( ( ))cos( ( ))( / 2) ci ut pytetxtutut p x iii θαθ − −−−−− =− − +pp x, (10) (11) (0) (3) tan( ( ))sin( ( )) (2cos( ( ))) ci ut petxtutut yt p yii i αθ θ − −−−− − =−pp y+ . (11) Combining (9), (10) and (11), we rewrite (7) and (8) as follows: (11) (0) (3) (0) (3) (0) (3) cos (2 cos( ( ))) tan( ( ))cos( ( ))sin (2 cos( ( ))) tan( ( )) tan( ( )) cos( ( ))( / 2) ut ut ut et yt xt ut ut et yt et xt ut ut p x iii iii x λαθθλα θαθ − −− − − − −− − −−−−− =⋅− ⋅− −−+ +pp pp pp i , (12) (11) (0) (3) (0) (3) sin( ( ))sin (2cos( ( ))cos( ( ))) tan( ( ))sin( ( )) (2cos( ( ))) ut xt ut ut xt et et xt ut ut yt p yi ii ii i θλ θ α αθ θ − −−−−−−−−−− =− ⋅ − + − +pp pp y (13 ) 8 Qinran Chen et al. / Electronic Letters on Computer Vision and Image Analysis 6(1):13-26, 2007 Then we try to locate ai p and bi p on ei Θ on the image plane. One of (,) ai ai ai p p x xp and (,) bi bi bi p p x xp is represented by symbol (,) ii ip p x xp. We have: ()cos( ())cos ()sin( ())cos( ())sin ii ici p et yt p et yt xt p p x ri i ri i i x θλ θθλ −− −− − =+ + ci , (14) ()sin( ())sin ii p et xt p p yri i y θλ −− =− + , (15) 22 ()()(( 2 )) pp pp et x ici ici xyyri − −+−= = pyt xt ii λθ θ −− = . (16) Thus the following equation can be conduced from (14), (15) and (16): 22 2 (cos( ( ))cos( ( ))) (tan( )) sin(2 ( ))cos( ( ))tan( ) (sin( ( ))) 0 i i xt yt p yt xt p yt ii ii i θθ λ θθλθ −− −− − −− . (17) tan( ) tan(( ())/cos( ()) i, ( ) 1 tan tan(( ())/cos( ()) ai and pytxt ii λθθ − −− = ( ) 1 tan tan(( ( ))/cos( ( )) bi pyt ii tx θθπ − −− λ = + are hence obtained from (17). Referring (12) and (13), we found that the pair of locations of ai p and bi p in image plane are determined by et− and yt−. According to statistics obtained from some 100 face images, et− should be in range of 35° to 65°. In this paper we assume yt−. Thus based on dense sampling et− and , we can compute K possible pairs of ai p and bi p which are represented by kai bi Φ. In the following, Φ= will be used to approximate all possible pairs of ai ()i α i θ i α i θ ∈− oo i α ∈oo kkk K=pp K () () () [ 30,30] () [35,65] () [ 30,30] yt i θ −∈− oo ( ( ), ( )), 1,..., ,1,..., kk p and bi p for et−and yt . Let et−, yt− represent the sample values of et− and yt− corresponding to the k Φ. The smaller the sampling interval is, the larger K is and the more efficient the approximation is. For our application, we set the interval at 1 for sampling both et− and , so as a result K is equal to 1800. According to the geometric character of our face model, it is given that among all possible pairs of ai () [35,65]i α ∈oo i θ −∈− oo ik ik θ i α i θ i () [ 30,30] (, ) α (, ) () () o() α () yt i θ − p and bi p , only the true one will appear on the face contour in the test image and the true one will definitely appear on the face contour. Let ai d represent the shortest distance from ai to the detected face contour and bi dk represent the shortest distance from bi p to the detected face contour. d is the sum of and . Then we estimate ai () cp kkpcp k() () () () i cp k () ai cp dk () bi cp dk p and bi p ,et−,et− and yt− by ai , bi , ()ri i α c cp() ()i θ ()p() (() ())2 ai bi cc−pp , and respectively, where the d is the smallest one among i cp , 1. We can estimate the corresponding mt r −, mt− and yt− in the same way but using the geometric information about two outside mouth corners and the face contour in the test face image. We average the two potential sidespin rotation angles estimated by using the eye corners information and the mouth corners information respectively as the final potential sidespin rotation angle . (, ) et ic α −(, ) yt ic θ −c dkk K=ii α i θ M= () i cp (), 1,..., cK≤≤ () () () () yt i θ − In this paper, the proposed algorithm is employed to estimates the pose and the 3D face model of the test face, assuming that the test image is an image of the m-th candidate in the gallery for m. Thus we obtain M potential poses and M potential models for the test face image referring to all candidates in the gallery. Among these potential poses, the i-th one will be regarded as the matched potential pose, if the test face and the i-th gallery face really belong to one person. The matched potential pose consists of matched potential seesaw rotation angle 1,... y mt θ − and matched potential sidespin rotation angle y mt−. θ (b)(a) Fig.9 Some test faces in (b) and their matched potential poses obtained by the proposed algorithm referring to the corresponding gallery faces in (a). The poses of the test faces and the gallery faces estimated by La m and Yan’s algorithm are also given in parentheses for comparison. 5 , 20.5 ym t xm t θθ −− == oo (3.6,0) yt xt θθ −− == oo 8 , 16.1 ym t xm t θθ −− =− = oo (6.3,0) yt xt θθ −− =− = oo 18 , 8.9 ym t xm t θθ −− =− = oo ( 15.2 , 0 ) yt xt θθ −− =− = oo (1.4,0) yt xt θθ −− == oo (3.1,0) yt xt θθ −− =− = oo (0.8,0) yt xt θθ −− == oo Fig. 9 shows some examples to illuminate the performance of the proposed algorithm for pose estimation. The images in Fig. 9.(a) show some test faces and the images in Fig. 9.(b) show the gallery faces belonging Qinran Chen et al. / Electronic Letters on Computer Vision and Image Analysis 6(1):13-26, 2007 9 to the same person. The matched potential poses of these test faces, which are estimated by the proposed algorithm, are shown in Fig. 9(b). For comparison, the poses of these test faces and gallery faces estimated by Lam and Yan’s pose estimation algorithm [6] are also given in the parentheses. The result shows that the proposed pose estimation algorithm can obtain effective and general estimate of poses (including both seesaw and sidespin rotations) for test faces. 5 Pose Invariant Feature Extraction and Classification Based on estimated pose and 3D face model of each gallery faces, we reconstruct a set of facial feature points’ locations , on the frontal view of each gallery face from ug ug ug p p −− −, for m = 1,…,M, using affine transform. Similarly, based on the M potential poses and corresponding potential models of the test face, M sets of facial feature points’ locations , m = 1,…,M on the corresponding M potential frontal views of the test face are reconstructed from . For m = 1,…,M, point set re normalized and aligned to (, , besides and are normalized to and . We define as follows: (,) ( ( , ), ( , )), 1,...,8 us g us g mj us g p p xmjymjj −− −=p (,) ( ( , ), ( , )), 1,...,8 mj xmjymjj=p 8a 8 (,) ( (,), (,)), 1,...,8 us t us t mj us t p p xmjymjj −− −=p() ( ( ), ( )), 1,...,8 ut ut j ut p p xjyjj −− −=p(,) ,1,..., mj us t j −=p ) ,1,..., mj us g j −=p() et rm −() mt rm −() eg rm −() mg rm − (), 1,... w Dmm M= 8(,) (,) 910 0 j= Corresponding geometric features matching is then preformed by computing w D, for m = 1,…,M to measure the similarity between the test face and each gallery face. N( () () () () () mj mj wjustusgetegmtmg Dm k kr m r m kr m r m −− − − − − =−+−+− ∑pp . (18) m NM ≤ ) faces in the gallery which satisfy wD m D δ ≤ ( D δ is a threshold) are chosen as qualified gallery faces and passed on for the further recognition process. Corresponding to these N qualified faces in the gallery, N potential poses and potential models are remained for the test face. Based on the n-th remained potential pose and corresponding potential model, for n = 1,…,N, the eyes and eyebrows template et T −, the nose template nt T − and the mouth template mt T − in the n-th potential frontal view of the test face are synthesized by affine transform. We also synthesized the eyes and eyebrows template eg T −, the nose template ng T − and the mouth template in the frontal view of the n-th qualified gallery face, for n = 1,…,N. Then et T −, nt T −and mt− compared to eg T −, ng T −and mg T − by correlations, for n = 1,…,N. The larger the correlations are, the more similar to the test face the corresponding qualified face in the gallery is. This comparison is referred to as corresponding template matching. Combining the results of the corresponding geometric feature matching and corresponding template matching, we can achieve final classification and find a qualified face in the gallery which is most like the test face. ()nn nnnnnTn nnn Fig.10. (a) A test face image. (b) The corresponding gallery face image which captures the same person as the test face image. The facial feature templates of the test face synthesized using Lam’s method (c), o () () () () () mg Tn −() () () () () () In this section, the operations of normalization, alignment, and correlation in our method are same as those in Lam and Yan’s method. In our method, the operations of synthesizing facial feature points’ locaitons on frontal view and synthesizing facial feature templates in frontal view are similar to corresponding operations in Lam and Yan’s method except for considering extra estimated seesaw rotation. In the operation of geometric feature matching, we use 9 facial feature points and 2 estimated model parameters to replace the 15 facial feature points used in Lam and Yan’s method, because the 6 facial feature points on the face contour are not pose invariant features in our case when the seesaw rotation is considered. The implementation of these operations can be referenced to [6]. Fig. 10 (a) and (b) show a test face image and a gallery face image of the same person respectively. Fig. 10 (c) and (e) show facial feature templates of the test face and the gallery face synthesized by using Lam and Yan’s method. Fig. 10(d) shows facial feature templates of the test face synthesized by using our method based on the matched potential pose and the matched potential face model. Compared to the facial feature templates shown in Fig. 10(c), the facial feature templates shown in Fig. 10(d) are more like in frontal view and have more similarity to those shown in Fig. 10(e). ( a ) ( b ) ( c ) ( d ) ( e ) f the test face synthesized using our method based on the matched potential pose (d), and of the gallery face s y nthesized usin g Lam and Yan’s method ( e ) .