scieee AI-readable full text Open interactive document viewer

A Comparison Framework for Walking Performances using aSpaces

Villanueva Pipaón, Juan José; Gonzàlez, Jordi; Varona, Javier; Roca, F. Xavier

Abstract

In this paper, we address the analysis of human actions by comparing different performances of the same action executed by different actors. Specifically, we present a comparison procedure applied to the walking action, but the scheme can be applied to other different actions, such as bending, running, etc. To achieve fair comparison results, we define a novel human body model based on joint angles, which maximizes the differences between human postures and, moreover, reflects the anatomical structure of human beings. Subsequently, a human action space, called aSpace, is built in order to represent each performance (i.e., each predefined sequence of postures) as a parametric manifold. The final human action representation is called p-action, which is based on the most characteristic human body postures found during several walking performances. These postures are found automatically by means of a predefined distance function, and they are called key-frames. By using key-frames, we synchronize any performance with respect to the p- action. Furthermore, by considering an arc length parameterization, independence from the speed at which performances are played is attained. As a result, the style of human walking can be successfully analysed by establishing the differences of the joints between a male and a female walkers.

Full text

Electronic Letters on Computer Vision and Image Analysis 5(3):105-116,2005 A Comparison Framework for Walking Performances using aSpaces Jordi Gonz`alez∗, Javier Varona+, F. Xavier Roca∗and Juan J. Villanueva∗ ∗Centre de Visi´ o per Computador & Dept. d’Inform` atica, Universitat Aut` onoma de Barcelona (UAB), 08193 Bellaterra, Spain +Dept. Matem` atiques i Inform` atica & Unitat de Gr` afics i Visi´ o, Universitat de les Illes Balears (UIB), 07071 Palma de Mallorca, Spain Received 16 December 2004; accepted 6 June 2005 Abstract In this paper, we address the analysis of humanactions by comparingdifferentperformances ofthe same action executed by different actors. Specifically, we present a comparison procedure applied to the walking action, but the scheme can be applied to other different actions, such as bending, running, etc. To achieve fair comparison results, we define a novel human body model based on joint angles, which maximizes the differences between human postures and, moreover, reflects the anatomical structure of human beings. Subsequently, a human action space, called aSpace, is built in order to represent each performance (i.e., each predefined sequence of postures) as a parametric manifold. The final human action representation is called p–action, which is based on the most characteristic human body postures found during several walking performances. These postures are found automatically by means of a predefined distance function, and they are calledkey-frames. By usingkey-frames,we synchronizeanyperformancewith respect to the p– action. Furthermore, by considering an arc length parameterization, independence from the speed at which performances are played is attained. As a result, the style of human walking can be successfully analysed by establishing the differences of the joints between a male and a female walkers. Key Words: Human Motion Modeling, Human Body Modeling, Synchronization, Key-frames. 1 Introduction Computational models of action style are relevant to several important application areas [10]. On the one hand, it helps to enhance the qualitative description provided by a human action recognition module. Thus, for example, it is important to generate style descriptions which best characterize an specific agent for identification purposes. Also, the style of a performance can help to establish ergonomic evaluation and athletic training procedures. Another application domain is to enhance the human action library by training different action models for different action styles, using the data acquired from a motion capture system. Thus, it should be possible to re-synthesize human performances exhibiting different postures. Correspondence to: <[email protected].es> Recommended for acceptance by <Perales F., Draper B.> ELCVIA ISSN:1577-5097 Published by Computer Vision Center / Universitat Aut`onoma de Barcelona, Barcelona, Spain 106 Gonz` alez et al. / Electronic Letters on Computer Vision and Image Analysis 5(3):105-116,2005 In the literature, the most studied human action is walking. Human walking is a complex, structured, and constrained action, which involves to maintain the balance of the human body while transporting the figure from one place to another. The most exploited characteristic is the cyclic nature of walking, because it provides uniformity to the observed performance. In this paper, we propose to use a human action model in the study of the style inherent in human walking performances, such as the gender, the walking pace, or the effects of carrying load, for example. Specifically, we show how to use the aSpace representation presented in [16] to establish a characterization of the walking style in terms of the gender of the walker. The resulting characterization will consist of a description of the variation of specific limb angles during several performances played by agents of different gender. The aim is to compare performances to derive motion differences between female and male walkers. 2 Related Work Motion capture is the process of recording live movement and translating it into usable mathematical terms by tracking a number of key points or regions/segments in space over time and combining them to obtain a 3-D representation of the performance [20]. By reviewing the literature, we distinguish between two different strategies for human action modeling based on motion capture data, namely data-driven and model-driven. Data-driven approaches build detailed descriptions of recorded actions, and develop procedures for their adaption and adjustment to different characters [14]. Model-driven strategies search for parameterized representations controlled by few parameters [6]: computational models provide compactness and facilities for an easy edition and manipulation. Both approaches are reviewed next. Data-driven procedures do care of specific details of motion: accurate movement descriptions are obtained by means of motion capture systems, usually optical. As a result, a large quantity of unstructured data is obtained, which is difficult to be modified while maintaining the essence of motion [24, 26]. Inverse Kinematics (IK) is a well-known technique for the correction of one human posture [7, 12]. However, it is difficult to apply IK over a whole action sequence while obeying spatial constraints and avoiding motion discontinuities. Consequently, current effort is centered on Motion Retargetting Problem [13, 22], i.e. the development of new methods for the edition of recorded movements. Model-driven methods search for the main properties of motion: the aim is to develop computational models controlled by a reduced set of parameters [2]. Thus, human action representations can be easily manipulated for its re-use. Unfortunately, the development of action models is a difficult task, and complex motions are hard to be composed [21]. Human action modeling can be based on Principal Component Analysis (PCA) [4, 8, 11, 16, 25]. PCA computes an orthogonal basis of the samples, the so-called eigenvectors, which control the variation along the maximum variance directions. Each principal component is associated to a mode of variation of the shape, and the training data can be described as a linear combination of the eigenvectors. The basic assumption is that the training data generates a single cluster in the shape eigenspace [18]. Following this strategy, we use a PCA-based space to emphasize similarities between the input data, in order to describe motion according to the gender of the performer. In fact, this space of reduced dimensionality will provide discriminative descriptions about style characteristics of motion. 3 Defining the Training Samples In our experiments, an optical system was used to provide real training data to our algorithms. The system is based on six synchronized video cameras to record images, which incorporates all the elements and equipment necessary for the automatic control of cameras and lights during the capture process. It also includes an advanced software pack for the reconstruction of movements and the effective treatment of occlusions. Gonz` alez et al. / Electronic Letters on Computer Vision and Image Analysis 5(3):105-116,2005 107 (a) (b) Figure 1: Procedure for data acquisition. Figs. (a) and (b) shows the agent with the 19 markers on the joints and other characteristic points of its body. (a) (b) Figure 2: (a) Generic human body model represented using a stick figure similar to [9], here composed of twelve limbs and fifteen joints. (b) Hierarchy of the joints of the human body model. Consequently, the subject first placed a set of 19 reflective markers on the joints and other characteristic points of the body, see Fig. 1.(a) and 1.(b). These markers are small round pieces of plastic covered in reflective material. Subsequently, the agent is placed in a controlled environment (i.e., controlled illumination and reflective noise), where the capture will be carried out. As a result, the accurate 3-D positions of the markers are obtained for each recorded posture p s, 30 postures per second: ps=(x1,y 1,z 1, ..., x19,y 19,z 19)T.(1) An action will be represented as a sequence of postures, so a proper body model is required. In our experiments, not all the 19 markers are considered to model human actions. In fact, we only process those markers which correspond to the joints of a predefined human body model. The body model considered is composed of twelve rigid body parts (hip, torso, shoulder, neck, two thighs, two legs, two arms and two forearms) and fifteen joints, see Fig. 2.(a). These joints are structured in a hierarchical manner, where the root is located at the hips, see Fig. 2.(b). We next represent the human body by describing the elevation and orientation of each limb using three different angles which are more natural to be used for limb movement description [3]. We consider the 3-D 108 Gonz` alez et al. / Electronic Letters on Computer Vision and Image Analysis 5(3):105-116,2005 Figure 3: The polar space coordinate system describes a limb in terms of the elevation φ l, latitude θl, and longitude ψl. polar space coordinate system which describes the orientation of a limb in terms of its elevation, latitude and longitude, see Fig. 3. As a result, the twelve independently moving limbs in the 3-D polar space have a total of twenty-four rotational DOFs which correspond to thirty-six absolute angles. So we compute the 3-D polar angles of a limb (i.e., elevation φ l, latitude θl, and longitude ψl) as: φl=tan −1⎛ ⎝ yi−yj (xi−xj)2+(zi−zj)2 ⎞ ⎠, θl=tan −1⎛ ⎝ xi−xj (yi−yj)2+(zi−zj)2 ⎞ ⎠, ψl=tan −1⎛ ⎝ zi−zj (xi−xj)2+(yi−yj)2 ⎞ ⎠,(2) where denominators are also prevented to be equal to zero. Using this description, angle values lie between the range of −π 2,π 2, and the angle discontinuity problem is avoided. Note that human actions are constrained movement patterns which involve to move the limbs of the body in a particular manner. That means, there is a relationship between the movement of different limbs while performing an action. In order to incorporate this relationship into the human action representation, we consider the hierarchy of Fig. 2.(b) in order to describe each limb with respect to its parent. That means, the relative angles between two adjacent limbs are next computed using the absolute angles of Eq. (2). Consequently, by describing the the human body using the relative angles of the limbs, we actually model the body as a hierarchical and articulated figure. As a result, the model of the human body consists of thirty-six relative angles: Δs=(φ 1,θ 1,ψ 1,φ  2,θ 2,ψ 2, ..., φ 12,θ 12,ψ 12)T.(3) Using this definition, we measure the relative motion of the human body. In order to measure the global motion of the agent within the scene, the variation of the (normalized) height of the hip u sover time is included in the model definition: xs=(us,Δs)T.(4) Gonz` alez et al. / Electronic Letters on Computer Vision and Image Analysis 5(3):105-116,2005 109 Figure 4: The three most important modes of variation of the aWalk aSpace. Therefore, our training data set Ais composed of rsequences A={H 1,H2, ..., Hr}, each one corresponding to a cycle or stride of the aWalk action. Three males and three females were recorded, each one walking five times in circles. Each performance Hjof the action Acorresponds to fjhuman body configurations: Hj={x1,x2, ..., xfj},(5) where each xiof dimensionality n×1stands for the 37 values of the human body model described previously. Consequently, our human performance analysis is restricted to be applied to the variation of these twelve limbs. 4 The aWalk aSpace Once the learning samples are available, we compute the aSpace representation Ωof the aWalk action, as detailed in [16]. In our experiments, the walking performances of three females and two males were captured to collect the training data set. For each walker, near 50 aWalk cycles have been recorded. As a result, the training data is composed of near 1500 human posture configurations per agent, thus resulting 7500 3D body postures for building the aWalk aSpace. From Eq. (5), the training data set Ais composed of the acquired human postures of the rperformances: A={x1,x2, ..., xf},(6) where frefers to the overall number of training postures for this action: f= r  j=1 fj.(7) The mean human posture ¯ xand the covariance matrix Σof Aare calculated. Subsequently, the eigenvalues Λand eigenvectors Eof Σare found by solving the eigenvector decomposition equation. We preserve major linear correlations by considering the eigenvectors e icorresponding to the largest eigenvalues λi. Fig. 4 shows the three eigenvectors associated to the three largest eigenvalues, which correspond to the most relevant modes of change of the human posture in the aWalk aSpace. As expected, these modes of variation are mainly related to the movement of legs and arms. So, by selecting the first meigenvectors, {e1,e2, ..., em}, we determine the most important modes of variation of human body during the aWalk action [8]. The value for mis commonly determined by eigenvalue thresholding. Consider the overall variance of the training samples, computed as the sum of the eigenvalues: 110 Gonz` alez et al. / Electronic Letters on Computer Vision and Image Analysis 5(3):105-116,2005 λT= n  k=1 λk.(8) If we need to guarantee that the first meigenvectors actually model, for example, 95% of the overall variance of the samples, we choose mso that: m k=1 λk λT ≥0.95.(9) The individual contribution of each eigenvector determines that 95% of the variation of the training data is captured by the thirteen eigenvectors associated to the thirteen largest eigenvalues. So the resulting aWalk aSpace Ωis defined as the combination of the eigenvectors E, the eigenvalues Λand the mean posture¯ x: Ω=(E,Λ,¯ x).(10) 5 Parametric Action Representation: the p–action Using the aWalk aSpace, each performance is represented as a set of points, each point corresponding to the projection of a learning human posture xi: yi=[e1, ..., em]T(xi−¯ x).(11) Thus, we obtain a set of discrete points yiin the action space that represents the action class Ω. By projecting the set of human postures of an aWalk performance H j, we obtain a cloud of points wich corresponds to the projections of the postures exhibited during such a performance. We consider the projections of each performance as the control values for an interpolating curve g j(p), which is computed using a standard cubic-spline interpolation algorithm [23]. The parameter prefers to the temporal variation of the posture, which is normalized for each performance, that is, p∈[0,1]. Thus, by varying p,we actually move along the manifold. This process is repeated for each performance of the learning set, thus obtaining rmanifolds: gj(p),p∈[0,1],j =1, ..., r. (12) Afterwards, the mean manifold g(p)is obtained by interpolating between these means for each index p. This performance representation is not influenced by its duration, expressed in seconds or number of frames. Unfortunately, this resulting parametric manifold is influenced by the fact that any subject performs an action in the way he or she is used to. That is to say, the extreme variability of human posture configurations recorded during different performances of the aWalk action affects the mean calculation for each index p. As a result, the manifold may comprise abrupt changes of direction. A similar problem can be found in the computer animation domain, where the goal is to generate virtual figures exhibiting smooth and realistic movement. Commonly, animators define and draw a set of specific frames, called key frames or extremes, which assist the task of drawing the intermediate frames of the animated sequence. Likewise, our goal is set to the extract the most characteristic body posture configurations which will correspond to the set of key-frames for that action. From a probabilistic point of view, we define characteristic postures as the least likely body postures exhibited during the action performances. As the aSpace is built based on PCA, such a space can also be used to compute the action class conditional density P(x j|Ω). We assume that the Mahalanobis distance is a sufficient statistic for characterizing the likelihood: d(xj)=(xj−¯ x)TΣ(xj−¯ x).(13) Gonz` alez et al. / Electronic Letters on Computer Vision and Image Analysis 5(3):105-116,2005 111 Figure 5: Distance measure after pose ordering applied to the points of the mean manifold in the aWalk aSpace. Maxima (i.e., the key-frames) also correspond to important changes of direction of the manifold. So, once the mean manifold g(p)is established, we compute the likelihood values for the sequence of poseordered projections that lie in such a manifold [5, 19]. That is, we apply Eq. (13) for each component of the manifold g(p). Local maxima of this function correspond to locally maximal distances or, in other words, to the least likely samples, see Fig. 5. Since each maximum of the distance function corresponds to a key-frame k i, the number of key-frames kis determined by the number of maxima. Thus, we obtain the set of time-ordered key-frames for the aWalk action: K={k1,k2, ..., kk},ki∈g(p).(14) Once the key-frame set Kis found, the final human action model is represented as a parametric manifold f(p), called p–action, which is built by interpolation between the peaks of the distance function defined in Eq. (13). We refer the reader to [16] for additional details. Fig. 6 shows the final aWalk model Γ, defined as the combination of the aWalk aSpace Ω, the key-frames Kand the p–action f: Γ=(Ω,K,f).(15) 6 Human Performance Comparison In order to compare performances played by male and female agents, we define two different training sets: HWM={x1,x2, ..., xfM}, HWF={x1,x2, ..., xfF},(16) that is, the set human postures exhibited during several aWalk performances for a male and a female agent, respectively. Next, we project the human postures of HWMand HWFin the aWalk aSpace, as shown in Fig. 7. The cyclic nature of the aWalk action explains the resulting circular clouds of projections. Also, note that both performances do not intersect, that is, they do not exhibit the same set of human postures. This is due to the high variability inherent in human performances. Consequently, we can identify a posture as belonging to a male or female walker. 112 Gonz` alez et al. / Electronic Letters on Computer Vision and Image Analysis 5(3):105-116,2005 Figure 6: Prototypical performance manifold, or p–action, in the aWalk aSpace. Depicted human postures correspond to the key-frame set. (a) (b) Figure 7: Male and female postures projected in the aWalk aSpace, by considering two (a) and three (b) eigenvectors for the aSpace representation. However, the scope of this paper is not centered on determining a discriminative procedure between generic male and female walkers. Instead, we look for a comparison procedure to subsequently evaluate the variation of the angles of specific agents while performing the same action, in order to derive a characterization of the action style. Following the procedure described in the last section, we use the projections of each walker to compute the performance representation for the male ΓWMand female ΓWFagents: ΓWM=(Ω,KWM,fWM), ΓWF=(Ω,KWF,fWF),(17) where fWMand fWFrefer to the male and female p–actions, respectively. These manifolds have been obtained by interpolation between the key-frames of their respective key-frame set, i.e., K WMand KWF. Fig. 8 shows the resulting p–action representations in the aWalk aSpace Ω. Gonz` alez et al. / Electronic Letters on Computer Vision and Image Analysis 5(3):105-116,2005 113 Figure 8: Male and female p–action representations in the aWalk aSpace. 7 Arc length parameterization of p–actions In order to compare the human posture variation for both performances, we sample both p–actions to describe each manifold as a sequence of projections: fWM(p)=[ yWM 1,yWM 2, ..., yWM qM], fWF(p)=[ yWF 1,yWF 2, ..., yWF qF],(18) where qMand qFrefer to the number of projections considered for performance comparison. Subsequently, the sampling rate of both p–actions should be established in order to attain independence from the speed at which both performances have been played. That means, synchronization of recorded performances is compulsory to allow comparison. Speed control is achieved by considering the distance along a curve of interpolation or, in other words, by establishing a reparameterization of the curve by arc length [17]. Thus, once the aWalk p–action is parameterized by arc length, it is possible to control the speed at which the manifold is traversed. Subsequently, the key-frames will be exploited for synchronization: the idea of synchronization arises from the assumption that any performance of a given action should present the key-frames of such an action. Therefore, the key-frame set is considered as the reference postures in order to adjust or synchronize any new performance to our action model. Subsequently, by considering the arc length parameterization, the aim is to sample the new performance and the p–action so that the key-frames are equally spaced in both manifolds. Therefore, both p–actions are parameterized by arc length and, subsequently, the synchronization procedure described in [15] is applied: once the key-frames establish the correspondences for f WMand fWF, we can modify the rate at which the male and female p–actions are sampled, so that their key-frames coincide in time with the key-frames of the aWalk p–action. 8 Experimental Results Once the male and female p–actions are synchronized, the angle variation for different limbs of the human body model can be analysed. Fig. 9.(a),(b),(c), and (d) show the evolution of the elevation angle for four limbs of the human body model, namely the shoulder, torso, left arm, and right thigh, respectively. By comparing the depicted angle variation values of both walkers, several differences can be observed. The female walker moves her shoulder in a higher degree than the male shoulder. That is, the swing movement of