Full text
2012 5 Grégory Rogez Advances in Monocular Exemplarbased Human Body Pose Analysis: Modeling, Detection and Tracking Departamento Director/es Ingeniería Electrónica y Comunicaciones Orrite Uruñuela, Carlos Director/es Tesis Doctoral Autor Repositorio de la Universidad de Zaragoza – Zaguan http://zaguan.unizar.es UNIVERSIDAD DE ZARAGOZA
Departamento Director/es Grégory Rogez ADVANCES IN MONOCULAR EXEMPLAR-BASED HUMAN BODY POSE ANALYSIS: MODELING, DETECTION AND TRACKING Director/es Ingeniería Electrónica y Comunicaciones Orrite Uruñuela, Carlos Tesis Doctoral Autor 2012 Repositorio de la Universidad de Zaragoza – Zaguan http://zaguan.unizar.es UNIVERSIDAD DE ZARAGOZA
Departamento Director/es Director/es Tesis Doctoral Autor Repositorio de la Universidad de Zaragoza – Zaguan http://zaguan.unizar.es UNIVERSIDAD DE ZARAGOZA
Advances in Monocular Exemplar-based Human Body Pose Analysis: Modeling, Detection and Tracking Gr´egory Rogez Ph.D. Dissertation Advisor Dr. Carlos Orrite Uru˜ nuela June 2012
ii
Advances in Monocular Exemplar-based Human Body Pose Analysis: Modeling, Detection and Tracking Gr´egory Rogez Ph.D. Thesis PhD Committee Members: Dr. D. Armando Roy Yarza Universidad de Zaragoza, Espa˜na Dr. D. Jos´e Mara Mart´ınez Montiel Universidad de Zaragoza, Espa˜na Dr. D. Dariu M. Gavrila University of Amsterdam, The Netherlands Dr. D. Dimitrios Makris Kingston University London, UK Dr. D. Nicol´as P´erez de la Blanca Capilla Universidad de Granada, Espa˜na External European Reviewers: Dr. D. Philip H. S. Torr Oxford Brookes University, UK Dr. D. Antonis Argyros University of Creete, Greece
iv
Acknowledgements This dissertation would not have been possible without the guidance and the help of several individuals who, in one way or another, contributed in the preparation and completion of this thesis. First and foremost, my gratitude goes to my supervisor and friend, Dr. Carlos Orrite, whose guidance, support and encouragements made this research possible. Thanks for giving me freedom to shape my research path and the time to finally submit the thesis I had in mind. Next, I would like to thank Prof. Armando Roy. This thesis would not have been possible without the FPU grant that I obtained with his help. During the preparation of the thesis, I have had the chance to collaborate with many people. I want to thank the members of the CVLab in Zaragoza, more concretely Dr J. Elias Herrero and Dr Jes´us Mart´ınez whose earlier phd work have been the basis of the first part of this thesis. I want to thank Dr Jose J. Guerrero for his help and advices on the elaboration of the second part of the thesis. The research work of the third part of the thesis has been done when I was visiting Oxford Brookes University. During the time I have spent working at Oxford Brookes, I have had the honour and the pleasure of meeting and working with many very talented and enthusiastic people. I want to thank Professor Philip Torr for giving me the opportunity to work in a great environment, and for his valuable advice, enthusiasm and support. I want to thank Dr Jon Rihan for the great collaboration. I will always remember our endless brainstorming sessions. I want to thank the external reviewers of the thesis Professor Antonis Argyros and Professor Philip Torr (bis) and the member of the PhD committee for their valuable observations and suggestions. I also want to thank the anonymous reviewers who helped to improve the quality of our conference and journal papers and consequently the quality of this thesis. Last but not least, I wish to thank my family and friends for their support, especially Dr Ruben Mart´ınez, Dr Antonio Miguel and Dr Ignacio Mart´ınez for their advices and help, and Jes´us Senar for his continuous encouragements. And most of all, thank you Marga for your patience and for the love you gave me. The completion of my dissertation has been a long journey. It would not have been possible without you. The author would like thank the I3A and the Department of Electronics and Communications Engineering, University of Zaragoza, for their confidence. He also wants to acknowledge support provided by: “Departamento de Ciencia, Tecnolog´ıa y Universidad del Gobierno de Arag´on”, “Fondo Social Europeo” and “Ministerio de Ciencia e Innovaci´on (TIN2010-20177)”. Part of this work was also supported by the Spanish Ministry of Education under the FPU grant AP2003-2257.
vi
xiii line sin ning´un coste extra, por lo tanto, la clasificaci´on de cada ventana puede realizarse con un conjunto diferente de cascadas. Este esquema de clasificaci´on adaptable permite una aceleraci´on considerable y una detecci´on de posturas a´un m´as eficiente que simplemente utilizando el mismo conjunto de tama˜no fijo sobre toda la imagen. Cada cascada puede votar por una o m´as clases por lo que el conjunto genera una distribuci´on que puede ser ´util para resolver ambig¨uedades entre posturas. Seguimiento de personas En esta tesis hemos abordado tanto el problema de seguimiento del sujeto en la imagen, como el del seguimiento en el espacio de posturas. La versi´on discretizada del Toroide ha sido empleada para limitar el espacio de modelos plausibles en el cap´ıtulo 3 o las clases plausibles en el cap´ıtulo 7, mediante simples restricciones espacio-temporales. La versi´on continua del Toroide ha sido utilizada en el cap´ıtulo 5 para muestrear posibles candidatos de postura-vista sobre la superficie del espacio basado en el filtro de part´ıculas. Este ´ultimo m´etodo ha demostrado ser m´as robusto a la hora de resolver ambig¨uedades en la postura, ya que permite mantener m´ultiples hip´otesis a lo largo del tiempo. El seguimiento en la imagen se ha realizado mediante un simple filtro de Kalman sobre la posici´on en la imagen, la escala y el ´angulo en el cap´ıtulo 3 y cap´ıtulo 7, con casos f´aciles donde s´olo un sujeto es seguido. En la segunda parte de esta tesis, el problema ha sido simplificado mediante la introducci´on de la calibraci´on de la c´amara respecto a la escena ya que el seguimiento se aplica en el plano del suelo, explorando as´ı un espacio 2D en lugar de un espacio de 4 dimensiones (posici´on, escala y ´angulo). En el cap´ıtulo 5, hemos propuesto un sistema eficiente de filtro de part´ıculas para seguimiento de posturas 3D en escenas de video vigilancia calibradas: el seguimiento se realiz´o conjuntamente en el plano del suelo y en la superficie de este subespacio Toroide que asocia punto de vista y postura. As´ı, s´olo cuatro dimensiones necesitan ser exploradas para realizar el seguimiento de posturas de personas caminando en el espacio 3D. Aportaciones y L´ıneas Futuras de Investigaci´on A continuaci´on, vamos a resumir las principales aportaciones de cada secci´on de la tesis y enumerar algunas de las posibles l´ıneas futuras de investigaci´on. Parte I: estimaci´on de la postura de un ´unico individuo con una c´amara est´atica paralela al suelo En la primera parte de la tesis hemos presentado un conjunto de modelos 2D espaciotemporales para an´alisis de movimiento humano. Para hacer frente a la restricci´on respecto al punto de vista, se han entrenado diferentes modelos locales 2D espacio-temporales correspondientes a varias vistas de la misma secuencia. Posteriormente, se han concatenado todos ellos orden´andolos en un espacio global. Al procesar una secuencia se han considerado las limitaciones temporales y espaciales para construir la matriz de transici´on probabil´ıstica (PTM) que da la predicci´on de un fotograma a otro de los modelos m´as probables del conjunto. Los experimentos llevados a cabo en secuencias tanto de interiores como de exteriores, han demostrado la capacidad de este m´etodo para segmentar y estimar la postura de peatones, independientemente de la direcci´on del movimiento. Tambi´en han demostrado que el m´etodo responde muy adecuadamente a cualquier cambio de direcci´on durante la secuencia.
xiv A pesar de que ha sido probado ´unicamente con el movimiento de andar, el enfoque presentado es gen´erico y puede aplicarse a cualquier otra acci´on. Para ello, ser´ıa necesario disponer de una base de datos m´as amplia, con m´as movimientos y mediante un modelo gr´afico 3D se podr´ıa sintetizar autom´aticamente un conjunto de entrenamiento de representaciones 2D y 3D. En esta tesis se ha proporcionado una forma de transici´on entre varias vistas de una misma acci´on. De igual modo, podr´ıan considerarse transiciones entre sub-espacios de diferentes actividades. Parte II: localizaci´on y seguimiento de m´ultiples sujetos mediante una c´amara de vigilancia fija En la segunda parte de la tesis hemos combinado los mejores componentes de sistemas de seguimiento de posturas humanas y nos hemos aprovechado de la geometr´ıa proyectiva para desarrollar, en escenas calibradas de vigilancia, un eficiente sistema de seguimiento de posturas 3D basado en el filtro de part´ıculas . Por medio de la geometr´ıa proyectiva hemos reemplazado la transformaci´on de similitud 2D, com´unmente empleada para relacionar los planos de la imagen y del modelo, por una alineaci´on basada en homograf´ıa. Asimismo, hemos propuesto un eficiente c´alculo de la probabilidad de aparici´on de cada part´ıcula bas´andonos ´unicamente en los bordes de la imagen y la sustracci´on respecto al fondo, resultando en un emparejamiento r´apido de la silueta humana. Tambi´en hemos introducido un nuevo estimador de Estado a partir del conjunto de part´ıculas. La eficiencia de nuestro algoritmo ha sido demostrada mediante el procesamiento de un conjunto de v´ıdeos de vigilancia particularmente dif´ıciles presentando una evaluaci´on num´erica para 2784 posturas manualmente etiquetadas que ser´an puestas a disposici´on de la comunidad cient´ıfica para futuras investigaciones. Los experimentos muestran que nuestro sistema es capaz de seguir correctamente varios peatones y estimar sus posturas 3D en casos de movimiento en grupo, con oclusiones y sombras. Como trabajo futuro planteamos las siguientes l´ıneas: 1. Una vez calibrada la c´amara respecto a la escena, la c´amara no puede moverse, lo que supone una limitaci´on de la propuesta. Se podr´ıa considerar un m´etodo autom´atico de detecci´on de los puntos de fuga que permitiera calcular las homograf´ıas de forma completamente autom´atica. 2. Para tratar el tema del seguimiento de m´ultiples sujetos interactuando hemos reponderado y reducido la influencia de las muestras usando un simple enfoque de ocupaci´on 3D, que ha demostrado ser eficaz con los v´ıdeos procesados en esta tesis. El problema del tema del seguimiento de m´ultiples sujetos en situaciones m´as complejas no entraba dentro de los objetivos de esta tesis y el uso de un filtro para varios objetos o un modelado m´as adecuado de las interacciones se plantean como trabajo futuro. 3. A´un que todos los experimentos son espec´ıficos para la actividad de caminar (debido a la mayor disponibilidad conjuntos de datos de entrenamiento y evaluaci´on), nuestro sistema es suficientemente general para extenderse a otras actividades. La baja dimensionalidad del espacio de b´usqueda, combinada con un limitado n´umero de vistas de entrenamiento, hacen que nuestro trabajo sea f´acilmente ampliable a m´as acciones y hace m´as factible el desarrollo de software de reconocimiento de actividades para aplicaciones de vigilancia reales. Para diferentes acciones podr´ıa aprenderse un modelo de baja dimensi´on, utilizando un mapeo para modelar las conmutaciones entre actividades. 4. De cara a encontrar la soluci´on ´optima en cada instante de tiempo en un modelo basado en el filtro de part´ıculas, se podr´ıan emplear t´ecnicas de b´usqueda por gradiente o un estudio
xv m´as complejo de la distribuci´on posterior. La adaptaci´on de nuestro m´etodo proyectivo para entornos no-calibrados ofrece otra l´ınea interesante para futuras investigaciones. Parte III: localizaci´on y estimaci´on de posturas en secuencia de v´ıdeo capturada con una c´amara en movimiento o simplemente en una sola imagen est´atica En la tercera parte de esta tesis, hemos abordado el problema conjunto de detecci´on de personas y estimaci´on de su postura formul´andolo como un problema de clasificaci´on. Hemos seguido una t´ecnica de ventana deslizante para localizar y clasificar posturas humanas mediante un r´apido clasificador multi-clase que combina los mejores componentes de clasificadores existentes incluyendo ´arboles jer´arquicos, cascadas o bosques aleatorios. Hemos validado nuestro m´etodo con una evaluaci´on num´erica para 3 niveles diferentes de an´alisis: detecci´on y localizaci´on de personas en im´agenes, clasificaci´on de posturas humanas y estimaci´on de la postura (con localizaci´on de las articulaciones del cuerpo). Si el espacio de b´usqueda (ubicaci´on en la imagen y escala) puede reducirse, por ejemplo, usando un algoritmo de seguimiento o limitando la distancia a la c´amara en una interfaz hombre-m´aquina, nuestro m´etodo puede actuar en tiempo real. Para mejorar el algoritmo planteamos varias alternativas: 1. El m´etodo actual de selecci´on de caracter´ısticas requiere una gran cantidad de im´agenes de entrenamiento alineadas. El trabajo futuro debe centrarse en desarrollar un algoritmo de aprendizaje que pueda manejar im´agenes d´ebilmente etiquetadas o peque˜nos conjuntos de entrenamiento. 2. Una implementaci´on computacionalmente eficiente del descriptor HOG, utilizando por ejemplo una implementaci´on para GPUs, podr´ıa acelerar la detecci´on a´un m´as. 3. Nuestro algoritmo busca una distribuci´on sobre las clases, una direcci´on interesante de trabajo futuro ser´ıa combinar esta b´usqueda con alg´un tipo de regresi´on con objeto de encontrar un m´etodo computacionalmente eficiente y m´as preciso en t´erminos de estimaci´on de postura. 4. Finalmente, este trabajo abre varias otras l´ıneas interesantes de trabajo futuro: por ejemplo, se podr´ıa intentar combinar diferentes tipos de caracter´ısticas (color, profundidad, etc.) dentro de nuestro clasificador, extender el algoritmo para una m´as amplia gama de posturas y acciones o aplicar el algoritmo a otros problemas m´as generales de aprendizaje autom´atico.
xvi
Contents 1 Introduction 1 1.1 Human Pose Analysis and its Applications ...................... 1 1.1.1 Surveillance Applications ........................... 2 1.1.2 Control Applications .............................. 3 1.1.3 Analysis Applications ............................. 4 1.2 The Problem and its Difficulties ............................ 5 1.2.1 The Problem .................................. 5 1.2.2 The Difficulties ................................. 5 1.3 State of the Art ..................................... 7 1.3.1 Detection .................................... 8 1.3.2 Pose Estimation ................................ 9 1.3.3 Tracking ..................................... 11 1.4 Goals and Hypotheses ................................. 13 1.5 Contributions and Organization ............................ 13 1.5.1 Part I: Segmentation and Pose Estimation with a Static Camera . . . . . 14 1.5.2 Part II: Pose Tracking in Video-Surveillance Environments ........ 15 1.5.3 Part III: Pose Estimation with a Moving Camera or in Static Images . . . 16 1.5.4 Part IV ..................................... 16 1.6 List of Relevant Publications ............................. 16 1.6.1 Part I ...................................... 16 1.6.2 Part II ...................................... 17 1.6.3 Part III ..................................... 17 1.6.4 Miscellaneous: ................................. 17 I Segmentation and Pose Estimation with a Static Camera 19 2 Dealing with Non-linearities in Shape Modeling 21 2.1 Introduction ....................................... 21 2.1.1 Related Work .................................. 21 2.1.2 Overview .................................... 23 2.2 Shape-Skeleton Training Database .......................... 23 2.2.1 Training Database Construction ....................... 23 2.2.2 Shapes Normalization ............................. 24 2.2.3 Shape-Skeleton Eigenspace - PCA Model .................. 24 2.3 View-based Shape-Skeleton Gaussian Mixture Models ............... 25 2.3.1 Structural clustering .............................. 25 2.3.2 View-based Non-Linear Models ........................ 29
xviii CONTENTS 2.3.3 Joint Estimation of Shape and Skeleton ................... 29 2.4 Non-Linear Models Testing .............................. 30 2.5 Conclusions ....................................... 32 3 A Framework of Spatio-temporal Models 33 3.1 Introduction ....................................... 33 3.1.1 Overview of the work ............................. 33 3.2 Framework Construction ................................ 35 3.3 Constraint-based search ................................ 37 3.3.1 Markov Chain for Modeling Temporal Constraint .............. 37 3.3.2 Modeling Spatial Constraint .......................... 37 3.3.3 Combining Spatial and Temporal Constraints ................ 38 3.4 Joint Segmentation and Pose Estimation ....................... 39 3.5 Experiments ....................................... 40 3.5.1 Framework Validation ............................. 41 3.5.2 Numerical Evaluation with HumanEva dataset ............... 45 3.6 Summary ........................................ 46 II Pose Tracking in Video-Surveillance Environments 49 4 View-invariant Motion Analysis using View-based Models 51 4.1 Introduction ....................................... 51 4.1.1 Related Work .................................. 52 4.1.2 Motivation and Overview of the Approach. ................. 54 4.2 Geometrical Considerations in Man-Made Environments .............. 55 4.2.1 Notations .................................... 55 4.2.2 Camera and Scene Calibration ........................ 56 4.3 Projection Image-Training View Through a Vertical Plane ............. 57 4.3.1 Qualitative Results using Ground Truth Data ................ 62 4.4 View-invariant Pose Analysis ............................. 63 4.5 Experiments ....................................... 65 4.5.1 Numerical Evaluation of the Effect of Noise ................. 65 4.6 Conclusions ....................................... 67 5 View-invariant 3D Pose Tracking 69 5.1 Introduction ....................................... 69 5.1.1 Related Work .................................. 69 5.1.2 Overview .................................... 70 5.2 Torus Manifold for Pose and Appearance Modeling ................. 71 5.2.1 Body Pose Modeling .............................. 71 5.2.2 Appearance Modeling ............................. 72 5.3 Recursive Bayesian Sampling ............................. 73 5.3.1 Formulation. .................................. 73 5.3.2 Dynamic Model. ................................ 74 5.3.3 Image Measurements - Observation Model. ................. 74 5.3.4 Tracking Multiple Pedestrians. ........................ 77 5.4 Experimental Results .................................. 80 5.4.1 Settings and Parameters. ........................... 81 5.4.2 Experiments. .................................. 81
CONTENTS xix 5.5 Conclusions ....................................... 88 III Pose Estimation with a Moving Camera or in Static Images 91 6 Multi-class Pose Classifier 93 6.1 Introduction ....................................... 93 6.1.1 Related Previous Work ............................ 93 6.1.2 Motivation and Overview of the Approach .................. 94 6.2 Sampling of Discriminative HOGs .......................... 96 6.2.1 Formulation ................................... 97 6.3 Randomized Cascades of Rejectors ..........................100 6.3.1 Bottom-up Hierarchical Tree Construction ..................100 6.3.2 Discriminative HOG Blocks Selection ....................102 6.3.3 Randomization .................................105 6.4 Class Definition .....................................107 6.4.1 HumanEVA Dataset ..............................107 6.4.2 MoBo Dataset .................................109 6.5 Experiments using MoBo Dataset ...........................112 6.6 Conclusions .......................................116 7 Human Localization and Pose Estimation 119 7.1 Introduction .......................................119 7.2 Human Pose Detection .................................119 7.3 Properties of the Random Cascades Classifiers for Pose Detection .........121 7.4 Experiments .......................................123 7.4.1 Localization Results using Mobo Dataset for Training ...........123 7.4.2 Pose Estimation in Video Sequences .....................134 7.5 Conclusions .......................................136 IV Discussions and Conclusions 139 8 Conclusions 141 8.1 Introduction .......................................141 8.2 Conclusions .......................................141 8.2.1 Modeling ....................................141 8.2.2 Detection ....................................142 8.2.3 Tracking .....................................143 8.3 Main Contributions and Future Lines of Work ....................143 8.3.1 Part I ......................................144 8.3.2 Part II ......................................144 8.3.3 Part III .....................................145 V Appendix 147 A Projection to a Vertical Plane 149
xx CONTENTS
List of Figures 1.1 Surveillance applications ................................ 2 1.2 Human computer interfaces .............................. 3 1.3 Pose analysis applications ............................... 4 1.4 Examples of human pose and appearance modeling ................. 10 1.5 Testing environments studied in this thesis ..................... 13 1.6 Viewpoint parameterization and training views ................... 14 2.1 MoBo training database ................................ 24 2.2 Gait cycles in pose eigenspace ............................. 26 2.3 Structural clustering in pose eigenspace ....................... 27 2.4 Gait cycle phases .................................... 28 2.5 View-based mixtures of PCA ............................. 29 2.6 Principal modes of variation of the gaussian models for each view-based GMM . 30 2.7 Non-linear models testing ............................... 31 2.8 Caviar sequences selected for testing ......................... 31 3.1 Framework of spatio-temporal shape-skeleton models ................ 34 3.2 Training views considered for the framework construction ............. 35 3.3 Multi-view Gaussian mixture model represented in the pose-silhouette eigenspace 36 3.4 Toroidal transition matrix ............................... 37 3.5 Shape model fitting ................................... 39 3.6 Examples of sequences processed with the framework of pose-silhouette models . 41 3.7 Challenging frames processed with the framework of pose-silhouette models . . . 41 3.8 Results obtained for the outdoor “Walkcircle” sequence .............. 42 3.9 Results obtained for the indoor “Elevator” sequence ................ 43 3.10 Segmentation results .................................. 44 3.11 Pose estimation results ................................. 45 3.12 Segmentation and 2D pose estimation obtained for the 4 HumanEva testing sequences ........................................ 46 3.13 Numerical results obtained for the 4 HumanEva testing sequences ......... 47 4.1 Training and testing video sequences considered in the chapter .......... 52 4.2 Viewpoint parameterization .............................. 53 4.3 Tracking system diagram ............................... 55 4.4 camera and scene calibration ............................. 57 4.5 Examples of projection on vertical planes ...................... 58 4.6 Schematical representation of the transformation between training and testing images through a vertical plane ............................ 59
xxii LIST OF FIGURES 4.7 Projection to vertical plane .............................. 61 4.8 Examples of projections to training planes ...................... 62 4.9 View invariant tracking based on head detector ................... 64 4.10 View-based shape registration and pose estimation ................. 65 4.11 Example of a processed sequence ........................... 66 4.12 Effect of noise on 2D pose estimation ......................... 68 5.1 Tracking system flowchart ............................... 71 5.2 Training 3D pose and shape and the pose-viewpoint torus manifold ........ 72 5.3 Likelihood function as a Laplacian distribution over a distance .......... 76 5.4 Foreground likelihood ................................. 76 5.5 Likelihood values in 2 sequences ............................ 77 5.6 Visualization of the likelihood for the sample set in several frames ........ 78 5.7 Percentage of lost tracks vs number of particles for similarity and homographic alignment ........................................ 82 5.8 Tracking results: percentage of valid localizations and 2D pose errors ....... 84 5.9 Detailed performances w.r.t. the distance to the camera .............. 85 5.10 Average 2D pose error varying the state estimator ................. 86 5.11 Qualitative 3D pose tracking results ......................... 87 5.12 Qualitative 3D pose tracking results for sequences with two interacting subjects . 88 5.13 Qualitative 3D pose tracking results for sequences with three interacting subjects 88 6.1 Random Forest preliminary results .......................... 95 6.2 Log-likelihood ratio for human pose ......................... 96 6.3 Log-likelihood ratio for face expressions ....................... 98 6.4 Selection of discriminative HOG for facial expressions classification ........ 99 6.5 Bottom-up hierarchical tree construction .......................100 6.6 Example of a selected HOG block ...........................103 6.7 Rejector branch decision ................................105 6.8 Diagram of image classification using an ensemble of ncrandomized cascades . . 107 6.9 Alignment of training images .............................108 6.10 Class definition - HumanEVA .............................109 6.11 Image classification using an ensemble of randomized cascades ...........110 6.12 Modified MoBo dataset ................................111 6.13 Class definition for MoBo dataset ...........................112 6.14 Set of classes obtained for the MoBo dataset ....................113 6.15 Pose classification results on MoBo dataset for PSH, Random Forest (1000 trees) and SVMs classifiers ..................................114 6.16 Pose classification rates on MoBo datset for 4 different ensembles .........114 6.17 Pose classification experiments on MoBo using randomized cascades .......116 7.1 Cascade classifier response for a dense scan using different ensembles . . . . . . . 120 7.2 Detection score with on-line randomization .....................122 7.3 Localization dataset ..................................123 7.4 HOG feature scaling ..................................124 7.5 Detection results on the localization dataset using different classifiers .......125 7.6 Detection results for different configurations of the cascades classifiers with a 3rd pass of hard negative retraining ............................127 7.7 Detection time for different configurations of the cascades classifiers after a 3rd pass of hard negative retraining ............................127
1 Introduction “Computer vision is the science and technology of machines that see.” Computer vision refers to a recent but broad field of research whose goal is to make the computers understand what they see. At the crossroads between computer sciences, electrical engineering and mathematics, it is closely related to other research areas such as pattern recognition or machine learning, and can be seen as a branch of the Artificial Intelligence (AI) field. The main concern of computer vision is to address the theory and technology for building artificial systems that can obtain information from images. The frames of a video sequence, the views from several cameras, or the multi-dimensional information from a medical scanner are some examples of the multiple forms that image data can take. Over the last decades, the rapid progresses in imaging sensor technologies, the advances in data storage and transmission, and the exponential growth of computational power have all contributed to convert video into an ubiquitous and unavoidable media in modern life. The increasing quantity of available data has led to the recent interest for computer vision because of the possibilities offered by automatic video analysis. Indeed, computer vision systems should be able to automatically and quickly extract information from an image, or a sequence of images, and generate a description of the objects and actions observed in the scene. Among all the possible objects observable in a video sequence, humans are of special interest since they play a major role in many activities. The part of computer vision dedicated to humans is usually known as human motion analysis. It intends to extract information such like presence/absence, position, posture, behaviors or activities from a single or multiple images. The monocular analysis of human motion presents a wide range of potential applications and is consequently one of the most active research areas in the field. 1.1 Human Pose Analysis and its Applications Full-body human pose analysis from monocular images constitutes one of the fundamental problems in Computer Vision as shown by the recent special issue of the International Journal of Computer Vision [Sigal and Black,2010]. It has a wide range of potential applications which can be classified into 3 main categories: surveillance, control and analysis.
2 Chapter 1. Introduction 1.1.1 Surveillance Applications Surveillance applications are certainly the first applications that come to mind when talking about human motion analysis. In recent years, the number of cameras deployed for surveillance and safety in urban environments, such as streets, airports, subways, train stations or commercial center, has increased considerably. The main reasons are the potential terrorist menaces and, certainly, their falling cost. While a constant and effective human monitoring of these numerous cameras seems difficult to achieve, automatic video understanding systems could enable a single operator to monitor many cameras and control wide areas more reliably. Such applications could, for instance, detect abnormal activities (fights, thefts or left-luggage) and provide timely alarm. They could also help to track eventual suspects, count people or analyze crowd flow. Some people even proposed to use gait analysis as biometric for people identification. A computer vision application can be considered to automate the acquisition and analysis of consumer shopping behavior in commercial centers. Taking full advantage of these video-surveillance networks, the processing of hours and hours of consumer behavior could provide a valuable feedback to the retailers with scarcely any additional costs. Finally, some researchers have been working on home care applications for elderly people. Population aging in developed countries causes changes in living arrangements resulting in increasing number of older people living alone. Yet vision-based home care systems have tended to focus on those elderly living alone, the main goal being to detect abnormal situations - mainly fall detection - and call for assistance when required. As the reader will notice soon, this last example provides a good transition with the next category of applications. Figure 1.1: Surveillance applications: (upperleft) results from research project ADVISOR (Annotated Digital Video for Surveillance and Optimised Retrieval) in which is developed an integrated visual surveillance and behaviour analysis system based on the people tracker from [Baumberg,1995]. (upper right) Multi-camera surveillance system in a shop. (bottom) A camera system for fall detection of elderly people living alone in their own private homes (image source: http://www.mobilab-khk.be).
1.1 Human Pose Analysis and its Applications 3 1.1.2 Control Applications These applications refer to the Human-Computer Interfaces, or Human-Computer Interaction (HCI), where the user interacts with a machine and controls it through particular gestures and movements. Video-games like EyeToy1or Kinect2and virtual reality are some of the few examples that one can encounter nowadays but new paradigms for interacting with computers are being investigated. They will enable in the future to communicate and interact effortlessly and intuitively with the machines. To interact seamlessly with people, HCI systems will need to understand their environment through vision and auditory sensors, and should learn how to adapt themselves and intelligently respond depending on the context. One can remember the HAL 9000 Computer, the non-human and central character in the futuristic film by Stanley Kubrick and Arthur C. Clarke - 2001: A Space Odyssey. Its purpose is to watch over the space craft “Discovery I” by controlling all of the relevant ship’s functions and by monitoring each event happening on the ship. HAL is able to localize the craft’s members, identify them, know their emotions and recognize their activities. This constitutes a good example of what an Intelligent Environment could be even if, in this case, perhaps HAL becomes a bit too smart... Figure 1.2: Human-Computer Interfaces: (upper left) Sony Eyetoy video game. (upper right) Family playing a MicroSoft Kinect video game. (bottom left) Example of a human computer interface using gesture recognition in the science fiction movie Minority Report. (bottom right) Virtual reality application. 1Sony’s EyeToy is a small digital camera that sits on top of the TV and plugs into the Sony’s PlayStation. The motion sensitive camera films the player as he stands in front of his TV, putting his image on screen in the middle of the action. The player can then use any part of his body to play the games. 2Microsoft Kinect, released in late 2010, is an interactive game which can track the joints of up to two people simultaneously using a 3D time-of-flight camera.
4 Chapter 1. Introduction 1.1.3 Analysis Applications This part basically regroups all the applications that do not belong to any of the previous two categories. First, there are a non-negligible number of existing applications which require motion capture data 3. These applications include gait analysis and diagnostic of orthopedic patients, study of athletes’s movements and performances, or character animation and special effect making for the film industry. Until now, the estimation of the motion has been performed by motion capture systems that provide the 3D position/orientation of a set of markers using mechanical, electro-magnetic, acoustic or optical features. One can easily imagine how these applications could perform, in a similar way, using motion capture data from computer vision systems. Such video-based marker-less systems would present the advantage to be much cheaper and less invasive than the current marker based ones. Moreover, no specific hardware would be required since video from ordinary cameras could be used as input. Some “content-based” applications also appear in this same category. They basically require the extraction and interpretation of the information present in the images for later processes. Automatic video annotation/indexing for fast content-based retrieval and image/video compression for efficient data storage and transmission are some clear examples. Finally, research is being done by the car industry for safety applications: inside the car, by controlling the driver behaviors (sleeping detection, attention control or hand position on the steering wheel), and outside the car to prevent eventual accidents by detecting pedestrians crossing the street. Figure 1.3: Pose analysis applications: (upper left) Diagnostic of an orthopedic patient with an optical motion capture system. (upper center) Performance analysis in golf. (upper right) Character animation. (bottom left) Sport event analysis (image source: http://www.cs.ubc.ca) (bottom right) Automatic video annotation. 3Motion capture is the process of recording real life movement of a subject and converting it into usable mathematical terms by tracking a number of key points in 3D space over time.
1.2 The Problem and its Difficulties 5 1.2 The Problem and its Difficulties As well as presenting a wide spectrum of potential applications, the monocular analysis of human motion also constitutes one of the fundamental but unsolved problems of computer vision. Even if great advances have been achieved with the recently released MicroSoft Kinect, it is still an open problem as this videogame assumes a front-facing subject at a restricted distance from the TV. Moreover, the use of a 3D time-of-flight camera makes the system very sensitive to adverse lighting conditions. In this thesis, we focus on monocular sequences captured by a standard 2D video camera. By restricting the problem to a monocular observation - supposing that only one camera view is available - we make it more attractive both in terms of research and development. Indeed, extracting 3D information from 2D data is a challenging theme of research while, for practical consideration, it is often the case where only a single camera is available. In this section, a brief overview of the problem will be given as well as a detailed description of the involved difficulties. 1.2.1 The Problem The problem addressed in this thesis is the consecutive detection, pose estimation and tracking of humans in monocular sequences, letting for future work the eventual recognition and interpretation of the observed activities. The first step in almost every system of vision-based human motion analysis is the human detection. It aims at localizing the possible humans present in the scene. The goal of the segmentation step is to extract the regions corresponding to detected people from the rest of the image. These two first steps can be associated in a foreground segmentation algorithm that basically detects the region belonging to some eventual moving objects. The pose estimation task then relies on recovering the joints position (or angles) from image features while the tracking stage brings some temporal consistency by establishing coherent relations between consecutive frames of a video sequence. Location in the image, position on a map or even the posture of the observed subject are some possible features that can be tracked over time. When an estimation can not be made robustly from direct evaluation of image features (e.g. noisy measurement), the tracking stage can help to solve pose ambiguities by using the estimates from the previous frames. Although the work is focused on sequence processing, the frames are considered to be available one at a time, in contrast to batch approaches that estimate human state at any given time using all the available images, prior and posterior to that time step. Such hypothesis constrains the tracking in some sense but allows a real-time implementation of the algorithms. In [Zhang,2006], the author divides the human body analysis into 3 sub-categories depending on the relative distance between camera and subject. It is in fact relative to the size of the humans in the image and the quantity of information available. The three sub-categories are far, medium and near fields. In this thesis, only an extended medium field will be considered. Depending on the application, i.e. typical video-surveillance or motion capture scenario, the full body of the observed humans will be from 30 to 300 pixels tall. Obviously, the expected accuracy in terms of segmentation, pose estimation or tracking will strongly depends on the application and the available image resolution. 1.2.2 The Difficulties The monocular analysis of human motion involves a series of difficulties that are listed and discussed below. Some are due to the hypothesis of a monocular observation, others are inherent
6 Chapter 1. Introduction to the problem of observing a human and others are somehow generic to any vision-based problem. 2D-3D Projection and Depth Information: the projection of the 3D world into a 2D image suppresses the depth information. This is one of the fundamental problem of computer vision. Part of the problem can be solved by considering various camera views adequately located. In monocular case, the depth information lost during the 2D projection must be replaced by learned knowledge information about the 3D scene and the objects (their shape, structure and dynamics). Indeed, when a human looks at a 2D picture or only sees with one eye, he/she is still able to interpret the scene and recognize millions of objects and can even accurately estimate their 3D position and 3D pose from this 2D projection. This is made possible thanks to the learned knowledge his/her brain has accumulated over the years. As stated in [Bowden,1999], providing a similar knowledge of a small subset of objects to a computer is the premise of model based vision. High Variability in Shape, Appearance and Pose: because of its articulated nature, the human body can take a very high number of different poses that directly introduce a high variability in people shape and appearance. The considerable variability of both shape and appearance observable in images also results from other parameters that must be taken into account: •shape and appearance may vary dramatically from person to person, depending on the morphology (short vs tall, fat vs slim) and clothing type (loose vs fitted clothes, dress vs pants). •shape and appearance of a same subject can also vary over time because of clothing and illumination changes. •finally, the shape and appearance of a same subject in the same pose usually present some remarkable differences when observed from different camera viewpoints. Dimensionality. Human motion analysis, and more particularly human pose estimation and tracking, raises the problem of what really needs to be estimated. What is the optimum representation of a human body that a computer can learn and accurately estimate? The simplest 3D modeling of the human body relies on representing the limbs as rigid elements connected to each other at the joints that are then parameterized by their angles or 3D position. Such limited representation explains reasonably well the human motion but in any case allows to recover the real flexibility of the body. Even so, it still requires a minimum of 30 parameters! The previous remark means that pose estimation has to be computed in a parameter space whose dimension is equal or higher than 30. Additionally, the previously emphasized high variability in shape and appearance introduces more complexity in modeling and estimation that rapidly make the problem computationally unsolvable. Some compromises must be made when representing the human body - appearance, pose and shape - to balance modeling complexity and computational feasibility. Non-linearity. As demonstrated before, a feature space resulting from human pose, shape and/or appearance would present an extremely high dimensionality. It would also show a relatively high nonlinearity, mainly caused by the complex rotational deformations inherent to the articulated structure of the human body. In such conditions, modeling, estimation and tracking are much more challenging than with a linear problem. Occlusions are most of the time a direct consequence of the monocular observation. Two types of occlusions may occur along a sequence. The first ones correspond to the occlusions of the observed subject, complete or local, provoked by some other elements of the scene that are located between the subject and the camera. Such occluding objects can be immobile - a wall,
1.3 State of the Art 7 a tree or a post - or mobile, like a moving vehicle, an interacting subject or a simple passer by. The second category of occlusions refers to the self-occlusions. Frequently observed, they are inherent to the articulated nature of the human body. In those cases, some parts of the body are basically not visible in the image because they are occluded by some other parts of the body itself. Note that, additionally, permanent occlusions can be due to the type of clothes the observed subject is wearing: for instance, one can easily imagine how the legs can be hidden by a long coat or a long dress. Once more, thanks to the learned knowledge about human body shape and structure, model based vision should be able to solve part of the occlusion issue. Image clutter, shadows and motion blur can be seen as the noise introduced by the environment settings, respectively the background, the lighting sources and the camera type. First, image clutter refers to the distracting objects and structures in the background that can present some similarities with part of the subject in terms of shape or appearance. For instance, the position and pose of the subject’s legs would be much more difficult to estimate if the color of the pants is similar to the color of the background or if there is an object with a similar shape. Those distracters can introduce uncertainty during the different steps of the motion analysis process. Cast shadows, reflections or lighting changes are some of the possible artifacts that the lighting conditions can provoke. They can be interpreted as moving objects and can also distract the algorithms from their target. Finally, motion blur appears when the shutter time of the camera is too long compared to the velocity of the captured motion. It introduces an additional difficulty when analyzing relatively fast motions. In controlled environment, the different noise effects described before can be minimized by selecting an adequate background (cf blue screening technique employed by the film industry 4), some optimal light sources and a camera adapted to the problem. In that case, the problem is substantially easier to solve since the researchers can focus their efforts on the remaining difficulties. But when the environment is imposed and not controlled by the operator or user, as in many video-surveillance or video game scenarios, some vision based solutions must be found to deal with the resulting noise. 1.3 State of the Art The large number of related papers in journals and conferences, the number of special issues in journals and the numerous workshops dedicated to human motion capture and analysis demonstrate that it is a very active area of research in computer vision. In this thesis, various different aspects of the human pose analysis will be dealt with and our work is closely related to several very active lines of research in human motion analysis such as shape-based analysis (part I and II), view-invariant understanding (part II), monocular 3D body pose tracking (part II), human detection, tracking-by-detection (part III). There has been a significant number of interesting papers dedicated to each one of them. Listing and explaining in this chapter all these publications would not be very useful to the reader. Instead, we will give now a brief overview of the current state of the art, explaining only the few main methodologies for each step of the human motion capture and pose analysis. We will present the global strategies, their advantages and drawbacks, and group the efforts that consider similar underlying assumptions. Then, at the beginning of each chapter, a short but detailed section will be dedicated to the previous works directly related to the line of research 4Blue-screening is the process of removing blue from a scene. An actor stands in front of a blue background. Then the blue is replaced with another scene or picture.
8 Chapter 1. Introduction investigated in that chapter. In this way, we aim at helping shed light on previous research, topic by topic, as well as positioning our own work in the context of related approaches. Readers interested in seeing an extensive description of all the different relevant works dedicated to human motion capture and analysis are invited to read the numerous review papers including [Gavrila,1999,Moeslund and Granum,2001,Wang et al.,2003,Moeslund et al.,2006,Poppe,2007] or the recent book by [Moeslund et al.,2011]. 1.3.1 Detection As mentioned above, the detection stage aims at localizing the possible humans present in the scene. A reliable and robust detection is essential since all the following steps strictly depend on it. An efficient human detector is very helpful for providing reliable inputs to both tracking and pose estimation algorithms. Indeed, it makes plausible the numerous pose estimation works which just consider that the bounding box of the subject is provided. In our opinion, most of the human detection techniques can be grouped into two main categories. In the first class, we can find the techniques that separate the foreground from the background and classify the resulting detected objects as human or non-human. The other principal category corresponds to the approaches which basically extract multiple windows of pixels from the image (varying location, scale and eventually rotation) and classify them as human or not. 1.3.1.1 Foreground Detection and Classification Many previous works proposed to use a foreground segmentation and object classification to detect eventual humans in the scene. [Baumberg and Hogg,1994,Haritaoglu et al.,2000,Isard and MacCormick,2001,Elgammal et al.,2002,Siebel and Maybank,2002a,Zhao and Nevatia, 2004,Orrite-Uru˜nuela et al.,2004] are some few examples. Most of the time, a background model is used to estimate the foreground regions: pixels that do not correspond to the background model are considered as foreground. The resulting “blobs” are then classified as human or not, based on different criterions, mainly their size and shape. This method consequently supposes that the scene and the possible size of the humans in the image are known. In controlled environment, this method provides a very good estimation of the human segmentation. Moreover, because of its relative ease of implementation and its quick computation time, it is usually preferred for real-time applications such as video-surveillance or gaming (Eyetoy). Even so, this method suffers from serious drawbacks. First of all, it is extremely sensitive to illumination artifacts like shadows, lighting changes and reflections. It also requires a fixed camera and is not directly applicable to moving camera applications (or static isolated images). Finally, the binary classification alone - background vs foreground - is not always a sufficient feature for solving partial occlusions between multiple moving objects. In other words, when two subjects are too close, their segmentations are merged into a unique “blob” that can be difficult to analyze. 1.3.1.2 Scanning for Humans There is an extensive and recent literature on this second type of human detection, including [Dalal and Triggs,2005,Viola et al.,2005,Wu et al.,2005,Zhu et al.,2006,Dimitrijevic et al.,2006,Wu and Yu,2006,Gavrila,2007,Enzweiler and Gavrila,2009,Breitenstein et al., 2011,Gall et al.,2011,Sabzmeydani and Mori,2007]. These methods usually scan the image, classifying each extracted box as human or non-human based on a learnt knowledge of the human appearance. Sometimes, a coarse to fine search is considered to run the algorithm in real time as in [Gavrila,2007]. These cited works mainly focus on pedestrian detection, i.e.
1.3 State of the Art 9 standing and walking people. Recently, Zhang et al proposed in [Zhang et al.,2007b] a novel model-based approach to detect humans performing different activities and showing complex postures. All these detection techniques present some non-negligible advantages. For instance, they can work with isolated images as well as with sequences and do not require any previous knowledge about the scene. They can also work with moving cameras [Dimitrijevic et al.,2006] or with a camera mounted on a moving vehicle [Gavrila,2007]. Even if recent work has focused on training these classifiers with weakly labelled data, their main drawback is the large data sets of manually labelled images that are needed to train them. Additionally, complex and long learning phases are usually required to select the relevant features that will allow to discriminate humans from cluttered background. Another drawback is that these techniques need to be very quick as they usually have to classify thousands of windows for each classified image. In [Gavrila,2007], Gavrila presents a probabilistic approach to hierarchical, exemplar-based shape matching. This method achieves a very good detection rate and real time performance. 1.3.2 Pose Estimation There are hundreds of papers dedicated to human pose estimation and motion capture, but there are basically two main strategies: methods which employ a datase of training exemplars, i.e. training images and the corresponding 2D/3D poses, and approaches that use a hand-made kinematic model of the human body. 1.3.2.1 Kinematic Models Many efficient systems are based on the use of a model which is, most of the time, a representation of the human body. The selection of the appropriate model is a critical issue and the use of an explicit body model is not simple, given the numerous degrees of freedom (DOF) of the human body. Kinematic Models attempt to describe the human body structure as a kinematic chain of segments (the limbs) connected by joints, each joint being parameterized by a series of degrees of freedom (translation or rotation). These models can be represented in 2D or 3D. To match the pose with the image, the individual limbs have been modelled as layered patches in the 2D image plane [Yacoob and Black,1999,Yacoob and Davis,2000,Agarwal and Triggs,2004], or in the 3D world as stick figure [Taylor,2000], cylinders [Sidenbladh et al.,2000,Bregler et al.,2004,Sigal et al.,2004], truncated cones [Deutscher et al.,2000,Deutscher and Reid, 2005], superquadrics [Gavrila and Davis,1996,Sminchisescu and Triggs,2003] or individually deformable shape models [Kakadiaris and Metaxas,2000,Plankers and Fua,2001]. In all these works, the authors make use of generative models for pose tracking, modeling the human kinematics more than the human appearance. Those methods do not explain the image evidence and do not always fit the image accurately which sometimes causes artifacts on joints estimation. Nowadays, exemplar based approaches are usually considered as they better represent the appearance of the human and closer match to image observation than kinematic models alone. 1.3.2.2 Exemplar based approaches Exemplar-based approaches have been very successful for human pose estimation and tracking. Some consist of comparing the observed image with a data base of stored samples as in [Shakhnarovich et al.,2003,Mori and Malik,2006,Ong et al.,2006]. In some other cases, the training examples are used to learn a mapping between image feature space and 3D pose
10 Chapter 1. Introduction Figure 1.4: Examples of human pose and appearance modeling: (from left to right) 3D stick figure from [Taylor,2000], 2D shape and 3D skeletal structure PDM from [Bowden et al.,2000], articulated 2D shape from [Zhang et al.,2005a], 3D truncated cones from [Deutscher et al.,2000], superquadrics from [Sminchisescu and Triggs,2003] and 3D mesh model from [Balan et al.,2007]. space [Agarwal and Triggs,2006,Elgammal and Lee,2009,Lee and Elgammal,2010,Jaeggli et al.,2009]. Such mappings can be used in a bottom-up discriminative way [Sminchisescu et al., 2007] to directly infer a pose from an appearance descriptor or in a top-down generative manner [Jaeggli et al.,2009] through a framework (e.g. a particle filter) where pose hypotheses are made and their appearances aligned with the image to evaluate the corresponding observation likelihood or cost function. The exemplars can also be used to train multi-class pose classifiers [Okada and Soatto,2008,Andriluka et al.,2010] or part-based detectors [Ramanan et al.,2007, Wu and Nevatia,2009,Bourdev et al.,2010,Felzenszwalb et al.,2010b,Lin and Davis,2010] that are later employed to scan images. Nearest Neighbor Search techniques have been very successful for pose recognition. This method consists in comparing the observed image with a data base of samples as in [Athitsos and Sclaroff,2003,Orrite-Uru˜nuela et al.,2004,Mori and Malik,2006,Ong et al.,2006] and find the most similar exemplar in the training set. However, in a scenario involving a wide range of viewpoints and poses, a large number of exemplars would be required. As a result, the computational time would be very high to recognize individual poses. One approach, based on efficient nearest neighbors search using histogram of gradient features, addressed the problem of quick retrieval in large set of exemplars by using Parameters Sensitive Hashing (PSH) [Shakhnarovich et al.,2003], a variant of the original Locality Sensitive Hashing algorithm (LSH) [Datar et al.,2004]. The final pose estimate is produced by applying locally-weighted regression to the neighbors found by PSH. In [Grauman et al.,2004], the authors show how Earth Movers Distance (EMD) and Locality-Sensitive Hashing (LSH) can be used for quick contour-based shape retrievals. In [Toyama and Blake,2002], an exemplar-based approach with dynamics is proposed for tracking pedestrians. In [Dimitrijevic et al.,2006], the authors present a template-based pose detector and solve the problem of large data set by detecting only human silhouette in a characteristic postures (sideways opened-legs walking postures in this case). Gavrila [2007] presents a probabilistic approach to hierarchical shape matching. These last four works basically look for the training human shape that best matches the input image but they do not infer any pose representations. Similar to Gavrila [2007] in spirit, Stenger [2004] uses a hierarchical Bayesian filter for real-time articulated hand tracking. Even if some techniques have been found for quick retrieval, the main drawback of all these exemplars based methods remains the large amount of memory needed to store the large data set. Learning Based Approaches. Instead of storing and performing a nearest neighbor
1.6 List of Relevant Publications 17 the 4th International Conference on Articulated Motion and Deformable Objects (AMDO), pages 175–184 •Rogez, G., Mart´ınez, J., and Orrite, C. (2007b). Dealing with non-linearity in shape modelling of articulated objects. In Proc. of the Third Iberian Conference on Pattern Recognition and Image Analysis (IbPria), pages 63–71 •Rogez, G., Orrite, C., and Mart´ınez, J. (2008a). A spatio-temporal 2d-models framework for human pose recovery in monocular sequences. Pattern Recognition, 41(9):2926–2944 1.6.2 Part II •Rogez, G., Guerrero, J., Mart´ınez, J., and Orrite, C. (2006a). Viewpoint independent human motion analysis in man-made environments. In Proc. of the 17th British Machine Vision Conference (BMVC), volume 2, pages 659–668, Edinburgh, UK •Rogez, G., Guerrero, J. J., and Orrite, C. (2007a). View-invariant human feature extraction for video-surveillance applications. In Proc. of the IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS), pages 324–329 •Rogez, G., Orrite, C., Rihan, J., Guerrero, J. J., and Torr, P. H. (under review). Viewinvariant shape-based 3d human pose tracking in monocular surveillance videos. submitted to the International Journal of Computer Vision 1.6.3 Part III •Rogez, G., Rihan, J., Ramalingam, S., Orrite, C., and Torr, P. H. (2008b). Randomized trees for human pose detection. In Proc. of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR) •Rogez, G., Rihan, J., Orrite, C., and Torr, P. H. (2012). Fast human pose detection using randomized hierarchical cascades of rejectors. International Journal of Computer Vision, 99(1):25–52 1.6.4 Miscellaneous: Several follow-on papers are still under preparation but some of the ideas developed in this thesis have already been used in the following publications: •Rogez, G., Rius, I., Mart´ınez, J., and Orrite, C. (2007c). Exploiting spatio-temporal constraints for robust 2d pose tracking. In Proc. of the Second Workshop of Human Motion - Understanding, Modeling, Capture and Animation, pages 58–73 •Mart´ınez, J., Orrite-Uru˜nuela, C., and Rogez, G. (2007). Rao-blackwellized particle filter for human appearance and position tracking. In Proc. of the Third Iberian Conference on Pattern Recognition and Image Analysis (IbPria), pages 201–208 •Ek, C. H., Rihan, J., Torr, P. H. S., Rogez, G., and Lawrence, N. D. (2008). Ambiguity modeling in latent spaces. In MLMI, pages 62–73 •Orrite, C., Ga˜n´an, A., and Rogez, G. (2009). Hog-based decision tree for facial expression classification. In IbPRIA, pages 176–183
18 Chapter 1. Introduction
Part I Segmentation and Pose Estimation with a Static Camera
2 Dealing with Non-linearities in Shape Modeling 2.1 Introduction In the first part of this thesis, we consider that the camera is fixed and that an estimate of the foreground silhouette can be obtained using a common static background subtraction algorithm. People are able to deduce the pose of a known articulated object (e.g. a person) from a simple binary silhouette. The possible ambiguities can be solved from dynamics when the object is moving. Following this statement, we propose in this first chapter to construct a human model that encapsulates within a Point Distribution Model (PDM) [Cootes and Taylor,1997] both body silhouette information provided by the 2D shape and structural information given by the 2D skeleton joints. In that way, the 2D pose could be inferred from the silhouette and vice versa. Due to the high non-linearity of the resulting feature space, mainly caused by the rotational deformations inherent to the articulated structure of the human body, the use of non-linear statistical models will be considered in this chapter. This approach will be compared to other two methods previously tested for solving non-linearity issue. Such non-linear statistical models have been previously proposed by Bowden [Bowden et al.,2000] that demonstrated how the 3D structure of an object can be reconstructed from a single view of its outline. While Bowden only considered the upper human body and the frontal view, in this work the complete body will be modelled and different viewpoint will be taken into account. 2.1.1 Related Work Even though many new types of image features have recently been developed, silhouette-based approaches are still receiving much attention. These approaches focus on the use of the binary silhouette of the human body as a feature for detection [Gavrila,2007,Lin and Davis,2010], tracking [Baumberg and Hogg,1994,Siebel and Maybank,2002b,Giebel et al.,2004,Toyama and Blake,2002,Cremers,2006], pose estimation [Jaeggli et al.,2009,Elgammal and Lee, 2009,Lee and Elgammal,2010,Li et al.,2010] or action recognition [Weinland et al.,2007, Abdelkader et al.,2011] to cite a few. They rely on the observation that most human gestures can be recognized using only the outline shape of the body silhouette. The most important advantage of these features is their ease of extraction from raw video frames using low-level processing tasks like background subtraction or edge detection algorithms.
22 Chapter 2. Dealing with Non-linearities in Shape Modeling 2.1.1.1 Human shape models Human shape models have appeared to be powerful tools for human motion analysis. When a 2D representation is employed, the outline of the silhouette is usually parameterized by a series of 2D landmarks. Baumberg and Hogg [1994] used active shape models to track pedestrians from a fixed camera. The same active shape tracker was considered by Siebel and Maybank [Siebel and Maybank,2002a] that extended it by a head detector and a region tracker, all integrated in the visual surveillance system ADVISOR. In [Fan et al.,2003], Fan et al. presented a compound structural and textural image model for pedestrian registration. In [Veeraraghavan et al.,2005], the authors even exploit the shape deformations of a person’s silhouette as a discriminative feature for gait recognition, indicating that methods based on shape perform better than methods based on kinematics alone. Wang et al. [2003] also propose an efficient gait recognition algorithm using view-dependent silhouette analysis. They concluded that the lack of generality of viewing angle is a limitation to most gait recognition algorithms. Munder et al. [2008] proposed a Bayesian framework for tracking pedestrians from a moving vehicle: a method for learning spatio-temporal shape representations from examples was outlined, involving a set of distinct linear subspaces models. The main problem with 2D models is their dependency to the viewpoint. Indeed, the accuracy strongly depends on the similarity between training and testing viewpoints. Recent approaches proposed the use of detailed 3D mesh models learnt from 3D laser scans [Corazza et al.,2006,Rosenhahn et al.,2006]. All these approaches represent the shape of the human but they ignore the rigid deformations inherent to the kinematic of the human body. One exception can be found in [Zhang et al.,2005a] where the authors introduced a statistical shape representation of non-rigid and articulated body contours. To accommodate large viewpoint changes, the authors proposed to employ a mixture of a finite number of view-dependent models. Even so, all those works on shape models present the same common drawback that, even if they give an idea of the pose, no real regression to 2D/3D joints is considered. 2.1.1.2 Shape+Structure models. Some few works have attempted to model the shape together with the structure of the human body. It is not a simple task since both types of feature belong to different metric spaces - pixels vs angles or pixels vs 3D position. Grauman et al. [Grauman et al.,2003] inferred 3D structure from multi-view contour using a probabilistic “shape+structure” model. This idea was first introduced by Bowden [Bowden et al.,2000] that demonstrated how the 3D structure of an object can be reconstructed from a single view of its outline using a model of movement and shape. In this work, 2D shape and 3D skeletal structure were encapsulated within a non-linear Point Distribution Model (PDM). Some very encouraging results have been shown by these two papers in controlled environment and with relatively “good” input silhouette. However, both suffer from the same disadvantage than 2D shape models: the need of manually segmented data. More recently, Balan et al. [Balan et al.,2007] also proposed something in between the two different approaches - kinematic models vs shape model - considering a parametric model of 3D shape and pose-dependent deformations to represent the body via SCAPE mesh model. One must say that the results are really impressive. They represent both articulated and non-rigid deformations of the human body as well as body shape variability. But this method requires 3D body scan of naked people. Another drawback is the computational cost required to fit the thousands of triangles of the mesh model.
2.2 Shape-Skeleton Training Database 23 2.1.2 Overview Thanks to the structural knowledge, people are able to deduce the pose of an articulated object (e.g. a person) from a simple binary silhouette. Following this statement, our idea is to construct a human model encapsulating within a point distribution model (PDM) body silhouette information given by the 2D shape (landmarks located along the contour) and the structural information given by the 2D skeleton joints. In that way, the 2D pose could be inferred from the silhouette. Due to the high non-linearity of the feature space, mainly caused by the rotational deformations inherent to the articulated structure of the human body, we consider in this work the necessity to use non-linear statistical models. They have been previously proposed by Bowden [Bowden et al.,2000] that demonstrated how the 3D structure of an object can be reconstructed from a single view of its outline. In a previous work [Orrite-Uru˜nuela et al.,2004], the problem of non-linear principal component analysis was partially resolved by applying a different PDM depending on previous pose estimation (4 views were considered: frontal, lateral, diagonal and back views) and the same procedure will be followed in this work. Additionally, results were obtained from measurement by selecting the closest shape from the training set by means of a Nearest Neighbour (NN) classifier. However, to cope with the remaining non-linearity, we consider in this thesis the use of Gaussian Mixture Models (GMM) as in [Enzweiler and Gavrila,2008]. A dynamic-based clustering is considered by partitioning the 3D pose space. The 4 viewbased GMM are then built in their respective shape space using this same labelling. Dynamic correspondences are then obtained between gaussian models of the 4 mixtures. In [Rogez et al., 2005], we proposed to use Independent Component Analysis (ICA) to cope with the problem of non-linearity in human shape modeling. Results obtained with our GMM will be compared with results from ICA and NN modeling. In Sect. 2.2, we introduce the training database construction. In Sect. 2.3, we detail the construction of our view-based GMM. Some results are presented in Sect. 2.4 and some conclusions are finally drawn in Sect. 2.5. 2.2 Shape-Skeleton Training Database The goal is to construct a statistical model which represents a human body and the possible ways in which it can deform. Point distribution models (PDM) are used to associate silhouettes (shapes) and the corresponding skeletal structures. 2.2.1 Training Database Construction The generation of the 2D deformable model follows a procedure similar to [Koschan et al., 2003]. The CMU MoBo database [Gross and Shi,2001] is considered for the training stage: good training shapes are extracted manually trying to get accurate and detailed approximations of human contours. Simultaneously, 13 fundamental points corresponding to a stick model are extracted: head center, shoulders, elbows, wrists, hips, knees and ankles. The skeleton vectors are then defined as: ki= [xk1, ..., xk13,yk1, ..., yk13]>∈R26,(2.1) with i= 1...Nv,Nvbeing the number of training vectors. Two gait cycles (low and high speed) and 4 viewpoints (frontal, lateral, diagonal and back views) are considered for each one of the 15 selected subjects. This manual process leads to the generation of a very precise database but without shape-to-shape landmarks correspondences.
24 Chapter 2. Dealing with Non-linearities in Shape Modeling 2.2.2 Shapes Normalization The good results obtained by a PDM depend critically on the way the data set has been normalized and on the correspondences that have been established between its members [Davies et al.,2003]. Human silhouette is a very difficult case since people can take a large number of different poses that affect the contour appearance. A big difficulty relies on establishing correspondences between landmarks and normalizing all the possible human shapes with the same number of points. In this chapter, we propose to use a large number of points for defining all the contours and “superpose” the points that are not useful (see Fig. 2.1). Figure 2.1: Training database. From left to right: MoBo Image, 2D skeleton and shape normalization: hand-labelled landmarks (A), rectangular grid (B), 120 normalized landmarks (C), part of them grouped at “repository points”: 24-26 at RP2, 46-74 at RP3 and 94-99 at RP1. A rectangular grid with horizontal lines equally spaced is applied to the contours database. This idea appears as a solution to the global verticality of the shapes and the global horizontality of the motion. The intersections between contours and grid are then considered. The shapes are then divided into 3 different zones delimited by three fixed points: the higher point of the head (FP1) and the intersections with a line located at 1/3 of the height (FP2 and FP3). A number of landmarks is thus assigned to each segment and a repository point (RP) is selected to concentrate all the points that have not been used. In this chapter, all the training shapes are made of 120 normalized landmarks: si= [xs1, ..., xs120,ys1, ..., ys120]>∈R240,(2.2) with i= 1...Nv. 2.2.3 Shape-Skeleton Eigenspace - PCA Model Shapes and Skeletons are now concatenated into Shape-Skeleton vectors: vi= [s> ik> i]>∈R266,(2.3) with i= 1...Nv. This training set is aligned using Procrustes analysis (each view being aligned independently) and Principal Component Analysis (PCA) is applied [Cootes and Taylor,1997]
2.3 View-based Shape-Skeleton Gaussian Mixture Models 25 for dimensionality reduction on the 4 view-based training sets. In that way, 4 view-dependent Shape-Skeleton models are constructed by extracting the mean vector and the variation modes: vi≃¯ vθ+Φθbi(2.4) where ¯ vθand Φθare respectively the mean Shape-Skeleton vector and the matrix of Eigenvectors for the training viewpoint θ.bis the projection of viin the corresponding Eigenspace i.e. a vector of weights bi= [b1, b2...bn]>. The main problem with this approach is that PCA assumes a Gaussian distribution of the input data. This supposition fails because of the inherent non-linearity of the feature space and leads to a wrong description of the data: the resulting model can consider as valid some implausible Shape-Skeleton combinations. Therefore, other approaches have to be taken into account to generate the “Shape-Skeleton” model and adequately represent the training set. 2.3 View-based Shape-Skeleton Gaussian Mixture Models Many researchers have proposed approaches to non-linear PDM [Cootes and Taylor,1997, Bowden et al.,2000]. The use of Gaussian Mixture Model (GMM) was first proposed by Cootes and Taylor [Cootes and Taylor,1997]. They suggested modeling non-linear data sets using a GMM fitted to the data using the Expectation Maximization (EM) algorithm. This provides a more reliable model since the feature space is limited by the bounds of each Gaussian that appear to be more precise local constraints. pmix(b) = m X j=1 ωjN(b:¯ bj,Sj),(2.5) where N(b:¯ b,S) is the p.d.f. of a gaussian with mean ¯ band covariance S. Bowden [Bowden et al.,2000] proposed first to compute linear Principal Component Analysis (PCA) and to project all shapes on PCA basis. Then do cluster analysis on projections and select an optimum number of clusters. Each data point is assigned to a cluster and separate local PCA are performed independently on each cluster. This results in the identification of local model’s modes of variation inside each Gaussian distribution of the mixture: b≃¯ bj+Φjr (see Eq.2.4). Thus, a more complex model is built to represent the statistical variations. Given the promising results described in [Bowden et al.,2000], a similar procedure is followed in this work, the main difference relying on the way the feature space is clustered: the proposed methodology consists in partitioning the complete Shape-Skeleton feature space using only the dynamical information provided by the pose parameters. The contour parameters are not taken into account for clustering since they do not provide any additional information on dynamics and can lead to ambiguities as stated in [Agarwal and Triggs,2006]. 2.3.1 Structural clustering While in [Munder et al.,2008], the clustering of the shape feature space was based on a similarity measure derived from the registration procedure, here it is proposed to use the structural information provided by the pose to cluster both shape and skeleton training sets, thus establishing dynamical correspondences between view-based data.
26 Chapter 2. Dealing with Non-linearities in Shape Modeling Figure 2.2: Low and high speed gait cycles represented on the 3 first modes of the Pose Eigenspace. 2.3.1.1 Pose Eigenspace for Clustering The information provided by the 3D poses is used for clustering: for each snapshot of the training set, the 3D skeleton is built from the corresponding 2D poses kiof the 4 views, by reconstructing the 3D position of the joints using the 4 2D-projections and Tsai’s algorithm [Tsai,1986]. The resulting set of 3D poses is then aligned using Procrustes and reduced by PCA obtaining the Pose Eigenspace (Fig. 2.2) where the dynamic-based clustering will be operated. 2.3.1.2 Principal components selection The non-linearities of the training set are mainly localized in the first components of the PCA which capture the dynamics, as shown in Fig. 2.2. These components are really influential during the partitioning step while the last ones, more linear, only model local variations (details) and do not provide so much information for clustering. Only the first non-linear components are thus selected to perform the clustering of the data in a lower dimensional space. For components selection, the non-gaussianity of the data is measured on each component. There are different methodologies to test whether the assumed normal probability distribution accurately characterizes the observed data or not. Skewness and kurtosis, are two classical measures of non-gaussianity. A more robust measure is given by the Negentropy, the classic information theory’s measure of non-Gaussianity, whose value is zero for Gaussian distribution [Hyvaerinen et al.,2001]. Fig. 2.3a shows how the Negentropy converges to 0 and oscillates when considering lower modes. This oscillation between 0 and 0.75 ×10−4starts from the 4th mode. It can be observed how the first 3 modes present a much higher Negentropy compared to the other modes. According to this analysis, we select the 3 first components for clustering. 2.3.1.3 Determining the number of clusters. K-means algorithm is used fairly frequently as a result of its ease of implementation. Kmeans clustering splits a set of objects into a selected number of groups by maximizing between variations relative to within variations. The main disadvantage of this algorithm is its extreme sensitivity to the initial seeds. A solution could be found by applying k-means several times,
3 A Framework of Spatio-temporal Models 3.1 Introduction One of the difficulties when employing 2D-models relies on dealing with the viewpoint issue. During many years, most of the related work has been based on the fundamental assumption of “in-plane” motion or only presented results obtained from data satisfying such condition [Zhang et al.,2004,Ning et al.,2004a]. Relatively few papers considered motion-in-depth and out-ofplane rotation of the tracked people. In the last years, some authors have proposed a common approach consisting in discretizing the camera viewpoint and considering a series of view-based 2D models [Zhang et al.,2005b,Lee and Elgammal,2006,Lan and Huttenlocher,2004]. This method gives some good results, but there are two main problems that need to be addressed: spatial discontinuities due to the viewpoint discretization and temporal discontinuities due to the difficulties of maintaining the dynamics of the motion when the view is switched. It appears as a challenging problem to establish motion correspondences between viewpoints without considering a mapping to a complex 3D model. Therefore, the goal of the work presented in this chapter is to construct 2D dynamical models 1) that can perform independently of the orientation of the person with respect to the camera and 2) that can respond robustly to any change of direction during the sequence. 3.1.1 Overview of the work We extend the model presented in the previous chapter by considering additional training viewpoints and complete the ring of possible viewpoints around the subject varying the azimuth angle of the camera. Again, we consider the walking action, but the methodology can be extended to any type of action. As in chapter 2, 2D pose and 2D appearance parameters are first extracted from training images of the same gait sequences observed from the considered viewpoints. The resulting set of 2D shape and skeletons is then clustered following both spatial and temporal criteria, the spatial clustering being directly provided by the training views and the temporal clustering resulting from motion-based partitioning, i.e. the steps of the gait cycle. A spatio-temporal clustering is thus obtained in the global Shape-Skeleton eigenspace: the different clusters correspond in terms of dynamic (temporal clusters) or viewpoint (spatial clusters). A local 2D-model is then built for each spatio-temporal cluster, generalizing well for a particular training viewpoint and state of the considered action. All those models are
34 Chapter 3. A Framework of Spatio-temporal Models concatenated and sorted, what leads directly to the construction of the global Spatio-Temporal 2D-Models Framework (STMF) presented in Fig. 3.1. Figure 3.1: Spatio-temporal Shape-Skeleton Models Framework: 1st Variation Modes of the 48 local Models that compone the framework. The columns of this matix correspond to the gait steps (temporal clusters) while the rows represent the 8 camera views (spatial clusters). These spatio-temporal models generalize well for a particular viewpoint and state of the tracked action but some spatio-temporal discontinuities can appear along a sequence, as a direct consequence of the discretization. Additionaly, an efficient search method is required to guide the exploration of a large feature space. To overcome these problems, we propose to consider temporal and spatial constraints and build a Probabilistic Transition Matrix (PTM). This matrix limits the search in the feature space by giving a frame to frame estimation of the most probable local models to be considered during the fitting procedure. Our approach is similar to [Lan and Huttenlocher,2004] in that we integrate spatial and temporal models into a common framework, but differs in that we consider a combined transition that takes into account simultaneous state and viewpoint changes. This constraint-based search is described in Section 3.3. Once the model has been generated (off-line), it can be applied (on-line) to real sequences. Given an input human blob provided by a background subtraction algorithm, the model is fitted to jointly segment the body silhouette and infer the posture. This model fitting is explained in Section 3.4. Experiments are presented in Section 3.5 where both segmentation and 2D pose estimation are tested. The main goal of this part is to test the robustness of the approach w.r.t. the viewpoint changes with realistic conditions: indoor, outdoor, cluttered background, shadows etc. In that way, the following hypothesis will be considered to select the different testing sequences: only one walking pedestrian per sequence, with no occlusions but with important viewpoint changes. Note that both training and testing sets comprise of hand-labelled data. The
3.2 Framework Construction 35 CMU MoBo database [Gross and Shi,2001] will be used for training and real video-surveillance sequences for testing [Caviar,2004]. The HumanEVA dataset, recently introduced by Sigal and Black [Sigal et al.,2010], will be considered for numerical evaluation of the pose estimation. Section 3.6 finally concludes with some discussions. 3.2 Framework Construction Recently, some authors have proposed a common approach consisting in discretizing the space considering a series of view-based 2D models [Zhang et al.,2005a,Lan and Huttenlocher,2004]. In the same way, 8 different viewpoints will be considered, uniformly distributed between 0 and 2π, thus discretizing the frontal view (vertical image plane) into 8 sectors. For each sequence, the 4 training viewpoints used up to that point are now completed by a 5th supplementary back view that is also manually labelled. Finally, the 3 last missing views are interpolated using the periodicity and symmetry of human walking. By this process, a complete training database is generated encompassing more than 20000 Shape-Skeleton vectors, SS-vector (more than 2500 vectors per viewpoint). The resulting 8 view-based Shape-Skeleton associations for a particular snapshot of the CMU MoBo database are presented in Fig. 3.2. Figure 3.2: Training views considered for Framework Construction. The complete set of SS-vectors is concatenated in a common space (the 8 views together) whose dimensionality is reduced using PCA, obtaining: vi≃¯ v+ Φvai,(3.1) where ¯ vis the mean SS-vector, Φ is the matrix of Eigenvectors and aiis the new vector represented in the Eigenspace. Let us call Athe Shape-Skeleton Eigenspace {ai}. A series of local dynamic motion models has been learnt by clustering the structural parameters subspace. As mentioned in the Section 2.3.1, the gait cycle is divided into 6 basic steps, providing the temporal clusters Cj, while the 8 training views directly provide the spatial clustering (clusters Rr). The different clusters correspond in terms of dynamics or viewpoint. Using this structure-based partitioning and the correspondences between training viewpoints, 48 spatio-temporal clusters {{Tj,r =Cj∩Rr}6 j=1}8 r=1 are obtained in the global shape-skeleton feature space where all the views considered are projected together. Thus, following [Bowden et al.,2000], a local linear model is learnt for each spatio-temporal cluster Tj,r and a mixture of PCA is fitted to the clustered Aspace, obtaining a new SpatioTemporal 2D Models Framework (STMF). For each cluster, the local PCA leads to the extraction of local modes of variation, in which both shape and skeleton simultaneously deform
36 Chapter 3. A Framework of Spatio-temporal Models (a) (b) (c) (d) Figure 3.3: Multi-view Gaussian mixture model represented in the pose-silhouette eigenspace. The 48-clusters Gaussian Mixture Model is represented together with the training data and projected onto the planes defined by (a) 1st and 2nd, (b) 3rd and 4th, (c) 5th and 6th and (d) 7th and 8th components of the Shape-Skeleton Eigenspace. (see Fig. 3.4). Parameters for the 48 Gaussian mixture models components are determined using EM algorithm. The prior Shape-Skeleton model probability is then expressed as: pmix(a) = X j,r ωj,r N(a:¯ aj,r, σj,r),(3.2) where ais the eigen-decomposition of the Shape-Skeleton vector, N(a:¯ a, σ) is the p.d.f. of a gaussian with mean ¯ aand covariance σand ωj,r is the mixing parameter corresponding to Tj,r. The figure 3.3 shows the mixture projected onto various planes of the Eigenspace space A. The 48 hyper-ellipsoids corresponding to the 48 local spatio-temporal models are also plotted. It can be appreciated how well the GMM delimits the subspace of valid SS-vectors. Given this huge amount of data, an efficient search method is required. In that way, temporal and spatial constraints will be considered to constrain the evolution through the STMF along a sequence and limit the feature space only to the most probable models of the framework.
3.3 Constraint-based search 37 3.3 Constraint-based search The total space has been clustered following temporal approach (clusters Cj) as well as spatial approach (clusters Rj) as described in the previous Section. The first one partitions the dynamics of the motion, and the second one, the viewpoint i.e. the direction of motion in the image. The purpose of the following probabilistic modeling is to obtain a transition matrix combining both spatial and temporal constraints. Figure 3.4: 3D (left) and 2D (right) representations of the toroidal Probabilistic Transition Matrix (PTM). The 1st Variation Modes of the 48 local Models that compone the framework are superposed with the 2D representation of the PTM: the 6 columns correspond to the 6 temporal clusters Ciwhile the 8 rows represent the 8 spatial clusters Ri. 3.3.1 Markov Chain for Modeling Temporal Constraint Following the standard formulation of probabilistic motion model [Sidenbladh et al.,2002], the temporal prior p(St|St−1) satisfies a first-order Markov assumption where the choice of the present state Stis made upon the basis of the previous state St−1. In the same way, if this state space is partitioned into Nclusters C={C1, ..., CN}, the conditional probability mass function defined as p(Ct j|Ct−1 k) corresponds to the probability of being in cluster jat time tconditional on being in cluster kat time t-1 [Bowden et al.,2000]. The NxNState Transition Matrix (STM) computed in the previous Section points out the probabilities density function (pdf). 3.3.2 Modeling Spatial Constraint In this chapter, a novel spatial prior p(Dt|Dt−1,...t−m) is introduced for modeling spatial constraint. It expresses the statement that Dt(the present direction of motion of the observed pedestrian in the image) can be predicted given its mprevious directions of motion (Dt−1,Dt−2, ..., Dt−m). In this approach, the continuous values of all possible camera viewpoints
38 Chapter 3. A Framework of Spatio-temporal Models are discretized. Consequently, the direction of motion in the image plane Dttakes a fixed set of values corresponding to the discrete set of Mtraining viewpoints and Mclusters R={R1, ..., RM}in the feature space. Let ∆t= [Rt k0,Rt-1 k1, ..., Rt-m km] be the m+1-dimensional vector representing the sequence of the m+1 cluster labels (denoted by ki) up to and containing the one at time t. It has to be noted that some of these kilabels might be the same. Consequently, p(Rt j|∆t-1) is the probability of being in Rjat time t, conditional on being in Rk1at time t-1, in Rk2at time t-2, etc. (i.e. conditional on the mpreceding clusters). In this chapter, a reasonable assumption is made that this direction of motion follows a normal distribution, with expected value equal to the local mean trajectory angle θtand, variance calculated as a function of the sampling rate: p(Rt j|∆t-1) = p(Rt j|Rt-1 k1,Rt-2 k2, ..., Rt-m km)∼ N(θt, σ), (3.3) where θt=1 m+1 Pt-m i=tθi, being ma function of the sampling frequency. 3.3.3 Combining Spatial and Temporal Constraints Let Tbe the NxMmatrix, whose columns represent the Ntemporal clusters and rows correspond to the Mspatial clusters. Thus the probability p(Ct j∩Rt r) = p(Tt j,r) denotes the unconditional probability of being in Cjand in Rrat time t. The conditional spatio temporal transition probability is therefore defined as p(Tt j,r|Ct−1 k,∆t-1), the probability of being in Cjand in Rrat time tconditional on being in temporal cluster kat time t-1 and conditional on the mpreceding spatial clusters. In this thesis, the assumption is made that the two considered events, state and direction changes, are independent, even if it is not strictly true. Some comments about this assumption will be made in Sections 3.5 and 3.6. This leads to the following simplified equation: pj,r =p(Tt j,r|Ct−1 k,∆t-1)∝p(Ct j|Ct−1 k).p(Rt r|∆t-1). (3.4) The resulting NxMmatrix is the Probabilistic Transition Matrix (PTM) that gives, at each time step, the probability density function that limits the region of interest in the STMF to the most probable models. The matrix is in fact a 2D manifold representation for viewpoint and action where the action (consecutive temporal clusters) is represented by a 1D manifold and the viewpoint by another 1D manifold. Because of the cyclic nature of the viewpoint parameter (circular distribution of the training viewpoints), if it is modeled with a circle the resulting manifold is in fact a cylindrical one. When the action is cyclic too, as in our case with gait, the resulting 2D manifold lies on a “closed cylinder” topologically equivalent to a torus. The resulting PTM is thus a toroidal matrix (Fig. 3.4) whose lines correspond to the M training view-based gait manifolds. Its 3D and 2D representations are illustrated in Fig. 3.4. All the different models can be ordered and classified according to their direction of motion and state, thus putting in evidence the correspondences with the PTM as shown in Fig. 3.4. Spatial and temporal relationships can be appreciated between local models from adjacent cells. The content of the PTM can be visualized by converting it to grey scale image as will be shown in next sections. To compute this PTM and constrain the evolution through the STMF along a sequence, only previous viewpoint and previous state are required at each time step. Note that our approach shares some similarities with the one proposed by Lv and Nevatia [Lv and Nevatia,2007]. In that paper, the authors model an action as a series of 2D Poses rendered from a wide range of viewpoints and represent the constraints on transition by a graph model where they assume a uniform transitional probability for each link.
3.4 Joint Segmentation and Pose Estimation 39 3.4 Joint Segmentation and Pose Estimation A discriminative detector as the one proposed in [Dimitrijevic et al.,2006] could be used to initialize the shape model-driven algorithm presented next. In this work, scene context information is considered to roughly limit the feature space only to the “logical” 2D-models from the framework. For example, if an object appears in the scene from the right side (rightto-left direction of motion), only the 3 first lines of the PTM will be considered. (a) (b) (c) (d) (e) (f) (g) (h) (i) Figure 3.5: Model Fitting: (a) original input image. (b) Silhouette resulting from background subtraction. (c) Silhouette after being processed by BlobsProcessing. (d) Contour extracted by silhouette erosion. (e) Corrected shape represented on the Silhouette and corresponding mask (f) used for finer background subtraction. (g) Resulting Silhouette after 6 iterations and (h) corresponding segmentation. (i) Resulting shape and pose plotted on the original input image. Once the system has been initialized, each frame of the sequence is processed individually by applying SegmentationPoseEstimation (Algorithm 1), taking advantage of previous information (Trajectory angle θ, State index, Background B) that is used to treat the current frame. Algorithm 1: (s,k,B,θ,index) = SegmentationPoseEstimation (B,I,θ,index,nIter) m[i] = ModelsSelection(θ,index); initialize shape s←0; initialize pose k←0; for n←1 to nIter do blobsList = AdaptiveBackgroundSubtraction(B,I,s,n,nIter); Silhouette = BlobsProcessing(blobsList); s=ContourExtraction(Silhouette); [s,k,index]= ShapeSkeletonCorrection(s,k,m[i]); end for B=BackgroundUpdate(B,I,Silhouette); The prediction of the most probable models from the GMM is estimated in ModelsSelection by means of the PTM. It allows a substantial reduction in computational cost and can solve some possible ambiguities by considering a limited number of models. In ShapeSkeletonCorrection, the extracted shape sand an estimate for the skeleton are concatenated in v= [s>k>]>and projected into the SS-Eigenspace obtaining a. Then, for each one of the most probable clusters given by the PTM pj,r, we update the parameters to best fit the “local model” defined by its mean, eigenvectors and eigenvalues, as done in Sect. 2.3.3, obtaining a∗. The distance between extracted and corrected shapes is then calculated for each one of the estimations in order to select the best estimation. We then project the vector a∗ back to the feature space obtaining v∗which contains a new estimation of both shape s∗and skeleton k∗:v∗= [s∗>k∗>]>.
40 Chapter 3. A Framework of Spatio-temporal Models Algorithm 2: blobsList = AdaptiveBackgroundSubtraction(B,I,s,n,nIter) if (n== 1) then D = BackgroundSubtraction(B,I,thr); else Mask = ShapeToMask(s); tlow =DecreaseThreshold(thr,n,nIter ); Dlow =BackgroundSubtraction(B,I,tlow); thigh =IncreaseThreshold(thr,n,nIter); Dhigh =BackgroundSubtraction(B,I,thigh); D = Dlow ×Mask +Dhigh ×Mask; end if blobsList = BlobsLabelling(D); Aside from the models and the constraint-based search proposed in this work, some novelties appear in the segmentation algorithm. The first one refers to the shape extraction task (ContourExtraction in Alg. 1): while it is usually extracted from the blob looking along straight lines through each model point as in [Baumberg and Hogg,1994], here the shape is directly obtained by eroding the human blob and normalizing the resulting contour following the shape normalization proposed in Sect. 2.2.2. This allows a direct, precise and faster registration of the shape in the image. The only drawback of this shape registration is that it requires an entire and non-fragmented silhouette. The BlobsProcessing function thus previously applies some common morphological operations to the result of AdaptiveBackgroundSubtraction (Alg. 2) and connect the possible fragments. Another novelty of the fitting process appears in AdaptiveBackgroundSubtraction (Alg. 2) that aims at reconstructing the binary silhouette resulting from the background subtraction using the “corrected” shape returned by ShapeSkeletonCorrection. It is achieved by decreasing/increasing the detection threshold inside/outside the shape. This leads to an accurate silhouette segmentation, improving considerably the results specially when there is no significant difference between background and foreground pixels. The last novelty relies on the way the background is updated: the final segmented silhouette, the foreground, is used to actualize the Background more finely, eliminating shadows from the foreground and improving the segmentation in next frames. The different steps of the process are depicted in Fig. 3.5 for a particular frame. 3.5 Experiments The model is now evaluated with a series of testing sequences illustrating different situations which may occur in the analysis of pedestrian motion: straight line walking, changes of direction, of speed, etc. Only model fitting and pose estimation will be tested in the first set of experiments, and not the tracking in the image, the system is then provided with the bounding-box taken from groundtruth avoiding the possible problems due to the tracking. In the PTM matrices from Fig. 3.6 ( as well as from Fig. 3.8 and 3.9), the colored cells represent the probability pj,r from (Eq. 3.4). The obscured cell is the “winning one”: the local model that best fits the silhouette. For each frame, the row of the “winning” model in the PTM indicates the orientation of the pedestrian with respect to the camera. Additionally, both trajectory and previous states are respectively plotted in the image/matrix with a white line. As illustrated in Fig. 3.6 (up), the resultant vectors from a pedestrian crossing the scene
3.5 Experiments 41 Figure 3.6: Examples of sequences processed with the framework of pose-silhouette models: (up) Outdoor straight line walking sequence at constant speed and (down) Caviar sequence with slight bend trajectory. straight ahead without stopping or turning towards anything all belong to models from the same row of the PTM. Any change of direction is observed as a progressive change of row (See Fig. 3.6 (down)). In Figure 3.7, the results obtained with two challenging frames are presented: in (a), the pedestrian is carrying a bag and in (b) he is partially occluded by the wall. In both cases, a plausible estimation is made of both shape and pose. (a) (b) Figure 3.7: Results obtained with two challenging frames. For each of the two examples, original image (left), segmentation (centre) and resulting pose and shape are represented (right). 3.5.1 Framework Validation To validate the Framework, the 2D poses and 2D shapes of 3 different sequences with different characteristics of interest are hand-labelled: an outdoor straight-line walking sequences at constant speed (Fig. 3.6up), an outdoor “Walkcircle” sequences with constant speed and constant viewpoint & scale evolution (Fig. 3.8) and an indoor sequence with viewpoint & speed variations (Fig. 3.9). Note that the subjects turn and move “in depth” so that both apparent
42 Chapter 3. A Framework of Spatio-temporal Models scale and viewpoint vary. A top-down estimation of depth is directly provided by the “winning” model that points out the motion direction in 3D space (see Fig. 3.6 and later Fig. 3.8 and 3.9). (a) (b) (c) Figure 3.8: Results obtained for the outdoor “Walkcircle” sequence with constant speed and constant viewpoint & scale changes : (a) Estimated shapes and poses represented on the original image for frames 1, 15, 25, 40, 50, 60, 75 90, 100, 115, 130 and 140. (b) 12 corresponding PTM matrices and (c) 2D Poses estimated along the complete sequence. 3.5.1.1 Segmentation Quantitative validation is performed by comparing with manually segmented solutions, both the segmentation obtained by simple background subtraction and the one resulting from the
Part II Pose Tracking in Video-Surveillance Environments
4 View-invariant Motion Analysis using View-based Models 4.1 Introduction The second part of this thesis is dedicated to the problem of pose estimation and tracking in video-surveillance scenarios. In recent years, the number of cameras deployed for surveillance and safety in urban environments has increased considerably in part due to their falling cost. The potential benefit of an automatic video understanding system in surveillance applications has stimulated much research in computer vision, especially in the areas related to human motion analysis. The hope is that an automatic video understanding system would enable a single operator to monitor many cameras over wide areas more reliably. As demonstrated in the introduction chapter, exemplar-based approaches have been very successful in the different stages of human motion analysis: detection, pose estimation and tracking. The main disadvantage of the techniques based on training exemplars is their direct dependence on the point of view: the accuracy of the result strongly depends on the similarity of the camera viewpoint between testing and training images. Ideally, to deal with viewpoint dependency, one could generate training data from infinitely many camera viewpoints, ensuring that any possible camera viewpoint could be handled. Unfortunately, this set-up is physically impossible and makes the use of real data infeasible. It could, however, be simulated by using synthetic data, but using a large number of views would drastically increase the size of the training data. This would make the analysis much more complicated; furthermore, the problem is exacerbated when considering more actions. In practice, roof-top cameras are widely used for video surveillance applications and are usually placed at a significant angle from the floor, which is different from typical training viewpoints as shown in the example in Fig. 4.1. Perspective effects can deform the human appearance (e.g. silhouette features) in ways that prevent traditional techniques from being applied correctly. Freeing algorithms from the viewpoint dependency and solving the problem of perspective deformations is an urgent requirement for further practical applications in videosurveillance. The goal of the work presented in this chapter is to track and estimate the pose of multiple walking people independently of the point of view from which the scene is observed (see Fig. 4.2a), even in cases of high tilt angles and perspective distortion. We have seen that our framework of view-based models performs decently when the camera viewing direction is parallel to the ground (ϕ≃0) and an estimate of both 2D pose and camera viewpoint can be made in spite of the discretization of the training viewpoint. But when using
52 Chapter 4. View-invariant Motion Analysis using View-based Models (a) (b) Figure 4.1: Video sequences considered in the chapter: (a) CMU Mobo database [Gross and Shi,2001] for training and (b) video-surveillance Caviar database [Caviar,2004] for testing. The CMU snapshot in (a) is represented from 4 different viewing angles (clockwise from upper left): frontal, diagonal-rear, lateral and diagonal. The difference of viewing angle can be observed between training and testing sequences. roof-top camera sequences, a pre-processing of the input image is necessary for perspective correction and correct view alignment. The challenge is then to make use of those models successfully on any possible sequence taken from a single fixed camera with an arbitrary viewing angle. A solution is proposed to the paradigm of “View-insensitive process using view-based tools” for video-surveillance applications in man-made environments: supposing that the observed person walks on a planar ground in a calibrated environment, we propose to compute the homography relating the image points to the training plane of the selected viewpoint. The input image is then warped to this training view and a pose is estimated using the corresponding view-based models. 4.1.1 Related Work In addition to the problem of viewpoint dependency of the model, we will have to overcome the classical difficulties that appear when tracking people in complex, but real, surveillance video-sequences. These difficulties are quite common: people moving in groups, occlusions, shadows/reflections on the ground/wall, low image resolution and especially the small size of the subjects that makes the pose estimation much more complicated. Many surveillance systems can be found in the literature: for example, W4[Haritaoglu et al., 2000], BraMBLe [Isard and MacCormick,2001] and ADVISOR [Siebel and Maybank,2002a]. But those systems only consider data where multiple people are distributed horizontally in the image. Promising tracking results were presented by Zhao and Nevatia [Zhao and Nevatia, 2004] in conditions similar to those we propose: perspective images with camera model known and the assumption that people walk on a known ground plane. They first locate people by detecting the head as in [Haritaoglu et al.,2000], then use a coarse 3D shape model (an ellipsoid) for global motion tracking as done by Isard with a cylinder [Isard and MacCormick, 2001]. Finally they employ a locomotion model to infer the 3D human posture. Even though
4.1 Introduction 53 (a) (b) Figure 4.2: (a) Viewing hemisphere: the position of the camera with respect to the observed subject (the view) can be parameterized as the combination of two angles: the elevation ϕ∈0,π 2(also called latitude or tilt angle) and the azimuth θ∈[−π, π] (also called longitude). A third angle γ∈[−π, π] can be considered to parameterize the rotation around the viewing axis. (b) Viewpoint discretization for training: in this work, we use the MoBo dataset [Gross and Shi,2001] and discretize the viewing hemisphere into 8 locations where θis uniformly distributed around the subject. An example is given in Fig. 4.1.a for front (F), rear-diagonal (RD2), lateral (L1) and diagonal (D1) views. our work shares similarity with [Zhao and Nevatia,2004], there are two major differences: 1) in our work, segmentation, tracking and pose estimation will be done all together using a more detailed silhouette-pose model and 2) we take into account the possibility of a very large tilt angles. Viewpoint dependence has been one of the bottlenecks for research development of human motion analysis as indicated in a recent survey [Ji and Liu,2010]. Some work has been done on solving the problem of viewpoint dependency. In [Cucchiara et al.,2005], a calibrated approach is used in order to avoid perspective distortion of the extracted features. Farhadi and Tabrizi [2008] propose a method to build features that are highly stable under change of camera viewpoint and recognize action from new views. Recently, Gong and Medioni [2011] achieved view-invariant action recognition on videos by associating a few motion capture examples using a novel Dynamic Manifold Warping (DMW) alignment algorithm. In [Parameswaran and Chellappa,2004], the authors present a method to calculate the 3D positions of various body landmarks given an uncalibrated perspective image and point correspondences in the image of the body landmarks. They also address the problem of view-invariance for action recognition in [Parameswaran and Chellappa,2006]. Grauman et al. [2004] propose a solution for inferring a 3D shape from a single input silhouette with an unknown camera viewpoint. The model is learnt by collecting multi-view silhouette examples from a calibrated camera ring and the visual hull inference consists in finding the shape hypotheses most likely to have generated the observed 2D contour. The concept of “virtual cameras”, introduced in [Rosales et al.,2001], allows for the reconstruction of synthetic 2D features from any camera location. The joint log-likelihood of body pose and camera parameters is maximized and results in an estimate of the 3D body pose. In [Kale et al.,2003], a method is proposed for view invariant gait recognition: considering a person walking along a straight line (making a constant angle with the image plane), a side-view is synthesized using a homography. Recently, Datta et al. [2011] described a motion estimation algorithm for projective cameras that explicitly enforces articulation constraints and presented
54 Chapter 4. View-invariant Motion Analysis using View-based Models pose tracking results for binocular sequences. In [Bouchrika et al.,2009] the authors propose a reconstruction method to rectify and normalize gait features recorded from different viewpoints into the side-view plane, exploiting such data for human recognition. The rectification method is based on the anthropometric properties of human limbs and the characteristics of the gait action [Goffredo et al.,2008]. In [Rogez et al.,2006a] we proposed an algorithm that uses a projective transformation between training and testing images to find viewpoint invariance. This paper will be discussed in details in the next section. In the same spirit, Li et al. [2008] later employed a homographic transformation to improve human detection in images presenting perspective distortion. They reported an improvement in detection rate from 38.3% to 87.2% using 3D scene information instead of scanning over 2D ( plus in-plane rotation) on the same testing dataset [Caviar,2004] we consider in this chapter. 4.1.2 Motivation and Overview of the Approach. As discussed in [Riklin-Raviv et al.,2007], in presence of perspective effect, the distortion will cause the parts of the subject that are closer to the lens to appear abnormally large, thus deforming the shape of the human contour in ways that can prevent a correct analysis. The basic idea is that projective geometry could be exploited when camera viewpoints in training and testing images are too different. In [Rogez et al.,2006a], we numerically demonstrated that the use of a projective transformation for shape registration, projecting both model and image in a canonical vertical view, improves silhouette-based pose estimation. In that case, the parameters required to estimate the homography, i.e. the subject’s location on the ground plane (X, Y ) and the viewpoint θ(Fig. 4.2a), were taken from ground truth data. In this chapter, a classical process of Detection-Segmentation-Tracking (see Fig.4.3up) is considered during the processing of a video-surveillance sequence with arbitrary viewpoint. The pedestrians are tracked using a Kalman filter to automatically estimate X,Yand θ: the tracking is applied on the ground floor and the orientation with respect to the camera (the viewpoint) is estimated by approximating it by the direction of walk. In a calibrated environment, a good estimation of the ground plane position can be obtained by projecting vertically the head location. An advantage of sequences taken by a rooftop camera, is that the head is usually less likely to be occluded and appears as the best feature to track. We thus develop a view-invariant head tracker to estimate X,Yand θ. For each frame, the nearest training view is selected and the homography that relates the image points to the corresponding training plane is considered. Supposing the person walks on a planar ground in a structured man-made environment, this homography can be computed using the dominant 3D directions of the scene in both training and input images. The projection of the input image onto the corresponding training image plane then compensates for the effect of both discretization along θand variations along ϕ. It also removes part of the perspective effect. Once the input image has been warped, the pose can then be estimated employing the view-based models corresponding to the selected training view from the framework introduced in the first Part of this thesis. The resulting silhouette and 2D pose can then be back-projected onto the original input image plane. The rest of the chapter is organized as follows. First, some geometrical considerations are explained in Section 4.2. The computation of the projective transformation is then described in Section 4.3 while the tracking framework is depicted in Section 4.4. Experiments and quantitative evaluation are presented in Section 4.5 and conclusions are drawn in Section 4.6.
4.2 Geometrical Considerations in Man-Made Environments 55 Figure 4.3: (up) The proposed tracking system diagram comprises 3 main blocks: (A) Detection, (B) Pose Estimation&Segmentation and (C) Tracking. The main contribution we make in this chapter appears in the Segmentation-Pose Estimation block detailed below. 4.2 Geometrical Considerations in Man-Made Environments We propose to exploit camera and scene knowledge when working in a man-made environments which is the case of most video-surveillance scenarios. 4.2.1 Notations In the following sections upper case letters, e.g. Xor X, will be used to indicate quantities in space whereas image quantities will be indicated with lower case letters, e.g. xor x. Euclidean vectors are denoted with upright boldface letters, e.g xor X, while slanted letters denote their cartesian coordinates, e.g. (x, y, z) or (X, Y ). Following notations used in [Sola et al.,2012], underlined fonts •indicate homogeneous
56 Chapter 4. View-invariant Motion Analysis using View-based Models coordinates in projective spaces, e.g xor X. A homogeneous point x∈Pnis composed of a vector m∈Rnand a scalar ρ(usually referred to as the homogeneous part): x=m ρ∈Pn⊂Rn+1,(4.1) where the choice ρ= 1 is the original Euclidean point representation while ρ= 0 defines the points at infinity. The homogeneous point xthus refers to the Euclidean point x∈Rn: x=m/ρ. (4.2) By definition, all the homogeneous points {[ρxT, ρ]T}ρ∈R∗represents the same Euclidean point x(see [Hartley and Zisserman,2004]) and, for homogeneous coordinates, “=” means an assignment or an equivalence up to a non-zero scale factor. 4.2.2 Camera and Scene Calibration Supposing observed humans are walking on a planar ground floor with a vertical posture, camera model and ground plane assumptions provide useful geometric constraints that help reducing the search space as in [Lin and Davis,2010,Zhao et al.,2008,Zhao and Nevatia, 2004], instead of searching for all scales, all orientations and all positions. During the scene calibration two 3×3 homography matrices are calculated: Hgwhich characterizes the mapping between the ground plane in the image and the real world ground plane Πgd and Hhrelating the head plane in the image with Πgd. In this work, the homography matrices are estimated by the least-squares method using four or more pairs of manually preannotated points in several frames. The 2 homography mappings are illustrated in Fig. 4.4. Note that when surveillance cameras with a high field of view are used (as with [Caviar,2004]), a previous lens calibration is required to correct the optical distortion. Given an estimate of the subject’s location (X, Y ) on the world ground plane Πgd, the planar homographies Hgand Hhare used to evaluate the location of the subject’s head xHand “feet” xFin the image I: xH=Hh·[X,Y,1]T,(4.3) xF=Hg·[X,Y,1]T,(4.4) where points in the projective space P2are expressed in homogeneous coordinates. In this work, we want to compensate for the difference of camera view between input and training images using the dominant 3D directions of the scenes. We suppose that the camera model is known and people walk in a structured man-made environment where straight lines and planar walls are plentiful. The transformation matrices introduced in the next section are calculated online using the vanishing points 1evaluated in an off-line stage: the positions of the vertical vanishing point vZand l, the vanishing line of the ground plane, are directly obtained after a manual annotation of the parallel lines (on the ground and walls) in the image. An example of vertical vanishing point localization is given in Fig. 4.4. This method makes sense only for man-made environments because of the presence of numerous easy-to-detect straight lines. Previous work for vanishing points detection [Lutton et al.,1994] could be used to automate the process. Once we have calibrated the camera in the scene, the camera cannot be moved, which is a limitation of the proposal. In practice, the orientation of the camera could change, for example, 1A vanishing point is the intersection of the projections in the image of a set of parallel world lines. Any set of parallel lines on a plane define a vanishing point and the union of all these vanishing points is the vanishing line of that plane [Criminisi et al.,2000].
4.3 Projection Image-Training View Through a Vertical Plane 57 Figure 4.4: Camera and Scene Calibration: 2 homography matrices are calculated from manual annotations: Hgcharacterizing the mapping between the ground plane in the image (red dashed line) and the real world ground plane Πgd (upper left) and Hhrelating the head plane in the image (blue solid line) with Πgd. The vertical vanishing point vZand the horizontal vanishing line are also computed using the straight lines from walls and floor observed in the scene. due to the lack of stability of the camera support. A little change in orientation has a great influence in the image coordinates, and therefore invalidates previous calibration. However, if the camera is not changed in position, or position change is small with respect to the depth of the observed scene, the homography can easily be re-calibrated automatically. An automatic method to compute homographies and line matching between image pairs like the one presented in [Guerrero and Sag¨u´es,2003] can then be used. At the moment, however, this has not been included in our system.. 4.3 Projection Image-Training View Through a Vertical Plane As demonstrated in [Kale et al.,2003], for objects far enough from the camera, we can approximate the actual 3D object as being represented by a planar object. In other words, a person can be approximated by a planar object if he or she is far enough from the camera 2. As shown in [Riklin-Raviv et al.,2007], in the presence of perspective distortion neither similarity nor affine model provide reasonable approximation for the transformation between a prior shape and a shape to segment. Riklin-Raviv et al. [2007] demonstrate that a planar projective transformation is a better approximation even though the object shape contour is roughly planar. Following these two observations, we propose to find a projective transformation, i.e. a homography, between training and testing camera views to compensate for the effect of both discretization along θand variations along ϕ, thus alleviating the effect of perspective distortion 2This hypothesis is obviously not strictly true as it does not depend solely on the distance to the camera but also on the pose and orientation of the person w.r.t. the camera
58 Chapter 4. View-invariant Motion Analysis using View-based Models on silhouette-based human motion analysis. The 3 ×3 transformation PI2ΠI1between two images I1and I2through a vertical plane Π observed in both images can be obtained as the product of 2 homographies defined up to a rotational ambiguity. The first one, HΠ←I1, projects the 2D image points in I1to the vertical plane Π and the other one, HI2←Π, relates this vertical plane to the image I2. We thus obtain the following equation that relates the points x1from I1with image points x2from I2: x2=PI2ΠI1·x1,(4.5) where x1,x2∈P2and with: PI2ΠI1=HI2←Π·HΠ←I1 =HI2←Π·(HI1←Π)−1.(4.6) The two homographies HI1←Πand HI2←Πcan be computed from the vanishing points of the 3D directions spanning the vertical plane Π i.e. the vertical Z-axis and a reference horizontal line G= Π ∧Πgd, intersection of Π and ground plane Πgd: HI←Π= [vGαvZo],(4.7) where vGand vZare the vanishing point along the horizontal and vertical axis in I,ois the origin of the coordinate system and αis a scale factor (see Appendix A). In the same way, we now want to relate 2 images, e.g. training and testing images, observing two different calibrated scenes with 2 different subjects performing the same action from two similar viewing angles. These images can potentially be related through a vertical plane centered in the human body following Eq. 4.5 . The problem is to select the vertical plane that will optimize the 2D shape correspondence between the 2 images. We choose to select this vertical plane in the training image, where the azimuth angle θis known and the camera is in an approximately horizontal position (i.e. elevation angle ϕ≈0), and consider the closest vertical plane centered on the human body: if a camera view Φ is defined by its azimuth and elevation angles (θ, ϕ) on the viewing hemisphere (Fig. 4.2a), the closest vertical plane Π is the plane defined as (θ, 0). (a) (b) Figure 4.5: Projection on the vertical plane: examples of original and warped images resulting from applying the homography HΠv←Φvfor frontal (a) and “rear-diagonal” (b) views of the Mobo dataset. Thus, considering a set of training views {Φv}Nv v=1, the associated homographies {HΦv←Πv}Nv v=1 relating each view and its closest vertical plane Πvcentered on the human
4.5 Experiments 65 (a) (b) (c) (d) (e) Figure 4.10: Two examples of view-based shape registration and pose estimation are presented for view RD1 (up) and F(down). In both cases, the original image (a) is warped to the corresponding plane (b). The foreground information (c) is used to apply the view-based pose-silhouette model leading to the estimation of both 2D shape and 2D skeleton in the projected image (d). Shape and 2D pose can then be back-projected to the original image plane (e). 4.5 Experiments For the evaluation of the framework, the gait sequences from Sec. 4.3.1 are processed. As we expected, since the orientation is estimated from the direction of motion, the system fails with stationary cases. In Fig.4.11, we present the silhouettes and 2D poses that have been extracted from the sequence presented in the Walk3: we can observe how the direction of motion slowly changes along the sequence and how the images are projected on the selected model plane. The resulting shapes are not perfect but, given the complexity of the task (low resolution and high perspective effect), we find them reasonably good. However, while the 2D pose is well estimated during most of the presented sequence (when the viewpoint is lateral or diagonal), we can see that the sequentiality of the motion is lost in the last third of the sequence. As in Sect. 3.5.1.2 where we observed a similar effect with the “WalkiCircle” sequence, this is due to the very low shape variability in the back view where it is very complicated to distinguish a state from another. We observe that acceptable results are obtained with single walking subjects. However, the reliability of the warping, and consequently the accuracy of the silhouette and pose estimate, seem to strongly depend on the precision with which both ground plane position (X,Y) and orientation θare estimated. 4.5.1 Numerical Evaluation of the Effect of Noise To numerically evaluate this dependence, we conduct a series of simulations using a set of testing ground truth poses {kGT 1···kGT NGT }and a set of sampled training poses {{kv i}NT i=1}Nv v=1 (i.e. NTposes for each training view Φv). Each pose is made of 13 hand-labelled 2D joints: k= [xk1, ...., xk13 ]∈R2×13. For each tested frame t∈ {1, NGT }, we compute the projective transformation PIΠvΦvusing ground truth location (Xt, Yt) and viewpoint θtfrom above with
66 Chapter 4. View-invariant Motion Analysis using View-based Models Figure 4.11: Shape and 2D pose estimated using our tracking framework for the Walk3 sequence. For each presented frame, we show (from right to left): foreground image projected on the training plane and processed with our view-based model, image projected on the training plane with estimated shape and 2D pose, image projected on the vertical plane with estimated shape and 2D pose and representation of shape and 2D pose backprojected in the original image.
4.6 Conclusions 67 additive Gaussian white noises (ηXY and ηθof variance σ2 XY and σ2 θrespectively) and align the NTtraining poses {kv 1···kv NT}from the selected viewpoint Φv, obtaining {kHom 1,t ···kHom NT,t } with ∀i∈ {1, NT}: xHom kj,i,t =PIΠvΦv·xkj,i,t,∀j∈ {1,13}.(4.19) We then compute the average pose error over the testing set taking the closest aligned pose for each frame t: Hom =1 NGT NGT X t=1 min i∈{1,NT}dk(kGT t,kHom i,t ),(4.20) where dk, defined as: dk(k,k0),1 13 13 X j=1 ||xkj−x0 kj|| (4.21) is the average Root Mean Square Error over the 13 2D-joints (called RMS 2D Pose Error from now on). We repeat the same operation considering a simple Euclidean 2D similarity transformation Tto align training poses to the tested images and compute: Sim =1 NGT NGT X t=1 min i∈{1,NT}dk(kGT t,kSim i,t ),(4.22) where kSim = [xSim k1, ...., xSim k13 ]∈R2×13 with: xSim kj=T·xkj,∀j∈ {1,13}.(4.23) The similarity is defined as: T·x=u+sR (γ)·x,∀x∈R2,(4.24)in which (u, γ, s) are offset, rotation angle and scaling factor respectively. These parameters are readily calculated using head center xHand “feet” location on the ground floor xFin training and testing images. The results obtained when varying σXY and σθare given in Fig. 4.12. The average pose error almost linearly increases with increasing localization noise ηXY for both alignment methods, slightly more for the proposed homographic alignment (Fig. 4.12a). A slight noise in the viewpoint estimation σθ≤π 16 does not seem to affect any of the 2 alignment methods (Fig. 4.12b). However, while the error with similarity seems to linearly increase with increasing viewpoint noise ηθfor higher noise levels, the effect is much more pronounced for the projective alignment. By augmenting σθ, we slowly increase the possibility of picking the wrong view which has more important consequences when a homography is employed between training and testing view planes instead of a simple Euclidean transformation. The benefit of using a homographic alignment rapidly decreases with the amount of added noise in viewpoint estimation σθ≥π 8and the error even gets larger than the one obtained with a similarity transformation for σθ≥π 6. 4.6 Conclusions The view-invariant approach we have proposed in this chapter can be applied to any type of 2D model or exemplar-based technique. The viewing angle is discretized into a finite number of training viewpoints and a framework of view-based models is constructed. Then, when processing a sequence, the adequate training view is selected by estimating the orientation on
68 Chapter 4. View-invariant Motion Analysis using View-based Models (a) (b) Figure 4.12: Effect of noise on 2D pose estimation: the average RMS 2D pose error (in pixels) is computed over a set of manually labelled testing poses and a set of training poses aligned using homographic (Hom) and similarity (Sim) alignments. The results are obtained varying the variance of the additive Gaussian white noise which has been added to (a) the ground truth location (X, Y ) (in cm) and (b) the viewpoint angle θ(in radians). the ground plane. The viewpoint correspondence is thus established by projecting the input image onto this training plane and finally the selected view-based model is employed for feature extraction in the warped image. To ensure a reliable estimation of both ground plane position and motion direction, indispensable for obtaining the right warping, we have developed a viewinvariant head-based tracker. Acceptable results have been obtained for sequences with a single walking subject but we identified two main drawbacks: 1) the employed model does not handle pose ambiguities and does not recover from drifting and, 2) more importantly, the result greatly depends on the accuracy achieved when estimating both location (X, Y ) and orientation θ. We have numerically evaluated the effect of noise on pose estimation. Our framework performs sufficiently well when an accurate estimation of both ground plane location and orientation (i.e. viewpoint) can be made but, with high levels of noise, the effect on pose estimation is much more pronounced for our proposed projective alignment. These results explain why the estimation of the viewpoint θ from the ground plane trajectory is not satisfactory for our purpose. Even if it gives interesting results in case of constant speed motions, this method is not accurate enough and too sensitive to noisy measurements (e.g. with partial occlusions). Moreover, the estimation of the orientation from the direction of motion does not allow working with stationary cases. Therefore, in the next chapter, we will propose a tracking framework with a stochastic approach for estimating both location and viewpoint, and search for the optimum projective transformation by sampling multiple possible values for θat multiple locations (X, Y ).
5 View-invariant 3D Pose Tracking 5.1 Introduction The goal in this chapter is to track and estimate the 3D pose of multiple walking people by means of view-based models independently of the point of view from which the scene is observed. As in the previous chapter, we will consider a discrete set of training views and exploit projective geometry to find view-invariance. We propose a stochastic approach for estimating both ground plane location and camera viewpoint, and improve the search of the optimum projective transformation for pose recognition by sampling multiple possible values for θat multiple locations (X, Y ). Applying a different projective transformation to the input image for each sampled triplet (X, Y, θ) and processing each resulting warped image in a bottomup manner as we did in the previous chapter would be computationally inefficient. We instead consider a stochastic top-down approach for body pose (and associated appearance descriptor). The idea is to learn mappings between a body pose manifold and the 2D silhouette features (shape). For each triplet, a pose is then sampled on this manifold and the corresponding shape is evaluated in the original input image using the inverse homographic transformation. This approach will make the system more robust to possible drifting compared to the model employed previously in this thesis as it can maintain multiple hypotheses through time. 5.1.1 Related Work 3D pose tracking. Stochastic models have come to be the dominant way of approaching the problem of articulated 3D human body tracking: an approximate inference technique, usually a particle filtering, is used to tractably estimate the high-dimensional posture space [Deutscher and Reid,2005,Canton-Ferrer et al.,2011,Li et al.,2010,Chang and Lin,2010,Elgammal and Lee,2009,Lee and Elgammal,2010,Jaeggli et al.,2009]. Particle filtering allows modeling non-Gaussian multi-modal distributions and can maintain multiple hypotheses through time. However, the number of particles required to achieve an acceptable result considerably increases with the dimensionality of the search space. The number of degrees-of-freedom (generally more than 30) and the high dimensionality of the state space (i.e. valid poses) make the tracking problem computationally difficult. The search space gets even larger when the tracking algorithm also has to estimate the location, orientation and scale of the subject in the image or in the scene as in [Jaeggli et al.,2009]. Some work has investigated the use of learnt models of human motion to constrain the search in state space by providing strong priors on motion [Ning et al.,2004b,Urtasun et al.,2006a]. Others have focused their research on the problem of dimensionality reduction for pose tracking and proposed to use low dimensional embedding of
70 Chapter 5. View-invariant 3D Pose Tracking human motion data: Gaussian process latent variable model (GPLVM) [Urtasun et al.,2006b, Ek et al.,2008,Andriluka et al.,2010], Locally Linear Embedding (LLE) [Jaeggli et al.,2009], supervised manifold learning [Elgammal and Lee,2009,Lee and Elgammal,2010] or coordinated mixture of factor analysers [Li et al.,2010] are some examples. Most existing systems [Elgammal and Lee,2009,Jaeggli et al.,2009,Andriluka et al.,2010] typically assume that the camera axis is parallel to the ground i.e. elevation angle ϕ= 0 (see Fig. 4.2a for angle definition) and that the observed people are vertically displayed i.e. rotation angle γ= 0. They discretize the viewpoint in a circle around the subjects, selecting a set of values for the azimuth θ: 36 orientations in [Jaeggli et al.,2009], 16 in [Rosales and Sclaroff,2006], 12 in [Elgammal and Lee,2009] and 8 in [Zhang et al.,2005b,Lan and Huttenlocher,2004,Andriluka et al.,2010]. Results have been presented using different testing datasets in laboratory environments like HumanEva [Sigal et al.,2010] or challenging street views as in [Andriluka et al.,2010,Jaeggli et al.,2009], but generally training and testing images are captured from a similar environment or with a similar camera tilt angle. Very few present numerical evaluation of human pose tracking on surveillance scenario with low resolution and high perspective distortion, and few pose tracking algorithms exploit the key constraints provided by scene calibration which is available in a large number of real surveillance system. Zhao et al. [2008] presented tracking results in crowded video-surveillance sequences using a coarse 3D model but no body pose was estimated. 5.1.2 Overview We tackle the problem of view-invariant 3D body pose tracking and explore the use of projective shape matching in a particle filtering framework which jointly explores a low dimensional pose-viewpoint manifold and the real world ground plane. Our approach is motivated by the encouraging preliminary results for view-invariant human motion analysis obtained in the previous chapter and in our earlier work [Rogez et al.,2006a] and the recent advances in lowdimensional manifold learning for human pose tracking [Jaeggli et al.,2009,Elgammal and Lee,2009]. Given its proven effectiveness, we choose to model 3D walking poses using a low dimensional torus manifold for camera viewpoint and pose as in [Elgammal and Lee,2009]. We map this manifold to our view-based silhouette manifolds using kernel-based regressors, which are learnt using a Relevance Vector Machine (RVM). Given a point on the surface of the torus, the resulting generative model can regress the corresponding pose and view-based silhouette as illustrated in Fig. 5.2b. During the online stage, 3D body poses are thus tracked using a recursive Bayesian sampling conducted jointly over the scene’s ground plane and this pose-viewpoint torus manifold. For each sample, the homography that relates the corresponding training plane to the image points can be calculated using the dominant 3D directions of the scene, the sampled location on the ground plane and the sampled camera view as explained in the previous chapter. Each regressed silhouette shape is then projected using the projective transformation and matched in the image to estimate its observation likelihood. Our tracking framework is depicted in Fig. 5.1. The rest of the chapter is organized as follows: in Sect. 5.2, we introduce the torus manifold for pose and appearance modeling. In Sect. 5.3, we detail our tracking framework. Experimentations with qualitative and quantitative evaluations are presented in Sect. 5.4 and some conclusions are finally drawn in Sect. 5.5.
5.2 Torus Manifold for Pose and Appearance Modeling 71 Figure 5.1: System Flowchart: the 3D body poses are tracked using a recursive Bayesian sampling conducted jointly over the scene’s ground plane (X, Y ) and the pose-viewpoint (θ, µ) torus manifold ([Elgammal and Lee, 2009]). For each sample n, a projective transformation relating the corresponding training plane and the image points is calculated using the dominant 3D directions of the scene, the sampled location on the ground plane (X(n) t, Y (n) t) and the sampled camera view θ(n) t. Each regressed silhouette shape s(n) tis projected using this homographic transformation obtaining s0(n) twhich is later matched in the image to estimate its likelihood and consequently the importance weight. A state, i.e. an oriented 3D pose in 3D scene, is then estimated from the sample set. 5.2 Torus Manifold for Pose and Appearance Modeling Full body pose configurations are necessarily high dimensional; in our case, we use 13 3D-joint locations in a human-centered coordinate system for our representation which results in a 39dimensional pose configuration. To reduce the problem of high dimensionality in the learning stages, a dimensionality reduction step is needed to identify a low-dimensional embedding of the pose space. As we focus on the walking action which is cyclic, we consider a low-dimensional manifold embedding both camera viewing angle and body pose together and jointly model them by means of a torus manifold. Elgammal and Lee [2009] numerically demonstrated with experimental evaluation that the supervised torus embedding shows much better performances than unsupervised manifold representations (LLE, Isomap, GPLVM). If µ∈[0,1) is the body pose configuration on the torus and θ∈[−π, π] is the viewing angle 1, then the torus manifold illustrated in Fig. 5.2b can be defined parametrically in Euclidean space by: x= (R+rcos 2πµ) cos θ, y= (R+rcos 2πµ) sin θ, z=rsin 2πµ (5.1) where Ris the “major radius”, i.e. the distance from the center of the tube to the center of the torus, and ris the “minor radius”, i.e. the radius of the tube. 5.2.1 Body Pose Modeling The training sequences are mapped onto the surface of the torus. We refer the reader to [Elgammal and Lee,2009] for details. We learn the mapping back to the original data space from the torus manifold with a kernel regressor: K=fp(µ, θ) = WpΦp(x, y, z),(5.2) 1Note that for consistency with previous sections, we keep the viewpoint parameter as an angle while Elgammal and Lee [2009] define both viewpoint and action parameter in [0,1) space.
72 Chapter 5. View-invariant 3D Pose Tracking (a) (b) Figure 5.2: (a) Data used in this chapter: example of a training 3D pose and its 8 view-based 2D silhouettes and 2D poses extracted from the MoBo dataset. (b) Pose-viewpoint torus manifold (adapted from [Elgammal and Lee,2009]) learned using the Mobo dataset. The 2 dimensions of the surface represent gait cycle and camera viewpoint. We represent 8 different views of a same pose (blue circle), and 6 different poses from a same viewpoint (green circle). where K∈R3×13 is the orientated body pose configuration in the original 3D pose training space, Φpis a vector of kernel functions and Wpis a matrix of weights2. The matrix Wpis learnt using a Relevance Vector Machine (RVM). We use radial basis functions as the kernel functions in Φpcomputed at the training data locations. Any point (µ, θ)∈[0,1) ×[−π, π], on the surface of the torus, can be directly mapped to an oriented 3D pose. 5.2.2 Appearance Modeling Different shape representations have been used for human silhouettes in recent literature, including parametric B-splines [Isard and Blake,1998], shape context [Belongie et al.,2002, Mori and Malik,2006], level-sets [Cremers,2006], pose-adaptive shape descriptors [Lin and Davis,2010] and distance transform [Jaeggli et al.,2009,Elgammal and Lee,2009]. Again, we select the landmark parameterization [Baumberg and Hogg,1994,Siebel and Maybank,2002b, Giebel et al.,2004], i.e. a set of Nl2D-landmarks, to represent the silhouette: s= [xs1, ...., xsNl]∈R2×Nl.(5.3) Although non-linearity and normalization issues can appear during the training phase as we discussed in chapter 2, landmark-based shape representations are lower dimensional and much simpler to manipulate and transform. They also facilitate a very quick matching with the image making them ideal in a top down particle filtering framework. 2As done in [Elgammal and Lee,2009], we map from the Euclidean space where the torus lives and not from the coordinate system (µ, θ) since this coordinate system is not continuous at the boundary.
5.3 Recursive Bayesian Sampling 73 We now model the generative mapping from embedded pose µto silhouette descriptors s that allows us to predict image appearance given an hypothesis for the pose µand for the body orientation or camera viewpoint θ. In this work, the viewing hemisphere is discretized into a finite number Nvof training viewpoints {θv}Nv v=1 varying the azimuth angle (see example in Fig. 5.2a). For each training viewpoint a mapping is learnt from the torus manifolds to the corresponding view-based silhouette manifold, which are learnt using a Relevance Vector Machine (RVM): s=fs(µ, θ),∀µ∈[0,1) ,∀θ∈ {θv}Nv v=1,(5.4) with fs(µ, θ) = WsΦs(x, y, z).(5.5) Once again, the mapping fs(µ, θ) is learnt using RVM with weights Wsand kernel functions Φs(µ, θ). Given a point (µ, θ)∈[0,1) × {θ1,··· , θNv}on the torus manifold, the resulting generative model can generate the corresponding view-based silhouette. Note that in this work, the shape descriptor sis augmented with the 13 2D-joints k= [xk1, ...., xk13 ]∈R2×13 to facilitate the estimation of a 2D pose error in the experiment Section. 5.3 Recursive Bayesian Sampling 5.3.1 Formulation. At each time step, we simultaneously perform body pose estimation and image localisation since both processes can benefit from the coupling of the posture and image location as demonstrated in [Jaeggli et al.,2009]. As explained in Sect. 4.2.2, the advantage of assuming a calibrated environment and a planar ground plane is the considerable reduction of the search space as image location, scale and rotation can be recovered from the real world ground plane location. Thus, we define the state vector of the target as: χt= [XtYtθtµt],(5.6) consisting of the real-world ground plane location (Xt, Yt) and the embedding coordinates on the torus surface (µt, θt)∈[0,1) ×[−π, π]. The calibration of the scene and the torus embedding help us to face a much more tractable problem as the search has to be performed in a 4dimensional state space while, for instance, Jaeggli et al. [2009] explore a 10-dimensional space. We formulate the tracking problem as a Bayesian inference task, where the state of the tracked subject is recursively estimated at each time step given the evidence (image data) up to that moment. Formally, within the Bayesian filtering framework, we formulate the computation of the posterior distribution p(χt|It) of our model parameters χtover time as follows: p(χt|It)∝p(It|χt)p(χt|It−1),(5.7) where Itis the image sequence up to time tand p(It|χt) is the likelihood of observing the image Itgiven the parameterization χtof our model at time t, in other words the observation density. Finally p(χt|It−1) is the a priori density, which is the result of applying the dynamic model p(χt|χt−1) to the a posteriori density p(χt−1|It−1) of the previous time step: p(χt|It−1) = Zp(χt|χt−1)p(χt−1|It−1)dχt−1.(5.8) Unfortunately, when the involved distributions are non-Gaussian, Eq. (5.7) cannot be solved analytically. Instead, we use a particle filter [Isard and Blake,1998,Sidenbladh et al.,2002,
74 Chapter 5. View-invariant 3D Pose Tracking Deutscher and Reid,2005] in order to approximate the true posterior pdf p(χt|It) by means of a discrete weighted set of samples {χ(n) t, π(n) t}N n=1: p(χt|It)≈ N X n=1 π(n) tδ(χ(n) t),(5.9) where for each particle, δdenotes the Dirac delta and π(n) tis the normalized importance weight which is directly derived from measurement likelihood: π(n) t=p(It|χ(n) t) PN n0=1 p(It|χ(n0) t),(5.10) as defined in [Isard and Blake,1998].Hence, whilst the likelihood function decides which particles are worth propagating, the dynamic model is responsible for guiding the exploration through the state space. 5.3.2 Dynamic Model. Since a static camera is being considered in this work, and assuming the people face along the direction of motion, we model the dependence of viewpoint on ground plane location while we assume statistical independence between the remaining state variables3. The dynamic model p(χt|χt−1) is thus a product of four dynamic models, i.e. p(χt|χt−1) = p(Xt|Xt−1, θt−1)p(Yt|Yt−1, θt−1)p(θt|θt−1)p(µt|µt−1). Therefore, our state model has the following form on the torus manifold: θt=θt−1+nθ,(5.11) µt ˙µt=1δt 0 1 µt−1 ˙µt−1+nµ n˙µ,(5.12) where nθ,nµand n˙µare zero mean white Gaussian noises (whose variances are set to σθ=π 10 , σµ= 0.075 and σ˙µ= 0.0125 respectively) and δt the time interval between successive frames, and on the ground plane: Xt Yt=Xt−1 Yt−1+nVVX VY+nXY 1 1(5.13) where nXY and nVare zero mean white Gaussian noise (with variance set to σXY = 1cm and σV= 10cm in our experiments) and [VX, VY]T=V/||V|| is the unit orientation vector relating θand the camera viewing direction C(see Eq. 4.8 and Eq. 4.10). In this way, we model the fact that pedestrians are more likely to move in the facing direction and, after a stationary phase, we can predict in which way the subject is going to move based on his body orientation (i.e. viewpoint angle θ). 5.3.3 Image Measurements - Observation Model. The likelihood function p(I|χ) computes how likely it is to observe the image Igiven the unknown state χ. Given the viewpoint θand the location on the ground plane (X, Y ), a 3We choose not to model the dependencies between the gait parameter µand the ground plane location because stride length depends on the subject morphology and walking style.
5.4 Experimental Results 81 plane (i.e. γ= 0 in Eq. 4.5.1)5. Thus, a comparison of the performances of our framework replacing the projective transformation in Eq. 5.14 by a similarity transformation will provide a quantitative evaluation of the improvement achieved by our proposal w.r.t. state-of-the-art 6. 5.4.1 Settings and Parameters. We use training silhouettes and 2D/3D poses extracted from the MoBo dataset [Gross and Shi,2001] illustrated in Fig. 5.2a: for each one of the 8 training views, 15 walking cycles corresponding to 15 different subjects are temporally aligned, subsampled and averaged to compute a mean walking cycle made of 100 silhouettes and 2D poses. Thus, 800 silhouettes and associated 2D poses are used to learn the mapping between the torus manifold and the original data space. The same operation is performed with the 3D poses which are rotated around the Z-axis (azimuth θ) to cover the entire torus manifold. For each training view, we localize the horizontal and vertical vanishing points and compute the 8 homographies {HΦv←Πv}8 v=1. We remove lens distortion from Caviar testing images and calibrate the camera w.r.t. the scene by localizing the vertical vanishing point and the horizontal vanishing line, and compute the homographies Hgand Hhfrom manual annotations. In this work, we do not address the detection problem and take the ground plane location in the first frame from ground truth data, but the detector from [Li et al.,2008] would perfectly suit our framework as it can deal with perspective distortion and has shown significantly improved detection performance on the Caviar dataset. When a subject appears in the scene, we initialize a tracker by sampling in the entire space of possible poses and probable viewpoints. Supposing that the subject is facing in the direction of motion, the most probable viewpoints can defined based on the location in the first frame and the scene knowledge: e.g. if the subject is entering the scene by the right side, the viewpoint is most likely to be L1,D1or RD1and viewpoint should be sampled in the corresponding part of the torus. Note that, during tracking, the particles which fall in non-valid areas of the ground plane (such like walls, plants, etc) are assigned a 0 likelihood. When a subject is not moving, the likelihood is computed using only the shape landmarks corresponding to the upper part of the body. Since we model walking poses, our framework is not supposed to recognise standing poses. When motion is detected after a stationary phase, the tracker is reset by sampling in the entire space of possible poses. 5.4.2 Experiments. We ran a series of tests on the selected Caviar sequences varying the number of particles in the filter. Since randomness is involved in the re-sampling of the particles, to gain better statistical significance, we perform the same experiments 20 times and from now on we compute numerical result as the average over these 20 runs. We repeat the same operation using a similarity transformation instead of our homographic transformation. First we propose to evaluate the performance of the tracker, independently of the state estimator. We thus consider that a target has been lost and the localization is not valid if the minimum distance (in the particle set) to ground truth location exceeds a certain value. We believe that a pose estimation does not make sense if the nearest particle is 1 meter away from the target true location, 5Most of these techniques are based on a scheme where the images are scanned with a sliding window at different scales. 6The similarity transformation is computed using 2 reference points in image and model view: head center and ground floor location recovered from the real world ground plane coordinate (X, Y ) using Hhand Hgfrom Sect. 4.2.2.
82 Chapter 5. View-invariant 3D Pose Tracking i.e. minn∈{1,N}||bχGt t−χ(n) t||gd≥100 cm. A track is then considered lost when then the target has been lost during 20 frames or more and has not been recovered in the last frame of the sequence. Figure 5.7: Percentage of lost tracks vs number of particles for similarity and homographic alignment. we present the average performance over 20 runs of the tracking algorithm on the 11 sequences: a track is considered lost when the tracking has failed during 20 frames or more (the distance between the nearest particle and ground truth location is over 1 meter) and it has not recovered by the end of the sequence, i.e. in the last frame the subject is still one meter away from ground truth for the nearest particle. Results show that the proposed homographic alignment reduces the average percentage of lost tracks as can be observed in Fig. 5.7. The percentage of lost tracks decreases with the number of particles employed in the filter for both methods, but we reach 0% of lost tracks with 1000 particles and over while 5% of the tracks are still lost when considering 2000 particles and a similarity transformation. The perspective correction allows for better shape matching and consequently a more efficient shape-based tracking. If we look at the detailed results given in Tab. 5.2 and Tab. 5.3, we can observe how, on average, the homographic alignment outperforms the similarity alignment for 7 of the 11 tested tracks, i.e. it reaches 0% tracking failure with a smaller particle set (tracks 1, 2, 6, 7, 8, 9 and 10), while performing similarly for 3 other tracks (4, 5 and 11). Good tracking performance requires larger particle sets for the sequences presenting occlusions and multiple interacting people. We can also see that much fewer particles are required when a good foreground detection is available: a perfect result is obtained with our projective alignment method and only 20 particles in sequences where people wear dark clothes. If we compute the average number of valid localizations, i.e. the cases where the distance between the nearest particle and ground truth location is below 1 meter, the tracking usually looses less targets when a homographic alignment is used rather than a similarity alignment, (see Fig. 5.8a). We even reach an average of 99% of valid localizations above 1000 particles. We now evaluate the pose estimation performance and compute a RMS error between the 2D pose of the best particle and the ground truth 2D pose, thus evaluating the best pose that could be estimated from the posterior independently of the employed state estimator. In Fig. 5.8b, we see how the 2D pose error (RMSE between nearest particle and ground truth) decreases
5.4 Experimental Results 83 Table 5.2: Percentage of tracking failure using a similarity transformation for shape alignment. For each of the 11 selected tracks (see details in Tab. 5.1), we present the average performance over 20 runs of the tracking algorithm for different number of samples: a track is considered lost when tracking has failed during 20 frames or more (the distance between the nearest particle and ground truth location is over 1 meter) and it has not recovered by the end of the sequence. Alignment Similarity No. Particles 20 50 100 250 500 1000 2000 Track 1 5 0 0 0 0 0 0 (no 2 95 85 60 5 0 0 0 occlusion) 3 45 0 0 0 0 0 0 4 0 0 0 0 0 0 0 5 0 0 0 0 0 0 0 6 10 0 0 0 0 0 0 Track 7 65 40 25 40 30 25 20 (with 8 35 5 10 5 0 0 0 occlusions) 9 100 70 45 30 10 5 10 10 35 30 15 30 10 5 10 11 20 0 0 0 0 0 0 Average 37.27 21.36 14.09 9.55 4.55 3.18 3.64 Table 5.3: Percentage of tracking failure using the proposed homographic projection for shape alignment. For each of the 11 selected tracks (see details in Tab. 5.1), we present the average performance over 20 runs of the tracking algorithm for different number of samples: a track is considered lost when tracking has failed during 20 frames or more (the distance between the nearest particle and ground truth location is over 1 meter) and it has not recovered by the end of the sequence. Alignment Homography No. Particles 20 50 100 250 500 1000 2000 Track 1 0 0 0 0 0 0 0 (no 2 95 60 20 0 0 0 0 occlusion) 3 40 5 0 0 0 0 0 4 0 0 0 0 0 0 0 5 0 0 0 0 0 0 0 6 0 0 0 0 0 0 0 Track 7 50 45 10 0 0 0 0 (with 8 40 0 0 0 0 0 0 occlusions) 9 95 65 45 20 15 0 0 10 10 15 10 15 15 0 0 11 40 0 0 0 0 0 0 Average 33.64 16.82 7.73 3.18 2.73 0 0
84 Chapter 5. View-invariant 3D Pose Tracking (a) (b) (c) Figure 5.8: Tracking Results. Average Performance over 20 runs of the tracking algorithm on the 11 tracks for similarity and homographic alignment vs number of particles. (a) Percentage of valid localizations (valid if the distance between nearest particle and ground truth location is below 1 meter), (b) 2D pose error over all the poses and (c) 2D pose error computed using only valid poses from (a). The pose error in (b) and (c) is computed as the MRSE between nearest particle and ground truth using the 13 2D-joints in pixels. with the number of particles and how the framework, again, performs better when a projective transformation is used and allows for a more accurate pose estimation. If we compute the same error using only the valid localizations (from Fig. 5.8a,) we reach lower 2D pose errors, especially for small particle sets and for the similarity based approach (see Fig. 5.8c). This makes sense because of the larger amount of failed localizations which return a bad pose estimation and influence the average pose error. To aid the comparison of pose estimation performance and focus on the pose estimation when localisation is satisfactory, from now on, we exclude the non-valid poses and the different errors are computed over poses from valid localizations only. More frames are then considered for our homography based alignment because of its lower failure rate. We can see in Fig. 5.8c that the average 2D pose error obtained using our projective transformation is not too far from the result returned with a similarity transformation despite a qualitative improvement observed when watching the estimated poses and silhouettes. This last observation inspired us to carry out a deeper analysis of the different results, in particular visualize the different rates in function of the distance between the subject and the camera. In Fig. 5.9, we present the average percentage of valid localizations, the average 2D pose error and the average ground plane location error varying the maximum distance to the camera. Results are presented for different sized particle sets. In the middle row, we can observe that the average pose error globally decreases as we augment the maximum distance to the camera and add new poses further away. The opposite happens with the ground plane location error. This is expected because when people move away from the camera their size in the image gets smaller. Thus, the 2D pose gets smaller when moving away from the camera leading to a consecutive lower 2D pose error while an accurate localization on the ground plane becomes more difficult with the distance. We should also point out that the different errors are computed using a ground truth data obtained from manual labelling and the accuracy and reliability of this labelling also decrease with the distance to the camera. From Fig. 5.9, we can clearly observe that the improvement in terms of pose and ground plane localization is globally obtained when the subjects are close to the camera. This makes perfect sense since the viewpoint changes when a subject goes far away from the camera and tends to a tilt angle ϕ= 0 which is similar to the training viewpoint employed in this thesis. It seems that when the subject moves far away from the camera, a projective transformation is not required and a similarity transformation could be enough. In Fig. 5.10, we present the average 2D pose error obtained when estimating the state at
5.4 Experimental Results 85 N= 20 N= 50 N= 100 N= 250 N= 500 N= 1000 N= 2000 Figure 5.9: Detailed performances w.r.t. the distance to the camera: percentage of valid localizations (top row), 2D pose error (middle row) and (X, Y ) ground plane localization error (bottom row) are represented (from left to right) for 20, 50, 100, 250, 500, 1000 and 2000 particles. Note that the different values are computed using the poses from 0 meter up to the given distance to the camera. Only valid localizations from the top row have been used to compute the performances in middle and bottom rows. Again 2D pose error and (X, Y ) ground plane localization error are computed as the MRSE between nearest particle and ground truth using the 13 2D-joints locations in pixels and the 2D location in cm respectively. each time step using the Monte Carlo approximation (MC), the Maximum A Posteriori (MAP) criteria, the Viterbi path finding algorithm (Viterbi) and the proposed weighted sum around the Viterbi estimate (Viterbi WS)7. We present the results for the two different types of alignment. Corresponding numerical evaluation is given in Tab. 5.4. The first observation is that our homographic alignment clearly outperforms the similarity transformation independently of the employed state estimation technique. We can also observe that our proposed approach for state estimation outperforms all the other techniques for N≥500 particles while MC estimate is better for smaller particle sets. 7Note that the state estimate used to model each subject’s 3D occupancy on the ground floor in Eq. 5.22 is always computed using the Monte Carlo approximation as we want to compare the different state estimators from the same clouds of samples.
86 Chapter 5. View-invariant 3D Pose Tracking Figure 5.10: Average 2D pose error varying the state estimator: performance over 20 runs of the tracking algorithm on the 11 tracks for homographic and similarity alignments. We present the RMS 2D pose error when estimating the state at each time step using the Monte Carlo approximation (MC), the MAP criteria, Viterbi path finding algorithm (Viterbi) and the weighted sum around the Viterbi estimate (Viterbi WS). Table 5.4: Average 2D pose error: performance over 20 runs of the tracking algorithm on the 11 tracks for similarity and homographic alignments. We present the RMS 2D pose error (in pixels) when estimating the state at each time step using the Monte Carlo approximation cχtMC , the MAP criteria cχtMAP , Viterbi path finding algorithm cχtV it and the weighted sum around the Viterbi estimate cχtV it+W S. Alignment Similarity No. Particles 20 50 100 250 500 1000 2000 State bχtMC 7.59 6.53 5.97 5.50 5.13 4.98 5.14 Estimator bχtMAP 7.66 6.62 6.05 5.63 5.30 5.21 5.38 bχtV it 7.77 6.74 6.18 5.63 5.17 4.98 5.07 bχtV it+W S 7.72 6.65 6.08 5.53 5.08 4.89 5.00 Alignment Homography No. Particles 20 50 100 250 500 1000 2000 State bχtMC 7.52 5.71 5.37 4.96 4.75 4.61 4.72 Estimator bχtMAP 7.65 5.85 5.55 5.17 4.97 4.90 5.05 bχtV it 7.78 6.00 5.66 5.16 4.84 4.67 4.73 bχtV it+W S 7.70 5.88 5.52 5.02 4.71 4.54 4.61
5.4 Experimental Results 87 (a) (b) (c) Figure 5.11: Qualitative 3D pose tracking results for tracks 1 (a), 4 (b) and 5 (c) in Tab. 5.1 using our projective method for view-invariant pose tracking and 500 particles. For each sequence, we show from top to bottom: the tracked silhouettes for a few selected frames (same frames considered in Fig4.8), the estimated viewpoint θ (vs ground truth), the estimated gait parameter µ, the trajectory of the subject on the torus manifold (µand θtogether) and the estimated 3D poses corresponding to the silhouette in the first row. In Fig. 5.11, we present qualitative results for tracks 1, 4 and 5 from Tab. 5.1 using our projective method for view-invariant pose tracking and 500 particles. Note that the same sequences were considered in Fig. 4.8. For each sequence, we can observe the tracked silhouettes for a few frames and the trajectory of the subject in the image as well as the trajectory on the torus manifold and the estimated 3D poses which have been successfully tracked. If we look at the temporal evolution of the viewpoint, we can see that using the Viterbi algorithm and our approach, we achieve a smooth continuous estimation of the viewpoint angle θwhile using a model constructed from a discrete set of training views. As in [Elgammal and Lee,2009], we recover the typical sawtooth curve of the walking cycle but in our case with challenging perspective videos. We present more qualitative results for 2 sequences with multiple interacting subjects in Fig. 5.12 (tracks 7 and 8) and Fig. 5.13 (tracks 9, 10 and 11) using our proposed method and 1000 particles. For each sequence, we show the result for a few frames: the tracked silhouettes and the trajectories in the image, the trajectories on the torus manifold and the estimated 3D poses which have been successfully tracked despite the occlusions and the perspective effect.
88 Chapter 5. View-invariant 3D Pose Tracking (Frames 40) (120) (190) (230) (270) (310) Figure 5.12: Qualitative 3D pose tracking results for the Meet WalkTogether2 sequence (tracks 7 and 8 in Tab. 5.1) with 2 interacting subjects using our projective method for view-invariant pose tracking and 1000 particles. from top to bottom we show: the tracked silhouettes and the trajectories in the image, the trajectories on the torus manifold and the estimated 3D poses. (Frames 150) (260) (300) (330) (365) (415) Figure 5.13: Qualitative 3D pose tracking results for the Meet Split 3rdGuy sequence (tracks 9, 10 and 11 in Tab. 5.1) with 3 interacting subjects using our projective method for view-invariant pose tracking and 1000 particles. from top to bottom we show: the tracked silhouettes and the trajectories in the image, the trajectories on the torus manifold and the estimated 3D poses. 5.5 Conclusions In this chapter, we have presented a complete framework for view invariant shape based 3D body pose tracking in man-made environments from monocular surveillance videos with high perspective effect. We have assumed that the camera is calibrated w.r.t. the scene and that observed people move on a known ground plane, which are realistic assumptions in surveillance scenarios. We have demonstrated that exploiting projective geometry alleviates the problems caused by roof-top cameras with high tilt angles, and have shown that using a mapping from a low dimensional pose manifold to 8 training views was enough to produce acceptable results when using a projective alignment for silhouette matching: our framework is able to track 3D human walking poses in a 3D environment exploring only a 4 dimensional state space with a particle filter.
5.5 Conclusions 89 We have conducted a series of experiments to quantitatively and qualitatively evaluate our tracking framework for a wide variety of viewing angles and a variety of sequences, some with multiple interacting subjects and occlusions. In our experimental evaluation, we have demonstrated the significant improvements of the proposed projective alignment over a commonly used similarity alignment and have provided numerical pose tracking results for the monocular sequences with perspective effect from the CAVIAR dataset. Our results demonstrate that the incorporation of this perspective correction in the pose tracking framework results in a higher tracking rate and allows for a better estimation of body poses under wide viewpoint variations.
90 Chapter 5. View-invariant 3D Pose Tracking
6.2 Sampling of Discriminative HOGs 97 on log-likelihood gradient distribution. Collins and Liu [2003] also use a log-likelihood ratio based approach in the context of adaptive on-line tracking, and select the most discriminative features from a set of 49 RGB features that separate a foreground object from the background. In our case, we have a much higher set of possible features (histogram bins) if we consider all the possible configurations of HOG blocks at all locations exhaustively. So to make the problem more tractable, rather than selecting individual features in a dense HOG vector, entire HOG blocks are randomly selected using a log-likelihood ratio derived from the edge gradients of the training data. Dimitrijevic et al. [2006] use statistical learning techniques during the training phase to estimate and store the relevance of the different silhouette parts to the recognition task. We use a similar idea to learn relevant gradient features, although slightly different because of the absence of silhouette information. In what follows, we present our method to select the most discriminative and informative HOG blocks for human pose classification. The basic idea is to take advantage of accurate image alignment and study gradient distribution over the entire training set to favor locations that we expect to be more discriminative between different classes. Intra-class and inter-class probability density maps of gradient/edge distribution are used to select the best location for the HOG blocks. 6.2.1 Formulation Here we describe a simple Bayesian formulation to compute the log-likelihood ratios which can be used to determine the importance of different regions in the image when discriminating between different classes. Given a set of classes C, the probability that the classes represented by Ccould be explained by the observed edges Ecan be defined using a simple Bayes rule: p(C|E) = p(E|C)p(C) p(E).(6.1) The likelihood term p(E|C) of the edges being observed given classes C, can be estimated using the training data edges for the respective classes. Let T={(Ii, ci)}be a set of aligned training images, each with a corresponding class label. Let TC={(I, c)∈T|c∈C}be the set of training instances for the set of classes C={ci}. Then the likelihood of observing an edge given a set of classes Ccan be estimated as follows: p(E|C) = 1 |TC|X (I,c)∈TC ∇(I),(6.2) where ∇(·) calculates a normalized oriented gradient edge map for a given image I, with the value at any point being in the range [0,1]. Note that an accurate alignment of the positive samples is required to compute p(E|C). We refer the reader to the automatic alignment method proposed for human pose in Sect. 6.4.1.1 for a solution to this problem. Class specific information is represented by high values of p(E|C) from locations where edge gradients occur most frequently across the training instances. Edge gradients at locations that occur in only a few training instances (e.g. due to background or appearance) will tend to average out to low values. To increase robustness toward background noise the likelihood can be thresholded by a lower bound: p(E|C) = p(E|C) if p(E|C)> τ , 0 otherwise. (6.3) Suppose we have a subset of classes B⊂C. Discriminative edge gradients will be those that are strong across the instances within Bbut are not common across the instances within C.
98 Chapter 6. Multi-class Pose Classifier Figure 6.3: Log-likelihood ratio for face expressions. We used a subset of [Kanade et al.,2000] composed by 20 different individuals acting 5 basic emotions besides the neutral face: joy, anger, surprise, sadness and disgust. All images were normalized, i.e. cropped and manually rectified. Gradient probability map (left) and log-likelihood ratio (right) are represented for each one of the 6 classes. Hot colors indicate the discriminative areas for a given facial expression. Using the log-likelihood ratio between the two likelihoods p(E|B) and p(E|C) gives: L(B, C) = log p(E|B) p(E|C).(6.4) The log-likelihood distribution defines a gradient prior for the subset B. High values in this function give an indication of where informative gradient features may be located to discriminate between instances belonging to subset Band the rest of classes in C. For the example given in Fig. 6.2, we can see how the right knee is a very discriminative region for this particular class (see Fig. 6.2g). In Fig. 6.3, we present the log-likelihood distributions for 5 different facial expressions. Gradient orientation can be included by decomposing the gradient map into nθseparate orientation channels according to gradient orientation. The log-likelihood Lθ(B, C) is then computed separately for each channel, thereby increasing the discriminatory power of the likelihood function, especially in cases when there are many noisy edge points present in the images. Maximizing over the nθorientation channels, the log-likelihood gradient distribution for class Bthen becomes3: L(B, C) = max θ(Lθ(B, C)) .(6.5) We also obtain the corresponding orientation map: Θ(B, C) = argmax θ (Lθ(B, C)) .(6.6) Uninformative edges from a varied dataset will generally be present for only a few instances and not be common across instances from the same class, whereas common informative edges 3L(B, C) = L(B, C) if no separated orientation channels are considered.
6.2 Sampling of Discriminative HOGs 99 (a) (b) (c) (d) Figure 6.4: Selection of discriminative HOG for facial expressions classification: details of the database are given in Fig. 6.3. In (a), we show the log-likelihood map for 2 facial expressions together: joy and disgust. After thresholding (b), we obtain the most discriminative areas where HOG blocks will be extracted (c) to differentiate joy and disgust expressions from the other classes. In (d), we can see the areas covered by the sampled HOG blocks and how the resulting density follows the distribution from (b). Results using this feature selection scheme on facial expression recognition have been reported in [Orrite et al.,2009]. for pose will be reinforced across instances belonging to a subset of classes B, and be easier to discriminate from edges that are common between Band all classes in parent set C. Even if background edges were shared by many images (e.g. if positive training samples consist of images of standing actions shot by stationary cameras such as the gesture and the box actions in the HumanEVA dataset [Sigal et al.,2010]), then these edges become uninteresting as p(E|B)≈p(E|C) and a low log-likelihood value would be returned. However, when uninformative edges (from background or clothing) are not common across all training instances but occur in the images of the same class all the time by coincidence, edges can be falsely considered as potentially discriminative. We observed that particular case when trying to work with the Buffy dataset from [Ferrari et al.,2008]: all the images for the same pose (standing with arms folded) correspond to one unique character with the same clothes and same background scene. To address this issue, we create a varied dataset of instances which is discussed more detail in Sec.6.5. Given this log-likelihood gradient distribution L(B, C), we can randomly sample HOG blocks from positions (x, y) where they are expected to be informative, thus reducing the dimension of the feature space. We then use L(B, C) as distribution proposal to drive blocks sampling (x(i), y(i)): (x(i), y(i))∼ L(B, C).(6.7) Features are then extracted from areas of high gradient probability across our training set more than areas with low probability (see Fig. 6.2h). By using this information to sample features, the amount of useful information available to learn efficient classifiers is increased. In Fig. 6.4a we represent the log-likelihood for 2 facial expressions together: joy and disgust. After thresholding (Fig. 6.4b), we obtain the most significant areas where HOG blocks will be extracted (Fig. 6.4c) to differentiate joy and disgust expressions from the other classes. In Fig. 6.4d, we can see the areas covered by the selected HOG blocks and how the resulting density follows the distribution from Fig. 6.4b. Results using this feature selection scheme for facial expression recognition have been reported in [Orrite et al.,2009].
100 Chapter 6. Multi-class Pose Classifier (a) (b) (c) Figure 6.5: Bottom-up hierarchical tree construction: the structure Sis built using a bottom-up approach by recursively clustering and merging the classes at each level. We present an example of tree construction from 192 classes (see class definition in Sect. 6.4.1.2) on a torus manifold where the dimensions represent gait cycle and camera viewpoint. The matrix presented here (a) is built from the initial 192 classes and used to merge the classes at the very lowest level of the tree. The similarity matrix is then recomputed at each level of the tree with the resulting new classes. The resulting hierarchical tree-like structure Sis shown in (b) while the merging process on the torus manifold is depicted in (c). We can observe (b) how the first initial node acts as a viewpoint classifier. 6.3 Randomized Cascades of Rejectors The classifier is an ensemble of hierarchical cascade classifiers. The method takes inspiration from cascade approaches such as [Viola and Jones,2004,Zhu et al.,2006], hierarchical template trees such as [Gavrila,2007,Stenger,2004] and Random Forests [Breiman,2001,Lepetit and Fua,2006,Bosch et al.,2007]. 6.3.1 Bottom-up Hierarchical Tree Construction Tree structures are a very effective way to deal with large exemplar sets. Gavrila [2007] constructs hierarchical template trees using human shape exemplars and the chamfer distance between them. He recursively clusters together similar shape templates selecting at each node a single cluster prototype along with a chamfer similarity threshold calculated from all the templates that the cluster contains. Multiple branches can be explored if edges from a query image are considered to be similar to cluster exemplars for more than one branch in the tree. Stenger [2004] follows a similar approach for hierarchical template tree construction applied to articulated hand tracking, the main difference being that the tree is constructed by partitioning the state space. This state space includes pose parameters and viewpoint. Inspired by these two papers, Okada and Stenger [2008] present a method for human motion capture based on tree-based filtering using a hierarchy of body poses found by clustering the silhouette shapes. Although these existing template tree techniques are shown to have interesting qualities in terms of speed, they present some important drawbacks for the task we want to achieve. First, templates need to be stored for each node of the tree leading to memory issues when dealing with large sets of templates. The second limitation is that they require that a clean silhouette or template data is available from manual segmentation [Gavrila,2007] or generated synthetically from a 3D model [Stenger,2004,Okada and Stenger,2008]. Their methodology can not be directly applied to unsegmented image frames because of the presence of too many noisy edges
6.3 Randomized Cascades of Rejectors 101 from background and clothing of the individuals which dominate informative pose-related edges. By using only the silhouette outlines as image features, the approaches in [Gavrila,2007,Okada and Stenger,2008] ignore the non-negligible amount of information contained in the internal edges which are very informative for pose estimation applications. We thus propose a solution to adapt the construction of such a hierarchical tree structure for images. Algorithm 6: Class Hierarchy Construction. input : Labeled training images. output: Hierarchical structure S. while num. of classes of the level nl>1do for each class n do Compute new L(Cn, C) (cf. §6.2); Compute the nl×nlsimilarity matrix M(cf. Eq. 6.8); Set the number of merged classes nm= 0 ; Set the number of clusters ncl = 0 ; while nm< nldo Take the next 2 closest classes (C1, C2) in M; •case 1: C1and C2have not been merged yet. Create a new cluster Cncl+1 with C1and C2:Cncl+1 =C1∪C2; Update ncl =ncl + 1 and nm=nm+ 2; •case 2: C1and C2already merged together. Do nothing; •case 3: C1has already been merged in Cr. Merge C2in Cr:C0 r=Cr∪C2; nm=nm+ 1; •case 4: C2has already been merged in Cs. Merge C1in Cs:C0 s=Cs∪C1; nm=nm+ 1; •case 5: C1∈ Crand C2∈ Cs. Merge the 2 clusters Crand Cs:C0 r=Cr∪Cs; ncl =ncl −1; Create a new level lwith new hyper-classes {C0 n}ncl n=1 ; Update the number of classes for that level nl=ncl; Update the structure S0=S ∪{C0 n}nl n=1 ; Instead of successively partitioning the state space at each level of the tree [Stenger,2004] or clustering together similar shape templates from bottom-up [Gavrila,2007] or top-down [Okada and Stenger,2008], we propose a hybrid algorithm. Given that a parametric model of the human pose is available (3D or 2D joint locations), we first partition the state space into a
102 Chapter 6. Multi-class Pose Classifier series of classes4. Then, we construct the hierarchical tree structure by merging similar classes in a bottom-up manner as in [Gavrila,2007] but using for each class its gradient map. This process only requires that the cropped training images have been aligned without the need for clean silhouettes or templates. We thus recursively cluster and merge similar classes based on a similarity matrix that is recomputed at each level of the tree (Fig. 6.5a). The similarity matrix M={Mi,j }with i, j ∈ {1,··· , nl}(being nlthe number of classes at each level) is computed using the L2 distance between the log-likelihood ratios of the set of classes Cnthat represent the classes that fall below each node nof the current level and the global edge map Cconstructed from all the classes together: Mi,j =||L(Ci, C)−L(Cj, C)||.(6.8) Using the log-likelihood ratio to merge classes reduces the effect of uninformative edges on the hierarchy construction while at the same time increasing the influence of discriminative edges. At each level, classes are clustered by taking the values from the similarity matrix in ascending order and successively merge corresponding classes until they all get merged, then go to next level. The class hierarchy construction is depicted in Algorithm 6. This algorithm is fully automatic as it does not require any threshold to be tuned, and works well with continuous and symmetrical pose spaces like the ones considered in this thesis. However, we have observed that it fails with non-homogeneous training data or in presence of outliers. In that case, the use of a threshold on the similarity value in the second while loop should be considered to stop the clustering process before merging together classes which are too different. This process leads to a hierarchical structure S(Fig. 6.5b). The leaves of this tree-like structure Sdefine the partition of the state space while Sis constructed in the feature space: the similarity in term of image features between compared classes increases and the classification gets more difficult when going down the tree and reaching lower levels as in [Gavrila,2007]. But in our case each leaf represents a cluster in the state space as in [Stenger,2004] while in [Gavrila,2007], templates corresponding to completely different poses can end-up being merged in the same class making the regression to a pose difficult or even impossible. In [Gavrila,2007] the number of branches are selected before growing the tree, potentially forcing dissimilar templates to merge too early. while in our case, each node in the final hierarchy can have 2 or more branches as in [Stenger,2004,Okada and Stenger,2008]. Instead of storing and matching an entire template prototype at each node as in [Gavrila, 2007,Stenger,2004,Okada and Stenger,2008], we now propose a method to build a reduced list of discriminative HOG features, thus making the approach more scalable to the challenging size and complexity of human pose datasets. 6.3.2 Discriminative HOG Blocks Selection While other algorithms (PSH, RVMs, SVMs, etc) must extract the entire feature space for all training instances during learning, making them less practical when dealing with very large datasets, our method for learning hierarchical cascades only selects a small set of discriminative features extracted from a small subset of the training instances at a time. This makes it much more scalable for very large training sets. For each branch of our structure S, we use our algorithm for feature selection to build a list of potentially discriminative features, in our case a vector of HOG descriptors. HOG blocks need only to be placed near areas of high edge probability for a particular class. Feature sampling will then be concentrated in locations that are considered discriminative following the 4Two methods have been implemented to obtain the discrete set of classes: the torus manifold discretization (see Sect. 6.4.1.2) and 2D pose space clustering (see Sect. 6.4.2.1).
6.3 Randomized Cascades of Rejectors 103 discriminative log-likelihood maps (discussed in Sect. 6.2). For each node, let the set of classes that fall under this node be Cnand the subset of classes belonging to one of its branches bbe Cb. Then for each branch b,nHlocations are sampled from the distribution L(Cb, Cn) (as described in Sect. 6.2.1) to give a set of potentially discriminative locations for HOG descriptors: Hp={(x(i), y(i))}nH i=1 ∼ L(Cb, Cn).(6.9) For each of these positions we sample a corresponding HOG descriptor parameter: HΨ={ψ(i)}nH i=1 ∈Ψ (6.10) from a parameter space Ψ = (W×B×A) where W={16,24,32}is the width of the block in pixels, B={(2,2),(3,3)}are the cell configurations considered and A={(1,1),(1,2),(2,1)} are aspect ratios of the block. An example is proposed in Fig. 6.6a for a 16 ×16 HOG block with 2 ×2 cells. Figure 6.6: Example of a selected HOG block from B: we show the location of the 16×16 block (2×2 cells) on top of the log-likelihood map for the corresponding branch of the tree (a). We represent in (b) the 2 distributions obtained after training a binary classifier: the green distribution corresponds to the set of training images T+ b that should pass through that branch while the red one corresponds to the set T− bthat should not pass. In this example, fi(hi) is the projection of the block hion the hyperplane found by Support Vector Machines (SVM). Finally, we represent in (c) the ROC curve with True Positive (TP) vs False Positive (FP) rates varying the decision threshold. We give the precision and recall values for the selected threshold (red dotted line). Next, at each location (x(i), y(i))∈HpHOG features are extracted from all positive and negative training examples using corresponding parameters ψ(i)∈HΨ. For each location a positive set T+ bis created by sampling from instances belonging to Cband a negative set T− b
104 Chapter 6. Multi-class Pose Classifier Algorithm 7: HOG Blocks Selection input : Hierarchical structure S, and training images. Discriminative classifier g(·). output: List of discriminative HOG blocks B. for each level l do for each node n do Let Cn= set of classes under n; for each branch b do Let Cb= set of classes under branch b; Compute L(Cb, Cn) (cf. §6.2); Hp={(x(i), y(i))}nH i=1 ∼ L(Cb, Cn); HΨ={ψ(i)}nH i=1 ∈Ψ; for i= 1 to nHdo (x(i), y(i))∈Hp; ψi∈HΨ; Let T+ b∈Cb; Let T− b∈Cn−Cb; for all images under n do Extract HOG at (x(i), y(i)) using ψ(i); Let h+ i= HOG from T+ b; Let h− i= HOG from T− b; Train classifier gion 2 3of: h+ iand h− i; Test gion OOB set 1 3of: h+ iand h− i; Rank block (x(i), y(i), ψ(i), gi) Select nhbest blocks, Bb={(xj, yj, ψj, gj)}nh j=1 ; Update the list B0=B ∪Bb; is created by sampling from Cn−Cb. An out-of-bag (OOB) testing set is created by removing 1/3 of the instances from the positive and negative sets. A discriminative binary classifier gi (e.g. SVM) is trained using these examples. We then test this weak classifier on the OOB test instances to select a threshold τiand rank the block according to the actual True Positive (TP) and False Positive (FP) rates achieved: the overall performance of the weak classifier giis determined by selecting the point on its ROC curve lying closest to the upper-left hand corner that represents the best possible performance (i.e. 100% TP and 0% FP). We then rank the rejector using the Euclidean distance to that point. Each weak classifier gi(hi) thus consists of a function fi, a threshold τiand a parity term piindicating the direction of the inequality sign: gi(hi) = 1 if pifi(hi)< piτi, 0 otherwise. (6.11) Here hiis the HOG extracted at location (x(i), y(i)) using parameters ψ(i). In the example proposed in Fig. 6.6b, fi(hi) is the projection of the block hion the hyperplane found by SVM. The corresponding ROC curve is shown in Fig. 6.6c. For each branch, the nhbest blocks are kept in the list Bwhich is a bag/pool of HOG blocks and associated weak classifiers. If nB is the total number of branches in the tree over all levels, the final list Bhas nB×nhblock elements. The selection of discriminative HOG blocks is depicted in Alg. 7. By this process, features are extracted from areas of high edge probability across our training
6.3 Randomized Cascades of Rejectors 105 set more than areas with low probability. By using this information to sample features, the proportion of useful information available for random selection is increased. Figure 6.7: Rejector branch decision: shown here is a diagram of a rejector classifier. Note that the structure allows for any number of branches at each node and is not fixed at 2. The tree-like structure of the classifier is defined by Sand the blocks (and associated weak classifiers) stored during training for each branch for this structure are held in B. The random vector Φkdefines a cascade rkby selecting one of the rejector blocks in Bbat each branch b, being Bb∈ B the bag of HOG blocks for branch b. 6.3.3 Randomization Random Forests, as described in [Breiman,2001] are constructed as follows. Given a set of training examples T={(yi,xi)}, where yiare class labels and xithe corresponding Ddimensional feature vectors, a set of random trees Fis created such that for the k-th tree in the forest, a random vector Φkis used to grow the tree resulting in a classifier tk(x,Φk). Each element in Φkis a randomly selected feature index for a node. The resulting forest classifier Fis then used to classify a given feature vector xby taking the mode of all the classifications made by the tree classifiers t∈Fin the forest. Each vector Φkis generated independently of the past vectors Φ1..Φk−1, but with the same distribution: φ∼ U(1, D),∀φ∈Φk,(6.12) where U(1, D) is a discrete uniform distribution on the index space ID={1,··· , D}. The dimensionality of Φkdepends on its use in the tree construction (i.e. the number of branches in the tree). For each tree, a decision function f(·) splits the training data that reaches a node at a given level in the tree by selecting the best feature m* from a subset of m << D randomly sampled features (typically m=√Ddimensions). An advantage of this method over other tree based methods (e.g. single decision trees) is that since each tree is trained on a randomly sampled 2 3of the training examples [Breiman,1996] and that only a small random subset of the available dimensions are used to split the data, each tree makes a decision using a different view of the data. Each tree in the forest Flearns quite different decision boundaries, but when averaged together the boundaries end up reasonably fitting the training data. In Random Forests (RF), the use of randomization over feature dimensions makes the forest less prone to over-fitting, more robust to noisy data and better at handling outliers than single decision trees [Breiman,2001]. We exploit this random selection of features in our algorithm so that our classifier will also be less susceptible to over-fitting and more robust to noise compared to a single hierarchical decision tree that would use all the selected features together: a hierarchical tree-structured classifier rk(IN,Φk,S,B), with INbeing a normalized input image, is thus built by randomly sampling one of the HOG blocks in the list Bb∈ B at
106 Chapter 6. Multi-class Pose Classifier each branch bof the hierarchical structure S. This gives a random nB-dimensional vector Φk where each element corresponds to a branch in the structure S: Φk∈ InBwhere φ∼ U(1, nh),∀φ∈Φk. (6.13) Here U(1, nh) is a discrete uniform distribution on the index space I={1,··· , nh}. The value of each element in Φkis the index of the randomly selected HOG block from Bbfor its corresponding branch. An ensemble Rof nchierarchical tree-structured classifiers is grown by repeating this process nctimes: R(IN) = {rk(IN,Φk)}nc k=1,(6.14) where Sand Bare left out for brevity. Rthus has 2 design parameters ncand nhwhose main effects on performance will be evaluated in Sect. 6.4.2. Let Pωbe the path through the tree-structure Sfrom the top node to the class leaf ω.Pωis in fact an ordered sequence of nbω branches5:Pω=bω 1, bω 2,··· , bω nbω. For each classifier rk∈Rand for each class ω∈Ω, where Ω = {1,··· , nω}, we have the corresponding ordered sequence of HOG blocks and associated weak classifiers: Hω k={(xj, yj, ψj, gj)}nbω j=1,(6.15) with (xj, yj, ψj, gj) = Bb[Φk(b)], (6.16) where the branch index bis the jth element in Pω(i.e b=Pω[j]) and Bb∈ B. See Fig. 6.7a. For each tree-structured classifier rk, the decision to explore any branch in the hierarchy is based on a accept/reject decision of a simple binary classifier gjthat works in a similar way to a cascade decision (see Fig. 6.7b). The ensemble classifier Ris then, in fact, a series of Randomized Hierarchical Cascades of Rejectors. The decision at each node of a hierarchical cascade classifier rkis made in a one-vs-all manner, between the branch in question and its sibling branches. In this way multiple paths can be explored in the cascade that can potentially vote for more than one class. This is a useful attribute for classifying potentially ambiguous classes and allows the randomized cascades classifier to produce a distribution over multiple likely poses. Each cascade classifier rktherefore returns a vector of binary outputs (yes/no) ok=o1 k, o2 k,··· , onω kwhere a given oω kfrom ok for class index ω, takes a binary value: oω k=1 if ∀(xj, yj, ψj, gj)∈ Hω k,gj(hj)=1, 0 otherwise. (6.17) Here hjis the HOG extracted at location (xj, yj) using parameters ψj. Each output oω kis initialized to zero, which means that if no leaf is reached during classification, rkwill return a vector of zeros and will not contribute to the final classification. The uncertainty associated with each rejector during learning could be considered to provide a crisp output (i.e. each cascade voting for only one class) or to compute a soft confidence values of the cascade classifiers. Although other classifier combination techniques could be considered, here we choose to use a simple sum rule to combine the binary votes from the different cascade classifiers in the ensemble Rthat outputs the vector O: O=O1, O2,··· , Onω=R(I),(6.18) where each Oω∈[0, nc] represents a number of votes: Oω= nc X k=1 oω k,(6.19) 5Note that nbωcould be different for each class ω∈Ω
6.5 Experiments using MoBo Dataset 113 Figure 6.14: Classes (64) obtained with the class definition method presented in Fig.6.13. The 8 rows correspond to the 8 training views available in the dataset. For each class, we show the average gradient map and average pose. for training and 18,000 for testing. Benchmark experiments were performed on this dataset to get initial baseline results with three state-of-the-art multi-class (pose) classifiers: •Random Forests [Breiman,2001], inherently good for multi-class problems, that share similarities with our approach. •PSH [Shakhnarovich et al.,2003] which is a fast and effective way of finding the neighboring poses of a query image. We will take the class of the nearest pose as a classification. •A multi-SVMs classifier, similar in spirit to that of [Okada and Soatto,2008], which learns kone-vs-all linear SVMs to discriminate between kpredefined pose clusters10. 10As we do not perform pose-dependent feature selection and use linear SVMs instead of the ARD-Gaussian
114 Chapter 6. Multi-class Pose Classifier During classification the class is determined by choosing the cluster that has the highest probability p(Ck|x) from the SVMs. We will refer to this work as multi-SVMs or SVMs in the rest of the chapter. (a) (b) (c) Figure 6.15: Pose classification for PSH, Random Forest (1000 trees) and SVMs classifiers trained on a subset of the MoBo database containing 10 subjects and tested on the remaining 5 subjects. We compare the results using segmented images without background and images with a random background for 4 different grids of HOG descriptors. (a) (b) (c) Figure 6.16: Pose classification rates for 4 different ensembles (1, 10, 100 and 1000 cascades) trained on the subset of the MoBo database containing 10 subjects and tested on the remaining 5 subjects. We compare the results using segmented images without background and images with a random background for different number nhof sampled HOG blocks. We tuned the parameters of PSH and RF to get optimal performances on our 64 class MoBo dataset: we tailored PSH parameters to the number of images and used 200 18-bit hash functions and empirically validated that the optimum number mof randomly selected features for RF was near √D(as indicated in [Breiman,2001]). We trained 64 one-vs-all linear SVMs. Two groups of classifiers are trained using segmented images without background and images with a random background respectively. We run the same test for 3 different grids of HOG descriptors. Table 6.1 gives the corresponding training time and Fig. 6.15 shows the kernel SVMs, it is expected to perform a little lower than that of [Okada and Soatto,2008] but the training will be much faster. Since we use linear SVMs to train our cascade rejectors, we find it is a reasonable comparison.
6.5 Experiments using MoBo Dataset 115 performances11 when classifying images with or without background. RF results are given for a 1000-tree forest since convergence is reached for that number as observed in Fig. 6.1. Table 6.1: Classifiers training time on an Intel Core Quad Processor at 2.00GHz with 8Gb of RAM (see experiments in Fig. 6.15). HOG Grid 4x4 5x5 7x12 3 together Dimension 1152 1800 2688 5640 Backgrd RF 5h30 6h45 13h00 21h45 SVMs 2h00 4h00 8h00 17h30 PSH 1h50 2h30 3h00 5h30 No Backgrd RF 4h40 5h30 11h00 18h30 SVMs 45min 52min 1h00 2h20 PSH 1h20 2h20 3h08 5h30 The first observation we made is that all of the classifiers trained on segmented images perform poorly when classifying images with background (see Fig. 6.15b). This demonstrates the importance of considering a training dataset that includes a wider background appearance. The second observation is that using denser HOG feature grids improves pose classification accuracy but we are quickly facing memory issues that prevent us from working with denser grids. The third observation we can make is that PSH does not perform well in presence of cluttered background (Fig. 6.15c) while performing decently on segmented images (Fig. 6.15a). PSH tries to take a decision based on 1-bin splits. That could well be a reason as the histogram will be altered by the presence of background and affect that bin. In presence of cluttered background (Fig. 6.15c), SVM is the best classifier in terms of accuracy while Random Forest achieves the highest classification rate when working with segmented images (Fig. 6.15a). After constructing the tree structure S(112 branches using nθ= 1), two lists of HOG block rejectors Bare built using the two same training subsets of the MoBo database (with and without random background). The training took about 2 hours (testing 2500 HOG blocks at each branch) which is much faster than both SVMs and RF classifier training. Then, different ensembles are created varying the number ncof cascades and the number nhof HOG blocks that are considered for sampling. Corresponding classification rates on MoBo dataset are given in Fig. 6.16. We can observe that the accuracy increases with both ncand nhfor the 3 different tests and the cascades outperform the other classifiers when trained on images without background (Fig. 6.16 a and b). In particular, the cascades show better generalization performances when the testing data is significantly different from the training data (Fig. 6.16b). Performances seems to be lower than SVMs and RF when the classifiers are trained and tested on images with random background (see Fig. 6.16c vs Fig. 6.15c). Detailed results are reported in Fig. 6.17 for that concrete case. Fig. 6.17a shows that convergence is reached sooner when nhis low and the accuracy improves when increasing nhuntil nh= 400 for nc= 1000 (See Fig. 6.17b). The figure 6.17c compares the performances of a cascades classifier grown using the 400 best HOG blocks with the best RF classifier from Fig 6.1. The classification rate of the cascades classifier converges earlier (300 cascades) because the first cascades are already very informative compared to the first trees of the Random Forest: 30% of accuracy for the first cascade and 10% for the first tree of the RF. Actually, the accuracy of RF does not seem to converge to 11If neighboring classes are not considered as misclassification, some issues with images on the “boundaries between classes” are solved and the classification rates increase by about 15 −20%
116 Chapter 6. Multi-class Pose Classifier (a) (b) (c) Figure 6.17: Pose classification experiments on MoBo using randomized cascades: (a) we run the same experiment with our randomized cascades classifier (trained on the subset containing 10 subjects and tested the remaining 5 subjects from MoBo dataset) and compare the results obtained sampling from different numbers nhof selected HOG blocks at each branch of the tree. (b) We show the results varying nhwith nc = 1000. (c) We compare the performance of the best Random Forest vs our randomized cascades classifier (with nh= 400). a specific value as every new tree is informative. Our approach performs better than RF for nc≤1000 and is slightly less efficient than SVMs. 6.6 Conclusions In this chapter, we have proposed a new multi-class classifier that combines the best components of state-of-the-art classifiers including hierarchical trees, cascades of rejectors and randomized forests. Other algorithms require that the full feature space be available from an image during training. This can lead to very high dimensional feature vectors being extracted from an image
6.6 Conclusions 117 for large configurations of features and potentially leads to memory problems for training sets with a large number of examples. Our algorithm selects the features it considers to be the most informative during training, and can build a smaller more useful feature space by sampling a small set of features from a much larger configuration of feature descriptors. This approach is computationally efficient as cascades learning takes around 2 hours while SVM and RF approaches took respectively 17h30 and 21h45 with the same training set. Cascade approaches are efficient at quickly rejecting negative examples, and we have exploited this property by learning fast multi-class hierarchical cascades. By randomly sampling the feature, each cascade uses different sets of features to vote, it adds some robustness to noise that helps to prevent over fitting. Moreover, each cascade can vote for one or more class so the ensemble of random cascade classifiers outputs a distribution over possible poses that can be useful when combined with tracking algorithms for resolving ambiguities in pose. In this chapter, we have validated our approach for human pose classification with a numerical evaluation. In the next chapter, we will evaluate the performance of our cascade classifiers in a detection framework.
118 Chapter 6. Multi-class Pose Classifier
7 Human Localization and Pose Estimation 7.1 Introduction In this chapter, we continue the study of the proposed multi-class pose classifier and explore its use for joint human localization and pose estimation in a detection framework. Each hierarchical cascade in the ensemble can make a decision and efficiently reject negative candidates by only sampling a few features of the available feature space. This makes our classifier more suitable for sliding window detectors than Random Forests (RF) or Support Vector Machine (SVM) classifiers. In RF, all the trees have to be completely traversed to produce a vote while SVM classifiers require an entire feature vector to be extracted for each tested window. We thus expect a faster detection with our ensemble of cascades. Additionally, we will also analyze two of our classifier properties that allow a considerable speed-up of the pose detection. We have carried out an exhaustive experimentation to validate our approach with a numerical evaluation (using different publicly available training and testing datasets) and present a comparison with state-of-the-art methods for 2 different levels of analysis: human detection and human pose estimation (with body joints localization) performances with both fixed and moving cameras. This chapter is organized as follows. Section 7.2 shows how our pose classifier can be used for joint detection and pose estimation. In Section 7.3, we discuss the two particular properties of our algorithm while a numerical evaluation of our pose detector is presented in Section 7.4 for both localization and pose estimation. 7.2 Human Pose Detection Since we learn a classifier that is able to discriminate between very similar classes, we can also tackle localization. Given an image I, a sliding-window mechanism then localizes the individual within that image and, at each visited location (x, y) and scale s, a window Ipcan be extracted and classified by our multi-class pose classifier Robtaining Op=Oω pnω ω=1 where p= (x, y, s), is a location vector that defines the classified window. A multi-scale saliency map Mof the classifier response can be generated by taking the maximum value of the distribution for each classified window Ip: M(x, y, s) = max ω∈Ω(Op) = max ω∈Ω(Oω p)nω ω=1.(7.1)
120 Chapter 7. Human Localization and Pose Estimation A dense scan (trained on pose only) is shown in Fig. 7.1a where many isolated false positives appear where the classifier responds incorrectly. Zhang et al. [2007a] tackle joint object detection and pose estimation by using a graphstructured network that alternates the two tasks of foreground/background discrimination and pose estimation for rejecting negatives as quickly as possible. Instead of combining binary foreground/background and multi-class pose classifiers, we propose to perform the two tasks simultaneously by including hard background examples in the negative set T− b(see Fig. 6.6 and Sect. 6.3.2) during the training of our rejectors. To create the hard negatives dataset, the classifiers trained on pose images only can be run on negative examples (e.g. from the INRIA dataset [Dalal and Triggs,2005]) and strong positive classifications are incorporated as hard negative examples (see details in Sect. 7.4.1). Repeating the process of hard negative retraining several times helps to refine the classifiers as can be appreciated in Fig. 7.1 b and c. (a) (b) (c) (d) (e) (f) (g) (h) Figure 7.1: Cascade classifier response for a dense scan. Left side, from top to down: saliency map Mafter 0 (a), 1 (b) and 3 (c) passes of hard negatives retraining (i.e. adding hard negative examples in the cascade learning process). For visualization purpose we represent Mfor ground truth scale (i.e. s= 0.6) and show a zoom around the ground truth location. (d) input image with the pose corresponding to the peak in (c). Right side, from top to down: saliency map Mobtained with a 1-cascade (e), 1 cascade randomized on-line (f) and a 100-cascade ensemble randomized on-line (g), all after 3 passes of hard negatives retraining. In (h) we show the pose corresponding to the peak in (g). By generating dense scan saliency maps for many images, we have found that humans tend to have large “cores” of high confidence value as in Fig. 7.1c. This means that a coarse localization can be obtained with a sparse scan and a local search (e.g. gradient ascend method)
7.3 Properties of the Random Cascades Classifiers for Pose Detection 121 can then be used to find the individual accurately. In the example proposed in Fig. 7.1, taking the maximum classification value over the image, (after exploring all the possible positions and scales) results in reasonably good localization of a walking pedestrian. In that case, we have: p∗= (x∗, y∗, s∗) = argmax (x,y,s) (M(x, y, s)),(7.2) and the corresponding distribution over poses O∗=Op∗=Oω p∗nω ω=1. Once a human has been detected and classified, 3D joints for this image can be estimated by weighting the mean poses of the classes resulting from the distribution using the distribution values as weights, or regressors can be learnt as in [Okada and Soatto,2008]. The normalized 2D pose is computed in the same way and re-transformed from the normalized bounding box to the input image coordinate system obtaining the 2D joints location (see Fig. 7.1d). 7.3 Properties of the Random Cascades Classifiers for Pose Detection In addition to the advantages stated so far, our classifier exhibits some additional interesting properties for pose detection. Since the tree structure Sand HOG block rejectors list Bremain fixed after training, a new cascade classifier erk(IN,e Φk) can be constructed on-line by simply creating a new random vector e Φk. This means that for each classified window Ipin an image, a different ensemble Rpcan be regenerated instantly at no extra-cost: Rp(IN) = {erk(IN,f Φk)}nc k=1,(7.3) where p= (x, y, s),is the location vector that defines Ipand f Φkis sampled using Eq. 6.13. The qualitative effects of this on-line randomization can be appreciated in Fig. 7.1: a dense scan (1-pixel stride) with a 1-cascade classifier (Fig. 7.1e) produces responses which are grouped while randomizing Φ on-line (using a different random cascade at each location) produces a cloud of responses around the ground truth (see Fig. 7.1f). This property will favor an efficient localization at a higher search stride, as verified later by the numerical evaluation presented in Sect. 6.4.2 1. When randomizing a 100-cascade classifier (Fig. 7.1g), the saliency map becomes similar to the one obtained with a regular 100-cascade ensemble without on-line randomization (Fig. 7.1d). Performing construction on-line, each new random vector f Φkcontributes to the final distribution. The probability that any given randomly generated cascade classifier erkwill vote close to an object location increases as the position gets closer to the true location of the object. Even if the classification by the initial cascade for the pose class is wrong, it is generally close to an area where the object is. Subsequent classifications from other randomized rejectors push the distribution toward a stable result. As more classifications are made, the distribution converges and becomes stable as shown in Fig 7.2a and Fig 7.2b where we can observe the higher variability in classification with few cascades compared to the one obtained using bigger ensembles. This explains the similarity of the saliency map obtained with and without on-line randomization (Fig. 7.1) when considering a large number of cascades in the ensemble R. This convergence property can also be exploited for fast localization using cascade thresholding: a cascade rkis drawn at random (without replacement) from R until the 1This property may also be exploited to spread the computation of image classification and localization over time. They may also be readily parallelized due to the structure Sand Bbeing fixed once the classifier has been trained.
122 Chapter 7. Human Localization and Pose Estimation (a) (b) Figure 7.2: On-line Randomization: (a) Detection score (in percentage of votes) max(O) = max(100 ncPnc k=1 oω k),i.e. the peak of the distribution, obtained when classifying the same location (x, y, s) with 100 different randomized ensembles. When increasing nc, the number of cascades in the ensemble R, all the curves converge to 20%. In (b), we present the average and std across all these curves for 2 different locations in the image: ground truth location and 5 pixels away from ground truth. aggregated score reaches a sufficient confidence level. The confidence level can be selected based on a desired trade-off between speed and accuracy. Another possibility is to consider an adaptive cascade filtering by classifying with the entire ensemble R(Rpif on-line randomization is considered) only the locations p= (x, y, s) that yield a detection using a single or few cascades from the ensemble, i.e. filtering with the first nfcascades of the ensemble: M(x, y, s) = maxω∈ΩOω p∈Opif Snf>0, 0 otherwise, (7.4) where the detection score Snfis basically: Snf= max ω∈Ω nf X k=1 oω k!.(7.5) In other words, if a strong vote for background has been made with the first nfcascades (i.e. no vote for any human pose classes and Snf= 0), then the classifier stops classifying with the rest of the cascades in the ensemble. Localizing using this approximate approach means that a classifier with a few cascades can be used as a region of interest detector for a more dense classification. The adaptive cascade filtering combined with on-line randomization will produce an efficient and fast detection, as verified later in Sect. 7.4.1.1.