Video sequence calibration and applications
Abstract
842
Full text
VIDEO SEQUENCE CALIBRATION AND APPLICATIONS L. ÁLVAREZ, C. CUENCA. ASALGADO AND J.SANCHEZ Departamento de Infomxitica y Sistenias Universidad de Las Palnlas de Gran Canaria Campus Universitario de Tafira 3501 7 Las Palmas. SPAIN email: {lalvarez. ccuenca. jsanchez} @dis.ulpgc.es www: l~ttp://serdis.dis.~~lpgc.es/ariii ABSTRACT In this paper we propose a new technique to calibrate a video sequence. We will conlpute the location and rotation of the caniera for each franie of the video sequence. Classical n~ethods for canlera calibration work properly in the case of a small nuinber of caineras. When we take an iinage sequence with a video rate. the number of fraines is veiy large and classical methods do not work properly. To avoid this drawback, we focus the calibration process in the estiniation of the coordinates of a set of 30 points in the scene. so the unknown are these 3D point coordinates rather than the projection matrix associated to each fraine. The advantage of this approach is that with a few number of 3D point coordinates we can estiniate the prqjection matrix of any frame using the relation between the 3D points and its prqjection on each fran~e. In order to compute such set of 3D points. we will use a classical cainera calibration technique applied on a small subsequence of fraines taken froni the large original video sequence. We also perforni an iterative procedure to include the infomation of al1 the canleras in the calibration coniputation. To illustrate the capabilities of the proposed inethod, we will apply this technique to include virtual 3D objects in a real video sequence. KEYWORDS Augmented Reality, Geometric Algorithms, Tracking, Motion Control. 1. Introduction The inatheinatical aspects of multiple cainera calibration have been deeply studied in the literature [3],[4]. In the case of a few nunlber of carneras the classical calibration techniques work properly. However to calibrate a video sequence we have to deal with two important probleins: On one hand we have to calibrate a large nuinber of cameras (one canlera per video fran~e) and on the other hand the displacenient between consecutive franies is in general veiy small and the recovering of the camera parameters based o11 the tracking of singular points across the sequence is veiy unstable with a lot of local mininia configuration faiaway from the physical relevant solution. There are sophisticated techniques to deal with these probleins; however. probably due to the cominercial interest of such techniques. there is not nluch public inforn~ation about specific tools that deal with these probleins. In this paper we present a n~ethod to calibrate a video sequence properly. We do not claini that this method is better than the sophisticated tools presented in some cominercial software packages (In particular we assunle that the intrinsic parameters of the caniera are known, which is a restriction that could be avoided). However, the technique we propose in this paper seenls to work properly, as it is shown in the experimental results. and it could be easily impleinented for any doinain researcher. The reinainder of this paper is organized as follows: In section 2 we present a general overview of niultiple caniera calibration techniques. In section 3 we present the method we propose in this paper. In section 4. we present an application of this technique: the inclusion of virtual objets in a real video sequence. Finally, in section S we present the main conclusions of this paper. 2. Multiple Camera Calibration. A General Overview The problein of multiple camera calibration consists in recovering the camera positions and orientations with respect to the world coordinate system, using as input data tokens. such as pixels or lines. in correspondence in different iniages. Figure 1 shows this scenario for a systenl with three cameras. The specification of the i-th camera position is the 3D (u,orld) point Ci . where the superscript is the reference systeni in which the niagnitude is expressed. The orientation (u,orlrZ) specification is a rotation inatrix Ri or any equivalent representation, such as quateinions or Euler angles. When the image tokens in correspondence are projections of a set of 3D points {_21j)j =,,,-, -, where is the nuniber of points, it is possible to reconstruct each 3D point
Figure 1. Motion parameters derived from point inatches. expressed in the world coordinate systeiii by sinlply estimating the intersection point of the line set: where C: are the coordiilates of the optical center in the world reference system. and rn,$ are tlie coordiiiates of the projection of AIIj in the nomalized reference systeiii for the i-th cainera. A reference system is noimalized when the optical center is in the origin. the foca1 distance is 1 and the pixel is a square of size 1. We will assuine that the intrinsic paraiiieters of the canleras are known. which allows us to iloriiialize the reference systeiii. In order to estinlate the intersection 3D point of the line set it is necessary to know the position of the optical center and the rotation inatrix for each cainera. The conlputation of these parameters solves the probleiii of the iiiultiple caiiiera calibration. After estinlating these parameters, we can evaluate the solution accuracy by pro.jectiilg the reconstructed 3D points in each camera. and the best solution for the calibration problem is the one that miniinizes the energy function: (u o~ld), Ciuorld) f (CO (uorld) Riuorld) . .... Ro . . .) is the projection of the reconstructed There is no closed-foriii solution for the n~iniiiiization of the above energy fuction, and nonlinear ininimization methods must be used. A restriction to take into account in the application of these methods is that the solution must be not only ininiinum but also valid. (A solution is valid when firstly the rotation inatrixes are orthogonal, and secondly. the reconstructed 3D points are always beyoild the iiiiage plane. since the refereilce systenl in the caiiiera is noriiialized.) It is important to find a good initial approximation. close to the final solution. in order to supply as seed input to the nonlinear iiiinimization method which guarantees a fast convergence. This initial solution can be obtained by using linear iiiethods. In [2], the essential ~izatrix E (for n~otion parameters (t. R)) is defined, by: where T is the antisyinmetric matrix: Matrix T is such that T:r: = t A .r for al1 vectors x. The nine elements Eij of E are called esserltial pcrra~izeten. Since two sets of points {m j} and {mi}: i = 1: . ..: X can be interpreted as resulting froin a 3D camera inotion (t. R) if aild only if t. mj. Rm; = 0. then the last equation can be written as shown in [S]:
Equation (3) is linear aild honlogeneous with respect to the coefficient of E. Therefore, if eight of such equations are available. we can solve the system for those coefficieilts. The approach coilsists in first. estinlating the esseiltial nlatrix E. and then recoveriilg t and R froiii E. It is possible to rewrite equation (3) as: where X 1s the 9x1 vector [e;. e;. e:lT ( el is the i-th column vector of the essential matrix E), and n is the 9x1 vector [~nr'~~~n'~znr'~]~. If we have 7r points in correspondence, each one yields an equation like (4) and we can combine thein as follows: where A,, is an nx9 n~atrix: In the preseilce of iloise. equation (5) is oidy approxinlately satisfied, aild we can reforiiiulate the probleiii as that of fiilding the vector X that miniiiiizes the iloriii of --InX with the constraint that the nonn of X is d. It is well known that the solution to the last problem is the eigenvector of nornl 4 of the 9x9 iiiatrix z4*;z4r, corresponding to the snlallest eigenvalue. The conlputation of t aild R can also be perforiiied taking noise into account. The translation vector t is the solution of the following meansquare problem: with the constraint that the norm of It I2 = 1 In order to find the rotation inatrix R. we have to solve the following iiieansquare probleiii: subject to R~R = I aild RI = 1 With this inethod it is possible to calibrate a system with two caineras and without noise. When noise is present, the inethod inust be slightly changed because it is not possible to find a valid solution to the calibration probleiii. Theses chailges coilsist sinlply in iiltroduciilg heuristical rules to select the best solutioil. So extend the method to inore than two caineras it is enough to carry out the calibration for each couple of cameras (the first camera and the second. the second and the third. aild so 011) aild to fit a scale factor for each couple of carneras. In order to calculate the ~cale factor of two pairs of cameras, we reconmuct each 3D point, -UJ. from its respective coi-i-espoildence pairs aild then we nmiiiiize the expression: The analytical solution to this miniinization problem is given by: This method has the advantage of being linear and the disadvantage of being very iloise sensitive (hence the inlportance of a good estiiiiation of the poiilt coordiilates provided by a corner detection techilique). Moreover, the method does not take advantage of having multiple caineras in order to improve the result of the calibration. We can iilclude al1 canleras usiilg the linear solution as initial approxiiiiation for miniiiiization of equation (1). 111 each iteration, we nlust take into accouilt if the solution is still valid. A possible strategy consists in adding a heavy penalty terin to the equation (1) when the restrictions are violated. 3. Video Sequence Calibration The method we propose in this paper to calibrate a video sequence is divided into the following steps : Step 1 In the first step we conlpute autoiiiatically sequences of corresponding singular points across the video sequence. We can use as singular points corners detected using the classical Harris technique or a inore sophisticated one. as the one introduced in [l]. Step 2 In the second step we choose a siiiall subsequeilce of fran~es to recover a set of 3D points. The selected fran~es could be chosen by hand, or taking a fixed step between fraines (i.e. we take for instance frames 1-25-50-75- .....). We could also use a inore sophisticated way to choose the fran~es based on the robustness of the calibration inforiiiation between two cameras. Such robustiless is based on the two smallest eigenvalues Al < X2 of the matrix d;.-I,, (see (5)). The ideal case is that X1 = O and X2 >> O. SO we can choose the fraines by inaximizing some criteria associated to such robustness. for instance X2 - Xl X2 + t Once the video frame subsequence is obtained we conlpute a set of 3D poiilts -21, usiilg the techilique showed in the previous sectioil.
iteration 11 step 1 1 step 16 O 11 25 1 5.6 Table 1. Average reprojection error evolution Step 3 First we notice that from the tracking step we know for each 3D point M, the projection of the point m!: in the i-th camera. From these relations we can compute the projection matrix Pi associated to each camera (see [2] for details). We update the 3D point coordinates and the projection matrix using the following iterative scheme: 0 From M, and m!: we compute the projection matrix Pi for al1 frames in the video sequence. 0 From Pi and m$) we recompute the 3D points M, by intercepting the 3D lines going from m$) to the focus of Pi. o We update Mj with the new computed ones and we start a new iteration until convergence of the iterative scheme. 4. Experimental results. Inclusion of virtual objects in a real video sequence One of the main applications of video sequence calibration is the inclusion of virtual objects in a real video sequence. We will test our method in a real video sequence of 120 frames where we are going to include four artificial objets. In figure 2 we present four frames of the real video sequence (left column) and the same frames with the inclusion of the virtual objets using the calibration parameters obtained with the proposed method (right column). To illustrate the convergence behavior of the proposed iterative scheme we present in figure 3 the evolution across iterations of the average reprojection error in terms of image pixel values. We present the results using two different initial video subsequences. The first one is obtained using frames 1-17-33-49- ... (that is, we fit an step of 16 frames). The second one is obtained using al1 frames (step equal to 1). In table 1 we present the numerical values of the average reprojection error for the iterative scheme evolution presented in figure 3. C hilh Figure 3. Average reprojection error evolution 5. Conclusions In this paper we present a new method to perform video sequence calibration. The method is quite simple and seems to work properly. The experimental results are very prornising, as well as the convergence behavior of the iterative scheme. In a real video sequence we arrive to get an average reprojection error of 0.6 pixels which is a very good estimation. References [l] Alvarez L., Cuenca C., Mazorra L. 2001. Morphological Comer Detection. Application to Camera Calibration, Proceedings of ZASTED Znternational Conference SZGNAL PROCESSZNG, PATTERN RECOGNZTZON AND APPLZCATZONS, Rhodes, Greece , pages 21-26 [2] Faugeras 0. 1993.30 Computer Vision. A Geometric View Point. MIT Press. [3] Faugeras O., Luong Q-T, Papadopoulo T. 2001, The Geometry of Multiple Zmages, MIT Press [4] Hartley R. and Zisserman A. 2000. Multiple View Geometry in Computer Vision. Cambridge University Press. [5] Kanatani K. 1995. Geometric Computation for Machine Vision. Oxford University Press.
Figure 2. On the left, four frames of a real video sequence and on the right the same frames with the inclusion of the virtual objets (a glass, a mug, a teapot and a computer screen)