scieee AI-readable full text Open interactive document viewer

Dense disparity map estimation respecting image discontinuities

Álvarez León, Luis Miguel,Deriche, Rachid,Sánchez, Javier,Weickert, Joachim

Abstract

We present an energy based approach to estimate a dense disparity map from a set of two weakly calibrated stereoscopic images while preserving its discontinuities resulting from image boundaries. We first derive a simplified expression for the disparity that allows us to estimate it from a stereo pair of images using an energy minimization approach. We assume that the epipolar geometry is known, and we include this information in the energy model. Discontinuities are preserved by means of a regularization term based on the Nagel-Enkelmann operator. We investigate the associated Euler-Lagrange equation of the energy functional, and we approach the solution of the underlying partial differential equation (PDE) using a gradient descent method The resulting parabolic problem has a unique solution. In order to reduce the risk to be trapped within some irrelevant local minima during the iterations, we use a focusing strategy based on a linear scalespace. Experimental results on both synthetic and real images arere presented to illustrate the capabilities of this PDE and scale-space based method.

Full text

MVA2000 lAPR Workshop on Machine Vision Applications. Nov. 28-30. 2000, The University of Tokyo, Japan 1 1-1 Dense Disparity Map Estimation Respecting Image Discontinuities: A PDE and Scale-Space Based Approach Luis Alvarez * Rachid Deriche t Depart. de Informatica y Sistemas Comput. Vision Laboratory, Robot,vis Group Universidad de Las Palmas INRIA Sophia-Antipolis Javier Sanchez $ Joachim Weickert 5 Depart. de Informatica y Sistemas Depart. of Mathemat. and Comput. Science Universidad de Las Palmas University of Mannheim Abstract We present an energy based approach to estimate a dense disparity map from a set of two weakly calibrated stereoscopic images while preserving its discontinuities resulting from image boundaries. We first derive a simplified expression for the disparity that allows us to estimate it from a stereo pair of images using an energy minimization approach. We assume that the epipolar geometry is known, and we include this information in the energy model. Discontinuities are preserved by means of a regularization term based on the Nagel-Enkelmann operator. We investigate the associated Euler-Lagrange equation of the energy functional, and we approach the solution of the underlying partial differential equation (PDE) using a gradient descent method. The resulting parabolic problem has a unique solution. In order to reduce the risk to be trapped within some irrelevant local minima during the iterations, we use a focusing strategy based on a linear scalespace. Experimental results on both synthetic and real images are presented to illustrate the capabilities of this PDE and scale-space based method. 1 Introduction Energy based methods have been extensively used in the last years 14, 5, 6, 8, 9, 11, 121 for estimating the disparity map between images. The goal Address: Edificio de lnformiitica y Sistemas, Campus Universitario de Tafira. SP-35017 La.9 Palmas, Spain. E-mail: 1alvarezQdis.ulpgc.es Address: 2004 Route des Lucioles. F06902 Sophia-Antipolis, France. E-mail: Rachid.DericheQsophia.inria.fr Address: Edificio de lnformiitica y Sistemas, Campus Universitario de Tafira. SP-35017 La9 Palmas, Spain. E-mail: jsanchezQdis.ulpgc.es hddress: D-68131 Mannheim, Germany. E-mail: Joachim.WeickertQti.uni-mannheim.de of this paper is to present a variational approach to recover a dense disparity map from a set of two weakly calibrated stereoscopic images. To solve this problem, we first make full use of the knowledge of the so-called fundamental matrix [7] to derive the equations that relate corresponding pixels in the two views. The intrinsic and extrinsic parameters of the camera are not known. We directly compute the disparity map from the grey-level image intensities without dealing with any intermediate process such as rectification and we address the problem of accurately determining the dense disparity map while smoothing and regularizing this disparity map along the contours of the grey level image and inhibiting smoothing across the image discontinuities. The preservation of discontinuities in the disparity map is obtained using an anisotropic linear operator 13, 101 which allows to develop discontinuities in the disparity map across the edges of one of the 2 images. This important step is achieved by considering a well adapted regularization term that has already proven to be very useful in optical flow estimation. In order to avoid converging to irrelevant minima, a focusing strategy embedding our method in a linear scale-space is used [3], as it has already been successfully applied in optical flow estimation. The coarse-scale solution serves then as initial data for solving the problem at a finer scale. We have shown that our method leads to a mathematically correct concept and we have proven the existence and uniqueness of the solution of the parabolic equation which governs our method. Finally, our approach has been validated on a large set of synthetic and real stereo data. All these results are presented in a previous work 11). 2 Formalism of the matching process 2.1 Notation and Background In this paper we use a projective camera model. This model maps a 3D point M = [X, Y, ZIt to a 2D image point m = [x, ylt through a 3 x 4 projection matrix P via sm = PM, where s is a nonzero scale factor and the notation p is such that if p = [x, y,. . .It then p = [x, y,. . . ,lit. In the case of two images acquired by a binocular stereo system, every physical point M in space yields a pair of 2D projections ml and m2 on the two images. The 3 x 4 projection matrices are defined by the following relations: Assuming that the world coordinate system is associated with the first camera, the two projection matrices are given by where R and t represent the 3 x 3 rotation matrix and the 3 x 1 translation vector defining the rigid displacement between the two cameras, and 0 denotes the 3 x 1 null vector. The matrices A and A' are the 3 x 3 intrinsic parameters matrices of the two views, each depending on five parameters and having the following well-known form [7]: a, -a, cot8 U,O A= [ 0 a,,/ sin 8 vo 0 0 1 I All these matrices and parameters can be computed with good accuracy by means of a classical calibration method 171. In such a case, thc system is said to be calibrated. By eliminating the scalars $1 and .sz associated with the projection equations (1) as well as the point M, an equation relating the pair of projections of the same 3D point is obtained: For a point ml = [x, y]' in the first image Il, the fundamental matrix F = (f,,j) provides the epipolar line A of equation 1liZtFnil = 0 in the image I,. Let us introduce the notation We will use this equation in order to introduce our specific parameterization of the disparity, developed to yield a simple linear second order differential operator in the minimization part. 2.2 The Disparity Term Under the Lambertian assumption that corresponding pixels have equal grey values, the determination of the disparity from the stereo pair comes down to finding a function h(x, y) := (u.(x, y), v(x, y))' such that Il(z, y) = I,(x + u(x, y), y + V(Z, y)), where the point (x', y') = (x + ~(x, y), y + ~(x, y)) belongs to the epipolar line associated to (x, y). Let us denote by N (resp. T) the unitary normal (resp. tangential) vector of the epipolar line A given by the equation mztFnil = 0, and by D the unitary disparity vector associated to the point ml and m2 Then we have m2 = ml + 6D = ml - yN - AT (9) where 6 = represents the disparity. y represents the distance (modulus a sign) of the point ml to its epipolar line A in the second image and X represents the dist,ance (modulus a sign) of mo, the projection of the point ml on the epipolar line A, to the point m2 that lies along the epipolar line A. 2.3 The Energy Functional Let us now develop an approach to accurately estimate the A(x, y) function associated to a pair of stereoscopic images. The easiest possibility would be to proceed in a classical way and try to recover this important information using a simple correlation scheme. Unfortunately, this naive solution will not provide a correct and accurate result, in particular in the regions where the disparity map has discontinuities, as is often the case at image edges. It is well known that the disparity map of this classical method tends to be very smooth across the image boundaries. The idea we would like to formalize and develop here is to estimate a A(x, y) function which is smooth only along the image boundaries and not across them. This leads us to consider the minimization of the following energy functional: Using this notation, the epipolar line A can be writE(A) = In (I,(x, y) - I,(x + h(A(x, yen)))2 dz dy ten as a(x, y)xl + b(x, y)yl + c(x, y) = 0. (7) + C~VA~D(VI,)VA~~~~ (10) 424 Here, R denotes the image domain, C is a positive algorithm converges to a local minimum of the enconstant, and D(VIl) is a regularized projection maergy functional (10) that is located in the vicinity of trix perpendicular to VIl: the initial data. To avoid convergence to irrelevant where Id denotes the identity matrix. This projection has been introduced by Nagel and Enkelmann in the context of optic flow estimation. We use it here because of its simplicity (the underlying second order differential operator is linear) and because this method has demonstrated its performance numerous times in the context of optical flow estimations 131. 2.4 Minimizing the Energy In order to minimize our energy functional, we solve its associated Euler-Lagrange equation C div (D (VIl) VX) We obtain a solution of the above equation by calculating thr asymptotic state (t + co) of the parabolic equation local minima, we embed our method into a linear scale-space framework [13]. Considering the problem at a coarse scale avoids that the algorithm gets trapped in physically irrelevant local minima. The basic idea of embedding our method in linear scale-space is as follows : we replace the images I1 and I, by IF := G, * I( and I," := G, * I,, where * is the convolution operakor, and G, denotes a Gaussian with standard deviation a. We start with a large initial scale ao. Then we compute the disparity A,,, at scale a0 as the asymptotic state of the solution using some initial approximation (see below). Next, we choose a. number of scales u, < a,-, < .... < 00, and for each scale u, we compute the disparity A,, as the asymptotic state of the above PDE with initial data A,,-, The final disparity corresponds to the smallest scale a,. In accordance with the logarithmic sampling strategy in linear scale-space theory, we choose a, := qlaO with some decay rate q E (0,l). A detailed analysis of the usefulness of such a focusing strategy in the context of a related optic flow problem can be found in 131. Our method is governed by the evolution equation dX - = Cdiv (D (vI,) VX) ax, at - - at - Cdiv (D (VIP) VX,) X a () -b(%)l (12) a (%)" + (I1 - 1;) JFTP + (I; - dm (13) We observe that in this diffusion-reaction method the matrix D(VIl) plays the role of a diffusion tenFor the initial value X,""(x, y) we consider two possor. Its eigenvectors are vl := VIl and v2 := TI:, sibilities. The first one is to take a constant value and the corresponding eigenvalues are given by which depends, in general, on a rough a priori estimation of the expected disparity, following an esIn the interior of objects we have IVI,] -+ 0, and therefore X1 + 112 and X2 + 112. At ideal edges where ]VIlI + co, we obtain X1 -+ 0 arid X2 -, 1. Thus, we have isotropic behavior within regions, and at image boundaries the process smoothes anisotropically along thc edge. 2.5 A Linear Scale-Space Approach to Recover Large Disparities In general, the Enler-Lagrange equation (2.4) will have multiple solutions. As a consequence, the asymptotic state of the parabolic equation, which we use for approximating the disparity, will depend on the initial data. Typically, we may expect that the timation of the depth where the interesting objects are located. 3 Experimental Results In Figure 1 we present the calculated disparity (dw) using the classical correlation method (bottom left) and using our method with the correlation technique as initialization (bottom right). Two epipolar lines are depicted in the right image. They correspond to the points represented by a cross around the left eye and the nose in the left image. Figure 2 shows several views of the 3-D reconstruction of the face in the stereo pair, using the disparity map obtained by our method. The reconstruction looks very realistic. References Figure 1: Top: the original stereo pair. Bottom left: the computed disparity map using a correlation window of size 13 x 13. Bottom right: our method (a0 = 7, a, = 0.8, a = 1, s = 0.5, q = 0.95), with the correlation result as initialization. [lj L. Alvarez, R. Deriche, J. SAnchez, and J. Weickert, Dense Disparity Map Estimation Respecting Image Discontinuities: a PDE and ScaleSpace Based Approach, Technical Report INRIA 3874, Sophia-Antipolis, France, 2000. 121 L. Alvarez, P.-L. Lions, and J.-M. Morel, Image selective smoothing and edge detection by nonlinear diffusion. 11, SIAM J. Numer. Anal. 29, 1992, 845-866. [3] L. Alvarez, J. Weickert, and J. SBnchez, Reliable estimation of optical flow for large displacements, Int. J. Comput. Vision, in press. [4] S. T. Barnard, Stochastic stereo matching over scale. Int. J. Comput. Vision 3, 1989, 17-32. [5] R. Deriche, C. Bouvin, and 0. Faugeras, A level-set approach for stereo, in Investigative Image Processing, (L. I. Rudin and S. K. Bramble, Eds.), Boston, MA, Nov. 1996, SPIE Proceedings, Vol. 2942, pp. 150-161, 1997. [6] R. Deriche, S. Bouvin, and 0. Faugeras, Front propagation and level-set approach for geodesic active stereovision, in Third Asian Conference On Computer Vision, Hong Kong, January 1998. 171 0. Faugeras, Three-Dimensional Computer Vision: a Geometric Viewpoint, MIT Press, Carnbridge, MA, 1993. [8] 0. Faugeras and R. Keriven, Variational principles, surface evolution, PDE's, level set methods and the stereo problem, IEEE Trans. on Image Processing 7, 1998, 336-344. [9j F. Heitz, P. Perez, and P. Bouthemy, Multiscale minimization of global energy functions in some visual recovery problems. CVGIP : Image Understanding 59, 1994, 125-134. [lo] H.-H. Nagel, W. Enkelmann, An investigation of smoothness constraints for the estimation of displacement vector fields from images sequences, IEEE Trans. Pattern Anal. Mach. Intell. 8, 1986, 565-593. Ill] L. Robert and R. Deriche, Dense depth map reconstruction: a minimization and regularization approach which preserves discontinuities, in Computer Vision - ECCV '96 (B. Buxton, R. Cipolla, Eds.), Vol. I, pp. 439-451, Lecture Notes in Computer Science, Vol. 1064, Springer, Berlin, 1996. [12] J. Shah, A nonlinear diffusion model for discontinuous disparity and half-occlusions in stereo, in Proc. Int. Conf. on Computer Vision and Pattern Recognition, New York, June 15-17, 1993, Figure 2: Four views of the 3-D reconstruction of pp. 34-40, IEEE Computer Society Press. the stereo pair in Figure 1, using the disparity map 1131 J. Weickert, S. Ishikawa, and A. Imiya, Linear from our method. scale-space has first been proposed in Japan, J. Math. mag. Vision 10, 1999, 237-252.