scieee AI-readable full text Open interactive document viewer

Stereo Anywhere: Robust Zero-Shot Deep Stereo Matching Even Where Either Stereo or Mono Fail

Bartolomei, Luca; Tosi, Fabio; Poggi, Matteo; Mattoccia, Stefano

Abstract

We introduce Stereo Anywhere, a novel stereo-matching framework that combines geometric constraints with robust priors from monocular depth Vision Foundation Models (VFMs). By elegantly coupling these complementary worlds through a dual-branch architecture, we seamlessly integrate stereo matching with learned contextual cues. Following this design, our framework introduces novel cost volume fusion mechanisms that effectively handle critical challenges such as textureless regions, occlusions, and non-Lambertian surfaces. Through our novel optical illusion dataset, MonoTrap, and extensive evaluation across multiple benchmarks, we demonstrate that our synthetic-only trained model achieves state-of-the-art results in zero-shot generalization, significantly outperforming existing solutions while showing remarkable robustness to challenging cases such as mirrors and transparencies.

Full text

Stereo Anywhere: Robust Zero-Shot Deep Stereo Matching Even Where Either Stereo or Mono Fail Luca Bartolomei∗,†Fabio Tosi†Matteo Poggi∗,†Stefano Mattoccia∗,† ∗Advanced Research Center on Electronic System (ARCES) †Department of Computer Science and Engineering (DISI) University of Bologna, Italy https://stereoanywhere.github.io/ RGB Depth Anything v2 [121] RAFT-Stereo [55] Stereo Anywhere (Ours) Middlebury ✓ ✓ ✓ Booster ✓✗✓ MonoTrap ✗✓ ✓ Figure 1. Stereo Anywhere: Combining Monocular and Stereo Strenghts for Robust Depth Estimation. Our model achieves accurate results on standard conditions (on Middlebury [86]), while effectively handling non-Lambertian surfaces where stereo networks fail (on Booster [127]) and perspective illusions that deceive monocular depth foundation models (on MonoTrap, our novel dataset). Abstract We introduce Stereo Anywhere, a novel stereo-matching framework that combines geometric constraints with robust priors from monocular depth Vision Foundation Models (VFMs). By elegantly coupling these complementary worlds through a dual-branch architecture, we seamlessly integrate stereo matching with learned contextual cues. Following this design, our framework introduces novel cost volume fusion mechanisms that effectively handle critical challenges such as textureless regions, occlusions, and nonLambertian surfaces. Through our novel optical illusion dataset, MonoTrap, and extensive evaluation across multiple benchmarks, we demonstrate that our synthetic-only trained model achieves state-of-the-art results in zero-shot generalization, significantly outperforming existing solutions while showing remarkable robustness to challenging cases such as mirrors and transparencies. 1. Introduction Stereo is a fundamental task that computes depth from a synchronized, rectified image pair by finding pixel correspondences to measure their horizontal offset (disparity). Due to its effectiveness and minimal hardware requirements, stereo has become prevalent in numerous applications, from autonomous navigation to augmented reality. Although in principle single-image depth estimation [3] requires an even simpler acquisition setup, its ill-posed nature leads to scale ambiguity and perspective illusion issues that stereo methods inherently overcome through wellestablished geometric multi-view constraints. However, despite significant advances through deep learning [47,72], stereo models still face two main challenges: (i) limited generalization across different scenarios, and (ii) critical conditions that hinder matching or proper depth triangulation. Regarding (i), despite the initial success of synthetic datasets in enabling deep learnThis CVPR paper is the Open Access version, provided by the Computer Vision Foundation. Except for this watermark, it is identical to the accepted version; the final published version of the proceedings is available on IEEE Xplore. 1013 ing for stereo, their limited variety and simplified nature poorly reflect real-world complexity, and the scarcity of real training data further hinders the ability to handle heterogeneous scenarios. As for (ii), large textureless regions common in indoor environments make pixel matching highly ambiguous, while occlusions and non-Lambertian surfaces [76,115,127] violate the fundamental assumptions linking pixel correspondences to 3D geometry. We argue that both challenges are rooted in the underlying limitations of stereo training data. Indeed, while data has scaled up to millions - or even billions - for several computer vision tasks, stereo datasets are still constrained in quantity and variety. This is particularly evident for nonLambertian surfaces, which are severely underrepresented in existing datasets as their material properties prevent reliable depth measurements from active sensors (e.g. LiDAR). In contrast, single-image depth estimation has recently witnessed a significant scale-up in data availability, reaching the order of millions of samples and enabling the emergence of Vision Foundation Models (VFMs) [22,43,120, 121]. Such data abundance has influenced these models in different ways, either through direct training on largescale depth datasets [120,121] or indirectly by leveraging networks pre-trained on billions of images for diverse tasks [22,43]. Since these models rely on contextual cues for depth estimation, they show better capability in handling textureless regions and non-Lambertian materials [75,81,128,129] while being inherently immune to occlusions. Modern graphics engines have further accelerated this progress, enabling rapid generation of high-quality synthetic data with dense depth annotations. However, although synthetic datasets featuring non-Lambertian surfaces like HyperSim [81] have proven effective for monocular depth estimation [75,128,129], this data abundance has not translated to stereo. Despite efforts in generating stereo pairs via novel view synthesis [24,54,104], available data remains insufficient for robust stereo matching. In this paper, rather than focusing on costly real-world data collection or generating additional synthetic datasets, we propose to bridge this gap by leveraging existing VFMs for single-view depth estimation. To this end, we develop a novel dual-branch deep architecture that combines stereo matching principles with monocular depth cues. Specifically, while one branch of the proposed network constructs a cost volume from learned stereo image features, the other branch processes depth predictions from the VFM on both left and right images to build a second cost volume that incorporates depth priors to guide the disparity estimation process. These complementary signals are then iteratively combined [55], along with novel augmentation strategies applied to both cost volumes, to predict the final disparity map. Through this design, our network achieves robust performance on challenging cases like textureless regions, occlusions, and non-Lambertian surfaces, while requiring minimal synthetic stereo data. Importantly, while leveraging monocular cues, our approach preserves stereo matching geometric guarantees, effectively handling scenarios where monocular depth estimation typically fails, such as in the presence of perspective illusions. We validate this through our novel dataset of optical illusions, comprising 26 scenes with ground-truth depth maps. We dub our framework Stereo Anywhere, highlighting its ability to overcome the individual limitations of stereo and monocular approaches, as depicted in Fig. 1. To summarize, our main contributions are: • A novel deep stereo architecture leveraging monocular depth VFMs to achieve strong generalization capabilities and robustness to challenging conditions. • Novel data augmentation strategies designed to enhance the robustness of our model to textureless regions and non-Lambertian surfaces. • A challenging dataset with optical illusion, which is particularly challenging for monocular depth with VFMs. • Extensive experiments showing Stereo Anywhere’s superior generalization and robustness to conditions critical for either stereo or monocular approaches. 2. Related Works We briefly review the literature relevant to our work. Deep Stereo Matching. In the last decade, stereo matching has transitioned from classical hand-crafted algorithms [85] to deep learning solutions, leading to unprecedented accuracy in depth estimation. Early deep learning efforts focused on replacing individual components of the conventional pipeline [88,96,105,130,131]. Since DispNetC [61], end-to-end architectures have evolved into 2D [53,92, 125,125] and 3D [4,8,9,32,44,90,91,119,132,134] approaches, processing cost volumes through correlation layers or 3D convolutions respectively. More recent advances, thoroughly reviewed in [47,72,107], include recurrent architectures for stereo matching [13,27,40,50, 55,110,116,140] inspired by RAFT [99], Transformerbased solutions [31,52,59,97,113,117,138] for capturing long-range dependencies, and fully data-driven MRF models [28]. Among them, some methods specifically address temporal consistency in stereo videos [41,42,133,137]. Domain generalization remains a major challenge, with various approaches proposed including domain-invariant feature learning [17,56,80,93,135], hand-crafted matching costs [7,15], integration of additional geometric cues [2,66,105], and exploitation of sparse depth measurements from active sensors [5,49,69]. In parallel, selfsupervised approaches [25,57] have emerged as effective alternatives to supervised learning, even using pseudolabels from traditional algorithms [1,100] or deploying neural radiance fields [104]. Despite the numerous attempts to 1014 &RUUHODWLRQ 9ROXPH IURPQRUPDOV $JJUHJDWHG &RUUHODWLRQ 9ROXPH IURPQRUPDOV &RUUHODWLRQ 9ROXPH 7UXQFDWH )XQFWLRQ 7UXQFDWHG &RUUHODWLRQ 9ROXPH '+RXUJODVV (VWLPDWHG 1RUPDO0DSV &RQWH[W %DFNERQH )HDWXUHH[WUDFWLRQ %DFNERQH 6WHUHR3DLU 0RQRFXODU 'HSWK(VWLPDWLRQV 0'(V 'LIIHUHQWLDEOH 6FDOHU / / 6FDOHG0'( / / / / )LQDO 'LVSDULW\ &RUUHODWLRQ3\UDPLGV &RUUHODWLRQ3\UDPLGV IURPQRUPDOV        )HDWXUHV ([WUDFWLRQ &RUUHODWLRQ3\UDPLGV %XLOGLQJ ,WHUDWLYH'LVSDULW\ (VWLPDWLRQ Figure 2. Stereo Anywhere Architecture. Given a stereo pair, (1) a pre-trained backbone is used to extract features and then build a correlation volume. Such a volume is then truncated (2) to reject matching costs computed for disparity hypotheses being behind nonLambertian surfaces – glasses and mirrors. On a parallel branch, the two images are processed by a monocular VFM to obtain two depth maps (3): these are used to build a second correlation volume from retrieved normals (4). This volume is then aggregated through a 3D CNN to predict a new disparity map, used to align the original monocular depth to metric scale through a differentiable scaling module (5) for it. In parallel, the monocular depth map from left images is processed by another backbone (6) to extract context features. Finally, the two volumes and the context features from monocular depth guide the iterative disparity prediction (7). improve specific aspects through the aforementioned techniques, recent architectures achieve remarkable generalization by combining their architectural advances with the increasing availability of diverse training data, while online adaptation techniques enable further improvements during deployment through self-supervised learning [45,67, 71,101]. However, although progress on challenges like over-smoothing [103,118] and visually imbalanced stereo [2,11,58,105], handling non-Lambertian surfaces remains particularly challenging due to limited annotated data and complex appearance, with rare works like Depth4ToM [18] specifically addressing this through semantic guidance. Among all the aforementioned approaches, there have been limited attempts to integrate stereo with monocular cues [1,12,112], mostly in self-supervised settings or through loose coupling between modalities. Monocular Depth Estimation. Parallel to developments in stereo matching, single-image depth estimation has evolved from hand-crafted features [82] to deep learning methods [10,21,48,73,108], with self-supervised approaches [25,26,60,68,111,139,141] reframing the task as an image reconstruction problem. This led to multi-task approaches incorporating flow [79,102,124,142] and semantics [29,126], alongside advances in uncertainty estimation [34,70] and dynamic object handling [46,63,98]. Affine-invariant models [20,77,78,109,122] marked a breakthrough in cross-domain generalization, pioneered by MiDaS [78] and followed by works like DPT [77] and, more recently, the Depth Anything series [120]. These approaches used different data sources, from internet photos [51,94,95,122] to car sensors [23,62] and RGB-D devices [16,64], representing the first generation of VFMs for monocular depth estimation. Recent works have focused on metric depth estimation through camera parameter integration [30,35,123], diffusion models [19,22,33,38, 43,83,84], and temporal consistency [36,89]. Moreover, material-aware methods [18], diffusion models [106], and large-scale synthetic datasets have enabled robust monocular depth estimation for non-Lambertian surfaces [121]. Stereo methods, however, still struggle with these surfaces due to limited real-world and synthetic annotated data, affecting generalization. We address this by integrating robust monocular VFMs into a stereo architecture. Concurrent Works. Finally, we mention some solutions for stereo [14,39,114] and for multi-view stereo [37], developed in parallel with ours and sharing similar rationale. 3. Method Overview Given a rectified stereo pair IL,IR∈R3×H×W, we first obtain monocular depth estimates (MDEs) ML,MR∈ R1×H×Wusing a generic VFM ϕMfor monocular depth estimation. We aim to estimate a disparity map D= ϕS(IL,IR,ML,MR), incorporating VFM priors to provide accurate results even under challenging conditions, such as texture-less areas, occlusions, and non-Lambertian surfaces. At the same time, our stereo network ϕSis designed to avoid depth estimation errors that could arise from relying solely on contextual cues, which can be ambiguous, like in the presence of visual illusions. Following recent advances in iterative models [55], Stereo Anywhere comprises three main stages, as shown in Fig. 2: I) Feature Extraction, II) Correlation Pyramids Building, and III) Iterative Disparity Estimation. 1015 3.1. Feature Extraction Two distinct types of features are extracted [55]: image features and context features – (1) and (6) in Fig. 2. The image features are obtained through a feature encoder processing the stereo pair, yielding feature maps FL,FR∈ RD×H 4×W 4, which are used to build a stereo correlation volume at 1 4of the original input resolution. These encoders are initialized with pre-trained weights [55] and the image encoder is kept frozen during training. For context features, we employ a context encoder with identical architecture to the feature encoder, but processing the monocular depth estimate aligned with the reference image ML– (3) in Fig. 2 – instead of ILto capture strong geometry priors. Accordingly, during training the context encoder is optimized to extract meaningful features from these depth maps. 3.2. Correlation Pyramids Building As a standard practice in stereo matching, the cost volume is the data structure encoding the similarity between pixels across two images. Accordingly, our model utilizes cost volumes—specifically Correlation Pyramids [55]—but in a novel manner. Indeed, Stereo Anywhere constructs two correlation pyramids: a stereo correlation volume derived from IL,IRto encode image similarities, and a monocular correlation volume from ML,MRto encode geometric similarities—(2) and (4) in Fig. 2. Unlike the former, the latter remains unaffected by non-Lambertian surfaces, assuming a robust ϕM. Stereo Correlation Volume. Given FL,FR, we construct a 3D correlation volume VSusing dot product between feature maps: (VS)ijk =X h (FL)hij ·(FR)hik,VS∈RH 4×W 4×W 4(1) Monocular Correlation Volume. Given ML,MR, we downsample them to 1/4, compute their normals ∇L,∇R, and construct a 3D correlation volume VMusing dot product between normal maps: (VM)ijk =X h (∇L)hij ·(∇R)hik,VM∈RH 4×W 4×W 4 (2) Given the absence of texture in ∇Land ∇R, the resulting monocular volume VMwill be less informative. To alleviate this problem we segment VMusing the relative depth priors from MLand MR: to do so, we generate left and right segmentation masks ML∈ {0,1}H 4×W 4×1, MR∈ {0,1}H 4×1×W 4. We refer the reader to the supplementary material for a detailed description. Given the segmentation masks, we can generate masked volumes as: (VMn)ijk = (MLn)ij ·(MRn)ik ·(VM)ijk (3) Next, we insert a 3D Convolutional Regularization module ϕAto aggregate VMn, resulting in V′M= ϕA(VM 1,...,VMN,ML,MR), with N= 8. The architecture of ϕAfollows the one in [116], with a simple permutation to match the structure of the correlation volumes. We propose an adapted version of CoEx [4] correlation volume excitation that exploits both views. The resulting feature volumes V′M∈RF×H 4×W 4×W 4are fed to two different shallow 3D conv layers ϕDand ϕCto obtain two aggregated volumes VD M=ϕD(V′M)and VC M=ϕC(V′M) with VD M,VC M∈RH 4×W 4×W 4. Differentiable Monocular Scaling. Volume VD Mwill be used not only as a monocular guide for the iterative refinement unit but also to estimate the coarse disparity maps ˆ DLˆ DR, while VC Mis used to estimate confidence maps ˆ CL ˆ CR. These maps are then used to scale both MLand MR – (5) in Fig. 2. To estimate left disparity from a correlation volume, we first perform a softargmax on the last Wdimension of VD Mto extract the correlated pixel x-coordinate. Then, given the relationship between left disparity and correlation dL=jL−jR, we obtain a coarse disparity map ˆ DL: (ˆ DL)ij =j−softargmaxL(VD M)ij (4) Similarly, we estimate ˆ DRfrom VD M. We refer the reader to the supplementary for details. We also estimate a pair of confidence maps ˆ CL,ˆ CR∈[0,1]H×Wto classify outliers and perform robust scaling. Inspired by information entropy, we measure the chaos within correlation curves: clear monomodal-like cost curves—those with low entropy—are reliable, while chaotic curves with high entropy indicate uncertainty. To estimate the left confidence map, we perform asoftmax operation on the last Wdimension of VC M, then ˆ CLis obtained as follows: (ˆ CL)ij = 1 + PW 4 d e(VC M)ijd P W 4 fe(VC M)ijf ·log2 e(VC M)ijd P W 4 fe(VC M)ijf ! log2(W 4) (5) In the same way, we estimate ˆ CR. To further reduce outliers, we mask out occluded pixels from ˆ CLand ˆ CRusing aSoftLRC operator – see the supplementary material for details. Finally, we estimate the scale ˆsand shift ˆ tusing a differentiable weighted least-square approach: min ˆs,ˆ t L,R X  pˆ C⊙hˆsM+ˆ t−ˆ Di  F(6) where ∥·∥Fdenotes the Frobenius norm. Using the scaling coefficients, we obtain two disparity maps ˆ ML,ˆ MR: ˆ ML= ˆsML+ˆ t, ˆ MR= ˆsMR+ˆ t(7) 1016 Image Ground-Truth Depth Anything v2 [121] Image Ground-Truth Depth Anything v2 [121] Figure 3. Samples from MonoTrap Dataset. We report two scenes featured in our dataset, showing the left image, the ground-truth depth, and the predictions by Depth Anything v2 [121], highlighting how it fails in the presence of visual illusions. It is crucial to optimize both left and right scaling jointly to obtain consistency between ˆ MLand ˆ MR. Volume Augmentations. Unfortunately, Stereo Anywhere cannot properly learn when to choose stereo or mono information from [61] alone. Hence, we propose three volume augmentations and a monocular augmentation to overcome this issue: 1) Volume Rolling: we randomly apply a rolling operation to the last Wdimension of VDMor VS; 2) Volume Noising: we apply random noise sampled from the interval [0,1) using a uniform distribution; 3) Volume Zeroing: we apply a Gaussian-like curve with the peak where disparity equals zero. Furthermore, we randomly substitute the monocular depth with ground truth normalized between [0,1] as an additional augmentation. We apply only one volume augmentation to VDMor VSand only for a section of the volume, randomly selecting an Mn Lmask. Volume Truncation. To further help Stereo Anywhere to handle mirror surfaces, we introduce a hand-crafted volume truncation operation on VS. Firstly, we extract left confidence CM=softLRCL(ˆ ML,ˆ MR)to classify reliable monocular predictions. Then, we create a truncate mask T∈[0,1]H 4×W 4using the following logic condition: (T)ij =h(ˆ ML)ij >(ˆ DL)ij∧(CM)iji∨ h(CM)ij ∧ ¬(ˆ CL)iji. We implement this logic using fuzzy operators (more details in the supplementary material). The rationale is that stereo predicts farther depths on mirror surfaces: the mirror is perceived as a window into a new environment, specular to the real one. Finally, for values of T> Tm= 0.98, we truncate VSusing a sigmoid curve centered at the correlation value predicted by ˆ ML– i.e., the real disparity of mirror surfaces – preserving only the stereo correlation curve not “piercing” mirrors. 3.3. Iterative Disparity Estimation We aim to estimate a series of refined disparity maps {D1= ˆ ML,D2, . . . Dl, . . . }exploiting the guidance from both stereo and mono branches. Starting from the Multi-GRU update operator by [55], we introduce a second lookup operator that extracts correlation features GMfrom the additional volume VD M– (7) in Fig. 2. The two sets of correlation features from GSand GMare processed by the same two-layer encoder and concatenated with features derived from the current disparity estimation Dl. This concatenation is further processed by a 2D conv layer, and then by the ConvGRU operator. We inherit the convex upsampling module [55] to upsample final disparity to full resolution. 3.4. Training Supervision We supervise the iterative module using the well-known L1 loss with exponentially increasing weights [55], then ˆ DL, ˆ DR,ˆ MLand ˆ MRusing the L1 loss, finally ˆ CLand ˆ CR using the Binary Cross Entropy loss. We invite the reader to read the supplementary material for additional details. 4. The MonoTrap Dataset Monocular depth estimation is known for possibly failing in the presence of perspective illusions. The reader may wonder how Stereo Anywhere would behave in such cases: would it blindly trust the monocular VFM or rely on the stereo geometric principles to maintain robustness? To answer these questions, we introduce MonoTrap, a novel stereo dataset specifically designed to challenge monocular depth estimation. Our dataset comprises 26 scenes featuring perspective illusions, captured with a calibrated stereo setup and annotated with ground-truth depth from an Intel Realsense L515 LiDAR. The scenes contain carefully designed planar patterns that create visual illusions, such as apparent holes in walls or floors and simulated transparent surfaces that reveal content behind them. Figure 3shows examples from our dataset that illustrate how these visual illusions easily fool monocular methods. 5. Experiments We describe our implementation details, datasets, and evaluation protocols, followed by experiments. We also refer the reader to the supplementary material for more results. 5.1. Implementation and Experimental Settings We implement Stereo Anywhere using PyTorch, starting from RAFT-Stereo codebase [55]. We use Depth Anything v2 [121] as the VFM fueling our model, using the Large weights provided by the authors, trained on ground-truth labels from the HyperSim synthetic dataset [81] only. Starting from the Sceneflow RAFT-Stereo checkpoint, we train Stereo Anywhere on a single A100 GPU for 3 epochs, with learning rate 1e-4 and AdamW optimizer, on 1017 Booster (Q) Middlebury 2014 (H) Experiment bad Avg. bad >2Avg. >2>4>6>8(px) All Noc Occ (px) (A) Baseline [55] 17.84 13.06 10.76 9.24 3.59 11.15 8.06 29.06 1.55 (B) (A) + Monocular Context w/o re-train 15.85 10.98 8.89 7.69 3.05 14.96 11.70 34.38 2.82 (C) (A) + Monocular Context w/ re-train 14.94 10.40 8.61 7.63 3.03 9.62 6.98 25.39 1.13 (D) (C) + Normals Correlation Volume / Scaled Depth 11.33 6.88 5.32 4.59 1.87 7.67 5.24 21.51 0.96 (E) (D) + Volume augmentation / truncation 9.01 5.40 4.12 3.34 1.21 6.96 4.75 20.34 0.94 Table 1. Ablation Studies. We measure the impact of different design strategies. Networks trained on SceneFlow [61]. Middlebury 2014 (H) Middlebury 2021 ETH3D KITTI 2012 KITTI 2015 Model bad >2Avg. bad >2Avg. bad >1Avg. bad >3Avg. bad >3Avg. All Noc Occ (px) All Noc Occ (px) All Noc Occ (px) All Noc Occ (px) All Noc Occ (px) RAFT-Stereo [55] 11.15 8.06 29.06 1.55 12.05 9.38 37.89 1.81 2.59 2.24 8.78 0.25 4.80 4.23 29.21 0.89 5.44 5.21 14.09 1.16 PSMNet [8] 18.79 13.80 53.22 4.63 23.67 20.61 53.75 5.70 19.75 18.62 42.05 0.94 6.73 5.81 46.24 1.22 6.78 6.40 24.85 1.38 GMStereo [117] 15.63 10.98 46.04 1.87 25.43 22.43 54.70 2.86 6.22 5.58 19.97 0.42 5.68 4.87 38.84 1.10 5.72 5.44 17.33 1.21 ELFNet [59] 24.48 16.94 77.06 8.61 27.08 21.77 85.56 11.01 25.61 24.50 46.06 5.65 10.52 8.67 88.21 2.30 9.61 8.22 85.64 2.16 PCVNet [132] 16.79 13.54 35.66 2.96 12.92 10.19 40.23 2.18 4.24 3.61 14.01 0.41 4.44 3.92 27.70 0.89 5.08 4.88 13.72 1.24 DLNR [140] 9.46 6.20 28.75 1.45 8.44 5.88 32.71 1.24 23.12 22.94 26.93 9.89 9.45 8.83 36.75 1.59 15.74 15.41 34.32 2.83 Selective-RAFT [110] 12.05 9.46 27.42 2.35 15.69 13.86 36.32 5.92 4.36 3.81 10.23 0.34 5.71 5.16 30.54 1.08 6.50 6.22 18.44 1.27 Selective-IGEV [110] 9.98 7.09 27.62 1.60 8.89 6.34 32.88 1.60 6.42 5.71 18.71 1.73 6.22 5.54 34.78 1.09 5.87 5.66 14.99 1.42 IGEV-Stereo [116] 9.91 7.08 26.26 1.84 9.15 6.43 34.88 1.53 4.30 3.86 12.65 0.38 5.65 4.43 33.38 1.03 5.87 5.13 14.31 1.34 NMRF [28] 14.08 10.87 34.62 2.91 23.36 21.69 42.51 8.57 4.34 3.66 17.15 0.42 4.62 4.05 30.65 0.92 5.24 5.07 12.28 1.16 Stereo Anywhere (ours) 6.96 4.75 20.34 0.94 7.97 5.71 29.52 1.08 1.66 1.43 5.29 0.24 3.90 3.52 21.65 0.83 3.93 3.79 11.01 0.97 Table 2. Zero-shot Generalization. Comparison with state-of-the-art deep stereo models. Networks trained on SceneFlow [61]. batches of 2 images. We extract random crops of size 320×640 from images and apply standard color and spatial augmentations [55]. The VFM is used only to source monocular depth maps, remaining frozen during training. The number of iterations for GRUs is fixed to 12 during training and increased to 32 at inference time. 5.2. Evaluation Datasets & Protocol Datasets. We utilize SceneFlow [61] as our sole training dataset, comprising about 39k synthetic stereo pairs with dense ground-truth disparities. For evaluation, we employ several benchmarks: Middlebury 2014 [86] and its 2021 extension [65] provide high-resolution indoor scenes with semi-dense labels (15 and 24 stereo pairs), KITTI 2012 [23] and 2015 [62] feature outdoor driving scenarios (∼200 pairs each at 1280 ×384 with sparse LiDAR ground truth), and ETH3D [87] contributes 27 low-resolution indoor/outdoor scenes. For non-Lambertian surfaces, we primarily use Booster [127], containing 228 high-resolution (12 Mpx) indoor pairs with its 191-pair online benchmark, and LayeredFlow [115], featuring 400 pairs with transparent objects and sparse ground truth (∼50 points per pair). Additionally, we include our newly proposed MonoTrap dataset focusing on optical illusions. For zero-shot evaluation, we test on KITTI 2015, Middlebury v3 at half (H) resolution, Middlebury 2021, and ETH3D, while non-Lambertian zero-shot testing relies on Booster at quarter (Q) resolution and LayeredFlow at eight (E) resolution. Evaluation Metrics. We evaluate our method using two standard metrics: the average pixel error (Avg.), which computes the absolute difference between predicted and ground truth disparities averaged over all pixels, and the bad> τ error, which measures the percentage of pixels with a disparity error greater than τpixels – for the latter, we compute it considering all pixels or either non-occluded or occluded pixels, referred to as All,Noc or Occ respectively. We evaluate on MonoTrap through standard monocular depth metrics [25] - Absolute relative error (AbsRel), RMSE, and δ < 1.05 score. 5.3. Ablation Study We start our analysis by evaluating how individual components of our model contribute to the overall accuracy. All model variants are trained solely on the synthetic SceneFlow dataset and tested on Booster and Middlebury 2014, allowing us to examine their effectiveness on nonLambertian surfaces and general scenes. Table 1summarizes our findings. In (A), we report the performance of our baseline model, upon which we build Stereo Anywhere– i.e., RAFT-stereo [55]. On the one hand, by adding monocular context from an off-the-shelf monocular depth network to the pre-trained context backbone (B), we observe improved performance on non-Lambertian surfaces, though at the expense of a general drop in accuracy on Middlebury. On the other hand, by re-training the context backbone to process depth maps obtained from the monocular network on SceneFlow (C), we can appreciate a consistent improvement in both datasets. Introducing the normals correlation volume with subsequent differentiable depth scaling (D) significantly enhances the accuracy on non-Lambertian surfaces, also showing improvements on indoor scenes. Finally, cost volume augmentations and truncation (E) demonstrate positive effects on transparent surfaces and mirrors present in the Booster dataset by further reducing the bad-2 metric by approximately 1.5% and Avg. by 0.7 pixels, with minimal influence on Middlebury. According to these results, from now on, we will adopt (E) as the default setting for Stereo Anywhere. 1018 RGB RAFT-Stereo [55] DLNR [140] NMRF [28] Selective-IGEV [110]Stereo Anywhere KITTI 15 Middlebury ETH3D Figure 4. Qualitative Results – Zero-Shot Generalization. Predictions by state-of-the-art models and Stereo Anywhere. Booster (Q) LayeredFlow (E) Model Error Rate (%) Avg. Error Rate (%) Avg. >2>4>6>8(px) >1>3>5(px) RAFT-Stereo [55] 17.84 13.06 10.76 9.24 3.59 89.21 79.02 71.61 19.27 PSMNet [8] 34.47 24.83 20.46 17.77 7.26 91.85 79.84 70.04 21.18 GMStereo [117] 32.44 22.52 17.96 15.02 5.29 92.95 83.68 74.76 20.91 ELFNet [59] 45.52 35.79 30.72 27.33 14.04 93.08 82.24 70.41 20.19 PCVNet [132] 22.63 16.51 13.81 12.08 4.70 88.27 76.65 66.79 18.19 DLNR [140] 18.56 14.55 12.61 11.22 3.97 89.90 79.46 72.72 18.97 Selective-RAFT [110] 20.01 15.08 12.52 10.88 4.12 92.69 86.32 78.82 20.18 Selective-IGEV [110] 18.52 14.24 12.14 10.77 4.38 91.31 81.72 74.74 19.65 IGEV-Stereo [116] 16.90 13.23 11.40 10.20 3.94 87.28 80.07 72.91 19.07 NMRF [28] 27.08 19.06 15.43 13.21 5.02 89.08 79.13 70.51 20.17 Stereo Anywhere (ours) 9.01 5.40 4.12 3.34 1.21 81.83 57.66 45.12 11.20 Booster (Q) Online Benchmark LayeredFlow (E) DKT-RAFT [136] (*) 10.32 7.13 5.65 4.36 1.70 66.05 46.95 37.77 8.72 Stereo Anywhere (ours) (*) 6.52 2.82 1.77 1.27 0.73 51.24 25.63 15.65 4.84 Table 3. Zero-shot Non-Lambertian Generalization. Comparison with state-of-the-art models. Networks trained on SceneFlow [61]. (*) means fine-tuned on Booster training set. 5.4. Zero-Shot Generalization We now compare our Stereo Anywhere model against stateof-the-art deep stereo networks, assessing zero-shot generalization capability when transferred from synthetic to real images. Purposely, we follow a well-established benchmark in the literature [55,104], evaluating on real datasets models pre-trained exclusively on SceneFlow [61]. Table 2compares Stereo Anywhere with off-the-shelf stereo networks using authors’ provided weights. Considering All, Noc, and Avg. metrics, we can notice how Stereo Anywhere achieves consistently better results across most datasets, achieving almost 3% lower bad-2 All on Middlebury 2014 versus the second-best method DLNR [140], and breaking the 4% barrier on KITTI’s bad-3 All metric. The Occ metric further demonstrates how Stereo Anywhere consistently outperforms other stereo models on any dataset, with substantial margins over the second-best – i.e., approximately 6% on Middlebury 2014 and KITTI 2012, and 3% on ETH3D. This confirms that leveraging priors from VFMs for monocular depth estimation effectively improve the stereo matching estimation accuracy in challenging conditions where stereo matching is ill-posed, such as at occluded regions. Figure 4shows predictions on KITTI 2015, Middlebury 2014, and ETH3D samples. In particular, the first row shows an extremely challenging case for SceneFlow-trained models, where Stereo Anywhere achieves accurate disparity maps thanks to VFM priors. 5.5. Zero-Shot Non-Lambertian Generalization We now assess the generalization capabilities of Stereo Anywhere and existing stereo models when dealing with non-Lambertian materials, such as transparent surfaces or mirrors. To this end, we conduct a zero-shot generalization evaluation experiment on the Booster [74] and LayeredFlow [115] datasets, once again using models pre-trained on SceneFlow [61] – with weights provided by the authors. Table 3shows the outcome of this evaluation. This time, we can perceive even more clearly how Stereo Anywhere is the absolute winner, demonstrating unprecedented robustness in the presence of non-Lambertian surfaces despite being trained only on synthetic stereo data, not even featuring such objects. These results further validate how leveraging strong priors from existing VFMs for monocular depth esti1019 RGB RAFT-Stereo [55] DLNR [140] NMRF [28] Selective-IGEV [110]Stereo Anywhere Booster LayeredFlow Figure 5. Qualitative results – Zero-Shot non-Lambertian Generalization. Predictions by state-of-the-art models and Stereo Anywhere. MonoTrap Model AbsRel RMSE σ < 1.05 (%)↓(m)↓(%)↑ Depth Anything v2 [121] 53.46 0.36 15.21 Depth Anything v2 [121]†27.92 0.27 19.43 DepthPro [6] 47.77 0.32 21.90 DepthPro [6]†20.82 0.22 22.88 RAFT-Stereo [55] 5.01 0.09 77.05 Stereo Anywhere 3.50 0.06 80.27 Table 4. MonoTrap Benchmark. Comparison with state-of-theart monocular depth estimation models and RAFT-Stereo. Both RAFT-Stereo and Stereo Anywhere are trained on SceneFlow [61]. †refers to robust scaling through RANSAC. mation can play a game-changing role in stereo matching as well, especially when lacking training data explicitly targeting critical conditions such as non-Lambertian surfaces. At the bottom, we report results achieved by fine-tuning Stereo Anywhere on the Booster training set and evaluating on the online benchmark. Our model ranks first when evaluated at quarter resolution. Figure 5shows examples from Booster and LayeredFlow, where Stereo Anywhere is the only stereo model correctly perceiving the mirror and transparent railing. 5.6. MonoTrap Benchmark We conclude our evaluation by running experiments on our newly collected MonoTrap dataset to prove the robustness of Stereo Anywhere in the presence of critical conditions harming the accuracy of monocular depth predictors. Table 4collects the results achieved by state-of-the-art monocular depth estimation models, the baseline stereo model over which we built our framework (RAFT-Stereo) and Stereo Anywhere. Regarding the former models, as they predict affine-invariant depth maps, following the literature [78] we use least square errors to align them to the ground-truth. As these models are fooled by the visual illusions, this scaling procedure is likely to yield sub-optimal scale and shift parameters. Therefore, we alternatively align to ground-truth depth through a more robust RANSAC fitting – denoted with †in the table. On the one hand, by comparing monocular and stereo methods, we notice how the failures of the former negatively impact their evaluation metrics. Once again, we reRGB D. Anything v2 [121]Stereo Anywhere Figure 6. Qualitative results – MonoTrap. Stereo Anywhere is not fooled by erroneous predictions by its monocular engine [121]. mark that a direct comparison across the two families of methods is not the main goal of this experiment. On the other hand, we focus on the comparison between RAFTStereo and Stereo Anywhere, with our model performing slightly better than its baseline. This fact proves that despite its strong reliance on the priors retrieved from VFMs for monocular depth estimation, Stereo Anywhere can properly ignore such priors when unreliable. Figure 6shows three samples where Depth Anything v2 fails while Stereo Anywhere does not. 6. Conclusion In this paper, we introduced Stereo Anywhere, a novel stereo matching framework that leverages monocular depth VFMs to overcome traditional stereo matching limitations. Combining stereo geometric constraints with monocular priors, our approach demonstrates superior zero-shot generalization and robustness to challenging conditions like textureless regions, occlusions, and non-Lambertian surfaces. Furthermore, through our novel MonoTrap dataset, we showed that Stereo Anywhere effectively combines the best of both worlds - maintaining stereo matching’s geometric accuracy where monocular methods fail, while leveraging monocular priors to handle challenging stereo scenarios. Extensive comparisons against state-of-the-art networks in zero-shot settings validate these findings. 1020 Acknowledgement. This study was carried out within the MOST – Sustainable Mobility National Research Center and received funding from the European Union Next-GenerationEU – PIANO NAZIONALE DI RIPRESA E RESILIENZA (PNRR) – MISSIONE 4 COMPONENTE 2, INVESTIMENTO 1.4 – D.D. 1033 17/06/2022, CN00000023. This manuscript reflects only the authors’ views and opinions, neither the European Union nor the European Commission can be considered responsible for them. This study was funded by the European Union – Next Generation EU within the framework of the National Recovery and Resilience Plan NRRP – Mission 4 “Education and Research” – Component 2 - Investment 1.1 “National Research Program and Projects of Significant National Interest Fund (PRIN)” (Call D.D. MUR n. 104/2022) – PRIN2022 – Project reference: “RiverWatch: a citizen-science approach to river pollution monitoring” (ID: 2022MMBA8X, CUP: J53D23002260006). We also acknowledge the CINECA award under the ISCRA initiative, for the availability of high-performance computing resources and support. References [1] Filippo Aleotti, Fabio Tosi, Li Zhang, Matteo Poggi, and Stefano Mattoccia. Reversing the cycle: self-supervised deep stereo through enhanced monocular distillation. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XI 16, pages 614–632. Springer, 2020. 2,3 [2] Filippo Aleotti, Fabio Tosi, Pierluigi Zama Ramirez, Matteo Poggi, Samuele Salti, Stefano Mattoccia, and Luigi Di Stefano. Neural disparity refinement for arbitrary resolution stereo. In 2021 International Conference on 3D Vision (3DV), pages 207–217. IEEE, 2021. 2,3 [3] Vasileios Arampatzakis, George Pavlidis, Nikolaos Mitianoudis, and Nikos Papamarkos. Monocular depth estimation: A thorough review. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023. 1 [4] Antyanta Bangunharcana, Jae Won Cho, Seokju Lee, In So Kweon, Kyung-Soo Kim, and Soohyun Kim. Correlateand-excite: Real-time stereo matching via guided cost volume excitation. In IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2021. 2,4 [5] Luca Bartolomei, Matteo Poggi, Fabio Tosi, Andrea Conti, and Stefano Mattoccia. Active stereo without pattern projector. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 18470–18482, 2023. 2 [6] Aleksei Bochkovskii, Ama¨ el Delaunoy, Hugo Germain, Marcel Santos, Yichao Zhou, Stephan R. Richter, and Vladlen Koltun. Depth pro: Sharp monocular metric depth in less than a second. arXiv, 2024. 8 [7] Changjiang Cai, Matteo Poggi, Stefano Mattoccia, and Philippos Mordohai. Matching-space stereo networks for cross-domain generalization. In 2020 International Conference on 3D Vision (3DV), pages 364–373, 2020. 2 [8] Jia-Ren Chang and Yong-Sheng Chen. Pyramid stereo matching network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 5410– 5418, 2018. 2,6,7 [9] Liyan Chen, Weihan Wang, and Philippos Mordohai. Learning the distribution of errors in stereo matching for joint disparity and uncertainty estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17235–17244, 2023. 2 [10] Weifeng Chen, Zhao Fu, Dawei Yang, and Jia Deng. Single-image depth perception in the wild. In Proceedings of the 30th International Conference on Neural Information Processing Systems, page 730–738, Red Hook, NY, USA, 2016. Curran Associates Inc. 3 [11] Xihao Chen, Zhiwei Xiong, Zhen Cheng, Jiayong Peng, Yueyi Zhang, and Zheng-Jun Zha. Degradation-agnostic correspondence from resolution-asymmetric stereo. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 12962–12971, 2022. 3 [12] Zhi Chen, Xiaoqing Ye, Wei Yang, Zhenbo Xu, Xiao Tan, Zhikang Zou, Errui Ding, Xinming Zhang, and Liusheng Huang. Revealing the reciprocal relations between selfsupervised stereo and monocular depth estimation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 15529–15538, 2021. 3 [13] Ziyang Chen, Wei Long, He Yao, Yongjun Zhang, Bingshu Wang, Yongbin Qin, and Jia Wu. Mocha-stereo: Motif channel attention network for stereo matching. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024. 2 [14] Junda Cheng, Longliang Liu, Gangwei Xu, Xianqi Wang, Zhaoxing Zhang, Yong Deng, Jinliang Zang, Yurui Chen, Zhipeng Cai, and Xin Yang. Monster: Marry monodepth to stereo unleashes power. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025. 3 [15] Kelvin Cheng, Tianfu Wu, and Christopher Healey. Revisiting non-parametric matching cost volumes for robust and generalizable stereo matching. Advances in Neural Information Processing Systems, 35:16305–16318, 2022. 2 [16] Jaehoon Cho, Dongbo Min, Youngjung Kim, and Kwanghoon Sohn. Diml/cvl rgb-d dataset: 2m rgb-d images of natural indoor and outdoor scenes. arXiv preprint arXiv:2110.11590, 2021. 3 [17] WeiQin Chuah, Ruwan Tennakoon, Reza Hoseinnezhad, Alireza Bab-Hadiashar, and David Suter. Itsa: An information-theoretic approach to automatic shortcut avoidance and domain generalization in stereo matching networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 13022–13032, 2022. 2 [18] Alex Costanzino, Pierluigi Zama Ramirez, Matteo Poggi, Fabio Tosi, Stefano Mattoccia, and Luigi Di Stefano. Learning depth estimation for transparent and mirror surfaces. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9244–9255, 2023. 3 [19] Yiqun Duan, Xianda Guo, and Zheng Zhu. DiffusionDepth: Diffusion denoising approach for monocular depth estimation. arXiv preprint arXiv:2303.05021, 2023. 3 [20] Ainaz Eftekhar, Alexander Sax, Jitendra Malik, and Amir Zamir. Omnidata: A scalable pipeline for making multi1021