Loop Closure from Two Views: Revisiting PGO for Scalable Trajectory Estimation through Monocular Priors

1 DSO National Laboratories    2 ETH Zurich (CVG)    3 Microsoft    4 University of Bonn (RPL)
* Work done while at CVG
IEEE Robotics and Automation Letters (RA-L), 2026
2GO refining a 1.1 hour, 2.7 km quadruped robot trajectory overlaid on a 3D reconstruction, with loop-closure edges shown

2GO refines a 1.1 hr, 2.7 km trajectory from a quadruped robot using only two-view loop closures, without 3D reconstruction or Bundle Adjustment. Absolute and scale-free loop-closure edges are shown in yellow and blue; edges with both constraints are shown in red.

TL;DR

Bundle Adjustment (BA) is standard in Visual SLAM, but its cubic complexity in the number of poses scales poorly to long trajectories. 2GO is a reconstruction-free approach that refines trajectories using only two-view loop closures, leveraging recent advances in Visual Place Recognition, image matching, and Monocular Depth Estimation.
2GO:

  • supports real-time performance,
  • scales well in map size and trajectory duration,
  • and has competitive trajectory accuracy.

Abstract

(Visual) Simultaneous Localization and Mapping (SLAM) remains a fundamental challenge in enabling autonomous systems to navigate and understand large-scale environments. Traditional SLAM approaches struggle to balance efficiency and accuracy, particularly in large-scale settings where extensive computational resources are required for scene reconstruction and Bundle Adjustment (BA). However, this scene reconstruction, in the form of sparse point-clouds of visual landmarks, is often only used within the SLAM system because navigation and planning methods require different map representations. In this work, we therefore investigate a more scalable Visual SLAM (VSLAM) approach without reconstruction, mainly based on approaches for two-view loop closures. By restricting the map to a sparse keyframed pose graph without dense geometry representations, our '2GO' system achieves efficient optimization with competitive absolute trajectory accuracy. In particular, we find that recent advancements in image matching and monocular depth priors enable very accurate trajectory optimization without BA. We conduct extensive experiments on diverse datasets, including large-scale scenarios, and provide a detailed analysis of the trade-offs between runtime, accuracy, and map size. Our results demonstrate that this streamlined approach supports realtime performance, scales well in map size and trajectory duration, and effectively broadens the capabilities of VSLAM for long duration deployments to large environments.

System Overview

2GO system overview: keyframing, image retrieval, two-view pose estimation, and pose graph optimization

2GO takes a stream of posed, calibrated images as input, for example from any existing odometry source.

  1. Keyframing: new keyframes are selected based on distance from the previous keyframe. Because there is no landmark map, consecutive keyframes do not need covisible keypoints, which allows a sparse pose graph.
  2. Image Retrieval identifies loop-closure (LC) candidates via visual similarity (global descriptors from the BoQ VPR model) and spatial proximity (previous keyframes that are close in the current pose estimate).
  3. Two-view Pose Estimation establishes LC-Edges: pose constraints between the current and previous keyframes, both up-to-scale (Scale-free) and metric (Absolute). Candidates are filtered by image-matching feasibility, geometric feasibility, and pose-graph consistency.
  4. Pose Graph Optimization (iSAM2 in GTSAM, with a robust Cauchy loss) jointly optimizes odometry and LC-Edges to refine the trajectory.

Loop Closures from Two Views

Two-view pose estimation with the Correspond-Lift (CL) and Multi-View (MV) variants

From 2D-2D correspondences alone, the relative pose between two views is only known up to scale. 2GO adds these as Scale-free LC-Edges, and additionally lifts keypoints to 3D so that PnP on 2D-3D correspondences gives a metric-scale Absolute LC-Edge. We investigate two variants:

  • CL (Correspond-Lift) finds correspondences between views (DISK + LightGlue), then uses metric monocular depth (Metric3Dv2) to lift keypoints from one view into 3D.
  • MV (Multi-View) uses a multi-view model such as MASt3R to jointly predict 2D keypoint matches and 2D-3D correspondences with a single model.

Both variants are modular: as newer two-view models become available, they can be swapped in easily. With state-of-the-art matchers, 2GO incorporates loop closures even under extreme viewpoint changes.

Results

Accuracy

2GO is evaluated on KITTI and the Vision Benchmark in Rome (VBR), including trajectories longer than 20 minutes. We compare against classical VSLAM run in stereo-only mode, and Maplab 2.0, a multi-sensor mapping framework that performs global BA. 2GO achieves the best mean accuracy on the large-scale VBR sequences.

Mean RMSE ATE (m) ↓
Dataset VINS-Fusion VINS-Fusion + PGO ORB-SLAM3 Maplab-BIN Maplab-SP 2GO (Ours)
KITTI 7.18 5.04 3.29 16.55 11.55 3.31
VBR (Driving) 68.31 49.47 25.47 66.25 63.23 14.73
VBR (Handheld) 20.32 16.86 11.61 20.17 20.07 9.90

Scalability

On spagna_train0, the longest handheld VBR sequence (1.56 km, 23 min 34 s), 2GO needs only 0.83× real-time to process the sequence. It stores the smallest map (146 MB) and reaches the lowest trajectory error of all methods.

Runtime and map size on VBR spagna_train0 (1414 s real-time duration)
VINS-Fusion VINS-PGO ORB-SLAM3 Maplab-BIN Maplab-SP 2GO (Ours)
Runtime (s) ↓ 1564 6206 1755 110 (Map)
29 (Opt)
14139 (Map)
271 (Opt)
1167
Map Size (MB) ↓ / / 1136 319 572 146
RMSE ATE (m) ↓ 21.14 13.22 5.62 20.83 20.35 3.73

Real-World Robot Deployment

We record a 1.1 hr, 2.7 km trajectory with a Boston Dynamics Spot robot, covering multiple floors of a building and a short outdoor segment. Using Spot's own odometry as input, 2GO reduces RMSE ATE from 2.44 m to 0.66 m. It needs a cumulative runtime of only 15.0 min and saves a 162.5 MB map.

RMSE ATE (m) on the 1.1 hr real-world Spot trajectory
Spot Odom ZED Odom ZED w/ LC 2GO (Ours)
2.44 26.18 23.05 0.66

Qualitative Results

Conclusion and Outlook

  • 2GO is both scalable and accurate on long-duration trajectories, and introduces a map-free approach to global trajectory consistency that uses only two-view loop closures.
  • As two-view methods keep improving, 2GO's modularity allows components (e.g. VPR, MDE) to be easily upgraded.
  • 2GO's two-view architecture is deployable on resource-constrained systems, especially those with limited GPU resources.

Poster

BibTeX

@ARTICLE{11425787,
  author={Lim, Tian Yi and Sun, Boyang and Pollefeys, Marc and Blum, Hermann},
  journal={IEEE Robotics and Automation Letters}, 
  title={Loop Closure From Two Views: Revisiting PGO for Scalable Trajectory Estimation Through Monocular Priors}, 
  year={2026},
  volume={11},
  number={5},
  pages={5326-5333},
  keywords={Simultaneous localization and mapping;Trajectory;Accuracy;Image reconstruction;Odometry;Optimization;Visualization;Robots;Image retrieval;Image edge detection;SLAM;localization;mapping},
  doi={10.1109/LRA.2026.3671555}}