Loop Closure from Two Views: Revisiting PGO for Scalable Trajectory Estimation through Monocular Priors
TL;DR
Bundle Adjustment (BA) is standard in Visual SLAM, but its cubic complexity in the number of poses
scales
poorly to long trajectories. 2GO is a reconstruction-free approach that refines
trajectories using only two-view loop closures, leveraging recent advances in Visual
Place Recognition, image matching, and Monocular Depth Estimation.
2GO:
- supports real-time performance,
- scales well in map size and trajectory duration,
- and has competitive trajectory accuracy.
Abstract
(Visual) Simultaneous Localization and Mapping (SLAM) remains a fundamental challenge in enabling autonomous systems to navigate and understand large-scale environments. Traditional SLAM approaches struggle to balance efficiency and accuracy, particularly in large-scale settings where extensive computational resources are required for scene reconstruction and Bundle Adjustment (BA). However, this scene reconstruction, in the form of sparse point-clouds of visual landmarks, is often only used within the SLAM system because navigation and planning methods require different map representations. In this work, we therefore investigate a more scalable Visual SLAM (VSLAM) approach without reconstruction, mainly based on approaches for two-view loop closures. By restricting the map to a sparse keyframed pose graph without dense geometry representations, our '2GO' system achieves efficient optimization with competitive absolute trajectory accuracy. In particular, we find that recent advancements in image matching and monocular depth priors enable very accurate trajectory optimization without BA. We conduct extensive experiments on diverse datasets, including large-scale scenarios, and provide a detailed analysis of the trade-offs between runtime, accuracy, and map size. Our results demonstrate that this streamlined approach supports realtime performance, scales well in map size and trajectory duration, and effectively broadens the capabilities of VSLAM for long duration deployments to large environments.
System Overview
2GO takes a stream of posed, calibrated images as input, for example from any existing odometry source.
- Keyframing: new keyframes are selected based on distance from the previous keyframe. Because there is no landmark map, consecutive keyframes do not need covisible keypoints, which allows a sparse pose graph.
- Image Retrieval identifies loop-closure (LC) candidates via visual similarity (global descriptors from the BoQ VPR model) and spatial proximity (previous keyframes that are close in the current pose estimate).
- Two-view Pose Estimation establishes LC-Edges: pose constraints between the current and previous keyframes, both up-to-scale (Scale-free) and metric (Absolute). Candidates are filtered by image-matching feasibility, geometric feasibility, and pose-graph consistency.
- Pose Graph Optimization (iSAM2 in GTSAM, with a robust Cauchy loss) jointly optimizes odometry and LC-Edges to refine the trajectory.
Loop Closures from Two Views
From 2D-2D correspondences alone, the relative pose between two views is only known up to scale. 2GO adds these as Scale-free LC-Edges, and additionally lifts keypoints to 3D so that PnP on 2D-3D correspondences gives a metric-scale Absolute LC-Edge. We investigate two variants:
- CL (Correspond-Lift) finds correspondences between views (DISK + LightGlue), then uses metric monocular depth (Metric3Dv2) to lift keypoints from one view into 3D.
- MV (Multi-View) uses a multi-view model such as MASt3R to jointly predict 2D keypoint matches and 2D-3D correspondences with a single model.
Both variants are modular: as newer two-view models become available, they can be swapped in easily. With state-of-the-art matchers, 2GO incorporates loop closures even under extreme viewpoint changes.
Results
Accuracy
2GO is evaluated on KITTI and the Vision Benchmark in Rome (VBR), including trajectories longer than 20 minutes. We compare against classical VSLAM run in stereo-only mode, and Maplab 2.0, a multi-sensor mapping framework that performs global BA. 2GO achieves the best mean accuracy on the large-scale VBR sequences.
| Dataset | VINS-Fusion | VINS-Fusion + PGO | ORB-SLAM3 | Maplab-BIN | Maplab-SP | 2GO (Ours) |
|---|---|---|---|---|---|---|
| KITTI | 7.18 | 5.04 | 3.29 | 16.55 | 11.55 | 3.31 |
| VBR (Driving) | 68.31 | 49.47 | 25.47 | 66.25 | 63.23 | 14.73 |
| VBR (Handheld) | 20.32 | 16.86 | 11.61 | 20.17 | 20.07 | 9.90 |
Scalability
On spagna_train0, the longest handheld VBR sequence (1.56 km, 23 min 34 s), 2GO needs only 0.83× real-time to process the sequence. It stores the smallest map (146 MB) and reaches the lowest trajectory error of all methods.
| VINS-Fusion | VINS-PGO | ORB-SLAM3 | Maplab-BIN | Maplab-SP | 2GO (Ours) | |
|---|---|---|---|---|---|---|
| Runtime (s) ↓ | 1564 | 6206 | 1755 | 110 (Map) 29 (Opt) |
14139 (Map) 271 (Opt) |
1167 |
| Map Size (MB) ↓ | / | / | 1136 | 319 | 572 | 146 |
| RMSE ATE (m) ↓ | 21.14 | 13.22 | 5.62 | 20.83 | 20.35 | 3.73 |
Real-World Robot Deployment
We record a 1.1 hr, 2.7 km trajectory with a Boston Dynamics Spot robot, covering multiple floors of a building and a short outdoor segment. Using Spot's own odometry as input, 2GO reduces RMSE ATE from 2.44 m to 0.66 m. It needs a cumulative runtime of only 15.0 min and saves a 162.5 MB map.
| Spot Odom | ZED Odom | ZED w/ LC | 2GO (Ours) |
|---|---|---|---|
| 2.44 | 26.18 | 23.05 | 0.66 |
Qualitative Results
VBR ciampino_train0 (9.0 km): 2GO refines VINS-Fusion odometry, reducing RMSE ATE from 131.33 m to 22.84 m.
Scale-free LC-Edges from image pairs with large baselines (top: VBR diag_train0, bottom: our real-world recording). Both contribute to trajectory refinement.
Conclusion and Outlook
- 2GO is both scalable and accurate on long-duration trajectories, and introduces a map-free approach to global trajectory consistency that uses only two-view loop closures.
- As two-view methods keep improving, 2GO's modularity allows components (e.g. VPR, MDE) to be easily upgraded.
- 2GO's two-view architecture is deployable on resource-constrained systems, especially those with limited GPU resources.
Poster
BibTeX
@ARTICLE{11425787,
author={Lim, Tian Yi and Sun, Boyang and Pollefeys, Marc and Blum, Hermann},
journal={IEEE Robotics and Automation Letters},
title={Loop Closure From Two Views: Revisiting PGO for Scalable Trajectory Estimation Through Monocular Priors},
year={2026},
volume={11},
number={5},
pages={5326-5333},
keywords={Simultaneous localization and mapping;Trajectory;Accuracy;Image reconstruction;Odometry;Optimization;Visualization;Robots;Image retrieval;Image edge detection;SLAM;localization;mapping},
doi={10.1109/LRA.2026.3671555}}