Back to all papers

RC-EndoDepth: Two-stage self-supervised depth and motion estimation with improved scale consistency in endoscopic videos.

September 23, 2026pubmed logopapers

Authors

Han J,Li M,Zhang Z,Tan W,Wang Y

Affiliations (5)

  • School of Computer Science and Engineering, Northeastern University, No. 169, Baoyuan Street, Shenyang 110819, Liaoning, China; Key Laboratory of Medical Image Computing, Ministry of Education, Northeastern University, No. 169, Baoyuan Street, Shenyang 110819, Liaoning, China. Electronic address: [email protected].
  • School of Computer Science and Engineering, Northeastern University, No. 169, Baoyuan Street, Shenyang 110819, Liaoning, China; Key Laboratory of Medical Image Computing, Ministry of Education, Northeastern University, No. 169, Baoyuan Street, Shenyang 110819, Liaoning, China. Electronic address: [email protected].
  • School of Computer Science and Engineering, Northeastern University, No. 169, Baoyuan Street, Shenyang 110819, Liaoning, China; Key Laboratory of Medical Image Computing, Ministry of Education, Northeastern University, No. 169, Baoyuan Street, Shenyang 110819, Liaoning, China. Electronic address: [email protected].
  • School of Computer Science and Engineering, Northeastern University, No. 169, Baoyuan Street, Shenyang 110819, Liaoning, China; Key Laboratory of Medical Image Computing, Ministry of Education, Northeastern University, No. 169, Baoyuan Street, Shenyang 110819, Liaoning, China. Electronic address: [email protected].
  • School of Computer Science and Engineering, Northeastern University, No. 169, Baoyuan Street, Shenyang 110819, Liaoning, China; Key Laboratory of Medical Image Computing, Ministry of Education, Northeastern University, No. 169, Baoyuan Street, Shenyang 110819, Liaoning, China. Electronic address: [email protected].

Abstract

Maintaining a stable depth scale across endoscopic video frames remains challenging, which limits downstream applications such as surgical navigation and 3D reconstruction. Existing self-supervised methods often produce accurate relative depth for individual frames but still suffer from scale drift across a video. We present RC-EndoDepth, a two-stage self-supervised framework for endoscopic depth and motion estimation. Stage One learns scale-ambiguous relative depth, pose, and intrinsics from a large-scale endoscopic dataset. Stage Two performs sequence-level refinement by enforcing multi-frame depth consistency through pose-based projection, with the goal of improving within-sequence scale stability. Applied on top of two other self-supervised backbones, Stage Two improves most of their depth metrics as well, whereas its effect on pose depends on the pose head of the backbone. We also employ a Vision Transformer (ViT) based intrinsics network to obtain more stable camera parameter estimates when calibration is unavailable. Experiments show strong relative-depth performance on SCARED, SERV-CT, and C3VD, and improved sequence-level scale consistency on SCARED sequences, without ground-truth depth supervision; the sequence-level gains are largest within the temporal span each refinement adapted on. These results suggest that the proposed training strategy is a practical option for self-supervised 3D perception in endoscopic videos. Code, partial model checkpoints, and the project page are publicly available at https://han12580.github.io/RC-EndoDepth/index.html.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAISlice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.