{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,16]],"date-time":"2026-03-16T14:29:00Z","timestamp":1773671340057,"version":"3.50.1"},"reference-count":31,"publisher":"Springer Science and Business Media LLC","issue":"4","license":[{"start":{"date-parts":[[2025,2,28]],"date-time":"2025-02-28T00:00:00Z","timestamp":1740700800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,2,28]],"date-time":"2025-02-28T00:00:00Z","timestamp":1740700800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"JST CREST","award":["JPMJCR20D5"],"award-info":[{"award-number":["JPMJCR20D5"]}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J CARS"],"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:sec>\n            <jats:title>Purpose<\/jats:title>\n            <jats:p>Depth estimation is a powerful tool for navigation in laparoscopic surgery. Previous methods utilize predicted depth maps and the relative poses of the camera to accomplish self-supervised depth estimation. However, the smooth surfaces of organs with textureless regions and the laparoscope\u2019s complex rotations make depth and pose estimation difficult in laparoscopic scenes. Therefore, we propose a novel and effective self-supervised monocular depth estimation method with self-attention-guided pose estimation and a joint depth-pose loss function for laparoscopic images.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Methods<\/jats:title>\n            <jats:p>We extract feature maps and calculate the minimum re-projection error as a feature-metric loss to establish constraints based on feature maps with more meaningful representations. Moreover, we introduce the self-attention block in the pose estimation network to predict rotations and translations of the relative poses. In addition, we minimize the difference between predicted relative poses as the pose loss. We combine all of the losses as a joint depth-pose loss.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Results<\/jats:title>\n            <jats:p>The proposed method is extensively evaluated using SCARED and Hamlyn datasets. Quantitative results show that the proposed method achieves improvements of about 18.07<jats:inline-formula>\n                <jats:alternatives>\n                  <jats:tex-math>$$\\%$$<\/jats:tex-math>\n                  <mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                    <mml:mo>%<\/mml:mo>\n                  <\/mml:math>\n                <\/jats:alternatives>\n              <\/jats:inline-formula> and 14.00<jats:inline-formula>\n                <jats:alternatives>\n                  <jats:tex-math>$$\\%$$<\/jats:tex-math>\n                  <mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                    <mml:mo>%<\/mml:mo>\n                  <\/mml:math>\n                <\/jats:alternatives>\n              <\/jats:inline-formula> in the absolute relative error when combining all of the proposed components for depth estimation on SCARED and Hamlyn datasets. The qualitative results show that the proposed method produces smooth depth maps with low error in various laparoscopic scenes. The proposed method also exhibits a trade-off between computational efficiency and performance.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Conclusion<\/jats:title>\n            <jats:p>This study considers the characteristics of laparoscopic datasets and presents a simple yet effective self-supervised monocular depth estimation. We propose a joint depth-pose loss function based on the extracted feature for depth estimation on laparoscopic images guided by a self-attention block. The experimental results prove that all of the proposed components contribute to the proposed method. Furthermore, the proposed method strikes an efficient balance between computational efficiency and performance.<\/jats:p>\n          <\/jats:sec>","DOI":"10.1007\/s11548-025-03332-1","type":"journal-article","created":{"date-parts":[[2025,2,28]],"date-time":"2025-02-28T18:45:16Z","timestamp":1740768316000},"page":"775-785","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Enhanced self-supervised monocular depth estimation with self-attention and joint depth-pose loss for laparoscopic images"],"prefix":"10.1007","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4486-7207","authenticated-orcid":false,"given":"Wenda","family":"Li","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuichiro","family":"Hayashi","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7714-422X","authenticated-orcid":false,"given":"Masahiro","family":"Oda","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Takayuki","family":"Kitasaka","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kazunari","family":"Misawa","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0100-4797","authenticated-orcid":false,"given":"Kensaku","family":"Mori","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,2,28]]},"reference":[{"key":"3332_CR1","doi-asserted-by":"crossref","unstructured":"Takada C, Afifi A, Suzuki T, Nakaguchi T (2017) An enhanced hybrid tracking-mosaicking approach for surgical view expansion. In: 2017 39th annual international conference of the IEEE engineering in medicine and biology society (EMBC), pp 3692\u20133695","DOI":"10.1109\/EMBC.2017.8037659"},{"issue":"1","key":"3332_CR2","first-page":"87","volume":"42","author":"R Vecchio","year":"2000","unstructured":"Vecchio R, MacFayden B, Palazzo F (2000) History of laparoscopic surgery. Panminerva Med 42(1):87\u201390","journal-title":"Panminerva Med"},{"key":"3332_CR3","doi-asserted-by":"crossref","unstructured":"Qian L, Zhang X, Deguet A, Kazanzides P (2019) ARAMIS: augmented reality assistance for minimally invasive surgery using a head-mounted display. In: Medical image computing and computer assisted intervention\u2013MICCAI 2019: 22nd international conference, LNCS, vol 11768, pp 74\u201382","DOI":"10.1007\/978-3-030-32254-0_9"},{"key":"3332_CR4","doi-asserted-by":"crossref","unstructured":"Godard C, Mac\u00a0Aodha O, Firman M, Brostow GJ (2019) Digging into self-supervised monocular depth estimation. In: Proceedings of the IEEE\/CVF international conference on computer vision, pp 3828\u20133838","DOI":"10.1109\/ICCV.2019.00393"},{"key":"3332_CR5","doi-asserted-by":"crossref","unstructured":"Zhou T, Brown M, Snavely N, Lowe DG (2017) Unsupervised learning of depth and ego-motion from video. In: 2017 IEEE conference on computer vision and pattern recognition (CVPR), pp 6612\u20136619","DOI":"10.1109\/CVPR.2017.700"},{"key":"3332_CR6","doi-asserted-by":"crossref","unstructured":"Lyu X, Liu L, Wang M, Kong X, Liu L, Liu Y, Chen X, Yuan Y (2021) HR-depth: high resolution self-supervised monocular depth estimation. In: Proceedings of the AAAI conference on artificial intelligence, vol 35, pp 2294\u20132301","DOI":"10.1609\/aaai.v35i3.16329"},{"key":"3332_CR7","doi-asserted-by":"crossref","unstructured":"Zhao C, Poggi M, Tosi F, Zhou L, Sun Q, Tang Y, Mattoccia S (2023) GasMono: geometry-aided self-supervised monocular depth estimation for indoor scenes. In: Proceedings of the IEEE\/CVF international conference on computer vision, pp 16209\u201316220","DOI":"10.1109\/ICCV51070.2023.01485"},{"key":"3332_CR8","doi-asserted-by":"crossref","unstructured":"Saunders K, Vogiatzis G, Manso LJ (2023) Self-supervised monocular depth estimation: let\u2019s talk about the weather. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision, pp 8907\u20138917","DOI":"10.1109\/ICCV51070.2023.00818"},{"key":"3332_CR9","doi-asserted-by":"crossref","unstructured":"Wang R, Yu Z, Gao S (2023) Planedepth: self-supervised depth estimation via orthogonal planes. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp. 21425\u201321434","DOI":"10.1109\/CVPR52729.2023.02052"},{"key":"3332_CR10","doi-asserted-by":"crossref","unstructured":"Shotton J, Glocker B, Zach C, Izadi S, Criminisi A, Fitzgibbon A (2013) Scene coordinate regression forests for camera relocalization in RGB-D images. In: 2013 IEEE conference on computer vision and pattern recognition, pp 2930\u20132937","DOI":"10.1109\/CVPR.2013.377"},{"key":"3332_CR11","doi-asserted-by":"crossref","unstructured":"Menze M, Geiger A (2015) Object scene flow for autonomous vehicles. In: 2015 IEEE conference on computer vision and pattern recognition (CVPR), pp 3061\u20133070","DOI":"10.1109\/CVPR.2015.7298925"},{"key":"3332_CR12","doi-asserted-by":"crossref","unstructured":"Ye M, Johns E, Handa A, Zhang L, Pratt P, Yang G-Z (2017) Self-supervised siamese learning on stereo image pairs for depth estimation in robotic surgery. In: The Hamlyn symposium on medical robotics, p 27","DOI":"10.31256\/HSMR2017.14"},{"key":"3332_CR13","unstructured":"Allan M, Mcleod J, Wang C, Rosenthal JC, Hu Z, Gard N, Eisert P, Fu KX, Zeffiro T, Xia W et al (2021) Stereo correspondence and reconstruction of endoscopic data challenge. arXiv preprint arXiv:2101.01133"},{"issue":"4","key":"3332_CR14","doi-asserted-by":"publisher","first-page":"7225","DOI":"10.1109\/LRA.2021.3095528","volume":"6","author":"D Recasens","year":"2021","unstructured":"Recasens D, Lamarca J, F\u00e1cil JM, Montiel J, Civera J (2021) Endo-depth-and-motion: reconstruction and tracking in endoscopic videos using depth networks and photometric constraints. IEEE Robot Autom Lett 6(4):7225\u20137232","journal-title":"IEEE Robot Autom Lett"},{"issue":"2","key":"3332_CR15","doi-asserted-by":"publisher","first-page":"91","DOI":"10.1023\/B:VISI.0000029664.99615.94","volume":"60","author":"DG Lowe","year":"2004","unstructured":"Lowe DG (2004) Distinctive image features from scale-invariant keypoints. Int J Comput Vis 60(2):91\u2013110","journal-title":"Int J Comput Vis"},{"issue":"3","key":"3332_CR16","doi-asserted-by":"publisher","first-page":"274","DOI":"10.1080\/21681163.2021.2015723","volume":"10","author":"W Li","year":"2022","unstructured":"Li W, Hayashi Y, Oda M, Kitasaka T, Misawa K, Mori K (2022) Spatially variant biases considered self-supervised depth estimation based on laparoscopic videos. Comput Methods Biomech Biomed Eng Imaging Vis 10(3):274\u2013282","journal-title":"Comput Methods Biomech Biomed Eng Imaging Vis"},{"key":"3332_CR17","doi-asserted-by":"crossref","unstructured":"Huang B, Zheng J-Q, Nguyen A, Xu C, Gkouzionis I, Vyas K, Tuch D, Giannarou S, Elson DS (2022) Self-supervised depth estimation in laparoscopic image using 3D geometric consistency. In: Medical image computing and computer assisted intervention\u2013MICCAI 2022: 25th international conference, LNCS, vol 13437, pp 13\u201322","DOI":"10.1007\/978-3-031-16449-1_2"},{"key":"3332_CR18","doi-asserted-by":"crossref","unstructured":"Li W, Hayashi Y, Oda M, Kitasaka T, Misawa K, Mori K (2022) Geometric constraints for self-supervised monocular depth estimation on laparoscopic images with dual-task consistency. In: Medical image computing and computer assisted intervention-MICCAI 2022: 25th international conference, Singapore, Sept 18\u201322, 2022, Proceedings, LNCS, vol. 13434, pp. 467\u2013477","DOI":"10.1007\/978-3-031-16440-8_45"},{"key":"3332_CR19","doi-asserted-by":"crossref","unstructured":"Li W, Hayashi Y, Oda M, Kitasaka T, Misawa K, Mori K (2023) Multi-view guidance for self-supervised monocular depth estimation on laparoscopic images via spatio-temporal correspondence. In: Medical image computing and computer assisted intervention\u2014MICCAI 2023: 26th international conference, LNCS, vol 14228, pp 429\u2013439","DOI":"10.1007\/978-3-031-43996-4_41"},{"key":"3332_CR20","doi-asserted-by":"crossref","unstructured":"Murshed MS, Murphy C, Hou D, Khan N, Ananthanarayanan G, Hussain F (2021) Machine learning at the network edge: a survey. ACM Comput Surv (CSUR) 54(8):1\u201337","DOI":"10.1145\/3469029"},{"issue":"4","key":"3332_CR21","doi-asserted-by":"publisher","first-page":"600","DOI":"10.1109\/TIP.2003.819861","volume":"13","author":"Z Wang","year":"2004","unstructured":"Wang Z, Bovik AC, Sheikh HR, Simoncelli EP (2004) Image quality assessment: from error visibility to structural similarity. IEEE Trans Image Process 13(4):600\u2013612","journal-title":"IEEE Trans Image Process"},{"key":"3332_CR22","doi-asserted-by":"crossref","unstructured":"Wang X, Girshick R, Gupta A, He K (2018) Non-local neural networks. In: 2018 IEEE\/CVF conference on computer vision and pattern recognition, pp 7794\u20137803","DOI":"10.1109\/CVPR.2018.00813"},{"key":"3332_CR23","doi-asserted-by":"crossref","unstructured":"Shu C, Yu K, Duan Z, Yang K (2020) Feature-metric loss for self-supervised learning of depth and egomotion. In: Computer vision\u2014ECCV 2020, pp 572\u2013588","DOI":"10.1007\/978-3-030-58529-7_34"},{"key":"3332_CR24","doi-asserted-by":"crossref","unstructured":"Watson J, Mac\u00a0Aodha O, Prisacariu V, Brostow G, Firman M (2021) The temporal opportunist: self-supervised multi-frame monocular depth. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 1164\u20131174","DOI":"10.1109\/CVPR46437.2021.00122"},{"key":"3332_CR25","unstructured":"Zhou H, Greenwood D, Taylor S (2021) Self-supervised monocular depth estimation with internal feature fusion. In: British machine vision conference (BMVC)"},{"key":"3332_CR26","doi-asserted-by":"crossref","unstructured":"Zhao C, Zhang Y, Poggi M, Tosi F, Guo X, Zhu Z, Huang G, Tang Y, Mattoccia S (2022) MonoViT: self-supervised monocular depth estimation with a vision transformer. In: 2022 International conference on 3d vision (3DV), pp 668\u2013678","DOI":"10.1109\/3DV57658.2022.00077"},{"key":"3332_CR27","doi-asserted-by":"crossref","unstructured":"Han W, Yin J, Jin X, Dai X, Shen J (2022) BRNet: exploring comprehensive features for monocular depth estimation. In: Computer vision\u2014ECCV 2022, pp 586\u2013602","DOI":"10.1007\/978-3-031-19839-7_34"},{"key":"3332_CR28","doi-asserted-by":"publisher","DOI":"10.1016\/j.media.2021.102058","volume":"71","author":"KB Ozyoruk","year":"2021","unstructured":"Ozyoruk KB, Gokceler GI, Bobrow TL, Coskun G, Incetan K, Almalioglu Y, Mahmood F, Curto E, Perdigoto L, Oliveira M et al (2021) EndoSLAM dataset and an unsupervised monocular visual odometry and depth estimation approach for endoscopic videos. Med Image Anal 71:102058","journal-title":"Med Image Anal"},{"key":"3332_CR29","doi-asserted-by":"publisher","DOI":"10.1016\/j.media.2021.102338","volume":"77","author":"S Shao","year":"2022","unstructured":"Shao S, Pei Z, Chen W, Zhu W, Wu X, Sun D, Zhang B (2022) Self-supervised monocular depth and ego-motion estimation in endoscopy: appearance flow to the rescue. Med Image Anal 77:102338","journal-title":"Med Image Anal"},{"key":"3332_CR30","unstructured":"Paszke A, Gross S, Chintala S, Chanan G, Yang E, DeVito Z, Lin Z, Desmaison A, Antiga L, Lerer A (2017) Automatic differentiation in pytorch. In: NIPS 2017 workshop on autodiff"},{"key":"3332_CR31","unstructured":"Kingma DP, Ba J (2015) Adam: a method for stochastic optimization. In: 3rd International conference on learning representations, ICLR 2015, San Diego, CA, USA, May 7\u20139, 2015, Conference track proceedings"}],"container-title":["International Journal of Computer Assisted Radiology and Surgery"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11548-025-03332-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11548-025-03332-1\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11548-025-03332-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,4,27]],"date-time":"2025-04-27T10:03:08Z","timestamp":1745748188000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11548-025-03332-1"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,2,28]]},"references-count":31,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2025,4]]}},"alternative-id":["3332"],"URL":"https:\/\/doi.org\/10.1007\/s11548-025-03332-1","relation":{},"ISSN":["1861-6429"],"issn-type":[{"value":"1861-6429","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,2,28]]},"assertion":[{"value":"11 January 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"3 February 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"28 February 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors report there are no conflict of interest to declare.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}