{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,4]],"date-time":"2025-12-04T14:46:59Z","timestamp":1764859619924,"version":"3.37.3"},"reference-count":22,"publisher":"Springer Science and Business Media LLC","issue":"11","license":[{"start":{"date-parts":[[2024,2,27]],"date-time":"2024-02-27T00:00:00Z","timestamp":1708992000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,2,27]],"date-time":"2024-02-27T00:00:00Z","timestamp":1708992000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J CARS"],"abstract":"<jats:title>Abstract<\/jats:title><jats:sec>\n                <jats:title>\n                           <jats:bold>Purpose<\/jats:bold>\n                        <\/jats:title>\n                <jats:p>Analysis of operative fields is expected to aid in estimating procedural workflow and evaluating surgeons\u2019 procedural skills by considering the temporal transitions during the progression of the surgery. This study aims to propose an automatic recognition system for the procedural workflow by employing machine learning techniques to identify and distinguish elements in the operative field, including body tissues such as fat, muscle, and dermis, along with surgical tools.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>\n                           <jats:bold>Methods<\/jats:bold>\n                        <\/jats:title>\n                <jats:p>We conducted annotations on approximately 908 first-person-view images of breast surgery to facilitate segmentation. The annotated images were used to train a pixel-level classifier based on Mask R-CNN. To assess the impact on procedural workflow recognition, we annotated an additional 43,007 images. The network, structured on the Transformer architecture, was then trained with surgical images incorporating masks for body tissues and surgical tools.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>\n                           <jats:bold>Results<\/jats:bold>\n                        <\/jats:title>\n                <jats:p>The instance segmentation of each body tissue in the segmentation phase provided insights into the trend of area transitions for each tissue. Simultaneously, the spatial features of the surgical tools were effectively captured. In regard to the accuracy of procedural workflow recognition, accounting for body tissues led to an average improvement of 3 % over the baseline. Furthermore, the inclusion of surgical tools yielded an additional increase in accuracy by 4 % compared to the baseline.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>\n                           <jats:bold>Conclusion<\/jats:bold>\n                        <\/jats:title>\n                <jats:p>In this study, we revealed the contribution of the temporal transition of the body tissues and surgical tools spatial features to recognize procedural workflow in first-person-view surgical videos. Body tissues, especially in open surgery, can be a crucial element. This study suggests that further improvements can be achieved by accurately identifying surgical tools specific to each procedural workflow step.<\/jats:p>\n              <\/jats:sec>","DOI":"10.1007\/s11548-024-03074-6","type":"journal-article","created":{"date-parts":[[2024,2,27]],"date-time":"2024-02-27T13:02:24Z","timestamp":1709038944000},"page":"2195-2202","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["An analysis on the effect of body tissues and surgical tools on workflow recognition in first person surgical videos"],"prefix":"10.1007","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-7955-4148","authenticated-orcid":false,"given":"Hisako","family":"Tomita","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Naoto","family":"Ienaga","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hiroki","family":"Kajita","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tetsu","family":"Hayashida","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Maki","family":"Sugimoto","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2024,2,27]]},"reference":[{"key":"3074_CR1","doi-asserted-by":"publisher","first-page":"27","DOI":"10.1186\/1757-7241-21-27","volume":"21","author":"S Matsumoto","year":"2013","unstructured":"Matsumoto S, Sekine K, Yamazaki M, Funabiki T, Orita T, Shimizu M, Kitano M (2013) Digital video recording in trauma surgery using commercially available equipment. Scand J Trauma Resusc Emerg Med 21:27\u201331","journal-title":"Scand J Trauma Resusc Emerg Med"},{"key":"3074_CR2","doi-asserted-by":"publisher","first-page":"599","DOI":"10.1177\/1553350619853099","volume":"26","author":"TJ Saun","year":"2019","unstructured":"Saun TJ, Zuo KJ, Grantcharov TP (2019) Video technologies for recording open surgery: a systematic review. Surg Innov 26:599\u2013612","journal-title":"Surg Innov"},{"key":"3074_CR3","doi-asserted-by":"crossref","unstructured":"Avellino I, Nozari S, Canlorbe G, Jansen Y (2021) Surgical video summarization: Multifarious uses, summarization process and ad-hoc coordination, Proceedings of the ACM on Human-Computer Interaction, 5(CSCW1), pp 1-23","DOI":"10.1145\/3449214"},{"issue":"3","key":"3074_CR4","doi-asserted-by":"publisher","first-page":"201664","DOI":"10.1001\/jamanetworkopen.2020.1664","volume":"3","author":"S Khalid","year":"2020","unstructured":"Khalid S, Goldenberg M, Grantcharov T, Taati B, Rudzicz F (2020) Evaluation of deep learning models for identifying surgical actions and measuring performance. JAMA Netw Open 3(3):201664\u2013201664","journal-title":"JAMA Netw Open"},{"key":"3074_CR5","doi-asserted-by":"crossref","unstructured":"Qadir HA, Shin Y, Solhusvik J, Bergsland J, Aabakken L, Balasingham Ix(2019) Polyp detection and segmentation using mask r-cnn: Does a deeper feature extractor cnn always perform better? In: 2019 13th International Symposium on Medical Information and Communication Technology (ISMICT), pp. 1\u20136","DOI":"10.1109\/ISMICT.2019.8743694"},{"key":"3074_CR6","doi-asserted-by":"crossref","unstructured":"He K, Gkioxari G, Dollar P, Girshick R (2017) Mask r-cnn. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV)","DOI":"10.1109\/ICCV.2017.322"},{"key":"3074_CR7","doi-asserted-by":"crossref","unstructured":"Abibouraguimane I, Hagihara K, Higuchi K, Itoh Y, Sato Y, Hayashida T, Sugimoto M (2019) Cosummary: Adaptive fast-forwarding for surgical videos by detecting collaborative scenes using hand regions and gaze positions. In: Proceedings of the 24th International Conference on Intelligent User Interfaces. IUI \u201919, pp 580\u2013590","DOI":"10.1145\/3301275.3302284"},{"issue":"1","key":"3074_CR8","doi-asserted-by":"publisher","first-page":"2141001","DOI":"10.1142\/S2424905X21410014","volume":"7","author":"K Yoshida","year":"2022","unstructured":"Yoshida K, Hachiuma R, Tomita H, Pan J, Kitani K, Kajita H, Hayashida T, Sugimoto M (2022) Spatiotemporal video highlight by neural network considering gaze and hands of surgeon in egocentric surgical videos. J Med Robot Res 7(1):2141001","journal-title":"J Med Robot Res"},{"key":"3074_CR9","doi-asserted-by":"publisher","first-page":"437","DOI":"10.1007\/s11548-022-02559-6","volume":"17","author":"A Goldbraikh","year":"2022","unstructured":"Goldbraikh A, D\u2019Angelo A-L, Pugh CM, Laufer S (2022) Video-based fully automatic assessment of open surgery suturing skills. Int J Comput Assist Radiol Surg 17:437\u2013448","journal-title":"Int J Comput Assist Radiol Surg"},{"key":"3074_CR10","doi-asserted-by":"crossref","unstructured":"Hochreiter JS (1997) Long short-term memory. Neural Comput 9:1735\u20131780","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"3074_CR11","doi-asserted-by":"publisher","first-page":"685","DOI":"10.1007\/s11548-018-1882-8","volume":"14","author":"H Nakawala","year":"2019","unstructured":"Nakawala H, Bianchi R, Pescatori LE, De Cobelli O, Ferrigno G, De Momi E (2019) \u201cDeep-Onto\" network for surgical workflow and context recognition. Int J Comput Assist Radiol Surg 14:685\u2013696","journal-title":"Int J Comput Assist Radiol Surg"},{"key":"3074_CR12","doi-asserted-by":"publisher","unstructured":"Nakawala H (2017) Nephrec9. https:\/\/doi.org\/10.5281\/zenodo.1066831","DOI":"10.5281\/zenodo.1066831"},{"key":"3074_CR13","doi-asserted-by":"crossref","unstructured":"Liu Z, Lin Y, Cao Y, Hu H, Wei Y, Zhang Zg, Lin S, Guo B (2021) Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), pp 10012\u201310022","DOI":"10.1109\/ICCV48922.2021.00986"},{"issue":"1","key":"3074_CR14","doi-asserted-by":"publisher","first-page":"86","DOI":"10.1109\/TMI.2016.2593957","volume":"36","author":"AP Twinanda","year":"2017","unstructured":"Twinanda AP, Shehata S, Mutter D, Marescaux J, de Mathelin M, Padoy N (2017) EndoNet: a deep architecture for recognition tasks on laparoscopic videos. IEEE Trans Med Imag 36(1):86\u201397","journal-title":"IEEE Trans Med Imag"},{"key":"3074_CR15","unstructured":"Tobii pro Glasses. https:\/\/www.tobii.com\/"},{"key":"3074_CR16","unstructured":"VGG Image Annotator (VIA). https:\/\/www.robots.ox.ac.uk\/~vgg\/software\/via\/"},{"key":"3074_CR17","unstructured":"Max Planck Institute\u00a0for Psycholinguistics, T.L.A.: ELAN. https:\/\/archive.mpi.nl\/tla\/elan"},{"key":"3074_CR18","doi-asserted-by":"crossref","unstructured":"Deng J, Dong W, Socher R, Li L, Li K, Fei-Fei L (2009) Imagenet : a large-scale hierarchical image database. In: Proceedings IEEE Conference Computer Vision and Pattern Recognition, 2009","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"3074_CR19","unstructured":"Allan M, Shvets A, Kurmann T, Zhang Z, Duggal R, Su YH, Rieke N, Laina I, Kalavakonda N, Bodenstedt S, Herrera L (2019) 2017 robotic instrument segmentation challenge"},{"key":"3074_CR20","doi-asserted-by":"crossref","unstructured":"Gao X, Jin Y, Long Y, Dou Q, Heng P-A (2021) Trans-svnet: accurate phase recognition from surgical videos via hybrid embedding aggregation transformer. In: Medical Image Computing and Computer Assisted Intervention \u2013 MICCAI 2021, pp. 593\u2013603","DOI":"10.1007\/978-3-030-87202-1_57"},{"key":"3074_CR21","doi-asserted-by":"crossref","unstructured":"He K, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","DOI":"10.1109\/CVPR.2016.90"},{"key":"3074_CR22","doi-asserted-by":"publisher","first-page":"11","DOI":"10.1016\/j.jss.2018.09.015","volume":"235","author":"JL Green","year":"2019","unstructured":"Green JL, Suresh V, Bittar P, Ledbetter L, Mithani SK, Allori A (2019) The utilization of video technology in surgical education: a systematic review. J Surg Res 235:11\u2013180","journal-title":"J Surg Res"}],"container-title":["International Journal of Computer Assisted Radiology and Surgery"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11548-024-03074-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11548-024-03074-6\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11548-024-03074-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,11,6]],"date-time":"2024-11-06T16:23:19Z","timestamp":1730910199000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11548-024-03074-6"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,2,27]]},"references-count":22,"journal-issue":{"issue":"11","published-online":{"date-parts":[[2024,11]]}},"alternative-id":["3074"],"URL":"https:\/\/doi.org\/10.1007\/s11548-024-03074-6","relation":{},"ISSN":["1861-6429"],"issn-type":[{"type":"electronic","value":"1861-6429"}],"subject":[],"published":{"date-parts":[[2024,2,27]]},"assertion":[{"value":"15 March 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"9 February 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"27 February 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}