{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,23]],"date-time":"2026-02-23T20:58:59Z","timestamp":1771880339860,"version":"3.50.1"},"reference-count":26,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2025,11,11]],"date-time":"2025-11-11T00:00:00Z","timestamp":1762819200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,11,11]],"date-time":"2025-11-11T00:00:00Z","timestamp":1762819200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100018693","name":"Horizon Europe programme","doi-asserted-by":"crossref","award":["101059903"],"award-info":[{"award-number":["101059903"]}],"id":[{"id":"10.13039\/100018693","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J CARS"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:sec>\n                    <jats:title>\n                      <jats:bold>Purpose<\/jats:bold>\n                    <\/jats:title>\n                    <jats:p>The recognition of surgical instrument-tissue interactions can enhance the surgical workflow analysis, improve automated safety systems and enable skill assessment in minimally invasive surgery. However, current deep learning methods for surgical instrument-tissue interaction recognition often rely on static images or coarse temporal sampling, limiting their ability to capture rapid surgical dynamics. Therefore, this study systematically investigates the impact of incorporating fine-grained temporal context into deep learning models for interaction recognition.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>\n                      <jats:bold>Methods<\/jats:bold>\n                    <\/jats:title>\n                    <jats:p>We conduct extensive experiments with multiple curated video-based datasets to investigate the influence of fine-grained temporal context for the task of instrument-tissue interaction recognition using video transformer with spatio-temporal feature extraction capabilities. Additionally, we propose a multi-task-attention module that utilizes cross-attention and a gating mechanism to improve communication between the subtasks of identifying the surgical instrument, atomic action, and anatomical target.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>\n                      <jats:bold>Results<\/jats:bold>\n                    <\/jats:title>\n                    <jats:p>\n                      Our study demonstrates the benefit of utilizing the fine-grained temporal context for recognition of instrument-tissue interactions, with an optimal sampling rate of 6-8 Hz identified for the examined datasets. Furthermore, our proposed MTAM significantly outperforms state-of-the-art multi-task video transformer on the CholecT45-Vid and GraSP-Vid datasets, achieving relative increases of\n                      <jats:inline-formula>\n                        <jats:alternatives>\n                          <jats:tex-math>$$4.8 \\%$$<\/jats:tex-math>\n                          <mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                            <mml:mrow>\n                              <mml:mn>4.8<\/mml:mn>\n                              <mml:mo>%<\/mml:mo>\n                            <\/mml:mrow>\n                          <\/mml:math>\n                        <\/jats:alternatives>\n                      <\/jats:inline-formula>\n                      and\n                      <jats:inline-formula>\n                        <jats:alternatives>\n                          <jats:tex-math>$$5.9 \\%$$<\/jats:tex-math>\n                          <mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                            <mml:mrow>\n                              <mml:mn>5.9<\/mml:mn>\n                              <mml:mo>%<\/mml:mo>\n                            <\/mml:mrow>\n                          <\/mml:math>\n                        <\/jats:alternatives>\n                      <\/jats:inline-formula>\n                      in surgical instrument-tissue interaction recognition, respectively.\n                    <\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>\n                      <jats:bold>Conclusions<\/jats:bold>\n                    <\/jats:title>\n                    <jats:p>\n                      In this work, we demonstrate the benefits of using a fine-grained temporal context rather than static images or coarse temporal context for the task of surgical instrument-tissue interaction recognition. We also show that leveraging cross-attention with spatio-temporal features from various subtasks leads to improved surgical instrument-tissue interaction recognition performance. The project is available at:\n                      <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/lennart-maack.github.io\/InstrTissRec-MTAM\" ext-link-type=\"uri\">https:\/\/lennart-maack.github.io\/InstrTissRec-MTAM<\/jats:ext-link>\n                      .\n                    <\/jats:p>\n                  <\/jats:sec>","DOI":"10.1007\/s11548-025-03546-3","type":"journal-article","created":{"date-parts":[[2025,11,11]],"date-time":"2025-11-11T01:37:48Z","timestamp":1762825068000},"page":"21-30","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Surgical instrument-tissue interaction recognition with multi-task-attention video transformer"],"prefix":"10.1007","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0009-0002-0097-0157","authenticated-orcid":false,"given":"Lennart","family":"Maack","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Berk","family":"Cam","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sarah","family":"Latus","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tobias","family":"Maurer","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Alexander","family":"Schlaefer","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,11,11]]},"reference":[{"key":"3546_CR1","doi-asserted-by":"publisher","DOI":"10.1016\/j.media.2025.103726","volume":"106","author":"N Ayobi","year":"2025","unstructured":"Ayobi N, Rodr\u00edguez S, P\u00e9rez A, Hern\u00e1ndez I, Aparicio N, Dessevres E, Pe\u00f1a S, Santander J, Caicedo JI, Fern\u00e1ndez N, Arbel\u00e1ez P (2025) Pixel-wise recognition for holistic surgical scene understanding. Med Image Anal 106:103726. https:\/\/doi.org\/10.1016\/j.media.2025.103726","journal-title":"Med Image Anal"},{"key":"3546_CR2","doi-asserted-by":"publisher","DOI":"10.1007\/s00464-022-09296-6","author":"S Balvardi","year":"2022","unstructured":"Balvardi S, Kammili A, Hanson M, Mueller C, Vassiliou M, Lee L, Schwartzman K, Fiore JF, Feldman LS (2022) The association between video-based assessment of intraoperative technical performance and patient outcomes: a systematic review. Surg Endosc. https:\/\/doi.org\/10.1007\/s00464-022-09296-6","journal-title":"Surg Endosc"},{"key":"3546_CR3","doi-asserted-by":"publisher","DOI":"10.1056\/nejmsa1300625","author":"JD Birkmeyer","year":"2013","unstructured":"Birkmeyer JD, Finks JF, O\u2019Reilly A, Oerline M, Carlin AM, Nunn AR, Dimick J, Banerjee M, Birkmeyer NJ (2013) Surgical skill and complication rates after bariatric surgery. N Engl J Med. https:\/\/doi.org\/10.1056\/nejmsa1300625","journal-title":"N Engl J Med"},{"key":"3546_CR4","volume-title":"Modeling and segmentation of surgical workflow from laparoscopic video in medical image computing and computer-assisted intervention","author":"T Blum","year":"2010","unstructured":"Blum T, Feu\u00dfner H, Navab N (2010) Modeling and segmentation of surgical workflow from laparoscopic video in medical image computing and computer-assisted intervention. Springer, Berlin"},{"key":"3546_CR5","volume-title":"Tail-enhanced representation learning for surgical triplet recognition in proceedings of medical image computing and computer assisted intervention","author":"S Gui","year":"2024","unstructured":"Gui S, Wang Z (2024) Tail-enhanced representation learning for surgical triplet recognition in proceedings of medical image computing and computer assisted intervention. Springer, Berlin"},{"issue":"4","key":"3546_CR6","doi-asserted-by":"publisher","first-page":"1628","DOI":"10.1109\/TMI.2023.3345736","volume":"43","author":"S Gui","year":"2024","unstructured":"Gui S, Wang Z, Chen J, Zhou X, Zhang C, Cao Y (2024) Mt4mtl-kd: a multi-teacher knowledge distillation framework for triplet recognition. IEEE Trans Med Imaging 43(4):1628\u20131639. https:\/\/doi.org\/10.1109\/TMI.2023.3345736","journal-title":"IEEE Trans Med Imaging"},{"key":"3546_CR7","doi-asserted-by":"crossref","unstructured":"He K, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition. In: proceedings of the IEEE conference on computer vision and pattern recognition, pp 770\u2013778","DOI":"10.1109\/CVPR.2016.90"},{"issue":"8","key":"3546_CR8","doi-asserted-by":"publisher","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","volume":"9","author":"S Hochreiter","year":"1997","unstructured":"Hochreiter S, Schmidhuber J (1997) Long short-term memory. Neural Comput 9(8):1735\u20131780","journal-title":"Neural Comput"},{"issue":"10","key":"3546_CR9","doi-asserted-by":"publisher","first-page":"2962","DOI":"10.1097\/JS9.0000000000000595","volume":"109","author":"FR Kolbinger","year":"2023","unstructured":"Kolbinger FR, Rinner FM, Jenke AC, Carstens M, Krell S, Leger S, Distler M, Weitz J, Speidel S, Bodenstedt S (2023) Anatomy segmentation in laparoscopic surgery: comparison of machine learning and human expertise-an experimental study. Int J Surg 109(10):2962\u20132974","journal-title":"Int J Surg"},{"key":"3546_CR10","doi-asserted-by":"crossref","unstructured":"Li Y, Wu CY, Fan H, Mangalam K, Xiong B, Malik J, Feichtenhofer C (2022) Mvitv2: Improved multiscale vision transformers for classification and detection. In: proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 4804\u20134814","DOI":"10.1109\/CVPR52688.2022.00476"},{"key":"3546_CR11","doi-asserted-by":"crossref","unstructured":"Liu Z, Lin Y, Cao Y, Hu H, Wei Y, Zhang Z, Lin S, Guo B (2021) Swin transformer: Hierarchical vision transformer using shifted windows. In: proceedings of the IEEE\/CVF international conference on computer vision, pp 10012\u201310022","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"3546_CR12","doi-asserted-by":"crossref","unstructured":"Liu Z, Ning J, Cao Y, Wei Y, Zhang Z, Lin S, Hu H (2022) Video swin transformer. In: proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 3202\u20133211","DOI":"10.1109\/CVPR52688.2022.00320"},{"key":"3546_CR13","unstructured":"Maack L, Behrendt F, Bhattacharya D, Latus S, Schlaefer A (2024) Efficient anatomy segmentation in laparoscopic surgery using multi-teacher knowledge distillation. In: medical imaging with deep learning, PMLR, pp 937\u2013948"},{"issue":"1","key":"3546_CR14","doi-asserted-by":"publisher","first-page":"163","DOI":"10.1038\/s41746-022-00707-5","volume":"5","author":"P Mascagni","year":"2022","unstructured":"Mascagni P, Alapatt D, Sestini L, Altieri MS, Madani A, Watanabe Y, Alseidi A, Redan JA, Alfieri S, Costamagna G, Bo\u0161koski I (2022) Computer vision in surgery from potential to clinical value. Npj Digital Med 5(1):163","journal-title":"Npj Digital Med"},{"key":"3546_CR15","volume-title":"Recognition of instrument-tissue interactions in endoscopic videos via action triplets in medical image computing and computer assisted intervention","author":"CI Nwoye","year":"2020","unstructured":"Nwoye CI, Gonzalez C, Yu T, Mascagni P, Mutter D, Marescaux J, Padoy N (2020) Recognition of instrument-tissue interactions in endoscopic videos via action triplets in medical image computing and computer assisted intervention. Springer, Berlin"},{"key":"3546_CR16","doi-asserted-by":"publisher","DOI":"10.1016\/j.media.2022.102433","volume":"78","author":"CI Nwoye","year":"2022","unstructured":"Nwoye CI, Yu T, Gonzalez C, Seeliger B, Mascagni P, Mutter D, Marescaux J, Padoy N (2022) Rendezvous: attention mechanisms for the recognition of surgical action triplets in endoscopic videos. Med Image Anal 78:102433. https:\/\/doi.org\/10.1016\/j.media.2022.102433","journal-title":"Med Image Anal"},{"key":"3546_CR17","doi-asserted-by":"publisher","unstructured":"Padoy N, Blum T, Ahmadi SA, Feussner H, Berger MO, Navab N (2012) Statistical modeling and recognition of surgical workflow. Med Image Anal 16(3):632\u2013641. https:\/\/doi.org\/10.1016\/j.media.2010.10.001","DOI":"10.1016\/j.media.2010.10.001"},{"key":"3546_CR18","doi-asserted-by":"publisher","DOI":"10.1109\/TMI.2025.3590457","author":"J Pei","year":"2025","unstructured":"Pei J, Zhang J, Qin G, Wang K, Jin Y, Heng PA (2025) Instrument-tissue-guided surgical action triplet detection via textual-temporal trail exploration. IEEE Trans Med Imaging. https:\/\/doi.org\/10.1109\/TMI.2025.3590457","journal-title":"IEEE Trans Med Imaging"},{"key":"3546_CR19","doi-asserted-by":"publisher","unstructured":"Richards MK, McAteer JP, Drake FT, Goldin AB, Khandelwal S, Gow KW (2015) A national review of the frequency of minimally invasive surgery among general surgery residents: assessment of acgme case logs during 2 decades of general surgery resident training. JAMA Surg 150:169\u2013172. https:\/\/doi.org\/10.1001\/jamasurg.2014.1791","DOI":"10.1001\/jamasurg.2014.1791"},{"key":"3546_CR20","doi-asserted-by":"publisher","DOI":"10.1016\/j.media.2020.101920","volume":"70","author":"T Ro\u00df","year":"2021","unstructured":"Ro\u00df T, Reinke A, Full PM, Wagner M, Kenngott H, Apitz M, Hempe H, Mindroc-Filimon D, Scholz P, Tran TN et al (2021) Comparative validation of multi-instance instrument segmentation in endoscopy: results of the robust-mis 2019 challenge. Med Image Anal 70:101920","journal-title":"Med Image Anal"},{"key":"3546_CR21","doi-asserted-by":"publisher","unstructured":"Sharma S, Nwoye CI, Mutter D, Padoy N (2023) Rendezvous in time: an attention-based temporal fusion approach for surgical triplet recognition. Int J Comput Assist Radiol Surg 18(6):1053\u20131059. https:\/\/doi.org\/10.1007\/s11548-023-02914-1","DOI":"10.1007\/s11548-023-02914-1"},{"issue":"7","key":"3546_CR22","doi-asserted-by":"publisher","first-page":"1267","DOI":"10.1007\/s11548-024-03163-6","volume":"19","author":"Y Sheng","year":"2024","unstructured":"Sheng Y, Bano S, Clarkson MJ, Islam M (2024) Surgical-desam: decoupling sam for instrument segmentation in robotic surgery. Int J Comput Assist Radiol Surg 19(7):1267\u20131271","journal-title":"Int J Comput Assist Radiol Surg"},{"key":"3546_CR23","doi-asserted-by":"publisher","first-page":"86","DOI":"10.1109\/TMI.2016.2593957","volume":"36","author":"AP Twinanda","year":"2017","unstructured":"Twinanda AP, Shehata S, Mutter D, Marescaux J, de Mathelin M, Padoy N (2017) Endonet: a deep architecture for recognition tasks on laparoscopic videos. IEEE Trans Med Imaging 36:86\u201397. https:\/\/doi.org\/10.1109\/TMI.2016.2593957","journal-title":"IEEE Trans Med Imaging"},{"issue":"7","key":"3546_CR24","doi-asserted-by":"publisher","first-page":"1677","DOI":"10.1109\/TMI.2022.3147640","volume":"41","author":"B Van Amsterdam","year":"2022","unstructured":"Van Amsterdam B, Funke I, Edwards E, Speidel S, Collins J, Sridhar A, Kelly J, Clarkson MJ, Stoyanov D (2022) Gesture recognition in robotic surgery with multimodal attention. IEEE Trans Med Imaging 41(7):1677\u20131687","journal-title":"IEEE Trans Med Imaging"},{"key":"3546_CR25","unstructured":"Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser \u0141, Polosukhin I (2017) Attention is all you need. Advances in neural information processing systems 30"},{"key":"3546_CR26","doi-asserted-by":"crossref","unstructured":"Yamlahi A, Tran TN, Godau P, Schellenberg M, Michael D, Smidt FH, N\u00f6lke JH, Adler TJ, Tizabi MD, Nwoye CI (2023) Self-distillation for surgical action recognition. In: international conference on medical image computing and computer-assisted intervention, Springer, pp 637\u2013646","DOI":"10.1007\/978-3-031-43996-4_61"}],"container-title":["International Journal of Computer Assisted Radiology and Surgery"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11548-025-03546-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11548-025-03546-3","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11548-025-03546-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,23]],"date-time":"2026-02-23T20:03:52Z","timestamp":1771877032000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11548-025-03546-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,11]]},"references-count":26,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2026,1]]}},"alternative-id":["3546"],"URL":"https:\/\/doi.org\/10.1007\/s11548-025-03546-3","relation":{},"ISSN":["1861-6429"],"issn-type":[{"value":"1861-6429","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,11,11]]},"assertion":[{"value":"2 July 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"30 October 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"11 November 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Lennart Maack, Berk Cam, Sarah Latus, Tobias Maurer and Alexander Schlaefer declare that they have no Conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"This study exclusively utilized publicly available datasets.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval"}},{"value":"This study exclusively utilized publicly available datasets.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Informed consent"}}]}}