{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,2,21]],"date-time":"2025-02-21T13:25:24Z","timestamp":1740144324317,"version":"3.37.3"},"reference-count":27,"publisher":"Springer Science and Business Media LLC","issue":"12","license":[{"start":{"date-parts":[[2022,7,28]],"date-time":"2022-07-28T00:00:00Z","timestamp":1658966400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,7,28]],"date-time":"2022-07-28T00:00:00Z","timestamp":1658966400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100005713","name":"Technische Universit\u00e4t M\u00fcnchen","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100005713","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J CARS"],"abstract":"<jats:title>Abstract<\/jats:title><jats:sec>\n                <jats:title>Purpose:<\/jats:title>\n                <jats:p>Automatic surgical instruction generation is a crucial part for intra-operative surgical assistance. However, understanding and translating surgical activities into human-like sentences are particularly challenging due to the complexity of surgical environment and the modal gap between images and natural languages. To this end, we introduce <jats:bold>SIG-Former<\/jats:bold>, a transformer-backboned generation network to predict surgical instructions from monocular RGB images.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Methods:<\/jats:title>\n                <jats:p>Taking a surgical image as input, we first extract its visual attentive feature map with a fine-tuned ResNet-101 model, followed by transformer attention blocks to correspondingly model its visual representation, text embedding and visual\u2013textual relational feature. To tackle the loss-metric inconsistency between training and inference in sequence generation, we additionally apply a self-critical reinforcement learning approach to directly optimize the CIDEr score after regular training.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Results:<\/jats:title>\n                <jats:p>We validate our proposed method on DAISI dataset, which contains 290 clinical procedures from diverse medical subjects. Extensive experiments demonstrate that our method outperforms the baselines and achieves promising performance on both quantitative and qualitative evaluations.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Conclusion:<\/jats:title>\n                <jats:p>Our experiments demonstrate that SIG-Former is capable of mapping dependencies between visual feature and textual information. Besides, surgical instruction generation is still at its preliminary stage. Future works include collecting large clinical dataset, annotating more reference instructions and preparing pre-trained models on medical images.<\/jats:p>\n              <\/jats:sec>","DOI":"10.1007\/s11548-022-02718-9","type":"journal-article","created":{"date-parts":[[2022,7,28]],"date-time":"2022-07-28T13:17:10Z","timestamp":1659014230000},"page":"2203-2210","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["SIG-Former: monocular surgical instruction generation with transformers"],"prefix":"10.1007","volume":"17","author":[{"given":"Jinglu","family":"Zhang","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7023-6797","authenticated-orcid":false,"given":"Yinyu","family":"Nie","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jian","family":"Chang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jian Jun","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2022,7,28]]},"reference":[{"key":"2718_CR1","doi-asserted-by":"crossref","unstructured":"Anderson P, Fernando B, Johnson M, Gould S (2016) Spice: semantic propositional image caption evaluation. In: European conference on computer vision, pp 382\u2013398. Springer","DOI":"10.1007\/978-3-319-46454-1_24"},{"issue":"8","key":"2718_CR2","doi-asserted-by":"publisher","first-page":"2111","DOI":"10.1007\/s00464-012-2175-x","volume":"26","author":"SA Antoniou","year":"2012","unstructured":"Antoniou SA, Antoniou GA, Franzen J, Bollmann S, Koch OO, Pointner R, Granderath FA (2012) A comprehensive review of telementoring applications in laparoscopic general surgery. Surg Endosc 26(8):2111\u20132116","journal-title":"Surg Endosc"},{"key":"2718_CR3","unstructured":"Banerjee S, Lavie A (2005) Meteor: an automatic metric for MT evaluation with improved correlation with human judgments. In: Proceedings of the ACL workshop on intrinsic and extrinsic evaluation measures for machine translation and\/or summarization, pp 65\u201372"},{"issue":"4","key":"2718_CR4","doi-asserted-by":"publisher","first-page":"379","DOI":"10.1177\/1553350617708725","volume":"24","author":"E Bilgic","year":"2017","unstructured":"Bilgic E, Turkdogan S, Watanabe Y, Madani A, Landry T, Lavigne D, Feldman LS, Vassiliou MC (2017) Effectiveness of telementoring in surgery compared with on-site mentoring: a systematic review. Surg Innov 24(4):379\u2013385","journal-title":"Surg Innov"},{"key":"2718_CR5","doi-asserted-by":"publisher","DOI":"10.1016\/j.media.2020.101797","volume":"66","author":"A Bustos","year":"2020","unstructured":"Bustos A, Pertusa A, Salinas JM, de la Iglesia-Vay\u00e1 M (2020) Padchest: a large chest X-ray image dataset with multi-label annotated reports. Med Image Anal 66:101797","journal-title":"Med Image Anal"},{"key":"2718_CR6","doi-asserted-by":"crossref","unstructured":"Chen Z, Song Y, Chang TH, Wan X (2020) Generating radiology reports via memory-driven transformer. In: Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP), pp 1439\u20131449","DOI":"10.18653\/v1\/2020.emnlp-main.112"},{"key":"2718_CR7","doi-asserted-by":"crossref","unstructured":"Cornia M, Stefanini M, Baraldi L, Cucchiara R (2020) Meshed-memory transformer for image captioning. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 10578\u201310587","DOI":"10.1109\/CVPR42600.2020.01059"},{"key":"2718_CR8","doi-asserted-by":"crossref","unstructured":"Czempiel T, Paschali M, Keicher M, Simson W, Feussner H, Kim ST, Navab N (2020) Tecno: surgical phase recognition with multi-stage temporal convolutional networks. In: International conference on medical image computing and computer-assisted intervention, pp 343\u2013352. Springer","DOI":"10.1007\/978-3-030-59716-0_33"},{"key":"2718_CR9","doi-asserted-by":"crossref","unstructured":"Deng J, Dong W, Socher R, Li LJ, Li K, Fei-Fei L (2009) Imagenet: a large-scale hierarchical image database. In: 2009 IEEE conference on computer vision and pattern recognition, pp 248\u2013255. IEEE","DOI":"10.1109\/CVPR.2009.5206848"},{"issue":"1","key":"2718_CR10","doi-asserted-by":"publisher","first-page":"95","DOI":"10.1177\/1553350618813250","volume":"26","author":"S Erridge","year":"2019","unstructured":"Erridge S, Yeung DK, Patel HR, Purkayastha S (2019) Telementoring of surgeons: a systematic review. Surg Innov 26(1):95\u2013111","journal-title":"Surg Innov"},{"key":"2718_CR11","doi-asserted-by":"crossref","unstructured":"Funke I, Bodenstedt S, Oehme F, Bechtolsheim Fv, Weitz J, Speidel S (2019) Using 3D convolutional neural networks to learn spatiotemporal features for automatic surgical gesture recognition in video. In: International conference on medical image computing and computer-assisted intervention, pp 467\u2013475. Springer","DOI":"10.1007\/978-3-030-32254-0_52"},{"key":"2718_CR12","doi-asserted-by":"crossref","unstructured":"He K, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 770\u2013778","DOI":"10.1109\/CVPR.2016.90"},{"key":"2718_CR13","doi-asserted-by":"crossref","unstructured":"Jing B, Xie P, Xing E (2018) On the automatic generation of medical imaging reports. In: Proceedings of the 56th annual meeting of the association for computational linguistics (volume 1: long papers), pp 2577\u20132586","DOI":"10.18653\/v1\/P18-1240"},{"key":"2718_CR14","unstructured":"Lin CY (2004) Rouge: A package for automatic evaluation of summaries. In: Text summarization branches out, pp 74\u201381"},{"key":"2718_CR15","doi-asserted-by":"crossref","unstructured":"Lin TY, Maire M, Belongie S, Hays J, Perona P, Ramanan D, Doll\u00e1r P, Zitnick CL (2014) Microsoft coco: common objects in context. In: European conference on computer vision, pp 740\u2013755. Springer","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"2718_CR16","doi-asserted-by":"crossref","unstructured":"Papineni K, Roukos S, Ward T, Zhu WJ (2002) Bleu: a method for automatic evaluation of machine translation. In: Proceedings of the 40th annual meeting of the association for computational linguistics, pp 311\u2013318","DOI":"10.3115\/1073083.1073135"},{"key":"2718_CR17","unstructured":"Pascanu R, Mikolov T, Bengio Y (2013) On the difficulty of training recurrent neural networks. In: International conference on machine learning, pp 1310\u20131318"},{"key":"2718_CR18","unstructured":"Ranzato M, Chopra S, Auli M, Zaremba W (2016) Sequence level training with recurrent neural networks. In: 4th international conference on learning representations, ICLR 2016"},{"key":"2718_CR19","doi-asserted-by":"crossref","unstructured":"Rennie SJ, Marcheret E, Mroueh Y, Ross J, Goel V (2017) Self-critical sequence training for image captioning. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 7008\u20137024","DOI":"10.1109\/CVPR.2017.131"},{"key":"2718_CR20","unstructured":"Rojas-Mu\u00f1oz E, Couperus K, Wachs J (2020) DAISI: database for AI surgical instruction. arXiv preprint arXiv:2004.02809"},{"key":"2718_CR21","unstructured":"Simonyan K, Zisserman A (2015) Very deep convolutional networks for large-scale image recognition"},{"issue":"1","key":"2718_CR22","doi-asserted-by":"publisher","first-page":"86","DOI":"10.1109\/TMI.2016.2593957","volume":"36","author":"AP Twinanda","year":"2016","unstructured":"Twinanda AP, Shehata S, Mutter D, Marescaux J, De Mathelin M, Padoy N (2016) Endonet: a deep architecture for recognition tasks on laparoscopic videos. IEEE Trans Med Imaging 36(1):86\u201397","journal-title":"IEEE Trans Med Imaging"},{"key":"2718_CR23","unstructured":"Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser \u0141, Polosukhin I (2017) Attention is all you need. In: Advances in neural information processing systems, vol 30"},{"key":"2718_CR24","doi-asserted-by":"crossref","unstructured":"Vedantam R, Lawrence\u00a0Zitnick C, Parikh D (2015) CIDEr: consensus-based image description evaluation. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 4566\u20134575","DOI":"10.1109\/CVPR.2015.7299087"},{"key":"2718_CR25","doi-asserted-by":"crossref","unstructured":"Vinyals O, Toshev A, Bengio S, Erhan D (2015) Show and tell: a neural image caption generator. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 3156\u20133164","DOI":"10.1109\/CVPR.2015.7298935"},{"key":"2718_CR26","unstructured":"Xu K, Ba J, Kiros R, Cho K, Courville A, Salakhudinov R, Zemel R, Bengio Y (2015) Show, attend and tell: neural image caption generation with visual attention. In: International conference on machine learning, pp 2048\u20132057"},{"key":"2718_CR27","doi-asserted-by":"crossref","unstructured":"Zhang J, Nie Y, Chang J, Zhang JJ (2021) Surgical instruction generation with transformers. In: International conference on medical image computing and computer-assisted intervention, pp 290\u2013299. Springer","DOI":"10.1007\/978-3-030-87202-1_28"}],"container-title":["International Journal of Computer Assisted Radiology and Surgery"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11548-022-02718-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11548-022-02718-9\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11548-022-02718-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,11,11]],"date-time":"2022-11-11T14:15:16Z","timestamp":1668176116000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11548-022-02718-9"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,7,28]]},"references-count":27,"journal-issue":{"issue":"12","published-online":{"date-parts":[[2022,12]]}},"alternative-id":["2718"],"URL":"https:\/\/doi.org\/10.1007\/s11548-022-02718-9","relation":{},"ISSN":["1861-6429"],"issn-type":[{"type":"electronic","value":"1861-6429"}],"subject":[],"published":{"date-parts":[[2022,7,28]]},"assertion":[{"value":"1 April 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"4 July 2022","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"28 July 2022","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"This article does not contain any studies with human participants or animals performed by any of the authors.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical approval"}},{"value":"This article does not contain patient data.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Informed consent"}}]}}