{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,28]],"date-time":"2026-07-28T01:04:44Z","timestamp":1785200684988,"version":"3.55.0"},"reference-count":30,"publisher":"Springer Science and Business Media LLC","issue":"7","license":[{"start":{"date-parts":[[2021,5,19]],"date-time":"2021-05-19T00:00:00Z","timestamp":1621382400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,5,19]],"date-time":"2021-05-19T00:00:00Z","timestamp":1621382400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100010665","name":"H2020 Marie Sklodowska-Curie Actions","doi-asserted-by":"publisher","award":["813782"],"award-info":[{"award-number":["813782"]}],"id":[{"id":"10.13039\/100010665","id-type":"DOI","asserted-by":"publisher"}]},{"name":"BPI France"},{"name":"Agence nationale de la recherche","award":["ANR-16-CE33-0009"],"award-info":[{"award-number":["ANR-16-CE33-0009"]}]},{"name":"Agence nationale de la recherche","award":["ANR-10-IAHU-02"],"award-info":[{"award-number":["ANR-10-IAHU-02"]}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J CARS"],"published-print":{"date-parts":[[2021,7]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:sec>\n                <jats:title>Purpose<\/jats:title>\n                <jats:p>Automatic segmentation and classification of surgical activity is crucial for providing advanced support in computer-assisted interventions and autonomous functionalities in robot-assisted surgeries. Prior works have focused on recognizing either coarse activities, such as phases, or fine-grained activities, such as gestures. This work aims at jointly recognizing two complementary levels of granularity directly from videos, namely phases and steps.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Methods<\/jats:title>\n                <jats:p>We introduce two correlated surgical activities, phases and steps, for the laparoscopic gastric bypass procedure. We propose a multi-task multi-stage temporal convolutional network (MTMS-TCN) along with a multi-task convolutional neural network (CNN) training setup to jointly predict the phases and steps and benefit from their complementarity to better evaluate the execution of the procedure. We evaluate the proposed method on a large video dataset consisting of 40 surgical procedures (<jats:italic>Bypass40<\/jats:italic>).<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Results<\/jats:title>\n                <jats:p>We present experimental results from several baseline models for both phase and step recognition on the <jats:italic>Bypass40<\/jats:italic>. The proposed MTMS-TCN method outperforms single-task methods in both phase and step recognition by 1-2% in accuracy, precision and recall. Furthermore, for step recognition, MTMS-TCN achieves a superior performance of 3-6% compared to LSTM-based models on all metrics.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Conclusion<\/jats:title>\n                <jats:p>In this work, we present a multi-task multi-stage temporal convolutional network for surgical activity recognition, which shows improved results compared to single-task models on a gastric bypass dataset with multi-level annotations. The proposed method shows that the joint modeling of phases and steps is beneficial to improve the overall recognition of each type of activity.<\/jats:p>\n              <\/jats:sec>","DOI":"10.1007\/s11548-021-02388-z","type":"journal-article","created":{"date-parts":[[2021,5,20]],"date-time":"2021-05-20T07:08:20Z","timestamp":1621494500000},"page":"1111-1119","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":84,"title":["Multi-task temporal convolutional networks for joint recognition of surgical phases and steps in gastric bypass procedures"],"prefix":"10.1007","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3452-4764","authenticated-orcid":false,"given":"Sanat","family":"Ramesh","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Diego","family":"Dall\u2019Alba","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Cristians","family":"Gonzalez","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tong","family":"Yu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Pietro","family":"Mascagni","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Didier","family":"Mutter","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jacques","family":"Marescaux","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Paolo","family":"Fiorini","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Nicolas","family":"Padoy","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2021,5,19]]},"reference":[{"key":"2388_CR1","unstructured":"Obesity: preventing and managing the global epidemic. Report of a WHO consultation. World Health Organ Tech Rep Ser 894, 1\u2013253 (2000)"},{"issue":"9","key":"2388_CR2","doi-asserted-by":"publisher","first-page":"2025","DOI":"10.1109\/TBME.2016.2647680","volume":"64","author":"N Ahmidi","year":"2017","unstructured":"Ahmidi N, Tao L, Sefati S, Gao Y, Lea C, Haro BB, Zappella L, Khudanpur S, Vidal R, Hager GD (2017) A dataset and benchmarks for segmentation and recognition of gestures in robotic surgery. IEEE Trans Biomed Eng 64(9):2025\u20132041","journal-title":"IEEE Trans Biomed Eng"},{"key":"2388_CR3","doi-asserted-by":"crossref","unstructured":"Angrisani L, Santonicola A, Iovino P, Formisano G, Buchwald H, Scopinaro N (2015) Bariatric surgery worldwide 2013. Obes Surg 25(10):1822\u20131832","DOI":"10.1007\/s11695-015-1657-z"},{"key":"2388_CR4","doi-asserted-by":"publisher","unstructured":"Birkmeyer JD, Finks JF, OReilly A, Oerline M, Carlin AM, Nunn AR, Dimick J, Banerjee M, Birkmeyer NJ, (2013) Surgical skill and complication rates after bariatric surgery. New Engl J Med 369(15):1434\u20131442. https:\/\/doi.org\/10.1056\/nejmsa1300625","DOI":"10.1056\/nejmsa1300625"},{"key":"2388_CR5","doi-asserted-by":"crossref","unstructured":"Bricon-Souf N, Newman CR (2007) Context awareness in health care: A review. Int J Med Inf 76(1):2\u201312","DOI":"10.1016\/j.ijmedinf.2006.01.003"},{"issue":"5","key":"2388_CR6","doi-asserted-by":"publisher","first-page":"495","DOI":"10.1089\/lap.2005.15.495","volume":"15","author":"K Cleary","year":"2005","unstructured":"Cleary K, Kinsella A (2005) OR 2020: The operating room of the future - workshop report. J Laparoendosc Adv Surg Tech - Part A 15(5):495\u2013573","journal-title":"J Laparoendosc Adv Surg Tech - Part A"},{"key":"2388_CR7","doi-asserted-by":"crossref","unstructured":"Czempiel T, Paschali M, Keicher M, Simson W, Feussner H, Kim ST, Navab N (2020) Tecno: Surgical phase recognition with multi-stage temporal convolutional networks. In: MICCAI","DOI":"10.1007\/978-3-030-59716-0_33"},{"key":"2388_CR8","doi-asserted-by":"publisher","unstructured":"Eigen D, Fergus R (2015) Predicting depth, surface normals and semantic labels with a common multi-scale convolutional architecture. In: 2015 IEEE International Conference on Computer Vision (ICCV), pp. 2650\u20132658. https:\/\/doi.org\/10.1109\/ICCV.2015.304","DOI":"10.1109\/ICCV.2015.304"},{"key":"2388_CR9","doi-asserted-by":"crossref","unstructured":"Farha YA, Gall J (2019) MS-TCN: Multi-stage temporal convolutional network for action segmentation. In: CVPR","DOI":"10.1109\/CVPR.2019.00369"},{"key":"2388_CR10","doi-asserted-by":"crossref","unstructured":"Funke I, Bodenstedt S, Oehme F, von Bechtolsheim F, Weitz J, Speidel S (2019) Using 3d convolutional neural networks to learn spatiotemporal features for automatic surgical gesture recognition in video. In: MICCAI","DOI":"10.1007\/978-3-030-32254-0_52"},{"key":"2388_CR11","doi-asserted-by":"publisher","first-page":"203","DOI":"10.1016\/j.media.2018.05.001","volume":"47","author":"HA Hajj","year":"2018","unstructured":"Hajj HA, Lamard M, Conze PH, Cochener B, Quellec G (2018) Monitoring tool usage in surgery videos using boosted convolutional and recurrent neural networks. Med Image Anal 47:203\u2013218","journal-title":"Med Image Anal"},{"key":"2388_CR12","doi-asserted-by":"crossref","unstructured":"He K, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition. In: CVPR","DOI":"10.1109\/CVPR.2016.90"},{"key":"2388_CR13","doi-asserted-by":"crossref","unstructured":"He K, Zhang X, Ren S, Sun J (2016) Identity mappings in deep residual networks. In: Computer Vision \u2013 ECCV 2016, pp. 630\u2013645. Springer International Publishing","DOI":"10.1007\/978-3-319-46493-0_38"},{"key":"2388_CR14","doi-asserted-by":"crossref","unstructured":"Jin A, Yeung S, Jopling J, Krause J, Azagury D, Milstein A, Fei-Fei L (2018) Tool detection and operative skill assessment in surgical videos using region-based convolutional neural networks. 2018 IEEE Winter Conference on Applications of Computer Vision (WACV) pp. 691\u2013699","DOI":"10.1109\/WACV.2018.00081"},{"issue":"5","key":"2388_CR15","doi-asserted-by":"publisher","first-page":"1114","DOI":"10.1109\/TMI.2017.2787657","volume":"37","author":"Y Jin","year":"2018","unstructured":"Jin Y, Dou Q, Chen H, Yu L, Qin J, Fu CW, Heng PA (2018) SV-RCNet: Workflow recognition from surgical videos using recurrent convolutional network. IEEE Trans Med Imaging 37(5):1114\u20131126","journal-title":"IEEE Trans Med Imaging"},{"key":"2388_CR16","doi-asserted-by":"publisher","DOI":"10.1016\/j.media.2019.101572","volume":"59","author":"Y Jin","year":"2020","unstructured":"Jin Y, Li H, Dou Q, Chen H, Qin J, Fu C, Heng P (2020) Multi-task recurrent convolutional network with correlation loss for surgical video analysis. Medical image analysis 59:","journal-title":"Medical image analysis"},{"issue":"9","key":"2388_CR17","doi-asserted-by":"publisher","first-page":"2634","DOI":"10.1007\/s11695-018-3219-7","volume":"28","author":"MA Kaijser","year":"2018","unstructured":"Kaijser MA, van Ramshorst GH, Emous M, Veeger NJGM, van Wagensveld BA, Pierie JPEN (2018) A delphi consensus of the crucial steps in gastric bypass and sleeve gastrectomy procedures in the netherlands. Obesity Surg 28(9):2634\u20132643","journal-title":"Obesity Surg"},{"issue":"9","key":"2388_CR18","doi-asserted-by":"publisher","first-page":"1427","DOI":"10.1007\/s11548-015-1222-1","volume":"10","author":"D Kati\u0107","year":"2015","unstructured":"Kati\u0107 D, Julliard C, Wekerle AL, Kenngott H, M\u00fcller-Stich BP, Dillmann R, Speidel S, Jannin P, Gibaud B (2015) LapOntoSPM: an ontology for laparoscopic surgeries and its application to surgical phase recognition. Int J Comput Assisted Radiol Surg 10(9):1427\u20131434","journal-title":"Int J Comput Assisted Radiol Surg"},{"issue":"5","key":"2388_CR19","doi-asserted-by":"publisher","first-page":"1681","DOI":"10.1007\/s00464-012-2656-y","volume":"27","author":"M Kranzfelder","year":"2012","unstructured":"Kranzfelder M, Staub C, Fiolka A, Schneider A, Gillen S, Wilhelm D, Friess H, Knoll A, Feussner H (2012) Toward increased autonomy in the surgical OR: needs, requests, and expectations. Surg Endoscopy 27(5):1681\u20131688","journal-title":"Surg Endoscopy"},{"key":"2388_CR20","doi-asserted-by":"crossref","unstructured":"Lea C, Vidal R, Reiter A, Hager GD (2016) Temporal convolutional networks: A unified approach to action segmentation. In: Lecture Notes in Computer Science, pp. 47\u201354. Springer International Publishing","DOI":"10.1007\/978-3-319-49409-8_7"},{"issue":"9","key":"2388_CR21","doi-asserted-by":"publisher","first-page":"691","DOI":"10.1038\/s41551-017-0132-7","volume":"1","author":"L Maier-Hein","year":"2017","unstructured":"Maier-Hein L, Vedula SS, Speidel S, Navab N, Kikinis R, Park A, Eisenmann M, Feussner H, Forestier G, Giannarou S, Hashizume M, Katic D, Kenngott H, Kranzfelder M, Malpani A, M\u00e4rz K, Neumuth T, Padoy N, Pugh C, Schoch N, Stoyanov D, Taylor R, Wagner M, Hager GD, Jannin P (2017) Surgical data science for next-generation interventions. Nat Biomed Eng 1(9):691\u2013696. https:\/\/doi.org\/10.1038\/s41551-017-0132-7","journal-title":"Nat Biomed Eng"},{"key":"2388_CR22","doi-asserted-by":"publisher","first-page":"1059","DOI":"10.1007\/s11548-019-01958-6","volume":"14","author":"CI Nwoye","year":"2019","unstructured":"Nwoye CI, Mutter D, Marescaux J, Padoy N (2019) Weakly supervised convolutional lstm approach for tool tracking in laparoscopic videos. Int J Comput Assisted Radiol Surg 14:1059\u20131067","journal-title":"Int J Comput Assisted Radiol Surg"},{"key":"2388_CR23","unstructured":"van\u00a0den Oord A, Dieleman S, Zen H, Simonyan K, Vinyals O, Graves A, Kalchbrenner N, Senior A, Kavukcuoglu K (2016) WaveNet: A generative model for raw audio. In: Arxiv"},{"key":"2388_CR24","unstructured":"Twinanda AP (2017) Vision-based approaches for surgical activity recognition using laparoscopic and rbgd videos. In: PhD thesis"},{"issue":"1","key":"2388_CR25","doi-asserted-by":"publisher","first-page":"86","DOI":"10.1109\/TMI.2016.2593957","volume":"36","author":"AP Twinanda","year":"2017","unstructured":"Twinanda AP, Shehata S, Mutter D, Marescaux J, de Mathelin M, Padoy N (2017) EndoNet: A deep architecture for recognition tasks on laparoscopic videos. IEEE Trans Med Imaging 36(1):86\u201397","journal-title":"IEEE Trans Med Imaging"},{"key":"2388_CR26","doi-asserted-by":"crossref","unstructured":"Varadarajan B, Reiley C, Lin H, Khudanpur S, Hager G (2009) Data-derived models for segmentation with application to surgical assessment and training. In: G.Z. Yang, D.\u00a0Hawkes, D.\u00a0Rueckert, A.\u00a0Noble, C.\u00a0Taylor (eds.) MICCAI, pp. 426\u2013434","DOI":"10.1007\/978-3-642-04268-3_53"},{"issue":"1","key":"2388_CR27","doi-asserted-by":"publisher","first-page":"198","DOI":"10.1109\/JPROC.2019.2946993","volume":"108","author":"T Vercauteren","year":"2020","unstructured":"Vercauteren T, Unberath M, Padoy N, Navab N (2020) Cai4cai: The rise of contextual artificial intelligence in computer-assisted interventions. Proc IEEE 108(1):198\u2013214","journal-title":"Proc IEEE"},{"key":"2388_CR28","unstructured":"Yu T, Mutter D, Marescaux J, Padoy N (2019) Learning from a tiny dataset of manual annotations: a teacher\/student approach for surgical phase recognition"},{"issue":"7","key":"2388_CR29","doi-asserted-by":"publisher","first-page":"732","DOI":"10.1016\/j.media.2013.04.007","volume":"17","author":"L Zappella","year":"2013","unstructured":"Zappella L, B\u00e9jar B, Hager G, Vidal R (2013) Surgical gesture classification from video and kinematic data. Med Image Anal 17(7):732\u2013745","journal-title":"Med Image Anal"},{"key":"2388_CR30","doi-asserted-by":"crossref","unstructured":"Zisimopoulos O, Flouty E, Luengo I, Giataganas P, Nehme J, Chow A, Stoyanov D (2018) DeepPhase: Surgical phase recognition in cataracts videos. In: MICCAI","DOI":"10.1007\/978-3-030-00937-3_31"}],"container-title":["International Journal of Computer Assisted Radiology and Surgery"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11548-021-02388-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11548-021-02388-z\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11548-021-02388-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,7,6]],"date-time":"2021-07-06T14:32:57Z","timestamp":1625581977000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11548-021-02388-z"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,5,19]]},"references-count":30,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2021,7]]}},"alternative-id":["2388"],"URL":"https:\/\/doi.org\/10.1007\/s11548-021-02388-z","relation":{},"ISSN":["1861-6410","1861-6429"],"issn-type":[{"value":"1861-6410","type":"print"},{"value":"1861-6429","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,5,19]]},"assertion":[{"value":"22 January 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"27 April 2021","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"19 May 2021","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of Interest"}},{"value":"The research was conducted in accordance with the 1964 Helsinki Declaration. The surgical videos were recorded and collected in an anonymized manner following informed consent of patients. The local medical research and ethical committee cleared the present study from the Research Involving Human Subjects Act since the study did not imply any deviation from standard of care.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical approval"}},{"value":"The patients consented to data recording","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Informed consent"}},{"value":"The source code is public available at .","order":5,"name":"Ethics","group":{"name":"EthicsHeading","label":"Code availability"}}]}}