{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,19]],"date-time":"2026-06-19T20:21:40Z","timestamp":1781900500228,"version":"3.54.5"},"reference-count":39,"publisher":"Oxford University Press (OUP)","issue":"2","license":[{"start":{"date-parts":[[2022,2,26]],"date-time":"2022-02-26T00:00:00Z","timestamp":1645833600000},"content-version":"vor","delay-in-days":1,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key Research and Development Program of China","doi-asserted-by":"publisher","award":["2019YFB1311300"],"award-info":[{"award-number":["2019YFB1311300"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2022,2,25]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Laparoscopic surgery, as a representative minimally invasive surgery (MIS), is an active research area of clinical practice. Automatic surgical phase recognition of laparoscopic videos is a vital task with the potential to improve surgeons\u2019 efficiency and has gradually become an integral part of computer-assisted intervention systems in MIS. However, the performance of most methods currently employed for surgical phase recognition is deteriorated by optimization difficulties and inefficient computation, which hinders their large-scale practical implementation. This study proposes an efficient and novel surgical phase recognition method using an attention-based spatial\u2013temporal neural network consisting of a spatial model and a temporal model for accurate recognition by end-to-end training. The former subtly incorporates the attention mechanism to enhance the model\u2019s ability to focus on the key regions in video frames and efficiently capture more informative visual features. In the temporal model, we employ independently recurrent long short-term memory (IndyLSTM) and non-local block to extract long-term temporal information of video frames. We evaluated the performance of our method on the publicly available Cholec80 dataset. Our attention-based spatial\u2013temporal neural network purely produces the phase predictions without any post-processing strategies, achieving excellent recognition performance and outperforming other state-of-the-art phase recognition methods.<\/jats:p>","DOI":"10.1093\/jcde\/qwac011","type":"journal-article","created":{"date-parts":[[2022,1,17]],"date-time":"2022-01-17T12:06:41Z","timestamp":1642421201000},"page":"406-416","source":"Crossref","is-referenced-by-count":14,"title":["Attention-based spatial\u2013temporal neural network for accurate phase recognition in minimally invasive surgery: feasibility and efficiency verification"],"prefix":"10.1093","volume":"9","author":[{"given":"Pan","family":"Shi","sequence":"first","affiliation":[{"name":"School of Control Science and Engineering, Shandong University, Jinan 250061, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zijian","family":"Zhao","sequence":"additional","affiliation":[{"name":"School of Control Science and Engineering, Shandong University, Jinan 250061, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kaidi","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Control Science and Engineering, Shandong University, Jinan 250061, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Feng","family":"Li","sequence":"additional","affiliation":[{"name":"Department of General Surgery, Qilu Hospital of Shandong University, Jinan 250012, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2022,2,25]]},"reference":[{"key":"2022022603110896400_bib1","doi-asserted-by":"crossref","first-page":"101224","DOI":"10.1016\/j.csl.2021.101224","article-title":"Enhancing Arabic aspect-based sentiment analysis using deep learning models","volume":"69","author":"Al-Dabet","year":"2021","journal-title":"Computer Speech & Language"},{"key":"2022022603110896400_bib2","article-title":"Unsupervised temporal context learning using convolutional neural networks for laparoscopic workflow analysis","author":"Bodenstedt","year":"2017"},{"key":"2022022603110896400_bib3","doi-asserted-by":"crossref","first-page":"633","DOI":"10.1016\/j.media.2016.09.003","article-title":"Vision-based and marker-less surgical tool detection and tracking: A review of the literature","volume":"35","author":"Bouget","year":"2017","journal-title":"Medical Image Analysis"},{"key":"2022022603110896400_bib4","first-page":"343","article-title":"TeCNO: Surgical phase recognition with multi-stage temporal convolutional networks","volume-title":"International Conference on Medical Image Computing and Computer-Assisted Intervention","author":"Czempiel","year":"2020"},{"key":"2022022603110896400_bib5","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-030-87202-1_58","article-title":"OperA: Attention-regularized transformers for surgical phase recognition","author":"Czempiel","year":"2021"},{"issue":"4","key":"2022022603110896400_bib6","doi-asserted-by":"crossref","first-page":"677","DOI":"10.1109\/TPAMI.2016.2599174","article-title":"Long-term recurrent convolutional networks for visual recognition and description","volume":"39","author":"Donahue","year":"2017","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2022022603110896400_bib7","first-page":"3575","article-title":"MS-TCN: Multi-stage temporal convolutional network for action segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Farha","year":"2019"},{"issue":"6","key":"2022022603110896400_bib8","doi-asserted-by":"crossref","first-page":"361","DOI":"10.1007\/s10151-016-1444-4","article-title":"Application of objective clinical human reliability analysis (OCHRA) in assessment of technical performance in laparoscopic rectal cancer surgery","volume":"20","author":"Foster","year":"2016","journal-title":"Techniques in Coloproctology"},{"issue":"4","key":"2022022603110896400_bib9","doi-asserted-by":"crossref","first-page":"684","DOI":"10.1097\/SLA.0000000000004425","article-title":"Machine learning for surgical phase recognition: A systematic review","volume":"273","author":"Garrow","year":"2021","journal-title":"Annals of Surgery"},{"key":"2022022603110896400_bib10","first-page":"3352","article-title":"IndyLSTMs: Independently recurrent LSTMs","volume-title":"IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Gonnet","year":"2019"},{"key":"2022022603110896400_bib11","doi-asserted-by":"crossref","first-page":"691","DOI":"10.1109\/WACV.2018.00081","article-title":"Tool detection and operative skill assessment in surgical videos using region-based convolutional neural networks","volume-title":"2018 IEEE Winter Conference on Applications of Computer Vision (WACV)","author":"Jin","year":"2018"},{"issue":"5","key":"2022022603110896400_bib12","doi-asserted-by":"crossref","first-page":"1114","DOI":"10.1109\/TMI.2017.2787657","article-title":"SV-RCNet: Workflow recognition from surgical videos using recurrent convolutional network","volume":"37","author":"Jin","year":"2018","journal-title":"IEEE Transactions on Medical Imaging"},{"key":"2022022603110896400_bib13","doi-asserted-by":"crossref","first-page":"101572","DOI":"10.1016\/j.media.2019.101572","article-title":"Multi-task recurrent convolutional network with correlation loss for surgical video analysis","volume":"59","author":"Jin","year":"2019","journal-title":"Medical Image Analysis"},{"issue":"7","key":"2022022603110896400_bib14","doi-asserted-by":"crossref","first-page":"1911","DOI":"10.1109\/TMI.2021.3069471","article-title":"Temporal memory relation network for workflow recognition from surgical video","volume":"40","author":"Jin","year":"2021","journal-title":"IEEE Transactions on Medical Imaging"},{"key":"2022022603110896400_bib15","doi-asserted-by":"crossref","first-page":"109216","DOI":"10.1016\/j.jcp.2019.109216","article-title":"Deep unsupervised learning of turbulence for inflow generation at various Reynolds numbers","volume":"406","author":"Kim","year":"2020","journal-title":"Journal of Computational Physics"},{"key":"2022022603110896400_bib16","doi-asserted-by":"crossref","first-page":"928","DOI":"10.1109\/ICCES48766.2020.9137855","article-title":"Deep learning-based surgical workflow recognition from laparoscopic videos","volume-title":"2020 5th International Conference on Communication and Electronics Systems (ICCES)","author":"Kurian","year":"2020"},{"issue":"12","key":"2022022603110896400_bib18","doi-asserted-by":"crossref","first-page":"6073","DOI":"10.1109\/TNNLS.2018.2817538","article-title":"Rank-constrained spectral clustering with flexible embedding","volume":"29","author":"Li","year":"2018","journal-title":"IEEE Transactions on Neural Networks and Learning Systems"},{"issue":"12","key":"2022022603110896400_bib19","doi-asserted-by":"crossref","first-page":"6323","DOI":"10.1109\/TNNLS.2018.2829867","article-title":"Dynamic affinity graph construction for spectral clustering using multiple features","volume":"29","author":"Li","year":"2018","journal-title":"IEEE Transactions on Neural Networks and Learning Systems"},{"key":"2022022603110896400_bib20","doi-asserted-by":"crossref","first-page":"595","DOI":"10.1016\/j.patcog.2018.12.010","article-title":"Zero-shot event detection via event-adaptive concept relevance mining","volume":"88","author":"Li","year":"2019","journal-title":"Pattern Recognition"},{"key":"2022022603110896400_bib17","first-page":"1","article-title":"A combined loss-based multiscale fully convolutional network for high-resolution remote sensing image change detection","volume":"19","author":"Li","year":"2021","journal-title":"IEEE Geoscience and Remote Sensing Letters"},{"issue":"2","key":"2022022603110896400_bib21","doi-asserted-by":"crossref","first-page":"1","DOI":"10.14500\/aro.10827","article-title":"Efficient Kinect sensor-based Kurdish sign language recognition using echo system network","volume":"9","author":"Mirza","year":"2021","journal-title":"ARO-The Scientific Journal of Koya University"},{"key":"2022022603110896400_bib22","article-title":"Multitask learning of temporal connectionism in convolutional networks using a joint distribution loss function to simultaneously identify tools and phase in surgical videos","author":"Mondal","year":"2019"},{"key":"2022022603110896400_bib23","first-page":"124","article-title":"Automatic detection of surgical phases in laparoscopic videos","volume-title":"Proceedings on the International Conference on Artificial Intelligence (ICAI)","author":"Namazi","year":"2018"},{"issue":"2","key":"2022022603110896400_bib24","doi-asserted-by":"crossref","first-page":"82","DOI":"10.1080\/13645706.2019.1584116","article-title":"Machine and deep learning for workflow recognition during surgery","volume":"28","author":"Padoy","year":"2019","journal-title":"Minimally Invasive Therapy & Allied Technologies"},{"key":"2022022603110896400_bib25","doi-asserted-by":"crossref","first-page":"241","DOI":"10.1007\/978-3-319-73603-7_20","article-title":"Frame-based classification of operation phases in cataract surgery videos","volume-title":"International Conference on Multimedia Modeling","author":"Primus","year":"2018"},{"issue":"7","key":"2022022603110896400_bib26","doi-asserted-by":"crossref","first-page":"1111","DOI":"10.1007\/s11548-021-02388-z","article-title":"Multi-task temporal convolutional networks for joint recognition of surgical phases and steps in gastric bypass procedures","volume":"16","author":"Ramesh","year":"2021","journal-title":"International Journal of Computer Assisted Radiology and Surgery"},{"issue":"9","key":"2022022603110896400_bib27","doi-asserted-by":"crossref","first-page":"1573","DOI":"10.1007\/s11548-020-02198-9","article-title":"LRTD: Long-range temporal dependency-based active learning for surgical workflow recognition","volume":"15","author":"Shi","year":"2020","journal-title":"International Journal of Computer Assisted Radiology and Surgery"},{"key":"2022022603110896400_bib28","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/I2CT51068.2021.9418138","article-title":"Face recognition from image patches using an ensemble of CNN-local mesh pattern networks","volume-title":"2021 6th International Conference for Convergence in Technology (I2CT)","author":"Thomas","year":"2021"},{"key":"2022022603110896400_bib30","article-title":"Single- and multi-task architectures for tool presence detection challenge at M2CAI 2016","author":"Twinanda","year":"2016"},{"key":"2022022603110896400_bib29","volume-title":"Vision-based approaches for surgical activity recognition using laparoscopic and rbgd videos","author":"Twinanda","year":"2017"},{"issue":"1","key":"2022022603110896400_bib31","doi-asserted-by":"crossref","first-page":"86","DOI":"10.1109\/TMI.2016.2593957","article-title":"EndoNet: A deep architecture for recognition tasks on laparoscopic videos","volume":"36","author":"Twinanda","year":"2017","journal-title":"IEEE Transactions on Medical Imaging"},{"issue":"4","key":"2022022603110896400_bib32","doi-asserted-by":"crossref","first-page":"1069","DOI":"10.1109\/TMI.2018.2878055","article-title":"RSDNet: Learning to predict remaining surgery duration from laparoscopic videos without manual annotations","volume":"38","author":"Twinanda","year":"2019","journal-title":"IEEE Transactions on Medical Imaging"},{"key":"2022022603110896400_bib33","doi-asserted-by":"crossref","first-page":"1","DOI":"10.23919\/FUSION45008.2020.9190449","article-title":"Uncertainty based active learning with deep neural networks for inertial gait analysis","volume-title":"2020 IEEE 23rd International Conference on Information Fusion (FUSION)","author":"Vaith","year":"2020"},{"key":"2022022603110896400_bib34","first-page":"7794","article-title":"Non-local neural networks","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Wang","year":"2018"},{"key":"2022022603110896400_bib35","first-page":"22","article-title":"Instrument tracking with rigid part mixtures model","volume-title":"Computer-assisted and robotic endoscopy","author":"Wesierski","year":"2015"},{"key":"2022022603110896400_bib36","first-page":"3","article-title":"CBAM: Convolutional block attention module","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"Woo","year":"2018"},{"issue":"5","key":"2022022603110896400_bib37","doi-asserted-by":"crossref","first-page":"839","DOI":"10.1007\/s11548-021-02382-5","article-title":"Against spatial\u2013temporal discrepancy: Contrastive learning-based network for surgical workflow recognition","volume":"16","author":"Xia","year":"2021","journal-title":"International Journal of Computer Assisted Radiology and Surgery"},{"key":"2022022603110896400_bib38","first-page":"449","article-title":"Hard frame detection and online mapping for surgical phase recognition","volume-title":"International Conference on Medical Image Computing and Computer-Assisted Intervention","author":"Yi","year":"2019"},{"key":"2022022603110896400_bib39","first-page":"409","article-title":"Symmetric dilated convolution for surgical gesture recognition","volume-title":"International Conference on Medical Image Computing and Computer-Assisted Intervention","author":"Zhang","year":"2020"}],"container-title":["Journal of Computational Design and Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/jcde\/article-pdf\/9\/2\/406\/42616845\/qwac011.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/jcde\/article-pdf\/9\/2\/406\/42616845\/qwac011.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,2,26]],"date-time":"2022-02-26T03:12:30Z","timestamp":1645845150000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/jcde\/article\/9\/2\/406\/6537180"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,2,25]]},"references-count":39,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2022,2,25]]}},"URL":"https:\/\/doi.org\/10.1093\/jcde\/qwac011","relation":{},"ISSN":["2288-5048"],"issn-type":[{"value":"2288-5048","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2022,4]]},"published":{"date-parts":[[2022,2,25]]}}}