{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,12]],"date-time":"2026-06-12T06:54:17Z","timestamp":1781247257061,"version":"3.54.1"},"reference-count":32,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2026,4,15]],"date-time":"2026-04-15T00:00:00Z","timestamp":1776211200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,5,8]],"date-time":"2026-05-08T00:00:00Z","timestamp":1778198400000},"content-version":"vor","delay-in-days":23,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["npj Digit. Med."],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Surgical procedures unfold in complex environments demanding coordination between surgical teams, tools, imaging and increasingly, intelligent robotic systems. While AI solutions like ChatGPT and Gemini have revolutionized language understanding and seen early adaptions in clinical diagnosis, they fall short in the safety-critical, multimodal setting of surgery. Ensuring safety and efficiency in ORs of the future requires intelligent systems, like surgical robots, smart instruments and digital copilots, capable of understanding complex activities and hazards. We introduce ORQA, a multimodal foundation model unifying visual, auditory, and structured data for holistic surgical understanding. ORQA\u2019s question-answering framework empowers diverse tasks, serving as an intelligence core for surgical technologies. We benchmark ORQA against generalist vision-language models, and show that while they struggle to perceive surgical scenes, ORQA delivers substantially stronger, consistent performance. To meet diverse deployment needs, we design, and release a family of smaller ORQA models tailored to different computational requirements. This work establishes a foundation for the next wave of intelligent surgical solutions, enabling surgical teams and medical technology providers to create smarter and safer operating rooms.<\/jats:p>","DOI":"10.1038\/s41746-026-02631-4","type":"journal-article","created":{"date-parts":[[2026,4,15]],"date-time":"2026-04-15T11:46:30Z","timestamp":1776253590000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Specialized foundation models for intelligent operating rooms"],"prefix":"10.1038","volume":"9","author":[{"given":"Ege","family":"\u00d6zsoy","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chantal","family":"Pellegrini","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"David","family":"Bani-Harouni","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kun","family":"Yuan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Matthias","family":"Keicher","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Nassir","family":"Navab","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,4,15]]},"reference":[{"key":"2631_CR1","doi-asserted-by":"publisher","DOI":"10.1038\/s41746-024-01225-2","volume":"7","author":"S Protserov","year":"2024","unstructured":"Protserov, S. et al. Development, deployment and scaling of operating room-ready artificial intelligence for real-time surgical decision support. NPJ Digit. Med. 7, 231 (2024).","journal-title":"NPJ Digit. Med."},{"key":"2631_CR2","doi-asserted-by":"publisher","first-page":"1976","DOI":"10.1007\/s00464-020-08231-x","volume":"35","author":"F Kanji","year":"2021","unstructured":"Kanji, F. et al. Work-system interventions in robotic-assisted surgery: a systematic review exploring the gap between challenges and solutions. Surg. Endosc. 35, 1976\u20131989 (2021).","journal-title":"Surg. Endosc."},{"key":"2631_CR3","doi-asserted-by":"publisher","first-page":"495","DOI":"10.1007\/s11548-013-0940-5","volume":"9","author":"F Lalys","year":"2014","unstructured":"Lalys, F. & Jannin, P. Surgical process modelling: a review. Int. J. Comput. Assist. Radiol. Surg. Springe. Verl. 9, 495\u2013511 (2014).","journal-title":"Int. J. Comput. Assist. Radiol. Surg. Springe. Verl."},{"key":"2631_CR4","doi-asserted-by":"publisher","first-page":"691","DOI":"10.1038\/s41551-017-0132-7","volume":"1","author":"L Maier-Hein","year":"2017","unstructured":"Maier-Hein, L. et al. Surgical data science for next-generation interventions. Nat. Biomed. Eng. 1, 691\u2013696 (2017).","journal-title":"Nat. Biomed. Eng."},{"key":"2631_CR5","doi-asserted-by":"publisher","first-page":"86","DOI":"10.1109\/TMI.2016.2593957","volume":"36","author":"AP Twinanda","year":"2016","unstructured":"Twinanda, A. P. et al. Endonet: a deep architecture for recognition tasks on laparoscopic videos. IEEE Trans. Med. imaging 36, 86\u201397 (2016).","journal-title":"IEEE Trans. Med. imaging"},{"key":"2631_CR6","doi-asserted-by":"publisher","first-page":"101572","DOI":"10.1016\/j.media.2019.101572","volume":"59","author":"Y Jin","year":"2020","unstructured":"Jin, Y. et al. Multi-task recurrent convolutional network with correlation loss for surgical video analysis. Med. image Anal. 59, 101572 (2020).","journal-title":"Med. image Anal."},{"key":"2631_CR7","doi-asserted-by":"publisher","first-page":"102433","DOI":"10.1016\/j.media.2022.102433","volume":"78","author":"CI Nwoye","year":"2022","unstructured":"Nwoye, C. I. et al. Rendezvous: attention mechanisms for the recognition of surgical action triplets in endoscopic videos. Med. Image Anal. 78, 102433 (2022).","journal-title":"Med. Image Anal."},{"key":"2631_CR8","doi-asserted-by":"crossref","unstructured":"Ding, H., Zhang, J., Kazanzides, P., Wu, J. Y. & Unberath, M. Carts: Causality-driven robot tool segmentation from vision and kinematics data. In Medical Image Computing and Computer Assisted Intervention\u2013MICCAI 2022: 25th International Conference, Singapore, September 18\u201322, 2022, Proceedings, Part VII, 387\u2013398 (Springer, 2022).","DOI":"10.1007\/978-3-031-16449-1_37"},{"key":"2631_CR9","doi-asserted-by":"crossref","unstructured":"Pei, J. et al. S 2 former-or: Single-stage bi-modal transformer for scene graph generation in or. IEEE Trans. Med. Imaging 44, 361\u2013372 (2024).","DOI":"10.1109\/TMI.2024.3444279"},{"key":"2631_CR10","doi-asserted-by":"crossref","unstructured":"Guo, D. et al. Tri-modal confluence with temporal dynamics for scene graph generation in operating rooms. In International Conference on Medical Image Computing and Computer-Assisted Intervention, 714\u2013724 (Springer, 2024).","DOI":"10.1007\/978-3-031-72089-5_67"},{"key":"2631_CR11","doi-asserted-by":"crossref","unstructured":"He, R. et al. Pitvqa: Image-grounded text embedding LLM for visual question answering in pituitary surgery. In International Conference on Medical Image Computing and Computer-Assisted Intervention, 488\u2013498 (Springer, 2024).","DOI":"10.1007\/978-3-031-72089-5_46"},{"key":"2631_CR12","doi-asserted-by":"publisher","first-page":"1409","DOI":"10.1007\/s11548-024-03141-y","volume":"19","author":"K Yuan","year":"2024","unstructured":"Yuan, K. et al. Advancing surgical VQA with scene graph knowledge. Int. J. Comput. Assist. Radiol. Surg. 19, 1409\u20131417 (2024).","journal-title":"Int. J. Comput. Assist. Radiol. Surg."},{"key":"2631_CR13","unstructured":"Srivastav, V. K. et al. Mvor: A multi-view RGB-D operating room dataset for 2d and 3d human pose estimation. In MICCAI 2018 Satellite Workshop, Granada, Spain, September 16-20 2018 (Springer, 2018)."},{"key":"2631_CR14","doi-asserted-by":"crossref","unstructured":"Fujii, R., Hatano, M., Saito, H. & Kajita, H. Egosurgery-phase: a dataset of surgical phase recognition from egocentric open surgery videos. In International Conference on Medical Image Computing and Computer-Assisted Intervention, 187\u2013196 (Springer, 2024).","DOI":"10.1007\/978-3-031-72089-5_18"},{"key":"2631_CR15","doi-asserted-by":"crossref","unstructured":"Wu, Y. et al. Teleor: Real-time telemedicine system for full-scene operating room. In International Conference on Medical Image Computing and Computer-Assisted Intervention, 628\u2013638 (Springer, 2024).","DOI":"10.1007\/978-3-031-72089-5_59"},{"key":"2631_CR16","doi-asserted-by":"crossref","unstructured":"\u00d6zsoy, E. et al. 4d-or: Semantic scene graphs for or domain modeling. In Medical Image Computing and Computer Assisted Intervention\u2013MICCAI 2022: 25th International Conference, Singapore, September 18\u201322, 2022, Proceedings, Part VII (Springer, 2022).","DOI":"10.1007\/978-3-031-16449-1_45"},{"key":"2631_CR17","doi-asserted-by":"crossref","unstructured":"\u00d6zsoy, E., Czempiel, T., Holm, F., Pellegrini, C. & Navab, N. Labrad-or: Lightweight memory scene graphs for accurate bimodal reasoning in dynamic operating rooms. In International Conference on Medical Image Computing and Computer-Assisted Intervention (Springer, 2023).","DOI":"10.1007\/978-3-031-43996-4_29"},{"key":"2631_CR18","doi-asserted-by":"crossref","unstructured":"\u00d6zsoy, E., Pellegrini, C., Keicher, M. & Navab, N. Oracle: Large vision-language models for knowledge-guided holistic or domain modeling. In International Conference on Medical Image Computing and Computer-Assisted Intervention, 455\u2013465 (Springer, 2024).","DOI":"10.1007\/978-3-031-72089-5_43"},{"key":"2631_CR19","doi-asserted-by":"crossref","unstructured":"\u00d6zsoy, E. et al. Mm-or: A large multimodal operating room dataset for semantic understanding of high-intensity surgical environments. In CVPR (2025).","DOI":"10.1109\/CVPR52734.2025.01805"},{"key":"2631_CR20","doi-asserted-by":"publisher","first-page":"7866","DOI":"10.1038\/s41467-025-62385-7","volume":"16","author":"C Wu","year":"2025","unstructured":"Wu, C. et al. Towards generalist foundation model for radiology by leveraging web-scale 2d&3d medical data. Nature Communications 16, 7866 (2025).","journal-title":"Nature Communications"},{"key":"2631_CR21","unstructured":"Pellegrini, C. et al. Radialog: Large vision-language models for x-ray reporting and dialog-driven assistance. In Medical Imaging with Deep Learning (2025)."},{"key":"2631_CR22","doi-asserted-by":"publisher","first-page":"466","DOI":"10.1038\/s41586-024-07618-3","volume":"634","author":"MY Lu","year":"2024","unstructured":"Lu, M. Y. et al. A multimodal generative AI copilot for human pathology. Nature 634, 466\u2013473 (2024).","journal-title":"Nature"},{"key":"2631_CR23","unstructured":"Li, J. et al. Llava-surg: towards multimodal surgical assistant via structured surgical video learning. arXiv preprint arXiv:2408.07981 (2024)."},{"key":"2631_CR24","unstructured":"Shukor, M., Dancette, C., Rame, A. & Cord, M. UnIVAL: Unified model for image, video, audio and language tasks. Trans. Mach. Learn. Res. https:\/\/openreview.net\/forum?id=4uflhObpcp (2023)."},{"key":"2631_CR25","doi-asserted-by":"crossref","unstructured":"Lu, J. et al. Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 26439\u201326455 (2024).","DOI":"10.1109\/CVPR52733.2024.02497"},{"key":"2631_CR26","doi-asserted-by":"crossref","unstructured":"Wu, X. et al. Point transformer v3: Simpler faster stronger. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 4840\u20134851 (2024).","DOI":"10.1109\/CVPR52733.2024.00463"},{"key":"2631_CR27","doi-asserted-by":"crossref","unstructured":"Elizalde, B., Deshmukh, S., Al Ismail, M. & Wang, H. Clap learning audio concepts from natural language supervision. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 1\u20135 (IEEE, 2023).","DOI":"10.1109\/ICASSP49357.2023.10095889"},{"key":"2631_CR28","unstructured":"Radford, A. et al. Robust speech recognition via large-scale weak supervision. In International conference on machine learning, 28492\u201328518 (PMLR, 2023)."},{"key":"2631_CR29","unstructured":"Wang, P. et al. Qwen2-vl: Enhancing vision-language model\u2019s perception of the world at any resolution. arXiv preprint arXiv:2409.12191 (2024)."},{"key":"2631_CR30","unstructured":"Radford, A. et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, 8748\u20138763 (PMLR, 2021)."},{"key":"2631_CR31","unstructured":"Hu, E. J. et al. Lora: Low-rank adaptation of large language models. Iclr 1, 3 (2022)."},{"key":"2631_CR32","unstructured":"Hinton, G., Vinyals, O. & Dean, J. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531 (2015)."}],"container-title":["npj Digital Medicine"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.nature.com\/articles\/s41746-026-02631-4","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.nature.com\/articles\/s41746-026-02631-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.nature.com\/articles\/s41746-026-02631-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,8]],"date-time":"2026-05-08T06:16:13Z","timestamp":1778220973000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.nature.com\/articles\/s41746-026-02631-4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,15]]},"references-count":32,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2026,12]]}},"alternative-id":["2631"],"URL":"https:\/\/doi.org\/10.1038\/s41746-026-02631-4","relation":{},"ISSN":["2398-6352"],"issn-type":[{"value":"2398-6352","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,15]]},"assertion":[{"value":"7 August 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"4 April 2026","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 April 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The authors declare no competing interests.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"362"}}