{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,20]],"date-time":"2026-07-20T13:54:47Z","timestamp":1784555687674,"version":"3.55.0"},"reference-count":59,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2026,3,16]],"date-time":"2026-03-16T00:00:00Z","timestamp":1773619200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,5,4]],"date-time":"2026-05-04T00:00:00Z","timestamp":1777852800000},"content-version":"vor","delay-in-days":49,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["npj Digit. Med."],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    Electrocardiograms (ECGs) are essential, non-invasive diagnostic tools for assessing cardiac conditions. Existing methods often have limited generalizability, focus on narrow condition sets, and rely on raw physiological signals, which may be unavailable in resource-limited settings where only printed or digital ECG images are accessible. Recent advances in multimodal large language models (MLLMs) offer new opportunities, yet ECG image interpretation remains challenging due to the lack of instruction-tuning data and standardized benchmarks. To address these gaps, we introduce , the first large-scale ECG image instruction-tuning dataset with over one million samples, covering diverse tasks including feature recognition, rhythm analysis, morphology assessment, and clinical report generation. We develop , a fully open-source MLLM for ECG image interpretation trained on . We further curate , a human expert-developed benchmark spanning four core ECG interpretation tasks across nine datasets, incorporating both synthesized and real-world ECG images to enable clinically realistic evaluation. Our experiments demonstrate that establishes a new state of the art, outperforming general-purpose MLLMs by 21% to 33% in average accuracy. These results highlight the potential of to improve ECG image interpretation in clinical practice. All code, data and models are available at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/aimedlab.github.io\/PULSE\/\" ext-link-type=\"uri\">https:\/\/aimedlab.github.io\/PULSE\/<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1038\/s41746-026-02551-3","type":"journal-article","created":{"date-parts":[[2026,3,16]],"date-time":"2026-03-16T11:09:01Z","timestamp":1773659341000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["Teaching multimodal LLMs to comprehend 12-lead electrocardiographic images"],"prefix":"10.1038","volume":"9","author":[{"given":"Ruoqi","family":"Liu","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yuelin","family":"Bai","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiang","family":"Yue","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ping","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,3,16]]},"reference":[{"key":"2551_CR1","doi-asserted-by":"publisher","first-page":"65","DOI":"10.1038\/s41591-018-0268-3","volume":"25","author":"AY Hannun","year":"2019","unstructured":"Hannun, A. Y. et al. Cardiologist-level arrhythmia detection and classification in ambulatory electrocardiograms using a deep neural network. Nat. Med. 25, 65\u201369 (2019).","journal-title":"Nat. Med."},{"key":"2551_CR2","doi-asserted-by":"publisher","DOI":"10.1038\/s41467-020-15432-4","volume":"11","author":"AH Ribeiro","year":"2020","unstructured":"Ribeiro, A. H. et al. Automatic diagnosis of the 12-lead ecg using a deep neural network. Nat. Commun. 11, 1760 (2020).","journal-title":"Nat. Commun."},{"key":"2551_CR3","doi-asserted-by":"publisher","first-page":"1285","DOI":"10.1001\/jamacardio.2021.2746","volume":"6","author":"JW Hughes","year":"2021","unstructured":"Hughes, J. W. et al. Performance of a convolutional neural network and explainability technique for 12-lead electrocardiogram interpretation. JAMA Cardiol 6, 1285\u20131295 (2021).","journal-title":"JAMA Cardiol"},{"key":"2551_CR4","doi-asserted-by":"publisher","first-page":"465","DOI":"10.1038\/s41569-020-00503-2","volume":"18","author":"KC Siontis","year":"2021","unstructured":"Siontis, K. C., Noseworthy, P. A., Attia, Z. I. & Friedman, P. A. Artificial intelligence-enhanced electrocardiography in cardiovascular disease management. Nat. Rev. Cardiol. 18, 465\u2013478 (2021).","journal-title":"Nat. Rev. Cardiol."},{"key":"2551_CR5","doi-asserted-by":"publisher","DOI":"10.1038\/s41467-022-29153-3","volume":"13","author":"V Sangha","year":"2022","unstructured":"Sangha, V. et al. Automated multilabel diagnosis on electrocardiographic images and signals. Nat. Commun. 13, 1583 (2022).","journal-title":"Nat. Commun."},{"key":"2551_CR6","doi-asserted-by":"publisher","first-page":"765","DOI":"10.1161\/CIRCULATIONAHA.122.062646","volume":"148","author":"V Sangha","year":"2023","unstructured":"Sangha, V. et al. Detection of left ventricular systolic dysfunction from electrocardiographic images. Circulation 148, 765\u2013777 (2023).","journal-title":"Circulation"},{"key":"2551_CR7","unstructured":"OpenAI. Gpt-4v(ision) technical work and authors (2023). https:\/\/cdn.openai.com\/contributions\/gpt-4v.pdf (2023)."},{"key":"2551_CR8","unstructured":"Li, B. et al. Llava-onevision: Easy visual task transfer. Trans. Mach. Learn. Res. (2025)."},{"key":"2551_CR9","first-page":"28541","volume":"36","author":"C Li","year":"2023","unstructured":"Li, C. et al. Llava-med: Training a large language-and-vision assistant for biomedicine in one day. Adv. Neural Inform. Process. Syst. 36, 28541\u201328564 (2023).","journal-title":"Adv. Neural Inform. Process. Syst."},{"key":"2551_CR10","unstructured":"Liu, H. et al. Llava-next: Improved reasoning, ocr, and world knowledge (2024). https:\/\/llava-vl.github.io\/blog\/2024-01-30-llava-next\/."},{"key":"2551_CR11","doi-asserted-by":"publisher","first-page":"11941","DOI":"10.3390\/ijerph191911941","volume":"19","author":"D Cuevas-Gonz\u00e1lez","year":"2022","unstructured":"Cuevas-Gonz\u00e1lez, D. et al. Ecg standards and formats for interoperability between mhealth and healthcare information systems: a scoping review. Int. J. Environ. Res. Public Health 19, 11941 (2022).","journal-title":"Int. J. Environ. Res. Public Health"},{"key":"2551_CR12","first-page":"46595","volume":"36","author":"L Zheng","year":"2024","unstructured":"Zheng, L. et al. Judging llm-as-a-judge with mt-bench and chatbot arena. Adv. Neural Inform. Process. Syst. 36, 46595\u201346623 (2024).","journal-title":"Adv. Neural Inform. Process. Syst."},{"key":"2551_CR13","unstructured":"Li, J., Liu, C., Cheng, S., Arcucci, R. & Hong, S. Frozen language model helps ecg zero-shot learning. In Medical Imaging with Deep Learning, 402\u2013415 (PMLR, 2024)."},{"key":"2551_CR14","unstructured":"Liu, C. et al. Zero-shot ecg classification with multimodal learning and test-time clinical knowledge enhancement. In Proc. 41st Int. Conf. Mach. Learn. 31949\u201331963 (2024)."},{"key":"2551_CR15","unstructured":"Na, Y., Park, M., Tae, Y. & Joo, S. Guiding masked representation learning to capture spatio-temporal relationship of electrocardiogram. In The Twelfth International Conference on Learning Representations (2023)."},{"key":"2551_CR16","doi-asserted-by":"publisher","first-page":"103451","DOI":"10.1016\/j.media.2024.103451","volume":"101","author":"\u00d6 Turgut","year":"2025","unstructured":"Turgut, \u00d6. et al. Unlocking the diagnostic potential of electrocardiograms through information transfer from cardiac magnetic resonance imaging. Med. Image Anal. 101, 103451 (2025).","journal-title":"Med. Image Anal."},{"key":"2551_CR17","unstructured":"Goswami, M. et al. Moment: A family of open time-series foundation models. In Proc. 41st Int. Conf. Mach. Learn. 16115\u201316152 (2024)."},{"key":"2551_CR18","unstructured":"Khunte, A. et al. Automated diagnostic reports from images of electrocardiograms at the point-of-care. medRxiv (2024)."},{"key":"2551_CR19","unstructured":"OpenAI. Gpt-4o contributions (2024). https:\/\/openai.com\/gpt-4o-contributions\/ (2024)."},{"key":"2551_CR20","unstructured":"Reid, M. et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. Google DeepMind technical report (2024)."},{"key":"2551_CR21","unstructured":"Anthropic. Claude 3.5 sonnet (2024). https:\/\/www.anthropic.com\/news\/claude-3-5-sonnet. Accessed: September 24, 2024."},{"key":"2551_CR22","first-page":"34892","volume":"36","author":"H Liu","year":"2023","unstructured":"Liu, H., Li, C., Wu, Q. & Lee, Y. J. Visual instruction tuning. Adv. Neural Inform. Process. Syst. 36, 34892\u201334916 (2023).","journal-title":"Adv. Neural Inform. Process. Syst."},{"key":"2551_CR23","doi-asserted-by":"crossref","unstructured":"Liu, H., Li, C., Li, Y. & Lee, Y. J. Improved baselines with visual instruction tuning. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 26296\u201326306 (2024).","DOI":"10.1109\/CVPR52733.2024.02484"},{"key":"2551_CR24","unstructured":"Abdin, M. et al. Phi-3 technical report: A highly capable language model locally on your phone. Microsoft Research technical report (2024)."},{"key":"2551_CR25","first-page":"87874","volume":"37","author":"H Lauren\u00e7on","year":"2024","unstructured":"Lauren\u00e7on, H., Tronchon, L., Cord, M. & Sanh, V. What matters when building vision-language models? Adv. Neural Inform. Process. Syst. 37, 87874\u201387907 (2024).","journal-title":"Adv. Neural Inform. Process. Syst."},{"key":"2551_CR26","unstructured":"Lu, H. et al. Deepseek-vl: Towards real-world vision-language understanding (2024). 2403.05525."},{"key":"2551_CR27","unstructured":"Jiang, D. et al. Mantis: Interleaved multi-image instruction tuning. Trans. Mach. Learn. Res. (2024)."},{"key":"2551_CR28","unstructured":"Yao, Y. et al. Minicpm-v: A gpt-4v level mllm on your phone. arXiv preprint arXiv:2408.01800 (2024)."},{"key":"2551_CR29","doi-asserted-by":"crossref","unstructured":"Chen, Z. et al. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. In Proc. IEEE\/CVF Conf. Comput. Vis. Pattern Recognit. 24185\u201324198 (2024).","DOI":"10.1109\/CVPR52733.2024.02283"},{"key":"2551_CR30","doi-asserted-by":"publisher","first-page":"220101","DOI":"10.1007\/s11432-024-4231-5","volume":"67","author":"Z Chen","year":"2024","unstructured":"Chen, Z. et al. How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites. Sci. China Inf. Sci. 67, 220101 (2024).","journal-title":"Sci. China Inf. Sci."},{"key":"2551_CR31","unstructured":"Wang, P. et al. Qwen2-vl: Enhancing vision-language model\u2019s perception of the world at any resolution. arXiv preprint arXiv:2409.12191 (2024)."},{"key":"2551_CR32","unstructured":"Bao, H., Dong, L., Piao, S. & Wei, F. Beit: Bert pre-training of image transformers. In Proc. Int. Conf. Learn. Represent. (2022)."},{"key":"2551_CR33","first-page":"9","volume":"1","author":"A Radford","year":"2019","unstructured":"Radford, A. et al. Language models are unsupervised multitask learners. OpenAI blog 1, 9 (2019).","journal-title":"OpenAI blog"},{"key":"2551_CR34","doi-asserted-by":"publisher","first-page":"172","DOI":"10.1038\/s41586-023-06291-2","volume":"620","author":"K Singhal","year":"2023","unstructured":"Singhal, K. et al. Large language models encode clinical knowledge. Nature 620, 172\u2013180 (2023).","journal-title":"Nature"},{"key":"2551_CR35","doi-asserted-by":"publisher","first-page":"943","DOI":"10.1038\/s41591-024-03423-7","volume":"31","author":"K Singhal","year":"2025","unstructured":"Singhal, K. et al. Towards expert-level medical question answering with large language models. Nat. Med. 31, 943\u2013950 (2025).","journal-title":"Nat. Med."},{"key":"2551_CR36","unstructured":"Saab, K. et al. Capabilities of gemini models in medicine. arXiv preprint arXiv:2404.18416 (2024)."},{"key":"2551_CR37","unstructured":"Lu, M. Y. et al. A multimodal generative ai copilot for human pathology. Nature 1\u20133 (2024)."},{"key":"2551_CR38","unstructured":"Xu, H. et al. A whole-slide foundation model for digital pathology from real-world data. Nature 1\u20138 (2024)."},{"key":"2551_CR39","doi-asserted-by":"publisher","DOI":"10.1038\/s41467-025-62385-7","volume":"16","author":"C Wu","year":"2025","unstructured":"Wu, C., Zhang, X., Zhang, Y., Wang, Y. & Xie, W. Towards generalist foundation model for radiology by leveraging web-scale 2D & 3D medical data. Nat. Commun. 16, 7866 (2025).","journal-title":"Nat. Commun."},{"key":"2551_CR40","unstructured":"Zhu, D., Chen, J., Shen, X., Li, X. & Elhoseiny, M. Minigpt-4: Enhancing vision-language understanding with advanced large language models. In Proc. Int. Conf. Learn. Represent. (2024)."},{"key":"2551_CR41","doi-asserted-by":"crossref","unstructured":"Dai, W. et al. InstructBLIP: Towards general-purpose vision-language models with instruction tuning. In Thirty-seventh Conference on Neural Information Processing Systems (2023). https:\/\/openreview.net\/forum?id=vvoWPYqZJA.","DOI":"10.52202\/075280-2142"},{"key":"2551_CR42","doi-asserted-by":"crossref","unstructured":"Wan, Z. et al. MEIT: Multimodal electrocardiogram instruction tuning on large language models for report generation. In Findings Assoc. Comput. Linguist. ACL 14510\u201314527 (2025).","DOI":"10.18653\/v1\/2025.findings-acl.749"},{"key":"2551_CR43","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1038\/s41597-020-0495-6","volume":"7","author":"P Wagner","year":"2020","unstructured":"Wagner, P. et al. Ptb-xl, a large publicly available electrocardiography dataset. Sci. Data 7, 1\u201315 (2020).","journal-title":"Sci. Data"},{"key":"2551_CR44","unstructured":"Stearns, M. Q., Price, C., Spackman, K. A. & Wang, A. Y. Snomed clinical terms: overview of the development process and project status. In Proceedings of the AMIA Symposium, 662 (American Medical Informatics Association, 2001)."},{"key":"2551_CR45","doi-asserted-by":"publisher","first-page":"055019","DOI":"10.1088\/1361-6579\/ad4954","volume":"45","author":"KK Shivashankara","year":"2024","unstructured":"Shivashankara, K. K. et al. Ecg-image-kit: a synthetic image generation toolbox to facilitate deep learning-based electrocardiogram digitization. Physiol. Meas. 45, 055019 (2024).","journal-title":"Physiol. Meas."},{"key":"2551_CR46","doi-asserted-by":"publisher","DOI":"10.1038\/s41597-023-02153-8","volume":"10","author":"N Strodthoff","year":"2023","unstructured":"Strodthoff, N. et al. Ptb-xl+, a comprehensive electrocardiographic feature dataset. Sci. Data 10, 279 (2023).","journal-title":"Sci. Data"},{"key":"2551_CR47","doi-asserted-by":"publisher","unstructured":"Gow, B. et al. Mimic-iv-ecg: Diagnostic electrocardiogram matched subset (2023). https:\/\/doi.org\/10.13026\/4nqg-sb35 (2023).","DOI":"10.13026\/4nqg-sb35"},{"key":"2551_CR48","doi-asserted-by":"publisher","DOI":"10.1038\/s41597-022-01899-x","volume":"10","author":"AE Johnson","year":"2023","unstructured":"Johnson, A. E. et al. Mimic-iv, a freely accessible electronic health record dataset. Sci. Data 10, 1 (2023).","journal-title":"Sci. Data"},{"key":"2551_CR49","first-page":"10","volume":"9","author":"AH Ribeiro","year":"2021","unstructured":"Ribeiro, A. H. et al. Code-15%: A large scale annotated dataset of 12-lead ecgs. Zenodo 9, 10\u20135281 (2021).","journal-title":"Zenodo"},{"key":"2551_CR50","doi-asserted-by":"publisher","first-page":"S75","DOI":"10.1016\/j.jelectrocard.2019.09.008","volume":"57","author":"ALP Ribeiro","year":"2019","unstructured":"Ribeiro, A. L. P. et al. Tele-electrocardiography and bigdata: the code (clinical outcomes in digital electrocardiography) study. J. Electrocardiol. 57, S75\u2013S78 (2019).","journal-title":"J. Electrocardiol."},{"key":"2551_CR51","first-page":"66277","volume":"36","author":"J Oh","year":"2023","unstructured":"Oh, J., Lee, G., Bae, S., Kwon, J.-M. & Choi, E. Ecg-qa: A comprehensive question answering dataset combined with electrocardiogram. Adv. Neural Inform. Process. Syst. 36, 66277\u201366288 (2023).","journal-title":"Adv. Neural Inform. Process. Syst."},{"key":"2551_CR52","unstructured":"Xu, C. et al. Wizardlm: Empowering large pre-trained language models to follow complex instructions. In The Twelfth International Conference on Learning Representations (2024)."},{"key":"2551_CR53","unstructured":"Chiang, W.-L. et al. Chatbot arena: An open platform for evaluating llms by human preference. In Proc. 41st Int. Conf. Mach. Learn. 8359\u20138388 (2024)."},{"key":"2551_CR54","unstructured":"Meta. Introducing meta llama 3: The most capable openly available llm to date (2024). https:\/\/ai.meta.com\/blog\/meta-llama-3\/. Accessed: 2024-10-01."},{"key":"2551_CR55","first-page":"1368","volume":"8","author":"F Liu","year":"2018","unstructured":"Liu, F. et al. An open access database for evaluating the algorithms of electrocardiogram rhythm and morphology abnormality detection. J. Med. Imaging Health Inform. 8, 1368\u20131373 (2018).","journal-title":"J. Med. Imaging Health Inform."},{"key":"2551_CR56","doi-asserted-by":"publisher","DOI":"10.1038\/s41598-020-59821-7","volume":"10","author":"J Zheng","year":"2020","unstructured":"Zheng, J. et al. Optimal multi-stage arrhythmia classification approach. Sci. Rep. 10, 2898 (2020).","journal-title":"Sci. Rep."},{"key":"2551_CR57","doi-asserted-by":"publisher","DOI":"10.1038\/s41597-020-0386-x","volume":"7","author":"J Zheng","year":"2020","unstructured":"Zheng, J. et al. A 12-lead electrocardiogram database for arrhythmia research covering more than 10,000 patients. Sci. data 7, 48 (2020).","journal-title":"Sci. data"},{"key":"2551_CR58","doi-asserted-by":"crossref","unstructured":"Yue, X. et al. Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 9556\u20139567 (2024).","DOI":"10.1109\/CVPR52733.2024.00913"},{"key":"2551_CR59","unstructured":"Ohio Supercomputer Center. Ohio supercomputer center (1987). https:\/\/ror.org\/01apna436 (1987)."}],"container-title":["npj Digital Medicine"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.nature.com\/articles\/s41746-026-02551-3","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.nature.com\/articles\/s41746-026-02551-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.nature.com\/articles\/s41746-026-02551-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,3]],"date-time":"2026-05-03T23:13:40Z","timestamp":1777850020000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.nature.com\/articles\/s41746-026-02551-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,16]]},"references-count":59,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2026,12]]}},"alternative-id":["2551"],"URL":"https:\/\/doi.org\/10.1038\/s41746-026-02551-3","relation":{},"ISSN":["2398-6352"],"issn-type":[{"value":"2398-6352","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,16]]},"assertion":[{"value":"9 July 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"5 March 2026","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"16 March 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The authors declare no competing interests.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"349"}}