{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T13:22:09Z","timestamp":1783084929327,"version":"3.54.6"},"reference-count":65,"publisher":"Springer Science and Business Media LLC","issue":"5","license":[{"start":{"date-parts":[[2026,3,26]],"date-time":"2026-03-26T00:00:00Z","timestamp":1774483200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T00:00:00Z","timestamp":1783036800000},"content-version":"vor","delay-in-days":99,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62361027"],"award-info":[{"award-number":["62361027"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62161011"],"award-info":[{"award-number":["62161011"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J. King Saud Univ. Comput. Inf. Sci."],"published-print":{"date-parts":[[2026,7]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    Histopathology in whole slide images (WSIs) serves as the gold standard for cancer diagnosis, with clinical reports playing a critical role in decision-making. However, the time-consuming nature of conventional pathological examination has driven increasing and urgent demand for automated report generation. Deep learning methods offer a certain potential to revolutionize this requirement by Histopathology Report Generation (HRG). Nevertheless, existing HRG approaches suffer from low-quality generation results due to ineffective exploration of multi-scale visual context in gigapixel WSIs and the inherent semantic gap between heterogeneous vision-language modalities. To address these challenges, we propose HC-Gen, a novel framework which synergistically combines hierarchical context modeling with prototype-mediate cross-modal alignment for HRG. Inspired by pathologists\u2019 anatomically-grounded diagnostic logic, we design a hierarchical context fusion module to integrate multi-scale visual-semantic context and implicit hierarchy prior in WSIs. Furthermore, we propose a cross-modal prototypical memory module to establish learnable semantic prototypes as intermediate bridges to achieve unified and efficient vision-language alignment. Model performance was assessed through natural language generation metrics and human evaluation, extensive experiments on two benchmark datasets demonstrate that HC-Gen outperforms state-of-the-art methods. Extra visualization provides crucial support for the interpretability of the decision process. Our code is available at:\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/github.com\/Modaoshuangming\/HC-Gen\" ext-link-type=\"uri\">https:\/\/github.com\/Modaoshuangming\/HC-Gen<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1007\/s44443-026-00672-z","type":"journal-article","created":{"date-parts":[[2026,3,26]],"date-time":"2026-03-26T17:15:31Z","timestamp":1774545331000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Hierarchical context modeling and prototype-mediated cross-modal alignment for histopathology report generation"],"prefix":"10.1007","volume":"38","author":[{"given":"Chengxin","family":"Ye","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jingqin","family":"Lv","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Guangli","family":"Li","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Renzhong","family":"Wu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shiying","family":"Zeng","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Nan","family":"Jiang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Boyang","family":"Liu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jianguo","family":"Wu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Donghong","family":"Ji","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hongbin","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,3,26]]},"reference":[{"key":"672_CR1","doi-asserted-by":"crossref","unstructured":"Ahmed F, Yang L, Jaroensri T, Sellergren A, Matias Y, Hassidim A, Corrado GS, Webster DR, Shetty S, Prabhakara S et al (2025) Polypath: Adapting a large multimodal model for multi-slide pathology report generation. arXiv:2502.10536","DOI":"10.1016\/j.modpat.2025.100886"},{"issue":"4","key":"672_CR2","doi-asserted-by":"publisher","first-page":"1412","DOI":"10.1109\/TMI.2023.3337549","volume":"43","author":"G Bontempo","year":"2023","unstructured":"Bontempo G, Bolelli F, Porrello A, Calderara S, Ficarra E (2023) A graph-based multi-scale approach with knowledge distillation for wsi classification. IEEE Trans Med Imaging 43(4):1412\u20131421","journal-title":"IEEE Trans Med Imaging"},{"key":"672_CR3","doi-asserted-by":"crossref","unstructured":"Bu S, Li T, Yang Y, Dai Z (2024) Instance-level expert knowledge and aggregate discriminative attention for radiology report generation. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 14194\u201314204","DOI":"10.1109\/CVPR52733.2024.01346"},{"key":"672_CR4","doi-asserted-by":"crossref","unstructured":"Chen RJ, Chen C, Li Y, Chen TY, Trister AD, Krishnan RG, Mahmood F (2022) Scaling vision transformers to gigapixel images via hierarchical self-supervised learning. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 16144\u201316155","DOI":"10.1109\/CVPR52688.2022.01567"},{"key":"672_CR5","doi-asserted-by":"crossref","unstructured":"Chen P, Li H, Zhu C, Zheng S, Shui Z, Yang L (2024) Wsicaption: Multiple instance generation of pathology reports for gigapixel whole-slide images. In: International conference on medical image computing and computer-assisted intervention. Springer, pp 546\u2013556","DOI":"10.1007\/978-3-031-72083-3_51"},{"key":"672_CR6","doi-asserted-by":"crossref","unstructured":"Chen X, Ma L, Jiang W, Yao J, Liu W (2018) Regularizing rnns for caption generation by reconstructing the past with the present. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 7995\u20138003","DOI":"10.1109\/CVPR.2018.00834"},{"key":"672_CR7","doi-asserted-by":"crossref","unstructured":"Chen Z, Shen Y, Song Y, Wan X (2021) Cross-modal memory networks for radiology report generation. In: Proceedings of the 59th annual meeting of the association for computational linguistics and the 11th international joint conference on natural language processing (Volume 1: Long Papers), pp 5904\u20135914","DOI":"10.18653\/v1\/2021.acl-long.459"},{"key":"672_CR8","doi-asserted-by":"crossref","unstructured":"Chen Z, Song Y, Chang T-H, Wan X (2020) Generating radiology reports via memory-driven transformer. In: Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP), pp 1439\u20131449","DOI":"10.18653\/v1\/2020.emnlp-main.112"},{"key":"672_CR9","doi-asserted-by":"crossref","unstructured":"Chen Y, Wang G, Ji Y, Li Y, Ye J, Li T, Hu M, Yu R, Qiao Y, He J (2025) Slidechat: A large vision-language assistant for whole-slide pathology image understanding. In: Proceedings of the computer vision and pattern recognition conference, pp 5134\u20135143","DOI":"10.1109\/CVPR52734.2025.00484"},{"issue":"3","key":"672_CR10","doi-asserted-by":"publisher","first-page":"850","DOI":"10.1038\/s41591-024-02857-3","volume":"30","author":"RJ Chen","year":"2024","unstructured":"Chen RJ, Ding T, Lu MY, Williamson DF, Jaume G, Song AH, Chen B, Zhang A, Shao D, Shaban M et al (2024) Towards a general-purpose foundation model for computational pathology. Nat Med 30(3):850\u2013862","journal-title":"Nat Med"},{"key":"672_CR11","doi-asserted-by":"crossref","unstructured":"Cornia M, Stefanini M, Baraldi L, Cucchiara R (2020) Meshed-memory transformer for image captioning. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 10578\u201310587","DOI":"10.1109\/CVPR42600.2020.01059"},{"key":"672_CR12","doi-asserted-by":"crossref","unstructured":"Ding J, Ma S, Dong L, Zhang X, Huang S, Wang W, Zheng N, Wei F (2023) Longnet: Scaling transformers to 1,000,000,000 tokens. arXiv:2307.02486","DOI":"10.14218\/JCTH.2024.00317"},{"key":"672_CR13","unstructured":"Dosovitskiy A, Beyer L, Kolesnikov A, Weissenborn D, Zhai X, Unterthiner T, Dehghani M, Minderer M, Heigold G, Gelly S et al (2020) An image is worth 16x16 words: Transformers for image recognition at scale. arXiv:2010.11929"},{"key":"672_CR14","doi-asserted-by":"crossref","unstructured":"Fei M, Song Z, Shen Z, Liu M, Wang Q, Zhang L (2025) Weakly semi-supervised cervical lesion cell detection via twin-memory augmented multiple instance learning. In: International conference on medical image computing and computer-assisted intervention. Springer, pp 637\u2013647","DOI":"10.1007\/978-3-032-04984-1_61"},{"issue":"3","key":"672_CR15","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3617592","volume":"56","author":"T Ghandi","year":"2023","unstructured":"Ghandi T, Pourreza H, Mahyar H (2023) Deep learning approaches on image captioning: A review. ACM Comput Surv 56(3):1\u201339","journal-title":"ACM Comput Surv"},{"key":"672_CR16","unstructured":"Guevara BC, Marini N, Marchesin S, Aswolinskiy W, Schlimbach R-J, Podareanu D, Ciompi F (2023) Caption generation from histopathology whole-slide images using pre-trained transformers. In: Medical imaging with deep learning, short paper track"},{"key":"672_CR17","doi-asserted-by":"crossref","unstructured":"Guo Z, Ma J, Xu Y, Wang Y, Wang L, Chen H (2024) Histgen: Histopathology report generation via local-global feature encoding and cross-modal context interaction. In: International conference on medical image computing and computer-assisted intervention. Springer, pp 189\u2013199","DOI":"10.1007\/978-3-031-72083-3_18"},{"key":"672_CR18","doi-asserted-by":"crossref","unstructured":"He K, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 770\u2013778","DOI":"10.1109\/CVPR.2016.90"},{"issue":"1","key":"672_CR19","doi-asserted-by":"publisher","first-page":"175","DOI":"10.1049\/ipr2.12941","volume":"18","author":"K He","year":"2024","unstructured":"He K, Zhu J, Li L (2024) Two-stage coarse-to-fine method for pathological images in medical decision-making systems. IET Image Process 18(1):175\u2013193","journal-title":"IET Image Process"},{"issue":"9","key":"672_CR20","doi-asserted-by":"publisher","first-page":"2307","DOI":"10.1038\/s41591-023-02504-3","volume":"29","author":"Z Huang","year":"2023","unstructured":"Huang Z, Bianchi F, Yuksekgonul M, Montine TJ, Zou J (2023) A visual-language foundation model for pathology image analysis using medical twitter. Nat Med 29(9):2307\u20132316","journal-title":"Nat Med"},{"key":"672_CR21","doi-asserted-by":"publisher","DOI":"10.1016\/j.media.2025.103468","volume":"101","author":"H Jin","year":"2025","unstructured":"Jin H, Shen J, Cui L, Shi X, Li K, Zhu X (2025) Dynamic graph based weakly supervised deep hashing for whole slide image classification and retrieval. Med Image Anal 101:103468","journal-title":"Med Image Anal"},{"key":"672_CR22","doi-asserted-by":"crossref","unstructured":"Jing B, Xie P, Xing E (2018) On the automatic generation of medical imaging reports. In: Proceedings of the 56th annual meeting of the association for computational linguistics (volume 1: Long Papers), pp 2577\u20132586","DOI":"10.18653\/v1\/P18-1240"},{"key":"672_CR23","doi-asserted-by":"crossref","unstructured":"Li G, Huang C, Zhou X, Ji D, Zhang H (2025) Report is a mixture of topics: Topic-guided radiology report generation. Med Image Anal 103586","DOI":"10.1016\/j.media.2025.103586"},{"key":"672_CR24","doi-asserted-by":"crossref","unstructured":"Li B, Li Y, Eliceiri KW (2021) Dual-stream multiple instance learning network for whole slide image classification with self-supervised contrastive learning. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 14318\u201314328","DOI":"10.1109\/CVPR46437.2021.01409"},{"key":"672_CR25","unstructured":"Li J, Li D, Savarese S, Hoi S (2023) Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In: International conference on machine learning. PMLR, pp 19730\u201319742"},{"key":"672_CR26","doi-asserted-by":"crossref","unstructured":"Li CY, Liang X, Hu Z, Xing EP (2019) Knowledge-driven encode, retrieve, paraphrase for medical image report generation. In: Proceedings of the AAAI conference on artificial intelligence, vol 33, pp 6666\u20136673","DOI":"10.1609\/aaai.v33i01.33016666"},{"key":"672_CR27","doi-asserted-by":"crossref","unstructured":"Li M, Lin B, Chen Z, Lin H, Liang X, Chang X (2023) Dynamic graph enhanced contrastive learning for chest x-ray report generation. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 3334\u20133343","DOI":"10.1109\/CVPR52729.2023.00325"},{"issue":"6","key":"672_CR28","doi-asserted-by":"publisher","first-page":"4809","DOI":"10.1007\/s10462-021-10121-0","volume":"55","author":"X Li","year":"2022","unstructured":"Li X, Li C, Rahaman MM, Sun H, Li X, Wu J, Yao Y, Grzegorzek M (2022) A comprehensive review of computer-aided whole-slide image analysis: from datasets to feature extraction, segmentation, classification and detection approaches. Artif Intell Rev 55(6):4809\u20134878","journal-title":"Artif Intell Rev"},{"key":"672_CR29","doi-asserted-by":"crossref","unstructured":"Liang Y, Lyu X, Chen W, Ding M, Zhang J, He X, Wu S, Xing X, Yang S, Wang X et al (2025) Wsi-llava: A multimodal large language model for whole slide image. In: Proceedings of the IEEE\/CVF international conference on computer vision, pp 22718\u201322727","DOI":"10.1109\/ICCV51701.2025.02109"},{"key":"672_CR30","doi-asserted-by":"crossref","unstructured":"Liu H, Li C, Li Y, Lee YJ (2024) Improved baselines with visual instruction tuning. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 26296\u201326306","DOI":"10.1109\/CVPR52733.2024.02484"},{"key":"672_CR31","doi-asserted-by":"crossref","unstructured":"Liu F, Wu X, Ge S, Fan W, Zou Y (2021) Exploring and distilling posterior and prior knowledge for radiology report generation. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 13753\u201313762","DOI":"10.1109\/CVPR46437.2021.01354"},{"issue":"6","key":"672_CR32","doi-asserted-by":"publisher","first-page":"555","DOI":"10.1038\/s41551-020-00682-w","volume":"5","author":"MY Lu","year":"2021","unstructured":"Lu MY, Williamson DF, Chen TY, Chen RJ, Barbieri M, Mahmood F (2021) Data-efficient and weakly supervised computational pathology on whole-slide images. Nature Biomed Eng 5(6):555\u2013570","journal-title":"Nature Biomed Eng"},{"issue":"3","key":"672_CR33","doi-asserted-by":"publisher","first-page":"863","DOI":"10.1038\/s41591-024-02856-4","volume":"30","author":"MY Lu","year":"2024","unstructured":"Lu MY, Chen B, Williamson DF, Chen RJ, Liang I, Ding T, Jaume G, Odintsov I, Le LP, Gerber G et al (2024) A visual-language foundation model for computational pathology. Nat Med 30(3):863\u2013874","journal-title":"Nat Med"},{"key":"672_CR34","doi-asserted-by":"crossref","unstructured":"Matsui I, Matsumoto A, Imai A, Okushima H, Niioka H, Abe M, Tamai N, Nagasu H, Kanda E, Uchino E et al (2025) Domain-adaptive semi-supervised learning for efficient rare pathological lesion detection with minimal annotation. NPJ Digit Med","DOI":"10.1038\/s41746-025-02160-6"},{"key":"672_CR35","unstructured":"Mokady R, Hertz A, Bermano AH (2021) Clipcap: Clip prefix for image captioning. arXiv:2111.09734"},{"key":"672_CR36","unstructured":"Oquab M, Darcet T, Moutakanni T, Vo H, Szafraniec M, Khalidov V, Fernandez P, Haziza D, Massa F, El-Nouby A et al (2024) Dinov2: Learning robust visual features without supervision. Trans Mach Learn Res J 1\u201331"},{"key":"672_CR37","doi-asserted-by":"crossref","unstructured":"Pham V, Bluche T, Kermorvant C, Louradour J (2014) Dropout improves recurrent neural networks for handwriting recognition. In: 2014 14th International conference on frontiers in handwriting recognition. IEEE, pp 285\u2013290","DOI":"10.1109\/ICFHR.2014.55"},{"key":"672_CR38","unstructured":"Qin W, Xu R, Huang P, Wu X, Zhang H, Luo L (2023) What a whole slide image can tell? subtype-guided masked transformer for pathological image captioning. arXiv:2310.20607"},{"key":"672_CR39","doi-asserted-by":"crossref","unstructured":"Qin W, Xu R, Jiang S, Jiang T, Luo L (2022) Pathtr: Context-aware memory transformer for tumor localization in gigapixel pathology images. In: Proceedings of the asian conference on computer vision, pp 3603\u20133619","DOI":"10.1007\/978-3-031-26351-4_8"},{"key":"672_CR40","unstructured":"Qwen, Yang A, Yang B, Zhang B, Hui B, Zheng B, Yu B, Li C, Liu D, Huang F, Wei H, Lin H, Yang J, Tu J, Zhang J, Yang J, Yang J, Zhou J, Lin J, Dang K, Lu K, Bao K, Yang K, Yu L, Li M, Xue M, Zhang P, Zhu Q, Men R, Lin R, Li T, Tang T, Xia T, Ren X, Ren X, Fan Y, Su Y, Zhang Y, Wan Y, Liu Y, Cui Z, Zhang Z, Qiu Z (2025) Qwen2.5 technical report. arxiv:2412.15115"},{"key":"672_CR41","unstructured":"Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, Agarwal S, Sastry G, Askell A, Mishkin P, Clark J et al (2021) Learning transferable visual models from natural language supervision. In: International conference on machine learning. PMLR, pp 8748\u20138763"},{"key":"672_CR42","doi-asserted-by":"crossref","unstructured":"Sengupta S, Brown DE (2024) Automatic report generation for histopathology images using pre-trained vision transformers and bert. In: 2024 IEEE international symposium on biomedical imaging (ISBI). IEEE, pp 1\u20135","DOI":"10.1109\/ISBI56570.2024.10635175"},{"key":"672_CR43","unstructured":"Shaikovski G, Casson A, Severson K, Zimmermann E, Wang YK, Kunz JD, Retamero JA, Oakley G, Klimstra D, Kanan C et al (2024) Prism: A multi-modal generative foundation model for slide-level histopathology. arXiv:2405.10254"},{"issue":"5","key":"672_CR44","doi-asserted-by":"publisher","first-page":"71","DOI":"10.1007\/s44443-025-00039-w","volume":"37","author":"MWU Sima","year":"2025","unstructured":"Sima MWU, Wang C, Arshad M, Shaikh JA, Alkanhel RI, Hassan DS, Muthanna A (2025) Improving kidney segmentation in pathological images: a multiscale approach to resolve fragmentation and incomplete boundaries. J King Saud Univ Comput Inf Sci 37(5):71","journal-title":"J King Saud Univ Comput Inf Sci"},{"issue":"1","key":"672_CR45","doi-asserted-by":"publisher","first-page":"539","DOI":"10.1109\/TPAMI.2022.3148210","volume":"45","author":"M Stefanini","year":"2022","unstructured":"Stefanini M, Cornia M, Baraldi L, Cascianelli S, Fiameni G, Cucchiara R (2022) From show to tell: A survey on deep learning-based image captioning. IEEE Trans Pattern Anal Mach Intell 45(1):539\u2013559","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"672_CR46","doi-asserted-by":"crossref","unstructured":"Sun L, Zhao JJ, Han W, Xiong C (2025) Fact-aware multimodal retrieval augmentation for accurate medical radiology report generation. In: Proceedings of the 2025 conference of the nations of the Americas chapter of the association for computational linguistics: Human language technologies (Volume 1: Long Papers), pp 643\u2013655","DOI":"10.18653\/v1\/2025.naacl-long.28"},{"key":"672_CR47","doi-asserted-by":"crossref","unstructured":"Tan JW, Kim S, Kim E, Lee SH, Ahn S, Jeong W-K (2024) Clinical-grade multi-organ pathology report generation for multi-scale whole slide images via a semantically guided medical text foundation model. In: International conference on medical image computing and computer-assisted intervention. Springer, pp 25\u201335","DOI":"10.1007\/978-3-031-72083-3_3"},{"key":"672_CR48","doi-asserted-by":"crossref","unstructured":"Tanida T, M\u00fcller P, Kaissis G, Rueckert D (2023) Interactive and explainable region-guided radiology report generation. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 7433\u20137442","DOI":"10.1109\/CVPR52729.2023.00718"},{"issue":"2","key":"672_CR49","doi-asserted-by":"publisher","first-page":"599","DOI":"10.1038\/s41591-024-03302-1","volume":"31","author":"R Tanno","year":"2025","unstructured":"Tanno R, Barrett DG, Sellergren A, Ghaisas S, Dathathri S, See A, Welbl J, Lau C, Tu T, Azizi S et al (2025) Collaboration between clinicians and vision-language models in radiology report generation. Nat Med 31(2):599\u2013608","journal-title":"Nat Med"},{"key":"672_CR50","unstructured":"Tsuneki M, Kanavati F (2022) Inference of captions from histopathological patches. In: International conference on medical imaging with deep learning. PMLR, pp 1235\u20131250"},{"key":"672_CR51","unstructured":"Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser \u0141, Polosukhin I (2017) Attention is all you need. Adv Neural Inf Process Syst 30"},{"key":"672_CR52","doi-asserted-by":"crossref","unstructured":"Vinyals O, Toshev A, Bengio S, Erhan D (2015) Show and tell: A neural image caption generator. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 3156\u20133164","DOI":"10.1109\/CVPR.2015.7298935"},{"issue":"10","key":"672_CR53","doi-asserted-by":"publisher","first-page":"2803","DOI":"10.1109\/TMI.2022.3171661","volume":"41","author":"Z Wang","year":"2022","unstructured":"Wang Z, Han H, Wang L, Li X, Zhou L (2022) Automated radiographic report generation purely on transformer: A multicriteria supervised approach. IEEE Trans Med Imaging 41(10):2803\u20132813","journal-title":"IEEE Trans Med Imaging"},{"issue":"6","key":"672_CR54","doi-asserted-by":"publisher","first-page":"4651","DOI":"10.1007\/s00521-024-10884-x","volume":"37","author":"Y Wang","year":"2025","unstructured":"Wang Y, Li D, Li X, Guo Y, Zuo Y, Shen L (2025) Mlpformer: Mlp-integrated transformer for colorectal histopathology whole slide image segmentation. Neural Comput Appl 37(6):4651\u20134661","journal-title":"Neural Comput Appl"},{"key":"672_CR55","doi-asserted-by":"crossref","unstructured":"Xiang J, Wang X, Zhang X, Xi Y, Eweje F, Chen Y, Li Y, Bergstrom C, Gopaulchan M, Kim T et al (2025) A vision-language foundation model for precision oncology. Nature 1\u201310","DOI":"10.1038\/s41586-024-08378-w"},{"key":"672_CR56","doi-asserted-by":"crossref","unstructured":"Xu Q, Adam A, Abdullah A, Bariyah N (2025) A review of advanced deep learning methods of multi-target segmentation for breast cancer wsis. IEEE Access","DOI":"10.1109\/ACCESS.2025.3565648"},{"key":"672_CR57","doi-asserted-by":"crossref","unstructured":"Xu Y, Wang Y, Zhou F, Ma J, Jin C, Yang S, Li J, Zhang Z, Zhao C, Zhou H et al (2024) A multimodal knowledge-enhanced whole-slide pathology foundation model. arXiv:2407.15362","DOI":"10.1038\/s41467-025-66220-x"},{"issue":"8015","key":"672_CR58","doi-asserted-by":"publisher","first-page":"181","DOI":"10.1038\/s41586-024-07441-w","volume":"630","author":"H Xu","year":"2024","unstructured":"Xu H, Usuyama N, Bagga J, Zhang S, Rao R, Naumann T, Wong C, Gero Z, Gonz\u00e1lez J, Gu Y et al (2024) A whole-slide foundation model for digital pathology from real-world data. Nature 630(8015):181\u2013188","journal-title":"Nature"},{"key":"672_CR59","doi-asserted-by":"crossref","unstructured":"Zhang J, Nguyen AT, Han X, Trinh VQ-H, Qin H, Samaras D, Hosseini MS (2025) 2dmamba: Efficient state space model for image representation with applications on giga-pixel whole slide image classification. In: Proceedings of the computer vision and pattern recognition conference, pp 3583\u20133592","DOI":"10.1109\/CVPR52734.2025.00339"},{"key":"672_CR60","doi-asserted-by":"crossref","unstructured":"Zhang Y, Wang X, Xu Z, Yu Q, Yuille A, Xu D (2020) When radiology report generation meets knowledge graph. In: Proceedings of the AAAI conference on artificial intelligence, vol 34, pp 12910\u201312917","DOI":"10.1609\/aaai.v34i07.6989"},{"issue":"5","key":"672_CR61","doi-asserted-by":"publisher","first-page":"236","DOI":"10.1038\/s42256-019-0052-1","volume":"1","author":"Z Zhang","year":"2019","unstructured":"Zhang Z, Chen P, McGough M, Xing F, Wang C, Bui M, Xie Y, Sapkota M, Cui L, Dhillon J et al (2019) Pathologist-level interpretable whole-slide cancer diagnosis with deep learning. Nat Mach Intell 1(5):236\u2013245","journal-title":"Nat Mach Intell"},{"issue":"8","key":"672_CR62","doi-asserted-by":"publisher","first-page":"5689","DOI":"10.1007\/s00371-024-03746-z","volume":"41","author":"X Zhang","year":"2025","unstructured":"Zhang X, Yan B, Xing Z, Gao F, Tao Y, Han Z, Wang W, Zhu L (2025) Hadiff: hierarchy aggregated diffusion model for pathology image segmentation. Vis Comput 41(8):5689\u20135700","journal-title":"Vis Comput"},{"key":"672_CR63","doi-asserted-by":"crossref","unstructured":"Zhao J, Li X, Yang F, Zhai Q, Luo A, Zhao Y, Cheng H, Fu H (2025) Mexd: An expert-infused diffusion model for whole-slide image classification. In: Proceedings of the computer vision and pattern recognition conference, pp 20789\u201320799","DOI":"10.1109\/CVPR52734.2025.01936"},{"key":"672_CR64","doi-asserted-by":"crossref","unstructured":"Zheng T, Jiang K, Xiao Y, Zhao S, Yao H (2025) M3amba: Memory mamba is all you need for whole slide image classification. In: Proceedings of the computer vision and pattern recognition conference, pp 15601\u201315610","DOI":"10.1109\/CVPR52734.2025.01454"},{"key":"672_CR65","doi-asserted-by":"crossref","unstructured":"Zhou Q, Zhong W, Guo Y, Xiao M, Ma H, Huang J (2024) Pathm3: A multimodal multi-task multiple instance learning framework for whole slide image classification and captioning. In: International conference on medical image computing and computer-assisted intervention. Springer, pp 373\u2013383","DOI":"10.1007\/978-3-031-72083-3_35"}],"container-title":["Journal of King Saud University Computer and Information Sciences"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s44443-026-00672-z","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44443-026-00672-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44443-026-00672-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T12:54:01Z","timestamp":1783083241000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s44443-026-00672-z"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,26]]},"references-count":65,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2026,7]]}},"alternative-id":["672"],"URL":"https:\/\/doi.org\/10.1007\/s44443-026-00672-z","relation":{},"ISSN":["1319-1578","2213-1248"],"issn-type":[{"value":"1319-1578","type":"print"},{"value":"2213-1248","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,26]]},"assertion":[{"value":"23 January 2026","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 March 2026","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"26 March 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing Interests"}}],"article-number":"274"}}