{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,28]],"date-time":"2026-08-28T14:04:27Z","timestamp":1787925867173,"version":"build-2784847793"},"reference-count":217,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2026,5,29]],"date-time":"2026-05-29T00:00:00Z","timestamp":1780012800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Computation"],"abstract":"<jats:p>This survey provides a comprehensive guide to Multimodal Large Language Models (MLLMs) with a focus on vision\u2013language tasks, including image captioning, visual question answering, cross-modal retrieval, visual grounding, multi-image reasoning, long-video understanding, and embodied AI. We examine architectures, training pipelines, and practical applications, covering visual encoders, language model backbones, connector modules, contrastive pre-training, instruction tuning, and preference alignment. We also foreground first-principles constraints\u2014information bottlenecks, data-processing limits, and statistical co-occurrence bias\u2014that shape architecture, robustness, and evaluation. This survey centers on vision\u2013language systems and does not cover audio-only models or code-generation tools without visual inputs. Through task-level analysis and system-level case studies, we examine prominent MLLM implementations while addressing key challenges in scalability, memory, energy use, inference cost, robustness, and cross-modal learning. We present a unified taxonomy of the MLLM design space, a comparative overview of representative models and evaluation benchmarks, and a discussion of open problems. Concluding with ethical considerations and responsible AI development, this survey offers theoretical frameworks and practical insights for researchers, practitioners, and students working at the intersection of natural language processing and computer vision.<\/jats:p>","DOI":"10.3390\/computation14060125","type":"journal-article","created":{"date-parts":[[2026,5,29]],"date-time":"2026-05-29T10:06:56Z","timestamp":1780049216000},"page":"125","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["A Comprehensive Survey and Guide to Multimodal Large Language Models in Vision\u2013Language Tasks"],"prefix":"10.3390","volume":"14","author":[{"given":"Chia Xin","family":"Liang","sequence":"first","affiliation":[{"name":"JTB Technology CO., Ltd., Tainan 701020, Taiwan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Pu","family":"Tian","sequence":"additional","affiliation":[{"name":"Computer Science Program, School of Business, Stockton University, Galloway, NJ 08205, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1950-0737","authenticated-orcid":false,"given":"Caitlyn Heqi","family":"Yin","sequence":"additional","affiliation":[{"name":"Department of Computer Sciences, School of Computer, Data & Information Sciences, University of Wisconsin\u2013Madison, Madison, WI 53706, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yao","family":"Yua","sequence":"additional","affiliation":[{"name":"AppCubic, Miami, FL 33138, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"An-Hou","family":"Wei","sequence":"additional","affiliation":[{"name":"Nomad Sustaintech Ltd., Auckland 2019, New Zealand"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ming","family":"Li","sequence":"additional","affiliation":[{"name":"College of Computing, Georgia Institute of Technology, Atlanta, GA 30332, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-4209-5671","authenticated-orcid":false,"given":"Xinyuan","family":"Song","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Emory University, Atlanta, GA 30322, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tianyang","family":"Wang","sequence":"additional","affiliation":[{"name":"Department of Computer Science, University of Liverpool, Liverpool L69 3BX, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ziqian","family":"Bi","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Luddy School of Informatics, Computing, and Engineering, Indiana University Bloomington, Bloomington, IN 47408, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ming","family":"Liu","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Purdue University, West Lafayette, IN 47907, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Riyang","family":"Bao","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Emory University, Atlanta, GA 30322, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-6774-4296","authenticated-orcid":false,"given":"Pengbin","family":"Feng","sequence":"additional","affiliation":[{"name":"Department of Mathematics, University of Southern California, Los Angeles, CA 90089, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2026,5,29]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"nwae403","DOI":"10.1093\/nsr\/nwae403","article-title":"A survey on multimodal large language models","volume":"11","author":"Yin","year":"2024","journal-title":"Natl. Sci. Rev."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Zhang, D., Yu, Y., Li, C., Dong, J., Su, D., Chu, C., and Yu, D. (2024). MM-LLMs: Recent Advances in MultiModal Large Language Models. arXiv.","DOI":"10.18653\/v1\/2024.findings-acl.738"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Caffagni, D., Cocchi, F., Barsellotti, L., Moratelli, N., Sarto, S., Baraldi, L., Corsini, M., and Cucchiara, R. (2024). The Revolution of Multimodal Large Language Models: A Survey. arXiv.","DOI":"10.18653\/v1\/2024.findings-acl.807"},{"key":"ref_4","unstructured":"Xu, K., Ba, J.L., Kiros, R., Cho, K., Courville, A., Salakhutdinov, R., Zemel, R., and Bengio, Y. (2015, January 6\u201311). Show, Attend and Tell: Neural Image Caption Generation with Visual Attention. Proceedings of the International Conference on Machine Learning (ICML), Lille, France."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C.L., and Parikh, D. (2015, January 7\u201313). VQA: Visual Question Answering. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.279"},{"key":"ref_6","unstructured":"Lu, J., Batra, D., Parikh, D., and Lee, S. (2019, January 8\u201314). ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Vancouver, BC, Canada."},{"key":"ref_7","unstructured":"Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., and Clark, J. (2021, January 18\u201324). Learning Transferable Visual Models from Natural Language Supervision. Proceedings of the International Conference on Machine Learning (ICML), Virtual."},{"key":"ref_8","unstructured":"Tong, S., Brown, E., Wu, P., Woo, S., Middepogu, M., Akula, S.C., Yang, J., Yang, S., Iyer, A., and Pan, X. (2024). Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs. arXiv."},{"key":"ref_9","unstructured":"Wang, W., Lv, Q., Yu, W., Hong, W., Qi, J., Wang, Y., Ji, J., Yang, Z., Zhao, L., and Song, X. (2024). CogVLM: Visual Expert for Pretrained Language Models. arXiv."},{"key":"ref_10","unstructured":"Yao, Y., Yu, T., Zhang, A., Wang, C., Cui, J., Zhu, H., Cai, T., Li, H., Zhao, W., and He, Z. (2024). MiniCPM-V: A GPT-4V Level MLLM on Your Phone. arXiv."},{"key":"ref_11","unstructured":"01.AI, Young, A., Chen, B., Li, C., Huang, C., Zhang, G., Zhang, G., Li, H., Zhu, J., and Chen, J. (2024). Yi: Open Foundation Models by 01.AI. Includes Yi-VL vision-language variants. arXiv."},{"key":"ref_12","unstructured":"Polanyi, L., Culy, C., Van Den Berg, M., Thione, G.L., and Ahn, D. (May, January 30). A Rule Based Approach to Discourse Parsing. Proceedings of the 5th SIGdial Workshop on Discourse and Dialogue at HLT-NAACL 2004, Cambridge, MA, USA."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Koehn, P. (2009). Statistical Machine Translation, Cambridge University Press.","DOI":"10.1017\/CBO9780511815829"},{"key":"ref_14","unstructured":"Mikolov, T., Chen, K., Corrado, G., and Dean, J. (2013). Efficient Estimation of Word Representations in Vector Space. arXiv."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Pennington, J., Socher, R., and Manning, C.D. (2014, January 25\u201329). GloVe: Global Vectors for Word Representation. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar.","DOI":"10.3115\/v1\/D14-1162"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long Short-Term Memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Comput."},{"key":"ref_17","unstructured":"Bahdanau, D., Cho, K., and Bengio, Y. (2015, January 7\u20139). Neural Machine Translation by Jointly Learning to Align and Translate. Proceedings of the International Conference on Learning Representations (ICLR), San Diego, CA, USA."},{"key":"ref_18","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017, January 4\u20139). Attention is all you need. Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA."},{"key":"ref_19","unstructured":"Tishby, N., Pereira, F.C., and Bialek, W. (2000). The Information Bottleneck Method. arXiv."},{"key":"ref_20","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2021, January 3\u20137). An image is worth 16 \u00d7 16 words: Transformers for image recognition at scale. Proceedings of the International Conference on Learning Representations, Virtual."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"23716","DOI":"10.52202\/068431-1723","article-title":"Flamingo: A Visual Language Model for Few-Shot Learning","volume":"35","author":"Alayrac","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_22","unstructured":"Li, J., Li, D., Savarese, S., and Hoi, S. (2023, January 23\u201329). BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models. Proceedings of the International Conference on Machine Learning (ICML), Honolulu, HI, USA."},{"key":"ref_23","first-page":"34892","article-title":"Visual Instruction Tuning","volume":"36","author":"Liu","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D. (2017, January 21\u201326). Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.670"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Hudson, D.A., and Manning, C.D. (2019, January 16\u201320). GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00686"},{"key":"ref_26","unstructured":"Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., Parikh, D., and Rohrbach, M. (2019). TextVQA: Towards Reasoning about Text in Images. arXiv."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Liu, Y., Duan, H., Zhang, Y., Li, B., Zhang, S., Zhao, W., Yuan, Y., Wang, J., He, C., and Liu, Z. (2024). MMBench: Is Your Multi-modal Model an All-around Player?. arXiv.","DOI":"10.1007\/978-3-031-72658-3_13"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Yue, X., Ni, Y., Zhang, K., Zheng, T., Liu, R., Zhang, G., Stevens, S., Jiang, D., Ren, W., and Sun, Y. (2024). MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI. arXiv.","DOI":"10.1109\/CVPR52733.2024.00913"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Dai, W., Li, J., Li, D., Tiong, A.M.H., Zhao, J., Wang, W., Li, B., Fung, P., and Hoi, S. (2023). InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning. arXiv.","DOI":"10.52202\/075280-2142"},{"key":"ref_30","unstructured":"Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M. (2023). MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models. arXiv."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Liu, H., Li, C., Li, Y., and Lee, Y.J. (2024). Improved Baselines with Visual Instruction Tuning. arXiv.","DOI":"10.1109\/CVPR52733.2024.02484"},{"key":"ref_32","unstructured":"Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., and Zhou, J. (2023). Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond. arXiv."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Chen, Z., Wang, W., Tian, H., Ye, S., Gao, Z., Cui, E., Tong, W., Hu, K., Luo, J., and Ma, Z. (2024). How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites. arXiv.","DOI":"10.1007\/s11432-024-4231-5"},{"key":"ref_34","unstructured":"Lu, H., Liu, W., Zhang, B., Wang, B., Dong, K., Liu, B., Sun, J., Ren, T., Li, Z., and Sun, Y. (2024). DeepSeek-VL: Towards Real-World Vision-Language Understanding. arXiv."},{"key":"ref_35","unstructured":"Li, B., Zhang, Y., Guo, D., Zhang, R., Li, F., Zhang, H., Zhang, K., Li, Y., Liu, Z., and Li, C. (2024). LLaVA-OneVision: Easy Visual Task Transfer. arXiv."},{"key":"ref_36","unstructured":"Wang, P., Bai, S., Tan, S., Wang, S., Fan, Z., Bai, J., Chen, K., Liu, X., Wang, J., and Ge, W. (2024). Qwen2-VL: Enhancing Vision-Language Model\u2019s Perception of the World at Any Resolution. arXiv."},{"key":"ref_37","unstructured":"OpenGVLab Team (2026, May 10). InternVL2: Better than the Best\u2014Expanding Performance Boundaries of Open-Source Multimodal Models with the Progressive Scaling Strategy. Available online: https:\/\/internvl.github.io\/blog\/2024-07-02-InternVL-2.0\/."},{"key":"ref_38","unstructured":"Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F.L., Almeida, D., Altenschmidt, J., Altman, S., and Anadkat, S. (2023). GPT-4 Technical Report. arXiv."},{"key":"ref_39","unstructured":"OpenAI, Hurst, A., Lerer, A., Goucher, A.P., Perelman, A., Ramesh, A., Clark, A., Ostrow, A.J., Welihinda, A., and Hayes, A. (2024). GPT-4o System Card. arXiv."},{"key":"ref_40","unstructured":"Gemini Team, Anil, R., Borgeaud, S., Alayrac, J.B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A.M., Hauth, A., and Millican, K. (2023). Gemini: A Family of Highly Capable Multimodal Models. arXiv."},{"key":"ref_41","unstructured":"Anthropic (2024, September 29). Claude. Available online: https:\/\/www.anthropic.com\/."},{"key":"ref_42","unstructured":"Yan, S., Zhu, T., Wang, Z., Cao, Y., Zhang, M., Ghosh, S., Wu, Y., and Yu, J. (2022). VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners. arXiv."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Zhang, H., Li, X., and Bing, L. (2023). Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding. arXiv.","DOI":"10.18653\/v1\/2023.emnlp-demo.49"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Lin, B., Ye, Y., Zhu, B., Cui, J., Ning, M., Jin, P., and Yuan, L. (2023). Video-LLaVA: Learning United Visual Representation by Alignment Before Projection. arXiv.","DOI":"10.18653\/v1\/2024.emnlp-main.342"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Kahatapitiya, K., Ranasinghe, K., Park, J., and Ryoo, M.S. (2025). Language Repository for Long Video Understanding. Proceedings of the Findings of the Association for Computational Linguistics: ACL 2025, Vienna, Austria, 27 July\u20131 August 2025, Association for Computational Linguistics.","DOI":"10.18653\/v1\/2025.findings-acl.294"},{"key":"ref_46","unstructured":"Wu, H., Li, D., Chen, B., and Li, J. (2024, January 10\u201315). LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding. Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada."},{"key":"ref_47","unstructured":"Yu, L., Shi, B., Pasunuru, R., Muller, B., Golovneva, O., Wang, T., Babu, A., Tang, B., Karrer, B., and Sheynin, S. (2023). Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning. arXiv."},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Xu, Z., Shen, Y., and Huang, L. (2023, January 9\u201314). MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL), Toronto, ON, Canada.","DOI":"10.18653\/v1\/2023.acl-long.641"},{"key":"ref_49","unstructured":"Li, B., Zhang, Y., Chen, L., Wang, J., Pu, F., Cahyono, J.A., Yang, J., and Liu, Z. (2023). Otter: A Multi-Modal Model with In-Context Instruction Tuning. arXiv."},{"key":"ref_50","unstructured":"Wang, J., Jiang, H., Liu, Y., Ma, C., Zhang, X., Pan, Y., Liu, M., Gu, P., Xia, S., and Li, W. (2024). A comprehensive review of multimodal large language models: Performance and challenges across different tasks. arXiv."},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Jiang, C., Xu, H., Dong, M., Chen, J., Ye, W., Yan, M., Ye, Q., Zhang, J., Huang, F., and Zhang, S. (2024, January 17\u201321). Hallucination Augmented Contrastive Learning for Multimodal Large Language Model. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.02553"},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Jiao, Q., Chen, D., Huang, Y., Ding, B., Li, Y., and Shen, Y. (2024). Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models. arXiv.","DOI":"10.1109\/CVPR52734.2025.00868"},{"key":"ref_53","doi-asserted-by":"crossref","first-page":"102888","DOI":"10.1016\/j.inffus.2024.102888","article-title":"A Comprehensive Survey of Large Language Models and Multimodal Large Language Models in Medicine","volume":"117","author":"Xiao","year":"2025","journal-title":"Inf. Fusion"},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Li, X., Zhang, M., Geng, Y., Geng, H., Long, Y., Shen, Y., Zhang, R., Liu, J., and Dong, H. (2024, January 17\u201321). ManipLLM: Embodied Multimodal Large Language Model for Object-Centric Robotic Manipulation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.01710"},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Li, P., Liu, G., He, J., Zhao, Z., and Zhong, S. (2023). Masked Vision and Language Pre-training with Unimodal and Multimodal Contrastive Losses for Medical Visual Question Answering. Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), Springer Nature.","DOI":"10.1007\/978-3-031-43907-0_36"},{"key":"ref_56","unstructured":"Yuan, Z., Jin, Q., Tan, C., Zhao, Z., Yuan, H., Huang, F., and Huang, S. (November, January 29). RAMM: Retrieval-augmented Biomedical Visual Question Answering with Multi-modal Pre-training. Proceedings of the 31st ACM International Conference on Multimedia (ACM MM), Ottawa, ON, Canada."},{"key":"ref_57","doi-asserted-by":"crossref","unstructured":"Sun, R., Li, Z., Ding, Y., Wang, Q., Wang, J., Zheng, H., Wu, W., and Xian, Y. (2023, January 9\u201314). Fusion or Defusion? Flexible Vision-and-Language Pre-Training. Proceedings of the Findings of the Association for Computational Linguistics: ACL 2023, Toronto, ON, Canada.","DOI":"10.18653\/v1\/2023.findings-acl.316"},{"key":"ref_58","unstructured":"Shuai, Z., and Shen, L. (2024). Mitigating Heterogeneity in Federated Multimodal Learning with Biomedical Vision-Language Pre-training. arXiv."},{"key":"ref_59","doi-asserted-by":"crossref","unstructured":"Zhou, X., He, J., Ke, Y., Zhu, G., Guti\u00e9rrez-Basulto, V., and Pan, J.Z. (2024, January 11\u201316). An Empirical Study on Parameter-Efficient Fine-Tuning for MultiModal Large Language Models. Proceedings of the Findings of the Association for Computational Linguistics: ACL, Bangkok, Thailand.","DOI":"10.18653\/v1\/2024.findings-acl.598"},{"key":"ref_60","doi-asserted-by":"crossref","unstructured":"Long, Z., Killick, G., McCreadie, R., and Aragon-Camarasa, G. (2023). MultiWay-Adapter: Adapting Large-Scale Multi-Modal Models for Scalable Image-Text Retrieval. arXiv.","DOI":"10.1109\/ICASSP48485.2024.10446792"},{"key":"ref_61","unstructured":"Huang, J., Zhang, J., Jiang, K., Qiu, H., and Lu, S. (2023). Visual Instruction Tuning towards General-Purpose Multimodal Model: A Survey. arXiv."},{"key":"ref_62","doi-asserted-by":"crossref","first-page":"27","DOI":"10.1007\/s44267-025-00099-6","article-title":"Efficient Multimodal Large Language Models: A Survey","volume":"3","author":"Jin","year":"2025","journal-title":"Vis. Intell."},{"key":"ref_63","unstructured":"Baydin, A.G., Cornish, R., Rubio, D.M., Schmidt, M., and Wood, F. (2018). Online Learning Rate Adaptation with Hypergradient Descent. arXiv."},{"key":"ref_64","doi-asserted-by":"crossref","unstructured":"Liu, B., Chen, C., Liao, C., Gong, Z., Wang, H., Lei, Z., Liang, M., Chen, D., Shen, M., and Zhou, H. (2024, January 25\u201329). MFTCoder: Boosting Code LLMs with Multitask Fine-Tuning. Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Barcelona, Spain.","DOI":"10.1145\/3637528.3671609"},{"key":"ref_65","unstructured":"Mahabadi, R.K., Ruder, S., Dehghani, M., and Henderson, J. (2021, January 1\u20136). Parameter-Efficient Multi-Task Fine-Tuning for Transformers via Shared Hypernetworks. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics (ACL), Online."},{"key":"ref_66","unstructured":"Bari, M.S., Zhang, A., Zheng, S., Shi, X., Zhu, Y., Joty, S., and Li, M. (2022). SPT: Semi-Parametric Prompt Tuning for Multitask Prompted Learning. arXiv."},{"key":"ref_67","unstructured":"Liu, X., Liu, T., Huang, S., Xin, Y., Hu, Y., Yin, Q., Wang, D., Wu, Y., and Chen, H. (2024). M2IST: Multi-Modal Interactive Side-Tuning for Efficient Referring Expression Comprehension. arXiv."},{"key":"ref_68","unstructured":"Li, J., He, X., Wei, L., Qian, L., Zhu, L., Xie, L., Zhuang, Y., Tian, Q., and Tang, S. (December, January 28). Fine-Grained Semantically Aligned Vision-Language Pre-Training. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), New Orleans, LA, USA."},{"key":"ref_69","unstructured":"Hu, E.J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. (2022, January 25\u201329). LoRA: Low-Rank Adaptation of Large Language Models. Proceedings of the International Conference on Learning Representations (ICLR), Virtual."},{"key":"ref_70","unstructured":"Tsimpoukelli, M., Menick, J.L., Cabi, S., Eslami, S.M.A., Vinyals, O., and Hill, F. (2021, January 6\u201314). Multimodal Few-Shot Learning with Frozen Pretrained Language Models. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Virtual."},{"key":"ref_71","unstructured":"Huang, S., Dong, L., Wang, W., Hao, Y., Singhal, S., Ma, S., Lv, T., Cui, L., Mohammed, O.K., and Patra, B. (2023, January 10\u201316). Language Is Not All You Need: Aligning Perception with Language Models. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), New Orleans, LA, USA."},{"key":"ref_72","unstructured":"Hu, J., Yao, Y., Wang, C., Wang, S., Pan, Y., Chen, Q., Yu, T., Wu, H., Zhao, Y., and Zhang, H. (2023). Large Multilingual Models Pivot Zero-Shot Multimodal Learning across Languages. arXiv."},{"key":"ref_73","unstructured":"Peng, B., Li, C., He, P., Galley, M., and Gao, J. (2023). Instruction Tuning with GPT-4. arXiv."},{"key":"ref_74","unstructured":"Gupta, H., Sawant, S.A., Mishra, S., Nakamura, M., Mitra, A., Mashetty, S., and Baral, C. (2023). Instruction tuned models are quick learners. arXiv."},{"key":"ref_75","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014, January 6\u201312). Microsoft COCO: Common Objects in Context. Proceedings of the European Conference on Computer Vision (ECCV), Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_76","doi-asserted-by":"crossref","first-page":"67","DOI":"10.1162\/tacl_a_00166","article-title":"From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions","volume":"2","author":"Young","year":"2014","journal-title":"Trans. Assoc. Comput. Linguist."},{"key":"ref_77","doi-asserted-by":"crossref","unstructured":"Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., Parikh, D., and Rohrbach, M. (2019, January 16\u201320). Towards VQA Models That Can Read. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00851"},{"key":"ref_78","doi-asserted-by":"crossref","unstructured":"Marino, K., Rastegari, M., Farhadi, A., and Mottaghi, R. (2019, January 16\u201320). OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00331"},{"key":"ref_79","unstructured":"Agrawal, H., Desai, K., Wang, Y., Chen, X., Jain, R., Johnson, M., Batra, D., Parikh, D., Lee, S., and Anderson, P. (November, January 27). nocaps: Novel Object Captioning at Scale. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea."},{"key":"ref_80","unstructured":"Suhr, A., Zhou, S., Zhang, A., Zhang, I., Bai, H., and Artzi, Y. (August, January 28). A Corpus for Reasoning About Natural Language Grounded in Photographs. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL), Florence, Italy."},{"key":"ref_81","doi-asserted-by":"crossref","unstructured":"Pezeshkpour, P., and Hruschka, E. (2023). Large Language Models Sensitivity to the Order of Options in Multiple-Choice Questions. arXiv.","DOI":"10.18653\/v1\/2024.findings-naacl.130"},{"key":"ref_82","unstructured":"Zheng, C., Zhou, H., Meng, F., Zhou, J., and Huang, M. (2023). Large Language Models Are Not Robust Multiple Choice Selectors. arXiv."},{"key":"ref_83","doi-asserted-by":"crossref","unstructured":"Cornia, M., Stefanini, M., Baraldi, L., and Cucchiara, R. (2020, January 14\u201319). Meshed-Memory Transformer for Image Captioning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Virtual.","DOI":"10.1109\/CVPR42600.2020.01059"},{"key":"ref_84","doi-asserted-by":"crossref","unstructured":"Li, X., Yin, X., Li, C., Zhang, P., Hu, X., Zhang, L., Wang, L., Hu, H., Dong, L., and Wei, F. (2020). OSCAR: Object-Semantics Aligned Pre-training for Vision-Language Tasks. Proceedings of the European Conference on Computer Vision (ECCV), Springer.","DOI":"10.1007\/978-3-030-58577-8_8"},{"key":"ref_85","unstructured":"Johnson, J., Karpathy, A., and Fei-Fei, L. (July, January 26). DenseCap: Fully Convolutional Localization Networks for Dense Captioning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA."},{"key":"ref_86","doi-asserted-by":"crossref","unstructured":"Hu, X., Yin, X., Lin, K., Wang, L., Zhang, L., Gao, J., and Liu, Z. (2021, January 2\u20139). VIVO: Visual Vocabulary Pre-Training for Novel Object Captioning. Proceedings of the AAAI Conference on Artificial Intelligence, Virtual.","DOI":"10.1609\/aaai.v35i2.16249"},{"key":"ref_87","unstructured":"Chen, C., Mu, S., Xiao, W., Ye, Z., Wu, L., and Ju, Q. (February, January 27). Improving Image Captioning with Conditional Generative Adversarial Nets. Proceedings of the AAAI Conference on Artificial Intelligence, Honolulu, HI, USA."},{"key":"ref_88","unstructured":"Li, N., Chen, Z., and Liu, S. (February, January 27). Meta Learning for Image Captioning. Proceedings of the AAAI Conference on Artificial Intelligence, Honolulu, HI, USA."},{"key":"ref_89","unstructured":"de Faria, A.C.A.M., Bastos, F.d.C., da Silva, J.V.N.A., Fabris, V.L., Uchoa, V.d.S., Neto, D.G.d.A., and Santos, C.F.G.d. (2023). Visual Question Answering: A Survey on Techniques and Common Trends in Recent Literature. arXiv."},{"key":"ref_90","doi-asserted-by":"crossref","unstructured":"Yu, Z., Yu, J., Cui, Y., Tao, D., and Tian, Q. (2019, January 16\u201320). Deep Modular Co-Attention Networks for Visual Question Answering. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00644"},{"key":"ref_91","doi-asserted-by":"crossref","unstructured":"Lan, Y., Li, X., Liu, X., Li, Y., Qin, W., and Qian, W. (2023). Improving Zero-shot Visual Question Answering via Large Language Models with Reasoning Question Prompts. arXiv.","DOI":"10.1145\/3581783.3612389"},{"key":"ref_92","doi-asserted-by":"crossref","unstructured":"Bigham, J.P., Jayant, C., Ji, H., Little, G., Miller, A., Miller, R.C., Miller, R., Tatarowicz, A., White, B., and White, S. (2010, January 3\u20136). VizWiz: Nearly Real-time Answers to Visual Questions. Proceedings of the 23rd Annual ACM Symposium on User Interface Software and Technology (UIST), New York, NY, USA.","DOI":"10.1145\/1866029.1866080"},{"key":"ref_93","doi-asserted-by":"crossref","unstructured":"Li, B., Jia, G., Gao, X., and Ma, C. (2024, January 19\u201321). Multidimensional Semantic Augmented Visual Storytelling. Proceedings of the 2024 4th International Conference on Neural Networks, Information and Communication (NNICE), Guangzhou, China.","DOI":"10.1109\/NNICE61279.2024.10498935"},{"key":"ref_94","doi-asserted-by":"crossref","unstructured":"Hong, X., Shetty, R., Demberg, V., and Schiele, B. (2020, January 19\u201320). Diverse and Relevant Visual Storytelling with Scene Graph Embeddings. Proceedings of the 24th Conference on Computational Natural Language Learning (CoNLL), Virtual.","DOI":"10.18653\/v1\/2020.conll-1.34"},{"key":"ref_95","doi-asserted-by":"crossref","unstructured":"Fu, R., Liu, J., Chen, X., Nie, Y., and Xiong, W. (2024). Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning. arXiv.","DOI":"10.1109\/WACV61041.2025.00220"},{"key":"ref_96","doi-asserted-by":"crossref","unstructured":"Parde, N. (2020, January 12). And, Action! Towards Leveraging Multimodal Patterns for Storytelling and Content Analysis. Proceedings of the 2nd International Workshop on AI for Smart TV Content Production, Access and Delivery (AI4TV) at ACM MM, Virtual.","DOI":"10.1145\/3422839.3423060"},{"key":"ref_97","doi-asserted-by":"crossref","unstructured":"Wang, S., Yu, Z., Jiang, X., Lan, S., Shi, M., Chang, N., Kautz, J., Li, Y., and Alvarez, J.M. (2024). OmniDrive: A Holistic LLM-Agent Framework for Autonomous Driving with 3D Perception, Reasoning and Planning. arXiv.","DOI":"10.1109\/CVPR52734.2025.02090"},{"key":"ref_98","doi-asserted-by":"crossref","first-page":"164","DOI":"10.1016\/j.patrec.2021.06.011","article-title":"Beyond Visual Semantics: Exploring the Role of Scene Text in Image Understanding","volume":"149","author":"Dey","year":"2021","journal-title":"Pattern Recognit. Lett."},{"key":"ref_99","unstructured":"Zang, C., Tang, J., Zhang, R., Zhao, Z., Lv, T., Pei, M., and Liang, W. (2024). Let Storytelling Tell Vivid Stories: An Expressive and Fluent Multimodal Storyteller. arXiv."},{"key":"ref_100","doi-asserted-by":"crossref","unstructured":"Yang, B., He, L., Liu, K., and Yan, Z. (2024). VIAssist: Adapting Multi-modal Large Language Models for Users with Visual Impairments. arXiv.","DOI":"10.1109\/FMSys62467.2024.00010"},{"key":"ref_101","doi-asserted-by":"crossref","first-page":"884","DOI":"10.1109\/5.664278","article-title":"Next-generation content representation, creation, and searching for new-media applications in education","volume":"86","author":"Chang","year":"1998","journal-title":"Proc. IEEE"},{"key":"ref_102","doi-asserted-by":"crossref","unstructured":"Mart\u00edn, \u00c1., Iribas, H., Alberdi, I., and Aginako, N. (2012, January 10\u201313). Dynamic Multimedia Creation Using Knowledge Content Driven Database. Proceedings of the IEEE International Symposium on Parallel and Distributed Processing with Applications (ISPA), Legan\u00e9s, Madrid, Spain.","DOI":"10.1109\/ISPA.2012.114"},{"key":"ref_103","doi-asserted-by":"crossref","unstructured":"Tilekbay, B., Yang, S., Lewkowicz, M.C., Suryapranata, A., and Kim, J. (2024, January 18\u201321). ExpressEdit: Video Editing with Natural Language and Sketching. Proceedings of the Companion Proceedings of the 29th International Conference on Intelligent User Interfaces (IUI Companion), Greenville, SC, USA.","DOI":"10.1145\/3640544.3645226"},{"key":"ref_104","doi-asserted-by":"crossref","unstructured":"Kubicek, R., Zak, P., Zem\u010d\u00edk, P., and Herout, A. (2008, January 10\u201312). Automatic Video Editing for Multimodal Meetings. Proceedings of the International Conference on Computer Vision and Graphics (ICCVG), Warsaw, Poland.","DOI":"10.1007\/978-3-642-02345-3_26"},{"key":"ref_105","doi-asserted-by":"crossref","first-page":"62","DOI":"10.1109\/MMUL.2004.1261109","article-title":"A model-driven approach to content repurposing","volume":"11","author":"Obrenovic","year":"2004","journal-title":"IEEE Multimed."},{"key":"ref_106","doi-asserted-by":"crossref","unstructured":"Wieschebrink, S. (2011, January 19\u201322). Collaborative editing of multimodal annotation data. Proceedings of the ACM Symposium on Document Engineering (DocEng), Mountain View, CA, USA.","DOI":"10.1145\/2034691.2034706"},{"key":"ref_107","doi-asserted-by":"crossref","unstructured":"Santos, A. (1995). Multimedia and Groupware for Editing. Computer Graphics: Systems and Applications, Springer.","DOI":"10.1007\/978-3-642-79865-8"},{"key":"ref_108","doi-asserted-by":"crossref","unstructured":"Sauer, S., Osswald, K., Wielemans, X., and Stifter, M. (2006, January 4\u20136). U-Create: Creative Authoring Tools for Edutainment Applications. Proceedings of the 3rd International Conference on Technologies for Interactive Digital Storytelling and Entertainment (TIDSE), Darmstadt, Germany.","DOI":"10.1007\/11944577_16"},{"key":"ref_109","doi-asserted-by":"crossref","unstructured":"Jokela, T., Lehikoinen, J., and Korhonen, H. (2008, January 5\u201310). Mobile multimedia presentation editor: Enabling creation of audio-visual stories on mobile devices. Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (CHI), Florence, Italy.","DOI":"10.1145\/1357054.1357066"},{"key":"ref_110","doi-asserted-by":"crossref","unstructured":"Bateman, J. (2008). Multimodality and Genre: A Foundation for the Systematic Analysis of Multimodal Documents, Palgrave Macmillan.","DOI":"10.1057\/9780230582323_5"},{"key":"ref_111","doi-asserted-by":"crossref","unstructured":"Ranjan, V., Rasiwasia, N., and Jawahar, C. (2015, January 7\u201313). Multi-label Cross-Modal Retrieval. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.466"},{"key":"ref_112","unstructured":"Gomez, F.P., Sanabria, R., Sung, Y.h., Cer, D., Dalmia, S., and Hern\u00e1ndez Abrego, G. (2024, January 15\u201316). Transforming LLMs into Cross-modal and Cross-lingual Retrieval Systems. Proceedings of the 21st International Conference on Spoken Language Translation (IWSLT), Bangkok, Thailand."},{"key":"ref_113","doi-asserted-by":"crossref","unstructured":"Chen, H., Cooper, M.L., Joshi, D., and Girod, B. (2014, January 3\u20137). Multi-modal Language Models for Lecture Video Retrieval. Proceedings of the ACM International Conference on Multimedia (MM), Orlando, FL, USA.","DOI":"10.1145\/2647868.2654964"},{"key":"ref_114","doi-asserted-by":"crossref","first-page":"52","DOI":"10.1109\/MSP.2018.2868887","article-title":"Cross-Modal Music Retrieval and Applications: An Overview of Key Methodologies","volume":"36","author":"Arzt","year":"2019","journal-title":"IEEE Signal Process. Mag."},{"key":"ref_115","unstructured":"Chen, B., Xia, F., Ichter, B., and Rao, K. (June, January 29). Open-vocabulary Queryable Scene Representations for Real World Planning. Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), London, UK."},{"key":"ref_116","unstructured":"Dayma, B., Patil, S., Cuenca, P., Saifullah, K., Abraham, T., Le Khac, P., Melas, L., and Ghosh, R. (2024, September 29). DALL-E Mini. Available online: https:\/\/huggingface.co\/spaces\/dalle-mini\/dalle-mini."},{"key":"ref_117","unstructured":"OpenAI (2026, May 10). ChatGPT Retrieval Plugin. Available online: https:\/\/github.com\/openai\/chatgpt-retrieval-plugin."},{"key":"ref_118","unstructured":"Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M. (2022). Hierarchical Text-Conditional Image Generation with CLIP Latents. arXiv."},{"key":"ref_119","doi-asserted-by":"crossref","unstructured":"Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. (2022, January 19\u201324). High-Resolution Image Synthesis with Latent Diffusion Models. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"ref_120","unstructured":"Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E.L., Ghasemipour, S.K.S., Gontijo-Lopes, R., Karagol Ayan, B., and Salimans, T. (December, January 28). Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), New Orleans, LA, USA."},{"key":"ref_121","unstructured":"Borji, A. (2022). Generated Faces in the Wild: Quantitative Comparison of Stable Diffusion, Midjourney and DALL-E 2. arXiv."},{"key":"ref_122","unstructured":"Midjourney, Inc. (2024, September 29). Midjourney. Available online: https:\/\/www.midjourney.com\/."},{"key":"ref_123","unstructured":"Betker, J., Goh, G., Jing, L., Brooks, T., Wang, J., Li, L., Ouyang, L., Zhuang, J., Lee, J., and Guo, Y. (2023). Improving Image Generation with Better Captions, OpenAI. Technical Report."},{"key":"ref_124","unstructured":"Stability AI (2024, September 29). Stable Diffusion. Available online: https:\/\/stability.ai\/stable-diffusion."},{"key":"ref_125","unstructured":"Black Forest Labs (2024, September 29). FLUX.1: A High-Resolution Image Generation Model. Available online: https:\/\/github.com\/black-forest-labs\/flux."},{"key":"ref_126","doi-asserted-by":"crossref","unstructured":"Liu, R., Garrette, D., Saharia, C., Chan, W., Roberts, A., Narang, S., Blok, I., Mical, R., Norouzi, M., and Constant, N. (2023, January 9\u201314). Character-Aware Models Improve Visual Text Rendering. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL), Toronto, ON, Canada.","DOI":"10.18653\/v1\/2023.acl-long.900"},{"key":"ref_127","unstructured":"Carlini, N., Hayes, J., Nasr, M., Jagielski, M., Sehwag, V., Tram\u00e8r, F., Balle, B., Ippolito, D., and Wallace, E. (2023, January 9\u201311). Extracting Training Data from Diffusion Models. Proceedings of the 32nd USENIX Security Symposium, Anaheim, CA, USA."},{"key":"ref_128","unstructured":"Google (2024, September 29). Google Lens. Available online: https:\/\/lens.google.com\/."},{"key":"ref_129","unstructured":"Microsoft (2024, September 29). Bing Visual Search. Available online: https:\/\/www.bing.com\/visualsearch."},{"key":"ref_130","unstructured":"You.com (2024, September 29). You.com: The AI Search Engine You Control. Available online: https:\/\/you.com\/."},{"key":"ref_131","unstructured":"Perplexity AI (2024, September 29). Perplexity AI. Available online: https:\/\/www.perplexity.ai\/."},{"key":"ref_132","doi-asserted-by":"crossref","unstructured":"Zhou, T., Mei, S., Li, X., Liu, Z., Xiong, C., Liu, Z., Gu, Y., and Yu, G. (2024, January 11\u201316). MARVEL: Unlocking the Multi-Modal Capability of Dense Retrieval via Visual Module Plugin. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL), Bangkok, Thailand.","DOI":"10.18653\/v1\/2024.acl-long.783"},{"key":"ref_133","doi-asserted-by":"crossref","unstructured":"Feng, J., Tao, C., Geng, X., Shen, T., Xu, C., Long, G., Zhao, D., and Jiang, D. (2023). Synergistic Interplay between Search and Large Language Models for Information Retrieval. arXiv.","DOI":"10.18653\/v1\/2024.acl-long.517"},{"key":"ref_134","doi-asserted-by":"crossref","unstructured":"Cohan, A., Feldman, S., Beltagy, I., Downey, D., and Weld, D. (2020, January 5\u201310). SPECTER: Document-level Representation Learning using Citation-informed Transformers. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), Virtual.","DOI":"10.18653\/v1\/2020.acl-main.207"},{"key":"ref_135","unstructured":"LangChain (2024, September 29). LangChain. Available online: https:\/\/langchain.com\/."},{"key":"ref_136","unstructured":"Pinecone Systems Inc. (2024, September 29). Pinecone: Vector Database for Machine Learning. Available online: https:\/\/www.pinecone.io\/."},{"key":"ref_137","unstructured":"Chroma (2024, September 29). Chroma: The AI-Native Open-Source Embedding Database. Available online: https:\/\/www.trychroma.com\/."},{"key":"ref_138","doi-asserted-by":"crossref","first-page":"535","DOI":"10.1109\/TBDATA.2019.2921572","article-title":"Billion-Scale Similarity Search with GPUs","volume":"7","author":"Johnson","year":"2019","journal-title":"IEEE Trans. Big Data"},{"key":"ref_139","unstructured":"Weaviate (2024, September 29). Weaviate: Vector Database. Available online: https:\/\/weaviate.io\/."},{"key":"ref_140","unstructured":"Qdrant (2024, September 29). Qdrant: Vector Database for the Next Generation of AI Applications. Available online: https:\/\/qdrant.tech\/."},{"key":"ref_141","unstructured":"Vespa (2024, September 29). Vespa: The Open Big Data Serving Engine. Available online: https:\/\/vespa.ai\/."},{"key":"ref_142","doi-asserted-by":"crossref","first-page":"117","DOI":"10.1109\/TPAMI.2010.57","article-title":"Product Quantization for Nearest Neighbor Search","volume":"33","author":"Douze","year":"2011","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_143","doi-asserted-by":"crossref","first-page":"824","DOI":"10.1109\/TPAMI.2018.2889473","article-title":"Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs","volume":"42","author":"Malkov","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_144","doi-asserted-by":"crossref","unstructured":"Li, Y., Du, Y., Zhou, K., Wang, J., Zhao, W.X., and Wen, J.R. (2023, January 6\u201310). Evaluating Object Hallucination in Large Vision-Language Models. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), Singapore.","DOI":"10.18653\/v1\/2023.emnlp-main.20"},{"key":"ref_145","unstructured":"Wang, Z., Wang, Z., Le, L., Zheng, H.S., Mishra, S., Perot, V., Zhang, Y., Mattapalli, A., Taly, A., and Shang, J. (2025, January 24\u201328). Speculative RAG: Enhancing Retrieval Augmented Generation through Drafting. Proceedings of the Thirteenth International Conference on Learning Representations (ICLR), Singapore."},{"key":"ref_146","unstructured":"Yang, Z., Li, L., Lin, K., Wang, J., Lin, C.C., Liu, Z., and Wang, L. (2023). The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision). arXiv."},{"key":"ref_147","unstructured":"Runway AI, Inc. (2024, September 29). Runway ML. Available online: https:\/\/runwayml.com\/."},{"key":"ref_148","unstructured":"Wang, J., Yuan, H., Chen, D., Zhang, Y., Wang, X., and Zhang, S. (2023). ModelScope Text-to-Video Technical Report. arXiv."},{"key":"ref_149","first-page":"86","article-title":"A Comparative Study on the Features and Applications of AI Tools-Focus on PIKA Labs and RUNWAY","volume":"16","author":"Guo","year":"2024","journal-title":"Int. J. Internet Broadcast. Commun."},{"key":"ref_150","unstructured":"Technology, K. (2024, November 04). Kling AI: Advanced Text-to-Video Generation Model. Available online: https:\/\/kling.kuaishou.com\/en."},{"key":"ref_151","unstructured":"Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Chen, X., Choromanski, K., Ding, T., Driess, D., Dubey, A., and Finn, C. (2023, January 6\u20139). RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control. Proceedings of the 7th Conference on Robot Learning (CoRL), PMLR, Atlanta, GA, USA."},{"key":"ref_152","unstructured":"Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Fu, C., Gopalakrishnan, K., and Hausman, K. (2022, January 14\u201318). Do As I Can, Not As I Say: Grounding Language in Robotic Affordances. Proceedings of the 6th Conference on Robot Learning (CoRL), PMLR, Auckland, New Zealand."},{"key":"ref_153","unstructured":"Driess, D., Xia, F., Sajjadi, M.S., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., and Yu, T. (2023). PaLM-E: An Embodied Multimodal Language Model. arXiv."},{"key":"ref_154","unstructured":"Huang, W., Wang, C., Zhang, R., Li, Y., Wu, J., and Li, F.-F. (2023). VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models. arXiv."},{"key":"ref_155","unstructured":"Huang, C., Mees, O., Zeng, A., and Burgard, W. (June, January 29). Visual Language Maps for Robot Navigation. Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), London, UK."},{"key":"ref_156","unstructured":"Chase, H. (2026, May 10). LangChain: Building Applications with LLMs Through Composability. Available online: https:\/\/github.com\/langchain-ai\/langchain."},{"key":"ref_157","unstructured":"Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., K\u00fcttler, H., Lewis, M., Yih, W.t., and Rockt\u00e4schel, T. (2020, January 6\u201312). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Virtual."},{"key":"ref_158","unstructured":"Choromanski, K., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J., Mohiuddin, A., and Kaiser, L. (2021, January 3\u20137). Rethinking Attention with Performers. Proceedings of the International Conference on Learning Representations (ICLR), Virtual."},{"key":"ref_159","unstructured":"Kim, W., Son, B., and Kim, I. (2021, January 18\u201324). ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision. Proceedings of the International Conference on Machine Learning (ICML), PMLR, Virtual."},{"key":"ref_160","unstructured":"Beltagy, I., Peters, M.E., and Cohan, A. (2020). Longformer: The Long-Document Transformer. arXiv."},{"key":"ref_161","unstructured":"Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I. (2021, January 18\u201324). Zero-Shot Text-to-Image Generation. Proceedings of the International Conference on Machine Learning (ICML), PMLR, Virtual."},{"key":"ref_162","unstructured":"Lu, J., Clark, C., Zellers, R., Mottaghi, R., and Kembhavi, A. (2023, January 1\u20135). Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks. Proceedings of the International Conference on Learning Representations (ICLR), Kigali, Rwanda."},{"key":"ref_163","unstructured":"Kaplan, J., McCandlish, S., Henighan, T., Brown, T.B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. (2020). Scaling Laws for Neural Language Models. arXiv."},{"key":"ref_164","doi-asserted-by":"crossref","first-page":"1061","DOI":"10.1162\/tacl_a_00413","article-title":"Compressing Large-Scale Transformer-Based Models: A Case Study on BERT","volume":"9","author":"Ganesh","year":"2021","journal-title":"Trans. Assoc. Comput. Linguist."},{"key":"ref_165","unstructured":"Bengio, Y., L\u00e9onard, N., and Courville, A. (2013). Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation. arXiv."},{"key":"ref_166","unstructured":"Jia, C., Yang, Y., Xia, Y., Chen, Y.T., Parekh, Z., Pham, H., Le, Q., Sung, Y.H., Li, Z., and Duerig, T. (2021, January 18\u201324). Scaling Up Visual and Vision-Language Representation Learning with Noisy Text Supervision. Proceedings of the International Conference on Machine Learning (ICML), PMLR, Virtual."},{"key":"ref_167","unstructured":"Chen, T., Xu, B., Zhang, C., and Guestrin, C. (2016). Training Deep Nets with Sublinear Memory Cost. arXiv."},{"key":"ref_168","unstructured":"Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F. (2020, January 13\u201318). Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention. Proceedings of the International Conference on Machine Learning (ICML), PMLR, Virtual."},{"key":"ref_169","unstructured":"Micikevicius, P., Narang, S., Alben, J., Diamos, G., Elsen, E., Garcia, D., Ginsburg, B., Houston, M., Kuchaiev, O., and Venkatesh, G. (May, January 30). Mixed Precision Training. Proceedings of the International Conference on Learning Representations (ICLR), Vancouver, BC, Canada."},{"key":"ref_170","unstructured":"Ngiam, J., Khosla, A., Kim, M., Nam, J., Lee, H., and Ng, A.Y. (July, January 28). Multimodal Deep Learning. Proceedings of the International Conference on Machine Learning (ICML), Bellevue, WA, USA."},{"key":"ref_171","unstructured":"Chen, L., Gan, Z., Cheng, Y., Li, L., Carin, L., and Liu, J. (2020, January 13\u201318). Graph Optimal Transport for Cross-Domain Alignment. Proceedings of the International Conference on Machine Learning (ICML), PMLR, Virtual. Available online: https:\/\/proceedings.mlr.press\/v119\/chen20e.html."},{"key":"ref_172","doi-asserted-by":"crossref","unstructured":"Liu, Z., Ning, J., Cao, Y., Wei, Y., Zhang, Z., Lin, S., and Hu, H. (2022, January 19\u201324). Video Swin Transformer. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00320"},{"key":"ref_173","unstructured":"Child, R., Gray, S., Radford, A., and Sutskever, I. (2019). Generating Long Sequences with Sparse Transformers. arXiv."},{"key":"ref_174","unstructured":"Peng, X., Bai, Q., Xia, X., Huang, Z., Saenko, K., and Wang, B. (November, January 27). Moment Matching for Multi-Source Domain Adaptation. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea."},{"key":"ref_175","doi-asserted-by":"crossref","first-page":"3521","DOI":"10.1073\/pnas.1611835114","article-title":"Overcoming Catastrophic Forgetting in Neural Networks","volume":"114","author":"Kirkpatrick","year":"2017","journal-title":"Proc. Natl. Acad. Sci. USA"},{"key":"ref_176","unstructured":"Cohen, J., Rosenfeld, E., and Kolter, Z. (2019, January 9\u201315). Certified Adversarial Robustness via Randomized Smoothing. Proceedings of the International Conference on Machine Learning (ICML), PMLR, Long Beach, CA, USA. Available online: https:\/\/proceedings.mlr.press\/v97\/cohen19c.html."},{"key":"ref_177","unstructured":"Naseer, M.M., Khan, S., Khan, M.H., Khan, F.S., and Porikli, F. (2019, January 8\u201314). Cross-Domain Transferability of Adversarial Perturbations. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Vancouver, BC, Canada. Available online: https:\/\/proceedings.neurips.cc\/paper\/2019\/hash\/99cd3843754d20ec3c5885d805db8a32-Abstract.html."},{"key":"ref_178","unstructured":"Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A. (May, January 30). Towards Deep Learning Models Resistant to Adversarial Attacks. Proceedings of the International Conference on Learning Representations (ICLR), Vancouver, BC, Canada."},{"key":"ref_179","unstructured":"Lazaridou, A., Kuncoro, A., Gribovskaya, E., Agrawal, D., Liska, A., Terzi, T., Gimenez, M., de Masson d\u2019Autume, C., Kocisky, T., and Ruder, S. (2021, January 6\u201314). Mind the Gap: Assessing Temporal Generalization in Neural Language Models. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Virtual."},{"key":"ref_180","unstructured":"Guo, C., Pleiss, G., Sun, Y., and Weinberger, K.Q. (2017, January 6\u201311). On Calibration of Modern Neural Networks. Proceedings of the International Conference on Machine Learning (ICML), PMLR, Sydney, Australia."},{"key":"ref_181","doi-asserted-by":"crossref","first-page":"859","DOI":"10.1080\/01621459.2017.1285773","article-title":"Variational Inference: A Review for Statisticians","volume":"112","author":"Blei","year":"2017","journal-title":"J. Am. Stat. Assoc."},{"key":"ref_182","unstructured":"Tack, J., Mo, S., Jeong, J., and Shin, J. (2020, January 6\u201312). CSI: Novelty Detection via Contrastive Learning on Distributionally Shifted Instances. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Virtual."},{"key":"ref_183","doi-asserted-by":"crossref","unstructured":"Abnar, S., and Zuidema, W. (2020, January 5\u201310). Quantifying Attention Flow in Transformers. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), Virtual.","DOI":"10.18653\/v1\/2020.acl-main.385"},{"key":"ref_184","unstructured":"Sundararajan, M., Taly, A., and Yan, Q. (2017, January 6\u201311). Axiomatic Attribution for Deep Networks. Proceedings of the International Conference on Machine Learning (ICML), PMLR, Sydney, Australia."},{"key":"ref_185","doi-asserted-by":"crossref","unstructured":"Ribeiro, M.T., Singh, S., and Guestrin, C. (2016, January 13\u201317). \u201cWhy Should I Trust You?\u201d: Explaining the Predictions of Any Classifier. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD), San Francisco, CA, USA.","DOI":"10.1145\/2939672.2939778"},{"key":"ref_186","doi-asserted-by":"crossref","unstructured":"Rebuffi, S.A., Fong, R., Ji, X., and Vedaldi, A. (2020, January 14\u201319). There and Back Again: Revisiting Backpropagation Saliency Methods. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Virtual.","DOI":"10.1109\/CVPR42600.2020.00886"},{"key":"ref_187","unstructured":"Kim, B., Wattenberg, M., Gilmer, J., Cai, C., Wexler, J., Viegas, F., and Sayres, R. (2018, January 10\u201315). Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV). Proceedings of the International Conference on Machine Learning (ICML), PMLR, Stockholm, Sweden."},{"key":"ref_188","unstructured":"Mao, J., Gan, C., Kohli, P., Tenenbaum, J.B., and Wu, J. (2019, January 6\u20139). The Neuro-Symbolic Concept Learner: Interpreting Scenes, Words, and Sentences from Natural Supervision. Proceedings of the International Conference on Learning Representations (ICLR), New Orleans, LA, USA."},{"key":"ref_189","unstructured":"Ellis, K., Ritchie, D., Solar-Lezama, A., and Tenenbaum, J.B. (2018, January 2\u20138). Learning to Infer Graphics Programs from Hand-Drawn Images. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Montr\u00e9al, QC, Canada."},{"key":"ref_190","doi-asserted-by":"crossref","unstructured":"Johnson, J., Hariharan, B., van der Maaten, L., Fei-Fei, L., Zitnick, C.L., and Girshick, R. (2017, January 21\u201326). CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.215"},{"key":"ref_191","unstructured":"Lake, B.M., and Baroni, M. (2018, January 10\u201315). Generalization without Systematicity: On the Compositional Skills of Sequence-to-Sequence Recurrent Networks. Proceedings of the International Conference on Machine Learning (ICML), PMLR, Stockholm, Sweden."},{"key":"ref_192","unstructured":"Abdin, M., Aneja, J., Awadalla, H., Awadallah, A., Awan, A.A., Bach, N., Bahree, A., Bakhtiari, A., Bao, J., and Behl, H. (2024). Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone. Phi-3-Vision described in the technical report. arXiv."},{"key":"ref_193","doi-asserted-by":"crossref","first-page":"41","DOI":"10.47672\/ejt.1890","article-title":"Ethical Considerations in the Development and Deployment of AI Systems","volume":"8","author":"Konidena","year":"2024","journal-title":"Eur. J. Technol."},{"key":"ref_194","unstructured":"Peng, B., Chen, K., Li, M., Feng, P., Bi, Z., Liu, J., Song, X., and Niu, Q. (2024). Securing Large Language Models: Addressing Bias, Misinformation, and Prompt Attacks. arXiv."},{"key":"ref_195","unstructured":"Chang, Y., Chang, Y., and Wu, Y. (2024). BA-LoRA: Bias-Alleviating Low-Rank Adaptation to Mitigate Catastrophic Inheritance in Large Language Models. arXiv."},{"key":"ref_196","doi-asserted-by":"crossref","unstructured":"He, F., Zhu, T., Ye, D., Liu, B., Zhou, W., and Yu, P.S. (2024). The Emerged Security and Privacy of LLM Agent: A Survey with Case Studies. arXiv.","DOI":"10.1145\/3773080"},{"key":"ref_197","doi-asserted-by":"crossref","first-page":"5799","DOI":"10.1109\/OJCOMS.2024.3456549","article-title":"LLM-Based Edge Intelligence: A Comprehensive Survey on Architectures, Applications, Security and Trustworthiness","volume":"5","author":"Friha","year":"2024","journal-title":"IEEE Open J. Commun. Soc."},{"key":"ref_198","unstructured":"Chen, C., Gong, X., Liu, Z., Jiang, W., Goh, S.Q., and Lam, K.Y. (2024). Trustworthy, Responsible, and Safe AI: A Comprehensive Architectural Framework for AI Safety with Challenges and Mitigations. arXiv."},{"key":"ref_199","doi-asserted-by":"crossref","unstructured":"Rosenstrauch, D., Mangla, U., Gupta, A., and Masau, C.T. (2023). Artificial Intelligence and Ethics. Digital Health Entrepreneurship, Springer.","DOI":"10.1007\/978-3-031-33902-8_16"},{"key":"ref_200","doi-asserted-by":"crossref","first-page":"121","DOI":"10.1016\/j.iotcps.2023.04.003","article-title":"ChatGPT: A Comprehensive Review on Background, Applications, Key Challenges, Bias, Ethics, Limitations and Future Scope","volume":"3","author":"Ray","year":"2023","journal-title":"Internet Things Cyber Phys. Syst."},{"key":"ref_201","doi-asserted-by":"crossref","unstructured":"Xu, Y., Hu, L., Zhao, J., Qiu, Z., Ye, Y., and Gu, H. (2024). A Survey on Multilingual Large Language Models: Corpora, Alignment, and Bias. arXiv.","DOI":"10.1007\/s11704-024-40579-4"},{"key":"ref_202","unstructured":"Basta, C.R.S. (2022). Gender Bias in Natural Language Processing. [Ph.D. Thesis, Universitat Polit\u00e8cnica de Catalunya]."},{"key":"ref_203","unstructured":"Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C.D., and Ho, D.E. (2024). Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv."},{"key":"ref_204","doi-asserted-by":"crossref","unstructured":"Brown, H., Lee, K., Mireshghallah, F., Shokri, R., and Tram\u00e8r, F. (2022, January 21\u201324). What Does it Mean for a Language Model to Preserve Privacy?. Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (FAccT), Seoul, Republic of Korea.","DOI":"10.1145\/3531146.3534642"},{"key":"ref_205","doi-asserted-by":"crossref","first-page":"100211","DOI":"10.1016\/j.hcc.2024.100211","article-title":"A Survey on Large Language Model (LLM) Security and Privacy: The Good, the Bad, and the Ugly","volume":"4","author":"Yao","year":"2024","journal-title":"High-Confid. Comput."},{"key":"ref_206","doi-asserted-by":"crossref","unstructured":"Pan, X., Zhang, M., Ji, S., and Yang, M. (2020). Privacy Risks of General-Purpose Language Models. Proceedings of the 2020 IEEE Symposium on Security and Privacy (SP), Virtual, 18\u201320 May 2020, IEEE.","DOI":"10.1109\/SP40000.2020.00095"},{"key":"ref_207","unstructured":"Weidinger, L., Mellor, J., Rauh, M., Griffin, C., Uesato, J., Huang, P.S., Cheng, M., Glaese, M., Balle, B., and Kasirzadeh, A. (2021). Ethical and social risks of harm from language models. arXiv."},{"key":"ref_208","doi-asserted-by":"crossref","first-page":"2445","DOI":"10.1007\/s43681-024-00573-9","article-title":"Right to be Forgotten in the Era of Large Language Models: Implications, Challenges, and Solutions","volume":"5","author":"Zhang","year":"2025","journal-title":"Ai Ethics"},{"key":"ref_209","doi-asserted-by":"crossref","unstructured":"Weidinger, L., Uesato, J., Rauh, M., Griffin, C., Huang, P.S., Mellor, J., Glaese, A., Cheng, M., Balle, B., and Kasirzadeh, A. (2022, January 21\u201324). Taxonomy of Risks Posed by Language Models. Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (FAccT), Seoul, Republic of Korea.","DOI":"10.1145\/3531146.3533088"},{"key":"ref_210","doi-asserted-by":"crossref","first-page":"171","DOI":"10.1109\/TTS.2023.3257303","article-title":"AI Ethics Principles in Practice: Perspectives of Designers and Developers","volume":"4","author":"Sanderson","year":"2023","journal-title":"IEEE Trans. Technol. Soc."},{"key":"ref_211","doi-asserted-by":"crossref","unstructured":"Phattanaviroj, T., Moslehpour, M., and Walawalkar, A.M. (2024). Data Ethics and Privacy. Challenges in Large Language Model Development and AI Ethics, IGI Global.","DOI":"10.4018\/979-8-3693-3860-5.ch010"},{"key":"ref_212","doi-asserted-by":"crossref","first-page":"109698","DOI":"10.1016\/j.compeleceng.2024.109698","article-title":"Privacy Issues in Large Language Models: A Survey","volume":"120","author":"Kibriya","year":"2024","journal-title":"Comput. Electr. Eng."},{"key":"ref_213","doi-asserted-by":"crossref","first-page":"1","DOI":"10.4236\/jsea.2024.171001","article-title":"Whispered Tuning: Data Privacy Preservation in Fine-Tuning LLMs through Differential Privacy","volume":"17","author":"Singh","year":"2024","journal-title":"J. Softw. Eng. Appl."},{"key":"ref_214","unstructured":"Charles, Z., Ganesh, A., McKenna, R., McMahan, H.B., Mitchell, N., Pillutla, K., and Rush, K. (2024). Fine-Tuning Large Language Models with User-Level Differential Privacy. arXiv."},{"key":"ref_215","doi-asserted-by":"crossref","unstructured":"Kuang, W., Qian, B., Li, Z., Chen, D., Gao, D., Pan, X., Xie, Y., Li, Y., Ding, B., and Zhou, J. (2024, January 25\u201329). FederatedScope-LLM: A Comprehensive Package for Fine-Tuning Large Language Models in Federated Learning. Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD), Barcelona, Spain.","DOI":"10.1145\/3637528.3671573"},{"key":"ref_216","doi-asserted-by":"crossref","unstructured":"Wiest, I.C., Lessmann, M.E., Wolf, F., Ferber, D., Van Treeck, M., Zhu, J., Ebert, M.P., Westphalen, C.B., Wermke, M., and Kather, J.N. (2024). Anonymizing Medical Documents with Local, Privacy Preserving Large Language Models: The LLM-Anonymizer. medRxiv.","DOI":"10.1101\/2024.06.11.24308355"},{"key":"ref_217","doi-asserted-by":"crossref","unstructured":"Li, Q., Hong, J., Xie, C., Tan, J., Xin, R., Hou, J., Yin, X., Wang, Z., Hendrycks, D., and Wang, Z. (2024). LLM-PBE: Assessing Data Privacy in Large Language Models. arXiv.","DOI":"10.14778\/3681954.3681994"}],"container-title":["Computation"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2079-3197\/14\/6\/125\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,29]],"date-time":"2026-05-29T10:28:02Z","timestamp":1780050482000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2079-3197\/14\/6\/125"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,29]]},"references-count":217,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2026,6]]}},"alternative-id":["computation14060125"],"URL":"https:\/\/doi.org\/10.3390\/computation14060125","relation":{},"ISSN":["2079-3197"],"issn-type":[{"value":"2079-3197","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,5,29]]}}}