{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T13:48:04Z","timestamp":1782308884299,"version":"3.54.5"},"reference-count":283,"publisher":"Association for Computing Machinery (ACM)","issue":"13","license":[{"start":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T00:00:00Z","timestamp":1782259200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc-nd\/4.0\/legalcode"}],"funder":[{"name":"National Research Foundation, Singapore, and Cyber Security Agency of Singapore under its National Cybersecurity R&D Programme and CyberSG R&D Cyber Research Programme Office"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2026,10,31]]},"abstract":"<jats:p>The rapid deployment of large language and vision models in real-world applications has intensified the need to address hallucinations\u2014instances where models generate incorrect or incoherent outputs. These failures can spread misinformation and degrade workflows, causing financial and operational harm. Despite extensive research efforts, our understanding of hallucinations remains limited and fragmented. Without clear understanding, solutions risk addressing disparate symptoms rather than root causes, which undermines their effectiveness and generalisability during deployment. To address this, we first introduce a unified, multi-level framework to characterise both image and text hallucinations across broad applications, helping reduce conceptual fragmentation. Then, we trace their root causes to identifiable mechanisms within a model\u2019s lifecycle in a task-modality interleaved manner, fostering a deeper and more holistic understanding. Our investigations reveal hallucinations as predictable consequences of underlying distributions and biases. By enhancing our understanding of hallucinations, this survey lays the groundwork for more effective solutions to hallucinations in generative AI systems.<\/jats:p>","DOI":"10.1145\/3811409","type":"journal-article","created":{"date-parts":[[2026,4,27]],"date-time":"2026-04-27T11:26:17Z","timestamp":1777289177000},"page":"1-36","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Understanding Hallucinations in Large Visual and Language Models"],"prefix":"10.1145","volume":"58","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-2116-8133","authenticated-orcid":false,"given":"Zheng Yi","family":"Ho","sequence":"first","affiliation":[{"name":"Generative AI Lab, College of Computing and Data Science, Nanyang Technological University","place":["Singapore, Singapore"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6154-0233","authenticated-orcid":false,"given":"Siyuan","family":"Liang","sequence":"additional","affiliation":[{"name":"Generative AI Lab, College of Computing and Data Science, Nanyang Technological University","place":["Singapore, Singapore"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7225-5449","authenticated-orcid":false,"given":"Dacheng","family":"Tao","sequence":"additional","affiliation":[{"name":"Generative AI Lab, College of Computing and Data Science, Nanyang Technological University","place":["Singapore, Singapore"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,24]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-4278"},{"key":"e_1_3_2_3_2","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Akter Syeda Nahida","year":"2025","unstructured":"Syeda Nahida Akter, Shrimai Prabhumoye, John Kamalu, Sanjeev Satheesh, Eric Nyberg, Mostofa Patwary, Mohammad Shoeybi, and Bryan Catanzaro. 2025. MIND: Math informed synthetic dialogues for pretraining LLMs. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=TuOTSAiHDn"},{"key":"e_1_3_2_4_2","volume-title":"The Twelfth International Conference on Learning Representations","author":"Alemohammad Sina","year":"2024","unstructured":"Sina Alemohammad, Josue Casco-Rodriguez, Lorenzo Luzi, Ahmed Imtiaz Humayun, Hossein Babaei, Daniel LeJeune, Ali Siahkoohi, and Richard Baraniuk. 2024. Self-consuming generative models go MAD. In The Twelfth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=ShjMHfmPs0"},{"key":"e_1_3_2_5_2","volume-title":"The Thirty-Ninth Annual Conference on Neural Information Processing Systems","author":"Alido Jeffrey","year":"2025","unstructured":"Jeffrey Alido, Tongyu Li, Yu Sun, and Lei Tian. 2025. Whitened score diffusion: A structured prior for imaging inverse problems. In The Thirty-Ninth Annual Conference on Neural Information Processing Systems. Retrieved from https:\/\/openreview.net\/forum?id=FT6VHWYdvd"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","unstructured":"Wenbin An Feng Tian Sicong Leng Jiahao Nie Haonan Lin QianYing Wang Ping Chen Xiaoqin Zhang and Shijian Lu. 2024. Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention. DOI:10.48550\/ARXIV.2406.12718","DOI":"10.48550\/ARXIV.2406.12718"},{"key":"e_1_3_2_7_2","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Ananthram Amith","year":"2025","unstructured":"Amith Ananthram, Elias Stengel-Eskin, Mohit Bansal, and Kathleen McKeown. 2025. See it from my perspective: How language affects cultural bias in image understanding. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=Xbl6t6zxZs"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.findings-acl.58"},{"key":"e_1_3_2_9_2","unstructured":"Rim Assouel Pietro Astolfi Florian Bordes Michal Drozdzal and Adriana Romero-Soriano. 2025. Object-centric Binding in Contrastive Language-Image Pretraining. arxiv:2502.14113 [cs.CV]. https:\/\/arxiv.org\/abs\/2502.14113"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41598-022-20918-w"},{"key":"e_1_3_2_11_2","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Azizi Seyedarmin","year":"2025","unstructured":"Seyedarmin Azizi, Souvik Kundu, Mohammad Erfan Sadeghi, and Massoud Pedram. 2025. MambaExtend: A training-free approach to improve long context extension of mamba. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=LgzRo1RpLS"},{"key":"e_1_3_2_12_2","unstructured":"Yunpeng Bai Haoxiang Li and Qixing Huang. 2025. Positional Encoding Field. arxiv:2510.20385 [cs.CV]. https:\/\/arxiv.org\/abs\/2510.20385"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","unstructured":"Zechen Bai Pichao Wang Tianjun Xiao Tong He Zongbo Han Zheng Zhang and Mike Zheng Shou. 2024. Hallucination of Multimodal Large Language Models: A Survey. DOI:10.48550\/ARXIV.2404.18930","DOI":"10.48550\/ARXIV.2404.18930"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.eacl-long.5"},{"key":"e_1_3_2_15_2","volume-title":"The Twelfth International Conference on Learning Representations","author":"Bansal Hritik","year":"2024","unstructured":"Hritik Bansal, John Dang, and Aditya Grover. 2024. Peering through preferences: Unraveling feedback acquisition for aligning large language models. In The Twelfth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=dKl6lMwbCy"},{"key":"e_1_3_2_16_2","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Ben-Kish Assaf","year":"2025","unstructured":"Assaf Ben-Kish, Itamar Zimerman, Shady Abu-Hussein, Nadav Cohen, Amir Globerson, Lior Wolf, and Raja Giryes. 2025. DeciMamba: Exploring the length extrapolation potential of mamba. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=iWSl5Zyjjw"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i10.28981"},{"key":"e_1_3_2_18_2","volume-title":"The Twelfth International Conference on Learning Representations","author":"Berglund Lukas","year":"2024","unstructured":"Lukas Berglund, Meg Tong, Maximilian Kaufmann, Mikita Balesni, Asa Cooper Stickland, Tomasz Korbak, and Owain Evans. 2024. The reversal curse: LLMs trained on \u201cA is B\u201d fail to learn \u201cB is A\u201d. In The Twelfth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=GPKTIktA0k"},{"key":"e_1_3_2_19_2","volume-title":"The Twelfth International Conference on Learning Representations","author":"Bertrand Quentin","year":"2024","unstructured":"Quentin Bertrand, Joey Bose, Alexandre Duplessis, Marco Jiralerspong, and Gauthier Gidel. 2024. On the stability of iterative retraining of generative models on their own data. In The Twelfth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=JORAfH2xFd"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.52202\/075280-1219"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.52202\/075280-1169"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","unstructured":"Anna Bodonhelyi Efe Bozkir Shuo Yang Enkelejda Kasneci and Gjergji Kasneci. 2024. User Intent Recognition and Satisfaction with Large Language Models: A User Study with ChatGPT. DOI:10.48550\/ARXIV.2402.02136","DOI":"10.48550\/ARXIV.2402.02136"},{"key":"e_1_3_2_23_2","series-title":"ICML\u201924","volume-title":"Proceedings of the 41st International Conference on Machine Learning","author":"Bombari Simone","year":"2024","unstructured":"Simone Bombari and Marco Mondelli. 2024. How spurious features are memorized: precise analysis for random and NTK features. In Proceedings of the 41st International Conference on Machine Learning (Vienna, Austria) (ICML\u201924). JMLR.org, Article 171, 33 pages."},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","unstructured":"Yelysei Bondarenko Markus Nagel and Tijmen Blankevoort. 2023. Quantizable Transformers: Removing Outliers by Helping Attention Heads Do Nothing. DOI:10.48550\/ARXIV.2306.12929","DOI":"10.48550\/ARXIV.2306.12929"},{"key":"e_1_3_2_25_2","volume-title":"The Thirty-Ninth Annual Conference on Neural Information Processing Systems","author":"Bonnaire Tony","year":"2025","unstructured":"Tony Bonnaire, Rapha\u00ebl Urfin, Giulio Biroli, and Marc Mezard. 2025. Why diffusion models don\u2019t Memorize: The role of implicit dynamical regularization in training. In The Thirty-Ninth Annual Conference on Neural Information Processing Systems. Retrieved from https:\/\/openreview.net\/forum?id=BSZqpqgqM0"},{"key":"e_1_3_2_26_2","unstructured":"Sebastian Bordt Suraj Srinivas Valentyn Boreiko and Ulrike von Luxburg. 2025. How Much Can We Forget about Data Contamination?arxiv:2410.03249 [cs.LG]. https:\/\/arxiv.org\/abs\/2410.03249"},{"key":"e_1_3_2_27_2","volume-title":"Microsoft FY24 Fourth Quarter Earnings Conference Call","author":"Nadella Amy Hood Brett Iversen, Satya","year":"2024","unstructured":"Amy Hood Brett Iversen, Satya Nadella. 2024. Microsoft FY24 Fourth Quarter Earnings Conference Call. Retrieved from https:\/\/view.officeapps.live.com\/op\/view.aspx?src=https:\/\/cdn-dynmedia-1.microsoft.com\/is\/content\/microsoftcorp\/TranscriptFY24Q4. Accessed: 1 Jan 2025."},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2311.16822"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-2024"},{"key":"e_1_3_2_30_2","unstructured":"Zhongteng Cai Yaxuan Wang Yang Liu and Xueru Zhang. 2025. Stabilizing Self-Consuming Diffusion Models with Latent Space Filtering. arxiv:2511.12742 [cs.LG]. https:\/\/arxiv.org\/abs\/2511.12742"},{"key":"e_1_3_2_31_2","first-page":"5253","volume-title":"32nd USENIX Security Symposium (USENIX Security 23)","author":"Carlini Nicolas","year":"2023","unstructured":"Nicolas Carlini, Jamie Hayes, Milad Nasr, Matthew Jagielski, Vikash Sehwag, Florian Tramer, Borja Balle, Daphne Ippolito, and Eric Wallace. 2023. Extracting training data from diffusion models. In 32nd USENIX Security Symposium (USENIX Security 23). 5253\u20135270."},{"key":"e_1_3_2_32_2","volume-title":"Forty-First International Conference on Machine Learning","author":"Chakraborty Souradip","year":"2024","unstructured":"Souradip Chakraborty, Jiahao Qiu, Hui Yuan, Alec Koppel, Dinesh Manocha, Furong Huang, Amrit Bedi, and Mengdi Wang. 2024. MaxMin-RLHF: Alignment with diverse human preferences. In Forty-First International Conference on Machine Learning. Retrieved from https:\/\/openreview.net\/forum?id=8tzjEMF0Vq"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-1939"},{"key":"e_1_3_2_34_2","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Chang Yingshan","year":"2025","unstructured":"Yingshan Chang and Yonatan Bisk. 2025. Language models need inductive biases to count inductively. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=s3IBHTTDYl"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.findings-emnlp.665"},{"key":"e_1_3_2_36_2","unstructured":"Cong Chen Mingyu Liu Chenchen Jing Yizhou Zhou Fengyun Rao Hao Chen Bo Zhang and Chunhua Shen. 2025. PerturboLLaVA: Reducing Multimodal Hallucinations with Perturbative Visual Training. arxiv:2503.06486 [cs.CV]. https:\/\/arxiv.org\/abs\/2503.06486"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-1844"},{"key":"e_1_3_2_38_2","series-title":"ICML\u201924","volume-title":"Proceedings of the 41st International Conference on Machine Learning","author":"Chen Dongping","year":"2024","unstructured":"Dongping Chen, Ruoxi Chen, Shilin Zhang, Yaochen Wang, Yinuo Liu, Huichi Zhou, Qihui Zhang, Yao Wan, Pan Zhou, and Lichao Sun. 2024. MLLM-as-a-judge: Assessing multimodal LLM-as-a-judge with vision-language benchmark. In Proceedings of the 41st International Conference on Machine Learning (Vienna, Austria) (ICML\u201924). JMLR.org, Article 254, 34 pages."},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.emnlp-main.474"},{"key":"e_1_3_2_40_2","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Chen Haoxian","year":"2025","unstructured":"Haoxian Chen, Hanyang Zhao, Henry Lam, David Yao, and Wenpin Tang. 2025. MallowsPO: Fine-tune your LLM with preference dispersions. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=d8cnezVcaW"},{"key":"e_1_3_2_41_2","volume-title":"Forty-Second International Conference on Machine Learning","author":"Chen Jinpeng","year":"2025","unstructured":"Jinpeng Chen, Runmin Cong, Yuzhi Zhao, Hongzheng Yang, Guangneng Hu, Horace Ip, and Sam Kwong. 2025. SEFE: Superficial and essential forgetting eliminator for multimodal continual instruction tuning. In Forty-Second International Conference on Machine Learning. Retrieved from https:\/\/openreview.net\/forum?id=teJdFzLnKh"},{"key":"e_1_3_2_42_2","unstructured":"Junzhe Chen Tianshu Zhang Shiyu Huang Yuwei Niu Linfeng Zhang Lijie Wen and Xuming Hu. 2024. ICT: Image-Object Cross-Level Trusted Intervention for Mitigating Object Hallucination in Large Vision-Language Models. arxiv:2411.15268 [cs.CV]. https:\/\/arxiv.org\/abs\/2411.15268"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-0850"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","unstructured":"Shiqi Chen Tongyao Zhu Ruochen Zhou Jinghan Zhang Siyang Gao Juan Carlos Niebles Mor Geva Junxian He Jiajun Wu and Manling Li. 2025. Why Is Spatial Reasoning Hard for VLMs? An Attention Mechanism Perspective on Focus Areas. DOI:10.48550\/ARXIV.2503.01773","DOI":"10.48550\/ARXIV.2503.01773"},{"key":"e_1_3_2_45_2","volume-title":"The Twelfth International Conference on Learning Representations","author":"Chen Weize","year":"2024","unstructured":"Weize Chen, Yusheng Su, Jingwei Zuo, Cheng Yang, Chenfei Yuan, Chi-Min Chan, Heyang Yu, Yaxi Lu, Yi-Hsin Hung, Chen Qian, et\u00a0al. 2024. AgentVerse: Facilitating multi-agent collaboration and exploring emergent behaviors. In The Twelfth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=EHg5GDnyq1"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-1409"},{"key":"e_1_3_2_47_2","volume-title":"ICML 2024 Workshop on Foundation Models in the Wild","author":"Chen Zhaorun","year":"2024","unstructured":"Zhaorun Chen, Yichao Du, Zichen Wen, Yiyang Zhou, Chenhang Cui, Zhenzhen Weng, Haoqin Tu, Chaoqi Wang, Zhengwei Tong, Leria HUANG, et\u00a0al. 2024. MJ-bench: Is your multimodal reward model really a good judge?. In ICML 2024 Workshop on Foundation Models in the Wild. Retrieved from https:\/\/openreview.net\/forum?id=H6eELDnYvd"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-3064"},{"key":"e_1_3_2_49_2","unstructured":"Ziheng Chi Yifan Hou Chenxi Pang Shaobo Cui Mubashara Akhtar and Mrinmaya Sachan. 2025. Chimera: Diagnosing Shortcut Learning in Visual-Language Understanding. arxiv:2509.22437 [cs.CL]. https:\/\/arxiv.org\/abs\/2509.22437"},{"key":"e_1_3_2_50_2","volume-title":"Forty-Second International Conference on Machine Learning","author":"Choi Hyeong Kyu","year":"2025","unstructured":"Hyeong Kyu Choi, Maxim Khanov, Hongxin Wei, and Yixuan Li. 2025. How contaminated is your benchmark? Measuring dataset leakage in large language models with kernel divergence. In Forty-Second International Conference on Machine Learning. Retrieved from https:\/\/openreview.net\/forum?id=wVDR2qmE28"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCE-Asia63397.2024.10773636"},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-2942"},{"key":"e_1_3_2_53_2","series-title":"Proceedings of Machine Learning Research","first-page":"9346","volume-title":"Proceedings of the 41st International Conference on Machine Learning","volume":"235","author":"Conitzer Vincent","year":"2024","unstructured":"Vincent Conitzer, Rachel Freedman, Jobst Heitzig, Wesley H. Holliday, Bob M. Jacobs, Nathan Lambert, Milan Mosse, Eric Pacuit, Stuart Russell, Hailey Schoelkopf, et\u00a0al. 2024. Position: Social choice should guide AI alignment in dealing with diverse human feedback. In Proceedings of the 41st International Conference on Machine Learning(Proceedings of Machine Learning Research, Vol. 235), Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp (Eds.). PMLR, 9346\u20139360. Retrieved from https:\/\/proceedings.mlr.press\/v235\/conitzer24a.html"},{"key":"e_1_3_2_54_2","volume-title":"The Twelfth International Conference on Learning Representations","author":"Darcet Timoth\u00e9e","year":"2024","unstructured":"Timoth\u00e9e Darcet, Maxime Oquab, Julien Mairal, and Piotr Bojanowski. 2024. Vision transformers need registers. In The Twelfth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=2dnO3LLiJ1"},{"key":"e_1_3_2_55_2","unstructured":"Shounak Datta and Dhanasekar Sundararaman. 2025. Evaluating Hallucination in Large Vision-Language Models based on Context-Aware Object Similarities. arxiv:2501.15046 [cs.CV]. https:\/\/arxiv.org\/abs\/2501.15046"},{"key":"e_1_3_2_56_2","series-title":"ICML\u201924","volume-title":"Proceedings of the 41st International Conference on Machine Learning","author":"Ding Meng","year":"2025","unstructured":"Meng Ding, Kaiyi Ji, Di Wang, and Jinhui Xu. 2025. Understanding forgetting in continual learning with linear regression. In Proceedings of the 41st International Conference on Machine Learning (Vienna, Austria) (ICML\u201924). JMLR.org, Article 436, 24 pages."},{"key":"e_1_3_2_57_2","volume-title":"The Twelfth International Conference on Learning Representations","author":"Ding Nan","year":"2024","unstructured":"Nan Ding, Tomer Levinboim, Jialin Wu, Sebastian Goodman, and Radu Soricut. 2024. CausalLM is not optimal for in-context learning. In The Twelfth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=guRNebwZBb"},{"key":"e_1_3_2_58_2","volume-title":"Forty-First International Conference on Machine Learning","author":"Dohmatob Elvis","year":"2024","unstructured":"Elvis Dohmatob, Yunzhen Feng, Pu Yang, Francois Charton, and Julia Kempe. 2024. A tale of tails: Model collapse as a change of scaling laws. In Forty-First International Conference on Machine Learning. Retrieved from https:\/\/openreview.net\/forum?id=KVvku47shW"},{"key":"e_1_3_2_59_2","unstructured":"Mischa Dombrowski Weitong Zhang Sarah Cechnicka Hadrien Reynaud and Bernhard Kainz. 2024. Image Generation Diversity Issues and How to Tame Them. arxiv:2411.16171 [cs.CV]. https:\/\/arxiv.org\/abs\/2411.16171"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2025.acl-long.1595"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","unstructured":"Hongyuan Dong Jiawen Li Bohong Wu Jiacong Wang Yuan Zhang and Haoyuan Guo. 2024. Benchmarking and Improving Detail Image Caption. DOI:10.48550\/ARXIV.2405.19092","DOI":"10.48550\/ARXIV.2405.19092"},{"key":"e_1_3_2_62_2","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Dorner Florian E.","year":"2025","unstructured":"Florian E. Dorner, Vivian Yvonne Nastl, and Moritz Hardt. 2025. Limits to scalable evaluation at the frontier: LLM as judge won\u2019t beat twice the data. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=NO6Tv6QcDs"},{"key":"e_1_3_2_63_2","volume-title":"Forty-Second International Conference on Machine Learning","author":"Dosi Muskan","year":"2025","unstructured":"Muskan Dosi, Chiranjeev Chiranjeev, Kartik Thakral, Mayank Vatsa, and Richa Singh. 2025. Harmonizing geometry and uncertainty: Diffusion with hyperspheres. In Forty-Second International Conference on Machine Learning. Retrieved from https:\/\/openreview.net\/forum?id=A82tIFgJaK"},{"key":"e_1_3_2_64_2","unstructured":"Aleksandr Dremov Alexander H\u00e4gele Atli Kosson and Martin Jaggi. 2025. Training Dynamics of the Cooldown Stage in Warmup-Stable-Decay Learning Rate Scheduler. arxiv:2508.01483 [cs.LG]. https:\/\/arxiv.org\/abs\/2508.01483"},{"key":"e_1_3_2_65_2","volume-title":"First Conference on Language Modeling","author":"Dubois Yann","year":"2024","unstructured":"Yann Dubois, Percy Liang, and Tatsunori Hashimoto. 2024. Length-controlled alpacaeval: A simple debiasing of automatic evaluators. In First Conference on Language Modeling. Retrieved from https:\/\/openreview.net\/forum?id=CybBmzWBX0"},{"key":"e_1_3_2_66_2","volume-title":"The Thirty-Ninth Annual Conference on Neural Information Processing Systems","author":"Dunlop Connor","year":"2025","unstructured":"Connor Dunlop, Matthew Zheng, Kavana Venkatesh, and Pinar Yanardag. 2025. Personalized image editing in text-to-image diffusion models via collaborative direct preference optimization. In The Thirty-Ninth Annual Conference on Neural Information Processing Systems. Retrieved from https:\/\/openreview.net\/forum?id=BBZEcVu1nA"},{"key":"e_1_3_2_67_2","article-title":"Faith and fate: Limits of transformers on compositionality","volume":"36","author":"Dziri Nouha","year":"2024","unstructured":"Nouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li, Liwei Jiang, Bill Yuchen Lin, Sean Welleck, Peter West, Chandra Bhagavatula, Ronan Le Bras, et\u00a0al. 2024. Faith and fate: Limits of transformers on compositionality. Advances in Neural Information Processing Systems 36 (2024).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_68_2","doi-asserted-by":"publisher","DOI":"10.3390\/computers14010019"},{"key":"e_1_3_2_69_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-0911"},{"key":"e_1_3_2_70_2","doi-asserted-by":"publisher","DOI":"10.1109\/cvpr52734.2025.02353"},{"key":"e_1_3_2_71_2","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Fang Lizhe","year":"2025","unstructured":"Lizhe Fang, Yifei Wang, Zhaoyang Liu, Chenheng Zhang, Stefanie Jegelka, Jinyang Gao, Bolin Ding, and Yisen Wang. 2025. What is wrong with perplexity for long-context language modeling?. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=fL4qWkSmtM"},{"key":"e_1_3_2_72_2","volume-title":"Forty-second International Conference on Machine Learning","author":"Farquhar Sebastian","year":"2025","unstructured":"Sebastian Farquhar, Vikrant Varma, David Lindner, David Elson, Caleb Biddulph, Ian Goodfellow, and Rohin Shah. 2025. MONA: Myopic optimization with non-myopic approval can mitigate multi-step reward hacking. In Forty-second International Conference on Machine Learning. Retrieved from https:\/\/openreview.net\/forum?id=fyi34BxCwq"},{"key":"e_1_3_2_73_2","doi-asserted-by":"publisher","DOI":"10.1090\/jams\/852"},{"key":"e_1_3_2_74_2","volume-title":"Forty-Second International Conference on Machine Learning","author":"Feng Jiahai","year":"2025","unstructured":"Jiahai Feng, Stuart Russell, and Jacob Steinhardt. 2025. Extractive structures learned in pretraining enable generalization on finetuned facts. In Forty-Second International Conference on Machine Learning. Retrieved from https:\/\/openreview.net\/forum?id=W0GrWqqTJo"},{"key":"e_1_3_2_75_2","volume-title":"Forty-Second International Conference on Machine Learning","author":"Feng Tao","year":"2025","unstructured":"Tao Feng, Wei Li, Didi Zhu, Hangjie Yuan, Wendi Zheng, Dan Zhang, and Jie Tang. 2025. ZeroFlow: Overcoming catastrophic forgetting is easier than you think. In Forty-Second International Conference on Machine Learning. Retrieved from https:\/\/openreview.net\/forum?id=iPDw3O6u3T"},{"key":"e_1_3_2_76_2","series-title":"ICML\u201924","volume-title":"Proceedings of the 41st International Conference on Machine Learning","author":"Fernando Chrisantha","year":"2024","unstructured":"Chrisantha Fernando, Dylan Banarse, Henryk Michalewski, Simon Osindero, and Tim Rockt\u00e4schel. 2024. Promptbreeder: Self-referential self-improvement via prompt evolution. In Proceedings of the 41st International Conference on Machine Learning (Vienna, Austria) (ICML\u201924). JMLR.org, Article 541, 64 pages."},{"key":"e_1_3_2_77_2","doi-asserted-by":"publisher","unstructured":"Shi Fu Sen Zhang Yingjie Wang Xinmei Tian and Dacheng Tao. 2024. Towards Theoretical Understandings of Self-Consuming Generative Models. DOI:10.48550\/ARXIV.2402.11778","DOI":"10.48550\/ARXIV.2402.11778"},{"key":"e_1_3_2_78_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-0528"},{"key":"e_1_3_2_79_2","volume-title":"Workshop on Spurious Correlation and Shortcut Learning: Foundations and Solutions","author":"Gaudi Sachit","year":"2025","unstructured":"Sachit Gaudi, Gautam Sreekumar, and Vishnu Boddeti. 2025. Spurious correlations in diffusion models and how to fix them. In Workshop on Spurious Correlation and Shortcut Learning: Foundations and Solutions. Retrieved from https:\/\/openreview.net\/forum?id=lQqYpU70ul"},{"key":"e_1_3_2_80_2","volume-title":"First Conference on Language Modeling","author":"Gerstgrasser Matthias","year":"2024","unstructured":"Matthias Gerstgrasser, Rylan Schaeffer, Apratim Dey, Rafael Rafailov, Tomasz Korbak, Henry Sleight, Rajashree Agrawal, John Hughes, Dhruv Bhandarkar Pai, Andrey Gromov, et\u00a0al. 2024. Is model collapse inevitable? Breaking the curse of recursion by accumulating real and synthetic data. In First Conference on Language Modeling. Retrieved from https:\/\/openreview.net\/forum?id=5B2K4LRgmz"},{"key":"e_1_3_2_81_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-main.751"},{"key":"e_1_3_2_82_2","series-title":"Proceedings of Machine Learning Research","first-page":"15559","volume-title":"Proceedings of the 41st International Conference on Machine Learning","volume":"235","author":"Ghosh Sreyan","year":"2024","unstructured":"Sreyan Ghosh, Chandra Kiran Reddy Evuru, Sonal Kumar, Ramaneswaran S., Deepali Aneja, Zeyu Jin, Ramani Duraiswami, and Dinesh Manocha. 2024. A closer look at the limitations of instruction tuning. In Proceedings of the 41st International Conference on Machine Learning(Proceedings of Machine Learning Research, Vol. 235), Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp (Eds.). PMLR, 15559\u201315589. Retrieved from https:\/\/proceedings.mlr.press\/v235\/ghosh24a.html"},{"key":"e_1_3_2_83_2","volume-title":"First Conference on Language Modeling","author":"Golovneva Olga","year":"2024","unstructured":"Olga Golovneva, Zeyuan Allen-Zhu, Jason E. Weston, and Sainbayar Sukhbaatar. 2024. Reverse training to nurse the reversal curse. In First Conference on Language Modeling. Retrieved from https:\/\/openreview.net\/forum?id=HDkNbfLQgu"},{"key":"e_1_3_2_84_2","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Gong Zixuan","year":"2025","unstructured":"Zixuan Gong, Xiaolin Hu, Huayi Tang, and Yong Liu. 2025. Towards auto-regressive next-token prediction: In-context learning emerges from generalization. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=gK1rl98VRp"},{"key":"e_1_3_2_85_2","doi-asserted-by":"publisher","DOI":"10.1098\/rsta.2017.0237"},{"key":"e_1_3_2_86_2","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Goyal Sachin","year":"2025","unstructured":"Sachin Goyal, Christina Baek, J. Zico Kolter, and Aditi Raghunathan. 2025. Context-parametric inversion: Why instruction finetuning may not actually improve context reliance. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=SPS6HzVzyt"},{"key":"e_1_3_2_87_2","article-title":"Studying large language model generalization with influence functions","author":"Grosse Roger","year":"2023","unstructured":"Roger Grosse, Juhan Bae, Cem Anil, Nelson Elhage, Alex Tamkin, Amirhossein Tajdini, Benoit Steiner, Dustin Li, Esin Durmus, Ethan Perez, et\u00a0al. 2023. Studying large language model generalization with influence functions. arXiv preprint arXiv:arxiv 2308.03296 (2023).","journal-title":"arXiv preprint"},{"key":"e_1_3_2_88_2","doi-asserted-by":"publisher","DOI":"10.3390\/jtaer19030108"},{"key":"e_1_3_2_89_2","volume-title":"The Thirty-ninth Annual Conference on Neural Information Processing Systems","author":"Guo Ping","year":"2025","unstructured":"Ping Guo, Yubing Ren, BINBINLIU, Fengze Liu, Haobin Lin, Yifan Zhang, Bingni Zhang, Taifeng Wang, and Yin Zheng. 2025. Exploring polyglot harmony: On multilingual data allocation for large language models pretraining. In The Thirty-ninth Annual Conference on Neural Information Processing Systems. Retrieved from https:\/\/openreview.net\/forum?id=mHHrnCWwrD"},{"key":"e_1_3_2_90_2","unstructured":"Wenqi Marshall Guo Qingyun Qian Khalad Hasan and Shan Du. 2026. Position: Universal Aesthetic Alignment Narrows Artistic Expression. arxiv:2512.11883 [cs.CY]. https:\/\/arxiv.org\/abs\/2512.11883"},{"key":"e_1_3_2_91_2","doi-asserted-by":"publisher","DOI":"10.3389\/frai.2025.1697139"},{"key":"e_1_3_2_92_2","doi-asserted-by":"publisher","DOI":"10.1109\/CogSIMA61085.2024.10553755"},{"key":"e_1_3_2_93_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-2515"},{"key":"e_1_3_2_94_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-0-387-84858-7"},{"key":"e_1_3_2_95_2","volume-title":"The Thirty-Ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track","author":"Hayes Kevin David","year":"2025","unstructured":"Kevin David Hayes, Micah Goldblum, Vikash Sehwag, Gowthami Somepalli, Ashwinee Panda, and Tom Goldstein. 2025. FineGRAIN: Evaluating failure modes of text-to-image models with vision language model judges. In The Thirty-Ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track. Retrieved from https:\/\/openreview.net\/forum?id=qlZI9Bgxpy"},{"key":"e_1_3_2_96_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-0979"},{"key":"e_1_3_2_97_2","volume-title":"The Twelfth International Conference on Learning Representations","author":"Hermann Katherine","year":"2024","unstructured":"Katherine Hermann, Hossein Mobahi, Thomas FEL, and Michael Curtis Mozer. 2024. On the foundations of shortcut learning. In The Twelfth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=Tj3xLVuE9f"},{"key":"e_1_3_2_98_2","volume-title":"The Twelfth International Conference on Learning Representations","author":"Hosking Tom","year":"2024","unstructured":"Tom Hosking, Phil Blunsom, and Max Bartolo. 2024. Human feedback is not gold standard. In The Twelfth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=7W3GLNImfS"},{"key":"e_1_3_2_99_2","volume-title":"Workshop on Spurious Correlation and Shortcut Learning: Foundations and Solutions","author":"Hosseini Parsa","year":"2025","unstructured":"Parsa Hosseini, Sumit Nawathe, Mazda Moayeri, Sriram Balasubramanian, and Soheil Feizi. 2025. SpurLens: Spurious correlations in multimodal LLMs. In Workshop on Spurious Correlation and Shortcut Learning: Foundations and Solutions. Retrieved from https:\/\/openreview.net\/forum?id=SFtrsxKH79"},{"key":"e_1_3_2_100_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.findings-acl.890"},{"key":"e_1_3_2_101_2","unstructured":"Yihan Hu Jianing Peng Yiheng Lin Ting Liu Xiaochao Qu Luoqi Liu Yao Zhao and Yunchao Wei. 2025. DCEdit: Dual-Level Controlled Image Editing via Precisely Localized Semantics. arxiv:2503.16795 [cs.CV]. https:\/\/arxiv.org\/abs\/2503.16795"},{"key":"e_1_3_2_102_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-4423"},{"key":"e_1_3_2_103_2","volume-title":"Forty-Second International Conference on Machine Learning","author":"Huang Zizheng","year":"2025","unstructured":"Zizheng Huang, Haoxing Chen, Jiaqi Li, jun lan, Huijia Zhu, Weiqiang Wang, and Limin Wang. 2025. Stochastic layer-wise shuffle for improving vision mamba training. In Forty-Second International Conference on Machine Learning. Retrieved from https:\/\/openreview.net\/forum?id=6vP20U55bA"},{"key":"e_1_3_2_104_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10994-021-05946-3"},{"key":"e_1_3_2_105_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2025.emnlp-main.1182"},{"key":"e_1_3_2_106_2","unstructured":"Noam Issachar Guy Yariv Sagie Benaim Yossi Adi Dani Lischinski and Raanan Fattal. 2026. DyPE: Dynamic Position Extrapolation for Ultra High Resolution Diffusion. arxiv:2510.20766 [cs.CV]. https:\/\/arxiv.org\/abs\/2510.20766"},{"key":"e_1_3_2_107_2","volume-title":"The Thirty-Ninth Annual Conference on Neural Information Processing Systems","author":"Izadi Amirmohammad","year":"2025","unstructured":"Amirmohammad Izadi, Mohammadali Banayeeanzade, Fatemeh Askari, Ali Rahimiakbar, Mohammad Mahdi Vahedi, Hosein Hasani, and Mahdieh Soleymani Baghshah. 2025. Visual structures help visual reasoning: Addressing the binding problem in LVLMs. In The Thirty-Ninth Annual Conference on Neural Information Processing Systems. Retrieved from https:\/\/openreview.net\/forum?id=T52hZeT7rn"},{"key":"e_1_3_2_108_2","volume-title":"The Thirty-Eighth Annual Conference on Neural Information Processing Systems","author":"Jayaraman Bargav","year":"2024","unstructured":"Bargav Jayaraman, Chuan Guo, and Kamalika Chaudhuri. 2024. D\u00e9j\u00e0 Vu memorization in vision\u2013language models. In The Thirty-Eighth Annual Conference on Neural Information Processing Systems. Retrieved from https:\/\/openreview.net\/forum?id=SFCZdXDyNs"},{"key":"e_1_3_2_109_2","doi-asserted-by":"publisher","unstructured":"Sadeep Jayasumana Srikumar Ramalingam Andreas Veit Daniel Glasner Ayan Chakrabarti and Sanjiv Kumar. 2024. Rethinking FID: Towards a Better Evaluation Metric for Image Generation. DOI:10.48550\/ARXIV.2401.09603","DOI":"10.48550\/ARXIV.2401.09603"},{"key":"e_1_3_2_110_2","volume-title":"Forty-Second International Conference on Machine Learning","author":"Jeon Dongjae","year":"2025","unstructured":"Dongjae Jeon, Dueun Kim, and Albert No. 2025. Understanding and mitigating memorization in generative models via sharpness of probability landscapes. In Forty-Second International Conference on Machine Learning. Retrieved from https:\/\/openreview.net\/forum?id=EW2JR5aVLm"},{"key":"e_1_3_2_111_2","doi-asserted-by":"publisher","DOI":"10.1145\/3571730"},{"key":"e_1_3_2_112_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-0200"},{"key":"e_1_3_2_113_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-4306"},{"key":"e_1_3_2_114_2","unstructured":"Jaehun Jung Faeze Brahman and Yejin Choi. 2024. Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement. arxiv:2407.18370 [cs.LG]. https:\/\/arxiv.org\/abs\/2407.18370"},{"key":"e_1_3_2_115_2","volume-title":"The Twelfth International Conference on Learning Representations","author":"Kadkhodaie Zahra","year":"2024","unstructured":"Zahra Kadkhodaie, Florentin Guth, Eero P. Simoncelli, and St\u00e9phane Mallat. 2024. Generalization in diffusion models arises from geometry-adaptive harmonic representations. In The Twelfth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=ANvmVS2Yr0"},{"key":"e_1_3_2_116_2","doi-asserted-by":"publisher","unstructured":"Negar Kamali Karyn Nakamura Aakriti Kumar Angelos Chatzimparmpas Jessica Hullman and Matthew Groh. 2025. Characterizing Photorealism and Artifacts in Diffusion Model-Generated Images. DOI:10.48550\/ARXIV.2502.11989","DOI":"10.48550\/ARXIV.2502.11989"},{"key":"e_1_3_2_117_2","first-page":"15696","volume-title":"International Conference on Machine Learning","author":"Kandpal Nikhil","year":"2023","unstructured":"Nikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace, and Colin Raffel. 2023. Large language models struggle to learn long-tail knowledge. In International Conference on Machine Learning. PMLR, 15696\u201315707."},{"key":"e_1_3_2_118_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.02282"},{"key":"e_1_3_2_119_2","doi-asserted-by":"publisher","DOI":"10.52202\/075280-1082"},{"key":"e_1_3_2_120_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-3387"},{"key":"e_1_3_2_121_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-73004-7_6"},{"key":"e_1_3_2_122_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-3342"},{"key":"e_1_3_2_123_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.inlg-main.45"},{"key":"e_1_3_2_124_2","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Maninis Kevis kokitsi","year":"2025","unstructured":"Kevis kokitsi Maninis, Kaifeng Chen, Soham Ghosh, Arjun Karpur, Koert Chen, Ye Xia, Bingyi Cao, Daniel Salz, Guangxing Han, Jan Dlabal, et\u00a0al. 2025. TIPS: Text-image pretraining with spatial awareness. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=DaA0wAcTY7"},{"key":"e_1_3_2_125_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.findings-acl.29"},{"key":"e_1_3_2_126_2","doi-asserted-by":"publisher","unstructured":"Arjun Krishna Erick Galinkin Leon Derczynski and Jeffrey Martin. 2025. Importing Phantoms: Measuring LLM Package Hallucination Vulnerabilities. DOI:10.48550\/ARXIV.2501.19012","DOI":"10.48550\/ARXIV.2501.19012"},{"key":"e_1_3_2_127_2","volume-title":"Forty-Second International Conference on Machine Learning","author":"Kuang Yilun","year":"2025","unstructured":"Yilun Kuang, Noah Amsel, Sanae Lotfi, Shikai Qiu, Andres Potapczynski, and Andrew Gordon Wilson. 2025. Customizing the inductive biases of softmax attention using structured matrices. In Forty-Second International Conference on Machine Learning. Retrieved from https:\/\/openreview.net\/forum?id=Roc5O1ECEt"},{"key":"e_1_3_2_128_2","volume-title":"Customers are Putting Gemini to Work","author":"Kurian Thomas","year":"2024","unstructured":"Thomas Kurian. 2024. Customers are Putting Gemini to Work. Retrieved from https:\/\/blog.google\/products\/google-cloud\/gemini-at-work-ai-agents\/. Accessed: 1 Jan 2025."},{"key":"e_1_3_2_129_2","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Laidlaw Cassidy","year":"2025","unstructured":"Cassidy Laidlaw, Shivam Singhal, and Anca Dragan. 2025. Correlated proxies: A new definition and improved mitigation for reward hacking. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=msEr27EejF"},{"key":"e_1_3_2_130_2","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Laidlaw Cassidy","year":"2025","unstructured":"Cassidy Laidlaw, Shivam Singhal, and Anca Dragan. 2025. Correlated proxies: A new definition and improved mitigation for reward hacking. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=msEr27EejF"},{"key":"e_1_3_2_131_2","series-title":"Proceedings of Machine Learning Research","first-page":"26043","volume-title":"Proceedings of the 41st International Conference on Machine Learning","volume":"235","author":"Lavie Itay","year":"2024","unstructured":"Itay Lavie, Guy Gur-Ari, and Zohar Ringel. 2024. Towards understanding inductive bias in transformers: A view from infinity. In Proceedings of the 41st International Conference on Machine Learning(Proceedings of Machine Learning Research, Vol. 235), Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp (Eds.). PMLR, 26043\u201326069. Retrieved from https:\/\/proceedings.mlr.press\/v235\/lavie24a.html"},{"key":"e_1_3_2_132_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.findings-emnlp.249"},{"key":"e_1_3_2_133_2","article-title":"Enhancing diversity in text-to-image generation without compromising fidelity","author":"Li Jiazhi","year":"2025","unstructured":"Jiazhi Li, Mi Zhou, Mahyar Khayatkhoei, Jingyu Shi, Xiang Gao, Jiageng Zhu, Hanchen Xie, Xiyun Song, Zongfang Lin, Heather Yu, et\u00a0al. 2025. Enhancing diversity in text-to-image generation without compromising fidelity. Transactions on Machine Learning Research (2025). Retrieved from https:\/\/openreview.net\/forum?id=180S4tOpmx","journal-title":"Transactions on Machine Learning Research"},{"key":"e_1_3_2_134_2","unstructured":"Mingcheng Li Xiaolu Hou Ziyang Liu Dingkang Yang Ziyun Qian Jiawei Chen Jinjie Wei Yue Jiang Qingyao Xu and Lihua Zhang. 2025. MCCD: Multi-Agent Collaboration-based Compositional Diffusion for Complex Text-to-Image Generation. arxiv:2505.02648 [cs.CV]. https:\/\/arxiv.org\/abs\/2505.02648"},{"key":"e_1_3_2_135_2","volume-title":"The Thirty-Ninth Annual Conference on Neural Information Processing Systems","author":"Li Tanzhe","year":"2025","unstructured":"Tanzhe Li, Caoshuo Li, Jiayi Lyu, Hongjuan Pei, Baochang Zhang, Taisong Jin, and Rongrong Ji. 2025. DAMamba: Vision state space model with dynamic adaptive scan. In The Thirty-Ninth Annual Conference on Neural Information Processing Systems. Retrieved from https:\/\/openreview.net\/forum?id=KYFTBpKIOr"},{"key":"e_1_3_2_136_2","volume-title":"Workshop on Spurious Correlation and Shortcut Learning: Foundations and Solutions","author":"Li Yanshu","year":"2025","unstructured":"Yanshu Li. 2025. Unveiling and mitigating short-cut learning in multimodal in-context learning. In Workshop on Spurious Correlation and Shortcut Learning: Foundations and Solutions. Retrieved from https:\/\/openreview.net\/forum?id=RVVARgLdTT"},{"key":"e_1_3_2_137_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW69036.2025.00290"},{"key":"e_1_3_2_138_2","volume-title":"The Thirty-Ninth Annual Conference on Neural Information Processing Systems","author":"Lin Beibei","year":"2025","unstructured":"Beibei Lin, Tingting Chen, and Robby T. Tan. 2025. GeoComplete: Geometry-aware diffusion for reference-driven image completion. In The Thirty-Ninth Annual Conference on Neural Information Processing Systems. Retrieved from https:\/\/openreview.net\/forum?id=1EnpXg8s4v"},{"key":"e_1_3_2_139_2","volume-title":"The Eleventh International Conference on Learning Representations","author":"Liu Bingbin","year":"2023","unstructured":"Bingbin Liu, Jordan T. Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang. 2023. Transformers learn shortcuts to automata. In The Eleventh International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=De4FYqjFueZ"},{"key":"e_1_3_2_140_2","series-title":"NIPS\u201923","volume-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems","author":"Liu Bingbin","year":"2024","unstructured":"Bingbin Liu, Jordan T. Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang. 2024. Exposing attention glitches with flip-flop language modeling. In Proceedings of the 37th International Conference on Neural Information Processing Systems (New Orleans, LA, USA) (NIPS\u201923). Curran Associates Inc., Red Hook, NY, USA, Article 1112, 35 pages."},{"key":"e_1_3_2_141_2","unstructured":"Jingren Liu Shuning Xu Yun Wang Zhong Ji and Xiangyu Chen. 2025. CCD: Continual Consistency Diffusion for Lifelong Generative Modeling. arxiv:2505.11936 [cs.LG]. https:\/\/arxiv.org\/abs\/2505.11936"},{"key":"e_1_3_2_142_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00638"},{"key":"e_1_3_2_143_2","volume-title":"The Twelfth International Conference on Learning Representations","author":"Liu Qihao","year":"2024","unstructured":"Qihao Liu, Adam Kortylewski, Yutong Bai, Song Bai, and Alan Yuille. 2024. Discovering failure modes of text-guided diffusion models via adversarial search. In The Twelfth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=TOWdQQgMJY"},{"key":"e_1_3_2_144_2","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Liu Siqi","year":"2025","unstructured":"Siqi Liu, Ian Gemp, Luke Marris, Georgios Piliouras, Nicolas Heess, and Marc Lanctot. 2025. Re-evaluating open-ended evaluation of large language models. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=kbOAIXKWgx"},{"key":"e_1_3_2_145_2","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Liu Yexiang","year":"2025","unstructured":"Yexiang Liu, Jie Cao, Zekun Li, Ran He, and Tieniu Tan. 2025. Breaking mental set to improve reasoning through diverse multi-agent debate. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=t6QHYUOQL7"},{"key":"e_1_3_2_146_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.findings-acl.753"},{"key":"e_1_3_2_147_2","volume-title":"Forty-Second International Conference on Machine Learning","author":"Liu Zirui","year":"2025","unstructured":"Zirui Liu, Jiatong Li, Yan Zhuang, Qi Liu, Shuanghong Shen, Jie Ouyang, Mingyue Cheng, and Shijin Wang. 2025. am-ELO: A stable framework for arena-based LLM evaluation. In Forty-Second International Conference on Machine Learning. Retrieved from https:\/\/openreview.net\/forum?id=EUH4VUCXay"},{"key":"e_1_3_2_148_2","unstructured":"Zhongxin Liu Zhiwei Wang Jun Niu Ying Li Hongyu Sun Meng Xu He Wang Gaofei Wu and Yuqing Zhang. 2025. KSHSeek: Data-Driven Approaches to Mitigating and Detecting Knowledge-Shortcut Hallucinations in Generative Models. arxiv:2503.19482 [cs.CL]. https:\/\/arxiv.org\/abs\/2503.19482"},{"key":"e_1_3_2_149_2","volume-title":"The Thirty-Ninth Annual Conference on Neural Information Processing Systems","author":"Lu Peng","year":"2025","unstructured":"Peng Lu, Jerry Huang, QIUHAO Zeng, Xinyu Wang, Boxing Chen, Philippe Langlais, and Yufei Cui. 2025. Mamba modulation: On the length generalization of mamba models. In The Thirty-Ninth Annual Conference on Neural Information Processing Systems. Retrieved from https:\/\/openreview.net\/forum?id=QEU047bE8p"},{"key":"e_1_3_2_150_2","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Lu Rui","year":"2025","unstructured":"Rui Lu, Runzhe Wang, Kaifeng Lyu, Xitai Jiang, Gao Huang, and Mengdi Wang. 2025. Towards understanding text hallucination of diffusion models via local generation bias. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=SKW10XJlAI"},{"key":"e_1_3_2_151_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.emnlp-main.428"},{"key":"e_1_3_2_152_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51701.2025.01400"},{"key":"e_1_3_2_153_2","volume-title":"The Thirty-Ninth Annual Conference on Neural Information Processing Systems","author":"Malakouti Sina","year":"2025","unstructured":"Sina Malakouti and Adriana Kovashka. 2025. Role bias in diffusion models: Diagnosing and mitigating through intermediate decomposition. In The Thirty-Ninth Annual Conference on Neural Information Processing Systems. Retrieved from https:\/\/openreview.net\/forum?id=xpkJiQNC0E"},{"key":"e_1_3_2_154_2","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pone.0326729"},{"key":"e_1_3_2_155_2","doi-asserted-by":"publisher","unstructured":"Gonzalo Mart\u00ednez Lauren Watson Pedro Reviriego Jos\u00e9 Alberto Hern\u00e1ndez Marc Juarez and Rik Sarkar. 2023. Combining Generative Artificial Intelligence (AI) and the Internet: Heading towards Evolution or Degradation?DOI:10.48550\/ARXIV.2303.01255","DOI":"10.48550\/ARXIV.2303.01255"},{"key":"e_1_3_2_156_2","volume-title":"The Thirty-Ninth Annual Conference on Neural Information Processing Systems","author":"Masry Ahmed","year":"2025","unstructured":"Ahmed Masry, Juan A. Rodriguez, Tianyu Zhang, Suyuchen Wang, Chao Wang, Aarash Feizi, Akshay Kalkunte Suresh, Abhay Puri, Xiangru Jian, Pierre-Andre Noel, et\u00a0al. 2025. AlignVLM: Bridging vision and language latent spaces for multimodal document understanding. In The Thirty-Ninth Annual Conference on Neural Information Processing Systems. Retrieved from https:\/\/openreview.net\/forum?id=vAxGuGmshO"},{"key":"e_1_3_2_157_2","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.2322420121"},{"key":"e_1_3_2_158_2","doi-asserted-by":"publisher","DOI":"10.52202\/068431-1262"},{"key":"e_1_3_2_159_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-4270"},{"key":"e_1_3_2_160_2","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Mirzadeh Seyed Iman","year":"2025","unstructured":"Seyed Iman Mirzadeh, Keivan Alizadeh, Hooman Shahrokhi, Oncel Tuzel, Samy Bengio, and Mehrdad Farajtabar. 2025. GSM-symbolic: Understanding the limitations of mathematical reasoning in large language models. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=AjXkRZIvjB"},{"key":"e_1_3_2_161_2","volume-title":"The Twelfth International Conference on Learning Representations","author":"Moskovitz Ted","year":"2024","unstructured":"Ted Moskovitz, Aaditya K. Singh, D. J. Strouse, Tuomas Sandholm, Ruslan Salakhutdinov, Anca Dragan, and Stephen Marcus McAleer. 2024. Confronting reward model overoptimization with constrained RLHF. In The Twelfth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=gkfUvn0fLU"},{"key":"e_1_3_2_162_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-2336"},{"key":"e_1_3_2_163_2","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Movahedi Sajad","year":"2025","unstructured":"Sajad Movahedi, Antonio Orvieto, and Seyed-Mohsen Moosavi-Dezfooli. 2025. Geometric inductive biases of deep networks: The role of data and architecture. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=cmXWYolrlo"},{"key":"e_1_3_2_164_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51701.2025.01527"},{"key":"e_1_3_2_165_2","volume-title":"A Letter from Our Chair and CEO","author":"Narayen Shantanu","year":"2024","unstructured":"Shantanu Narayen. 2024. A Letter from Our Chair and CEO. Retrieved from https:\/\/www.adobe.com\/cc-shared\/assets\/pdf\/corporate\/investor-relations\/adbe-2024-stockholder-letter.pdf. Accessed: 1 Jan 2025."},{"key":"e_1_3_2_166_2","volume-title":"The Twelfth International Conference on Learning Representations","author":"Ning Mang","year":"2024","unstructured":"Mang Ning, Mingxiao Li, Jianlin Su, Albert Ali Salah, and Itir Onal Ertugrul. 2024. Elucidating the exposure bias in diffusion models. In The Twelfth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=xEJMoj1SpX"},{"key":"e_1_3_2_167_2","volume-title":"The Thirty-Ninth Annual Conference on Neural Information Processing Systems","author":"Niu Mengjia","year":"2025","unstructured":"Mengjia Niu, Hamed Haddadi, and Guansong Pang. 2025. Robust hallucination detection in LLMs via adaptive token selection. In The Thirty-Ninth Annual Conference on Neural Information Processing Systems. Retrieved from https:\/\/openreview.net\/forum?id=gOwqPdBlRB"},{"key":"e_1_3_2_168_2","unstructured":"Valentin No\u00ebl. 2026. Geometry of Reason: Spectral Signatures of Valid Mathematical Reasoning. arxiv:2601.00791 [cs.LG]. https:\/\/arxiv.org\/abs\/2601.00791"},{"key":"e_1_3_2_169_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-2459"},{"key":"e_1_3_2_170_2","unstructured":"Trevine Oorloff Yaser Yacoob and Abhinav Shrivastava. 2025. Mitigating Hallucinations in Diffusion Models through Adaptive Attention Modulation. arxiv:2502.16872 [cs.CV]. https:\/\/arxiv.org\/abs\/2502.16872"},{"key":"e_1_3_2_171_2","volume-title":"The Twelfth International Conference on Learning Representations","author":"Oren Yonatan","year":"2024","unstructured":"Yonatan Oren, Nicole Meister, Niladri S. Chatterji, Faisal Ladhak, and Tatsunori Hashimoto. 2024. Proving test set contamination in black-box language models. In The Twelfth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=KS8mIvetg2"},{"key":"e_1_3_2_172_2","article-title":"Attention overlap is responsible for the entity missing problem in text-to-image diffusion models!","author":"Oriyad Arash Mari","year":"2025","unstructured":"Arash Mari Oriyad, Mohammadali Banayeeanzade, Reza Abbasi, Mohammad Hossein Rohban, and Mahdieh Soleymani Baghshah. 2025. Attention overlap is responsible for the entity missing problem in text-to-image diffusion models! Transactions on Machine Learning Research (2025). Retrieved from https:\/\/openreview.net\/forum?id=Xv3ZrFayIO","journal-title":"Transactions on Machine Learning Research"},{"key":"e_1_3_2_173_2","doi-asserted-by":"publisher","DOI":"10.3390\/app15031079"},{"key":"e_1_3_2_174_2","series-title":"ICML\u201924","volume-title":"Proceedings of the 41st International Conference on Machine Learning","author":"Papadopoulos Vassilis","year":"2024","unstructured":"Vassilis Papadopoulos, J\u00e9r\u00e9mie Wenger, and Cl\u00e9ment Hongler. 2024. Arrows of time for large language models. In Proceedings of the 41st International Conference on Machine Learning (Vienna, Austria) (ICML\u201924). JMLR.org, Article 1600, 20 pages."},{"key":"e_1_3_2_175_2","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Park Core Francisco","year":"2025","unstructured":"Core Francisco Park, Andrew Lee, Ekdeep Singh Lubana, Yongyi Yang, Maya Okawa, Kento Nishi, Martin Wattenberg, and Hidenori Tanaka. 2025. ICLR: In-context learning of representations. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=pXlmOmlHJZ"},{"key":"e_1_3_2_176_2","unstructured":"Dongmin Park Zhaofang Qian Guangxing Han and Ser-Nam Lim. 2024. Mitigating Dialogue Hallucination for Large Vision Language Models via Adversarial Instruction Tuning. arxiv:2403.10492 [cs.CV]. https:\/\/arxiv.org\/abs\/2403.10492"},{"key":"e_1_3_2_177_2","volume-title":"The Fourteenth International Conference on Learning Representations","author":"Park Jaden","year":"2026","unstructured":"Jaden Park, Mu Cai, Feng Yao, Jingbo Shang, Soochahn Lee, and Yong Jae Lee. 2026. Contamination detection for VLMs using multi-modal semantic perturbations. In The Fourteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=gk6OC3XIZW"},{"key":"e_1_3_2_178_2","volume-title":"The Thirty-Ninth Annual Conference on Neural Information Processing Systems","author":"Park Seongheon","year":"2025","unstructured":"Seongheon Park and Sharon Li. 2025. GLSim: Detecting object hallucinations in LVLMs via global-local similarity. In The Thirty-Ninth Annual Conference on Neural Information Processing Systems. Retrieved from https:\/\/openreview.net\/forum?id=ZO8LyCizx9"},{"key":"e_1_3_2_179_2","volume-title":"The Thirty-Ninth Annual Conference on Neural Information Processing Systems","author":"Peng Liang","year":"2025","unstructured":"Liang Peng, Boxi Wu, Haoran Cheng, Yibo Zhao, and Xiaofei He. 2025. Self-supervised direct preference optimization for text-to-image diffusion models. In The Thirty-Ninth Annual Conference on Neural Information Processing Systems. Retrieved from https:\/\/openreview.net\/forum?id=COifgrjzXR"},{"key":"e_1_3_2_180_2","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Pham Kha","year":"2025","unstructured":"Kha Pham, Hung Le, Man Ngo, and Truyen Tran. 2025. Rapid selection and ordering of in-context demonstrations via prompt embedding clustering. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=1Iu2Yte5N6"},{"key":"e_1_3_2_181_2","unstructured":"Weiguo Pian Shijian Deng Shentong Mo Mingrui Liu Yunhui Guo and Yapeng Tian. 2025. Modality-Inconsistent Continual Learning of Multimodal Large Language Models. Retrieved from https:\/\/openreview.net\/forum?id=l13qyPJyUF"},{"key":"e_1_3_2_182_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-3548"},{"key":"e_1_3_2_183_2","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Prashanth USVSN Sai","year":"2025","unstructured":"USVSN Sai Prashanth, Alvin Deng, Kyle O\u2019Brien, Jyothir S. V., Mohammad Aflah Khan, Jaydeep Borkar, Christopher A. Choquette-Choo, Jacob Ray Fuehne, Stella Biderman, Tracy Ke, et\u00a0al. 2025. Recite, reconstruct, recollect: Memorization in LMs as a multifaceted phenomenon. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=3E8YNv1HjU"},{"key":"e_1_3_2_184_2","unstructured":"Jingyang Qiao Zhizhong Zhang Xin Tan Yanyun Qu Shouhong Ding and Yuan Xie. 2025. Large Continual Instruction Assistant. arxiv:2410.10868 [cs.LG]. https:\/\/arxiv.org\/abs\/2410.10868"},{"key":"e_1_3_2_185_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.emnlp-main.473"},{"key":"e_1_3_2_186_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-4009"},{"key":"e_1_3_2_187_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW69036.2025.00661"},{"key":"e_1_3_2_188_2","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Rashidinejad Paria","year":"2025","unstructured":"Paria Rashidinejad and Yuandong Tian. 2025. Sail into the headwind: Alignment via robust rewards and dynamic labels against reward hacking. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=I8af9JdQTy"},{"key":"e_1_3_2_189_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.findings-emnlp.59"},{"key":"e_1_3_2_190_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-0027"},{"key":"e_1_3_2_191_2","doi-asserted-by":"publisher","unstructured":"Weijieying Ren Xinlong Li Lei Wang Tianxiang Zhao and Wei Qin. 2024. Analyzing and Reducing Catastrophic Forgetting in Parameter Efficient Tuning. DOI:10.48550\/ARXIV.2402.18865","DOI":"10.48550\/ARXIV.2402.18865"},{"key":"e_1_3_2_192_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-1220"},{"key":"e_1_3_2_193_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-3049"},{"key":"e_1_3_2_194_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.eacl-demo.17"},{"key":"e_1_3_2_195_2","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Ross Brendan Leigh","year":"2025","unstructured":"Brendan Leigh Ross, Hamidreza Kamkari, Tongzi Wu, Rasa Hosseinzadeh, Zhaoyan Liu, George Stein, Jesse C. Cresswell, and Gabriel Loaiza-Ganem. 2025. A geometric framework for understanding memorization in generative models. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=aZ1gNJu8wO"},{"key":"e_1_3_2_196_2","volume-title":"First Conference on Language Modeling","author":"Ross Candace","year":"2024","unstructured":"Candace Ross, Melissa Hall, Adriana Romero-Soriano, and Adina Williams. 2024. What makes a good metric? Evaluating automatic metrics for text-to-image consistency. In First Conference on Language Modeling. Retrieved from https:\/\/openreview.net\/forum?id=LFfktMPAci"},{"key":"e_1_3_2_197_2","unstructured":"Subhadeep Roy Gagan Bhatia and Steffen Eger. 2026. Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics. arxiv:2601.04946 [cs.CV]. https:\/\/arxiv.org\/abs\/2601.04946"},{"key":"e_1_3_2_198_2","unstructured":"William Rudman Michal Golovanevsky Dana Arad Yonatan Belinkov Ritambhara Singh Carsten Eickhoff and Kyle Mahowald. 2026. Mechanisms of Prompt-Induced Hallucination in Vision-Language Models. arxiv:2601.05201 [cs.CV]. https:\/\/arxiv.org\/abs\/2601.05201"},{"key":"e_1_3_2_199_2","doi-asserted-by":"publisher","unstructured":"Kuniaki Saito Kihyuk Sohn Chen-Yu Lee and Yoshitaka Ushiku. 2024. Where is the answer? Investigating Positional Bias in Language Model Knowledge Extraction. DOI:10.48550\/ARXIV.2402.12170","DOI":"10.48550\/ARXIV.2402.12170"},{"key":"e_1_3_2_200_2","unstructured":"Shreyas N. Samaga Gilberto Gonzalez Arroyo and Tamal K. Dey. 2026. HalluZig: Hallucination Detection using Zigzag Persistence. arxiv:2601.01552 [cs.CL]. https:\/\/arxiv.org\/abs\/2601.01552"},{"key":"e_1_3_2_201_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i5.28270"},{"key":"e_1_3_2_202_2","volume-title":"Forty-Second International Conference on Machine Learning","author":"Sanyal Sunny","year":"2025","unstructured":"Sunny Sanyal, Hayden Prairie, Rudrajit Das, Ali Kavis, and Sujay Sanghavi. 2025. Upweighting easy samples in fine-tuning mitigates forgetting. In Forty-Second International Conference on Machine Learning. Retrieved from https:\/\/openreview.net\/forum?id=13HPTmZKbM"},{"key":"e_1_3_2_203_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-2719"},{"key":"e_1_3_2_204_2","article-title":"Are emergent abilities of large language models a mirage?","author":"Schaeffer Rylan","year":"2024","unstructured":"Rylan Schaeffer, Brando Miranda, and Sanmi Koyejo. 2024. Are emergent abilities of large language models a mirage? Advances in Neural Information Processing Systems 36, Article 2425 (2024), 55565\u201355581.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_205_2","first-page":"9573","article-title":"The pitfalls of simplicity bias in neural networks","volume":"33","author":"Shah Harshay","year":"2020","unstructured":"Harshay Shah, Kaustav Tamuly, Aditi Raghunathan, Prateek Jain, and Praneeth Netrapalli. 2020. The pitfalls of simplicity bias in neural networks. Advances in Neural Information Processing Systems 33, Article 803 (2020), 9573\u20139585.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_206_2","volume-title":"The Thirty-Ninth Annual Conference on Neural Information Processing Systems","author":"Shechter Mikey","year":"2025","unstructured":"Mikey Shechter and Yair Carmon. 2025. Filter like you test: Data-driven data filtering for CLIP pretraining. In The Thirty-Ninth Annual Conference on Neural Information Processing Systems. Retrieved from https:\/\/openreview.net\/forum?id=lYOOHqfM46"},{"key":"e_1_3_2_207_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-2210"},{"key":"e_1_3_2_208_2","volume-title":"The Twelfth International Conference on Learning Representations","author":"Shukor Mustafa","year":"2024","unstructured":"Mustafa Shukor, Alexandre Rame, Corentin Dancette, and Matthieu Cord. 2024. Beyond task performance: Evaluating and reducing the flaws of large multimodal models with in-context-learning. In The Twelfth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=mMaQvkMzDi"},{"key":"e_1_3_2_209_2","volume-title":"Advances in Neural Information Processing Systems","author":"Skalse Joar Max Viktor","year":"2022","unstructured":"Joar Max Viktor Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, and David Krueger. 2022. Defining and characterizing reward gaming. In Advances in Neural Information Processing Systems, Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho (Eds.). Retrieved from https:\/\/openreview.net\/forum?id=yb3HOXO3lX2"},{"key":"e_1_3_2_210_2","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Slocum Stewart","year":"2025","unstructured":"Stewart Slocum, Asher Parker-Sartori, and Dylan Hadfield-Menell. 2025. Diverse preference learning for capabilities and alignment. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=pOq9vDIYev"},{"key":"e_1_3_2_211_2","doi-asserted-by":"crossref","unstructured":"Eric Slyman Mehrab Tanjim Kushal Kafle and Stefan Lee. 2025. Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles. arxiv:2509.08777 [cs.CV]. https:\/\/arxiv.org\/abs\/2509.08777","DOI":"10.1109\/ICCV51701.2025.01600"},{"key":"e_1_3_2_212_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00586"},{"key":"e_1_3_2_213_2","doi-asserted-by":"publisher","DOI":"10.52202\/075280-2071"},{"key":"e_1_3_2_214_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2025.findings-emnlp.556"},{"key":"e_1_3_2_215_2","volume-title":"OpenAI Hits More Than 1 Million Paid Business Users","author":"Sophia Deborah","year":"2024","unstructured":"Deborah Sophia, Zaheer Kachwala, and Harshita Mary Varghese. 2024. OpenAI Hits More Than 1 Million Paid Business Users. Retrieved from https:\/\/www.reuters.com\/technology\/artificial-intelligence\/openai-considers-pricier-subscriptions-its-chatbot-ai-information-reports-2024-09-05\/. Accessed: 1-1-25."},{"key":"e_1_3_2_216_2","volume-title":"ICML","author":"Sorensen Taylor","year":"2024","unstructured":"Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, et\u00a0al. 2024. Position: A roadmap to pluralistic alignment. In ICML (Vienna, Austria). JMLR.org, Article 1882, 23 pages."},{"key":"e_1_3_2_217_2","unstructured":"Vighnesh Subramaniam David Mayo Colin Conwell Tomaso Poggio Boris Katz Brian Cheung and Andrei Barbu. 2025. Training the Untrainable: Introducing Inductive Bias via Representational Alignment. arxiv:2410.20035 [cs.LG]. https:\/\/arxiv.org\/abs\/2410.20035"},{"key":"e_1_3_2_218_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.62"},{"key":"e_1_3_2_219_2","unstructured":"Wangtao Sun Xiang Cheng Xing Yu Haotian Xu Zhao Yang Shizhu He Jun Zhao and Kang Liu. 2025. Probabilistic Uncertain Reward Model. arxiv:2503.22480 [cs.LG]. https:\/\/arxiv.org\/abs\/2503.22480"},{"key":"e_1_3_2_220_2","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Sun ZhongXiang","year":"2025","unstructured":"ZhongXiang Sun, Xiaoxue Zang, Kai Zheng, Jun Xu, Xiao Zhang, Weijie Yu, Yang Song, and Han Li. 2025. ReDeEP: Detecting hallucination in retrieval-augmented generation via mechanistic interpretability. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=ztzZDzgfrh"},{"key":"e_1_3_2_221_2","unstructured":"Chaodong Tong Qi Zhang Chen Li Lei Jiang and Yanbing Liu. 2026. FaithSCAN: Model-Driven Single-Pass Hallucination Detection for Faithful Visual Question Answering. arxiv:2601.00269. https:\/\/arxiv.org\/abs\/2601.00269"},{"key":"e_1_3_2_222_2","unstructured":"Kostas Triaridis Alexandros Graikos Aggelina Chatziagapi Grigorios G. Chrysos and Dimitris Samaras. 2025. Mitigating Diffusion Model Hallucinations with Dynamic Guidance. arxiv:2510.05356 [cs.CV]. https:\/\/arxiv.org\/abs\/2510.05356"},{"key":"e_1_3_2_223_2","series-title":"Proceedings of Machine Learning Research","first-page":"48728","volume-title":"Proceedings of the 41st International Conference on Machine Learning","volume":"235","author":"Tsoy Nikita","year":"2024","unstructured":"Nikita Tsoy and Nikola Konstantinov. 2024. Simplicity bias of two-layer networks beyond linearly separable data. In Proceedings of the 41st International Conference on Machine Learning(Proceedings of Machine Learning Research, Vol. 235), Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp (Eds.). PMLR, 48728\u201348767. Retrieved from https:\/\/proceedings.mlr.press\/v235\/tsoy24a.html"},{"key":"e_1_3_2_224_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.01847"},{"key":"e_1_3_2_225_2","volume-title":"The Thirty-Eighth Annual Conference on Neural Information Processing Systems","author":"Udandarao Vishaal","year":"2024","unstructured":"Vishaal Udandarao, Ameya Prabhu, Adhiraj Ghosh, Yash Sharma, Philip Torr, Adel Bibi, Samuel Albanie, and Matthias Bethge. 2024. No \u201czero-shot\u201d without exponential data: Pretraining concept frequency determines multimodal model performance. In The Thirty-Eighth Annual Conference on Neural Information Processing Systems. Retrieved from https:\/\/openreview.net\/forum?id=9VbGjXLzig"},{"key":"e_1_3_2_226_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2025.3602733"},{"key":"e_1_3_2_227_2","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Vastola John","year":"2025","unstructured":"John Vastola. 2025. Generalization through variance: How noise shapes inductive biases in diffusion models. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=7lUdo8Vuqa"},{"key":"e_1_3_2_228_2","article-title":"Attention is all you need","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in Neural Information Processing Systems 30 (2017), 6000\u20136010.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_229_2","unstructured":"Chaoyang Wang Yangfan He Yiyang Zhou Yixuan Wang Jiaqi Liu Peng Xia Zhengzhong Tu Mohit Bansal and Huaxiu Yao. 2025. Knowing the Answer Isn\u2019t Enough: Fixing Reasoning Path Failures in LVLMs. arxiv:2512.06258 [cs.CV]. https:\/\/arxiv.org\/abs\/2512.06258"},{"key":"e_1_3_2_230_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2025.naacl-long.399"},{"key":"e_1_3_2_231_2","doi-asserted-by":"publisher","unstructured":"Jiayin Wang Weizhi Ma Peijie Sun Min Zhang and Jian-Yun Nie. 2024. Understanding User Experience in Large Language Model Interactions. DOI:10.48550\/ARXIV.2401.08329","DOI":"10.48550\/ARXIV.2401.08329"},{"key":"e_1_3_2_232_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.acl-long.511"},{"key":"e_1_3_2_233_2","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Wang Qixun","year":"2025","unstructured":"Qixun Wang, Yifei Wang, Xianghua Ying, and Yisen Wang. 2025. Can in-context learning really generalize to out-of-distribution tasks?. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=INe4otjryz"},{"key":"e_1_3_2_234_2","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Wang Wenhao","year":"2025","unstructured":"Wenhao Wang, Adam Dziedzic, Grace C. Kim, Michael Backes, and Franziska Boenisch. 2025. Captured by captions: On memorization and its mitigation in CLIP models. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=5V0f8igznO"},{"key":"e_1_3_2_235_2","unstructured":"Xiao Wang Ibrahim Alabdulmohsin Daniel Salz Zhe Li Keran Rong and Xiaohua Zhai. 2025. Scaling Pre-training to One Hundred Billion Data for Vision Language Models. arxiv:2502.07617 [cs.CV]. https:\/\/arxiv.org\/abs\/2502.07617"},{"key":"e_1_3_2_236_2","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Wang Xinyi","year":"2025","unstructured":"Xinyi Wang, Antonis Antoniades, Yanai Elazar, Alfonso Amayuelas, Alon Albalak, Kexun Zhang, and William Yang Wang. 2025. Generalization v.s. memorization: Tracing language models\u2019 capabilities back to pretraining data. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=IQxBDLmVpT"},{"key":"e_1_3_2_237_2","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Wang Xiaolei","year":"2025","unstructured":"Xiaolei Wang, Xinyu Tang, Junyi Li, Xin Zhao, and Ji-Rong Wen. 2025. Investigating the pre-training dynamics of in-context learning: Task recognition vs. task learning. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=htDczodFN5"},{"key":"e_1_3_2_238_2","series-title":"SEC\u201925","volume-title":"Proceedings of the 34th USENIX Conference on Security Symposium","author":"Wang Yining","year":"2025","unstructured":"Yining Wang, Mi Zhang, Junjie Sun, Chenyue Wang, Min Yang, Hui Xue, Jialing Tao, Ranjie Duan, and Jiexi Liu. 2025. Mirage in the eyes: Hallucination attack on multi-modal large language models with only attention sink. In Proceedings of the 34th USENIX Conference on Security Symposium (Seattle, WA, USA) (SEC\u201925). USENIX Association, USA, Article 191, 20 pages."},{"key":"e_1_3_2_239_2","unstructured":"Ziqi Wang Chang Che Qi Wang Yangyang Li Zenglin Shi and Meng Wang. 2025. SMoLoRA: Exploring and Defying Dual Catastrophic Forgetting in Continual Visual Instruction Tuning. arxiv:2411.13949 [cs.CV]. https:\/\/arxiv.org\/abs\/2411.13949"},{"key":"e_1_3_2_240_2","volume-title":"The Fourteenth International Conference on Learning Representations","author":"Wang Zekun","year":"2026","unstructured":"Zekun Wang, Anant Gupta, Zihan Dong, and Christopher J. MacLellan. 2026. Avoid catastrophic forgetting with rank-1 fisher from diffusion models. In The Fourteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=zCZcbRsc4g"},{"key":"e_1_3_2_241_2","unstructured":"Xiwen Wei Mustafa Munir and Radu Marculescu. 2025. Mitigating Intra- and Inter-modal Forgetting in Continual Learning of Unified Multimodal Models. arxiv:2512.03125 [cs.LG]. https:\/\/arxiv.org\/abs\/2512.03125"},{"key":"e_1_3_2_242_2","unstructured":"Penghao Wu and Saining Xie. 2023. V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs. arxiv:2312.14135 [cs.CV]. https:\/\/arxiv.org\/abs\/2312.14135"},{"key":"e_1_3_2_243_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.naacl-long.102"},{"key":"e_1_3_2_244_2","volume-title":"The Twelfth International Conference on Learning Representations","author":"Xiao Guangxuan","year":"2024","unstructured":"Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. 2024. Efficient streaming language models with attention sinks. In The Twelfth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=NG7sS51zVF"},{"key":"e_1_3_2_245_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-2920"},{"key":"e_1_3_2_246_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-2773"},{"key":"e_1_3_2_247_2","doi-asserted-by":"publisher","unstructured":"Ruijie Xu Zengzhi Wang Run-Ze Fan and Pengfei Liu. 2024. Benchmarking Benchmark Leakage in Large Language Models. DOI:10.48550\/ARXIV.2404.18824","DOI":"10.48550\/ARXIV.2404.18824"},{"key":"e_1_3_2_248_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.findings-acl.905"},{"key":"e_1_3_2_249_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.01435"},{"key":"e_1_3_2_250_2","volume-title":"The Twelfth International Conference on Learning Representations","author":"Yan Jianhao","year":"2024","unstructured":"Jianhao Yan, Jin Xu, Chiyu Song, Chenming Wu, Yafu Li, and Yue Zhang. 2024. Understanding in-context learning from repetitions. In The Twelfth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=bGGYcvw8mp"},{"key":"e_1_3_2_251_2","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Yang Suorong","year":"2025","unstructured":"Suorong Yang, Peng Ye, Wanli Ouyang, Dongzhan Zhou, and Furao Shen. 2025. A CLIP-powered framework for robust and generalizable data selection. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=9bMZ29SPVx"},{"key":"e_1_3_2_252_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-2727"},{"key":"e_1_3_2_253_2","unstructured":"Yunxiang Yang Ningning Xu and Jidong J. Yang. 2025. Multi-Agent Visual-Language Reasoning for Comprehensive Highway Scene Understanding. arxiv:2508.17205 [cs.CV]. https:\/\/arxiv.org\/abs\/2508.17205"},{"key":"e_1_3_2_254_2","volume-title":"The Twelfth International Conference on Learning Representations","author":"Yang Zhantao","year":"2024","unstructured":"Zhantao Yang, Ruili Feng, Han Zhang, Yujun Shen, Kai Zhu, Lianghua Huang, Yifei Zhang, Yu Liu, Deli Zhao, Jingren Zhou, and Fan Cheng. 2024. Lipschitz singularities in diffusion models. In The Twelfth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=WNkW0cOwiz"},{"key":"e_1_3_2_255_2","unstructured":"Binwei Yao Zefan Cai Yun-Shiuan Chuang Shanglin Yang Ming Jiang Diyi Yang and Junjie Hu. 2025. No Preference Left Behind: Group Distributional Preference Optimization. arxiv:2412.20299 [cs.CL]. https:\/\/arxiv.org\/abs\/2412.20299"},{"key":"e_1_3_2_256_2","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Yao Yuzhe","year":"2025","unstructured":"Yuzhe Yao, Jun Chen, Zeyi Huang, Haonan Lin, Mengmeng Wang, Guang Dai, and Jingdong Wang. 2025. Manifold constraint reduces exposure bias in accelerated diffusion sampling. In The Thirteenth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=5xmXUwDxep"},{"key":"e_1_3_2_257_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v39i21.34363"},{"key":"e_1_3_2_258_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACVW65960.2025.00120"},{"key":"e_1_3_2_259_2","volume-title":"The Thirty-Ninth Annual Conference on Neural Information Processing Systems","author":"Yu Annan","year":"2025","unstructured":"Annan Yu and N. Benjamin Erichson. 2025. Block-biased mamba for long-range sequence processing. In The Thirty-Ninth Annual Conference on Neural Information Processing Systems. Retrieved from https:\/\/openreview.net\/forum?id=5WKEH9LhAQ"},{"key":"e_1_3_2_260_2","volume-title":"The Eleventh International Conference on Learning Representations","author":"Yuksekgonul Mert","year":"2023","unstructured":"Mert Yuksekgonul, Federico Bianchi, Pratyusha Kalluri, Dan Jurafsky, and James Zou. 2023. When and why vision-language models behave like bags-of-words, and what to do about it?. In The Eleventh International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=KRLUvxh8uaX"},{"key":"e_1_3_2_261_2","unstructured":"Shuangfei Zhai Tatiana Likhomanenko Etai Littwin Dan Busbridge Jason Ramapuram Yizhe Zhang Jiatao Gu and Josh Susskind. 2023. Stabilizing Transformer Training by Preventing Attention Entropy Collapse. arxiv:2303.06296 [cs.LG]. https:\/\/arxiv.org\/abs\/2303.06296"},{"key":"e_1_3_2_262_2","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Zhang Chi","year":"2025","unstructured":"Chi Zhang, Huaping Zhong, Kuan Zhang, Chengliang Chai, Rui Wang, Xinlin Zhuang, Tianyi Bai, Qiu Jiantao, Lei Cao, Ju Fan, et\u00a0al. 2025. Harnessing diversity for important data selection in pretraining large language models. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=bMC1t7eLRc"},{"key":"e_1_3_2_263_2","unstructured":"Feiran Zhang Yixin Wu Zhenghua Xiaohua Changze Xuanjing Huang and Xiaoqing Zheng. 2026. VIB-Probe: Detecting and Mitigating Hallucinations in Vision-Language Models via Variational Information Bottleneck. arxiv:2601.05547 [cs.CV]. https:\/\/arxiv.org\/abs\/2601.05547"},{"key":"e_1_3_2_264_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51701.2025.01415"},{"key":"e_1_3_2_265_2","doi-asserted-by":"crossref","unstructured":"Haozhuo Zhang Bin Zhu Yu Cao and Yanbin Hao. 2024. Hand1000: Generating Realistic Hands from Text with Only 1 000 Images. arxiv:2408.15461 [cs.CV]. https:\/\/arxiv.org\/abs\/2408.15461","DOI":"10.1609\/aaai.v39i9.33074"},{"key":"e_1_3_2_266_2","volume-title":"The Thirty-Ninth Annual Conference on Neural Information Processing Systems","author":"Zhang Jiayu","year":"2025","unstructured":"Jiayu Zhang, Changbang Li, Yinan Peng, Weihao Luo, Peilai Yu, and Xuan Zhang. 2025. Whose instructions count? Resolving preference bias in instruction fine-tuning. In The Thirty-Ninth Annual Conference on Neural Information Processing Systems. Retrieved from https:\/\/openreview.net\/forum?id=UFAKqq77e3"},{"key":"e_1_3_2_267_2","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Zhang Junyu","year":"2025","unstructured":"Junyu Zhang, Daochang Liu, Eunbyung Park, Shichao Zhang, and Chang Xu. 2025. Anti-exposure bias in diffusion models. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=MtDd7rWok1"},{"key":"e_1_3_2_268_2","unstructured":"Kewei Zhang Ye Huang Yufan Deng Jincheng Yu Junsong Chen Huan Ling Enze Xie and Daquan Zhou. 2026. MHLA: Restoring Expressivity of Linear Attention via Token-Level Multi-Head. arxiv:2601.07832 [cs.CV]. https:\/\/arxiv.org\/abs\/2601.07832"},{"key":"e_1_3_2_269_2","volume-title":"The Thirty-Ninth Annual Conference on Neural Information Processing Systems","author":"Zhang Shen","year":"2025","unstructured":"Shen Zhang, Siyuan Liang, Yaning Tan, Zhaowei Chen, Linze Li, Ge Wu, Yuhao Chen, Shuheng Li, Zhenyu Zhao, Caihua Chen, Jiajun Liang, and Yao Tang. 2025. LEDiT: Your length-extrapolatable diffusion transformer without positional encoding. In The Thirty-Ninth Annual Conference on Neural Information Processing Systems. Retrieved from https:\/\/openreview.net\/forum?id=8x5OmcFtJV"},{"key":"e_1_3_2_270_2","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Zhang Xingxuan","year":"2025","unstructured":"Xingxuan Zhang, Haoran Wang, Jiansheng Li, Yuan Xue, Shikai Guan, Renzhe Xu, Hao Zou, Han Yu, and Peng Cui. 2025. Understanding the generalization of in-context learning in transformers: An empirical study. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=yOhNLIqTEF"},{"key":"e_1_3_2_271_2","unstructured":"Yichi Zhang Jinlong Pang Zhaowei Zhu and Yang Liu. 2025. Evaluating LLM-Contaminated Crowdsourcing Data Without Ground Truth. arxiv:2506.06991 [cs.AI]. https:\/\/arxiv.org\/abs\/2506.06991"},{"key":"e_1_3_2_272_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2025.findings-acl.592"},{"key":"e_1_3_2_273_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-0451"},{"key":"e_1_3_2_274_2","volume-title":"The Twelfth International Conference on Learning Representations","author":"Zhao Bo","year":"2024","unstructured":"Bo Zhao, Robert M. Gower, Robin Walters, and Rose Yu. 2024. Improving convergence and generalization using parameter symmetries. In The Twelfth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=L0r0GphlIL"},{"key":"e_1_3_2_275_2","unstructured":"Hengyuan Zhao Ziqin Wang Qixin Sun Kaiyou Song Yilin Li Xiaolin Hu Qingpei Guo and Si Liu. 2025. LLaVA-CMoE: Towards Continual Mixture of Experts for Large Vision-Language Models. arxiv:2503.21227 [cs.CL]. https:\/\/arxiv.org\/abs\/2503.21227"},{"key":"e_1_3_2_276_2","unstructured":"Hongbo Zhao Fei Zhu Haiyang Guo Meng Wang Rundong Wang Gaofeng Meng and Zhaoxiang Zhang. 2025. MLLM-CL: Continual Learning for Multimodal Large Language Models. https:\/\/arxiv.org\/abs\/2506.05453"},{"key":"e_1_3_2_277_2","unstructured":"Chuanyang Zheng. 2025. The Linear Attention Resurrection in Vision Transformer. arxiv:2501.16182 [cs.CV]. https:\/\/arxiv.org\/abs\/2501.16182"},{"key":"e_1_3_2_278_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-2405"},{"key":"e_1_3_2_279_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2025.3540889"},{"key":"e_1_3_2_280_2","unstructured":"Feng Zhou Pu Cao Yiyang Ma Lu Yang and Jianqin Yin. 2025. Exploring Position Encoding in Diffusion U-Net for Training-free High-resolution Image Generation. arxiv:2503.09830 [cs.CV]. https:\/\/arxiv.org\/abs\/2503.09830"},{"key":"e_1_3_2_281_2","doi-asserted-by":"publisher","unstructured":"Xin Zhou Martin Weyssow Ratnadira Widyasari Ting Zhang Junda He Yunbo Lyu Jianming Chang Beiqi Zhang Dan Huang and David Lo. 2025. LessLeak-Bench: A First Investigation of Data Leakage in LLMs Across 83 Software Engineering Benchmarks. DOI:10.48550\/ARXIV.2502.06215","DOI":"10.48550\/ARXIV.2502.06215"},{"key":"e_1_3_2_282_2","volume-title":"The Thirty-Eighth Annual Conference on Neural Information Processing Systems","author":"Zhu Hanlin","year":"2024","unstructured":"Hanlin Zhu, Baihe Huang, Shaolun Zhang, Michael Jordan, Jiantao Jiao, Yuandong Tian, and Stuart Russell. 2024. Towards a theoretical understanding of the \u2018reversal curse\u2019 via training dynamics. In The Thirty-Eighth Annual Conference on Neural Information Processing Systems. Retrieved from https:\/\/openreview.net\/forum?id=QoWf3lo6m7"},{"key":"e_1_3_2_283_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-1820"},{"key":"e_1_3_2_284_2","unstructured":"Shaobin Zhuang Yiwei Guo Yanbo Ding Kunchang Li Xinyuan Chen Yaohui Wang Fangyikang Wang Ying Zhang Chen Li and Yali Wang. 2025. TimeStep Master: Asymmetrical Mixture of Timestep LoRA Experts for Versatile and Efficient Diffusion Models in Vision. arxiv:2503.07416 [cs.CV]. https:\/\/arxiv.org\/abs\/2503.07416"}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3811409","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T13:09:42Z","timestamp":1782306582000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3811409"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,24]]},"references-count":283,"journal-issue":{"issue":"13","published-print":{"date-parts":[[2026,10,31]]}},"alternative-id":["10.1145\/3811409"],"URL":"https:\/\/doi.org\/10.1145\/3811409","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,24]]},"assertion":[{"value":"2025-06-03","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-04-13","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-24","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}