{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,31]],"date-time":"2026-07-31T08:57:47Z","timestamp":1785488267119,"version":"3.56.0"},"publisher-location":"New York, NY, USA","reference-count":31,"publisher":"ACM","license":[{"start":{"date-parts":[[2025,12,17]],"date-time":"2025-12-17T00:00:00Z","timestamp":1765929600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2025,12,17]]},"DOI":"10.1145\/3774521.3774542","type":"proceedings-article","created":{"date-parts":[[2026,7,31]],"date-time":"2026-07-31T07:34:24Z","timestamp":1785483264000},"page":"1-9","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["HalluShift++: Bridging Language and Vision through Internal Representation Shifts for Hierarchical Hallucinations in MLLMs"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0003-9974-8018","authenticated-orcid":false,"given":"Sujoy","family":"Nath","sequence":"first","affiliation":[{"name":"Netaji Subhash Engineering College, Kolkata, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0136-0887","authenticated-orcid":false,"given":"Arkaprabha","family":"Basu","sequence":"additional","affiliation":[{"name":"TCG Crest, Kolkata, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-6385-243X","authenticated-orcid":false,"given":"Sharanya","family":"Dasgupta","sequence":"additional","affiliation":[{"name":"Indian Statistical Institute, Kolkata, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6843-4508","authenticated-orcid":false,"given":"Swagatam","family":"Das","sequence":"additional","affiliation":[{"name":"Electronics and Communication Sciences Unit, Indian Statistical Institute, Kolkata, India"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,7,31]]},"reference":[{"key":"e_1_3_3_1_2_2","unstructured":"Marah Abdin Jyoti Aneja Hany Awadalla and Ahmed\u00a0Awadallah et al.2024. Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone. arxiv:https:\/\/arXiv.org\/abs\/2404.14219\u00a0[cs.CL] https:\/\/arxiv.org\/abs\/2404.14219"},{"key":"e_1_3_3_1_3_2","unstructured":"Josh Achiam Steven Adler Sandhini Agarwal Lama Ahmad Ilge Akkaya Florencia\u00a0Leoni Aleman Diogo Almeida Janko Altenschmidt Sam Altman Shyamal Anadkat et\u00a0al. 2023. Gpt-4 technical report. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2303.08774 (2023)."},{"key":"e_1_3_3_1_4_2","unstructured":"Jinze Bai Shuai Bai Yunfei Chu and Zeyu Cui. 2023. Qwen Technical Report. arxiv:https:\/\/arXiv.org\/abs\/2309.16609\u00a0[cs.CL] https:\/\/arxiv.org\/abs\/2309.16609"},{"key":"e_1_3_3_1_5_2","unstructured":"Zechen Bai Pichao Wang Tianjun Xiao Tong He Zongbo Han Zheng Zhang and Mike\u00a0Zheng Shou. 2024. Hallucination of multimodal large language models: A survey. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2404.18930 (2024)."},{"key":"e_1_3_3_1_6_2","unstructured":"Lucas Beyer Andreas Steiner Andr\u00e9\u00a0Susano Pinto Alexander Kolesnikov Xiao Wang Daniel Salz Maxim Neumann Ibrahim Alabdulmohsin Michael Tschannen Emanuele Bugliarello Thomas Unterthiner Daniel Keysers Skanda Koppula Fangyu Liu Adam Grycner Alexey Gritsenko Neil Houlsby Manoj Kumar Keran Rong Julian Eisenschlos Rishabh Kabra Matthias Bauer Matko Bo\u0161njak Xi Chen Matthias Minderer Paul Voigtlaender Ioana Bica Ivana Balazevic Joan Puigcerver Pinelopi Papalampidi Olivier Henaff Xi Xiong Radu Soricut Jeremiah Harmsen and Xiaohua Zhai. 2024. PaliGemma: A versatile 3B VLM for transfer. arxiv:https:\/\/arXiv.org\/abs\/2407.07726\u00a0[cs.CV] https:\/\/arxiv.org\/abs\/2407.07726"},{"key":"e_1_3_3_1_7_2","unstructured":"Florian Bordes Richard\u00a0Yuanzhe Pang and Anurag Ajay. 2024. An Introduction to Vision-Language Modeling. arxiv:https:\/\/arXiv.org\/abs\/2405.17247\u00a0[cs.LG] https:\/\/arxiv.org\/abs\/2405.17247"},{"key":"e_1_3_3_1_8_2","unstructured":"Collin Burns Haotian Ye Dan Klein and Jacob Steinhardt. 2022. Discovering latent knowledge in language models without supervision. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2212.03827 (2022)."},{"key":"e_1_3_3_1_9_2","doi-asserted-by":"crossref","unstructured":"Nitesh\u00a0V Chawla Kevin\u00a0W Bowyer Lawrence\u00a0O Hall and W\u00a0Philip Kegelmeyer. 2002. SMOTE: synthetic minority over-sampling technique. Journal of artificial intelligence research 16 (2002) 321\u2013357.","DOI":"10.1613\/jair.953"},{"key":"e_1_3_3_1_10_2","unstructured":"Chao Chen Kai Liu Ze Chen Yi Gu Yue Wu Mingyuan Tao Zhihang Fu and Jieping Ye. 2024. INSIDE: LLMs\u2019 internal states retain the power of hallucination detection. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2402.03744 (2024)."},{"key":"e_1_3_3_1_11_2","unstructured":"Xiang Chen Chenxi Wang Yida Xue Ningyu Zhang Xiaoyan Yang Qiang Li Yue Shen Lei Liang Jinjie Gu and Huajun Chen. 2024. Unified hallucination detection for multimodal large language models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2402.03190 (2024)."},{"key":"e_1_3_3_1_12_2","unstructured":"Wenliang Dai Junnan Li Dongxu Li Anthony Meng\u00a0Huat Tiong Junqi Zhao Weisheng Wang Boyang Li Pascale Fung and Steven Hoi. 2023. InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning. arxiv:https:\/\/arXiv.org\/abs\/2305.06500\u00a0[cs.CV] https:\/\/arxiv.org\/abs\/2305.06500"},{"key":"e_1_3_3_1_13_2","doi-asserted-by":"crossref","unstructured":"Sharanya Dasgupta Sujoy Nath Arkaprabha Basu Pourya Shamsolmoali and Swagatam Das. 2025. Hallushift: Measuring distribution shifts towards hallucination detection in llms. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2504.09482 (2025).","DOI":"10.1109\/IJCNN64981.2025.11228484"},{"key":"e_1_3_3_1_14_2","doi-asserted-by":"crossref","unstructured":"Xuefeng Du Chaowei Xiao and Sharon Li. 2024. Haloscope: Harnessing unlabeled llm generations for hallucination detection. Advances in Neural Information Processing Systems 37 (2024) 102948\u2013102972.","DOI":"10.52202\/079017-3270"},{"key":"e_1_3_3_1_15_2","doi-asserted-by":"crossref","unstructured":"Ziwei Ji Nayeon Lee Rita Frieske Tiezheng Yu Dan Su Yan Xu Etsuko Ishii Ye\u00a0Jin Bang Andrea Madotto and Pascale Fung. 2023. Survey of hallucination in natural language generation. ACM computing surveys 55 12 (2023) 1\u201338.","DOI":"10.1145\/3571730"},{"key":"e_1_3_3_1_16_2","unstructured":"Liqiang Jing Ruosen Li Yunmo Chen and Xinya Du. 2023. FaithScore: Fine-grained Evaluations of Hallucinations in Large Vision-Language Models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2311.01477 (2023)."},{"key":"e_1_3_3_1_17_2","unstructured":"Imed Keraghel Stanislas Morbieu and Mohamed Nadif. 2024. Recent Advances in Named Entity Recognition: A Comprehensive Survey and Comparative Study. arxiv:https:\/\/arXiv.org\/abs\/2401.10825\u00a0[cs.CL] https:\/\/arxiv.org\/abs\/2401.10825"},{"key":"e_1_3_3_1_18_2","first-page":"19730","volume-title":"International conference on machine learning","author":"Li Junnan","year":"2023","unstructured":"Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning. PMLR, 19730\u201319742."},{"key":"e_1_3_3_1_19_2","doi-asserted-by":"crossref","unstructured":"Kenneth Li Oam Patel Fernanda Vi\u00e9gas Hanspeter Pfister and Martin Wattenberg. 2023. Inference-time intervention: Eliciting truthful answers from a language model. Advances in Neural Information Processing Systems 36 (2023) 41451\u201341530.","DOI":"10.52202\/075280-1797"},{"key":"e_1_3_3_1_20_2","unstructured":"Yifan Li Yifan Du Kun Zhou Jinpeng Wang Wayne\u00a0Xin Zhao and Ji-Rong Wen. 2023. Evaluating object hallucination in large vision-language models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2305.10355 (2023)."},{"key":"e_1_3_3_1_21_2","unstructured":"Tsung-Yi Lin Michael Maire Serge Belongie Lubomir Bourdev Ross Girshick James Hays Pietro Perona Deva Ramanan C.\u00a0Lawrence Zitnick and Piotr Doll\u00e1r. 2015. Microsoft COCO: Common Objects in Context. arxiv:https:\/\/arXiv.org\/abs\/1405.0312\u00a0[cs.CV] https:\/\/arxiv.org\/abs\/1405.0312"},{"key":"e_1_3_3_1_22_2","unstructured":"Fuxiao Liu Kevin Lin Linjie Li Jianfeng Wang Yaser Yacoob and Lijuan Wang. 2023. Mitigating hallucination in large multi-modal models via robust instruction tuning. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2306.14565 (2023)."},{"key":"e_1_3_3_1_23_2","unstructured":"Haotian Liu Chunyuan Li Qingyang Wu and Yong\u00a0Jae Lee. 2023. Visual Instruction Tuning. arxiv:https:\/\/arXiv.org\/abs\/2304.08485\u00a0[cs.CV] https:\/\/arxiv.org\/abs\/2304.08485"},{"key":"e_1_3_3_1_24_2","unstructured":"Hanchao Liu Wenyuan Xue Yifei Chen Dapeng Chen Xiutian Zhao Ke Wang Liping Hou Rongjun Li and Wei Peng. 2024. A survey on hallucination in large vision-language models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2402.00253 (2024)."},{"key":"e_1_3_3_1_25_2","unstructured":"Xichen Pan Li Dong Shaohan Huang Zhiliang Peng Wenhu Chen and Furu Wei. 2024. Kosmos-G: Generating Images in Context with Multimodal Large Language Models. arxiv:https:\/\/arXiv.org\/abs\/2310.02992\u00a0[cs.CV] https:\/\/arxiv.org\/abs\/2310.02992"},{"key":"e_1_3_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.02782"},{"key":"e_1_3_3_1_27_2","doi-asserted-by":"crossref","unstructured":"Anna Rohrbach Lisa\u00a0Anne Hendricks Kaylee Burns Trevor Darrell and Kate Saenko. 2018. Object hallucination in image captioning. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/1809.02156 (2018).","DOI":"10.18653\/v1\/D18-1437"},{"key":"e_1_3_3_1_28_2","unstructured":"Anna Rohrbach Atousa Torabi Marcus Rohrbach Niket Tandon Christopher Pal Hugo Larochelle Aaron Courville and Bernt Schiele. 2016. Movie Description. arxiv:https:\/\/arXiv.org\/abs\/1605.03705\u00a0[cs.CV] https:\/\/arxiv.org\/abs\/1605.03705"},{"key":"e_1_3_3_1_29_2","doi-asserted-by":"crossref","unstructured":"Gaurang Sriramanan Siddhant Bharti Vinu\u00a0Sankar Sadasivan Shoumik Saha Priyatham Kattakinda and Soheil Feizi. 2024. Llm-check: Investigating detection of hallucinations in large language models. Advances in Neural Information Processing Systems 37 (2024) 34188\u201334216.","DOI":"10.52202\/079017-1077"},{"key":"e_1_3_3_1_30_2","unstructured":"Hugo Touvron Thibaut Lavril Gautier Izacard Xavier Martinet Marie-Anne Lachaux Timoth\u00e9e Lacroix Baptiste Rozi\u00e8re Naman Goyal Eric Hambro Faisal Azhar Aurelien Rodriguez Armand Joulin Edouard Grave and Guillaume Lample. 2023. LLaMA: Open and Efficient Foundation Language Models. arxiv:https:\/\/arXiv.org\/abs\/2302.13971\u00a0[cs.CL] https:\/\/arxiv.org\/abs\/2302.13971"},{"key":"e_1_3_3_1_31_2","unstructured":"Junyang Wang Yiyang Zhou Guohai Xu Pengcheng Shi Chenlin Zhao Haiyang Xu Qinghao Ye Ming Yan Ji Zhang Jihua Zhu et\u00a0al. 2023. Evaluation and analysis of hallucination in large vision-language models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2308.15126 (2023)."},{"key":"e_1_3_3_1_32_2","unstructured":"Mingrui Wu Jiayi Ji Oucheng Huang Jiale Li Yuhang Wu Xiaoshuai Sun and Rongrong Ji. 2024. Evaluating and analyzing relationship hallucinations in large vision-language models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2406.16449 (2024)."}],"event":{"name":"ICVGIP 2025: Indian Conference on Computer Vision, Graphics, and Image Processing","location":"Mandi Himachal Pradesh India","acronym":"ICVGIP 2025"},"container-title":["Proceedings of the Sixteen Indian Conference on Computer Vision, Graphics and Image Processing"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3774521.3774542","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,31]],"date-time":"2026-07-31T08:01:26Z","timestamp":1785484886000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3774521.3774542"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,12,17]]},"references-count":31,"alternative-id":["10.1145\/3774521.3774542","10.1145\/3774521"],"URL":"https:\/\/doi.org\/10.1145\/3774521.3774542","relation":{},"subject":[],"published":{"date-parts":[[2025,12,17]]},"assertion":[{"value":"2026-07-31","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}