{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,25]],"date-time":"2026-04-25T16:21:36Z","timestamp":1777134096162,"version":"3.51.4"},"reference-count":69,"publisher":"Association for Computing Machinery (ACM)","issue":"5","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62302176, 62072046, and 62302181"],"award-info":[{"award-number":["62302176, 62072046, and 62302181"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Key R&D Program of Hubei Province","award":["2023BAB017 and 2023BAB079"],"award-info":[{"award-number":["2023BAB017 and 2023BAB079"]}]},{"name":"Knowledge Innovation Program of Wuhan-Basic Research","award":["2022010801010083"],"award-info":[{"award-number":["2022010801010083"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Softw. Eng. Methodol."],"published-print":{"date-parts":[[2026,5,31]]},"abstract":"<jats:p>\n                    Image captioning (IC) systems, including Microsoft Azure Cognitive Service, are commonly utilized to convert image content into descriptive natural language. However, inaccuracies in caption generation can lead to serious misinterpretations. Advanced testing techniques such as MetaIC and ROME have been developed to mitigate these issues, yet they encounter notable challenges. First, these strategies demand intensive labor, relying on detailed manual annotations like bounding box data of objects to create test cases. Second, the realism of the generated images is compromised, with MetaIC adding unrelated objects and ROME failing to remove objects effectively. Finally, the capability to generate diversified test suites is restricted. MetaIC is limited to only inserting specific objects to prevent overlap, whereas ROME can generate only\n                    <jats:inline-formula content-type=\"math\/tex\">\n                      <jats:tex-math notation=\"LaTeX\" version=\"MathJax\">\\(3^{n}-2^{n}\\)<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    variations of test cases from an original seed image containing\n                    <jats:inline-formula content-type=\"math\/tex\">\n                      <jats:tex-math notation=\"LaTeX\" version=\"MathJax\">\\( n \\)<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    objects.\n                  <\/jats:p>\n                  <jats:p>In this study, we present SPOLRE, a novel automated tool designed for semantic preserving object layout reconstruction in image captioning system testing. SPOLRE is based on the insight that modifying the arrangement of objects within an image does not alter its inherent semantics. We utilize four semantic preserving transformation techniques\u2014translation, rotation, mirroring, and scaling\u2014to modify object layouts autonomously, eliminating the need for manual annotation. This approach enables the creation of realistic and varied test suites for IC system testing. Our extensive testing demonstrates that more than 75% of survey respondents find the images produced by SPOLRE more realistic compared to those generated by SOTA methods. Additionally, SPOLRE exhibits outstanding performance in identifying caption errors, detecting 31,544 incorrect captions across seven IC systems with an average precision of 91.62%. This significantly outperforms other methods, which only achieve 85.65% accuracy on average and identify 17,160 incorrect captions. Notably, SPOLRE exposes 6,236 unique issues within Microsoft Azure Cognitive Service, highlighting its effectiveness against one of the most advanced IC systems available.<\/jats:p>","DOI":"10.1145\/3748306","type":"journal-article","created":{"date-parts":[[2025,7,11]],"date-time":"2025-07-11T14:31:17Z","timestamp":1752244277000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["SPOLRE: Semantic Preserving Object Layout Reconstruction for Image Captioning System Testing"],"prefix":"10.1145","volume":"35","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4978-127X","authenticated-orcid":false,"given":"Yi","family":"Liu","sequence":"first","affiliation":[{"name":"Huazhong University of Science and Technology, Wuhan, China and Quantstamp, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-0172-3950","authenticated-orcid":false,"given":"Guanyu","family":"Wang","sequence":"additional","affiliation":[{"name":"Beihang University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-5792-8897","authenticated-orcid":false,"given":"Xinyi","family":"Zheng","sequence":"additional","affiliation":[{"name":"Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0046-6674","authenticated-orcid":false,"given":"Gelei","family":"Deng","sequence":"additional","affiliation":[{"name":"Nanyang Technological University, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3977-6573","authenticated-orcid":false,"given":"Kailong","family":"Wang","sequence":"additional","affiliation":[{"name":"Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7300-9215","authenticated-orcid":false,"given":"Yang","family":"Liu","sequence":"additional","affiliation":[{"name":"Nanyang Technological University, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1100-8633","authenticated-orcid":false,"given":"Haoyu","family":"Wang","sequence":"additional","affiliation":[{"name":"Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,4,24]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"Yi Liu Guanyu Wang Xinyi Zheng Gelei Deng Kailong Wang Yang Liu Haoyu Wang. 2024. SPOLRE: Semantic Preserving Object Layout Reconstruction. Retrieved from https:\/\/sites.google.com\/view\/spolre4ics"},{"key":"e_1_3_1_3_2","unstructured":"Yi Liu Guanyu Wang Xinyi Zheng Gelei Deng Kailong Wang Yang Liu Haoyu Wang. 2024. User Study. Retrieved from https:\/\/forms.gle\/SrfPzhMxbr9xt4or9"},{"key":"e_1_3_1_4_2","unstructured":"James Betker Gabriel Goh Li Jing Tim Brooks Jianfeng Wang Linjie Li Long Ouyang Juntang Zhuang Joyce Lee Yufei Guo et al. 2023. Improving image generation with better captions. Computer Science (2023). Retrieved from https:\/\/cdn.openai.com\/papers\/dall-e-3.pdf"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/3600211.3604722"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00132"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.5555\/3620237.3620531"},{"key":"e_1_3_1_8_2","doi-asserted-by":"crossref","unstructured":"Jingye Chen Yupan Huang Tengchao Lv Lei Cui Qifeng Chen and Furu Wei. 2023. TextDiffuser-2: Unleashing the power of language models for text rendering. arXiv:2311.16465. Retrieved from https:\/\/arxiv.org\/abs\/2311.16465","DOI":"10.1007\/978-3-031-72652-1_23"},{"key":"e_1_3_1_9_2","doi-asserted-by":"crossref","unstructured":"Jingye Chen Yupan Huang Tengchao Lv Lei Cui Qifeng Chen and Furu Wei. 2023. TextDiffuser: Diffusion models as text painters. arXiv:2305.10855. Retrieved from https:\/\/arxiv.org\/abs\/2305.10855","DOI":"10.52202\/075280-0410"},{"key":"e_1_3_1_10_2","unstructured":"Tsong Y. Chen Shing C. Cheung and Shiu Ming Yiu. 2020. Metamorphic testing: A new approach for generating next test cases. arXiv:2002.12543. Retrieved from https:\/\/arxiv.org\/abs\/2002.12543"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00391"},{"key":"e_1_3_1_12_2","unstructured":"Sheng-Yen Chou Pin-Yu Chen and Tsung-Yi Ho. 2023. VillanDiffusion. A unified backdoor attack framework for diffusion models. arXiv:2306.06874. Retrieved from https:\/\/arxiv.org\/abs\/2306.06874"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1177\/001316446002000104"},{"key":"e_1_3_1_14_2","doi-asserted-by":"crossref","unstructured":"Gelei Deng Yi Liu Yuekang Li Kailong Wang Ying Zhang Zefeng Li Haoyu Wang Tianwei Zhang and Yang Liu. 2023. Jailbreaker: Automated jailbreak across multiple large language model chatbots. arXiv:2307.08715. Retrieved from https:\/\/arxiv.org\/abs\/2307.08715","DOI":"10.14722\/ndss.2024.24188"},{"key":"e_1_3_1_15_2","doi-asserted-by":"crossref","unstructured":"Nouha Dziri Sivan Milton Mo Yu Osmar Zaiane and Siva Reddy. 2022. On the origin of hallucinations in conversational models: Is it the datasets or the models?. arXiv:2204.07931. Retrieved from https:\/\/arxiv.org\/abs\/2204.07931","DOI":"10.18653\/v1\/2022.naacl-main.387"},{"key":"e_1_3_1_16_2","unstructured":"Mohamed Elaraby Mengyin Lu Jacob Dunn Xueying Zhang Yu Wang Shizhu Liu Pingchuan Tian Yuping Wang and Yuxuan Wang. 2023. Halo: Estimation and reduction of hallucinations in open-source weak large language models. arXiv:2308.11764. Retrieved from https:\/\/arxiv.org\/abs\/2308.11764"},{"key":"e_1_3_1_17_2","first-page":"27","article-title":"Generative adversarial nets","author":"Goodfellow Ian","year":"2014","unstructured":"Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. Generative adversarial nets. In Advances in Neural Information Processing Systems, Vol. 27.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW59228.2023.00087"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_1_21_2","unstructured":"Yihao Huang Qing Guo and Felix Juefei-Xu. 2023. Zero-day backdoor attack against text-to-image diffusion models via personalization. arXiv:2305.10701. Retrieved from https:\/\/arxiv.org\/abs\/2305.10701"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.632"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01428"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.4097\/kjae.2015.68.6.540"},{"key":"e_1_3_1_25_2","unstructured":"Ankur Kumar. 2022. The Illustrated Image Captioning using transformers. Retrieved from https:\/\/ankur3107.github.io\/blogs\/the-illustrated-image-captioning-using-transformers"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/5.726791"},{"key":"e_1_3_1_27_2","unstructured":"Nayeon Lee Wei Ping Peng Xu Mostofa Patwary Pascale Fung Mohammad Shoeybi and Bryan Catanzaro. 2023. Factuality enhanced language models for open-ended text generation. arXiv:2206.04624. Retrieved from https:\/\/arxiv.org\/abs\/2206.04624"},{"key":"e_1_3_1_28_2","doi-asserted-by":"crossref","unstructured":"Mike Lewis Yinhan Liu Naman Goyal Marjan Ghazvininejad Abdelrahman Mohamed Omer Levy Ves Stoyanov and Luke Zettlemoyer. 2019. Bart: Denoising sequence-to-sequence pre-training for natural language generation translation and comprehension. arXiv:1910.13461. Retrieved from https:\/\/arxiv.org\/abs\/1910.13461","DOI":"10.18653\/v1\/2020.acl-main.703"},{"key":"e_1_3_1_29_2","unstructured":"Junnan Li Dongxu Li Silvio Savarese and Steven Hoi. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv:2301.12597. Retrieved from https:\/\/arxiv.org\/abs\/2301.12597"},{"key":"e_1_3_1_30_2","first-page":"12888","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Li Junnan","year":"2022","unstructured":"Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In Proceedings of the International Conference on Machine Learning. PMLR, 12888\u201312900."},{"key":"e_1_3_1_31_2","unstructured":"Lei Li Yekun Chai Shuohuan Wang Yu Sun Hao Tian Ningyu Zhang and Hua Wu. 2023. Tool-augmented reward modeling. arXiv:2310.01045. Retrieved from https:\/\/arxiv.org\/abs\/2310.01045"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1145\/3689776"},{"key":"e_1_3_1_33_2","unstructured":"Shaobo Li Xiaoguang Li Lifeng Shang Zhenhua Dong Chengjie Sun Bingquan Liu Zhenzhou Ji Xin Jiang and Qun Liu. 2022. How pre-trained language models capture factual knowledge? a causal-inspired analysis. arXiv:2203.16747. Retrieved from https:\/\/arxiv.org\/abs\/2203.16747"},{"key":"e_1_3_1_34_2","unstructured":"Wenbo Li Xin Yu Kun Zhou Yibing Song Zhe Lin and Jiaya Jia. 2022. SDM: Spatial diffusion model for large hole image inpainting. arXiv:2212.02963. Retrieved from https:\/\/arxiv.org\/abs\/2212.02963"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01117"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1214\/aoms\/1177730491"},{"key":"e_1_3_1_38_2","unstructured":"Meta. 2021. Using AI to Improve Photo Descriptions for People Who Are Blind and Visually Impaired. Retrieved from https:\/\/about.fb.com\/news\/2021\/01\/using-ai-to-improve-photo-descriptions-for-blind-and-visually-impaired-people\/"},{"key":"e_1_3_1_39_2","unstructured":"Miscrosoft. 2021. Azure AI Vision. Retrieved from https:\/\/learn.microsoft.com\/en-us\/azure\/ai-services\/computer-vision\/concept-describing-images"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/LSP.2012.2227726"},{"key":"e_1_3_1_41_2","unstructured":"Guilherme Penedo Quentin Malartic Daniel Hesslow Ruxandra Cojocaru Alessandro Cappelli Hamza Alobeidli Baptiste Pannier Ebtesam Almazrouei and Julien Launay. 2023. The RefinedWeb dataset for Falcon LLM: Outperforming curated corpora with web data and web data only. arXiv:2306.01116. Retrieved from https:\/\/arxiv.org\/abs\/2306.01116"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-demos.14"},{"key":"e_1_3_1_43_2","first-page":"8748","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning. PMLR, 8748\u20138763."},{"issue":"8","key":"e_1_3_1_44_2","first-page":"9","article-title":"Language models are unsupervised multitask learners","volume":"1","author":"Radford Alec","year":"2019","unstructured":"Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI Blog 1, 8 (2019), 9.","journal-title":"OpenAI Blog"},{"key":"e_1_3_1_45_2","unstructured":"Aditya Ramesh Prafulla Dhariwal Alex Nichol Casey Chu and Mark Chen. 2022. Hierarchical text-conditional image generation with clip latents. arXiv:2204.06125. Retrieved from https:\/\/arxiv.org\/abs\/2204.06125"},{"key":"e_1_3_1_46_2","doi-asserted-by":"crossref","unstructured":"Vipula Rawte Swagata Chakraborty Agnibh Pathak Anubhav Sarkar S. M. Towhidul Islam Tonmoy Aman Chadha Amit P. Sheth and Amitava Das. 2023. The troubling emergence of hallucination in large language models \u2013 An extensive definition quantification and prescriptive remediations. arXiv:2310.04988. Retrieved from https:\/\/arxiv.org\/abs\/2310.04988","DOI":"10.18653\/v1\/2023.emnlp-main.155"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/3528233.3530757"},{"issue":"4","key":"e_1_3_1_49_2","first-page":"4713","article-title":"Image super-resolution via iterative refinement","volume":"45","author":"Saharia Chitwan","year":"2022","unstructured":"Chitwan Saharia, Jonathan Ho, William Chan, Tim Salimans, David J. Fleet, and Mohammad Norouzi. 2022. Image super-resolution via iterative refinement. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 4 (2022), 4713\u20134726.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1238"},{"key":"e_1_3_1_51_2","doi-asserted-by":"crossref","unstructured":"Roman Suvorov Elizaveta Logacheva Anton Mashikhin Anastasia Remizova Arsenii Ashukha Aleksei Silvestrov Naejin Kong Harshith Goka Kiwoong Park and Victor Lempitsky. 2021. Resolution-robust large mask inpainting with Fourier convolutions. arXiv:2109.07161. Retrieved from https:\/\/arxiv.org\/abs\/2109.07161","DOI":"10.1109\/WACV51458.2022.00323"},{"key":"e_1_3_1_52_2","first-page":"11","article-title":"Visualizing data using t-SNE","volume":"9","author":"Van der Maaten Laurens","year":"2008","unstructured":"Laurens Van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-SNE. Journal of Machine Learning Research 9 (2008), 11.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/NCC.2015.7084843"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1145\/3650105.3652297"},{"key":"e_1_3_1_55_2","unstructured":"Jianfeng Wang Zhengyuan Yang Xiaowei Hu Linjie Li Kevin Lin Zhe Gan Zicheng Liu Ce Liu and Lijuan Wang. 2022. Git: A generative image-to-text transformer for vision and language. arXiv:2205.14100. Retrieved from https:\/\/arxiv.org\/abs\/2205.14100"},{"key":"e_1_3_1_56_2","first-page":"23318","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Wang Peng","year":"2022","unstructured":"Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang. 2022. Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework. In Proceedings of the International Conference on Machine Learning. PMLR, 23318\u201323340."},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1145\/3324884.3416584"},{"key":"e_1_3_1_58_2","unstructured":"Tengfei Wang Ting Zhang Bo Zhang Hao Ouyang Dong Chen Qifeng Chen and Fang Wen. 2022. Pretraining is all you need for image-to-image translation. arXiv:2205.12952. Retrieved from https:\/\/arxiv.org\/abs\/2205.12952"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE48619.2023.00200"},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1145\/3510003.3510212"},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.1145\/3489048.3522655"},{"key":"e_1_3_1_62_2","unstructured":"Boxi Yu Zhiqing Zhong Jiaqi Li Yixing Yang Shilin He and Pinjia He. 2023. ROME: Testing image captioning systems via recursive object melting. arXiv:2306.02228. Retrieved from https:\/\/arxiv.org\/abs\/2306.02228"},{"key":"e_1_3_1_63_2","doi-asserted-by":"publisher","DOI":"10.1145\/3533767.3534389"},{"key":"e_1_3_1_64_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01663"},{"key":"e_1_3_1_65_2","unstructured":"Hao Zhang Feng Li Xueyan Zou Shilong Liu Chunyuan Li Jianfeng Gao Jianwei Yang and Lei Zhang. 2023. A simple framework for open-vocabulary segmentation and detection. arXiv:2303.08131. Retrieved from https:\/\/arxiv.org\/abs\/2303.08131"},{"key":"e_1_3_1_66_2","doi-asserted-by":"crossref","unstructured":"Pengchuan Zhang Xiujun Li Xiaowei Hu Jianwei Yang Lei Zhang Lijuan Wang Yejin Choi and Jianfeng Gao. 2021. Vinvl: Making visual representations matter in vision-language models. arXiv:2101.00529. Retrieved from https:\/\/arxiv.org\/abs\/2101.00529","DOI":"10.1109\/CVPR46437.2021.00553"},{"key":"e_1_3_1_67_2","first-page":"10146","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhang Yuxin","unstructured":"Yuxin Zhang, Nisha Huang, Fan Tang, Haibin Huang, and Chongyang Ma. Weiming dong, and changsheng Xu. 2023. Inversion-based style transfer with diffusion models. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 10146\u201310156."},{"key":"e_1_3_1_68_2","unstructured":"Shen Zheng Jie Huang and Kevin Chen-Chuan Chang. 2023. Why does ChatGPT fall short in providing truthful answers? arXiv:2304.10513. Retrieved from https:\/\/arxiv.org\/abs\/2304.10513"},{"key":"e_1_3_1_69_2","unstructured":"Chunting Zhou Pengfei Liu Puxin Xu Srini Iyer Jiao Sun Yuning Mao Xuezhe Ma Avia Efrat Ping Yu Lili Yu et al. 2023. LIMA: Less is more for alignment. arXiv:2305.11206. Retrieved from https:\/\/arxiv.org\/abs\/2305.11206"},{"key":"e_1_3_1_70_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.244"}],"container-title":["ACM Transactions on Software Engineering and Methodology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3748306","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,25]],"date-time":"2026-04-25T06:33:51Z","timestamp":1777098831000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3748306"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,24]]},"references-count":69,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2026,5,31]]}},"alternative-id":["10.1145\/3748306"],"URL":"https:\/\/doi.org\/10.1145\/3748306","relation":{},"ISSN":["1049-331X","1557-7392"],"issn-type":[{"value":"1049-331X","type":"print"},{"value":"1557-7392","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,24]]},"assertion":[{"value":"2024-07-25","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-07-07","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-04-24","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}