{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,17]],"date-time":"2026-07-17T15:01:44Z","timestamp":1784300504609,"version":"3.55.0"},"reference-count":50,"publisher":"Association for Computing Machinery (ACM)","issue":"2","funder":[{"name":"Beijing Nova Program, the Natural Science Foundation of China","award":["62572475, 61902209, 62377044"],"award-info":[{"award-number":["62572475, 61902209, 62377044"]}]},{"name":"Beijing Outstanding Young Scientist Program","award":["NO.BJJWZYJH012019100020098"],"award-info":[{"award-number":["NO.BJJWZYJH012019100020098"]}]},{"DOI":"10.13039\/501100004260","name":"Renmin University of China","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100004260","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Inf. Syst."],"published-print":{"date-parts":[[2026,2,28]]},"abstract":"<jats:p>Recent studies have shown that large language models (LLMs) can assess relevance and support information retrieval (IR) tasks such as document ranking and relevance judgment generation. However, the internal mechanisms by which off-the-shelf LLMs understand and operationalize relevance remain largely unexplored. In this article, we systematically investigate how different LLM modules contribute to relevance judgment through the lens of mechanistic interpretability. Using activation patching techniques, we analyze the roles of various model components and identify a multi-stage, progressive process in generating either pointwise or pairwise relevance judgment. Specifically, LLMs first extract query and document information in the early layers, then process relevance information according to instructions in the middle layers, and finally utilize specific attention heads in the later layers to generate relevance judgments in the required format. Our findings provide insights into the mechanisms underlying relevance assessment in LLMs, offering valuable implications for future research on leveraging LLMs for IR tasks.<\/jats:p>","DOI":"10.1145\/3774942","type":"journal-article","created":{"date-parts":[[2025,11,11]],"date-time":"2025-11-11T14:45:36Z","timestamp":1762872336000},"page":"1-27","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["How Do Large Language Models Understand Relevance? A Mechanistic Interpretability Perspective"],"prefix":"10.1145","volume":"44","author":[{"ORCID":"https:\/\/orcid.org\/0009-0000-4144-938X","authenticated-orcid":false,"given":"Qi","family":"Liu","sequence":"first","affiliation":[{"name":"Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-2027-2568","authenticated-orcid":false,"given":"Haozhe","family":"Duan","sequence":"additional","affiliation":[{"name":"Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9257-5498","authenticated-orcid":false,"given":"Jiaxin","family":"Mao","sequence":"additional","affiliation":[{"name":"Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9777-9676","authenticated-orcid":false,"given":"Ji-Rong","family":"Wen","sequence":"additional","affiliation":[{"name":"Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,12,23]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","unstructured":"Zahra Abbasiantaeb Chuan Meng Leif Azzopardi and Mohammad Aliannejadi. 2024. Can we use large language models to fill relevance judgment holes? arXiv:2405.05600. DOI: 10.48550\/arXiv.2405.05600","DOI":"10.48550\/arXiv.2405.05600"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","unstructured":"Avishek Anand Lijun Lyu Maximilian Idahl Yumeng Wang Jonas Wallat and Zijian Zhang. 2022. Explainable information retrieval: A survey. arXiv:2211.02405. DOI: 10.48550\/arXiv.2211.02405","DOI":"10.48550\/arXiv.2211.02405"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","unstructured":"Payal Bajaj Daniel Campos Nick Craswell Li Deng Jianfeng Gao Xiaodong Liu Rangan Majumder Andrew McNamara Bhaskar Mitra Tri Nguyen et al. 2018. MS MARCO: A human generated machine reading comprehension dataset. arXiv:1611.09268. DOI: 10.48550\/arXiv.1611.09268","DOI":"10.48550\/arXiv.1611.09268"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/3626772.3657841"},{"key":"e_1_3_2_6_2","unstructured":"Shijie Chen Bernal Jim\u00e9nez Guti\u00e9rrez and Yu Su. 2024. Attention in large language models yields efficient zero-shot re-rankers. arXiv:2410.02642. Retrieved from https:\/\/arxiv.org\/abs\/2410.02642"},{"key":"e_1_3_2_7_2","unstructured":"Yiqun Chen Qi Liu Yi Zhang Weiwei Sun Daiting Shi Jiaxin Mao and Dawei Yin. 2024. TourRank: Utilizing large language models for documents ranking with a tournament-inspired strategy. arXiv:2406.11678. Retrieved from https:\/\/arxiv.org\/abs\/2406.11678"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","unstructured":"Hoagy Cunningham Aidan Ewart Logan Riggs Robert Huben and Lee Sharkey. 2023. Sparse autoencoders find highly interpretable features in language models. arXiv:2309.08600. DOI: 10.48550\/arXiv.2309.08600","DOI":"10.48550\/arXiv.2309.08600"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","unstructured":"Jacob Devlin Ming-Wei Chang Kenton Lee and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805. DOI: 10.48550\/arXiv.1810.04805","DOI":"10.48550\/arXiv.1810.04805"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","unstructured":"Nelson Elhage Tristan Hume Catherine Olsson Nicholas Schiefer Tom Henighan Shauna Kravec Zac Hatfield-Dodds Robert Lasenby Dawn Drain Carol Chen et al. 2022. Toy models of superposition. arXiv:2209.10652. DOI: 10.48550\/arXiv.2209.10652","DOI":"10.48550\/arXiv.2209.10652"},{"issue":"1","key":"e_1_3_2_11_2","first-page":"12","article-title":"A mathematical framework for transformer circuits","volume":"1","author":"Elhage Nelson","year":"2021","unstructured":"Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, et al. 2021. A mathematical framework for transformer circuits. Transformer Circuits Thread 1, 1 (2021), 12.","journal-title":"Transformer Circuits Thread"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3578337.3605136"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","unstructured":"Thibault Formal Benjamin Piwowarski and St\u00e9phane Clinchant. 2021. Match your words! A study of lexical matching in neural information retrieval. arXiv:2112.05662. DOI: 10.48550\/arXiv.2112.05662","DOI":"10.48550\/arXiv.2112.05662"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","unstructured":"Atticus Geiger Duligur Ibeling Amir Zur Maheep Chaudhary Sonakshi Chauhan Jing Huang Aryaman Arora Zhengxuan Wu Noah Goodman Christopher Potts et al. 2024. Causal abstraction: A theoretical foundation for mechanistic interpretability. arXiv:2301.04709. DOI: 10.48550\/arXiv.2301.04709","DOI":"10.48550\/arXiv.2301.04709"},{"key":"e_1_3_2_15_2","first-page":"9574","volume-title":"Advances in Neural Information Processing Systems","author":"Geiger Atticus","year":"2021","unstructured":"Atticus Geiger, Hanson Lu, Thomas Icard, and Christopher Potts. 2021. Causal abstractions of neural networks. In Advances in Neural Information Processing Systems, Vol. 34, Curran Associates, Inc., 9574\u20139586."},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","unstructured":"Mor Geva Avi Caciularu Kevin Ro Wang and Yoav Goldberg. 2022. Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space. arXiv:2203.14680. DOI: 10.48550\/arXiv.2203.14680","DOI":"10.48550\/arXiv.2203.14680"},{"key":"e_1_3_2_17_2","unstructured":"Aaron Grattafiori Abhimanyu Dubey Abhinav Jauhri Abhinav Pandey Abhishek Kadian Ahmad Al-Dahle Aiesha Letman Akhil Mathur Alan Schelten Alex Vaughan et al. 2024. The Llama 3 herd of models. arXiv:2407.21783. Retrieved from https:\/\/arxiv.org\/abs\/2407.21783"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","unstructured":"Stefan Heimersheim and Neel Nanda. 2024. How to use and interpret activation patching. arXiv:2404.15255. DOI: 10.48550\/arXiv.2404.15255","DOI":"10.48550\/arXiv.2404.15255"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","unstructured":"Albert Q. Jiang Alexandre Sablayrolles Arthur Mensch Chris Bamford Devendra Singh Chaplot Diego de las Casas Florian Bressand Gianna Lengyel Guillaume Lample Lucile Saulnier et al. 2023. Mistral 7B. arXiv:2310.06825. DOI: 10.48550\/arXiv.2310.06825","DOI":"10.48550\/arXiv.2310.06825"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","unstructured":"Omar Khattab and Matei Zaharia. 2020. ColBERT: Efficient and effective passage search via contextualized late interaction over BERT. arXiv:2004.12832. DOI: 10.48550\/arXiv.2004.12832","DOI":"10.48550\/arXiv.2004.12832"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00276"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","unstructured":"Percy Liang Rishi Bommasani Tony Lee Dimitris Tsipras Dilara Soylu Michihiro Yasunaga Yian Zhang Deepak Narayanan Yuhuai Wu Ananya Kumar et al. 2022. Holistic evaluation of language models. arXiv:2211.09110. DOI: 10.48550\/arXiv.2211.09110","DOI":"10.48550\/arXiv.2211.09110"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","unstructured":"Qi Liu Haozhe Duan Yiqun Chen Quanfeng Lu Weiwei Sun and Jiaxin Mao. 2025. LLM4Ranking: An easy-to-use framework of utilizing large language models for document reranking. arXiv:2504.07439. DOI: 10.48550\/arXiv.2504.07439","DOI":"10.48550\/arXiv.2504.07439"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3583780.3615282"},{"key":"e_1_3_2_25_2","unstructured":"Qi Liu Bo Wang Nan Wang and Jiaxin Mao. 2024. Leveraging passage embeddings for efficient listwise reranking with large language models. arXiv:2406.14848. Retrieved from https:\/\/arxiv.org\/abs\/2406.14848"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00457"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/3539618.3592032"},{"key":"e_1_3_2_28_2","unstructured":"Thomas McGrath Matthew Rahtz Janos Kramar Vladimir Mikulik and Shane Legg. 2023. The hydra effect: Emergent self-repair in language model computations. arXiv:2307.15771. Retrieved from https:\/\/arxiv.org\/abs\/2307.15771"},{"key":"e_1_3_2_29_2","first-page":"17359","article-title":"Locating and editing factual associations in GPT","volume":"35","author":"Meng Kevin","year":"2022","unstructured":"Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022. Locating and editing factual associations in GPT. In Advances in Neural Information Processing Systems, Vol. 35, 17359\u201317372.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","unstructured":"Neel Nanda Lawrence Chan Tom Lieberum Jess Smith and Jacob Steinhardt. 2023. Progress measures for grokking via mechanistic interpretability. arXiv:2301.05217. DOI: 10.48550\/arXiv.2301.05217","DOI":"10.48550\/arXiv.2301.05217"},{"key":"e_1_3_2_31_2","unstructured":"nostalgebraist. 2020. Interpreting GPT: The Logit Lens. Retrieved from https:\/\/www.lesswrong.com\/posts\/AcKRB8wDpdaN6v6ru\/interpreting-gpt-the-logit-lens"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.23915\/distill.00024.001"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.23915\/distill.00007"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","unstructured":"OpenAI. 2023. GPT-4 technical report. arXiv:2303.08774. DOI: 10.48550\/arXiv.2303.08774","DOI":"10.48550\/arXiv.2303.08774"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/3501714.3501736"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","unstructured":"Nikhil Prakash Tamar Rott Shaham Tal Haklay Yonatan Belinkov and David Bau. 2024. Fine-tuning enhances existing mechanisms: A case study on entity tracking. arXiv:2402.14811. DOI: 10.48550\/arXiv.2402.14811","DOI":"10.48550\/arXiv.2402.14811"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","unstructured":"Zhen Qin Rolf Jagerman Kai Hui Honglei Zhuang Junru Wu Jiaming Shen Tianqi Liu Jialu Liu Donald Metzler Wang Xuanhui et al. 2023. Large language models are effective text rankers with pairwise ranking prompting. arXiv:2306.17563. DOI: 10.48550\/arXiv.2306.17563","DOI":"10.48550\/arXiv.2306.17563"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","unstructured":"An Yang Baosong Yang Beichen Zhang Binyuan Hui Bo Zheng Bowen Yu Chengyuan Li Dayiheng Liu Fei Huang Haoran Wei et al. 2024. Qwen2.5 technical report. arXiv:2412.15115. DOI: 10.48550\/arXiv.2412.15115","DOI":"10.48550\/arXiv.2412.15115"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1561\/1500000019"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1002\/asi.4630260604"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","unstructured":"Alessandro Stolfo Yonatan Belinkov and Mrinmaya Sachan. 2023. A Mechanistic interpretation of arithmetic reasoning in language models using causal mediation analysis. arXiv:2305.15054. DOI: 10.48550\/arXiv.2305.15054","DOI":"10.48550\/arXiv.2305.15054"},{"key":"e_1_3_2_42_2","doi-asserted-by":"crossref","unstructured":"Weiwei Sun Lingyong Yan Xinyu Ma Pengjie Ren Dawei Yin and Zhaochun Ren. 2023. Is ChatGPT good at search? Investigating large language models as re-ranking agent. arXiv:2304.09542. Retrieved from https:\/\/arxiv.org\/abs\/2304.09542","DOI":"10.18653\/v1\/2023.emnlp-main.923"},{"key":"e_1_3_2_43_2","unstructured":"Paul Thomas Seth Spielman Nick Craswell and Bhaskar Mitra. 2023. Large language models can accurately predict searcher preferences. arXiv:2309.10621. Retrieved from https:\/\/arxiv.org\/abs\/2309.10621"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","unstructured":"Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N. Gomez Lukasz Kaiser and Illia Polosukhin. 2017. Attention is all you need. arXiv:1706.03762. DOI: 10.48550\/arXiv.1706.03762","DOI":"10.48550\/arXiv.1706.03762"},{"key":"e_1_3_2_45_2","first-page":"12388","volume-title":"Advances in Neural Information Processing Systems","author":"Vig Jesse","year":"2020","unstructured":"Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber. 2020. Investigating gender bias in language models using causal mediation analysis. In Advances in Neural Information Processing Systems, Vol. 33, Curran Associates, Inc., 12388\u201312401."},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-28238-6_17"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","unstructured":"Kevin Wang Alexandre Variengien Arthur Conmy Buck Shlegeris and Jacob Steinhardt. 2022. Interpretability in the wild: A circuit for indirect object identification in GPT-2 small. arXiv:2211.00593. DOI: 10.48550\/arXiv.2211.00593","DOI":"10.48550\/arXiv.2211.00593"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/1852102.1852106"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401325"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","unstructured":"Fred Zhang and Neel Nanda. 2024. Towards best practices of activation patching in language models: Metrics and methods. arXiv:2309.16042. DOI: 10.48550\/arXiv.2309.16042","DOI":"10.48550\/arXiv.2309.16042"},{"key":"e_1_3_2_51_2","doi-asserted-by":"crossref","unstructured":"Honglei Zhuang Zhen Qin Kai Hui Junru Wu Le Yan Xuanhui Wang and Michael Berdersky. 2023. Beyond yes and no: Improving zero-shot LLM rankers via scoring fine-grained relevance labels. arXiv:2310.14122. Retrieved from https:\/\/arxiv.org\/abs\/2310.14122","DOI":"10.18653\/v1\/2024.naacl-short.31"}],"container-title":["ACM Transactions on Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3774942","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,23]],"date-time":"2025-12-23T14:06:51Z","timestamp":1766498811000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3774942"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,12,23]]},"references-count":50,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,2,28]]}},"alternative-id":["10.1145\/3774942"],"URL":"https:\/\/doi.org\/10.1145\/3774942","relation":{},"ISSN":["1046-8188","1558-2868"],"issn-type":[{"value":"1046-8188","type":"print"},{"value":"1558-2868","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,12,23]]},"assertion":[{"value":"2025-04-26","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-11-03","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-12-23","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}