{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,15]],"date-time":"2025-10-15T00:40:43Z","timestamp":1760488843281,"version":"build-2065373602"},"reference-count":49,"publisher":"Association for Computing Machinery (ACM)","issue":"1","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62272467"],"award-info":[{"award-number":["62272467"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Fundamental Research Funds for the Central Universities, the Research Funds of Renmin University of China","award":["22XNKJ34"],"award-info":[{"award-number":["22XNKJ34"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Inf. Syst."],"published-print":{"date-parts":[[2026,1,31]]},"abstract":"<jats:p>Search result diversification focuses on providing relevant and diverse documents covering different users\u2019 intents. Intuitively, training an effective and stable search result diversification model needs a large amount of training data. Unfortunately, annotating such data that encompass real users\u2019 search intents is expensive and time-consuming, and most existing models are trained with limited training data, which might lead to unsatisfactory ranking results. Given that Wikipedia contains massive amounts of rigorous editorial and well-structured data, in this article, we propose a pre-training framework leveraging the large-scale Wikipedia data to build weak-supervised signals. Specifically, we introduce four strategies to extract paired supervised signals reflecting the subtopic coverage information from Wikipedia. We also propose a subtopic-disentangled negative sampling strategy to sample hard negative samples and enhance the model\u2019s ability to identify subtle subtopic differences. Four auxiliary tasks are devised to pre-train the Transformer model, which is further adopted as the representation generation model in the downstream diversified ranking. Experimental results demonstrate that our pre-trained model can significantly improve the performance of several existing models, which confirms the effectiveness and scalability of pre-training for search result diversification.<\/jats:p>","DOI":"10.1145\/3764662","type":"journal-article","created":{"date-parts":[[2025,8,26]],"date-time":"2025-08-26T15:08:06Z","timestamp":1756220886000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["A Model-agnostic Pre-training Framework for Search Result Diversification"],"prefix":"10.1145","volume":"44","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8952-7666","authenticated-orcid":false,"given":"Zhirui","family":"Deng","sequence":"first","affiliation":[{"name":"Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9781-948X","authenticated-orcid":false,"given":"Zhicheng","family":"Dou","sequence":"additional","affiliation":[{"name":"Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9432-3251","authenticated-orcid":false,"given":"Yutao","family":"Zhu","sequence":"additional","affiliation":[{"name":"Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9777-9676","authenticated-orcid":false,"given":"Ji-Rong","family":"Wen","sequence":"additional","affiliation":[{"name":"Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,10,14]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/290941.291025"},{"key":"e_1_3_2_3_2","volume-title":"Proceedings of the8th International Conference on Learning Representations (ICLR \u201920)","author":"Chang Wei-Cheng","year":"2020","unstructured":"Wei-Cheng Chang, Felix X. Yu, Yin-Wen Chang, Yiming Yang, and Sanjiv Kumar. 2020. Pre-training tasks for embedding-based large-scale retrieval. In Proceedings of the8th International Conference on Learning Representations (ICLR \u201920). OpenReview.net. Retrieved from https:\/\/openreview.net\/forum?id=rkg-mA4FDr"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/1645953.1646033"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.5555\/3524938.3525087"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/1390334.1390446"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-04417-5_17"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3331184.3331303"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/2348283.2348296"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/2484028.2484095"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/3652852"},{"issue":"1","key":"e_1_3_2_12_2","doi-asserted-by":"crossref","first-page":"9","DOI":"10.1007\/s10791-023-09427-0","article-title":"DeepQFM: A deep learning based query facets mining method","volume":"26","author":"Deng Zhirui","year":"2023","unstructured":"Zhirui Deng, Zhicheng Dou, and Ji-Rong Wen. 2023. DeepQFM: A deep learning based query facets mining method. Information Retrieval Journal 26, 1 (2023), 9.","journal-title":"Information Retrieval Journal"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/3626772.3657888"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1145\/3616855.3635851"},{"key":"e_1_3_2_15_2","unstructured":"Zhirui Deng Zhicheng Dou Yutao Zhu Ji-Rong Wen Ruibin Xiong Mang Wang and Weipeng Chen. 2024. From novice to expert: LLM agent policy optimization via step-wise reinforcement learning. arXiv:2411.03817. Retrieved from https:\/\/arxiv.org\/abs\/2411.03817"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/n19-1423"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/1242572.1242651"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.552"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/2806416.2806455"},{"key":"e_1_3_2_20_2","unstructured":"Gautier Izacard Mathilde Caron Lucas Hosseini Sebastian Riedel Piotr Bojanowski Armand Joulin and Edouard Grave. 2021. Unsupervised dense information retrieval with contrastive learning. arXiv:2112.09118. Retrieved from https:\/\/arxiv.org\/abs\/2112.09118"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0306-4573(99)00056-4"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3077136.3080805"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401084"},{"key":"e_1_3_2_24_2","volume-title":"Proceedings of the7th International Conference on Learning Representations (ICLR \u201919)","author":"Loshchilov Ilya","year":"2019","unstructured":"Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In Proceedings of the7th International Conference on Learning Representations (ICLR \u201919). OpenReview.net. Retrieved from https:\/\/openreview.net\/forum?id=Bkg6RiCqY7"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3437963.3441777"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/3404835.3462869"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/3459637.3482286"},{"key":"e_1_3_2_28_2","article-title":"From doc2query to docTTTTTquery","volume":"6","author":"Nogueira Rodrigo","year":"2019","unstructured":"Rodrigo Nogueira, Jimmy Lin, and AI Epistemic. 2019. From doc2query to docTTTTTquery. Online Preprint 6 (2019).","journal-title":"Online Preprint"},{"key":"e_1_3_2_29_2","unstructured":"Rodrigo Frassetto Nogueira and Kyunghyun Cho. 2019. Passage re-ranking with BERT. arXiv:1901.04085. Retrieved from http:\/\/arxiv.org\/abs\/1901.04085"},{"key":"e_1_3_2_30_2","unstructured":"Rodrigo Frassetto Nogueira Wei Yang Kyunghyun Cho and Jimmy Lin. 2019. Multi-stage document ranking with BERT. arXiv:1910.14424. Retrieved from http:\/\/arxiv.org\/abs\/1910.14424"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/n18-1202"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1145\/3340531.3411914"},{"key":"e_1_3_2_33_2","first-page":"61","volume-title":"Proceedings of the NTCIR-18","author":"Rus Clara","year":"2025","unstructured":"Clara Rus, Jasmin Kareem, Chen Xu, Yuanna Liu, Zhirui Deng, and Maria Heuss. 2025. AMS42 at the NTCIR-18 FairWeb-2 task. In Proceedings of the NTCIR-18, 61\u201367."},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/1772690.1772780"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/3409256.3409839"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/1242572.1242749"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3404835.3462872"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/3534678.3539459"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1145\/1390156.1390306"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/2766462.2767710"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/2911451.2911498"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/3726302.3730280"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/3442381.3449831"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2020.102356"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1145\/3583780.3615050"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3366423.3380126"},{"key":"e_1_3_2_47_2","first-page":"3189","volume-title":"Proceedings of the 29th ACM International on Conference on Information and Knowledge Management (CIKM \u201920)","author":"Zamani Hamed","year":"2020","unstructured":"Hamed Zamani, Gord Lueck, Everest Chen, Rodolfo Quispe, Flint Luu, and Nick Craswell. 2020. MIMICS: A large-scale data collection for search clarification. In Proceedings of the 29th ACM International on Conference on Information and Knowledge Management (CIKM \u201920), 3189\u20133196."},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.11"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/2600428.2609634"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1145\/3459637.3482243"}],"container-title":["ACM Transactions on Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3764662","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,14]],"date-time":"2025-10-14T18:31:48Z","timestamp":1760466708000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3764662"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,14]]},"references-count":49,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,1,31]]}},"alternative-id":["10.1145\/3764662"],"URL":"https:\/\/doi.org\/10.1145\/3764662","relation":{},"ISSN":["1046-8188","1558-2868"],"issn-type":[{"type":"print","value":"1046-8188"},{"type":"electronic","value":"1558-2868"}],"subject":[],"published":{"date-parts":[[2025,10,14]]},"assertion":[{"value":"2024-07-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-08-17","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-10-14","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}