{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,17]],"date-time":"2026-04-17T16:17:47Z","timestamp":1776442667956,"version":"3.51.2"},"reference-count":87,"publisher":"Association for Computing Machinery (ACM)","issue":"5","license":[{"start":{"date-parts":[[2024,4,29]],"date-time":"2024-04-29T00:00:00Z","timestamp":1714348800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Inf. Syst."],"published-print":{"date-parts":[[2024,9,30]]},"abstract":"<jats:p>\n            Retrieval with extremely long queries and documents is a well-known and challenging task in information retrieval and is commonly known as Query-by-Document (QBD) retrieval. Specifically designed Transformer models that can handle long input sequences have not shown high effectiveness in QBD tasks in previous work. We propose a Re-Ranker based on the novel Proportional Relevance Score (RPRS) to compute the relevance score between a query and the top-\n            <jats:italic>k<\/jats:italic>\n            candidate documents. Our extensive evaluation shows RPRS obtains significantly better results than the state-of-the-art models on five different datasets. Furthermore, RPRS is highly efficient, since all documents can be pre-processed, embedded, and indexed before query time that gives our re-ranker the advantage of having a complexity of\n            <jats:italic>O(N)<\/jats:italic>\n            , where\n            <jats:italic>N<\/jats:italic>\n            is the total number of sentences in the query and candidate documents. Furthermore, our method solves the problem of the low-resource training in QBD retrieval tasks as it does not need large amounts of training data and has only three parameters with a limited range that can be optimized with a grid search even if a small amount of labeled data is available. Our detailed analysis shows that RPRS benefits from covering the full length of candidate documents and queries.\n          <\/jats:p>","DOI":"10.1145\/3631938","type":"journal-article","created":{"date-parts":[[2023,11,11]],"date-time":"2023-11-11T09:38:32Z","timestamp":1699695512000},"page":"1-32","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":8,"title":["Retrieval for Extremely Long Queries and Documents with RPRS: A Highly Efficient and Effective Transformer-based Re-Ranker"],"prefix":"10.1145","volume":"42","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4712-832X","authenticated-orcid":false,"given":"Arian","family":"Askari","sequence":"first","affiliation":[{"name":"Leiden University, Leiden, The Netherlands"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9609-9505","authenticated-orcid":false,"given":"Suzan","family":"Verberne","sequence":"additional","affiliation":[{"name":"Leiden University, Leiden, The Netherlands"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-3725-7312","authenticated-orcid":false,"given":"Amin","family":"Abolghasemi","sequence":"additional","affiliation":[{"name":"Leiden University, Leiden, The Netherlands"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7797-619X","authenticated-orcid":false,"given":"Wessel","family":"Kraaij","sequence":"additional","affiliation":[{"name":"Leiden University, Leiden, The Netherlands"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6080-8170","authenticated-orcid":false,"given":"Gabriella","family":"Pasi","sequence":"additional","affiliation":[{"name":"University of Milano-Bicocca, Milano, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,4,29]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3539813.3545133"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-99739-7_1"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1352"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-72240-1_1"},{"key":"e_1_3_2_6_2","first-page":"162","volume-title":"Proceedings of the 2nd International Conference on Design of Experimental Search & Information REtrieval Systems","author":"Askari A.","year":"2021","unstructured":"A. Askari and S. Verberne. 2021. Combining lexical and neural retrieval with longformer-based summarization for effective case law retrieva. In Proceedings of the 2nd International Conference on Design of Experimental Search & Information REtrieval Systems. CEUR, 162\u2013170."},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i04.5722"},{"key":"e_1_3_2_8_2","unstructured":"Iz Beltagy Matthew E. Peters and Arman Cohan. 2020. Longformer: The Long-document transformer. CoRR abs\/2004.05150 (2020). Retrieved from https:\/\/arxiv.org\/abs\/2004.05150"},{"key":"e_1_3_2_9_2","volume-title":"Natural Language Processing with Python: Analyzing Text with the Natural Language Toolkit","author":"Bird Steven","year":"2009","unstructured":"Steven Bird, Ewan Klein, and Edward Loper. 2009. Natural Language Processing with Python: Analyzing Text with the Natural Language Toolkit. O\u2019Reilly Media, Inc."},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1002\/(SICI)1097-4571(199401)45:1<12::AID-ASI2>3.0.CO;2-L"},{"key":"e_1_3_2_11_2","volume-title":"International Conference on Learning Representations","author":"Carlsson Fredrik","year":"2020","unstructured":"Fredrik Carlsson, Amaru Cuba Gyllensten, Evangelia Gogoulou, Erik Ylip\u00e4\u00e4 Hellqvist, and Magnus Sahlgren. 2020. Semantic re-tuning with contrastive tension. In International Conference on Learning Representations."},{"key":"e_1_3_2_12_2","article-title":"LEGAL-BERT: The muppets straight out of law school","author":"Chalkidis Ilias","year":"2020","unstructured":"Ilias Chalkidis, Manos Fergadiotis, Prodromos Malakasiotis, Nikolaos Aletras, and Ion Androutsopoulos. 2020. LEGAL-BERT: The muppets straight out of law school. arXiv:2010.02559. Retrieved from https:\/\/arxiv.org\/abs\/2010.02559","journal-title":"arXiv:2010.02559"},{"key":"e_1_3_2_13_2","first-page":"1597","volume-title":"International Conference on Machine Learning","author":"Chen Ting","year":"2020","unstructured":"Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A simple framework for contrastive learning of visual representations. In International Conference on Machine Learning. PMLR, 1597\u20131607."},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-99736-6_8"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1108\/eb026538"},{"key":"e_1_3_2_16_2","article-title":"Specter: Document-level representation learning using citation-informed transformers","author":"Cohan Arman","year":"2020","unstructured":"Arman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey, and Daniel S. Weld. 2020. Specter: Document-level representation learning using citation-informed transformers. arXiv:2004.07180. Retrieved from https:\/\/arxiv.org\/abs\/2004.07180","journal-title":"arXiv:2004.07180"},{"key":"e_1_3_2_17_2","article-title":"langdetect: Language detection library ported from Google\u2019s language detection","author":"Danilak M.","year":"2014","unstructured":"M. Danilak. 2014. langdetect: Language detection library ported from Google\u2019s language detection. Retrieved January 19, 2014 from https:\/\/pypi.python.org\/pypi\/langdetect\/","journal-title":"Retrieved January 19, 2014 from https:\/\/pypi.python.org\/pypi\/langdetect\/"},{"key":"e_1_3_2_18_2","article-title":"Bert: Pre-training of deep bidirectional transformers for language understanding","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805. Retrieved from https:\/\/arxiv.org\/abs\/1810.04805","journal-title":"arXiv:1810.04805"},{"key":"e_1_3_2_19_2","volume-title":"CLEF (Notebook Papers\/Labs\/Workshop)","author":"D\u2019hondt Eva","year":"2011","unstructured":"Eva D\u2019hondt, Suzan Verberne, Wouter Alink, and Roberto Cornacchia. 2011. Combining document representations for prior-art retrieval. In CLEF (Notebook Papers\/Labs\/Workshop)."},{"key":"e_1_3_2_20_2","volume-title":"Proceedings of the NTCIR Workshop","author":"Fujii Atsushi","year":"2007","unstructured":"Atsushi Fujii, Makoto Iwayama, and Noriko Kando. 2007. Overview of the patent retrieval task at the NTCIR-6 workshop. In Proceedings of the NTCIR Workshop."},{"key":"e_1_3_2_21_2","article-title":"Simcse: Simple contrastive learning of sentence embeddings","author":"Gao Tianyu","year":"2021","unstructured":"Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021. Simcse: Simple contrastive learning of sentence embeddings. arXiv:2104.08821. Retrieved from http:\/\/arxiv.org\/abs\/2104.08821","journal-title":"arXiv:2104.08821"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.findings-acl.272"},{"key":"e_1_3_2_23_2","article-title":"Don\u2019t stop pretraining: Adapt language models to domains and tasks","author":"Gururangan Suchin","year":"2020","unstructured":"Suchin Gururangan, Ana Marasovi\u0107, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020. Don\u2019t stop pretraining: Adapt language models to domains and tasks. arXiv:2004.10964. Retrieved from https:\/\/arxiv.org\/abs\/2004.10964","journal-title":"arXiv:2004.10964"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2006.100"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3404835.3462889"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-15712-8_57"},{"key":"e_1_3_2_27_2","article-title":"Poly-encoders: Transformer architectures and pre-training strategies for fast and accurate multi-sentence scoring","author":"Humeau Samuel","year":"2019","unstructured":"Samuel Humeau, Kurt Shuster, Marie-Anne Lachaux, and Jason Weston. 2019. Poly-encoders: Transformer architectures and pre-training strategies for fast and accurate multi-sentence scoring. arXiv:1905.01969. Retrieved from https:\/\/arxiv.org\/abs\/19505.01969","journal-title":"arXiv:1905.01969"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.550"},{"key":"e_1_3_2_29_2","article-title":"Reformer: The efficient transformer","author":"Kitaev Nikita","year":"2020","unstructured":"Nikita Kitaev, \u0141ukasz Kaiser, and Anselm Levskaya. 2020. Reformer: The efficient transformer. arXiv:2001.04451. Retrieved from https:\/\/arxiv.org\/abs\/2001.04451","journal-title":"arXiv:2001.04451"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.simpa.2021.100058"},{"key":"e_1_3_2_31_2","unstructured":"Steven A. Lastres. 2015. Rebooting legal research in a digital age. Insights Paper (white paper funded by LexisNexis) (July 2013) https:\/\/www.lxisnexis.com"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.is.2021.101793"},{"key":"e_1_3_2_33_2","article-title":"On the sentence embeddings from pre-trained language models","author":"Li Bohan","year":"2020","unstructured":"Bohan Li, Hao Zhou, Junxian He, Mingxuan Wang, Yiming Yang, and Lei Li. 2020. On the sentence embeddings from pre-trained language models. arXiv:2011.05864. Retrieved from https:\/\/arxiv.org\/abs\/2011.05864","journal-title":"arXiv:2011.05864"},{"issue":"3","key":"e_1_3_2_34_2","first-page":"1","article-title":"The power of selecting key blocks with local pre-ranking for long document information retrieval","volume":"41","author":"Li Minghan","year":"2023","unstructured":"Minghan Li, Diana Nicoleta Popa, Johan Chagnon, Yagmur Gizem Cinar, and Eric Gaussier. 2023. The power of selecting key blocks with local pre-ranking for long document information retrieval. ACM Trans. Inf. Syst. 41, 3 (2023), 1\u201335.","journal-title":"ACM Trans. Inf. Syst."},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/2808194.2809486"},{"key":"e_1_3_2_36_2","article-title":"Roberta: A robustly optimized bert pretraining approach","author":"Liu Yinhan","year":"2019","unstructured":"Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv:1907.11692. Retrieved from https:\/\/arxiv.org\/abs\/1907.11692","journal-title":"arXiv:1907.11692"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-70145-5_14"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/1458082.1458139"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1145\/2063576.2063871"},{"key":"e_1_3_2_40_2","volume-title":"Proceedings of the 8th International Competition on Legal Information Extraction\/Entailment (COLIEE\u201921)","author":"Ma Yixiao","year":"2021","unstructured":"Yixiao Ma, Yunqiu Shao, Bulou Liu, Yiqun Liu, Min Zhang, and Shaoping Ma. 2021. Retrieving legal cases from a large-scale candidate corpus. In Proceedings of the 8th International Competition on Legal Information Extraction\/Entailment (COLIEE\u201921)."},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/2661829.2661899"},{"key":"e_1_3_2_42_2","article-title":"Multi-vector models with textual guidance for fine-grained scientific document similarity","author":"Mysore Sheshera","year":"2021","unstructured":"Sheshera Mysore, Arman Cohan, and Tom Hope. 2021. Multi-vector models with textual guidance for fine-grained scientific document similarity. arXiv:2111.08366. Retrieved from https:\/\/arxiv.org\/abs\/2111.08366","journal-title":"arXiv:2111.08366"},{"key":"e_1_3_2_43_2","article-title":"CSFCube\u2013A test collection of computer science research articles for faceted query by example","author":"Mysore Sheshera","year":"2021","unstructured":"Sheshera Mysore, Tim O\u2019Gorman, Andrew McCallum, and Hamed Zamani. 2021. CSFCube\u2013A test collection of computer science research articles for faceted query by example. arXiv:2103.12906. Retrieved from https:\/\/arxiv.org\/abs\/2103.12906","journal-title":"arXiv:2103.12906"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","unstructured":"Ha-Thanh Nguyen Manh-Kien Phi Xuan-Bach Ngo Vu Tran Le-Minh Nguyen and Minh-Phuong Tu. 2022. Attentive deep neural networks for legal document retrieval. Artificial Intelligence and Law (December 2022). DOI:10.1007\/s10506-022-09341-8","DOI":"10.1007\/s10506-022-09341-8"},{"key":"e_1_3_2_45_2","article-title":"JNLP team: Deep learning for legal processing in COLIEE 2020","author":"Nguyen Ha-Thanh","year":"2020","unstructured":"Ha-Thanh Nguyen, Hai-Yen Thi Vuong, Phuong Minh Nguyen, Binh Tran Dang, Quan Minh Bui, Sinh Trong Vu, Chau Minh Nguyen, Vu Tran, Ken Satoh, and Minh Le Nguyen. 2020. JNLP team: Deep learning for legal processing in COLIEE 2020. arXiv:2011.08071. Retrieved from https:\/\/arxiv.org\/abs\/2011.08071","journal-title":"arXiv:2011.08071"},{"key":"e_1_3_2_46_2","first-page":"660","article-title":"MS MARCO: A human generated machine reading comprehension dataset","volume":"2640","author":"Nguyen Tri","year":"2016","unstructured":"Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016. MS MARCO: A human generated machine reading comprehension dataset. Choice 2640 (2016), 660.","journal-title":"Choice"},{"key":"e_1_3_2_47_2","article-title":"Passage re-ranking with BERT","author":"Nogueira Rodrigo","year":"2019","unstructured":"Rodrigo Nogueira and Kyunghyun Cho. 2019. Passage re-ranking with BERT. arXiv:1901.04085. Retrieved from https:\/\/arxiv.org\/abs\/1901.04085","journal-title":"arXiv:1901.04085"},{"key":"e_1_3_2_48_2","article-title":"Investigating the successes and failures of BERT for passage re-ranking","author":"Padigela Harshith","year":"2019","unstructured":"Harshith Padigela, Hamed Zamani, and W. Bruce Croft. 2019. Investigating the successes and failures of BERT for passage re-ranking. arXiv:1905.01758. Retrieved from https:\/\/arxiv.org\/abs\/1905.01758","journal-title":"arXiv:1905.01758"},{"key":"e_1_3_2_49_2","first-page":"8026","article-title":"Pytorch: An imperative style, high-performance deep learning library","volume":"32","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et\u00a0al. 2019. Pytorch: An imperative style, high-performance deep learning library. Adv. Neural Inf. Process. Syst. 32 (2019), 8026\u20138037.","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-22948-1_15"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-40802-1_25"},{"key":"e_1_3_2_52_2","volume-title":"CLEF (Notebook Papers\/Labs\/Workshop)","author":"Piroi Florina","year":"2011","unstructured":"Florina Piroi, Mihai Lupu, Allan Hanbury, and Veronika Zenz. 2011. CLEF-IP 2011: Retrieval in the intellectual property domain. In CLEF (Notebook Papers\/Labs\/Workshop). Citeseer."},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-demos.14"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1007\/s12626-022-00105-z"},{"key":"e_1_3_2_55_2","first-page":"84","volume-title":"JSAI International Symposium on Artificial Intelligence","author":"Rabelo Juliano","year":"2022","unstructured":"Juliano Rabelo, Mi-Young Kim, and Randy Goebel. 2022. Semantic-based classification of relevant case law. In JSAI International Symposium on Artificial Intelligence. Springer, 84\u201395."},{"key":"e_1_3_2_56_2","doi-asserted-by":"crossref","unstructured":"Juliano Rabelo Mi-Young Kim Randy Goebel Masaharu Yoshioka Yoshinobu Kano and Ken Satoh. 2020. COLIEE 2020: Methods for Legal Document Retrieval and Entailment. Retrieved from https:\/\/sites.ualberta.ca\/rabelo\/COLIEE2021\/COLIEE_2020_summary.pdf","DOI":"10.1007\/978-3-030-79942-7_13"},{"key":"e_1_3_2_57_2","article-title":"Sentence-bert: Sentence embeddings using siamese bert-networks","author":"Reimers Nils","year":"2019","unstructured":"Nils Reimers and Iryna Gurevych. 2019. Sentence-bert: Sentence embeddings using siamese bert-networks. arXiv:1908.10084. Retrieved from https:\/\/arxiv.org\/abs\/1908.10084","journal-title":"arXiv:1908.10084"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.1561\/1500000019"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4471-2099-5_24"},{"key":"e_1_3_2_60_2","article-title":"Yes, BM25 is a strong baseline for legal case retrieval","author":"Rosa Guilherme Moraes","year":"2021","unstructured":"Guilherme Moraes Rosa, Ruan Chaves Rodrigues, Roberto Lotufo, and Rodrigo Nogueira. 2021. Yes, BM25 is a strong baseline for legal case retrieval. arXiv:2105.05686. Retrieved from https:\/\/arxiv.org\/abs\/2105.05686","journal-title":"arXiv:2105.05686"},{"key":"e_1_3_2_61_2","unstructured":"Seitz Rudi. 2020. Understanding TF-IDF and BM-25. Retrieved from https:\/\/kmwllc.com\/index.php\/2020\/03\/20\/understanding-tf-idf-and-bm-25\/"},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W19-2204"},{"key":"e_1_3_2_63_2","article-title":"DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter","author":"Sanh Victor","year":"2019","unstructured":"Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019. DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter. arXiv:1910.01108. Retrieved from https:\/\/arxiv.org\/abs\/1910.01108","journal-title":"arXiv:1910.01108"},{"key":"e_1_3_2_64_2","article-title":"Longformer for MS MARCO document re-ranking task","author":"Sekuli\u0107 Ivan","year":"2020","unstructured":"Ivan Sekuli\u0107, Amir Soleimani, Mohammad Aliannejadi, and Fabio Crestani. 2020. Longformer for MS MARCO document re-ranking task. arXiv:2009.09392. Retrieved from https:\/\/arxiv.org\/abs\/2009.09392","journal-title":"arXiv:2009.09392"},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10115-018-1322-7"},{"key":"e_1_3_2_66_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2020\/484"},{"key":"e_1_3_2_67_2","first-page":"16857","article-title":"Mpnet: Masked and permuted pre-training for language understanding","volume":"33","author":"Song Kaitao","year":"2020","unstructured":"Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu. 2020. Mpnet: Masked and permuted pre-training for language understanding. Adv. Neural Inf. Process. Syst. 33 (2020), 16857\u201316867.","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_3_2_68_2","volume-title":"Leveraging the BERT Algorithm for Patents with TensorFlow and Big Query","author":"Srebrovic Rob","year":"2020","unstructured":"Rob Srebrovic and Jay Yonamine. 2020. Leveraging the BERT Algorithm for Patents with TensorFlow and Big Query. Technical Report. Google."},{"key":"e_1_3_2_69_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10506-020-09262-4"},{"key":"e_1_3_2_70_2","first-page":"5998","volume-title":"Advances in Neural Information Processing Systems","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems. 5998\u20136008."},{"key":"e_1_3_2_71_2","volume-title":"Proceedings of the 1st International Workshop on Advances in Patent Information Retrieval (AsPIRe\u201910).","author":"Verberne Suzan","year":"2010","unstructured":"Suzan Verberne, E. K. L. D\u2019hondt, N. H. J. Oostdijk, and Cornelis H. A. Koster. 2010. Quantifying the challenges in parsing patent claims. In Proceedings of the 1st International Workshop on Advances in Patent Information Retrieval (AsPIRe\u201910)."},{"key":"e_1_3_2_72_2","first-page":"497","volume-title":"Workshop of the Cross-Language Evaluation Forum for European Languages","author":"Verberne Suzan","year":"2009","unstructured":"Suzan Verberne and Eva D\u2019hondt. 2009. Prior art retrieval using the claims section as a bag of words. In Workshop of the Cross-Language Evaluation Forum for European Languages. Springer, 497\u2013501."},{"key":"e_1_3_2_73_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10791-016-9286-2"},{"key":"e_1_3_2_74_2","article-title":"TSDAE: Using transformer-based sequential denoising auto-encoder for unsupervised sentence embedding learning","author":"Wang Kexin","year":"2021","unstructured":"Kexin Wang, Nils Reimers, and Iryna Gurevych. 2021. TSDAE: Using transformer-based sequential denoising auto-encoder for unsupervised sentence embedding learning. arXiv:2104.06979. Retrieved from https:\/\/arxiv.org\/abs\/2104.06979","journal-title":"arXiv:2104.06979"},{"key":"e_1_3_2_75_2","article-title":"GPL: Generative pseudo labeling for unsupervised domain adaptation of dense retrieval","author":"Wang Kexin","year":"2021","unstructured":"Kexin Wang, Nandan Thakur, Nils Reimers, and Iryna Gurevych. 2021. GPL: Generative pseudo labeling for unsupervised domain adaptation of dense retrieval. arXiv:2112.07577. Retrieved from https:\/\/arxiv.org\/abs\/2112.07577","journal-title":"arXiv:2112.07577"},{"key":"e_1_3_2_76_2","first-page":"5776","article-title":"Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers","volume":"33","author":"Wang Wenhui","year":"2020","unstructured":"Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou. 2020. Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers. Adv. Neural Inf. Process. Syst. 33 (2020), 5776\u20135788.","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_3_2_77_2","volume-title":"Proceedings of the 16th International Workshop on Juris-informatics (JURISIN 22)","author":"Wen J.","year":"2022","unstructured":"J. Wen, Z. Zhong, Y. Bai, X. Zhao, and M. Yang. 2022. Siat@ coliee-2022: Legal case retrieval with longformer-based contrastive learning. In Proceedings of the 16th International Workshop on Juris-informatics (JURISIN 22)."},{"key":"e_1_3_2_78_2","article-title":"Huggingface\u2019s transformers: State-of-the-art natural language processing","author":"Wolf Thomas","year":"2019","unstructured":"Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, R\u00e9mi Louf, Morgan Funtowicz, et\u00a0al. 2019. Huggingface\u2019s transformers: State-of-the-art natural language processing. arXiv:1910.03771. Retrieved from https:\/\/arxiv.org\/abs\/1910.03771","journal-title":"arXiv:1910.03771"},{"key":"e_1_3_2_79_2","volume-title":"Proceedings of the Text Retrieval Conference (TREC\u201919)","author":"Yan Ming","year":"2019","unstructured":"Ming Yan, Chenliang Li, Chen Wu, Bin Bi, Wei Wang, Jiangnan Xia, and Luo Si. 2019. IDST at TREC 2019 deep learning track: Deep cascade ranking with generation-based document expansion and pre-trained language modeling. In Proceedings of the Text Retrieval Conference (TREC\u201919)."},{"key":"e_1_3_2_80_2","unstructured":"Eugene Yang David D. Lewis Ophir Frieder David A. Grossman and Roman Yurchak. 2018. Retrieval and richness when querying by document. In Design of Experimental Search & Information REtrieval Systems (DESIRES) 68\u201375."},{"key":"e_1_3_2_81_2","doi-asserted-by":"publisher","DOI":"10.1145\/1498759.1498806"},{"key":"e_1_3_2_82_2","first-page":"19","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP\u201919): System Demonstrations","author":"Yilmaz Zeynep Akkalyoncu","year":"2019","unstructured":"Zeynep Akkalyoncu Yilmaz, Shengjin Wang, Wei Yang, Haotian Zhang, and Jimmy Lin. 2019. Applying BERT to document retrieval with birch. In Proceedings of the Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP\u201919): System Demonstrations. 19\u201324."},{"key":"e_1_3_2_83_2","first-page":"3490","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP\u201919)","author":"Yilmaz Zeynep Akkalyoncu","year":"2019","unstructured":"Zeynep Akkalyoncu Yilmaz, Wei Yang, Haotian Zhang, and Jimmy Lin. 2019. Cross-domain modeling of sentence-level evidence for document retrieval. In Proceedings of the Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP\u201919). 3490\u20133496."},{"key":"e_1_3_2_84_2","first-page":"196","volume-title":"New Frontiers in Artificial Intelligence: JSAI-isAI\u201920 Workshops, JURISIN, and LENLS\u201920 Workshops, Revised Selected Papers","author":"Yoshioka Masaharu Juliano","year":"2021","unstructured":"Masaharu Juliano Yoshioka. 2021. COLIEE 2020: Methods for legal document retrieval and entailment. In New Frontiers in Artificial Intelligence: JSAI-isAI\u201920 Workshops, JURISIN, and LENLS\u201920 Workshops, Revised Selected Papers, Vol. 12758. Springer Nature, 196."},{"key":"e_1_3_2_85_2","volume-title":"Proceedings of the Conference on Neural Information Processing Systems (NeurIPS\u201920)","author":"Zaheer Manzil","year":"2020","unstructured":"Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et\u00a0al. 2020. Big bird: Transformers for longer sequences. In Proceedings of the Conference on Neural Information Processing Systems (NeurIPS\u201920)."},{"key":"e_1_3_2_86_2","volume-title":"Proceedings of the Text Retrieval Conference (TREC\u201917)","author":"Zhang Haotian","year":"2017","unstructured":"Haotian Zhang, Mustafa Abualsaud, Nimesh Ghelani, Angshuman Ghosh, Mark D. Smucker, Gordon V. Cormack, and Maura R. Grossman. 2017. UWaterlooMDS at the TREC 2017 common core track. In Proceedings of the Text Retrieval Conference (TREC\u201917)."},{"key":"e_1_3_2_87_2","article-title":"Are larger pretrained language models uniformly better? comparing performance at the instance level","author":"Zhong Ruiqi","year":"2021","unstructured":"Ruiqi Zhong, Dhruba Ghosh, Dan Klein, and Jacob Steinhardt. 2021. Are larger pretrained language models uniformly better? comparing performance at the instance level. arXiv:2105.06020. Retrieved from https:\/\/arxiv.org\/abs\/2105.06020","journal-title":"arXiv:2105.06020"},{"key":"e_1_3_2_88_2","article-title":"Rethinking soft labels for knowledge distillation: A bias-variance tradeoff perspective","author":"Zhou Helong","year":"2021","unstructured":"Helong Zhou, Liangchen Song, Jiajie Chen, Ye Zhou, Guoli Wang, Junsong Yuan, and Qian Zhang. 2021. Rethinking soft labels for knowledge distillation: A bias-variance tradeoff perspective. arXiv:2102.00650. Retrieved from https:\/\/arxiv.org\/abs\/2102.00650","journal-title":"arXiv:2102.00650"}],"container-title":["ACM Transactions on Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3631938","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3631938","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T22:51:02Z","timestamp":1750287062000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3631938"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,4,29]]},"references-count":87,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2024,9,30]]}},"alternative-id":["10.1145\/3631938"],"URL":"https:\/\/doi.org\/10.1145\/3631938","relation":{},"ISSN":["1046-8188","1558-2868"],"issn-type":[{"value":"1046-8188","type":"print"},{"value":"1558-2868","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,4,29]]},"assertion":[{"value":"2023-02-28","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-10-16","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-04-29","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}