{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,6]],"date-time":"2026-08-06T10:42:48Z","timestamp":1786012968978,"version":"3.56.0"},"reference-count":42,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2022,2,7]],"date-time":"2022-02-07T00:00:00Z","timestamp":1644192000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,2,7]],"date-time":"2022-02-07T00:00:00Z","timestamp":1644192000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100002790","name":"Canadian Network for Research and Innovation in Machining Technology, Natural Sciences and Engineering Research Council of Canada","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100002790","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Rev Socionetwork Strat"],"published-print":{"date-parts":[[2022,4]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>We describe the techniques applied by the University of Alberta (UA) team in the most recent Competition on Legal Information Extraction and Entailment (COLIEE 2021). We participated in retrieval and entailment tasks for both case law and statute law; we applied a transformer-based approach for the case law entailment task, an information retrieval technique based on BM25 for legal information retrieval, and a natural language inference mechanism using semantic knowledge applied to statute law texts. This competition included 25 teams from 14 countries; our case law entailment approach was ranked no. 4 in Task 2, the BM25 technique for legal information retrieval was ranked no. 3 in Task 3, and the natural language inference technique incorporating semantic information was ranked no. 4 in Task 4. The combination of the latter two techniques on Task 5 was ranked no. 2. We also performed error analysis of our system in Task 4, which provides some insight into current state-of-the-art and research priorities for future directions.<\/jats:p>","DOI":"10.1007\/s12626-022-00103-1","type":"journal-article","created":{"date-parts":[[2022,2,7]],"date-time":"2022-02-07T08:03:00Z","timestamp":1644220980000},"page":"157-174","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":34,"title":["Legal Information Retrieval and Entailment Based on BM25, Transformer and Semantic Thesaurus Methods"],"prefix":"10.1007","volume":"16","author":[{"given":"Mi-Young","family":"Kim","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Juliano","family":"Rabelo","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kingsley","family":"Okeke","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Randy","family":"Goebel","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2022,2,7]]},"reference":[{"key":"103_CR1","doi-asserted-by":"crossref","unstructured":"Rabelo, J., Kim, M-Y., & Goebel, R. (2019). Combining similarity and transformer methods for case law entailment. In: Proceedings of the seventeenth international conference on artificial intelligence and law(Montreal, QC, Canada)(ICAIL \u201919). Association for computing machinery, New York, NY, USA, pp. 290\u2013296","DOI":"10.1145\/3322640.3326741"},{"key":"103_CR2","doi-asserted-by":"crossref","unstructured":"Rabelo, J., Kim, M-Y., & Goebel, R. (2020). Application of text entailment techniques in COLIEE 2020. In JURISIN","DOI":"10.1007\/978-3-030-79942-7_16"},{"key":"103_CR3","unstructured":"Abacha, A. B., & Demner-Fushman, D. (2019). A question-entailment approach to question answering. CoRR abs\/1901.08079 (2019). arXiv:1901.08079."},{"key":"103_CR4","unstructured":"Lloret, E., Ferr\u00e1ndez, \u00d3., Mu\u00f1oz, R., & Palomar, M. (2008). A text summarization approach under the influence of textual entailment. In: NLPCS -5th international workshop on natural language processing and cognitive science, pp. 22\u201331"},{"key":"103_CR5","doi-asserted-by":"crossref","unstructured":"Bowman, S.R., Angeli, G., Potts, C., Manning, C. D. (2015). A large annotated corpus for learning natural language inference. In Proceedings of the conference on empirical methods in natural language processing. ACL","DOI":"10.18653\/v1\/D15-1075"},{"key":"103_CR6","doi-asserted-by":"crossref","unstructured":"Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., & Bowman, S R. (2018). GLUE: A multi-task benchmark and analysis platform for natural language understanding. CoRRabs\/1804.07461. arXiv:1804.07461","DOI":"10.18653\/v1\/W18-5446"},{"key":"103_CR7","unstructured":"Androutsopoulos, I, & Malakasiotis, P. (2009). A survey of paraphrasing and textual entailment methods. CoRR abs\/0912.3747 (2009). arXiv:0912.3747"},{"key":"103_CR8","unstructured":"Matthew, E. (2018). Peters, Mark Neumann, Mohit Iyyer, Matt Gardner. ChristopherClark: Kenton Lee, and Luke Zettlemoyer. Deep contextualized wordrepresentations. In: Proc. of NAACL"},{"key":"103_CR9","unstructured":"Devlin, J., Chang, M-W., Lee, K., & Toutanova, K. (2018). BERT:pre-training of deep bidirectional transformers for language understanding. CoRRabs\/1810.04805. arXiv:1810.04805"},{"key":"103_CR10","unstructured":"Howard, J., & Ruder, S. (2018). Fine-tuned language models for text classification. CoRRabs\/1801.06146. arXiv:1801.06146"},{"key":"103_CR11","unstructured":"Dai, A M., & Le, Q V. (2015). Semi-supervised sequence learning. CoRR. arXiv:1511.01432"},{"issue":"8","key":"103_CR12","doi-asserted-by":"publisher","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","volume":"9","author":"S Hochreiter","year":"1997","unstructured":"Hochreiter, S., & Schmidhuber, J. (1997). Long short-term memory. Neural Computation, 9(8), 1735\u20131780. https:\/\/doi.org\/10.1162\/neco.1997.9.8.1735.","journal-title":"Neural Computation"},{"key":"103_CR13","doi-asserted-by":"crossref","unstructured":"Lai, G., Xie, Q., Liu, H., Yang, Y., & Hovy, E. H. (2017). RACE: Large-scale ReAding comprehension dataset from examinations. CoRRabs\/1704.04683. arXiv:1704.04683","DOI":"10.18653\/v1\/D17-1082"},{"key":"103_CR14","unstructured":"Roemmele, M., Bejan, C., & Andrew G. (2011). Choice of plausible alternatives: An evaluation of commonsense causal reasoning. In: AAAI Spring Symposium Series"},{"key":"103_CR15","doi-asserted-by":"crossref","unstructured":"Dagan, I., Glickman, O., & Magnini, B. (2005). The PASCAL recognising textual entailment challenge. In ML challenges workshop. Springer, pp. 177\u2013190","DOI":"10.1007\/11736790_9"},{"key":"103_CR16","doi-asserted-by":"crossref","unstructured":"Kano, Y., Kim, M-Y., Yoshioka, Mas., Lu, Y., Rabelo, J., Kiyota, N., Goebel, R., & Satoh, K. (2018). COLIEE-2018: evaluation of the competition on legal information extraction and entailment. In: 12th International workshop on juris-informatics","DOI":"10.1007\/978-3-030-31605-1_14"},{"key":"103_CR17","unstructured":"Chen, Y., Zhou, Y., Zhen, L., Sun, H., & Yang, W. (2018). In Twelfth international workshop on juris-informatics: legal in-formation retrieval by association rules"},{"key":"103_CR18","unstructured":"Mikolov, T., Sutskever, I., Chen, K., Corrado, G., & Dean, J. (2013). CoRR: Distributed representations of words and phrases and their compositionality"},{"key":"103_CR19","unstructured":"Le, Q V., & Mikolov, T. (2014). Distributed representations of sentences and documents. CoRRabs\/1405.4053. arXiv:1405.4053"},{"key":"103_CR20","unstructured":"Rabelo, J., Kim, M-Y., Babiker, H., Goebel, R., & Farruque, N. (2018). Legal information extraction and entailment for statute lawand case law. In: Twelfth international workshop on juris-informatics (JURISIN)"},{"issue":"1","key":"103_CR21","doi-asserted-by":"publisher","first-page":"5","DOI":"10.1023\/A:1010933404324","volume":"45","author":"L Breiman","year":"2001","unstructured":"Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5\u201332.","journal-title":"Machine Learning"},{"issue":"1","key":"103_CR22","doi-asserted-by":"publisher","first-page":"321","DOI":"10.1613\/jair.953","volume":"16","author":"NV Chawla","year":"2002","unstructured":"Chawla, N. V., Bowyer, K. W., Hall, L. O., & Kegelmeyer, W. P. (2002). SMOTE: synthetic minority over-sampling technique. The Journal of Artificial Intelligence Research, 16(1), 321\u2013357.","journal-title":"The Journal of Artificial Intelligence Research"},{"key":"103_CR23","unstructured":"Nguyen, H-T., Thi Vuong, H-Y., Nguyen, P M., Dang, B T., Bui, Q M., Vu, S T., Nguyen, C M., Tran, V., Satoh, K., Nguyen, M L. (2020). JNLP team: deep learning for legal processing in COLIEE 2020, COLIEE"},{"key":"103_CR24","doi-asserted-by":"crossref","unstructured":"Sugathadasa, K., Ayesha, B., de Silva, N., Perera, A S., Jayawardana, V., Lakmal, D., Perera, M. (2017). Synergistic Union of Word2Vec and lexicon for domain specific semantic similarity. IEEE international conference on industrial and information systems (ICIIS)","DOI":"10.1109\/ICIINFS.2017.8300343"},{"key":"103_CR25","doi-asserted-by":"crossref","unstructured":"Alberts, H., Ipek, A., Lucas, R., Wozny, P. (2020). COLIEE 2020: Legal information retrieval and entailment with legal embeddings and boosting, COLIEE","DOI":"10.1007\/978-3-030-79942-7_14"},{"key":"103_CR26","doi-asserted-by":"crossref","unstructured":"Jiang, N., de Marneffe, M C. (2019). Evaluating BERT for natural language inference: A case study on the CommitmentBank. In: Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP). pp. 6088\u20136093","DOI":"10.18653\/v1\/D19-1630"},{"key":"103_CR27","doi-asserted-by":"crossref","unstructured":"Yang, X., Zhu, X., Zhao, H., Zhang, Q., & Feng, Y. (2019). Enhancing unsupervised pretraining with external knowledge for natural language inference. In Proceeding of the Canadian conference on artificial intelligence. Springer, pp. 413\u2013419","DOI":"10.1007\/978-3-030-18305-9_38"},{"key":"103_CR28","volume-title":"MNew synonyms dictionary","author":"S Ohno","year":"1981","unstructured":"Ohno, S., & Hamanishi, M. (1981). MNew synonyms dictionary. Tokyo: Kadogawa Shoten."},{"key":"103_CR29","doi-asserted-by":"crossref","unstructured":"Shan, X., Liu, C., Xia, Y., Chen, Q., Zhang, Y., Ding, K., Liang, Y., Luo, A., & Luo, Y. (2020). GLOW : global weighted self-attention network for web search. arXiv:2007.05186","DOI":"10.1109\/BigData52589.2021.9671546"},{"key":"103_CR30","unstructured":"Williams, A., Nangia, N., & Bowman, S R. (2017). A broad-coverage challenge corpus for sentence understanding through inference. CoRRabs\/1704.05426. arXiv:1704.05426"},{"key":"103_CR31","unstructured":"Dolan, William B., & Brockett, C. (2005). Automatically constructing a corpus of sentential paraphrases. In: Proceeding of the 3rd international workshop on paraphrasing"},{"key":"103_CR32","doi-asserted-by":"crossref","unstructured":"Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., & Zettlemoyer, L. (2019). BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. arXiv preprint arXiv:1910.13461","DOI":"10.18653\/v1\/2020.acl-main.703"},{"key":"103_CR33","doi-asserted-by":"publisher","unstructured":"Chalkidis, I., Fergadiotis, M., Malakasiotis, P., Aletras, N., Androutsopoulos, I. (2020). LEGAL-BERT: the muppets straight out of law school. In Findings of the association for computational linguistics: EMNLP. Association for computational linguistics, Online, 2898\u20132904. https:\/\/doi.org\/10.18653\/v1\/2020.findings-emnlp.261","DOI":"10.18653\/v1\/2020.findings-emnlp.261"},{"key":"103_CR34","doi-asserted-by":"crossref","unstructured":"Zaragoza, H., & Robertson, S. (2009). The probabilistic relevance framework: BM25and beyond. In: Found. Trends Inf. Retr, pp. 333\u2013389","DOI":"10.1561\/1500000019"},{"key":"103_CR35","unstructured":"Robertson, S., Zaragoza, H., & Hiemstra, D. (2004). A language modeling approach to information retrieval. In: Parsimonious language models for information retrieval, pp.178\u2013185."},{"key":"103_CR36","doi-asserted-by":"crossref","unstructured":"Ponte, J M., & Croft, W B. (1998). A language modeling approach to information retrieval. In Proceedings of the 21st annual international ACM SIGIR conference on Research and development in information retrieval, pp. 275\u2013281","DOI":"10.1145\/290941.291008"},{"key":"103_CR37","doi-asserted-by":"crossref","unstructured":"Paik, J H. (2013). A novel TF-IDF weighting scheme for effective ranking. In: Proceedings of the 36th international ACM SIGIR conference on research and development in information retrieval, pp. 343\u2013352.","DOI":"10.1145\/2484028.2484070"},{"key":"103_CR38","doi-asserted-by":"crossref","unstructured":"Lafferty, J., & Zhai, C. (2004). A study of smoothing methods for language models applied to information retrieval. In: ACM Transactions on Information and Systems, pp. 179\u2013214","DOI":"10.1145\/984321.984322"},{"key":"103_CR39","doi-asserted-by":"crossref","unstructured":"Kang, S-J., & Lee, J-H. (2001). Semi-automatic practical ontology construction by using a thesaurus. In: Proceedings of the ACL 2001 workshop on human language technology and knowledge management, pp. 413\u2013419.","DOI":"10.3115\/1118220.1118226"},{"key":"103_CR40","unstructured":"Kim, M-Y., Kang, S-J., & Lee, J-H. (2001). Resolving ambiguity in inter-chunk dependency parsing. In: Proceedings of 6th natural language processing pacific rim symposium, pp. 263\u2013270"},{"key":"103_CR41","doi-asserted-by":"crossref","unstructured":"Parikh, A P., Oscar, T., Dipanjan, D., & Jakob, U. (2016). A decomposable attention model for natural language inference. arXiv preprint arXiv:1606.01933","DOI":"10.18653\/v1\/D16-1244"},{"key":"103_CR42","unstructured":"Liu, Y., Myle, O., Naman, G., Jingfei, D., Mandar, J., Danqi, C., Omer, L., Mike, L., Luke, Z., & Veselin S. (2019). Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692"}],"container-title":["The Review of Socionetwork Strategies"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s12626-022-00103-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s12626-022-00103-1\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s12626-022-00103-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,26]],"date-time":"2023-01-26T03:54:55Z","timestamp":1674705295000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s12626-022-00103-1"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,2,7]]},"references-count":42,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2022,4]]}},"alternative-id":["103"],"URL":"https:\/\/doi.org\/10.1007\/s12626-022-00103-1","relation":{},"ISSN":["2523-3173","1867-3236"],"issn-type":[{"value":"2523-3173","type":"print"},{"value":"1867-3236","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,2,7]]},"assertion":[{"value":"10 September 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"7 January 2022","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"7 February 2022","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"On behalf of all authors, the corresponding author states that there is no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}