{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,3]],"date-time":"2026-06-03T11:35:16Z","timestamp":1780486516399,"version":"3.54.1"},"reference-count":36,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2024,1,11]],"date-time":"2024-01-11T00:00:00Z","timestamp":1704931200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,1,11]],"date-time":"2024-01-11T00:00:00Z","timestamp":1704931200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100000038","name":"Natural Sciences and Engineering Research Council of Canada","doi-asserted-by":"publisher","award":["DGECR-2022-00369 and RGPIN-2022-03469"],"award-info":[{"award-number":["DGECR-2022-00369 and RGPIN-2022-03469"]}],"id":[{"id":"10.13039\/501100000038","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100000038","name":"Natural Sciences and Engineering Research Council of Canada","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100000038","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100013373","name":"Alberta Machine Intelligence Institute","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100013373","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100009192","name":"Alberta Innovates","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100009192","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Rev Socionetwork Strat"],"published-print":{"date-parts":[[2024,4]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>The challenge of information overload in the legal domain increases every day. The COLIEE competition has created four challenge tasks that are intended to encourage the development of systems and methods to alleviate some of that pressure: a case law retrieval (Task 1) and entailment (Task 2), and a statute law retrieval (Task 3) and entailment (Task 4). Here we describe our methods for Task 1 and Task 4. In Task 1, we used a sentence-transformer model to create a numeric representation for each case paragraph. We then created a histogram of the similarities between a query case and a candidate case. The histogram is used to build a binary classifier that decides whether a candidate case should be noticed or not. In Task 4, our approach relies on fine-tuning a pre-trained DeBERTa large language model (LLM) trained on SNLI and MultiNLI datasets. Our method for Task 4 was ranked third among eight participating teams in the COLIEE 2023 competition. For Task 4, We also compared the performance of the DeBERTa model with those of a knowledge distillation model and ensemble methods including Random Forest and Voting.<\/jats:p>","DOI":"10.1007\/s12626-023-00153-z","type":"journal-article","created":{"date-parts":[[2024,1,11]],"date-time":"2024-01-11T05:01:31Z","timestamp":1704949291000},"page":"101-121","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":16,"title":["Legal Information Retrieval and Entailment Using Transformer-based Approaches"],"prefix":"10.1007","volume":"18","author":[{"ORCID":"https:\/\/orcid.org\/0009-0006-4486-9738","authenticated-orcid":false,"given":"Mi-Young","family":"Kim","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Juliano","family":"Rabelo","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Housam Khalifa Bashier","family":"Babiker","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Md Abed","family":"Rahman","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Randy","family":"Goebel","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2024,1,11]]},"reference":[{"key":"153_CR1","unstructured":"Abolghasemi, A., Althammer, S., Hanbury, A., & Verberne, S. (2022). Dossier@coliee2022: Dense retrieval and neural re-ranking for legal case retrieval. In: Sixteenth international workshop on Juris-informatics (JURISIN)"},{"key":"153_CR2","doi-asserted-by":"crossref","unstructured":"Bowman, S.R., Angeli, G., Potts, C., & Manning, C.D. (2015). A large annotated corpus for learning natural language inference. arXiv preprint arXiv:1508.05326","DOI":"10.18653\/v1\/D15-1075"},{"key":"153_CR3","unstructured":"Bui, M.Q., Nguyen, C., Do, D.T., Le, N.K., Nguyen, D.H., & Nguyen, T.T.T. (2022). Using deep learning approaches for tackling legal\u2019s challenges (coliee 2022). In: Sixteenth international workshop on Juris-informatics (JURISIN)"},{"key":"153_CR4","doi-asserted-by":"publisher","unstructured":"Chalkidis, I., Fergadiotis, M., Malakasiotis, P., Aletras, N., & Androutsopoulos, I. LEGAL-BERT: The muppets straight out of law school. In: Findings of the association for computational linguistics: EMNLP 2020, pp. 2898\u20132904. Association for computational linguistics, online (2020). https:\/\/doi.org\/10.18653\/v1\/2020.findings-emnlp.261. https:\/\/www.aclweb.org\/anthology\/2020.findings-emnlp.261","DOI":"10.18653\/v1\/2020.findings-emnlp.261"},{"key":"153_CR5","unstructured":"Clark, K., Luong, M.T., Le, Q.V., & Manning, C.D. (2020). Electra: Pre-training text encoders as discriminators rather than generators. arXiv preprint arXiv:2003.10555"},{"key":"153_CR6","unstructured":"Devlin, J., Chang, M., Lee, K., & Toutanova, K. (2018). BERT: pre-training of deep bidirectional transformers for language understanding. CoRR. http:\/\/arxiv.org\/abs\/1810.04805"},{"key":"153_CR7","unstructured":"Fink, T., Recski, G., Kusa, W., & Hanbury, A. (2022). Statute-enhanced lexical retrieval of court cases for coliee 2022. In: Sixteenth international workshop on Juris-informatics (JURISIN)"},{"key":"153_CR8","doi-asserted-by":"crossref","unstructured":"Friedman, J.H. (2001). Greedy function approximation: a gradient boosting machine. Annals of statistics pp. 1189\u20131232","DOI":"10.1214\/aos\/1013203451"},{"key":"153_CR9","doi-asserted-by":"crossref","unstructured":"Fujita, M., Onaga, T., Ueyama, A., & Kano, Y. (2022). Legal textual entailment using ensemble of rule-based and bert-based method with data augmentations including generation without excess or deficiency. In: Sixteenth international workshop on Juris-informatics (JURISIN)","DOI":"10.1007\/978-3-031-29168-5_10"},{"key":"153_CR10","doi-asserted-by":"crossref","unstructured":"Geiger, A., Richardson, K., & Potts, C. (2020). Neural natural language inference models partially embed theories of lexical entailment and negation. In: Proceedings of the third BlackboxNLP workshop on analyzing and interpreting neural networks for NLP, pp. 163\u2013173","DOI":"10.18653\/v1\/2020.blackboxnlp-1.16"},{"key":"153_CR11","unstructured":"Gong, Y., Luo, H., & Zhang, J. (2017). Natural language inference over interaction space. arXiv preprint arXiv:1709.04348"},{"key":"153_CR12","unstructured":"He, P., Gao, J., & Chen, W. (2021). Debertav3: Improving deberta using electra-style pre-training with gradient-disentangled embedding sharing. arXiv preprint arXiv:2111.09543"},{"key":"153_CR13","unstructured":"He, P., Liu, X., Gao, J., & Chen, W. (2020). Deberta: Decoding-enhanced bert with disentangled attention. arXiv preprint arXiv:2006.03654"},{"key":"153_CR14","unstructured":"Ho, T.K. Random decision forests. (1995). In: Proceedings of 3rd international conference on document analysis and recognition, vol.\u00a01, pp. 278\u2013282. IEEE"},{"key":"153_CR15","doi-asserted-by":"crossref","unstructured":"Honnibal, M., & Johnson, M. (2015). An improved non-monotonic transition system for dependency parsing. In: Proceedings of the 2015 conference on empirical methods in natural language processing, pp. 1373\u20131378. Association for computational linguistics, Lisbon, Portugal. https:\/\/aclweb.org\/anthology\/D\/D15\/D15-1162","DOI":"10.18653\/v1\/D15-1162"},{"key":"153_CR16","doi-asserted-by":"crossref","unstructured":"Jiang, N., & Marneffe, M.d. (2019). Evaluating bert for natural language inference: A case study on the commitmentbank. In: Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP), pp. 6088\u20146093","DOI":"10.18653\/v1\/D19-1630"},{"key":"153_CR17","doi-asserted-by":"crossref","unstructured":"Kim, M.Y., Rabelo, J., Goebel, R., Yoshioka, M., Kano, Y., & Satoh, K. Coliee (2023). 2022 summary: Methods for legal document retrieval and entailment. New Frontiers in Artificial Intelligence. JSAI-isAI 2022. Lecture notes in computer science 13859","DOI":"10.1007\/978-3-031-29168-5_4"},{"key":"153_CR18","unstructured":"Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., & Soricut, R. (2020). Albert: A lite bert for self-supervised learning of language representations"},{"key":"153_CR19","unstructured":"Lin, M., Huang, S.C., & Shao, H.L. (2022). Rethinking attention: An attempting on revaluing attention weight with disjunctive union of longest uncommon subsequence for legal queries answering. In: Sixteenth international workshop on Juris-informatics (JURISIN)"},{"key":"153_CR20","unstructured":"Liu, X., Duh, K., & Gao, J. (2018). Stochastic answer networks for natural language inference. arXiv preprint arXiv:1804.07888"},{"key":"153_CR21","unstructured":"Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., & Stoyanov, V. (2019). Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692"},{"key":"153_CR22","unstructured":"Nakatani, S. Language detection library for java (2010). https:\/\/github.com\/shuyo\/language-detection"},{"key":"153_CR23","doi-asserted-by":"crossref","unstructured":"Nangia, N., & Bowman, S. (2019). Human vs. muppet: A conservative estimate of human performance on the glue benchmark. In: Proceedings of the 57th annual meeting of the association for computational linguistics, pp. 4566\u20134575","DOI":"10.18653\/v1\/P19-1449"},{"key":"153_CR24","doi-asserted-by":"crossref","unstructured":"Rabelo, J., Goebel, R., Kim, M.Y., Kano, Y., Yoshioka, M., & Satoh, K. (2022). Overview and discussion of the competition on legal information extraction\/entailment (coliee) 2021. Journal of Review of Socionetwork Strategies 16(1)","DOI":"10.1007\/s12626-022-00105-z"},{"key":"153_CR25","unstructured":"Radford, A., Narasimhan, K., Salimans, T., & Sutskever, I. (2018). Improving language understanding by generative pre-training"},{"key":"153_CR26","doi-asserted-by":"crossref","unstructured":"Ravichander, A., Naik, A., Rose, C., & Hovy, E. (2019). Equate: A benchmark evaluation framework for quantitative reasoning in natural language inference. In: Proceedings of the 23rd conference on computational natural language learning (CoNLL), pp. 349\u2013361","DOI":"10.18653\/v1\/K19-1033"},{"key":"153_CR27","doi-asserted-by":"crossref","unstructured":"Reimers, N., & Gurevych, I. (2019). Sentence-bert: Sentence embeddings using siamese bert-networks. In: Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP), pp. 3982\u20133992","DOI":"10.18653\/v1\/D19-1410"},{"key":"153_CR28","unstructured":"Rockt\u00e4schel, T., Grefenstette, E., Hermann, K.M., Ko\u010disk\u1ef3, T., & Blunsom, P. (2015). Reasoning about entailment with neural attention. arXiv preprint arXiv:1509.06664"},{"key":"153_CR29","unstructured":"Rosa, G.M., Rodrigues, R.C., Lotufo, R., & Nogueira, R. (2021). Yes, bm25 is a strong baseline for legal case retrieval. In: Proceedings of the 18th international conference on artificial intelligence and Law (ICAIL)"},{"key":"153_CR30","unstructured":"Schilder, F., Chinnappa, D., Madan, K., Harmouche, J., Vold, A., Bretz, H., & Hudzina, J. (2021). A pentapus grapples with legal reasoning. In: Proceedings of the COLIEE workshop in ICAIL"},{"key":"153_CR31","doi-asserted-by":"crossref","unstructured":"Tan, C., Wei, F., Wang, W., Lv, W., & Zhou, M. (2018). Multiway attention networks for modeling sentence pairs. In: IJCAI, pp. 4411\u20134417","DOI":"10.24963\/ijcai.2018\/613"},{"key":"153_CR32","doi-asserted-by":"crossref","unstructured":"Wang, W., Bao, H., Huang, S., Dong, L., & Wei, F. (2020). Minilmv2: Multi-head self-attention relation distillation for compressing pretrained transformers. arXiv preprint arXiv:2012.15828","DOI":"10.18653\/v1\/2021.findings-acl.188"},{"key":"153_CR33","doi-asserted-by":"crossref","unstructured":"Wang, Z., Hamza, W., & Florian, R. (2017). Bilateral multi-perspective matching for natural language sentences. In: Proceedings of the 26th International joint conference on artificial intelligence, pp. 4144\u20134150","DOI":"10.24963\/ijcai.2017\/579"},{"key":"153_CR34","doi-asserted-by":"crossref","unstructured":"Wehnert, S., Kutty, L., & Luca, E.W.D. (2022). Using textbook knowledge for statute retrieval and entailment classification. In: Sixteenth international workshop on Juris-informatics (JURISIN)","DOI":"10.1007\/978-3-031-29168-5_9"},{"key":"153_CR35","doi-asserted-by":"crossref","unstructured":"Williams, A., Nangia, N., & Bowman, S.R. (2017). A broad-coverage challenge corpus for sentence understanding through inference. arXiv preprint arXiv:1704.05426","DOI":"10.18653\/v1\/N18-1101"},{"key":"153_CR36","doi-asserted-by":"crossref","unstructured":"Yoshioka, M., Suzuki, Y., & Aoki, Y. (2022). Hukb at the coliee 2022 statute law task. In: Sixteenth International Workshop on Juris-informatics (JURISIN)","DOI":"10.1007\/978-3-031-29168-5_8"}],"container-title":["The Review of Socionetwork Strategies"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s12626-023-00153-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s12626-023-00153-z\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s12626-023-00153-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,4,18]],"date-time":"2024-04-18T10:33:00Z","timestamp":1713436380000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s12626-023-00153-z"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,1,11]]},"references-count":36,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2024,4]]}},"alternative-id":["153"],"URL":"https:\/\/doi.org\/10.1007\/s12626-023-00153-z","relation":{},"ISSN":["2523-3173","1867-3236"],"issn-type":[{"value":"2523-3173","type":"print"},{"value":"1867-3236","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,1,11]]},"assertion":[{"value":"31 August 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"28 November 2023","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"11 January 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"On behalf of all authors, the corresponding author states that there is no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}