{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T17:56:52Z","timestamp":1783101412886,"version":"3.54.6"},"reference-count":54,"publisher":"MDPI AG","issue":"12","license":[{"start":{"date-parts":[[2025,12,12]],"date-time":"2025-12-12T00:00:00Z","timestamp":1765497600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"the Italian Ministry of University and Research under the PRIN 2022 project \u201cDici-A: Dictionary of Italian Collocations for Learners\u201d","award":["2022HXZR5E"],"award-info":[{"award-number":["2022HXZR5E"]}]},{"name":"the Italian Ministry of University and Research under the PRIN 2022 project \u201cDici-A: Dictionary of Italian Collocations for Learners\u201d","award":["J53D23008060006"],"award-info":[{"award-number":["J53D23008060006"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Computers"],"abstract":"<jats:p>The automatic construction of learners\u2019 dictionaries requires robust methods for identifying non-literal word combinations, or collocations, which represent a significant challenge for second-language (L2) learners. This paper addresses the critical initial step of accurately extracting collocation candidates from corpora to build a learner\u2019s dictionary for Italian. The adopted method and the implemented application are significant for learning the Italian language. We present a comparative study of three methodologies for identifying these candidates within a 41.7-million-word Italian corpus: a Part-Of-Speech-based approach, a syntactic dependency-based approach, and a novel Hybrid method that integrates both. The analysis yielded 2,097,595 potential collocations. Results indicate that the Hybrid method achieves superior performance in terms of Recall and Benchmark Match, identifying the most significant portion of candidates, 42.35% of the total. We conducted an in-depth analysis to refine the extracted dataset, calculating multiple statistical metrics for each candidate, which are described in detail in the paper. Such analysis allows for the classification of collocations by relevance, difficulty, and frequency of use, forming the basis for the future development of a high-quality, web-based dictionary tailored to the proficiency levels of Italian learners.<\/jats:p>","DOI":"10.3390\/computers14120552","type":"journal-article","created":{"date-parts":[[2025,12,12]],"date-time":"2025-12-12T12:52:31Z","timestamp":1765543951000},"page":"552","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Hybrid Methods for Automatic Collocation Extraction in Building a Learners\u2019 Dictionary of Italian"],"prefix":"10.3390","volume":"14","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-6815-6659","authenticated-orcid":false,"given":"Damiano","family":"Perri","sequence":"first","affiliation":[{"name":"Department of Math and Computer Science, University of Perugia, Via Luigi Vanvitelli, 1, 06123 Perugia, Umbria, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4327-520X","authenticated-orcid":false,"given":"Osvaldo","family":"Gervasi","sequence":"additional","affiliation":[{"name":"Department of Math and Computer Science, University of Perugia, Via Luigi Vanvitelli, 1, 06123 Perugia, Umbria, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9174-9065","authenticated-orcid":false,"given":"Sergio","family":"Tasso","sequence":"additional","affiliation":[{"name":"Department of Math and Computer Science, University of Perugia, Via Luigi Vanvitelli, 1, 06123 Perugia, Umbria, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9957-3903","authenticated-orcid":false,"given":"Stefania","family":"Spina","sequence":"additional","affiliation":[{"name":"Department of Italian Language, Literature and Arts, University for Foreigners of Perugia, Piazza Fortebraccio, 4, 06123 Perugia, Umbria, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5182-9394","authenticated-orcid":false,"given":"Irene","family":"Fioravanti","sequence":"additional","affiliation":[{"name":"Department of Italian Language, Literature and Arts, University for Foreigners of Perugia, Piazza Fortebraccio, 4, 06123 Perugia, Umbria, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5183-9613","authenticated-orcid":false,"given":"Fabio","family":"Zanda","sequence":"additional","affiliation":[{"name":"Department of Italian Language, Literature and Arts, University for Foreigners of Perugia, Piazza Fortebraccio, 4, 06123 Perugia, Umbria, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5520-7795","authenticated-orcid":false,"given":"Luciana","family":"Forti","sequence":"additional","affiliation":[{"name":"Department of Languages, Literature and Modern Cultures, University of Chieti \u2018G. d\u2019Annunzio\u2019, Via dei Vestini, 31, 66100 Chieti, Abruzzo, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2025,12,12]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Dale, R., Wong, K.F., Su, J., and Kwong, O.Y. (2005). Relative Compositionality of Multi-word Expressions: A Study of Verb-Noun (V-N) Collocations. Natural Language Processing\u2014IJCNLP 2005, Springer.","DOI":"10.1007\/11562214"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Biber, D., and Reppen, R. (2015). Lexicographyand phraseology. The Cambridge Handbook of English Corpus Linguistics, Cambridge University Press. Cambridge Handbooks in Language and Linguistics.","DOI":"10.1017\/CBO9781139764377"},{"key":"ref_3","first-page":"35","article-title":"The role of Learner Corpus Research in the study of L2 phraseology: Main contributions and future directions","volume":"XX","author":"Spina","year":"2020","journal-title":"Rivista Psicolinguistica Appl.\u2014J. Appl. Psycholinguist."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Benson, M., Benson, E., and Ilson, R. (1986). The BBI Dictionary of English Word Combinations, Benjamins.","DOI":"10.1075\/z.bbi1(1st)"},{"key":"ref_5","unstructured":"McIntosh, C., Francis, B., and Poole, R. (2002). Oxford Collocations Dictionary for Students of English, Oxford University Press."},{"key":"ref_6","unstructured":"Rundell, M. (2010). Macmillan Collocations Dictionary for Learners of English, Macmillan Education."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Bandyopadhyay, S., Naskar, S.K., and Ekbal, A. (2013). Emerging Applications of Natural Language Processing: Concepts and New Research, IGI Global Scientific Publishing.","DOI":"10.4018\/978-1-4666-2169-5"},{"key":"ref_8","unstructured":"Urz\u00ec, F. (2009). Dizionario delle Combinazioni Lessicali, Convivium."},{"key":"ref_9","unstructured":"Tiberii, P. (2012). Dizionario Delle Collocazioni, Zanichelli."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Lo Cascio, V. (2013). Dizionario Combinatorio Italiano, Benjamins.","DOI":"10.1075\/z.178"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Granger, S., and Paquot, M. (2012). Corpus evidence and Electronic Lexicography. Electronic Lexicography, Oxford University Press.","DOI":"10.1093\/acprof:oso\/9780199654864.001.0001"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Seretan, V. (2011). Syntax-Based Collocation Extraction, Springer.","DOI":"10.1007\/978-94-007-0134-2"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"573","DOI":"10.1080\/14622416.2024.2429946","article-title":"Advancing pharmacogenomics research: Automated extraction of insights from PubMed using SpaCy NLP framework","volume":"25","author":"Caneppa","year":"2024","journal-title":"Pharmacogenomics"},{"key":"ref_14","unstructured":"Calzolari, N., Choukri, K., Declerck, T., Goggi, S., Grobelnik, M., Maegaard, B., Mariani, J., Mazo, H., Moreno, A., and Odijk, J. (2016, January 23\u201328). UDPipe: Trainable Pipeline for Processing CoNLL-U Files Performing Tokenization, Morphological Analysis, POS Tagging and Parsing. Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC\u201916), Portoro\u017e, Slovenia."},{"key":"ref_15","unstructured":"Pastor, G.C. (2016). POS-patterns or Syntax? Comparing methods for extracting Word Combinations. Computerised and Corpus-Based Approaches to Phraseology: Monolingual and Multilingual Perspectives, Tradulex."},{"key":"ref_16","unstructured":"Wu, H., and Zhou, M. (2003, January 7\u201312). Synonymous collocation extraction using translation information. Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL 2003), Sapporo, Japan."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Lin, D. (1999, January 20\u201326). Automatic identification of non-compositional phrases. Proceedings of the 37th Annual Meeting of the Association for Computational Linguistics on Computational Linguistics, Morristown, NJ, USA.","DOI":"10.3115\/1034678.1034730"},{"key":"ref_18","unstructured":"Celikyilmaz, A., and Wen, T.H. (2020, January 5\u201310). Stanza: A Python Natural Language Processing Toolkit for Many Human Languages. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, Online."},{"key":"ref_19","unstructured":"Bender, E.M., Derczynski, L., and Isabelle, P. (2018, January 20\u201326). Contextual String Embeddings for Sequence Labeling. Proceedings of the 27th International Conference on Computational Linguistics, Santa Fe, NM, USA."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"837","DOI":"10.1162\/COLI_a_00302","article-title":"Survey: Multiword Expression Processing: A Survey","volume":"43","author":"Constant","year":"2017","journal-title":"Comput. Linguist."},{"key":"ref_21","first-page":"143","article-title":"Retrieving collocations from text: Xtract","volume":"19","author":"Smadja","year":"1993","journal-title":"Comput. Linguist."},{"key":"ref_22","unstructured":"Choueka, Y. (1988, January 21\u201324). Looking for needles in a haystack or locating interesting collocational expressions in large textual databases. Proceedings of the User-Oriented Content-Based Text and Image Handling, Paris, France."},{"key":"ref_23","unstructured":"Evert, S. (2004). The Statistics of Word Cooccurrences: Word Pairs and Collocations. [Ph.D. Thesis, University of Stuttgart]."},{"key":"ref_24","unstructured":"Krenn, B. (2000, January 9\u201312). Collocation mining: Exploiting corpora for collocation idenfication and representation. Proceedings of the KONVENS 2000, Ilmenau, Germany."},{"key":"ref_25","unstructured":"Breidt, E. (1993, January 22). Extraction of V-N-collocations from text corpora: A feasibility study for German. Proceedings of the Workshop on Very Large Corpora: Academic and Industrial Perspectives, Columbus, OH, USA."},{"key":"ref_26","unstructured":"Ritz, J. (2006, January 3\u20137). Collocation Extraction: Needs, Feeds and Results of an Extraction System for German. Proceedings of the Workshop on Multi-Word-Expressions in a Multilingual Context at the 11th Conference of the European Chapter of the Association for Computational Linguistics, Trento, Italy."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Su, K.Y., Wu, M.W., and Chang, J.S. (1994, January 27\u201330). A corpus-based approach to automatic compound extraction. Proceedings of the 32nd Annual Meeting of the Association for Computational Linguistics, Las Cruces, NM, USA.","DOI":"10.3115\/981732.981765"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"L\u00fc, Y., and Zhou, M. (2004, January 21\u201326). Collocation Translation Acquisition Using Monolingual Corpora. Proceedings of the Annual Meeting of the Association for Computational Linguistics, Barcelona, Spain.","DOI":"10.3115\/1218955.1218977"},{"key":"ref_29","unstructured":"Orliac, B., and Dillinger, M. (2003, January 18\u201322). Collocation extraction for machine translation. Proceedings of the Machine Translation Summit IX: Papers, New Orleans, LA, USA."},{"key":"ref_30","unstructured":"Markantonatou, S., Ramisch, C., Savary, A., and Vincze, V. (2017, January 4). USzeged: Identifying Verbal Multiword Expressions with POS Tagging and Parsing Techniques. Proceedings of the 13th Workshop on Multiword Expressions (MWE 2017), Valencia, Spain."},{"key":"ref_31","unstructured":"Jurafsky, D., Chai, J., Schluter, N., and Tetreault, J. (2020, January 5\u201310). Extracting Headless MWEs from Dependency Parse Trees: Parsing, Tagging, and Joint Modeling Approaches. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Online."},{"key":"ref_32","unstructured":"Bhatia, A., Bouma, G., Do\u011fru\u00f6z, A.S., Evang, K., Garcia, M., Giouli, V., Han, L., Nivre, J., and Rademaker, A. (2024, January 25). Combining Grammatical and Relational Approaches. A Hybrid Method for the Identification of Candidate Collocations from Corpora. Proceedings of the Joint Workshop on Multiword Expressions and Universal Dependencies, MWE-UD at LREC-COLING 2024, Torino, Italy."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Vorvilas, G., Pantazi, D., Paxinou, E., Feretzakis, G., Kalles, D., Kameas, A., Karousos, N., and Verykios, V.S. (2024, January 17\u201319). An Automated Text Summarization Approach for Open-ended Responses in Student Online Surveys. Proceedings of the 2024 15th International Conference on Information, Intelligence, Systems & Applications (IISA), Chania Crete, Greece.","DOI":"10.1109\/IISA62523.2024.10786702"},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3434402","article-title":"On the Anatomy of Predictive Models for Accelerating GPU Convolution Kernels and beyond","volume":"18","author":"Labini","year":"2021","journal-title":"ACM Trans. Archit. Code Optim."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Perri, D., Simonetti, M., and Gervasi, O. (2022). Deploying Serious Games for Cognitive Rehabilitation. Computers, 11.","DOI":"10.20944\/preprints202205.0242.v1"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"23877","DOI":"10.1109\/ACCESS.2025.3536923","article-title":"A Novel Computational Framework for Visual Snow Syndrome","volume":"13","author":"Perri","year":"2025","journal-title":"IEEE Access"},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"745","DOI":"10.54103\/2037-3597\/29101","article-title":"FROM PEC TO PEC24: A NEW REFERENCE CORPUS FOR ITALIAN","volume":"17","author":"Spina","year":"2025","journal-title":"Ital. Linguadue"},{"key":"ref_38","unstructured":"Vilas, B.S. (2016). Learner corpus research and phraseology in Italian as a second language: The case of the DICI-A, a learner dictionary of Italian collocations. Collocations Cross-Linguistically. Corpora, Dictionaries and Language Teaching, Memoires de la Societe Neophilologique de Helsinki."},{"key":"ref_39","first-page":"255","article-title":"Universal Dependencies","volume":"47","author":"Manning","year":"2021","journal-title":"Comput. Linguist."},{"key":"ref_40","unstructured":"Schmid, H. (2013). Probabilistic part-of-speech tagging using decision trees. New Methods in Language Processing, Routledge."},{"key":"ref_41","first-page":"354","article-title":"Il Perugia Corpus: Una risorsa di riferimento per l\u2019italiano. Composizione, annotazione e valutazione","volume":"Volume 1","author":"Basili","year":"2014","journal-title":"Proceedings of the First Italian Conference on Computational Linguistics CLiC-it 2014"},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"380","DOI":"10.1075\/ijcl.17.3.04har","article-title":"CQPweb combining power, flexibility and usability in a corpus analysis tool","volume":"17","author":"Hardie","year":"2012","journal-title":"Int. J. Corpus Linguist."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Attardi, G., Saletti, S., and Simi, M. (2015, January 3\u20134). Evolution of Italian Treebank and Dependency Parsing towards Universal Dependencies. Proceedings of the Second Italian Conference on Computational Linguistics CLiC-it 2015, Trento, Italy.","DOI":"10.4000\/books.aaccademia.1295"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Perri, D., Gervasi, O., Tasso, S., Fioravanti, I., Zanda, F., Spina, S., and Forti, L. (July, January 30). Dockerized Architecture for a Progressive Web App: An Italian Collocations Dictionary. Proceedings of the Computational Science and Its Applications, ICCSA 2025 Workshops, Istanbul, Turkey.","DOI":"10.1007\/978-3-031-97617-9_3"},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"155","DOI":"10.1111\/lang.12225","article-title":"Collocations in Corpus-Based Language Learning Research: Identifying, Comparing, and Interpreting the Evidence","volume":"67","author":"Gablasova","year":"2017","journal-title":"Lang. Learn."},{"key":"ref_46","unstructured":"Evert, S., Uhrig, P., Bartsch, S., and Proisl, T. (2017, January 27\u201329). E-VIEW-affilation\u2014A large-scale evaluation study of association measures for collocation identification. Proceedings of the Electronic Lexicography in the 21st Century. Proceedings of the eLex 2017 Conference, Brno, Czech Republic."},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"191","DOI":"10.1075\/ijcl.19111.den","article-title":"A multi-dimensional comparison of the effectiveness and efficiency of association measures in collocation extraction","volume":"27","author":"Deng","year":"2022","journal-title":"Int. J. Corpus Linguist."},{"key":"ref_48","unstructured":"Rychl\u1ef3, P. (2008, January 5\u20137). A Lexicographer-Friendly Association Score. Proceedings of the RASLAN, Karlova Studanka, Czech Republic."},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"133","DOI":"10.1515\/cllt-2015-0030","article-title":"Log-likelihood and odds ratio: Keyness statistics for different purposes of keyword analysis","volume":"14","author":"Pojanapunya","year":"2018","journal-title":"Corpus Linguist. Linguist. Theory"},{"key":"ref_50","unstructured":"Juilland, A., Brodin, D., and Davidovitch, C. (1971). Frequency Dictionary of French Words; Romance Languages and Their Structures, Mouton."},{"key":"ref_51","doi-asserted-by":"crossref","first-page":"403","DOI":"10.1075\/ijcl.13.4.02gri","article-title":"Dispersions and adjusted frequencies in corpora","volume":"13","author":"Gries","year":"2008","journal-title":"Int. J. Corpus Linguist."},{"key":"ref_52","doi-asserted-by":"crossref","first-page":"1420","DOI":"10.1148\/rg.2021210025","article-title":"Bag-of-Words Technique in Natural Language Processing: A Primer for Radiologists","volume":"41","author":"Juluru","year":"2021","journal-title":"RadioGraphics"},{"key":"ref_53","unstructured":"McGill, M., Koll, M., and Noreault, T. (2025, June 11). An Evaluation of Factors Affecting Document Ranking by Information Retrieval Systems, Available online: https:\/\/eric.ed.gov\/?id=ED188587."},{"key":"ref_54","first-page":"19","article-title":"A Survey on Similarity Measures in Text Mining","volume":"3","author":"Vijaymeena","year":"2016","journal-title":"Mach. Learn. Appl. Int. J."}],"container-title":["Computers"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-431X\/14\/12\/552\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,12]],"date-time":"2025-12-12T12:56:40Z","timestamp":1765544200000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-431X\/14\/12\/552"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,12,12]]},"references-count":54,"journal-issue":{"issue":"12","published-online":{"date-parts":[[2025,12]]}},"alternative-id":["computers14120552"],"URL":"https:\/\/doi.org\/10.3390\/computers14120552","relation":{},"ISSN":["2073-431X"],"issn-type":[{"value":"2073-431X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,12,12]]}}}