{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,14]],"date-time":"2025-10-14T00:43:08Z","timestamp":1760402588314,"version":"build-2065373602"},"reference-count":39,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2020,4,15]],"date-time":"2020-04-15T00:00:00Z","timestamp":1586908800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100009509","name":"Kementerian Riset Teknologi Dan Pendidikan Tinggi Republik Indonesia","doi-asserted-by":"publisher","award":["215\/SP2H\/LT\/DRPM\/2019"],"award-info":[{"award-number":["215\/SP2H\/LT\/DRPM\/2019"]}],"id":[{"id":"10.13039\/501100009509","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>During the previous decades, intelligent identification of acronym and expansion pairs from a large corpus has garnered considerable research attention, particularly in the fields of text mining, entity extraction, and information retrieval. Herein, we present an improved approach to recognize the accurate acronym and expansion pairs from a large Indonesian corpus. Generally, an acronym can be either a combination of uppercase letters or a sequence of speech sounds (syllables). Our proposed approach can be computationally divided into four steps: (1) acronym candidate identification; (2) acronym and expansion pair collection; (3) feature generation; and (4) acronym and expansion pair recognition using supervised learning techniques. Further, we introduce eight numerical features and evaluate their effectiveness in representing the acronym and expansion pairs based on the precision, recall, and F-measure. Furthermore, we compare the k-nearest neighbors (K-NN), support vector machine (SVM), and bidirectional encoder representations from transformers (BERT) algorithms in terms of accurate acronym and expansion pair classification. The experimental results indicate that the SVM polynomial model that considers eight features exhibits the highest accuracy (97.93%), surpassing those of the SVM polynomial model that considers five features (90.45%), the K-NN algorithm with k = 3 that considers eight features (96.82%), the K-NN algorithm with k = 3 that considers five features (95.66%), BERT-Base model (81.64%), and BERT-Base Multilingual Cased model (88.10%). Moreover, we analyze the performance of the Hadoop technology using various numbers of data nodes to identify the acronym and expansion pairs and obtain their feature vectors. The results reveal that the Hadoop cluster containing a large number of data nodes is faster than that with fewer data nodes when processing from ten million to one hundred million pairs of acronyms and expansions.<\/jats:p>","DOI":"10.3390\/info11040210","type":"journal-article","created":{"date-parts":[[2020,4,15]],"date-time":"2020-04-15T09:19:50Z","timestamp":1586942390000},"page":"210","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["Recognizing Indonesian Acronym and Expansion Pairs with Supervised Learning and MapReduce"],"prefix":"10.3390","volume":"11","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3859-6706","authenticated-orcid":false,"given":"Taufik Fuadi","family":"Abidin","sequence":"first","affiliation":[{"name":"Department of Informatics, Universitas Syiah Kuala, Banda Aceh 23111, Indonesia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3107-7660","authenticated-orcid":false,"given":"Amir","family":"Mahazir","sequence":"additional","affiliation":[{"name":"Department of Informatics, Universitas Syiah Kuala, Banda Aceh 23111, Indonesia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9526-3409","authenticated-orcid":false,"given":"Muhammad","family":"Subianto","sequence":"additional","affiliation":[{"name":"Department of Informatics, Universitas Syiah Kuala, Banda Aceh 23111, Indonesia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7507-9476","authenticated-orcid":false,"given":"Khairul","family":"Munadi","sequence":"additional","affiliation":[{"name":"Department of Electrical and Computer Engineering, Universitas Syiah Kuala, Banda Aceh 23111, Indonesia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0370-3620","authenticated-orcid":false,"given":"Ridha","family":"Ferdhiana","sequence":"additional","affiliation":[{"name":"Department of Statistics, Universitas Syiah Kuala, Banda Aceh 23111, Indonesia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,4,15]]},"reference":[{"key":"ref_1","first-page":"431","article-title":"Big data technologies: A survey","volume":"30","author":"Oussous","year":"2018","journal-title":"J. King Saud Univ. Comput. Inf. Sci."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"314","DOI":"10.1016\/j.ins.2014.01.015","article-title":"Data-intensive applications, challenges, techniques and technologies: A survey on Big Data","volume":"275","author":"Chen","year":"2014","journal-title":"Inf. Sci."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"124","DOI":"10.1016\/j.jnca.2017.02.002","article-title":"Technologies and challenges in developing machine-to-machine applications: A survey","volume":"83","author":"Ali","year":"2017","journal-title":"J. Netw. Comput. Appl."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"684","DOI":"10.1016\/j.future.2015.09.021","article-title":"Integration of cloud computing and Internet of things: A survey","volume":"56","author":"Botta","year":"2016","journal-title":"Future Gener. Comput. Syst."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"1203","DOI":"10.1126\/science.1248506","article-title":"The parable of Google flu: Traps in big data analysis","volume":"343","author":"Lazer","year":"2014","journal-title":"Science"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"1012","DOI":"10.1038\/nature07634","article-title":"Detecting influenza epidemics using search engine query data","volume":"457","author":"Ginsberg","year":"2009","journal-title":"Nature"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"171","DOI":"10.1007\/s11036-013-0489-0","article-title":"Big data: A survey","volume":"19","author":"Chen","year":"2014","journal-title":"Mob. Netw. Appl."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"267","DOI":"10.1016\/j.future.2013.07.014","article-title":"Intelligent services for big data science","volume":"37","author":"Dobre","year":"2014","journal-title":"Future Gener. Comput. Syst."},{"key":"ref_9","unstructured":"Woetzel, J., Remes, J., Boland, B., Katrina, L.V., Sinha, S., Strube, G., Means, J., Law, J., Cadena, A., and Tann, V.V.D. (2018). Smart Cities: Digital Solutions for a More Livable Future, McKinsey Global Institute."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"293","DOI":"10.1016\/j.bushor.2017.01.004","article-title":"Big data: Dimensions, evolution, impacts, and challenges","volume":"60","author":"Lee","year":"2017","journal-title":"Bus. Horiz."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1186\/s40537-017-0077-4","article-title":"Analysis of agriculture data using data mining techniques: Application of big data","volume":"4","author":"Majumdar","year":"2017","journal-title":"J. Big Data"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Almada, M. (2019, January 17\u201321). Human intervention in automated decision-making: Toward the construction of contestable systems. Proceedings of the 17th International Conference on Artificial Intelligence and Law (ICAIL), Montreal, QC, Canada.","DOI":"10.1145\/3322640.3326699"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"191","DOI":"10.1007\/s100320050018","article-title":"Recognizing acronyms and their definitions","volume":"1","author":"Taghva","year":"1999","journal-title":"Int. J. Doc. Anal. Recognit."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Larkey, L.S., Ogilvie, P., Price, A., and Tamilio, B. (2000, January 2\u20137). Acrophile: An automated acronym extractor and server. Proceedings of the 5th ACM Conference on Digital Libraries, San Antonio, TX, USA.","DOI":"10.1145\/336597.336664"},{"key":"ref_15","unstructured":"Park, Y., and Byrd, R.J. (2001, January 3\u20134). Hybrid text mining for finding abbreviations and their definitions. Proceedings of the Conference on Empirical Methods in Natural Language Processing, Pittsburgh, PA, USA."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"612","DOI":"10.1197\/jamia.M1139","article-title":"Creating an online dictionary of abbreviations from MEDLINE","volume":"9","author":"Chang","year":"2002","journal-title":"J. Am. Med. Inform. Assoc."},{"key":"ref_17","first-page":"319","article-title":"A supervised learning approach to acronym identification","volume":"Volume 3501","author":"Lapalme","year":"2005","journal-title":"Advances in Artificial Intelligence"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"369","DOI":"10.1007\/s00500-006-0091-5","article-title":"Using SVM to extract acronym from text","volume":"11","author":"Xu","year":"2007","journal-title":"Soft Comput."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"371","DOI":"10.1007\/978-3-540-78849-2_38","article-title":"Mining, ranking, and using acronym patterns","volume":"4976","author":"Ji","year":"2008","journal-title":"Lect. Notes Comput. Sci."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"311","DOI":"10.1007\/s10489-009-0197-4","article-title":"Automatic extraction of acronym definitions from the web","volume":"34","author":"Sanchez","year":"2011","journal-title":"J. Appl. Intell."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"1073","DOI":"10.1002\/spe.2296","article-title":"Identifying the most appropriate expansion of acronyms used in wikipedia text","volume":"45","author":"Choi","year":"2015","journal-title":"Softw. Pract. Exp."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Jacobs, K., Itai, A., and Wintner, S. (2018). Acronyms: Identification, expansion and disambiguation. Ann. Math. Artif. Intell., 49.","DOI":"10.1007\/s10472-018-9608-8"},{"key":"ref_23","unstructured":"Wahyudi, J., and Abidin, T.F. (2011, January 10). Automatic determination of acronyms and their expansion from Indonesian texts data. Proceedings of the SNATIKA, Malang, Indonesia. (In Indonesian)."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Abidin, T.F., Adriman, R., and Ferdhiana, R. (2018, January 13\u201314). Performance analysis of Apache Hadoop for generating candidates of acronym and expansion pairs and their numerical features. Proceedings of the 3rd International Conference on Information Technology, Information System and Electrical Engineering, Yogyakarta, Indonesia.","DOI":"10.1109\/ICITISEE.2018.8721020"},{"key":"ref_25","unstructured":"Senthilkumar, R.M., and Jayanthi, V.E. (2018, January 27\u201328). A survey on acronym-expansion mining approaches from text and web. Proceedings of the 2nd International Conference on SCI, Vijayawada, India."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Boser, B.E., Guyon, I.M., and Vapnik, V.N. (1992, January 27\u201329). A training algorithm for optimal margin classifiers. Proceedings of the 5th Annual ACM Workshop on Computational Learning Theory, Pittsburgh, PA, USA.","DOI":"10.1145\/130385.130401"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"21","DOI":"10.1109\/TIT.1967.1053964","article-title":"Nearest neighbor pattern classification","volume":"13","author":"Cover","year":"1967","journal-title":"IEEE Trans. Inf. Theory"},{"key":"ref_28","unstructured":"Turc, I., Chang, M.-W., Lee, K., and Toutanova, K. (2019). Well-read students learn better: On the importance of pre-training compact models. arXiv."},{"key":"ref_29","unstructured":"Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2019, January 2\u20137). BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of the NAACL-HLT, Minneapolis, MN, USA."},{"key":"ref_30","first-page":"1","article-title":"MapReduce: Simplified data processing on large clusters","volume":"51","author":"Dean","year":"2004","journal-title":"Commun. ACM"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"7776","DOI":"10.1109\/ACCESS.2017.2696365","article-title":"Machine learning with big data: Challenges and approaches","volume":"5","author":"Grolinger","year":"2017","journal-title":"IEEE Access"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"832","DOI":"10.1007\/s10766-015-0395-0","article-title":"MapReduce parallel programming model: A state-of-the-art survey","volume":"44","author":"Li","year":"2016","journal-title":"Int. J. Parallel Program."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"45","DOI":"10.1016\/j.procs.2015.04.108","article-title":"Hadoop, mapreduce and HDFS: A developers perspective","volume":"48","author":"Ghazi","year":"2015","journal-title":"Procedia Comput. Sci."},{"key":"ref_34","first-page":"1","article-title":"Apriori versions based on MapReduce for mining frequent patterns on big data","volume":"47","author":"Luna","year":"2017","journal-title":"IEEE Trans. Cybern."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"313","DOI":"10.1109\/TSMC.2015.2437327","article-title":"FiDoop: Parallel mining of frequent itemsets using MapReduce","volume":"46","author":"Xun","year":"2016","journal-title":"IEEE Trans. Syst. Man Cybern. Syst."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Zhonghua, M. (2017, January 28\u201330). Seismic data attribute extraction based on Hadoop platform. Proceedings of the 2nd IEEE International Conference on Cloud Computing and Big Data Analysis, Chengdu, China.","DOI":"10.1109\/ICCCBDA.2017.7951907"},{"key":"ref_37","unstructured":"Scholkopf, B., Burges, C., and Smola, A. (1998). Making Large-Scale SVM Learning Practical, MIT Press."},{"key":"ref_38","unstructured":"Witten, I.H., Frank, E., and Hall, M. (2011). Data Mining: Practical Machine Learning Tools and Techniques, Morgan Kaufmann Publishers. [3rd ed.]."},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1186\/s13673-017-0098-1","article-title":"A novel lightweight URL phishing detection system using SVM and similarity index","volume":"7","author":"Zouina","year":"2017","journal-title":"Hum. Centric Comput. Inf. Sci."}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/11\/4\/210\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,13]],"date-time":"2025-10-13T13:45:01Z","timestamp":1760363101000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/11\/4\/210"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,4,15]]},"references-count":39,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2020,4]]}},"alternative-id":["info11040210"],"URL":"https:\/\/doi.org\/10.3390\/info11040210","relation":{},"ISSN":["2078-2489"],"issn-type":[{"type":"electronic","value":"2078-2489"}],"subject":[],"published":{"date-parts":[[2020,4,15]]}}}