{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,11]],"date-time":"2026-06-11T22:29:28Z","timestamp":1781216968009,"version":"3.54.1"},"reference-count":53,"publisher":"Association for Computing Machinery (ACM)","issue":"5","license":[{"start":{"date-parts":[[2022,9,23]],"date-time":"2022-09-23T00:00:00Z","timestamp":1663891200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Master, PhD Scholarship Programme of Vingroup Innovation Foundation"},{"name":"Institute of Big Data","award":["VINIF.2021.TS.026"],"award-info":[{"award-number":["VINIF.2021.TS.026"]}]},{"name":"Vietnam National University HoChiMinh City"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2022,9,30]]},"abstract":"<jats:p>Machine reading comprehension is a natural language understanding task where the computing system is required to read a text and then find the answer to a specific question posed by a human. Large-scale and high-quality corpora are necessary for evaluating machine reading comprehension models. Furthermore, machine reading comprehension (MRC) for the health sector has potential for practical applications; nevertheless, MRC research in this domain is currently scarce. This article presents UIT-ViNewsQA, a new corpus for the Vietnamese language to evaluate MRC models for the healthcare textual domain. The corpus consists of 22,057 human-generated question-answer pairs. Crowd-workers create the questions and answers on a collection of 4,416 online Vietnamese healthcare news articles, where the answers are textual spans extracted from the corresponding articles. We introduce a process for creating a high-quality corpus for the Vietnamese machine reading comprehension task. Linguistically, our corpus accommodates diversity in question and answer types. In addition, we conduct experiments and compare the effectiveness of different MRC methods based on the neural networks and transformer architectures. Experimental results on our corpus show that the MRC system based on ALBERT architecture outperforms the neural network architectures and the BERT-based approach, an exact match score of 65.26% and an F1-score of 84.89%. The best machine model achieves about 10.90% F1-score less efficiently than humans, which proves that exploring machine models on UIT-ViNewsQA to surpass humans is challenging for researchers in the future. Our corpus is publicly available on our website: http:\/\/nlp.uit.edu.vn\/datasets\n for research purposes.<\/jats:p>","DOI":"10.1145\/3527631","type":"journal-article","created":{"date-parts":[[2022,5,2]],"date-time":"2022-05-02T12:22:21Z","timestamp":1651494141000},"page":"1-28","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":19,"title":["New Vietnamese Corpus for Machine Reading Comprehension of Health News Articles"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8456-2742","authenticated-orcid":false,"given":"Kiet","family":"Van Nguyen","sequence":"first","affiliation":[{"name":"University of Information Technology, Ho Chi Minh City, Vietnam and Vietnam National University, Ho Chi Minh City, Vietnam"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4990-2868","authenticated-orcid":false,"given":"Tin","family":"Van Huynh","sequence":"additional","affiliation":[{"name":"University of Information Technology, Ho Chi Minh City, Vietnam and Vietnam National University, Ho Chi Minh City, Vietnam"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0072-2524","authenticated-orcid":false,"given":"Duc-Vu","family":"Nguyen","sequence":"additional","affiliation":[{"name":"University of Information Technology, Ho Chi Minh City, Vietnam and Vietnam National University, Ho Chi Minh City, Vietnam"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3606-4199","authenticated-orcid":false,"given":"Anh Gia-Tuan","family":"Nguyen","sequence":"additional","affiliation":[{"name":"University of Information Technology, Ho Chi Minh City, Vietnam and Vietnam National University, Ho Chi Minh City, Vietnam"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3931-849X","authenticated-orcid":false,"given":"Ngan Luu-Thuy","family":"Nguyen","sequence":"additional","affiliation":[{"name":"University of Information Technology, Ho Chi Minh City, Vietnam and Vietnam National University, Ho Chi Minh City, Vietnam"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,9,23]]},"reference":[{"key":"e_1_3_4_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/CANDARW.2019.00039"},{"key":"e_1_3_4_3_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17-1171"},{"key":"e_1_3_4_4_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.747"},{"key":"e_1_3_4_5_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1600"},{"key":"e_1_3_4_6_2","first-page":"1777","volume-title":"Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers","author":"Cui Yiming","year":"2016","unstructured":"Yiming Cui, Ting Liu, Zhipeng Chen, Shijin Wang, and Guoping Hu. 2016. Consensus attention-based neural networks for chinese reading comprehension. In Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers. 1777\u20131786."},{"key":"e_1_3_4_7_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.coling-main.589"},{"key":"e_1_3_4_8_2","first-page":"4171","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). 4171\u20134186."},{"key":"e_1_3_4_9_2","doi-asserted-by":"crossref","first-page":"1193","DOI":"10.18653\/v1\/2020.findings-emnlp.107","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2020","author":"d\u2019Hoffschmidt Martin","year":"2020","unstructured":"Martin d\u2019Hoffschmidt, Wacim Belblidia, Quentin Heinrich, Tom Brendl\u00e9, and Maxime Vidal. 2020. FQuAD: French question answering dataset. In Findings of the Association for Computational Linguistics: EMNLP 2020. Association for Computational Linguistics, Online, 1193\u20131208. https:\/\/www.aclweb.org\/anthology\/2020.findings-emnlp.107."},{"key":"e_1_3_4_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/3453651"},{"key":"e_1_3_4_11_2","first-page":"2368","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Dua Dheeru","year":"2019","unstructured":"Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019. DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). 2368\u20132378."},{"key":"e_1_3_4_12_2","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1007\/978-3-030-58219-7_1","volume-title":"Experimental IR Meets Multilinguality, Multimodality, and Interaction","author":"Efimov Pavel","year":"2020","unstructured":"Pavel Efimov, Andrey Chertok, Leonid Boytsov, and Pavel Braslavski. 2020. SberQuAD \u2013 Russian reading comprehension dataset: Description and analysis. In Experimental IR Meets Multilinguality, Multimodality, and Interaction, Avi Arampatzis, Evangelos Kanoulas, Theodora Tsikrika, Stefanos Vrochidis, Hideo Joho, Christina Lioma, Carsten Eickhoff, Aur\u00e9lie N\u00e9v\u00e9ol, Linda Cappellato, and Nicola Ferro (Eds.). Springer International Publishing, Cham, 3\u201315."},{"issue":"2","key":"e_1_3_4_13_2","first-page":"1","article-title":"A deep neural network framework for English Hindi question answering","volume":"19","author":"Gupta Deepak","year":"2019","unstructured":"Deepak Gupta, Asif Ekbal, and Pushpak Bhattacharyya. 2019. A deep neural network framework for English Hindi question answering. ACM Transactions on Asian and Low-Resource Language Information Processing (TALLIP) 19, 2 (2019), 1\u201322.","journal-title":"ACM Transactions on Asian and Low-Resource Language Information Processing (TALLIP)"},{"key":"e_1_3_4_14_2","first-page":"37","article-title":"DuReader: A Chinese machine reading comprehension dataset from real-world applications","author":"He Wei","year":"2018","unstructured":"Wei He, Kai Liu, Jing Liu, Yajuan Lyu, Shiqi Zhao, Xinyan Xiao, Yuan Liu, Yizhong Wang, Hua Wu, Qiaoqiao She, et\u00a0al. 2018. DuReader: A Chinese machine reading comprehension dataset from real-world applications. ACL 2018 (2018), 37.","journal-title":"ACL 2018"},{"key":"e_1_3_4_15_2","first-page":"1693","volume-title":"Advances in Neural Information Processing Systems","author":"Hermann Karl Moritz","year":"2015","unstructured":"Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015. Teaching machines to read and comprehend. In Advances in Neural Information Processing Systems. 1693\u20131701."},{"key":"e_1_3_4_16_2","article-title":"The Goldilocks principle: Reading children\u2019s books with explicit memory representations","author":"Hill Felix","year":"2015","unstructured":"Felix Hill, Antoine Bordes, Sumit Chopra, and Jason Weston. 2015. The Goldilocks principle: Reading children\u2019s books with explicit memory representations. arXiv preprint arXiv:1511.02301 (2015).","journal-title":"arXiv preprint arXiv:1511.02301"},{"key":"e_1_3_4_17_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1031"},{"key":"e_1_3_4_18_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Huang Hsin-Yuan","year":"2018","unstructured":"Hsin-Yuan Huang, Chenguang Zhu, Yelong Shen, and Weizhu Chen. 2018. FusionNet: Fusing via fully-aware attention with application to machine comprehension. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_4_19_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1259"},{"key":"e_1_3_4_20_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00300"},{"key":"e_1_3_4_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICTAI.2016.0128"},{"key":"e_1_3_4_22_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-2012"},{"key":"e_1_3_4_23_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D17-1082"},{"key":"e_1_3_4_24_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Lan Zhenzhong","year":"2019","unstructured":"Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019. ALBERT: A lite BERT for self-supervised learning of language representations. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_4_25_2","first-page":"1049","volume-title":"Companion Proceedings of the Web Conference 2018","author":"Le Phuong Hong","year":"2018","unstructured":"Phuong Hong Le and Duc-Thien Bui. 2018. A factoid question answering system for Vietnamese. In Companion Proceedings of the Web Conference 2018. International World Wide Web Conferences Steering Committee, 1049\u20131055."},{"key":"e_1_3_4_26_2","article-title":"Korquad1.0: Korean QA dataset for machine reading comprehension","author":"Lim Seungyoung","year":"2019","unstructured":"Seungyoung Lim, Myungji Kim, and Jooyoul Lee. 2019. Korquad1.0: Korean QA dataset for machine reading comprehension. arXiv preprint arXiv:1909.07005 (2019).","journal-title":"arXiv preprint arXiv:1909.07005"},{"key":"e_1_3_4_27_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.coling-main.233"},{"key":"e_1_3_4_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/KSE.2018.8573337"},{"key":"e_1_3_4_29_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-79457-6_49"},{"key":"e_1_3_4_30_2","first-page":"618","volume-title":"New Trends in Intelligent Software Methodologies, Tools and Techniques","author":"Nguyen Nhung Thi-Hong","year":"2021","unstructured":"Nhung Thi-Hong Nguyen, Phuong Phan-Dieu Ha, Luan Thanh Nguyen, Kiet Van Nguyen, and Ngan Luu-Thuy Nguyen. 2021. Vietnamese complaint detection on e-commerce websites. In New Trends in Intelligent Software Methodologies, Tools and Techniques. IOS Press, 618\u2013629."},{"key":"e_1_3_4_31_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10579-017-9398-3"},{"key":"e_1_3_4_32_2","doi-asserted-by":"publisher","DOI":"10.1145\/3365679"},{"key":"e_1_3_4_33_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N18-1202"},{"key":"e_1_3_4_34_2","unstructured":"Alec Radford Karthik Narasimhan Tim Salimans and Ilya Sutskever. 2018. Improving language understanding by generative pretraining. (2018). https:\/\/s3-us-west-2.amazonaws.com\/openai-assets\/research-covers\/language-unsupervised\/language_understanding_paper.pdf."},{"key":"e_1_3_4_35_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-2124"},{"key":"e_1_3_4_36_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D16-1264"},{"key":"e_1_3_4_37_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00266"},{"key":"e_1_3_4_38_2","first-page":"193","volume-title":"Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing","author":"Richardson Matthew","year":"2013","unstructured":"Matthew Richardson, Christopher J. C. Burges, and Erin Renshaw. 2013. MCTest: A challenge dataset for the open-domain machine comprehension of text. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing. 193\u2013203."},{"key":"e_1_3_4_39_2","article-title":"Bidirectional attention flow for machine comprehension","author":"Seo Minjoon","year":"2016","unstructured":"Minjoon Seo, Aniruddha Kembhavi, Ali Farhadi, and Hannaneh Hajishirzi. 2016. Bidirectional attention flow for machine comprehension. arXiv preprint arXiv:1611.01603 (2016).","journal-title":"arXiv preprint arXiv:1611.01603"},{"key":"e_1_3_4_40_2","article-title":"DRED: A Chinese machine reading comprehension dataset","author":"Shao Chih Chieh","year":"2018","unstructured":"Chih Chieh Shao, Trois Liu, Yuting Lai, Yiying Tseng, and Sam Tsai. 2018. DRED: A Chinese machine reading comprehension dataset. arXiv preprint arXiv:1806.00920 (2018).","journal-title":"arXiv preprint arXiv:1806.00920"},{"key":"e_1_3_4_41_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00264"},{"key":"e_1_3_4_42_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N18-1140"},{"key":"e_1_3_4_43_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W17-2623"},{"key":"e_1_3_4_44_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.findings-emnlp.364"},{"key":"e_1_3_4_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2020.3035701"},{"key":"e_1_3_4_46_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W18-2610"},{"key":"e_1_3_4_47_2","article-title":"Machine comprehension using match-lstm and answer pointer","author":"Wang Shuohang","year":"2016","unstructured":"Shuohang Wang and Jing Jiang. 2016. Machine comprehension using match-lstm and answer pointer. arXiv preprint arXiv:1608.07905 (2016).","journal-title":"arXiv preprint arXiv:1608.07905"},{"key":"e_1_3_4_48_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.12053"},{"key":"e_1_3_4_49_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17-1018"},{"key":"e_1_3_4_50_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/K17-1028"},{"key":"e_1_3_4_51_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1257"},{"key":"e_1_3_4_52_2","article-title":"End-to-end open-domain question answering with bertserini","author":"Yang Wei","year":"2019","unstructured":"Wei Yang, Yuqing Xie, Aileen Lin, Xingyu Li, Luchen Tan, Kun Xiong, Ming Li, and Jimmy Lin. 2019. End-to-end open-domain question answering with bertserini. arXiv preprint arXiv:1902.01718 (2019).","journal-title":"arXiv preprint arXiv:1902.01718"},{"key":"e_1_3_4_53_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Yu Adams Wei","year":"2018","unstructured":"Adams Wei Yu, David Dohan, Minh-Thang Luong, Rui Zhao, Kai Chen, Mohammad Norouzi, and Quoc V. Le. 2018. QANet: Combining local convolution with global self-attention for reading comprehension. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_4_54_2","volume-title":"Proceedings of the 32nd AAAI Conference on Artificial Intelligence","author":"Zhang Xiao","year":"2018","unstructured":"Xiao Zhang, Ji Wu, Zhiyang He, Xien Liu, and Ying Su. 2018. Medical exam question answering with large-scale reading comprehension. In Proceedings of the 32nd AAAI Conference on Artificial Intelligence."}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3527631","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3527631","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:18:53Z","timestamp":1750191533000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3527631"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,9,23]]},"references-count":53,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2022,9,30]]}},"alternative-id":["10.1145\/3527631"],"URL":"https:\/\/doi.org\/10.1145\/3527631","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"value":"2375-4699","type":"print"},{"value":"2375-4702","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,9,23]]},"assertion":[{"value":"2020-05-29","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-02-03","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-09-23","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}