{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:27:42Z","timestamp":1750220862542,"version":"3.41.0"},"reference-count":57,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2020,1,19]],"date-time":"2020-01-19T00:00:00Z","timestamp":1579392000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Korean government"},{"name":"Institute for Information 8 Communications Technology Planning 8 Evaluation (IITP"},{"name":"Development of Knowledge Evolutionary WiseQA Platform Technology for Human Knowledge Augmented Services","award":["2013-2-00131"],"award-info":[{"award-number":["2013-2-00131"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2020,5,31]]},"abstract":"<jats:p>\n            There has been growing interest among researchers in quality estimation (QE), which attempts to automatically predict the quality of machine translation (MT) outputs. Most existing works on QE are based on supervised approaches using quality-annotated training data. However, QE training data quality scores readily become\n            <jats:italic>imbalanced<\/jats:italic>\n            or\n            <jats:italic>skewed<\/jats:italic>\n            : QE data are mostly composed of high translation quality sentence pairs but the data lack low translation quality sentence pairs. The use of imbalanced data with an induced quality estimator tends to produce\n            <jats:italic>biased<\/jats:italic>\n            translation quality scores with \u201chigh\u201d translation quality scores assigned even to poorly translated sentences. To address the data imbalance, this article proposes a simple, efficient procedure called\n            <jats:italic>uniformly interpolated balancing<\/jats:italic>\n            to construct more\n            <jats:italic>balanced<\/jats:italic>\n            QE training data by inserting greater uniformness to training data. The proposed uniformly interpolated balancing procedure is based on the preparation of two different types of manually annotated QE data: (1)\n            <jats:italic>default skewed data<\/jats:italic>\n            and (2)\n            <jats:italic>near-uniform data<\/jats:italic>\n            . First, we obtain default skewed data in a naive manner without considering the imbalance by manually annotating qualities on MT outputs. Second, we obtain near-uniform data in a selective manner by manually annotating a subset only, which is selected from the automatically quality-estimated sentence pairs. Finally, we create\n            <jats:italic>uniformly interpolated balanced data<\/jats:italic>\n            by combining these two types of data, where one half originates from the default skewed data and the other half originates from the near-uniform data. We expect that uniformly interpolated balancing reflects the intrinsic skewness of the true quality distribution and manages the imbalance problem. Experimental results on an English-Korean quality estimation task show that the proposed uniformly interpolated balancing leads to robustness on both skewed and uniformly distributed quality test sets when compared to the test sets of other non-balanced datasets.\n          <\/jats:p>","DOI":"10.1145\/3365916","type":"journal-article","created":{"date-parts":[[2020,4,4]],"date-time":"2020-04-04T03:08:03Z","timestamp":1585969683000},"page":"1-27","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Uniformly Interpolated Balancing for Robust Prediction in Translation Quality Estimation"],"prefix":"10.1145","volume":"19","author":[{"given":"Hyun","family":"Kim","sequence":"first","affiliation":[{"name":"Electronics and Telecommunications Research Institute (ETRI), Yuseong-gu, Daejeon, Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Seung-Hoon","family":"Na","sequence":"additional","affiliation":[{"name":"Jeonbuk National University, Baekje-daero, deokjin-gu, Jeonju, Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2020,1,19]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/1007730.1007735"},{"volume-title":"Findings of the 2016 conference on machine translation. In Proceedings of the 1st Conference on Machine Translation. Association for Computational Linguistics, 131--198","year":"2016","author":"Bojar Ond\u0159ej","key":"e_1_2_1_2_1"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.5555\/1046920.1194898"},{"volume-title":"Ribeiro","year":"2016","author":"Branco Paula","key":"e_1_2_1_4_1"},{"volume-title":"Proceedings of the 1st International Workshop on Learning with Imbalanced Domains: Theory and Applications (LIDTA@PKDD\/ECML\u201917)","author":"Branco Paula","key":"e_1_2_1_5_1"},{"volume-title":"Proceedings of the 7th European Conference on Principles and Practice of Knowledge Discovery in Databases (PKDD\u201903)","author":"Chawla N. V.","key":"e_1_2_1_6_1"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.5555\/1622407.1622416"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.5555\/2390524.2390526"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/1321440.1321461"},{"volume-title":"Proceedings of the 30th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR\u201907)","author":"Ertekin Seyda","key":"e_1_2_1_10_1"},{"volume-title":"Proceedings of the 16th International Conference on Machine Learning (ICML\u201999)","author":"Fan Wei","key":"e_1_2_1_11_1"},{"volume-title":"Proceedings of the 7th Workshop on Statistical Machine Translation. Association for Computational Linguistics, 96--103","year":"2012","author":"Felice Mariano","key":"e_1_2_1_12_1"},{"key":"e_1_2_1_13_1","article-title":"SMOTE for learning from imbalanced data: Progress and challenges, marking the 15-year anniversary","volume":"61","author":"Fern\u00e1ndez Alberto","year":"2018","journal-title":"J. Artific. Intell. Res."},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1006\/inco.1995.1136"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSMCC.2011.2161285"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10590-013-9139-3"},{"volume-title":"Proceedings of the 7th Workshop on Statistical Machine Translation. Association for Computational Linguistics, 104--108","year":"2012","author":"Gonz\u00e1lez-Rubio Jes\u00fas","key":"e_1_2_1_17_1"},{"volume-title":"Proceedings of the IEEE International Joint Conference on Neural Networks (IEEE World Congress on Computational Intelligence). 1322--1328","year":"2008","author":"He Haibo","key":"e_1_2_1_18_1"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2008.239"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/TNN.2006.882812"},{"volume-title":"Class imbalances versus small disjuncts. ACM SIGKDD Expl. Newslett.\u2014Spec. Issue Learn. Imbal. Datas. 6, 1","year":"2004","author":"Jo Taeho","key":"e_1_2_1_21_1"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3109480"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W16-2384"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N16-1059"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W17-4763"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3321127"},{"key":"e_1_2_1_27_1","first-page":"545","article-title":"Quality estimation of English-Korean machine translation using neural network based predictor-estimator model","volume":"45","author":"Kim Hyun","year":"2018","journal-title":"J. Korean Inst. Inf. Sci. Eng."},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W16-2385"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1007\/s13748-016-0094-0"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W15-3037"},{"volume-title":"Proceedings of the European Conference on Artificial Intelligence (ECAI\u201998)","year":"1998","author":"Kukar Matjaz","key":"e_1_2_1_31_1"},{"volume-title":"Iterative nearest neighborhood oversampling in semisupervised learning from imbalanced data. Sci. World J","year":"2013","author":"Li Fengqi","key":"e_1_2_1_32_1"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSMCB.2008.2007853"},{"volume-title":"Class imbalance problem in data mining review. CoRR abs\/1305.1707","year":"2013","author":"Longadge Rushi","key":"e_1_2_1_34_1"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W16-2387"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00056"},{"volume-title":"Proceedings of the 25th International Conference on Computational Linguistics: Technical Papers (COLING\u201914)","year":"2014","author":"Moreau Erwan","key":"e_1_2_1_37_1"},{"volume-title":"Proceedings of the 16th Pacific-Asia Conference on Advances in Knowledge Discovery and Data Mining\u2014Volume Part I (PAKDD\u201912)","author":"Park Youngja","key":"e_1_2_1_38_1"},{"volume-title":"Proceedings of the 1st Conference on Machine Translation. Association for Computational Linguistics, 819--824","author":"Patel Raj Nath","key":"e_1_2_1_39_1"},{"volume-title":"Proceedings of the 14th Machine Translation Summit. 295--302","year":"2013","author":"Rubino Raphael","key":"e_1_2_1_40_1"},{"volume-title":"Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL\u201913)","year":"2013","author":"Sankaran Baskaran","key":"e_1_2_1_41_1"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10590-014-9164-x"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W15-3041"},{"volume-title":"Proceedings of the Association for Machine Translation in the Americas. 223--231","year":"2006","author":"Snover Matthew","key":"e_1_2_1_44_1"},{"volume-title":"Proceedings of the 48th Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, 612--621","year":"2010","author":"Soricut Radu","key":"e_1_2_1_45_1"},{"key":"e_1_2_1_46_1","unstructured":"Lucia Specia Varvara Logacheva and Carolina Scarton. 2016. WMT16 Quality Estimation Shared Task Training and Development Data. Retrieved from http:\/\/hdl.handle.net\/11372\/LRT-1646.  Lucia Specia Varvara Logacheva and Carolina Scarton. 2016. WMT16 Quality Estimation Shared Task Training and Development Data. Retrieved from http:\/\/hdl.handle.net\/11372\/LRT-1646."},{"volume-title":"Proceedings of the 51st Meeting of the Association for Computational Linguistics: System Demonstrations. Association for Computational Linguistics, 79--84","year":"2013","author":"Specia Lucia","key":"e_1_2_1_47_1"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2007.04.009"},{"volume-title":"Proceedings of the 17th International Conference on Machine Learning (ICML\u201900)","year":"2000","author":"Ting Kai Ming","key":"e_1_2_1_49_1"},{"key":"e_1_2_1_50_1","first-page":"769","article-title":"Two modifications of CNN","volume":"6","author":"Tomek Ivan","year":"1976","journal-title":"IEEE Trans. Syst., Man, Cyber."},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1111\/exsy.12081"},{"volume-title":"Proceedings of the 20th International Conference on International Conference on Machine Learning (ICML\u201903)","author":"Wu Gang","key":"e_1_2_1_52_1"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2005.95"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2008.06.108"},{"volume-title":"Proceedings of the International Conference on Machine Learning (ICML\u201903)","year":"2013","author":"Zhang Jianping","key":"e_1_2_1_55_1"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2006.17"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2012.08.010"}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3365916","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3365916","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T23:23:37Z","timestamp":1750202617000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3365916"}},"subtitle":["A Case Study of English-Korean Translation"],"short-title":[],"issued":{"date-parts":[[2020,1,19]]},"references-count":57,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2020,5,31]]}},"alternative-id":["10.1145\/3365916"],"URL":"https:\/\/doi.org\/10.1145\/3365916","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"type":"print","value":"2375-4699"},{"type":"electronic","value":"2375-4702"}],"subject":[],"published":{"date-parts":[[2020,1,19]]},"assertion":[{"value":"2019-09-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-10-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-01-19","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}