{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,25]],"date-time":"2026-04-25T21:52:40Z","timestamp":1777153960241,"version":"3.51.4"},"reference-count":29,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2023,6,16]],"date-time":"2023-06-16T00:00:00Z","timestamp":1686873600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["U21B2027, 61972186, 62266028, 62266027"],"award-info":[{"award-number":["U21B2027, 61972186, 62266028, 62266027"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Yunnan Provincial Major Science and Technology Special Plan Projects","award":["202103AA080015, 202202AD080003"],"award-info":[{"award-number":["202103AA080015, 202202AD080003"]}]},{"name":"General Projects of Basic Research in Yunnan Province","award":["202201AS070179, 202201AT070915"],"award-info":[{"award-number":["202201AS070179, 202201AT070915"]}]},{"name":"Kunming University of Science and Technology \u201cdouble first-class\u201d joint project","award":["202201BE070001-021"],"award-info":[{"award-number":["202201BE070001-021"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2023,6,30]]},"abstract":"<jats:p>Cross-lingual sentence embedding\u2019s goal is mapping sentences with similar semantics but in different languages close together and dissimilar sentences farther apart in the representation space. It is the basis of many downstream tasks such as cross-lingual document matching and cross-lingual summary extraction. At present, the works of cross-lingual sentence embedding tasks mainly focus on languages with large-scale corpus. But low-resource languages such as Chinese-Vietnamese are short of sentence-level parallel corpora and clear cross-lingual monitoring signals, and these works on low-resource languages have poor performances. Therefore, we propose a cross-lingual sentence embedding method based on contrastive learning and effectively fine-tune powerful pretraining mode by constructing sentence-level positive and negative samples to avoid the catastrophic forgetting problem of the traditional fine-tuning pre-trained model based only on small-scale aligned positive samples. First, we construct positive and negative examples by taking parallel Chinese Vietnamese sentences as positive examples and non-parallel sentences as negative examples. Second, we construct a siamese network to get contrastive loss by inputting positive and negative samples and fine-tuning our model. The experimental results show that our method can effectively improve the semantic alignment accuracy of cross-lingual sentence embedding in Chinese and Vietnamese contexts.<\/jats:p>","DOI":"10.1145\/3589341","type":"journal-article","created":{"date-parts":[[2023,4,18]],"date-time":"2023-04-18T12:41:36Z","timestamp":1681821696000},"page":"1-18","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":7,"title":["Cross-lingual Sentence Embedding for Low-resource Chinese-Vietnamese Based on Contrastive Learning"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1277-6212","authenticated-orcid":false,"given":"Yuxin","family":"Huang","sequence":"first","affiliation":[{"name":"Faculty of Information Engineering and Automation, Yunnan Key Laboratory ofArtificial Intelligence, Kunming University of Science and Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-4914-7347","authenticated-orcid":false,"given":"Yin","family":"Liang","sequence":"additional","affiliation":[{"name":"Faculty of Information Engineering and Automation, Yunnan Key Laboratory ofArtificial Intelligence, Kunming University of Science and Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-0528-5937","authenticated-orcid":false,"given":"Zhaoyuan","family":"Wu","sequence":"additional","affiliation":[{"name":"Faculty of Information Engineering and Automation, Yunnan Key Laboratory ofArtificial Intelligence, Kunming University of Science and Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1340-4368","authenticated-orcid":false,"given":"Enchang","family":"Zhu","sequence":"additional","affiliation":[{"name":"Faculty of Information Engineering and Automation, Yunnan Key Laboratory ofArtificial Intelligence, Kunming University of Science and Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4012-461X","authenticated-orcid":false,"given":"Zhengtao","family":"Yu","sequence":"additional","affiliation":[{"name":"Faculty of Information Engineering and Automation, Yunnan Key Laboratory ofArtificial Intelligence, Kunming University of Science and Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,6,16]]},"reference":[{"key":"e_1_3_1_2_2","article-title":"Distributed representations of sentences and documents","volume":"1405","author":"Le Quoc V.","year":"2014","unstructured":"Quoc V. Le and Tom\u00e1s Mikolov. 2014. Distributed representations of sentences and documents. CoRR abs\/1405.4053 (2014).","journal-title":"CoRR"},{"key":"e_1_3_1_3_2","article-title":"Unsupervised learning of sentence embeddings using compositional n-gram features","volume":"1703","author":"Pagliardini Matteo","year":"2017","unstructured":"Matteo Pagliardini, Prakhar Gupta, and Martin Jaggi. 2017. Unsupervised learning of sentence embeddings using compositional n-gram features. CoRR abs\/1703.02507 (2017).","journal-title":"CoRR"},{"key":"e_1_3_1_4_2","first-page":"616","volume-title":"Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing","author":"El-Kishky Ahmed","year":"2020","unstructured":"Ahmed El-Kishky and Francisco Guzm\u00e1n. 2020. Massively multilingual document alignment with cross-lingual sentence-mover\u2019s distance. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing. Association for Computational Linguistics, 616\u2013625. Retrieved from https:\/\/aclanthology.org\/2020.aacl-main.62\/."},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.538"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1410"},{"key":"e_1_3_1_7_2","volume-title":"Proceedings of the 7th International Conference on Learning Representations","author":"Gao Jun","year":"2019","unstructured":"Jun Gao, Di He, Xu Tan, Tao Qin, Liwei Wang, and Tie-Yan Liu. 2019. Representation degeneration problem in training natural language generation models. In Proceedings of the 7th International Conference on Learning Representations. OpenReview.net. Retrieved from https:\/\/openreview.net\/forum?id=SkEYojRqtm."},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.findings-emnlp.150"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.eacl-main.215"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/p19-1493"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00288"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/n19-1392"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/n19-1423"},{"key":"e_1_3_1_14_2","article-title":"Cross-lingual language model pretraining","volume":"1901","author":"Lample Guillaume","year":"2019","unstructured":"Guillaume Lample and Alexis Conneau. 2019. Cross-lingual language model pretraining. CoRR abs\/1901.07291 (2019).","journal-title":"CoRR"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.747"},{"key":"e_1_3_1_16_2","article-title":"Bridging subword gaps in pretrain-finetune paradigm for natural language generation","volume":"2106","author":"Liu Xin","year":"2021","unstructured":"Xin Liu, Baosong Yang, Dayiheng Liu, Haibo Zhang, Weihua Luo, Min Zhang, Haiying Zhang, and Jinsong Su. 2021. Bridging subword gaps in pretrain-finetune paradigm for natural language generation. CoRR abs\/2106.06125 (2021).","journal-title":"CoRR"},{"key":"e_1_3_1_17_2","article-title":"A simple framework for contrastive learning of visual representations","volume":"2002","author":"Chen Ting","year":"2020","unstructured":"Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey E. Hinton. 2020. A simple framework for contrastive learning of visual representations. CoRR abs\/2002.05709 (2020).","journal-title":"CoRR"},{"key":"e_1_3_1_18_2","article-title":"Supervised contrastive learning","volume":"2004","author":"Khosla Prannay","year":"2020","unstructured":"Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan. 2020. Supervised contrastive learning. CoRR abs\/2004.11362 (2020).","journal-title":"CoRR"},{"key":"e_1_3_1_19_2","article-title":"SimCSE: Simple contrastive learning of sentence embeddings","volume":"2104","author":"Gao Tianyu","year":"2021","unstructured":"Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021. SimCSE: Simple contrastive learning of sentence embeddings. CoRR abs\/2104.08821 (2021).","journal-title":"CoRR"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.naacl-main.311"},{"key":"e_1_3_1_21_2","article-title":"Towards user-driven neural machine translation","volume":"2106","author":"Lin Huan","year":"2021","unstructured":"Huan Lin, Liang Yao, Baosong Yang, Dayiheng Liu, Haibo Zhang, Weihua Luo, Degen Huang, and Jinsong Su. 2021. Towards user-driven neural machine translation. CoRR abs\/2106.06200 (2021).","journal-title":"CoRR"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1142\/S0218001493000339"},{"key":"e_1_3_1_23_2","article-title":"Chinese sentences similarity via cross-attention based siamese network","volume":"2104","author":"Wang Zhen","year":"2021","unstructured":"Zhen Wang, Xiangxie Zhang, and Yicong Tan. 2021. Chinese sentences similarity via cross-attention based siamese network. CoRR abs\/2104.08787 (2021).","journal-title":"CoRR"},{"key":"e_1_3_1_24_2","article-title":"Concatenated p-mean word embeddings as universal cross-lingual sentence representations","volume":"1803","author":"R\u00fcckl\u00e9 Andreas","year":"2018","unstructured":"Andreas R\u00fcckl\u00e9, Steffen Eger, Maxime Peyrard, and Iryna Gurevych. 2018. Concatenated p-mean word embeddings as universal cross-lingual sentence representations. CoRR abs\/1803.01400 (2018).","journal-title":"CoRR"},{"key":"e_1_3_1_25_2","first-page":"5998","volume-title":"Proceedings of the Annual Conference on Neural Information Processing Systems","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the Annual Conference on Neural Information Processing Systems. 5998\u20136008. Retrieved from https:\/\/proceedings.neurips.cc\/paper\/2017\/hash\/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html."},{"key":"e_1_3_1_26_2","article-title":"Improving neural networks by preventing co-adaptation of feature detectors","volume":"1207","author":"Hinton Geoffrey E.","year":"2012","unstructured":"Geoffrey E. Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2012. Improving neural networks by preventing co-adaptation of feature detectors. CoRR abs\/1207.0580 (2012).","journal-title":"CoRR"},{"key":"e_1_3_1_27_2","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"1910","author":"Raffel Colin","year":"2019","unstructured":"Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019. Exploring the limits of transfer learning with a unified text-to-text transformer. CoRR abs\/1910.10683 (2019).","journal-title":"CoRR"},{"key":"e_1_3_1_28_2","article-title":"mT5: A massively multilingual pre-trained text-to-text transformer","volume":"2010","author":"Xue Linting","year":"2020","unstructured":"Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2020. mT5: A massively multilingual pre-trained text-to-text transformer. CoRR abs\/2010.11934 (2020).","journal-title":"CoRR"},{"key":"e_1_3_1_29_2","article-title":"Unsupervised cross-lingual representation learning at scale","volume":"1911","author":"Conneau Alexis","year":"2019","unstructured":"Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzm\u00e1n, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Unsupervised cross-lingual representation learning at scale. CoRR abs\/1911.02116 (2019).","journal-title":"CoRR"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1356"}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3589341","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3589341","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:03:16Z","timestamp":1750291396000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3589341"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,6,16]]},"references-count":29,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2023,6,30]]}},"alternative-id":["10.1145\/3589341"],"URL":"https:\/\/doi.org\/10.1145\/3589341","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"value":"2375-4699","type":"print"},{"value":"2375-4702","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,6,16]]},"assertion":[{"value":"2022-09-22","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-03-13","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-06-16","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}