{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:14:51Z","timestamp":1750220091315,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":25,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,8,14]],"date-time":"2022-08-14T00:00:00Z","timestamp":1660435200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,8,14]]},"DOI":"10.1145\/3534678.3539091","type":"proceedings-article","created":{"date-parts":[[2022,8,12]],"date-time":"2022-08-12T19:06:41Z","timestamp":1660331201000},"page":"3187-3196","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Self-Supervised Augmentation and Generation for Multi-lingual Text Advertisements at Bing"],"prefix":"10.1145","author":[{"given":"Xiaoyu","family":"Kou","sequence":"first","affiliation":[{"name":"Microsoft Corporation, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tianqi","family":"Zhao","sequence":"additional","affiliation":[{"name":"Microsoft Corporation, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fan","family":"Zhang","sequence":"additional","affiliation":[{"name":"Microsoft Corporation, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Song","family":"Li","sequence":"additional","affiliation":[{"name":"Microsoft Corporation, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qi","family":"Zhang","sequence":"additional","affiliation":[{"name":"Microsoft Corporation, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,8,14]]},"reference":[{"key":"e_1_3_2_2_1_1","doi-asserted-by":"crossref","unstructured":"Ateret Anaby-Tavor Boaz Carmeli Esther Goldbraich Amir Kantor George Kour Segev Shlomov Naama Tepper and Naama Zwerdling. 2020. Do not have enough data? Deep learning to the rescue!. In AAAI. 7383--7390.","DOI":"10.1609\/aaai.v34i05.6233"},{"key":"e_1_3_2_2_2_1","volume-title":"Translation artifacts in cross-lingual transfer learning. arXiv preprint arXiv:2004.04721","author":"Artetxe Mikel","year":"2020","unstructured":"Mikel Artetxe, Gorka Labaka, and Eneko Agirre. 2020. Translation artifacts in cross-lingual transfer learning. arXiv preprint arXiv:2004.04721 (2020)."},{"key":"e_1_3_2_2_3_1","volume-title":"Unsupervised cross-lingual representation learning at scale. arXiv preprint arXiv:1911.02116","author":"Conneau Alexis","year":"2019","unstructured":"Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzm\u00e1n, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Unsupervised cross-lingual representation learning at scale. arXiv preprint arXiv:1911.02116 (2019)."},{"key":"e_1_3_2_2_4_1","volume-title":"Generative adversarial networks: An overview","author":"Creswell Antonia","year":"2018","unstructured":"Antonia Creswell, Tom White, Vincent Dumoulin, Kai Arulkumaran, Biswa Sengupta, and Anil A Bharath. 2018. Generative adversarial networks: An overview. IEEE Signal Processing Magazine (2018), 53--65."},{"key":"e_1_3_2_2_5_1","volume-title":"Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018)."},{"key":"e_1_3_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/2389376.2389401"},{"key":"e_1_3_2_2_7_1","unstructured":"Yingmei Guo Linjun Shou Jian Pei Ming Gong Mingxing Xu Zhiyong Wu and Daxin Jiang. 2021. Learning from multiple noisy augmented data sets for better cross-lingual spoken language understanding. In EMNLP ."},{"key":"e_1_3_2_2_8_1","volume-title":"Long short-term memory. Neural computation","author":"Hochreiter Sepp","year":"1997","unstructured":"Sepp Hochreiter and J\u00fcrgen Schmidhuber. 1997. Long short-term memory. Neural computation, Vol. 9, 8 (1997), 1735--1780."},{"key":"e_1_3_2_2_9_1","volume-title":"Unicoder: A universal language encoder by pre-training with multiple cross-lingual tasks. arXiv preprint arXiv:1909.00964","author":"Huang Haoyang","year":"2019","unstructured":"Haoyang Huang, Yaobo Liang, Nan Duan, Ming Gong, Linjun Shou, Daxin Jiang, and Ming Zhou. 2019. Unicoder: A universal language encoder by pre-training with multiple cross-lingual tasks. arXiv preprint arXiv:1909.00964 (2019)."},{"key":"e_1_3_2_2_10_1","doi-asserted-by":"crossref","unstructured":"J Weston Hughes Keng-hao Chang and Ruofei Zhang. 2019. Generating better search engine text advertisements with deep reinforcement learning. In SIGKDD .","DOI":"10.1145\/3292500.3330754"},{"key":"e_1_3_2_2_11_1","volume-title":"Data Augmentation Approaches in Natural Language Processing: A Survey. arXiv preprint arXiv:2110.01852","author":"Li Bohan","year":"2021","unstructured":"Bohan Li, Yutai Hou, and Wanxiang Che. 2021. Data Augmentation Approaches in Natural Language Processing: A Survey. arXiv preprint arXiv:2110.01852 (2021)."},{"key":"e_1_3_2_2_12_1","volume-title":"Unsupervised Cross-lingual Adaptation for Sequence Tagging and Beyond. arXiv preprint arXiv:2010.12405","author":"Li Xin","year":"2020","unstructured":"Xin Li, Lidong Bing, Wenxuan Zhang, Zheng Li, and Wai Lam. 2020. Unsupervised Cross-lingual Adaptation for Sequence Tagging and Beyond. arXiv preprint arXiv:2010.12405 (2020)."},{"key":"e_1_3_2_2_13_1","volume-title":"Reinforced Iterative Knowledge Distillation for Cross-Lingual Named Entity Recognition. arXiv preprint arXiv:2106.00241","author":"Liang Shining","year":"2021","unstructured":"Shining Liang, Ming Gong, Jian Pei, Linjun Shou, Wanli Zuo, Xianglin Zuo, and Daxin Jiang. 2021. Reinforced Iterative Knowledge Distillation for Cross-Lingual Named Entity Recognition. arXiv preprint arXiv:2106.00241 (2021)."},{"key":"e_1_3_2_2_14_1","volume-title":"Xglue: A new benchmark dataset for cross-lingual pre-training, understanding and generation. arXiv preprint arXiv:2004.01401","author":"Liang Yaobo","year":"2020","unstructured":"Yaobo Liang, Nan Duan, Yeyun Gong, Ning Wu, Fenfei Guo, Weizhen Qi, Ming Gong, Linjun Shou, Daxin Jiang, Guihong Cao, et al. 2020. Xglue: A new benchmark dataset for cross-lingual pre-training, understanding and generation. arXiv preprint arXiv:2004.01401 (2020)."},{"key":"e_1_3_2_2_15_1","volume-title":"Multilingual denoising pre-training for neural machine translation. ACL","author":"Liu Yinhan","year":"2020","unstructured":"Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020. Multilingual denoising pre-training for neural machine translation. ACL (2020), 726--742."},{"key":"e_1_3_2_2_16_1","volume-title":"Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692","author":"Liu Yinhan","year":"2019","unstructured":"Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692 (2019)."},{"key":"e_1_3_2_2_17_1","volume-title":"Aditya Srikanth Veerubhotla, and Manish Gupta","author":"Mitra Rajarshee","year":"2021","unstructured":"Rajarshee Mitra, Rhea Jain, Aditya Srikanth Veerubhotla, and Manish Gupta. 2021. Zero-shot Multi-lingual Interrogative Question Generation for\" People Also Ask\" at Bing. In SIGKDD. 3414--3422."},{"key":"e_1_3_2_2_18_1","volume-title":"Data augmentation for spoken language understanding via pretrained models. arXiv e-prints","author":"Peng Baolin","year":"2020","unstructured":"Baolin Peng, Chenguang Zhu, Michael Zeng, and Jianfeng Gao. 2020. Data augmentation for spoken language understanding via pretrained models. arXiv e-prints (2020), arXiv--2004."},{"key":"e_1_3_2_2_19_1","volume-title":"Prophetnet: Predicting future n-gram for sequence-to-sequence pre-training. In EMNLP: Findings. 2401--2410.","author":"Qi Weizhen","year":"2020","unstructured":"Weizhen Qi, Yu Yan, Yeyun Gong, Dayiheng Liu, Nan Duan, Jiusheng Chen, Ruofei Zhang, and Ming Zhou. 2020. Prophetnet: Predicting future n-gram for sequence-to-sequence pre-training. In EMNLP: Findings. 2401--2410."},{"key":"e_1_3_2_2_20_1","volume-title":"Control, generate, augment: A scalable framework for multi-attribute text generation. arXiv preprint arXiv:2004.14983","author":"Russo Giuseppe","year":"2020","unstructured":"Giuseppe Russo, Nora Hollenstein, Claudiu Musat, and Ce Zhang. 2020. Control, generate, augment: A scalable framework for multi-attribute text generation. arXiv preprint arXiv:2004.14983 (2020)."},{"key":"e_1_3_2_2_21_1","volume-title":"Mihir Sanjay Kale, and Linting Xue","author":"Shakeri Siamak","year":"2020","unstructured":"Siamak Shakeri, Noah Constant, Mihir Sanjay Kale, and Linting Xue. 2020. Multilingual Synthetic Question and Answer Generation for Cross-Lingual Reading Comprehension. arXiv e-prints (2020), arXiv--2010."},{"key":"e_1_3_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/2505515.2507876"},{"key":"e_1_3_2_2_23_1","doi-asserted-by":"crossref","unstructured":"Shyam Upadhyay Manaal Faruqui Gokhan T\u00fcr Hakkani-T\u00fcr Dilek and Larry Heck. 2018. (Almost) zero-shot cross-lingual spoken language understanding. In ICASSP. IEEE 6034--6038.","DOI":"10.1109\/ICASSP.2018.8461905"},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"crossref","unstructured":"Xiting Wang Xinwei Gu Jie Cao Zihua Zhao Yulan Yan Bhuvan Middha and Xing Xie. 2021. Reinforcing Pretrained Models for Generating Attractive Text Advertisements. In ACM SIGKDD. 3697--3707.","DOI":"10.1145\/3447548.3467105"},{"key":"e_1_3_2_2_25_1","volume-title":"mt5: A massively multilingual pre-trained text-to-text transformer. arXiv preprint arXiv:2010.11934","author":"Xue Linting","year":"2020","unstructured":"Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2020. mt5: A massively multilingual pre-trained text-to-text transformer. arXiv preprint arXiv:2010.11934 (2020)."}],"event":{"name":"KDD '22: The 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining","sponsor":["SIGMOD ACM Special Interest Group on Management of Data","SIGKDD ACM Special Interest Group on Knowledge Discovery in Data"],"location":"Washington DC USA","acronym":"KDD '22"},"container-title":["Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3534678.3539091","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3534678.3539091","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T18:09:51Z","timestamp":1750183791000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3534678.3539091"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,8,14]]},"references-count":25,"alternative-id":["10.1145\/3534678.3539091","10.1145\/3534678"],"URL":"https:\/\/doi.org\/10.1145\/3534678.3539091","relation":{},"subject":[],"published":{"date-parts":[[2022,8,14]]},"assertion":[{"value":"2022-08-14","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}