{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T11:29:50Z","timestamp":1782818990640,"version":"3.54.5"},"reference-count":65,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2023,6,13]],"date-time":"2023-06-13T00:00:00Z","timestamp":1686614400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2023,6,13]]},"abstract":"<jats:p>Embedding-based deep recommendation models (EDRMs), which contain small dense models and large embedding tables, are widely used in industry. Embedding communication constitutes the main cost for the distributed training of EDRMs, and thus we propose two strategies to improve its efficiency, i.e.,embedding tiering andpre-fetching. In particular, embedding tiering uses AllReduce to communicate popular embeddings that are accessed frequently. This is counter-intuitive as embeddings belong to the sparse embedding tables, but reasonable because the access pattern of popular embeddings resembles dense models. Pre-fetching starts communication early for embeddings that receive no updates such that they are removed from the critical path of training. We implement embedding tiering and pre-fetching in a system called FEC and compare it with the state-of-the-art systems on real datasets. The results show that FEC consistently outperforms the existing methods on all datasets, and its speed can be up to 6.65x and 2.42x in terms of embedding communication time and training throughput compared with the best performing baseline.<\/jats:p>","DOI":"10.1145\/3589310","type":"journal-article","created":{"date-parts":[[2023,6,20]],"date-time":"2023-06-20T20:26:45Z","timestamp":1687292805000},"page":"1-21","source":"Crossref","is-referenced-by-count":6,"title":["FEC: Efficient Deep Recommendation Model Training with Flexible Embedding Communication"],"prefix":"10.1145","volume":"1","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-3176-7816","authenticated-orcid":false,"given":"Kaihao","family":"Ma","sequence":"first","affiliation":[{"name":"The Chinese University of Hong Kong, Hong Kong, Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2122-915X","authenticated-orcid":false,"given":"Xiao","family":"Yan","sequence":"additional","affiliation":[{"name":"The Chinese University of Hong Kong, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0199-4866","authenticated-orcid":false,"given":"Zhenkun","family":"Cai","sequence":"additional","affiliation":[{"name":"The Chinese University of Hong Kong, Hong Kong, Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-6165-5559","authenticated-orcid":false,"given":"Yuzhen","family":"Huang","sequence":"additional","affiliation":[{"name":"Meta, Menlo Park, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3996-9316","authenticated-orcid":false,"given":"Yidi","family":"Wu","sequence":"additional","affiliation":[{"name":"Meta, Menlo Park, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6313-6288","authenticated-orcid":false,"given":"James","family":"Cheng","sequence":"additional","affiliation":[{"name":"The Chinese University of Hong Kong &amp; KASMA PTE, LTD., Hong Kong, Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,6,20]]},"reference":[{"key":"e_1_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.14778\/3485450.3485462"},{"key":"e_1_2_2_2_1","volume-title":"Proceedings of the 31st International Conference on Neural Information Processing Systems. 1707--1718","author":"Alistarh Dan","year":"2017","unstructured":"Dan Alistarh, Demjan Grubic, Jerry Z Li, Ryota Tomioka, and Milan Vojnovic. 2017. QSGD: communication-efficient SGD via gradient quantization and encoding. In Proceedings of the 31st International Conference on Neural Information Processing Systems. 1707--1718."},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2021.3132413"},{"key":"e_1_2_2_4_1","volume-title":"Revisiting Distributed Synchronous SGD. CoRR","author":"Chen Jianmin","year":"2016","unstructured":"Jianmin Chen, Rajat Monga, Samy Bengio, and Rafal J\u00f3 zefowicz. 2016. Revisiting Distributed Synchronous SGD. CoRR, Vol. abs\/1604.00981 (2016)."},{"key":"e_1_2_2_5_1","volume-title":"FLEN: Leveraging Field for Scalable CTR Prediction. CoRR","author":"Chen Wenqiang","year":"2019","unstructured":"Wenqiang Chen, Lizhang Zhan, Yuanlong Ci, and Chen Lin. 2019. FLEN: Leveraging Field for Scalable CTR Prediction. CoRR, Vol. abs\/1911.04690 (2019)."},{"key":"e_1_2_2_6_1","doi-asserted-by":"crossref","unstructured":"Heng-Tze Cheng Levent Koc Jeremiah Harmsen Tal Shaked Tushar Chandra Hrishi Aradhye Glen Anderson Greg Corrado Wei Chai Mustafa Ispir Rohan Anil Zakaria Haque Lichan Hong Vihan Jain Xiaobing Liu and Hemal Shah. 2016. Wide & Deep Learning for Recommender Systems. In DLRS@RecSys. ACM 7--10.","DOI":"10.1145\/2988450.2988454"},{"key":"e_1_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i04.5768"},{"key":"e_1_2_2_8_1","doi-asserted-by":"crossref","unstructured":"Paul Covington Jay Adams and Emre Sargin. 2016. Deep Neural Networks for YouTube Recommendations. In RecSys. ACM 191--198.","DOI":"10.1145\/2959100.2959190"},{"key":"e_1_2_2_9_1","unstructured":"CriteoLabs. 2014. Criteo display ad challenge. https:\/\/www.kaggle.com\/c\/criteo-display-ad-challenge"},{"key":"e_1_2_2_10_1","unstructured":"CriteoLabs. 2022. Criteo 1TB Click Logs dataset. https:\/\/ailab.criteo.com\/criteo-1tb-click-logs-dataset-for-mlperf\/"},{"key":"e_1_2_2_11_1","volume-title":"Proceedings of the 25th International Conference on Neural Information Processing Systems-Volume 1. 1223--1231","author":"Dean Jeffrey","year":"2012","unstructured":"Jeffrey Dean, Greg S Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Quoc V Le, Mark Z Mao, Marc'Aurelio Ranzato, Andrew Senior, Paul Tucker, et al. 2012. Large scale distributed deep networks. In Proceedings of the 25th International Conference on Neural Information Processing Systems-Volume 1. 1223--1231."},{"key":"e_1_2_2_12_1","volume-title":"P3: Distributed Deep Graph Learning at Scale","author":"Gandhi Swapnil","unstructured":"Swapnil Gandhi and Anand Padmanabha Iyer. 2021. P3: Distributed Deep Graph Learning at Scale. In OSDI. USENIX Association, 551--568."},{"key":"e_1_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10107-014-0846-1"},{"key":"e_1_2_2_14_1","doi-asserted-by":"crossref","unstructured":"Huifeng Guo Ruiming Tang Yunming Ye Zhenguo Li and Xiuqiang He. 2017. DeepFM: A Factorization-Machine based Neural Network for CTR Prediction. In IJCAI. ijcai.org 1725--1731.","DOI":"10.24963\/ijcai.2017\/239"},{"key":"e_1_2_2_15_1","doi-asserted-by":"crossref","unstructured":"Suyog Gupta Wei Zhang and Fei Wang. 2017. Model Accuracy and Runtime Tradeoff in Distributed Deep Learning: A Systematic Study. In IJCAI. ijcai.org 4854--4858.","DOI":"10.24963\/ijcai.2017\/681"},{"key":"e_1_2_2_16_1","volume-title":"The Architectural Implications of Facebook's DNN-Based Personalized Recommendation","author":"Gupta Udit","unstructured":"Udit Gupta, Carole-Jean Wu, Xiaodong Wang, Maxim Naumov, Brandon Reagen, David Brooks, Bradford Cottel, Kim M. Hazelwood, Mark Hempstead, Bill Jia, Hsien-Hsin S. Lee, Andrey Malevich, Dheevatsa Mudigere, Mikhail Smelyanskiy, Liang Xiong, and Xuan Zhang. 2020. The Architectural Implications of Facebook's DNN-Based Personalized Recommendation. In HPCA. IEEE, 488--501."},{"key":"e_1_2_2_17_1","volume-title":"Xiaohan Wei, Xing Wang, Yuzhen Huang, Arun Kejariwal, Kannan Ramchandran, and Michael W. Mahoney.","author":"Gupta Vipul","year":"2021","unstructured":"Vipul Gupta, Dhruv Choudhary, Ping Tak Peter Tang, Xiaohan Wei, Xing Wang, Yuzhen Huang, Arun Kejariwal, Kannan Ramchandran, and Michael W. Mahoney. 2021. Training Recommender Systems at Scale: Communication-Efficient Model and Data Parallelism. In KDD. ACM, 2928--2936."},{"key":"e_1_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3077136.3080777"},{"key":"e_1_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.14778\/3430915.3430932"},{"key":"e_1_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.14778\/3357377.3357381"},{"key":"e_1_2_2_21_1","volume-title":"Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019","author":"Huang Yanping","year":"2019","unstructured":"Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Xu Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, Yonghui Wu, and Zhifeng Chen. 2019. GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8--14, 2019, Vancouver, BC, Canada, Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d'Alch\u00e9 -Buc, Emily B. Fox, and Roman Garnett (Eds.). 103--112. https:\/\/proceedings.neurips.cc\/paper\/2019\/hash\/093f65e080a295f8076b1c5722a46aa2-Abstract.html"},{"key":"e_1_2_2_22_1","volume-title":"Hierarchical Training: Scaling Deep Recommendation Models on Large CPU Clusters. In KDD. ACM, 3050--3058.","author":"Huang Yuzhen","year":"2021","unstructured":"Yuzhen Huang, Xiaohan Wei, Xing Wang, Jiyan Yang, Bor-Yiing Su, Shivam Bharuka, Dhruv Choudhary, Zewei Jiang, Hai Zheng, and Jack Langman. 2021. Hierarchical Training: Scaling Deep Recommendation Models on Large CPU Clusters. In KDD. ACM, 3050--3058."},{"key":"e_1_2_2_23_1","volume-title":"Colin Taylor, Xing Liu, Will Feng, Rahul Kindi, Anirudh Sudarshan, and Shahin Sefati.","author":"Ivchenko Dmytro","year":"2022","unstructured":"Dmytro Ivchenko, Dennis Van Der Staay, Colin Taylor, Xing Liu, Will Feng, Rahul Kindi, Anirudh Sudarshan, and Shahin Sefati. 2022. TorchRec: a PyTorch Domain Library for Recommendation Systems. In RecSys. ACM, 482--483."},{"key":"e_1_2_2_24_1","volume-title":"Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsm\"a ssan","author":"Jia Zhihao","year":"2018","unstructured":"Zhihao Jia, Sina Lin, Charles R. Qi, and Alex Aiken. 2018. Exploring Hidden Dimensions in Parallelizing Convolutional Neural Networks. In Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsm\"a ssan, Stockholm, Sweden, July 10--15, 2018 (Proceedings of Machine Learning Research, Vol. 80), Jennifer G. Dy and Andreas Krause (Eds.). PMLR, 2279--2288. http:\/\/proceedings.mlr.press\/v80\/jia18a.html"},{"key":"e_1_2_2_25_1","unstructured":"Zhihao Jia James Thomas Todd Warszawski Mingyu Gao Matei Zaharia and Alex Aiken. 2019a. Optimizing DNN Computation with Relaxed Graph Substitutions. In MLSys. mlsys.org."},{"key":"e_1_2_2_26_1","volume-title":"Proceedings of Machine Learning and Systems 2019","author":"Jia Zhihao","year":"2019","unstructured":"Zhihao Jia, Matei Zaharia, and Alex Aiken. 2019b. Beyond Data and Model Parallelism for Deep Neural Networks. In Proceedings of Machine Learning and Systems 2019, MLSys 2019, Stanford, CA, USA, March 31 - April 2, 2019, Ameet Talwalkar, Virginia Smith, and Matei Zaharia (Eds.). mlsys.org. https:\/\/proceedings.mlsys.org\/book\/265.pdf"},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3326937.3341255"},{"key":"e_1_2_2_28_1","volume-title":"Exploring the Limits of Language Modeling. CoRR","author":"Rafal J\u00f3","year":"2016","unstructured":"Rafal J\u00f3 zefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu. 2016. Exploring the Limits of Language Modeling. CoRR, Vol. abs\/1602.02410 (2016)."},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/3035918.3064023"},{"key":"e_1_2_2_30_1","first-page":"1","article-title":"Parallax","volume":"43","author":"Kim Soojeong","year":"2019","unstructured":"Soojeong Kim, Gyeong-In Yu, Hojin Park, Sungwoo Cho, Eunji Jeong, Hyeonmin Ha, Sanha Lee, Joo Seong Jeong, and Byung-Gon Chun. 2019. Parallax: Sparsity-aware Data Parallel Training of Deep Neural Networks. In EuroSys. ACM, 43:1--43:15.","journal-title":"Sparsity-aware Data Parallel Training of Deep Neural Networks. In EuroSys. ACM"},{"key":"e_1_2_2_31_1","volume-title":"Alexander J. Smola, Amr Ahmed, Vanja Josifovski, James Long, Eugene J. Shekita, and Bor-Yiing Su.","author":"Li Mu","year":"2014","unstructured":"Mu Li, David G. Andersen, Jun Woo Park, Alexander J. Smola, Amr Ahmed, Vanja Josifovski, James Long, Eugene J. Shekita, and Bor-Yiing Su. 2014. Scaling Distributed Machine Learning with the Parameter Server. In OSDI. USENIX Association, 583--598."},{"key":"e_1_2_2_32_1","volume-title":"TaP: Table-based Prefetching for Storage Caches. In 6th USENIX Conference on File and Storage Technologies, FAST 2008","author":"Li Mingju","year":"2008","unstructured":"Mingju Li, Elizabeth Varki, Swapnil Bhatia, and Arif Merchant. 2008. TaP: Table-based Prefetching for Storage Caches. In 6th USENIX Conference on File and Storage Technologies, FAST 2008, February 26--29, 2008, San Jose, CA, USA, Mary Baker and Erik Riedel (Eds.). USENIX, 81--96. http:\/\/www.usenix.org\/events\/fast08\/tech\/li.html"},{"key":"e_1_2_2_33_1","doi-asserted-by":"crossref","unstructured":"Xiang Li Chao Wang Jiwei Tan Xiaoyi Zeng Dan Ou and Bo Zheng. 2020. Adversarial Multimodal Representation Learning for Click-Through Rate Prediction. In WWW. ACM \/ IW3C2 827--836.","DOI":"10.1145\/3366423.3380163"},{"key":"e_1_2_2_34_1","doi-asserted-by":"crossref","unstructured":"Jianxun Lian Xiaohuan Zhou Fuzheng Zhang Zhongxia Chen Xing Xie and Guangzhong Sun. 2018b. xDeepFM: Combining Explicit and Implicit Feature Interactions for Recommender Systems. In KDD. ACM 1754--1763.","DOI":"10.1145\/3219819.3220023"},{"key":"e_1_2_2_35_1","volume-title":"Proceedings of the 28th International Conference on Neural Information Processing Systems-Volume 2. 2737--2745","author":"Lian Xiangru","year":"2015","unstructured":"Xiangru Lian, Yijun Huang, Yuncheng Li, and Ji Liu. 2015. Asynchronous parallel stochastic gradient for nonconvex optimization. In Proceedings of the 28th International Conference on Neural Information Processing Systems-Volume 2. 2737--2745."},{"key":"e_1_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/3534678.3539070"},{"key":"e_1_2_2_37_1","volume-title":"ICML (Proceedings of Machine Learning Research","volume":"3058","author":"Lian Xiangru","year":"2018","unstructured":"Xiangru Lian, Wei Zhang, Ce Zhang, and Ji Liu. 2018a. Asynchronous Decentralized Parallel Stochastic Gradient Descent. In ICML (Proceedings of Machine Learning Research, Vol. 80). PMLR, 3049--3058."},{"key":"e_1_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/564691.564763"},{"key":"e_1_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/CLUSTR.2004.1392611"},{"key":"e_1_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.14778\/3489496.3489511"},{"key":"e_1_2_2_41_1","volume-title":"ICLR Workshop.","author":"Mikolov Tomas","year":"2013","unstructured":"Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space. In ICLR Workshop."},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/3534678.3539038"},{"key":"e_1_2_2_43_1","doi-asserted-by":"crossref","unstructured":"Dheevatsa Mudigere Yuchen Hao Jianyu Huang Zhihao Jia Andrew Tulloch Srinivas Sridharan Xing Liu Mustafa Ozdal Jade Nie Jongsoo Park Liang Luo Jie Amy Yang Leon Gao Dmytro Ivchenko Aarti Basant Yuxi Hu Jiyan Yang Ehsan K. Ardestani Xiaodong Wang Rakesh Komuravelli Ching-Hsiang Chu Serhat Yilmaz Huayu Li Jiyuan Qian Zhuobo Feng Yinbin Ma Junjie Yang Ellie Wen Hong Li Lin Yang Chonglin Sun Whitney Zhao Dimitry Melts Krishna Dhulipala K. R. Kishore Tyler Graf Assaf Eisenman Kiran Kumar Matam Adi Gangidi Guoqiang Jerry Chen Manoj Krishnan Avinash Nayak Krishnakumar Nair Bharath Muthiah Mahmoud khorashadi Pallab Bhattacharya Petr Lapukhov Maxim Naumov Ajit Mathews Lin Qiao Mikhail Smelyanskiy Bill Jia and Vijay Rao. 2022. Software-hardware co-design for fast and scalable training of deep learning recommendation models. In ISCA. ACM 993--1011.","DOI":"10.1145\/3470496.3533727"},{"key":"e_1_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3341301.3359646"},{"key":"e_1_2_2_45_1","unstructured":"Maxim Naumov Dheevatsa Mudigere Hao-Jun Michael Shi Jianyu Huang Narayanan Sundaraman Jongsoo Park Xiaodong Wang Udit Gupta Carole-Jean Wu Alisson G. Azzolini Dmytro Dzhulgakov Andrey Mallevich Ilia Cherniavskii Yinghai Lu Raghuraman Krishnamoorthi Ansha Yu Volodymyr Kondratenko Stephanie Pereira Xianjie Chen Wenlin Chen Vijay Rao Bill Jia Liang Xiong and Misha Smelyanskiy. 2019. Deep Learning Recommendation Model for Personalization and Recommendation Systems. CoRR Vol. abs\/1906.00091 (2019). https:\/\/arxiv.org\/abs\/1906.00091"},{"key":"e_1_2_2_46_1","doi-asserted-by":"crossref","unstructured":"Yabo Ni Dan Ou Shichen Liu Xiang Li Wenwu Ou Anxiang Zeng and Luo Si. 2018. Perceive Your Users in Depth: Learning Universal User Representations from Multiple E-commerce Tasks. In KDD. ACM 596--605.","DOI":"10.1145\/3219819.3219828"},{"key":"e_1_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2007.370405"},{"key":"e_1_2_2_48_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2008.09.002"},{"key":"e_1_2_2_49_1","volume-title":"GPU Technology Conference, NVIDIA.","author":"Schroeder Tim C","year":"2011","unstructured":"Tim C Schroeder. 2011. Peer-to-peer & unified virtual addressing. In GPU Technology Conference, NVIDIA."},{"key":"e_1_2_2_50_1","doi-asserted-by":"crossref","unstructured":"Geet Sethi Bilge Acun Niket Agarwal Christos Kozyrakis Caroline Trippel and Carole-Jean Wu. 2022. RecShard: statistical feature-based memory optimization for industry-scale neural recommendation. In ASPLOS. ACM 344--358.","DOI":"10.1145\/3503222.3507777"},{"key":"e_1_2_2_51_1","unstructured":"Karen Simonyan and Andrew Zisserman. 2015. Very Deep Convolutional Networks for Large-Scale Image Recognition. In ICLR."},{"key":"e_1_2_2_52_1","volume-title":"Rethinking the Inception Architecture for Computer Vision","author":"Szegedy Christian","unstructured":"Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna. 2016. Rethinking the Inception Architecture for Computer Vision. In CVPR. IEEE Computer Society, 2818--2826."},{"key":"e_1_2_2_53_1","doi-asserted-by":"publisher","unstructured":"Muhammad Tahir Rabia Enam and Syed Muhammad Nabeel Mustafa. 2021. E-commerce platform based on Machine Learning Recommendation System. https:\/\/doi.org\/10.1109\/IMTIC53841.2021.9719822","DOI":"10.1109\/IMTIC53841.2021.9719822"},{"key":"e_1_2_2_54_1","volume-title":"Communication-Efficient Distributed Deep Learning: A Comprehensive Survey. CoRR","author":"Tang Zhenheng","year":"2020","unstructured":"Zhenheng Tang, Shaohuai Shi, Xiaowen Chu, Wei Wang, and Bo Li. 2020. Communication-Efficient Distributed Deep Learning: A Comprehensive Survey. CoRR, Vol. abs\/2003.06307 (2020). showeprint[arXiv]2003.06307 https:\/\/arxiv.org\/abs\/2003.06307"},{"key":"e_1_2_2_55_1","first-page":"1","article-title":"Deep & Cross Network for Ad Click Predictions. In ADKDD@KDD","volume":"12","author":"Wang Ruoxi","year":"2017","unstructured":"Ruoxi Wang, Bin Fu, Gang Fu, and Mingliang Wang. 2017. Deep & Cross Network for Ad Click Predictions. In ADKDD@KDD. ACM, 12:1--12:7.","journal-title":"ACM"},{"key":"e_1_2_2_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2021.3064966"},{"key":"e_1_2_2_57_1","volume-title":"Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation. CoRR","author":"Wu Yonghui","year":"2016","unstructured":"Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Lukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean. 2016. Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation. CoRR, Vol. abs\/1609.08144 (2016)."},{"key":"e_1_2_2_58_1","doi-asserted-by":"crossref","unstructured":"Minhui Xie Youyou Lu Jiazhen Lin Qing Wang Jian Gao Kai Ren and Jiwu Shu. 2022. Fleche: an efficient GPU embedding cache for personalized recommendations. In EuroSys. ACM 402--416.","DOI":"10.1145\/3492321.3519554"},{"key":"e_1_2_2_59_1","unstructured":"Sixin Zhang Anna Choromanska and Yann LeCun. 2015. Deep learning with Elastic Averaging SGD. In NIPS. 685--693."},{"key":"e_1_2_2_60_1","volume-title":"Staleness-Aware Async-SGD for Distributed Deep Learning","author":"Zhang Wei","unstructured":"Wei Zhang, Suyog Gupta, Xiangru Lian, and Ji Liu. 2016. Staleness-Aware Async-SGD for Distributed Deep Learning. In IJCAI. IJCAI\/AAAI Press, 2350--2356."},{"key":"e_1_2_2_61_1","unstructured":"Weijie Zhao Deping Xie Ronglai Jia Yulei Qian Ruiquan Ding Mingming Sun and Ping Li. 2020. Distributed Hierarchical GPU Parameter Server for Massive Scale Deep Learning Ads Systems. In MLSys. mlsys.org."},{"key":"e_1_2_2_62_1","volume-title":"Chi","author":"Zhao Zhe","year":"2019","unstructured":"Zhe Zhao, Lichan Hong, Li Wei, Jilin Chen, Aniruddh Nath, Shawn Andrews, Aditee Kumthekar, Maheswaran Sathiamoorthy, Xinyang Yi, and Ed H. Chi. 2019. Recommending what video to watch next: a multitask ranking system. In RecSys. ACM, 43--51."},{"key":"e_1_2_2_63_1","doi-asserted-by":"crossref","unstructured":"Kaifu Zheng Lu Wang Yu Li Xusong Chen Hu Liu Jing Lu Xiwei Zhao Changping Peng Zhangang Lin and Jingping Shao. 2022. Implicit User Awareness Modeling via Candidate Items for CTR Prediction in Search Ads. In WWW. ACM 246--255.","DOI":"10.1145\/3485447.3511953"},{"key":"e_1_2_2_64_1","volume-title":"ShadowSync: Performing Synchronization in the Background for Highly Scalable Distributed Training. CoRR","author":"Zheng Qinqing","year":"2020","unstructured":"Qinqing Zheng, Bor-Yiing Su, Jiyan Yang, Alisson G. Azzolini, Qiang Wu, Ou Jin, Shri Karandikar, Hagay Lupesko, Liang Xiong, and Eric Zhou. 2020. ShadowSync: Performing Synchronization in the Background for Highly Scalable Distributed Training. CoRR, Vol. abs\/2003.03477 (2020)."},{"key":"e_1_2_2_65_1","volume-title":"Deep Interest Evolution Network for Click-Through Rate Prediction","author":"Zhou Guorui","unstructured":"Guorui Zhou, Na Mou, Ying Fan, Qi Pi, Weijie Bian, Chang Zhou, Xiaoqiang Zhu, and Kun Gai. 2019. Deep Interest Evolution Network for Click-Through Rate Prediction. In AAAI. AAAI Press, 5941--5948."}],"container-title":["Proceedings of the ACM on Management of Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3589310","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3589310","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:46:13Z","timestamp":1750178773000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3589310"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,6,13]]},"references-count":65,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2023,6,13]]}},"alternative-id":["10.1145\/3589310"],"URL":"https:\/\/doi.org\/10.1145\/3589310","relation":{},"ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,6,13]]}}}