{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:10:32Z","timestamp":1750219832506,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":31,"publisher":"ACM","license":[{"start":{"date-parts":[[2023,9,14]],"date-time":"2023-09-14T00:00:00Z","timestamp":1694649600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,9,14]]},"DOI":"10.1145\/3604915.3608882","type":"proceedings-article","created":{"date-parts":[[2023,9,14]],"date-time":"2023-09-14T22:40:23Z","timestamp":1694731223000},"page":"267-271","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Efficient Data Representation Learning in Google-scale Systems"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0000-7943-8328","authenticated-orcid":false,"given":"Derek Zhiyuan","family":"Cheng","sequence":"first","affiliation":[{"name":"Google DeepMind, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-4824-1316","authenticated-orcid":false,"given":"Ruoxi","family":"Wang","sequence":"additional","affiliation":[{"name":"Google DeepMind, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-8795-3665","authenticated-orcid":false,"given":"Wang-Cheng","family":"Kang","sequence":"additional","affiliation":[{"name":"Google DeepMind, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-8045-3717","authenticated-orcid":false,"given":"Benjamin","family":"Coleman","sequence":"additional","affiliation":[{"name":"Google DeepMind, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5169-9849","authenticated-orcid":false,"given":"Yin","family":"Zhang","sequence":"additional","affiliation":[{"name":"Google DeepMind, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6863-8073","authenticated-orcid":false,"given":"Jianmo","family":"Ni","sequence":"additional","affiliation":[{"name":"Google DeepMind, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-4587-1458","authenticated-orcid":false,"given":"Jonathan","family":"Valverde","sequence":"additional","affiliation":[{"name":"Google DeepMind, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-9563-554X","authenticated-orcid":false,"given":"Lichan","family":"Hong","sequence":"additional","affiliation":[{"name":"Google DeepMind, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3230-5338","authenticated-orcid":false,"given":"Ed","family":"Chi","sequence":"additional","affiliation":[{"name":"Google DeepMind, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,9,14]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"Rohan Anil Sandra Gadanho Da Huang Nijith Jacob Zhuoshu Li Dong Lin Todd Phillips Cristina Pop Kevin Regan Gil\u00a0I. Shamir Rakesh Shivanna and Qiqi Yan. 2022. On the Factory Floor: ML Engineering for Industrial-Scale Ads Recommendation Models. In Proceedings of the 5th Workshop on Online Recommender Systems and User Modeling co-located with the 16th ACM Conference on Recommender Systems ORSUM@RecSys Jo\u00e3o Vinagre Marie Al-Ghossein Al\u00edpio\u00a0M\u00e1rio Jorge Albert Bifet and Ladislav Peska (Eds.)."},{"key":"e_1_3_2_1_2_1","unstructured":"Yoshua Bengio R\u00e9jean Ducharme and Pascal Vincent. 2000. A Neural Probabilistic Language Model. In Advances in Neural Information Processing Systems. 932\u2013938."},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/3159652.3159727"},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/2988450.2988454"},{"key":"e_1_3_2_1_5_1","volume-title":"Unified Embedding: Battle-Tested Feature Representations for Web-Scale ML Systems. arxiv:2305.12102\u00a0[cs.LG]","author":"Coleman Benjamin","year":"2023","unstructured":"Benjamin Coleman, Wang-Cheng Kang, Matthew Fahrbach, Ruoxi Wang, Lichan Hong, Ed\u00a0H. Chi, and Derek\u00a0Zhiyuan Cheng. 2023. Unified Embedding: Battle-Tested Feature Representations for Web-Scale ML Systems. arxiv:2305.12102\u00a0[cs.LG]"},{"key":"e_1_3_2_1_6_1","first-page":"762","article-title":"Random Offset Block Embedding (ROBE) for compressed embedding tables in deep learning recommendation systems","volume":"4","author":"Desai Aditya","year":"2022","unstructured":"Aditya Desai, Li Chou, and Anshumali Shrivastava. 2022. Random Offset Block Embedding (ROBE) for compressed embedding tables in deep learning recommendation systems. Proceedings of Machine Learning and Systems 4 (2022), 762\u2013778.","journal-title":"Proceedings of Machine Learning and Systems"},{"key":"e_1_3_2_1_7_1","unstructured":"Aditya Desai and Anshumali Shrivastava. 2022. The trade-offs of model size in large recommendation models : 100GB to 10MB Criteo-tb DLRM model. In Advances in Neural Information Processing Systems Alice\u00a0H. Oh Alekh Agarwal Danielle Belgrave and Kyunghyun Cho (Eds.). https:\/\/openreview.net\/forum?id=c9I_NArDIjD"},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"crossref","unstructured":"Kaize Ding Albert\u00a0Jiongqian Liang Bryan Perrozi Ting Chen Ruoxi Wang Lichan Hong Ed\u00a0H. Chi Huan Liu and Derek\u00a0Zhiyuan Cheng. 2023. HyperFormer: Learning Expressive Sparse Feature Representations via Hypergraph Transformer. arxiv:2305.17386\u00a0[cs.IR]","DOI":"10.1145\/3539618.3591999"},{"key":"e_1_3_2_1_9_1","volume-title":"Learning Rate Schedules in the Presence of Distribution Shift. arXiv preprint arXiv:2303.15634","author":"Fahrbach Matthew","year":"2023","unstructured":"Matthew Fahrbach, Adel Javanmard, Vahab Mirrokni, and Pratik Worah. 2023. Learning Rate Schedules in the Presence of Distribution Shift. arXiv preprint arXiv:2303.15634 (2023)."},{"key":"e_1_3_2_1_10_1","unstructured":"Huifeng Guo Ruiming Tang Yunming Ye Zhenguo Li and Xiuqiang He. 2017. DeepFM: a factorization-machine based neural network for CTR prediction. arXiv preprint arXiv:1703.04247 (2017)."},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/3579371.3589350"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/3366424.3383416"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/3447548.3467304"},{"key":"e_1_3_2_1_15_1","unstructured":"Aditya Kusupati Gantavya Bhatt Aniket Rege Matthew Wallingford Aditya Sinha Vivek Ramanujan William Howard-Snyder Kaifeng Chen Sham Kakade Prateek Jain and Ali Farhadi. 2022. Matryoshka Representation Learning. arxiv:2205.13147\u00a0[cs.LG]"},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"crossref","unstructured":"Jianxun Lian Xiaohuan Zhou Fuzheng Zhang Zhongxia Chen Xing Xie and Guangzhong Sun. 2018. xDeepFM. In SIGKDD.","DOI":"10.1145\/3219819.3220023"},{"key":"e_1_3_2_1_17_1","unstructured":"Tomas Mikolov Ilya Sutskever Kai Chen Gregory\u00a0S. Corrado and Jeffrey Dean. 2013. Distributed Representations of Words and Phrases and their Compositionality. In Advances in Neural Information Processing Systems. 3111\u20133119."},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3470496.3533727"},{"key":"e_1_3_2_1_19_1","volume-title":"Deep learning recommendation model for personalization and recommendation systems. arXiv preprint arXiv:1906.00091","author":"Naumov Maxim","year":"2019","unstructured":"Maxim Naumov, Dheevatsa Mudigere, Hao-Jun\u00a0Michael Shi, Jianyu Huang, Narayanan Sundaraman, Jongsoo Park, Xiaodong Wang, Udit Gupta, Carole-Jean Wu, Alisson\u00a0G Azzolini, 2019. Deep learning recommendation model for personalization and recommendation systems. arXiv preprint arXiv:1906.00091 (2019)."},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.2015509117"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1038\/323533a0"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394486.3403059"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/3357384.3357925"},{"key":"e_1_3_2_1_24_1","volume-title":"Hash embeddings for efficient word representations. Advances in Neural Information Processing Systems","author":"Svenstrup Dan","year":"2017","unstructured":"Dan Svenstrup, Jonas Hansen, and Ole Winther. 2017. Hash embeddings for efficient word representations. Advances in Neural Information Processing Systems (2017)."},{"key":"e_1_3_2_1_25_1","volume-title":"Attention is all you need. Advances in neural information processing systems 30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan\u00a0N Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017)."},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3124749.3124754"},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3442381.3450078"},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/3523227.3546765"},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/1553374.1553516"},{"key":"e_1_3_2_1_30_1","volume-title":"Mixed Negative Sampling for Learning Two-Tower Neural Networks in Recommendations. In Companion Proceedings of the Web Conference 2020","author":"Yang Ji","year":"2020","unstructured":"Ji Yang, Xinyang Yi, Derek Zhiyuan\u00a0Cheng, Lichan Hong, Yang Li, Simon Xiaoming\u00a0Wang, Taibai Xu, and Ed\u00a0H. Chi. 2020. Mixed Negative Sampling for Learning Two-Tower Neural Networks in Recommendations. In Companion Proceedings of the Web Conference 2020 (Taipei, Taiwan) (WWW \u201920). 441\u2013447."},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/3298689.3346996"}],"event":{"name":"RecSys '23: Seventeenth ACM Conference on Recommender Systems","sponsor":["SIGWEB ACM Special Interest Group on Hypertext, Hypermedia, and Web","SIGAI ACM Special Interest Group on Artificial Intelligence","SIGKDD ACM Special Interest Group on Knowledge Discovery in Data","SIGIR ACM Special Interest Group on Information Retrieval","SIGCHI ACM Special Interest Group on Computer-Human Interaction","SIGecom Special Interest Group on Economics and Computation"],"location":"Singapore Singapore","acronym":"RecSys '23"},"container-title":["Proceedings of the 17th ACM Conference on Recommender Systems"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3604915.3608882","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3604915.3608882","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:46:35Z","timestamp":1750178795000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3604915.3608882"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,9,14]]},"references-count":31,"alternative-id":["10.1145\/3604915.3608882","10.1145\/3604915"],"URL":"https:\/\/doi.org\/10.1145\/3604915.3608882","relation":{},"subject":[],"published":{"date-parts":[[2023,9,14]]},"assertion":[{"value":"2023-09-14","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}