{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,10]],"date-time":"2026-02-10T18:46:26Z","timestamp":1770749186064,"version":"3.50.0"},"publisher-location":"New York, NY, USA","reference-count":15,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,4,25]],"date-time":"2022-04-25T00:00:00Z","timestamp":1650844800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,4,25]]},"DOI":"10.1145\/3487553.3524206","type":"proceedings-article","created":{"date-parts":[[2022,8,16]],"date-time":"2022-08-16T22:41:30Z","timestamp":1660689690000},"page":"110-115","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["DCAF-BERT: A Distilled Cachable Adaptable Factorized Model For Improved Ads CTR Prediction"],"prefix":"10.1145","author":[{"given":"Aashiq","family":"Muhamed","sequence":"first","affiliation":[{"name":"Amazon, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jaspreet","family":"Singh","sequence":"additional","affiliation":[{"name":"Amazon, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shuai","family":"Zheng","sequence":"additional","affiliation":[{"name":"Amazon Web Services, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Iman","family":"Keivanloo","sequence":"additional","affiliation":[{"name":"Amazon, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sujan","family":"Perera","sequence":"additional","affiliation":[{"name":"Amazon, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"James","family":"Mracek","sequence":"additional","affiliation":[{"name":"Amazon, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yi","family":"Xu","sequence":"additional","affiliation":[{"name":"Amazon, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qingjun","family":"Cui","sequence":"additional","affiliation":[{"name":"Amazon, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Santosh","family":"Rajagopalan","sequence":"additional","affiliation":[{"name":"Amazon, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Belinda","family":"Zeng","sequence":"additional","affiliation":[{"name":"Amazon, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Trishul","family":"Chilimbi","sequence":"additional","affiliation":[{"name":"Amazon, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,8,16]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"Tom\u00a0B Brown Benjamin Mann Nick Ryder Melanie Subbiah Jared Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell 2020. Language models are few-shot learners. arXiv preprint arXiv:2005.14165(2020).  Tom\u00a0B Brown Benjamin Mann Nick Ryder Melanie Subbiah Jared Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell 2020. Language models are few-shot learners. arXiv preprint arXiv:2005.14165(2020)."},{"key":"e_1_3_2_1_2_1","volume-title":"Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805(2018).","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2018 . Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805(2018). Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805(2018)."},{"key":"e_1_3_2_1_3_1","unstructured":"Jingfei Du Edouard Grave Beliz Gunel Vishrav Chaudhary Onur Celebi Michael Auli Ves Stoyanov and Alexis Conneau. 2020. Self-training improves pre-training for natural language understanding. arXiv preprint arXiv:2010.02194(2020).  Jingfei Du Edouard Grave Beliz Gunel Vishrav Chaudhary Onur Celebi Michael Auli Ves Stoyanov and Alexis Conneau. 2020. Self-training improves pre-training for natural language understanding. arXiv preprint arXiv:2010.02194(2020)."},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"crossref","unstructured":"Mitchell\u00a0A Gordon and Kevin Duh. 2020. Distill adapt distill: Training small in-domain models for neural machine translation. arXiv preprint arXiv:2003.02877(2020).  Mitchell\u00a0A Gordon and Kevin Duh. 2020. Distill adapt distill: Training small in-domain models for neural machine translation. arXiv preprint arXiv:2003.02877(2020).","DOI":"10.18653\/v1\/2020.ngt-1.12"},{"key":"e_1_3_2_1_5_1","unstructured":"Geoffrey Hinton Oriol Vinyals and Jeff Dean. 2015. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531(2015).  Geoffrey Hinton Oriol Vinyals and Jeff Dean. 2015. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531(2015)."},{"key":"e_1_3_2_1_6_1","volume-title":"Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. arXiv preprint arXiv:1808.06226(2018).","author":"Kudo Taku","year":"2018","unstructured":"Taku Kudo and John Richardson . 2018 . Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. arXiv preprint arXiv:1808.06226(2018). Taku Kudo and John Richardson. 2018. Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. arXiv preprint arXiv:1808.06226(2018)."},{"key":"e_1_3_2_1_7_1","unstructured":"Guillaume Lample and Alexis Conneau. 2019. Cross-lingual language model pretraining. arXiv preprint arXiv:1901.07291(2019).  Guillaume Lample and Alexis Conneau. 2019. Cross-lingual language model pretraining. arXiv preprint arXiv:1901.07291(2019)."},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/3336191.3371785"},{"key":"e_1_3_2_1_9_1","volume-title":"Twinbert: Distilling knowledge to twin-structured bert models for efficient retrieval. arXiv preprint arXiv:2002.06275(2020).","author":"Lu Wenhao","year":"2020","unstructured":"Wenhao Lu , Jian Jiao , and Ruofei Zhang . 2020 . Twinbert: Distilling knowledge to twin-structured bert models for efficient retrieval. arXiv preprint arXiv:2002.06275(2020). Wenhao Lu, Jian Jiao, and Ruofei Zhang. 2020. Twinbert: Distilling knowledge to twin-structured bert models for efficient retrieval. arXiv preprint arXiv:2002.06275(2020)."},{"key":"e_1_3_2_1_10_1","volume-title":"NeurIPS Efficient Natural Language and Speech Processing Workshop.","author":"Muhamed Aashiq","year":"2021","unstructured":"Aashiq Muhamed , Iman Keivanloo , Sujan Perera , James Mracek , Yi Xu , Qingjun Cui , Santosh Rajagopalan , Belinda Zeng , and Trishul Chilimbi . 2021 . CTR-BERT: Cost-effective knowledge distillation for billion-parameter teacher models . In NeurIPS Efficient Natural Language and Speech Processing Workshop. Aashiq Muhamed, Iman Keivanloo, Sujan Perera, James Mracek, Yi Xu, Qingjun Cui, Santosh Rajagopalan, Belinda Zeng, and Trishul Chilimbi. 2021. CTR-BERT: Cost-effective knowledge distillation for billion-parameter teacher models. In NeurIPS Efficient Natural Language and Speech Processing Workshop."},{"key":"e_1_3_2_1_11_1","unstructured":"Iulia Turc Ming-Wei Chang Kenton Lee and Kristina Toutanova. 2019. Well-read students learn better: On the importance of pre-training compact models. arXiv preprint arXiv:1908.08962(2019).  Iulia Turc Ming-Wei Chang Kenton Lee and Kristina Toutanova. 2019. Well-read students learn better: On the importance of pre-training compact models. arXiv preprint arXiv:1908.08962(2019)."},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i05.6451"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/3442381.3450078"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"crossref","unstructured":"Weinan Zhang Jiarui Qin Wei Guo Ruiming Tang and Xiuqiang He. 2021. Deep Learning for Click-Through Rate Estimation. arXiv preprint arXiv:2104.10584(2021).  Weinan Zhang Jiarui Qin Wei Guo Ruiming Tang and Xiuqiang He. 2021. Deep Learning for Click-Through Rate Estimation. arXiv preprint arXiv:2104.10584(2021).","DOI":"10.24963\/ijcai.2021\/636"},{"key":"e_1_3_2_1_15_1","unstructured":"Barret Zoph Golnaz Ghiasi Tsung-Yi Lin Yin Cui Hanxiao Liu Ekin\u00a0D Cubuk and Quoc\u00a0V Le. 2020. Rethinking pre-training and self-training. arXiv preprint arXiv:2006.06882(2020).  Barret Zoph Golnaz Ghiasi Tsung-Yi Lin Yin Cui Hanxiao Liu Ekin\u00a0D Cubuk and Quoc\u00a0V Le. 2020. Rethinking pre-training and self-training. arXiv preprint arXiv:2006.06882(2020)."}],"event":{"name":"WWW '22: The ACM Web Conference 2022","location":"Virtual Event, Lyon France","acronym":"WWW '22","sponsor":["SIGWEB ACM Special Interest Group on Hypertext, Hypermedia, and Web"]},"container-title":["Companion Proceedings of the Web Conference 2022"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3487553.3524206","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3487553.3524206","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:30:33Z","timestamp":1750188633000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3487553.3524206"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,4,25]]},"references-count":15,"alternative-id":["10.1145\/3487553.3524206","10.1145\/3487553"],"URL":"https:\/\/doi.org\/10.1145\/3487553.3524206","relation":{},"subject":[],"published":{"date-parts":[[2022,4,25]]},"assertion":[{"value":"2022-08-16","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}