{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,22]],"date-time":"2025-08-22T00:41:20Z","timestamp":1755823280093,"version":"3.44.0"},"publisher-location":"New York, NY, USA","reference-count":79,"publisher":"ACM","license":[{"start":{"date-parts":[[2024,5,13]],"date-time":"2024-05-13T00:00:00Z","timestamp":1715558400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2024,5,13]]},"DOI":"10.1145\/3589334.3645463","type":"proceedings-article","created":{"date-parts":[[2024,5,8]],"date-time":"2024-05-08T07:08:13Z","timestamp":1715152093000},"page":"3485-3496","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Helen: Optimizing CTR Prediction Models with Frequency-wise Hessian Eigenvalue Regularization"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-9639-5987","authenticated-orcid":false,"given":"Zirui","family":"Zhu","sequence":"first","affiliation":[{"name":"National University of Singapore, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6542-0626","authenticated-orcid":false,"given":"Yong","family":"Liu","sequence":"additional","affiliation":[{"name":"National University of Singapore, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1505-1535","authenticated-orcid":false,"given":"Zangwei","family":"Zheng","sequence":"additional","affiliation":[{"name":"National University of Singapore, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7393-8994","authenticated-orcid":false,"given":"Huifeng","family":"Guo","sequence":"additional","affiliation":[{"name":"Huawei Noah's Ark Lab, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2816-4384","authenticated-orcid":false,"given":"Yang","family":"You","sequence":"additional","affiliation":[{"name":"National University of Singapore, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,5,13]]},"reference":[{"key":"e_1_3_2_2_1_1","volume-title":"Proceedings of the 39th International Conference on Machine Learning (Proceedings of Machine Learning Research","volume":"32","author":"Abbas Momin","year":"2022","unstructured":"Momin Abbas, Quan Xiao, Lisha Chen, Pin-Yu Chen, and Tianyi Chen. 2022. Sharp-MAML: Sharpness-Aware Model-Agnostic Meta Learning. In Proceedings of the 39th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 162), Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvari, Gang Niu, and Sivan Sabato (Eds.). PMLR, 10--32."},{"key":"e_1_3_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/3109859.3109912"},{"key":"e_1_3_2_2_3_1","volume-title":"Proceedings of the RecSys Workshop on Recommendation in Multistakeholder Environments (RMSE).","author":"Abdollahpouri Himan","year":"2019","unstructured":"Himan Abdollahpouri, Masoud Mansoury, Robin Burke, and Bamshad Mobasher. 2019. The unfairness of popularity bias in recommendation. In Proceedings of the RecSys Workshop on Recommendation in Multistakeholder Environments (RMSE)."},{"key":"e_1_3_2_2_4_1","unstructured":"Avazu. 2015. Avazu Click-Through Rate Prediction."},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/3240323.3240360"},{"key":"e_1_3_2_2_6_1","first-page":"12","article-title":"Stochastic gradient learning in neural networks","volume":"91","author":"Bottou L\u00e9on","year":"1991","unstructured":"L\u00e9on Bottou. 1991. Stochastic gradient learning in neural networks. Proceedings of Neuro-Nimes 91, 8 (1991), 12.","journal-title":"Proceedings of Neuro-Nimes"},{"key":"e_1_3_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/3580305.3599884"},{"key":"e_1_3_2_2_8_1","volume-title":"The 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24--26, 2017, Conference Track Proceedings. OpenReview.net. https:\/\/openreview. net\/forum?id=B1YfAfcgl","author":"Chaudhari Pratik","year":"2017","unstructured":"Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina. 2017. Entropy-sgd: Biasing gradient descent into wide valleys. In The 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24--26, 2017, Conference Track Proceedings. OpenReview.net. https:\/\/openreview. net\/forum?id=B1YfAfcgl"},{"key":"e_1_3_2_2_9_1","volume-title":"International Conference on Learning Representations.","author":"Chen Xiangning","year":"2022","unstructured":"Xiangning Chen, Cho-Jui Hsieh, and Boqing Gong. 2022. When vision transformers outperform resnets without pre-training or strong data augmentations. In International Conference on Learning Representations."},{"key":"e_1_3_2_2_10_1","unstructured":"Xiangning Chen Chen Liang Da Huang Esteban Real Kaiyuan Wang Yao Liu Hieu Pham Xuanyi Dong Thang Luong Cho-Jui Hsieh et al. 2023. Symbolic discovery of optimization algorithms. arXiv preprint arXiv:2302.06675 (2023)."},{"key":"e_1_3_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/2988450.2988454"},{"key":"e_1_3_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/2959100.2959190"},{"key":"e_1_3_2_2_13_1","volume-title":"Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018)."},{"key":"e_1_3_2_2_14_1","volume-title":"International Conference on Machine Learning. PMLR, 1019--1028","author":"Dinh Laurent","year":"2017","unstructured":"Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio. 2017. Sharp minima can generalize for deep nets. In International Conference on Machine Learning. PMLR, 1019--1028."},{"key":"e_1_3_2_2_15_1","volume-title":"Proceedings of the 4th International Conference on Learning Representations. 1--4.","author":"Dozat Timothy","year":"2016","unstructured":"Timothy Dozat. 2016. Incorporating Nesterov Momentum into Adam. In Proceedings of the 4th International Conference on Learning Representations. 1--4."},{"key":"e_1_3_2_2_16_1","volume-title":"Liangli Zhen, Rick Siow Mong Goh, and Vincent YF Tan.","author":"Du Jiawei","year":"2021","unstructured":"Jiawei Du, Hanshu Yan, Jiashi Feng, Joey Tianyi Zhou, Liangli Zhen, Rick Siow Mong Goh, and Vincent YF Tan. 2021. Efficient sharpness-aware minimization for improved training of neural networks. arXiv preprint arXiv:2110.03141 (2021)."},{"key":"e_1_3_2_2_17_1","article-title":"Adaptive Subgradient Methods for Online Learning and Stochastic Optimization","volume":"12","author":"Duchi John C.","year":"2010","unstructured":"John C. Duchi, Elad Hazan, and Yoram Singer. 2010. Adaptive Subgradient Methods for Online Learning and Stochastic Optimization. Journal of Machine Learning Research 12, 7 (2010).","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_2_18_1","volume-title":"Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data. arXiv preprint arXiv:1703.11008","author":"Dziugaite Gintare Karolina","year":"2017","unstructured":"Gintare Karolina Dziugaite and Daniel M Roy. 2017. Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data. arXiv preprint arXiv:1703.11008 (2017)."},{"key":"e_1_3_2_2_19_1","volume-title":"Sharpness-aware Minimization for Efficiently Improving Generalization. In International Conference on Learning Representations. https:\/\/openreview.net\/forum? id=6Tm1mposlrM","author":"Foret Pierre","year":"2021","unstructured":"Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur. 2021. Sharpness-aware Minimization for Efficiently Improving Generalization. In International Conference on Learning Representations. https:\/\/openreview.net\/forum? id=6Tm1mposlrM"},{"key":"e_1_3_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2017\/239"},{"key":"e_1_3_2_2_21_1","volume-title":"Advances in Neural Information Processing Systems","volume":"32","author":"He Haowei","year":"2019","unstructured":"Haowei He, Gao Huang, and Yang Yuan. 2019. Asymmetric Valleys: Beyond Sharp and Flat Local Minima. In Advances in Neural Information Processing Systems, Vol. 32. Curran Associates, Inc. https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2019\/file\/01d8bae291b1e4724443375634ccfa0e-Paper.pdf"},{"key":"e_1_3_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_2_23_1","first-page":"2","article-title":"Neural networks for machine learning lecture 6a overview of mini-batch gradient descent","volume":"14","author":"Hinton Geoffrey","year":"2012","unstructured":"Geoffrey Hinton, Nitish Srivastava, and Kevin Swersky. 2012. Neural networks for machine learning lecture 6a overview of mini-batch gradient descent. Cited on 14, 8 (2012), 2.","journal-title":"Cited on"},{"key":"e_1_3_2_2_24_1","volume-title":"Coursera: Neural networks for machine learning. Lecture 9c: Using noise as a regularizer","author":"Hinton Geoffrey","year":"2012","unstructured":"Geoffrey Hinton, N Srivastava, K Swersky, T Tieleman, and A Mohamed. 2012. Coursera: Neural networks for machine learning. Lecture 9c: Using noise as a regularizer (2012)."},{"key":"e_1_3_2_2_25_1","volume-title":"Flat minima. Neural computation 9, 1","author":"Hochreiter Sepp","year":"1997","unstructured":"Sepp Hochreiter and Jrgen Schmidhuber. 1997. Flat minima. Neural computation 9, 1 (1997), 1--42."},{"key":"e_1_3_2_2_26_1","volume-title":"Averaging weights leads to wider optima and better generalization. arXiv preprint arXiv:1803.05407","author":"Izmailov Pavel","year":"2018","unstructured":"Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson. 2018. Averaging weights leads to wider optima and better generalization. arXiv preprint arXiv:1803.05407 (2018)."},{"key":"e_1_3_2_2_27_1","volume-title":"Adaptive mixtures of local experts. Neural computation 3, 1","author":"Jacobs Robert A","year":"1991","unstructured":"Robert A Jacobs, Michael I Jordan, Steven J Nowlan, and Geoffrey E Hinton. 1991. Adaptive mixtures of local experts. Neural computation 3, 1 (1991), 79--87."},{"key":"e_1_3_2_2_28_1","volume-title":"On the local minima of the empirical risk. arXiv preprint arXiv:1803.09357","author":"Jin Chi","year":"2018","unstructured":"Chi Jin, Lydia T Liu, Rong Ge, and Michael I Jordan. 2018. On the local minima of the empirical risk. arXiv preprint arXiv:1803.09357 (2018)."},{"key":"e_1_3_2_2_29_1","volume-title":"Correcting Popularity Bias by Enhancing Recommendation Neutrality. RecSys Posters 10","author":"Kamishima Toshihiro","year":"2014","unstructured":"Toshihiro Kamishima, Shotaro Akaho, Hideki Asoh, and Jun Sakuma. 2014. Correcting Popularity Bias by Enhancing Recommendation Neutrality. RecSys Posters 10 (2014)."},{"key":"e_1_3_2_2_30_1","volume-title":"Proceedings on \"I Can't Believe It's Not Better! - Understanding Deep Learning Through Empirical Falsification\" at NeurIPS 2022 Workshops (Proceedings of Machine Learning Research","volume":"65","author":"Kaur Simran","year":"2023","unstructured":"Simran Kaur, Jeremy Cohen, and Zachary Chase Lipton. 2023. On the Maximum Hessian Eigenvalue and Generalization. In Proceedings on \"I Can't Believe It's Not Better! - Understanding Deep Learning Through Empirical Falsification\" at NeurIPS 2022 Workshops (Proceedings of Machine Learning Research, Vol. 187), Javier Antor\u00e1n, Arno Blaas, Fan Feng, Sahra Ghalebikesabi, Ian Mason, Melanie F. Pradier, David Rohde, Francisco J. R. Ruiz, and Aaron Schein (Eds.). PMLR, 51--65."},{"key":"e_1_3_2_2_31_1","volume-title":"5th International Conference on Learning Representations, ICLR","author":"Keskar Nitish Shirish","year":"2017","unstructured":"Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang. 2017. On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24--26, 2017, Conference Track Proceedings. OpenReview.net. https:\/\/openreview.net\/forum?id=H1oyRlYgg"},{"key":"e_1_3_2_2_32_1","volume-title":"Kingma and Jimmy Ba","author":"Diederik","year":"2015","unstructured":"Diederik P. Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7--9, 2015, Conference Track Proceedings. http:\/\/arxiv.org\/abs\/1412.6980"},{"key":"e_1_3_2_2_33_1","volume-title":"International Conference on Machine Learning. PMLR, 5905-- 5914","author":"Kwon Jungmin","year":"2021","unstructured":"Jungmin Kwon, Jeongseop Kim, Hyunseo Park, and In Kwon Choi. 2021. Asam: Adaptive sharpness-aware minimization for scale-invariant learning of deep neural networks. In International Conference on Machine Learning. PMLR, 5905-- 5914."},{"key":"e_1_3_2_2_34_1","unstructured":"Criteo Labs. 2014. Display Advertising Challenge."},{"key":"e_1_3_2_2_35_1","volume-title":"Conference on learning theory. PMLR, 1246--1257","author":"Lee Jason D","year":"2016","unstructured":"Jason D Lee, Max Simchowitz, Michael I Jordan, and Benjamin Recht. 2016. Gradient descent only converges to minimizers. In Conference on learning theory. PMLR, 1246--1257."},{"key":"e_1_3_2_2_36_1","volume-title":"Visualizing the loss landscape of neural nets. arXiv preprint arXiv:1712.09913","author":"Li Hao","year":"2017","unstructured":"Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein. 2017. Visualizing the loss landscape of neural nets. arXiv preprint arXiv:1712.09913 (2017)."},{"key":"e_1_3_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/3336191.3371785"},{"key":"e_1_3_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/3041021.3054192"},{"key":"e_1_3_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401083"},{"key":"e_1_3_2_2_40_1","volume-title":"Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training. arXiv preprint arXiv:2305.14342","author":"Liu Hong","year":"2023","unstructured":"Hong Liu, Zhiyuan Li, David Hall, Percy Liang, and Tengyu Ma. 2023. Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training. arXiv preprint arXiv:2305.14342 (2023)."},{"key":"e_1_3_2_2_41_1","volume-title":"Proceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining. 2686--2696","author":"Liu Hu","year":"2020","unstructured":"Hu Liu, Jing Lu, Hao Yang, Xiwei Zhao, Sulong Xu, Hao Peng, Zehua Zhang, Wenjie Niu, Xiaokun Zhu, Yongjun Bao, et al. 2020. Category-Specific CNN for Visual-aware CTR Prediction at JD. com. In Proceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining. 2686--2696."},{"key":"e_1_3_2_2_42_1","volume-title":"On the Variance of the Adaptive Learning Rate and Beyond. In International Conference on Learning Representations.","author":"Liu Liyuan","year":"2020","unstructured":"Liyuan Liu, Haoming Jiang, Pengcheng He,Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han. 2020. On the Variance of the Adaptive Learning Rate and Beyond. In International Conference on Learning Representations."},{"key":"e_1_3_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01204"},{"key":"e_1_3_2_2_44_1","volume-title":"DecoupledWeight Decay Regularization. In International Conference on Learning Representations. https:\/\/openreview.net\/ forum?id=Bkg6RiCqY7","author":"Loshchilov Ilya","year":"2019","unstructured":"Ilya Loshchilov and Frank Hutter. 2019. DecoupledWeight Decay Regularization. In International Conference on Learning Representations. https:\/\/openreview.net\/ forum?id=Bkg6RiCqY7"},{"key":"e_1_3_2_2_45_1","volume-title":"Asian Conference on Machine Learning. PMLR, 325--340","author":"Lu Jing","year":"2013","unstructured":"Jing Lu, Steven Hoi, and Jialei Wang. 2013. Second order online collaborative filtering. In Asian Conference on Machine Learning. PMLR, 325--340."},{"key":"e_1_3_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/2487575.2488200"},{"key":"e_1_3_2_2_47_1","unstructured":"Maxim Naumov Dheevatsa Mudigere Hao-Jun Michael Shi Jianyu Huang Narayanan Sundaraman Jongsoo Park XiaodongWang Udit Gupta Carole-Jean Wu Alisson G. Azzolini Dmytro Dzhulgakov Andrey Mallevich Ilia Cherniavskii Yinghai Lu Raghuraman Krishnamoorthi Ansha Yu Volodymyr Kondratenko Stephanie Pereira Xianjie Chen Wenlin Chen Vijay Rao Bill Jia Liang Xiong and Mikhail Smelyanskiy. 2019. Deep Learning Recommendation Model for Personalization and Recommendation Systems. arXiv abs\/1906.00091 (2019)."},{"key":"e_1_3_2_2_48_1","first-page":"543","article-title":"A method for unconstrained convex minimization problem with the rate of convergence O (1\/k2)","volume":"269","author":"Nesterov Yurii","year":"1983","unstructured":"Yurii Nesterov. 1983. A method for unconstrained convex minimization problem with the rate of convergence O (1\/k2). In Dokl. Akad. Nauk. SSSR, Vol. 269. 543.","journal-title":"Dokl. Akad. Nauk. SSSR"},{"key":"e_1_3_2_2_49_1","volume-title":"Recommendation networks and the long tail of electronic commerce. Mis quarterly","author":"Oestreicher-Singer Gal","year":"2012","unstructured":"Gal Oestreicher-Singer and Arun Sundararajan. 2012. Recommendation networks and the long tail of electronic commerce. Mis quarterly (2012), 65--83."},{"key":"e_1_3_2_2_50_1","doi-asserted-by":"publisher","DOI":"10.1214\/aop\/1176990853"},{"key":"e_1_3_2_2_51_1","volume-title":"Some methods of speeding up the convergence of iteration methods. Ussr computational mathematics and mathematical physics 4, 5","author":"Polyak Boris T","year":"1964","unstructured":"Boris T Polyak. 1964. Some methods of speeding up the convergence of iteration methods. Ussr computational mathematics and mathematical physics 4, 5 (1964), 1--17."},{"key":"e_1_3_2_2_52_1","volume-title":"On the momentum term in gradient descent learning algorithms. Neural networks 12, 1","author":"Qian Ning","year":"1999","unstructured":"Ning Qian. 1999. On the momentum term in gradient descent learning algorithms. Neural networks 12, 1 (1999), 145--151."},{"key":"e_1_3_2_2_53_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2016.0151"},{"key":"e_1_3_2_2_54_1","volume-title":"On the convergence of adam and beyond. arXiv preprint arXiv:1904.09237","author":"Reddi Sashank J","year":"2019","unstructured":"Sashank J Reddi, Satyen Kale, and Sanjiv Kumar. 2019. On the convergence of adam and beyond. arXiv preprint arXiv:1904.09237 (2019)."},{"key":"e_1_3_2_2_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/1242572.1242643"},{"key":"e_1_3_2_2_56_1","volume-title":"A stochastic approximation method. The annals of mathematical statistics","author":"Robbins Herbert","year":"1951","unstructured":"Herbert Robbins and Sutton Monro. 1951. A stochastic approximation method. The annals of mathematical statistics (1951), 400--407."},{"key":"e_1_3_2_2_57_1","doi-asserted-by":"publisher","DOI":"10.1080\/02650487.2007.11073031"},{"key":"e_1_3_2_2_58_1","doi-asserted-by":"publisher","DOI":"10.1145\/2124295.2124333"},{"key":"e_1_3_2_2_59_1","volume-title":"International Conference on Learning Representations (Workshop Track).","author":"Sagun Levent","year":"2018","unstructured":"Levent Sagun, Utku Evci, V Ugur Guney, Yann Dauphin, and Leon Bottou. 2018. Empirical analysis of the hessian of over-parametrized neural networks. In International Conference on Learning Representations (Workshop Track)."},{"key":"e_1_3_2_2_60_1","volume-title":"Proceedings of The 33rd International Conference on Machine Learning (Proceedings of Machine Learning Research","volume":"1679","author":"Schnabel Tobias","year":"2016","unstructured":"Tobias Schnabel, Adith Swaminathan, Ashudeep Singh, Navin Chandak, and Thorsten Joachims. 2016. Recommendations as Treatments: Debiasing Learning and Evaluation. In Proceedings of The 33rd International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 48), Maria Florina Balcan and Kilian Q. Weinberger (Eds.). PMLR, New York, New York, USA, 1670--1679."},{"key":"e_1_3_2_2_61_1","volume-title":"Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. arXiv preprint arXiv:1701.06538","author":"Shazeer Noam","year":"2017","unstructured":"Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. arXiv preprint arXiv:1701.06538 (2017)."},{"key":"e_1_3_2_2_62_1","volume-title":"A bayesian perspective on generalization and stochastic gradient descent. arXiv preprint arXiv:1710.06451","author":"Smith Samuel L","year":"2017","unstructured":"Samuel L Smith and Quoc V Le. 2017. A bayesian perspective on generalization and stochastic gradient descent. arXiv preprint arXiv:1710.06451 (2017)."},{"key":"e_1_3_2_2_63_1","unstructured":"Taobao. 2018. Alibaba Ad Display\/Click Data Prediction."},{"key":"e_1_3_2_2_64_1","volume-title":"International Conference on Machine Learning. PMLR, 9636--9647","author":"Tsuzuku Yusuke","year":"2020","unstructured":"Yusuke Tsuzuku, Issei Sato, and Masashi Sugiyama. 2020. Normalized flat minima: Exploring scale invariant definition of flat minima for neural networks using pac-bayesian analysis. In International Conference on Machine Learning. PMLR, 9636--9647."},{"key":"e_1_3_2_2_65_1","volume-title":"Proceedings of the 11th International Workshop on Data Mining for Online Advertising (ADKDD). 12:1--12:7.","author":"Fu Bin","year":"2017","unstructured":"RuoxiWang, Bin Fu, Gang Fu, and MingliangWang. 2017. Deep & Cross Network for Ad Click Predictions. In Proceedings of the 11th International Workshop on Data Mining for Online Advertising (ADKDD). 12:1--12:7."},{"key":"e_1_3_2_2_66_1","volume-title":"Proceedings of the Web Conference","author":"Wang Ruoxi","year":"2021","unstructured":"Ruoxi Wang, Rakesh Shivanna, Derek Zhiyuan Cheng, Sagar Jain, Dong Lin, Lichan Hong, and Ed H. Chi. 2021. DCN V2: Improved Deep & Cross Network and Practical Lessons for Web-scale Learning to Rank Systems. In Proceedings of the Web Conference 2021."},{"key":"e_1_3_2_2_67_1","volume-title":"International Conference on Learning Representations.","author":"Wen Kaiyue","year":"2022","unstructured":"Kaiyue Wen, Tengyu Ma, and Zhiyuan Li. 2022. How Sharpness-Aware Minimization Minimizes Sharpness?. In International Conference on Learning Representations."},{"key":"e_1_3_2_2_68_1","volume-title":"International conference on machine learning. PMLR, 802--810","author":"Yan Ling","year":"2014","unstructured":"Ling Yan, Wu-Jun Li, Gui-Rong Xue, and Dingyi Han. 2014. Coupled group lasso for web-scale ctr prediction in display advertising. In International conference on machine learning. PMLR, 802--810."},{"key":"e_1_3_2_2_69_1","volume-title":"Positively scale-invariant flatness of relu neural networks. arXiv preprint arXiv:1903.02237","author":"Yi Mingyang","year":"2019","unstructured":"Mingyang Yi, Qi Meng,Wei Chen, Zhi-ming Ma, and Tie-Yan Liu. 2019. Positively scale-invariant flatness of relu neural networks. arXiv preprint arXiv:1903.02237 (2019)."},{"key":"e_1_3_2_2_70_1","doi-asserted-by":"publisher","DOI":"10.14778\/2311906.2311916"},{"key":"e_1_3_2_2_71_1","first-page":"15383","article-title":"Why are adaptive methods good for attention models","volume":"33","author":"Zhang Jingzhao","year":"2020","unstructured":"Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank Reddi, Sanjiv Kumar, and Suvrit Sra. 2020. Why are adaptive methods good for attention models? Advances in Neural Information Processing Systems 33 (2020), 15383--15393.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_2_72_1","doi-asserted-by":"publisher","DOI":"10.1145\/3404835.3462875"},{"key":"e_1_3_2_2_73_1","doi-asserted-by":"publisher","DOI":"10.1145\/3442381.3449788"},{"key":"e_1_3_2_2_74_1","doi-asserted-by":"crossref","unstructured":"Zangwei Zheng Pengtai Xu Xuan Zou Da Tang Zhen Li Chenguang Xi Peng Wu Leqi Zou Yijie Zhu Ming Chen et al. 2022. CowClip: Reducing CTR Prediction Model Training Time from 12 hours to 10 minutes on 1 GPU. arXiv preprint arXiv:2204.06240 (2022).","DOI":"10.1609\/aaai.v37i9.26347"},{"key":"e_1_3_2_2_75_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33015941"},{"key":"e_1_3_2_2_76_1","doi-asserted-by":"publisher","DOI":"10.1145\/3219819.3219823"},{"key":"e_1_3_2_2_77_1","volume-title":"BARS: Towards Open Benchmarking for Recommender Systems. In SIGIR '22: The 45th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Zhu Jieming","year":"2022","unstructured":"Jieming Zhu, Quanyu Dai, Liangcai Su, Rong Ma, Jinyang Liu, Guohao Cai, Xi Xiao, and Rui Zhang. 2022. BARS: Towards Open Benchmarking for Recommender Systems. In SIGIR '22: The 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, Madrid, Spain, July 11 - 15, 2022, Enrique Amig\u00f3, Pablo Castells, Julio Gonzalo, Ben Carterette, J. Shane Culpepper, and Gabriella Kazai (Eds.). ACM, 2912--2923. https:\/\/doi.org\/10.1145\/ 3477495.3531723"},{"key":"e_1_3_2_2_78_1","doi-asserted-by":"publisher","DOI":"10.1145\/3459637.3482486"},{"key":"e_1_3_2_2_79_1","volume-title":"Surrogate gap minimization improves sharpness-aware training. arXiv preprint arXiv:2203.08065","author":"Zhuang Juntang","year":"2022","unstructured":"Juntang Zhuang, Boqing Gong, Liangzhe Yuan, Yin Cui, Hartwig Adam, Nicha Dvornek, Sekhar Tatikonda, James Duncan, and Ting Liu. 2022. Surrogate gap minimization improves sharpness-aware training. arXiv preprint arXiv:2203.08065 (2022)."}],"event":{"name":"WWW '24: The ACM Web Conference 2024","sponsor":["SIGWEB ACM Special Interest Group on Hypertext, Hypermedia, and Web"],"location":"Singapore Singapore","acronym":"WWW '24"},"container-title":["Proceedings of the ACM Web Conference 2024"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3589334.3645463","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3589334.3645463","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,22]],"date-time":"2025-08-22T00:24:54Z","timestamp":1755822294000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3589334.3645463"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,5,13]]},"references-count":79,"alternative-id":["10.1145\/3589334.3645463","10.1145\/3589334"],"URL":"https:\/\/doi.org\/10.1145\/3589334.3645463","relation":{},"subject":[],"published":{"date-parts":[[2024,5,13]]},"assertion":[{"value":"2024-05-13","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}