{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,21]],"date-time":"2026-02-21T18:20:26Z","timestamp":1771698026994,"version":"3.50.1"},"publisher-location":"New York, NY, USA","reference-count":50,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,10,10]],"date-time":"2022-10-10T00:00:00Z","timestamp":1665360000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"the National Natural Science Foundation of China","award":["U19B2043"],"award-info":[{"award-number":["U19B2043"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,10,10]]},"DOI":"10.1145\/3503161.3547860","type":"proceedings-article","created":{"date-parts":[[2022,10,10]],"date-time":"2022-10-10T15:42:35Z","timestamp":1665416555000},"page":"4445-4455","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Unified Normalization for Accelerating and Stabilizing Transformers"],"prefix":"10.1145","author":[{"given":"Qiming","family":"Yang","sequence":"first","affiliation":[{"name":"Hikvision Research Institute, Hangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kai","family":"Zhang","sequence":"additional","affiliation":[{"name":"Hikvision Research Institute, Hangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chaoxiang","family":"Lan","sequence":"additional","affiliation":[{"name":"Hikvision Research Institute, Hangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhi","family":"Yang","sequence":"additional","affiliation":[{"name":"Hikvision Research Institute, Hangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zheyang","family":"Li","sequence":"additional","affiliation":[{"name":"Hikvision Research Institute &amp; Zhejiang University, Hangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wenming","family":"Tan","sequence":"additional","affiliation":[{"name":"Hikvision Research Institute, Hangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jun","family":"Xiao","sequence":"additional","affiliation":[{"name":"Zhejiang University, Hangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shiliang","family":"Pu","sequence":"additional","affiliation":[{"name":"Hikvision Research Institute, Hangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,10,10]]},"reference":[{"key":"e_1_3_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00013-012-0434-7"},{"key":"e_1_3_2_2_2_1","volume-title":"Jamie Ryan Kiros, and Geoffrey E. Hinton","author":"Ba Jimmy Lei","year":"2016","unstructured":"Jimmy Lei Ba , Jamie Ryan Kiros, and Geoffrey E. Hinton . 2016 . Layer Normalization . arXiv preprint arXiv:1607.06450 (2016). Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016. Layer Normalization. arXiv preprint arXiv:1607.06450 (2016)."},{"key":"e_1_3_2_2_3_1","unstructured":"Tom Brown Benjamin Mann Nick Ryder Melanie Subbiah Jared D Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell etal 2020. Language models are few-shot learners. Advances in neural information processing systems 33 (2020) 1877--1901.  Tom Brown Benjamin Mann Nick Ryder Melanie Subbiah Jared D Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell et al. 2020. Language models are few-shot learners. Advances in neural information processing systems 33 (2020) 1877--1901."},{"key":"e_1_3_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"e_1_3_2_2_5_1","unstructured":"Kai Chen Jiaqi Wang Jiangmiao Pang Yuhang Cao Yu Xiong Xiaoxiao Li Shuyang Sun Wansen Feng Ziwei Liu Jiarui Xu etal 2019. MMDetection: Open mmlab detection toolbox and benchmark. arXiv preprint arXiv:1906.07155 (2019).  Kai Chen Jiaqi Wang Jiangmiao Pang Yuhang Cao Yu Xiong Xiaoxiao Li Shuyang Sun Wansen Feng Ziwei Liu Jiarui Xu et al. 2019. MMDetection: Open mmlab detection toolbox and benchmark. arXiv preprint arXiv:1906.07155 (2019)."},{"key":"e_1_3_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00950"},{"key":"e_1_3_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00063"},{"key":"e_1_3_2_2_8_1","first-page":"613","article-title":"Tests for departure from normality. Empirical results for the distributions of b 2 and b","volume":"60","author":"Pearson RALPH","year":"1973","unstructured":"RALPH D'AGOSTINO and Egon S Pearson . 1973 . Tests for departure from normality. Empirical results for the distributions of b 2 and b . Biometrika 60 , 3 (1973), 613 -- 622 . RALPH D'AGOSTINO and Egon S Pearson. 1973. Tests for departure from normality. Empirical results for the distributions of b 2 and b. Biometrika 60, 3 (1973), 613--622.","journal-title":"Biometrika"},{"key":"e_1_3_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_2_2_10_1","volume-title":"Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2018 . Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018). Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018)."},{"key":"e_1_3_2_2_11_1","unstructured":"Mingyu Ding Bin Xiao Noel Codella Ping Luo Jingdong Wang and Lu Yuan. 2022. DaViT: Dual Attention Vision Transformers. https:\/\/doi.org\/10.48550\/ ARXIV.2204.03645  Mingyu Ding Bin Xiao Noel Codella Ping Luo Jingdong Wang and Lu Yuan. 2022. DaViT: Dual Attention Vision Transformers. https:\/\/doi.org\/10.48550\/ ARXIV.2204.03645"},{"key":"e_1_3_2_2_12_1","doi-asserted-by":"crossref","unstructured":"Xiaoyi Dong Jianmin Bao Dongdong Chen etal 2021. Cswin transformer: A general vision transformer backbone with cross-shaped windows. arXiv preprint arXiv:2107.00652 (2021).  Xiaoyi Dong Jianmin Bao Dongdong Chen et al. 2021. Cswin transformer: A general vision transformer backbone with cross-shaped windows. arXiv preprint arXiv:2107.00652 (2021).","DOI":"10.1109\/CVPR52688.2022.01181"},{"key":"e_1_3_2_2_13_1","unstructured":"Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly etal 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929 (2020).  Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929 (2020)."},{"key":"e_1_3_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01204"},{"key":"e_1_3_2_2_15_1","volume-title":"Conformer: Convolution-augmented transformer for speech recognition. arXiv preprint arXiv:2005.08100","author":"Gulati Anmol","year":"2020","unstructured":"Anmol Gulati , James Qin , Chung-Cheng Chiu , Niki Parmar , Yu Zhang , Jiahui Yu , Wei Han , Shibo Wang , Zhengdong Zhang , Yonghui Wu , 2020 . Conformer: Convolution-augmented transformer for speech recognition. arXiv preprint arXiv:2005.08100 (2020). Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, et al. 2020. Conformer: Convolution-augmented transformer for speech recognition. arXiv preprint arXiv:2005.08100 (2020)."},{"key":"e_1_3_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.322"},{"key":"e_1_3_2_2_17_1","volume-title":"Normalization techniques in training dnns: Methodology, analysis and application. arXiv preprint arXiv:2009.12836","author":"Huang Lei","year":"2020","unstructured":"Lei Huang , Jie Qin , Yi Zhou , Fan Zhu , Li Liu , and Ling Shao . 2020. Normalization techniques in training dnns: Methodology, analysis and application. arXiv preprint arXiv:2009.12836 ( 2020 ). Lei Huang, Jie Qin, Yi Zhou, Fan Zhu, Li Liu, and Ling Shao. 2020. Normalization techniques in training dnns: Methodology, analysis and application. arXiv preprint arXiv:2009.12836 (2020)."},{"key":"e_1_3_2_2_18_1","volume-title":"Shuffle Transformer: Rethinking Spatial Shuffle for Vision Transformer. arXiv preprint arXiv:2106.03650","author":"Huang Zilong","year":"2021","unstructured":"Zilong Huang , Youcheng Ben , Guozhong Luo , Pei Cheng , Gang Yu , and Bin Fu . 2021 . Shuffle Transformer: Rethinking Spatial Shuffle for Vision Transformer. arXiv preprint arXiv:2106.03650 (2021). Zilong Huang, Youcheng Ben, Guozhong Luo, Pei Cheng, Gang Yu, and Bin Fu. 2021. Shuffle Transformer: Rethinking Spatial Shuffle for Vision Transformer. arXiv preprint arXiv:2106.03650 (2021)."},{"key":"e_1_3_2_2_19_1","volume-title":"International conference on machine learning. PMLR, 448--456","author":"Ioffe Sergey","year":"2015","unstructured":"Sergey Ioffe and Christian Szegedy . 2015 . Batch normalization: Accelerating deep network training by reducing internal covariate shift . In International conference on machine learning. PMLR, 448--456 . Sergey Ioffe and Christian Szegedy. 2015. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International conference on machine learning. PMLR, 448--456."},{"key":"e_1_3_2_2_20_1","volume-title":"A survey of transformers. arXiv preprint arXiv:2106.04554","author":"Lin Tianyang","year":"2021","unstructured":"Tianyang Lin , Yuxin Wang , Xiangyang Liu , and Xipeng Qiu . 2021. A survey of transformers. arXiv preprint arXiv:2106.04554 ( 2021 ). Tianyang Lin, Yuxin Wang, Xiangyang Liu, and Xipeng Qiu. 2021. A survey of transformers. arXiv preprint arXiv:2106.04554 (2021)."},{"key":"e_1_3_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.106"},{"key":"e_1_3_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"e_1_3_2_2_23_1","volume-title":"Towards fully 8-bit integer inference for the transformer model. arXiv preprint arXiv:2009.08034","author":"Lin Ye","year":"2020","unstructured":"Ye Lin , Yanyang Li , Tengbo Liu , Tong Xiao , Tongran Liu , and Jingbo Zhu . 2020. Towards fully 8-bit integer inference for the transformer model. arXiv preprint arXiv:2009.08034 ( 2020 ). Ye Lin, Yanyang Li, Tengbo Liu, Tong Xiao, Tongran Liu, and Jingbo Zhu. 2020. Towards fully 8-bit integer inference for the transformer model. arXiv preprint arXiv:2009.08034 (2020)."},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.463"},{"key":"e_1_3_2_2_25_1","doi-asserted-by":"crossref","unstructured":"Ze Liu Han Hu Yutong Lin etal 2021. Swin Transformer V2: Scaling Up Capacity and Resolution. arXiv preprint arXiv:2111.09883 (2021).  Ze Liu Han Hu Yutong Lin et al. 2021. Swin Transformer V2: Scaling Up Capacity and Resolution. arXiv preprint arXiv:2111.09883 (2021).","DOI":"10.1109\/CVPR52688.2022.01170"},{"key":"e_1_3_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"e_1_3_2_2_27_1","volume-title":"Differentiable learning-to-normalize via switchable normalization. arXiv preprint arXiv:1806.10779","author":"Luo Ping","year":"2018","unstructured":"Ping Luo , Jiamin Ren , Zhanglin Peng , Ruimao Zhang , and Jingyu Li. 2018. Differentiable learning-to-normalize via switchable normalization. arXiv preprint arXiv:1806.10779 ( 2018 ). Ping Luo, Jiamin Ren, Zhanglin Peng, Ruimao Zhang, and Jingyu Li. 2018. Differentiable learning-to-normalize via switchable normalization. arXiv preprint arXiv:1806.10779 (2018)."},{"key":"e_1_3_2_2_28_1","volume-title":"Delight: Deep and light-weight transformer. arXiv preprint arXiv:2008.00623","author":"Mehta Sachin","year":"2020","unstructured":"Sachin Mehta , Marjan Ghazvininejad , Srinivasan Iyer , Luke Zettlemoyer , and Hannaneh Hajishirzi . 2020 . Delight: Deep and light-weight transformer. arXiv preprint arXiv:2008.00623 (2020). Sachin Mehta, Marjan Ghazvininejad, Srinivasan Iyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2020. Delight: Deep and light-weight transformer. arXiv preprint arXiv:2008.00623 (2020)."},{"key":"e_1_3_2_2_29_1","volume-title":"Do wide and deep networks learn the same things? uncovering how neural network representations vary with width and depth. arXiv preprint arXiv:2010.15327","author":"Nguyen Thao","year":"2020","unstructured":"Thao Nguyen , Maithra Raghu , and Simon Kornblith . 2020. Do wide and deep networks learn the same things? uncovering how neural network representations vary with width and depth. arXiv preprint arXiv:2010.15327 ( 2020 ). Thao Nguyen, Maithra Raghu, and Simon Kornblith. 2020. Do wide and deep networks learn the same things? uncovering how neural network representations vary with width and depth. arXiv preprint arXiv:2010.15327 (2020)."},{"key":"e_1_3_2_2_30_1","doi-asserted-by":"crossref","unstructured":"Myle Ott Sergey Edunov Alexei Baevski etal 2019. fairseq: A fast extensible toolkit for sequence modeling. arXiv preprint arXiv:1904.01038 (2019).  Myle Ott Sergey Edunov Alexei Baevski et al. 2019. fairseq: A fast extensible toolkit for sequence modeling. arXiv preprint arXiv:1904.01038 (2019).","DOI":"10.18653\/v1\/N19-4009"},{"key":"e_1_3_2_2_31_1","volume-title":"Proceedings of the 40th annual meeting of the Association for Computational Linguistics. 311--318","author":"Papineni Kishore","year":"2002","unstructured":"Kishore Papineni , Salim Roukos , Todd Ward , and Wei-Jing Zhu . 2002 . Bleu: a method for automatic evaluation of machine translation . In Proceedings of the 40th annual meeting of the Association for Computational Linguistics. 311--318 . Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics. 311--318."},{"key":"e_1_3_2_2_32_1","volume-title":"Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems 28","author":"Ren Shaoqing","year":"2015","unstructured":"Shaoqing Ren , Kaiming He , Ross Girshick , and Jian Sun . 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems 28 ( 2015 ). Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems 28 (2015)."},{"key":"e_1_3_2_2_33_1","volume-title":"Dynamic Token Normalization Improves Vision Transformer. arXiv preprint arXiv:2112.02624","author":"Shao Wenqi","year":"2021","unstructured":"Wenqi Shao , Yixiao Ge , Zhaoyang Zhang , Xuyuan Xu , XiaogangWang, Ying Shan , and Ping Luo . 2021. Dynamic Token Normalization Improves Vision Transformer. arXiv preprint arXiv:2112.02624 ( 2021 ). Wenqi Shao, Yixiao Ge, Zhaoyang Zhang, Xuyuan Xu, XiaogangWang, Ying Shan, and Ping Luo. 2021. Dynamic Token Normalization Improves Vision Transformer. arXiv preprint arXiv:2112.02624 (2021)."},{"key":"e_1_3_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00053"},{"key":"e_1_3_2_2_35_1","volume-title":"International Conference on Machine Learning. PMLR, 8741--8751","author":"Shen Sheng","year":"2020","unstructured":"Sheng Shen , Zhewei Yao , Amir Gholami , Michael Mahoney , and Kurt Keutzer . 2020 . Powernorm: Rethinking batch normalization in transformers . In International Conference on Machine Learning. PMLR, 8741--8751 . Sheng Shen, Zhewei Yao, Amir Gholami, Michael Mahoney, and Kurt Keutzer. 2020. Powernorm: Rethinking batch normalization in transformers. In International Conference on Machine Learning. PMLR, 8741--8751."},{"key":"e_1_3_2_2_36_1","volume-title":"Going deeper with image transformers. arXiv preprint arXiv:2103.17239","author":"Touvron Hugo","year":"2021","unstructured":"Hugo Touvron , Matthieu Cord , Alexandre Sablayrolles , Gabriel Synnaeve , and Herv\u00e9 J\u00e9gou . 2021. Going deeper with image transformers. arXiv preprint arXiv:2103.17239 ( 2021 ). Hugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve, and Herv\u00e9 J\u00e9gou. 2021. Going deeper with image transformers. arXiv preprint arXiv:2103.17239 (2021)."},{"key":"e_1_3_2_2_37_1","volume-title":"Instance normalization: The missing ingredient for fast stylization. arXiv preprint arXiv:1607.08022","author":"Ulyanov Dmitry","year":"2016","unstructured":"Dmitry Ulyanov , Andrea Vedaldi , and Victor Lempitsky . 2016. Instance normalization: The missing ingredient for fast stylization. arXiv preprint arXiv:1607.08022 ( 2016 ). Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky. 2016. Instance normalization: The missing ingredient for fast stylization. arXiv preprint arXiv:1607.08022 (2016)."},{"key":"e_1_3_2_2_38_1","volume-title":"Attention is all you need. Advances in neural information processing systems 30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani , Noam Shazeer , Niki Parmar , Jakob Uszkoreit , Llion Jones , Aidan N Gomez , Lukasz Kaiser , and Illia Polosukhin . 2017. Attention is all you need. Advances in neural information processing systems 30 ( 2017 ). Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017)."},{"key":"e_1_3_2_2_39_1","volume-title":"Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. arXiv preprint arXiv:2102.12122","author":"Xie Enze","year":"2021","unstructured":"WenhaiWang, Enze Xie , Xiang Li , Deng-Ping Fan , Kaitao Song , Ding Liang , Tong Lu , Ping Luo , and Ling Shao . 2021. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. arXiv preprint arXiv:2102.12122 ( 2021 ). WenhaiWang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. 2021. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. arXiv preprint arXiv:2102.12122 (2021)."},{"key":"e_1_3_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2018-1456"},{"key":"e_1_3_2_2_41_1","doi-asserted-by":"crossref","unstructured":"Yuxin Wu and Kaiming He. 2018. Group normalization. In ECCV. 3--19.  Yuxin Wu and Kaiming He. 2018. Group normalization. In ECCV. 3--19.","DOI":"10.1007\/978-3-030-01261-8_1"},{"key":"e_1_3_2_2_42_1","unstructured":"Yonghui Wu Mike Schuster Zhifeng Chen Quoc V Le Mohammad Norouzi Wolfgang Macherey Maxim Krikun Yuan Cao Qin Gao Klaus Macherey etal 2016. Google's neural machine translation system: Bridging the gap between human and machine translation. arXiv preprint arXiv:1609.08144 (2016).  Yonghui Wu Mike Schuster Zhifeng Chen Quoc V Le Mohammad Norouzi Wolfgang Macherey Maxim Krikun Yuan Cao Qin Gao Klaus Macherey et al. 2016. Google's neural machine translation system: Bridging the gap between human and machine translation. arXiv preprint arXiv:1609.08144 (2016)."},{"key":"e_1_3_2_2_43_1","volume-title":"International Conference on Machine Learning. PMLR, 10524--10533","author":"Xiong Ruibin","year":"2020","unstructured":"Ruibin Xiong , Yunchang Yang , Di He , Kai Zheng , Shuxin Zheng , Chen Xing , Huishuai Zhang , Yanyan Lan , Liwei Wang , and Tieyan Liu . 2020 . On layer normalization in the transformer architecture . In International Conference on Machine Learning. PMLR, 10524--10533 . Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu. 2020. On layer normalization in the transformer architecture. In International Conference on Machine Learning. PMLR, 10524--10533."},{"key":"e_1_3_2_2_44_1","volume-title":"Understanding and improving layer normalization. Advances in Neural Information Processing Systems 32","author":"Xu Jingjing","year":"2019","unstructured":"Jingjing Xu , Xu Sun , Zhiyuan Zhang , Guangxiang Zhao , and Junyang Lin . 2019. Understanding and improving layer normalization. Advances in Neural Information Processing Systems 32 ( 2019 ). Jingjing Xu, Xu Sun, Zhiyuan Zhang, Guangxiang Zhao, and Junyang Lin. 2019. Understanding and improving layer normalization. Advances in Neural Information Processing Systems 32 (2019)."},{"key":"e_1_3_2_2_45_1","volume-title":"Towards stabilizing batch statistics in backward propagation of batch normalization. arXiv preprint arXiv:2001.06838","author":"Yan Junjie","year":"2020","unstructured":"Junjie Yan , Ruosi Wan , Xiangyu Zhang , Wei Zhang , Yichen Wei , and Jian Sun . 2020. Towards stabilizing batch statistics in backward propagation of batch normalization. arXiv preprint arXiv:2001.06838 ( 2020 ). Junjie Yan, Ruosi Wan, Xiangyu Zhang, Wei Zhang, Yichen Wei, and Jian Sun. 2020. Towards stabilizing batch statistics in backward propagation of batch normalization. arXiv preprint arXiv:2001.06838 (2020)."},{"key":"e_1_3_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW54120.2021.00050"},{"key":"e_1_3_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2021-1983"},{"key":"e_1_3_2_2_48_1","volume-title":"SOIT: Segmenting Objects with Instance-Aware Transformers. arXiv preprint arXiv:2112.11037","author":"Yu Xiaodong","year":"2021","unstructured":"Xiaodong Yu , Dahu Shi , Xing Wei , Ye Ren , Tingqun Ye , and Wenming Tan . 2021 . SOIT: Segmenting Objects with Instance-Aware Transformers. arXiv preprint arXiv:2112.11037 (2021). Xiaodong Yu, Dahu Shi, Xing Wei, Ye Ren, Tingqun Ye, and Wenming Tan. 2021. SOIT: Segmenting Objects with Instance-Aware Transformers. arXiv preprint arXiv:2112.11037 (2021)."},{"key":"e_1_3_2_2_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00060"},{"key":"e_1_3_2_2_50_1","volume-title":"Root mean square layer normalization. Advances in Neural Information Processing Systems 32","author":"Zhang Biao","year":"2019","unstructured":"Biao Zhang and Rico Sennrich . 2019. Root mean square layer normalization. Advances in Neural Information Processing Systems 32 ( 2019 ). Biao Zhang and Rico Sennrich. 2019. Root mean square layer normalization. Advances in Neural Information Processing Systems 32 (2019)."}],"event":{"name":"MM '22: The 30th ACM International Conference on Multimedia","location":"Lisboa Portugal","acronym":"MM '22","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 30th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503161.3547860","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3503161.3547860","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:02:35Z","timestamp":1750186955000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503161.3547860"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,10]]},"references-count":50,"alternative-id":["10.1145\/3503161.3547860","10.1145\/3503161"],"URL":"https:\/\/doi.org\/10.1145\/3503161.3547860","relation":{},"subject":[],"published":{"date-parts":[[2022,10,10]]},"assertion":[{"value":"2022-10-10","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}