{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,25]],"date-time":"2026-07-25T19:01:13Z","timestamp":1785006073501,"version":"3.55.0"},"reference-count":256,"publisher":"Association for Computing Machinery (ACM)","issue":"10s","license":[{"start":{"date-parts":[[2022,1,31]],"date-time":"2022-01-31T00:00:00Z","timestamp":1643587200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2022,1,31]]},"abstract":"<jats:p>Astounding results from Transformer models on natural language tasks have intrigued the vision community to study their application to computer vision problems. Among their salient benefits, Transformers enable modeling long dependencies between input sequence elements and support parallel processing of sequence as compared to recurrent networks, e.g., Long short-term memory. Different from convolutional networks, Transformers require minimal inductive biases for their design and are naturally suited as set-functions. Furthermore, the straightforward design of Transformers allows processing multiple modalities (e.g., images, videos, text, and speech) using similar processing blocks and demonstrates excellent scalability to very large capacity networks and huge datasets. These strengths have led to exciting progress on a number of vision tasks using Transformer networks. This survey aims to provide a comprehensive overview of the Transformer models in the computer vision discipline. We start with an introduction to fundamental concepts behind the success of Transformers, i.e., self-attention, large-scale pre-training, and bidirectional feature encoding. We then cover extensive applications of transformers in vision including popular recognition tasks (e.g., image classification, object detection, action recognition, and segmentation), generative modeling, multi-modal tasks (e.g., visual-question answering, visual reasoning, and visual grounding), video processing (e.g., activity recognition, video forecasting), low-level vision (e.g., image super-resolution, image enhancement, and colorization), and three-dimensional analysis (e.g., point cloud classification and segmentation). We compare the respective advantages and limitations of popular techniques both in terms of architectural design and their experimental value. Finally, we provide an analysis on open research directions and possible future works. We hope this effort will ignite further interest in the community to solve current challenges toward the application of transformer models in computer vision.<\/jats:p>","DOI":"10.1145\/3505244","type":"journal-article","created":{"date-parts":[[2022,1,6]],"date-time":"2022-01-06T16:22:00Z","timestamp":1641486120000},"page":"1-41","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2965,"title":["Transformers in Vision: A Survey"],"prefix":"10.1145","volume":"54","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9502-1749","authenticated-orcid":false,"given":"Salman","family":"Khan","sequence":"first","affiliation":[{"name":"MBZUAI, UAE and Australian National University, Canberra, ACT, AU"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7663-7161","authenticated-orcid":false,"given":"Muzammal","family":"Naseer","sequence":"additional","affiliation":[{"name":"MBZUAI, UAE and Australian National University, Canberra, ACT, AU"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2706-5985","authenticated-orcid":false,"given":"Munawar","family":"Hayat","sequence":"additional","affiliation":[{"name":"Department of DSAI, Faculty of IT, Monash University, Clayton, Victoria, AU"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7198-0187","authenticated-orcid":false,"given":"Syed Waqas","family":"Zamir","sequence":"additional","affiliation":[{"name":"Inception Institute of Artificial Intelligence, Masdar City, Abu Dhabi, UAE"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4263-3143","authenticated-orcid":false,"given":"Fahad Shahbaz","family":"Khan","sequence":"additional","affiliation":[{"name":"MBZUAI, UAE and CVL, Link\u00f6ping University, Link\u00f6ping, Sweden"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6172-5572","authenticated-orcid":false,"given":"Mubarak","family":"Shah","sequence":"additional","affiliation":[{"name":"CRCV, University of Central Florida, Orlando, FL, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,9,13]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"https:\/\/www.youtube.com\/watch?v=UX8OubxsY8w AAAI 2020 Keynotes Turing Award Winners Event"},{"key":"e_1_3_2_3_2","unstructured":"https:\/\/lambdalabs.com\/blog\/demystifying-gpt-3\/ OpenAI\u2019s GPT-3 Language Model: A Technical Overview"},{"key":"e_1_3_2_4_2","unstructured":"https:\/\/ai.googleblog.com\/2017\/07\/revisiting-unr easonable-effectiveness.html Revisiting the Unreasonable Effectiveness of Data"},{"key":"e_1_3_2_5_2","article-title":"Quantifying attention flow in transformers","author":"Abnar Samira","year":"2020","unstructured":"Samira Abnar and Willem Zuidema. 2020. Quantifying attention flow in transformers. arXiv:2005.00928. Retrieved from https:\/\/arxiv.org\/abs\/2005.00928.","journal-title":"arXiv:2005.00928"},{"key":"e_1_3_2_6_2","volume-title":"WACV","author":"Ahsan Unaiza","year":"2019","unstructured":"Unaiza Ahsan, Rishi Madhok, and Irfan Essa. 2019. Video jigsaw: Unsupervised learning of spatiotemporal context for video action recognition. In WACV."},{"key":"e_1_3_2_7_2","volume-title":"EMNLP","author":"Alberti Chris","year":"2019","unstructured":"Chris Alberti, Jeffrey Ling, Michael Collins, and David Reitter. 2019. Fusion of detected objects in text for visual question answering. In EMNLP."},{"key":"e_1_3_2_8_2","volume-title":"ICLR","year":"2022","unstructured":"Anonymous. 2022. Patches are all you need? In ICLR."},{"key":"e_1_3_2_9_2","volume-title":"ICCV","author":"Antol Stanislaw","year":"2015","unstructured":"Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015. VQA: Visual question answering. In ICCV."},{"key":"e_1_3_2_10_2","volume-title":"ICCV","author":"Arandjelovic Relja","year":"2017","unstructured":"Relja Arandjelovic and Andrew Zisserman. 2017. Look, listen and learn. In ICCV."},{"key":"e_1_3_2_11_2","article-title":"Vivit: A video vision transformer","author":"Arnab Anurag","year":"2021","unstructured":"Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lu\u010di\u0107, and Cordelia Schmid. 2021. Vivit: A video vision transformer. arXiv:2103.15691. Retrieved from https:\/\/arxiv.org\/abs\/2103.15691.","journal-title":"arXiv:2103.15691"},{"key":"e_1_3_2_12_2","article-title":"Layer normalization","author":"Ba Jimmy Lei","year":"2016","unstructured":"Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016. Layer normalization. arXiv:1607.06450. Retrieved from https:\/\/arxiv.org\/abs\/1607.06450.","journal-title":"arXiv:1607.06450"},{"key":"e_1_3_2_13_2","volume-title":"NeurIPS","author":"Bachman Philip","year":"2019","unstructured":"Philip Bachman, R. Hjelm, and W. Buchwalter. 2019. Learning representations by maximizing mutual information across views. In NeurIPS."},{"key":"e_1_3_2_14_2","volume-title":"ICLR","author":"Bello Irwan","year":"2021","unstructured":"Irwan Bello. 2021. LambdaNetworks: Modeling long-range interactions without attention. In ICLR."},{"key":"e_1_3_2_15_2","volume-title":"ICCV","author":"Bello Irwan","year":"2019","unstructured":"Irwan Bello, Barret Zoph, Ashish Vaswani, Jonathon Shlens, and Quoc V. Le. 2019. Attention augmented convolutional networks. In ICCV."},{"key":"e_1_3_2_16_2","article-title":"Longformer: The long-document transformer","author":"Beltagy Iz","year":"2020","unstructured":"Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020. Longformer: The long-document transformer. arXiv:2004.05150. Retrieved from https:\/\/arxiv.org\/abs\/2004.05150.","journal-title":"arXiv:2004.05150"},{"key":"e_1_3_2_17_2","volume-title":"Deep Learning","author":"Bengio Yoshua","year":"2017","unstructured":"Yoshua Bengio, Ian Goodfellow, and Aaron Courville. 2017. Deep Learning. MIT Press."},{"key":"e_1_3_2_18_2","first-page":"9739","volume-title":"CVPR","author":"Bertasius Gedas","year":"2020","unstructured":"Gedas Bertasius and Lorenzo Torresani. 2020. Classifying, segmenting, and tracking object instances in video with mask propagation. In CVPR. 9739\u20139748."},{"key":"e_1_3_2_19_2","volume-title":"ICML","author":"Bertasius Gedas","year":"2021","unstructured":"Gedas Bertasius, Heng Wang, and Lorenzo Torresani. 2021. Is space-time attention all you need for video understanding? In ICML."},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patrec.2008.04.005"},{"key":"e_1_3_2_21_2","article-title":"Language models are few-shot learners","author":"Brown Tom B.","year":"2020","unstructured":"Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et\u00a0al. 2020. Language models are few-shot learners. arXiv:2005.14165. Retrieved from https:\/\/arxiv.org\/abs\/2005.14165.","journal-title":"arXiv:2005.14165"},{"key":"e_1_3_2_22_2","volume-title":"CVPR","author":"Buades Antoni","year":"2005","unstructured":"Antoni Buades, Bartomeu Coll, and J.-M. Morel. 2005. A non-local algorithm for image denoising. In CVPR."},{"key":"e_1_3_2_23_2","article-title":"End-to-end object detection with transformers","author":"Carion Nicolas","year":"2020","unstructured":"Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. 2020. End-to-end object detection with transformers. arXiv:2005.12872. Retrieved from https:\/\/arxiv.org\/abs\/2005.12872.","journal-title":"arXiv:2005.12872"},{"key":"e_1_3_2_24_2","article-title":"Emerging properties in self-supervised vision transformers","author":"Caron Mathilde","year":"2021","unstructured":"Mathilde Caron, Hugo Touvron, Ishan Misra, Herv\u00e9 J\u00e9gou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. 2021. Emerging properties in self-supervised vision transformers. arXiv:2104.14294. Retrieved from https:\/\/arxiv.org\/abs\/2104.14294.","journal-title":"arXiv:2104.14294"},{"key":"e_1_3_2_25_2","article-title":"A short note on the kinetics-700 human action dataset","author":"Carreira Joao","year":"2019","unstructured":"Joao Carreira, Eric Noland, Chloe Hillier, and A. Zisserman. 2019. A short note on the kinetics-700 human action dataset. arXiv:1907.06987. Retrieved from https:\/\/arxiv.org\/abs\/1907.06987.","journal-title":"arXiv:1907.06987"},{"key":"e_1_3_2_26_2","article-title":"An attentive survey of attention models","author":"Chaudhari Sneha","year":"2019","unstructured":"Sneha Chaudhari, Gungor Polatkan, Rohan Ramanath, and Varun Mithal. 2019. An attentive survey of attention models. arXiv:1904.02874. Retrieved from https:\/\/arxiv.org\/abs\/1904.02874.","journal-title":"arXiv:1904.02874"},{"key":"e_1_3_2_27_2","article-title":"Transformer interpretability beyond attention visualization","author":"Chefer Hila","year":"2020","unstructured":"Hila Chefer, Shir Gur, and Lior Wolf. 2020. Transformer interpretability beyond attention visualization. arXiv:2012.09838. Retrieved from https:\/\/arxiv.org\/abs\/2012.09838.","journal-title":"arXiv:2012.09838"},{"key":"e_1_3_2_28_2","article-title":"GLiT: Neural architecture search for global and local image transformer","author":"Chen Boyu","year":"2021","unstructured":"Boyu Chen, Peixia Li, Chuming Li, Baopu Li, Lei Bai, Chen Lin, Ming Sun, Wanli Ouyang, et\u00a0al. 2021. GLiT: Neural architecture search for global and local image transformer. arXiv:2107.02960 (2021).","journal-title":"arXiv:2107.02960"},{"key":"e_1_3_2_29_2","article-title":"Crossvit: Cross-attention multi-scale vision transformer for image classification","author":"Chen Chun-Fu","year":"2021","unstructured":"Chun-Fu Chen, Quanfu Fan, and Rameswar Panda. 2021. Crossvit: Cross-attention multi-scale vision transformer for image classification. arXiv:2103.14899. Retrieved from https:\/\/arxiv.org\/abs\/2103.14899.","journal-title":"arXiv:2103.14899"},{"key":"e_1_3_2_30_2","unstructured":"Chun-Fu Chen Rameswar Panda and Quanfu Fan. 2021. RegionViT: Regional-to-local attention for vision transformers. arxiv:cs.CV\/2106.02689. Retrieved from https:\/\/arxiv.org\/abs\/2106.02689."},{"key":"e_1_3_2_31_2","article-title":"Pre-trained image processing transformer","author":"Chen Hanting","year":"2020","unstructured":"Hanting Chen, Yunhe Wang, Tianyu Guo, Chang Xu, Yiping Deng, Zhenhua Liu, Siwei Ma, Chunjing Xu, Chao Xu, and Wen Gao. 2020. Pre-trained image processing transformer. arXiv:2012.00364. Retrieved from https:\/\/arxiv.org\/abs\/2012.00364.","journal-title":"arXiv:2012.00364"},{"key":"e_1_3_2_32_2","article-title":"Topological planning with transformers for vision-and-language navigation","author":"Chen Kevin","year":"2020","unstructured":"Kevin Chen, Junshen K. Chen, Jo Chuang, Marynel V\u00e1zquez, and Silvio Savarese. 2020. Topological planning with transformers for vision-and-language navigation. arXiv:2012.05292. Retrieved from https:\/\/arxiv.org\/abs\/2012.05292.","journal-title":"arXiv:2012.05292"},{"key":"e_1_3_2_33_2","article-title":"AutoFormer: Searching transformers for visual recognition","author":"Chen Minghao","year":"2021","unstructured":"Minghao Chen, Houwen Peng, Jianlong Fu, and Haibin Ling. 2021. AutoFormer: Searching transformers for visual recognition. arXiv:2107.00651. Retrieved from https:\/\/arxiv.org\/abs\/2107.00651.","journal-title":"arXiv:2107.00651"},{"key":"e_1_3_2_34_2","volume-title":"ICML","author":"Chen Mark","year":"2020","unstructured":"Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and I. Sutskever. 2020. Generative pretraining from pixels. In ICML."},{"key":"e_1_3_2_35_2","article-title":"A simple framework for contrastive learning of visual representations","author":"Chen Ting","year":"2020","unstructured":"Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A simple framework for contrastive learning of visual representations. arXiv:2002.05709. Retrieved from https:\/\/arxiv.org\/abs\/2002.05709.","journal-title":"arXiv:2002.05709"},{"key":"e_1_3_2_36_2","unstructured":"Ting Chen Saurabh Saxena Lala Li David J. Fleet and Geoffrey Hinton. 2021. Pix2seq: A Language Modeling Framework for Object Detection. arxiv:cs.CV\/2109.10852. Retrieved from https:\/\/arxiv.org\/abs\/2109.10852."},{"key":"e_1_3_2_37_2","article-title":"Improved baselines with momentum contrastive learning","author":"Chen Xinlei","year":"2020","unstructured":"Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He. 2020. Improved baselines with momentum contrastive learning. arXiv:2003.04297. Retrieved from https:\/\/arxiv.org\/abs\/2003.04297.","journal-title":"arXiv:2003.04297"},{"key":"e_1_3_2_38_2","article-title":"When vision transformers outperform ResNets without pretraining or strong data augmentations","author":"Chen Xiangning","year":"2021","unstructured":"Xiangning Chen, Cho-Jui Hsieh, and Boqing Gong. 2021. When vision transformers outperform ResNets without pretraining or strong data augmentations. arXiv:2106.01548. Retrieved from https:\/\/arxiv.org\/abs\/2106.01548.","journal-title":"arXiv:2106.01548"},{"key":"e_1_3_2_39_2","doi-asserted-by":"crossref","unstructured":"Xinlei Chen Saining Xie and Kaiming He. 2021. An empirical study of training self-supervised visual transformers. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV\u201921) . 9640\u20139649.","DOI":"10.1109\/ICCV48922.2021.00950"},{"key":"e_1_3_2_40_2","volume-title":"ECCV","author":"Chen Yen-Chun","year":"2020","unstructured":"Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020. UNITER: Universal image-text representation learning. In ECCV."},{"key":"e_1_3_2_41_2","article-title":"Generating long sequences with sparse transformers","author":"Child Rewon","year":"2019","unstructured":"Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019. Generating long sequences with sparse transformers. arXiv:1904.10509. Retrieved from https:\/\/arxiv.org\/abs\/1904.10509.","journal-title":"arXiv:1904.10509"},{"key":"e_1_3_2_42_2","volume-title":"ICLR","author":"Choromanski Krzysztof","year":"2021","unstructured":"Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et\u00a0al. 2021. Rethinking attention with performers. In ICLR."},{"key":"e_1_3_2_43_2","article-title":"Twins: Revisiting the design of spatial attention in vision transformers","author":"Chu Xiangxiang","year":"2021","unstructured":"Xiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang, Haibing Ren, Xiaolin Wei, Huaxia Xia, and Chunhua Shen. 2021. Twins: Revisiting the design of spatial attention in vision transformers. arXiv:2104.13840. Retrieved from https:\/\/arxiv.org\/abs\/2104.13840.","journal-title":"arXiv:2104.13840"},{"key":"e_1_3_2_44_2","unstructured":"Xiangxiang Chu Zhi Tian Bo Zhang Xinlong Wang Xiaolin Wei Huaxia Xia and Chunhua Shen. 2021. Conditional Positional Encodings for Vision Transformers. arxiv:cs.CV\/2102.10882. Retrieved from https:\/\/arxiv.org\/abs\/2102.10882."},{"key":"e_1_3_2_45_2","volume-title":"AISTATS","author":"Coates Adam","year":"2011","unstructured":"Adam Coates, Andrew Ng, and Honglak Lee. 2011. An analysis of single-layer networks in unsupervised feature learning. In AISTATS."},{"key":"e_1_3_2_46_2","volume-title":"ICLR","author":"Cordonnier Jean-Baptiste","year":"2019","unstructured":"Jean-Baptiste Cordonnier, Andreas Loukas, and Martin Jaggi. 2019. On the relationship between self-attention and convolutional layers. In ICLR."},{"key":"e_1_3_2_47_2","volume-title":"CVPR","author":"Cordts Marius","year":"2016","unstructured":"Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. 2016. The cityscapes dataset for semantic urban scene understanding. In CVPR."},{"key":"e_1_3_2_48_2","first-page":"764","volume-title":"ICCV","author":"Dai Jifeng","year":"2017","unstructured":"Jifeng Dai, Haozhi Qi, Yuwen Xiong, Yi Li, Guodong Zhang, Han Hu, and Yichen Wei. 2017. Deformable convolutional networks. In ICCV. 764\u2013773."},{"key":"e_1_3_2_49_2","volume-title":"CVPR","author":"Dai Tao","year":"2019","unstructured":"Tao Dai, Jianrui Cai, Yongbing Zhang, S. Xia, and L. Zhang. 2019. Second-order attention network for single image super-resolution. In CVPR."},{"key":"e_1_3_2_50_2","unstructured":"Zihang Dai Hanxiao Liu Quoc V. Le and Mingxing Tan. 2021. CoAtNet: Marrying convolution and attention for all data sizes. arxiv:cs.CV\/2106.04803. Retrieved from https:\/\/arxiv.org\/abs\/2106.04803."},{"key":"e_1_3_2_51_2","article-title":"Attention, please! A survey of neural attention models in deep learning","author":"Correia Alana de Santana","year":"2021","unstructured":"Alana de Santana Correia and Esther Luna Colombini. 2021. Attention, please! A survey of neural attention models in deep learning. arXiv:2103.16775. Retrieved from https:\/\/arxiv.org\/abs\/2103.16775.","journal-title":"arXiv:2103.16775"},{"key":"e_1_3_2_52_2","volume-title":"CVPR","author":"Deng Jia","year":"2009","unstructured":"Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. ImageNet: A large-scale hierarchical image database. In CVPR."},{"key":"e_1_3_2_53_2","doi-asserted-by":"crossref","unstructured":"Jiajun Deng Zhengyuan Yang Tianlang Chen Wengang Zhou and Houqiang Li. 2021. TransVG: End-to-end visual grounding with transformers. arxiv:cs.CV\/2104.08541. Retrieved from https:\/\/arxiv.org\/abs\/2104.08541.","DOI":"10.1109\/ICCV48922.2021.00179"},{"key":"e_1_3_2_54_2","article-title":"BERT: Pre-training of deep bidirectional transformers for language understanding","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805. Retrieved from https:\/\/arxiv.org\/abs\/1810.04805.","journal-title":"arXiv:1810.04805"},{"key":"e_1_3_2_55_2","article-title":"CrossTransformers: Spatially-aware few-shot transfer","author":"Doersch Carl","year":"2020","unstructured":"Carl Doersch, Ankush Gupta, and Andrew Zisserman. 2020. CrossTransformers: Spatially-aware few-shot transfer. NeurIPS.","journal-title":"NeurIPS"},{"key":"e_1_3_2_56_2","article-title":"Cswin transformer: A general vision transformer backbone with cross-shaped windows","author":"Dong Xiaoyi","year":"2021","unstructured":"Xiaoyi Dong, Jianmin Bao, Dongdong Chen, Weiming Zhang, Nenghai Yu, Lu Yuan, Dong Chen, and Baining Guo. 2021. Cswin transformer: A general vision transformer backbone with cross-shaped windows. arXiv:2107.00652. Retrieved from https:\/\/arxiv.org\/abs\/2107.00652.","journal-title":"arXiv:2107.00652"},{"key":"e_1_3_2_57_2","article-title":"An image is worth 16  \\( \\times \\)  16 words: Transformers for image recognition at scale","author":"Dosovitskiy Alexey","year":"2020","unstructured":"Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et\u00a0al. 2020. An image is worth 16 \\( \\times \\) 16 words: Transformers for image recognition at scale. arXiv:2010.11929. Retrieved from https:\/\/arxiv.org\/abs\/2010.11929.","journal-title":"arXiv:2010.11929"},{"key":"e_1_3_2_58_2","unstructured":"Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly Jakob Uszkoreit and Neil Houlsby. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arxiv:cs.CV\/2010.11929. Retrieved from https:\/\/arxiv.org\/abs\/2010.11929."},{"key":"e_1_3_2_59_2","article-title":"Visual grounding with transformers","author":"Du Ye","year":"2021","unstructured":"Ye Du, Zehua Fu, Qingjie Liu, and Yunhong Wang. 2021. Visual grounding with transformers. arXiv:2105.04281. Retrieved from https:\/\/arxiv.org\/abs\/2105.04281.","journal-title":"arXiv:2105.04281"},{"key":"e_1_3_2_60_2","unstructured":"Alaaeldin El-Nouby Hugo Touvron Mathilde Caron Piotr Bojanowski Matthijs Douze Armand Joulin Ivan Laptev Natalia Neverova Gabriel Synnaeve Jakob Verbeek and Herv\u00e9 Jegou. 2021. XCiT: Cross-Covariance Image Transformers. Advances in Neural Information Processing Systems 34 (2021)."},{"key":"e_1_3_2_61_2","article-title":"Taming transformers for high-resolution image synthesis","author":"Esser Patrick","year":"2020","unstructured":"Patrick Esser, Robin Rombach, and Bj\u00f6rn Ommer. 2020. Taming transformers for high-resolution image synthesis. arXiv:2012.09841. Retrieved from https:\/\/arxiv.org\/abs\/2012.09841.","journal-title":"arXiv:2012.09841"},{"key":"e_1_3_2_62_2","doi-asserted-by":"crossref","unstructured":"Haoqi Fan Bo Xiong Karttikeya Mangalam Yanghao Li Zhicheng Yan Jitendra Malik and Christoph Feichtenhofer. 2021. Multiscale vision transformers. arxiv:cs.CV\/2104.11227. Retrieved fromd https:\/\/arxiv.org\/abs\/2104.11227.","DOI":"10.1109\/ICCV48922.2021.00675"},{"key":"e_1_3_2_63_2","unstructured":"Yuxin Fang Bencheng Liao Xinggang Wang Jiemin Fang Jiyang Qi Rui Wu Jianwei Niu and Wenyu Liu. 2021. You only look at one sequence: Rethinking transformer in vision through object detection. arxiv:cs.CV\/2106.00666. Retrieved from https:\/\/arxiv.org\/abs\/2106.00666."},{"key":"e_1_3_2_64_2","article-title":"Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity","author":"Fedus William","unstructured":"William Fedus, Barret Zoph, and Noam Shazeer. [n.d.]. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. arXiv:2101.03961. Retrieved from https:\/\/arxiv.org\/abs\/2101.03961.","journal-title":"arXiv:2101.03961"},{"key":"e_1_3_2_65_2","article-title":"Sharpness-aware minimization for efficiently improving generalization","author":"Foret Pierre","year":"2020","unstructured":"Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur. 2020. Sharpness-aware minimization for efficiently improving generalization. arXiv:2010.01412. Retrieved from https:\/\/arxiv.org\/abs\/2010.01412.","journal-title":"arXiv:2010.01412"},{"key":"e_1_3_2_66_2","first-page":"5680","volume-title":"CVPR","author":"Gao Chen","year":"2020","unstructured":"Chen Gao, Yunpeng Chen, Si Liu, Zhenxiong Tan, and Shuicheng Yan. 2020. Adversarialnas: Adversarial neural architecture search for gans. In CVPR. 5680\u20135689."},{"key":"e_1_3_2_67_2","volume-title":"ICASSP","author":"Gemmeke Jort F.","year":"2017","unstructured":"Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter. 2017. Audio set: An ontology and human-labeled dataset for audio events. In ICASSP."},{"key":"e_1_3_2_68_2","volume-title":"ICLR","author":"Gidaris Spyros","year":"2018","unstructured":"Spyros Gidaris, Praveer Singh, and Nikos Komodakis. 2018. Unsupervised representation learning by predicting image rotations. In ICLR."},{"key":"e_1_3_2_69_2","article-title":"COOT: Cooperative hierarchical transformer for video-text representation learning","author":"Ging Simon","year":"2020","unstructured":"Simon Ging, Mohammadreza Zolfaghari, Hamed Pirsiavash, and Thomas Brox. 2020. COOT: Cooperative hierarchical transformer for video-text representation learning. arXiv:2011.00597. Retrieved from https:\/\/arxiv.org\/abs\/2011.00597.","journal-title":"arXiv:2011.00597"},{"key":"e_1_3_2_70_2","volume-title":"CVPR","author":"Girdhar Rohit","year":"2019","unstructured":"Rohit Girdhar, Joao Carreira, Carl Doersch, and Andrew Zisserman. 2019. Video action transformer network. In CVPR."},{"key":"e_1_3_2_71_2","volume-title":"ICCV","author":"Girshick Ross","year":"2015","unstructured":"Ross Girshick. 2015. Fast R-CNN. In ICCV."},{"key":"e_1_3_2_72_2","volume-title":"NeurIPS","author":"Goodfellow Ian","year":"2014","unstructured":"Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. Generative adversarial nets. In NeurIPS."},{"key":"e_1_3_2_73_2","doi-asserted-by":"crossref","unstructured":"Ben Graham Alaaeldin El-Nouby Hugo Touvron Pierre Stock Armand Joulin Herv\u00e9 J\u00e9gou and Matthijs Douze. 2021. LeViT: a vision transformer in ConvNet\u2019s clothing for faster inference. arxiv:cs.CV\/2104.01136. Retrieved from https:\/\/arxiv.org\/abs\/2104.01136.","DOI":"10.1109\/ICCV48922.2021.01204"},{"key":"e_1_3_2_74_2","article-title":"Bootstrap your own latent: A new approach to self-supervised learning","author":"Grill Jean-Bastien","year":"2020","unstructured":"Jean-Bastien Grill, Florian Strub, Florent Altch\u00e9, Corentin Tallec, Pierre H. Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, et\u00a0al. 2020. Bootstrap your own latent: A new approach to self-supervised learning. In NeurIPS.","journal-title":"NeurIPS"},{"key":"e_1_3_2_75_2","article-title":"Transformer in transformer","author":"Han Kai","year":"2021","unstructured":"Kai Han, An Xiao, Enhua Wu, Jianyuan Guo, Chunjing Xu, and Yunhe Wang. 2021. Transformer in transformer. arXiv:2103.00112. Retrieved from https:\/\/arxiv.org\/abs\/2103.00112.","journal-title":"arXiv:2103.00112"},{"key":"e_1_3_2_76_2","volume-title":"CVPR","author":"Han Wei","year":"2018","unstructured":"Wei Han, Shiyu Chang, Ding Liu, Mo Yu, M. Witbrock, and T. Huang. 2018. Image super-resolution via dual-state recurrent networks. In CVPR."},{"key":"e_1_3_2_77_2","volume-title":"CVPR","author":"Hao Weituo","year":"2020","unstructured":"Weituo Hao, Chunyuan Li, Xiujun Li, Lawrence Carin, and Jianfeng Gao. 2020. Towards learning a generic agent for vision-and-language navigation via pre-training. In CVPR."},{"key":"e_1_3_2_78_2","unstructured":"Ali Hassani Steven Walton Nikhil Shah Abulikemu Abuduweili Jiachen Li and Humphrey Shi. 2021. Escaping the big data paradigm with compact transformers. arxiv:cs.CV\/2104.05704. Retrieved from https:\/\/arxiv.org\/abs\/2104.05704."},{"key":"e_1_3_2_79_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00975"},{"key":"e_1_3_2_80_2","volume-title":"ICCV","author":"He Kaiming","year":"2017","unstructured":"Kaiming He, Georgia Gkioxari, Piotr Doll\u00e1r, and Ross Girshick. 2017. Mask R-CNN. In ICCV."},{"key":"e_1_3_2_81_2","volume-title":"CVPR","author":"He Kaiming","year":"2016","unstructured":"Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In CVPR."},{"key":"e_1_3_2_82_2","article-title":"TransReID: Transformer-based object re-identification","author":"He Shuting","year":"2021","unstructured":"Shuting He, Hao Luo, P. Wang, F. Wang, H. Li, and W. Jiang. 2021. TransReID: Transformer-based object re-identification. arXiv:2102.04378. Retrieved from https:\/\/arxiv.org\/abs\/2102.04378.","journal-title":"arXiv:2102.04378"},{"key":"e_1_3_2_83_2","article-title":"Data-efficient image recognition with contrastive predictive coding","author":"H\u00e9naff Olivier J.","year":"2019","unstructured":"Olivier J. H\u00e9naff, Aravind Srinivas, Jeffrey De Fauw, Ali Razavi, Carl Doersch, S. M. Eslami, and Aaron van den Oord. 2019. Data-efficient image recognition with contrastive predictive coding. arXiv:1905.09272. Retrieved from https:\/\/arxiv.org\/abs\/1905.09272.","journal-title":"arXiv:1905.09272"},{"key":"e_1_3_2_84_2","article-title":"Axial attention in multidimensional transformers","author":"Ho Jonathan","year":"2019","unstructured":"Jonathan Ho, Nal Kalchbrenner, Dirk Weissenborn, and Tim Salimans. 2019. Axial attention in multidimensional transformers. arXiv:1912.12180. Retrieved from https:\/\/arxiv.org\/abs\/1912.12180.","journal-title":"arXiv:1912.12180"},{"key":"e_1_3_2_85_2","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_2_86_2","volume-title":"IntelliSys","author":"Hu Dichao","year":"2019","unstructured":"Dichao Hu. 2019. An introductory survey on attention mechanisms in NLP problems. In IntelliSys."},{"key":"e_1_3_2_87_2","volume-title":"ICCV","author":"Hu Han","year":"2019","unstructured":"Han Hu, Zheng Zhang, Zhenda Xie, and Stephen Lin. 2019. Local relation networks for image recognition. In ICCV."},{"key":"e_1_3_2_88_2","unstructured":"Zilong Huang Youcheng Ben Guozhong Luo Pei Cheng Gang Yu and Bin Fu. 2021. Shuffle transformer: Rethinking spatial shuffle for vision transformer. arxiv:cs.CV\/2106.03650. Retrieved from https:\/\/arxiv.org\/abs\/2106.03650."},{"key":"e_1_3_2_89_2","volume-title":"ICCV","author":"Huang Zilong","year":"2019","unstructured":"Zilong Huang, Xinggang Wang, Lichao Huang, Chang Huang, Yunchao Wei, and Wenyu Liu. 2019. CCNet: Criss-cross attention for semantic segmentation. In ICCV."},{"key":"e_1_3_2_90_2","article-title":"Perceiver IO: A general architecture for structured inputs & outputs","author":"Jaegle Andrew","year":"2021","unstructured":"Andrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch, Catalin Ionescu, David Ding, Skanda Koppula, Daniel Zoran, Andrew Brock, Evan Shelhamer, et\u00a0al. 2021. Perceiver IO: A general architecture for structured inputs & outputs. arXiv:2107.14795. Retrieved from https:\/\/arxiv.org\/abs\/2107.14795.","journal-title":"arXiv:2107.14795"},{"key":"e_1_3_2_91_2","article-title":"Perceiver: General perception with iterative attention","author":"Jaegle Andrew","year":"2021","unstructured":"Andrew Jaegle, Felix Gimeno, Andrew Brock, Andrew Zisserman, Oriol Vinyals, and Joao Carreira. 2021. Perceiver: General perception with iterative attention. arXiv:2103.03206. Retrieved from https:\/\/arxiv.org\/abs\/2103.03206.","journal-title":"arXiv:2103.03206"},{"key":"e_1_3_2_92_2","unstructured":"Yifan Jiang Shiyu Chang and Zhangyang Wang. 2021. TransGAN: Two transformers can make one strong GAN. arxiv:cs.CV\/2102.07074. Retrieved from https:\/\/arxiv.org\/abs\/2102.07074."},{"key":"e_1_3_2_93_2","article-title":"All tokens matter: Token labeling for training better vision transformers","author":"Jiang Zihang","year":"2021","unstructured":"Zihang Jiang, Qibin Hou, Li Yuan, Daquan Zhou, Yujun Shi, Xiaojie Jin, Anran Wang, and Jiashi Feng. 2021. All tokens matter: Token labeling for training better vision transformers. arXiv:2104.10858. Retrieved from https:\/\/arxiv.org\/abs\/2104.10858.","journal-title":"arXiv:2104.10858"},{"key":"e_1_3_2_94_2","article-title":"Self-supervised visual feature learning with deep neural networks: A survey","author":"Jing Longlong","year":"2020","unstructured":"Longlong Jing and Yingli Tian. 2020. Self-supervised visual feature learning with deep neural networks: A survey. IEEE Trans. Pattern Anal. Mach. Intell. (2020).","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"e_1_3_2_95_2","article-title":"Self-supervised spatiotemporal feature learning via video rotation prediction","author":"Jing Longlong","year":"2018","unstructured":"Longlong Jing, Xiaodong Yang, Jingen Liu, and Yingli Tian. 2018. Self-supervised spatiotemporal feature learning via video rotation prediction. arXiv:1811.11387. Retrieved from https:\/\/arxiv.org\/abs\/1811.11387.","journal-title":"arXiv:1811.11387"},{"key":"e_1_3_2_96_2","volume-title":"ECCV","author":"Johnson Justin","year":"2016","unstructured":"Justin Johnson, Alexandre Alahi, and Li Fei-Fei. 2016. Perceptual losses for real-time style transfer and super-resolution. In ECCV."},{"key":"e_1_3_2_97_2","article-title":"MDETR\u2013Modulated detection for end-to-end multi-modal understanding","author":"Kamath Aishwarya","year":"2021","unstructured":"Aishwarya Kamath, Mannat Singh, Yann LeCun, Ishan Misra, Gabriel Synnaeve, and Nicolas Carion. 2021. MDETR\u2013Modulated detection for end-to-end multi-modal understanding. arXiv:2104.12763. Retrieved from https:\/\/arxiv.org\/abs\/2104.12763.","journal-title":"arXiv:2104.12763"},{"key":"e_1_3_2_98_2","first-page":"8110","volume-title":"CVPR","author":"Karras Tero","year":"2020","unstructured":"Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. 2020. Analyzing and improving the image quality of stylegan. In CVPR. 8110\u20138119."},{"key":"e_1_3_2_99_2","article-title":"The kinetics human action video dataset","author":"Kay Will","year":"2017","unstructured":"Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et\u00a0al. 2017. The kinetics human action video dataset. arXiv:1705.06950. Retrieved from https:\/\/arxiv.org\/abs\/1705.06950.","journal-title":"arXiv:1705.06950"},{"key":"e_1_3_2_100_2","volume-title":"EMNLP","author":"Kazemzadeh Sahar","year":"2014","unstructured":"Sahar Kazemzadeh, Vicente Ordonez, M. Matten, and T. Berg. 2014. Referitgame: Referring to objects in photographs of natural scenes. In EMNLP."},{"key":"e_1_3_2_101_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-01821-3"},{"key":"e_1_3_2_102_2","doi-asserted-by":"publisher","DOI":"10.1145\/10.1109\/wacv.2018.00092"},{"key":"e_1_3_2_103_2","article-title":"Auto-encoding variational bayes","author":"Kingma Diederik P.","year":"2013","unstructured":"Diederik P. Kingma and Max Welling. 2013. Auto-encoding variational bayes. arXiv:1312.6114. Retrieved from https:\/\/arxiv.org\/abs\/1312.6114.","journal-title":"arXiv:1312.6114"},{"key":"e_1_3_2_104_2","volume-title":"CVPR","author":"Kirillov Alexander","year":"2019","unstructured":"Alexander Kirillov, Kaiming He, Ross Girshick, Carsten Rother, and Piotr Doll\u00e1r. 2019. Panoptic segmentation. In CVPR."},{"key":"e_1_3_2_105_2","volume-title":"ICLR","author":"Kitaev Nikita","year":"2020","unstructured":"Nikita Kitaev, \u0141ukasz Kaiser, and Anselm Levskaya. 2020. Reformer: The efficient transformer. In ICLR."},{"key":"e_1_3_2_106_2","volume-title":"NeurIPS","author":"Korbar Bruno","year":"2018","unstructured":"Bruno Korbar, Du Tran, and Lorenzo T.2018. Cooperative learning of audio and video models from self-supervised synchronization. In NeurIPS."},{"key":"e_1_3_2_107_2","first-page":"706","volume-title":"ICCV","author":"Krishna Ranjay","year":"2017","unstructured":"Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles. 2017. Dense-captioning events in videos. In ICCV. 706\u2013715."},{"key":"e_1_3_2_108_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-016-0981-7"},{"key":"e_1_3_2_109_2","volume-title":"Learning Multiple Layers of Features from Tiny Images","author":"Krizhevsky Alex","year":"2009","unstructured":"Alex Krizhevsky. 2009. Learning Multiple Layers of Features from Tiny Images. Technical Report."},{"key":"e_1_3_2_110_2","volume-title":"ICLR","author":"Kumar Manoj","year":"2021","unstructured":"Manoj Kumar, Dirk Weissenborn, and Nal Kalchbrenner. 2021. Colorization transformer. In ICLR."},{"key":"e_1_3_2_111_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature14539"},{"key":"e_1_3_2_112_2","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1989.1.4.541"},{"key":"e_1_3_2_113_2","volume-title":"CVPR","author":"Ledig Christian","year":"2017","unstructured":"Christian Ledig, Lucas Theis, Ferenc Husz\u00e1r, Jose Caballero, Andrew Cunningham, Alejandro Acosta, Andrew Aitken, Alykhan Tejani, Johannes Totz, Zehan Wang, et\u00a0al. 2017. Photo-realistic single image super-resolution using a generative adversarial network. In CVPR."},{"key":"e_1_3_2_114_2","volume-title":"ICCV","author":"Lee Hsin-Ying","year":"2017","unstructured":"Hsin-Ying Lee, Jia-Bin Huang, Maneesh Singh, and Ming-Hsuan Yang. 2017. Unsupervised representation learning by sorting sequences. In ICCV."},{"key":"e_1_3_2_115_2","volume-title":"ECCV","author":"Lee Kuang-Huei","year":"2018","unstructured":"Kuang-Huei Lee, Xi Chen, Gang Hua, Houdong Hu, and Xiaodong He. 2018. Stacked cross attention for image-text matching. In ECCV."},{"key":"e_1_3_2_116_2","article-title":"Parameter efficient multimodal transformers for video representation learning","author":"Lee Sangho","year":"2020","unstructured":"Sangho Lee, Youngjae Yu, Gunhee Kim, Thomas Breuel, Jan Kautz, and Yale Song. 2020. Parameter efficient multimodal transformers for video representation learning. arXiv:2012.04124. Retrieved from https:\/\/arxiv.org\/abs\/2012.04124.","journal-title":"arXiv:2012.04124"},{"key":"e_1_3_2_117_2","article-title":"Gshard: Scaling giant models with conditional computation and automatic sharding","author":"Lepikhin Dmitry","year":"2020","unstructured":"Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. 2020. Gshard: Scaling giant models with conditional computation and automatic sharding. arXiv:2006.16668. Retrieved from https:\/\/arxiv.org\/abs\/2006.16668.","journal-title":"arXiv:2006.16668"},{"key":"e_1_3_2_118_2","article-title":"Bossnas: Exploring hybrid cnn-transformers with block-wisely self-supervised neural architecture search","author":"Li Changlin","year":"2021","unstructured":"Changlin Li, Tao Tang, Guangrun Wang, Jiefeng Peng, Bing Wang, Xiaodan Liang, and Xiaojun Chang. 2021. Bossnas: Exploring hybrid cnn-transformers with block-wisely self-supervised neural architecture search. arXiv:2103.12424. Retrieved from https:\/\/arxiv.org\/abs\/2103.12424.","journal-title":"arXiv:2103.12424"},{"key":"e_1_3_2_119_2","article-title":"Efficient Self-supervised vision transformers for representation learning","author":"Li Chunyuan","year":"2021","unstructured":"Chunyuan Li, Jianwei Yang, Pengchuan Zhang, Mei Gao, Bin Xiao, Xiyang Dai, Lu Yuan, and Jianfeng Gao. 2021. Efficient Self-supervised vision transformers for representation learning. arXiv:2106.09785. Retrieved from https:\/\/arxiv.org\/abs\/2106.09785.","journal-title":"arXiv:2106.09785"},{"key":"e_1_3_2_120_2","volume-title":"AAAI","author":"Li Gen","year":"2020","unstructured":"Gen Li, Nan Duan, Yuejian Fang, Ming Gong, Daxin Jiang, and Ming Zhou. 2020. Unicoder-VL: A universal encoder for vision and language by cross-modal pre-training. In AAAI."},{"key":"e_1_3_2_121_2","volume-title":"arXiv:1908.03557","author":"Li Liunian Harold","year":"2019","unstructured":"Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019. VisualBERT: A simple and performant baseline for vision and language. arXiv:1908.03557. Retrieved from https:\/\/arxiv.org\/abs\/1908.03557."},{"key":"e_1_3_2_122_2","article-title":"Referring transformer: A one-step approach to multi-task visual grounding","author":"Li Muchen","year":"2021","unstructured":"Muchen Li and Leonid Sigal. 2021. Referring transformer: A one-step approach to multi-task visual grounding. arXiv:2106.03089. Retrieved from https:\/\/arxiv.org\/abs\/2106.03089.","journal-title":"arXiv:2106.03089"},{"key":"e_1_3_2_123_2","volume-title":"ECCV","author":"Li Xiujun","year":"2020","unstructured":"Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et\u00a0al. 2020. Oscar: Object-semantics aligned pre-training for vision-language tasks. In ECCV."},{"key":"e_1_3_2_124_2","unstructured":"Yawei Li Kai Zhang Jiezhang Cao Radu Timofte and Luc Van Gool. 2021. LocalViT: Bringing locality to vision transformers. arxiv:cs.CV\/2104.05707. Retrieved from https:\/\/arxiv.org\/abs\/2104.05707."},{"key":"e_1_3_2_125_2","volume-title":"ICCVW","author":"Liang Jingyun","year":"2021","unstructured":"Jingyun Liang, Jiezhang Cao, Guolei Sun, Kai Zhang, Luc Van Gool, and Radu Timofte. 2021. SwinIR: Image restoration using swin transformer. In ICCVW."},{"key":"e_1_3_2_126_2","article-title":"Look into person: Joint body parsing & pose estimation network and a new benchmark","author":"Liang Xiaodan","year":"2018","unstructured":"Xiaodan Liang, Ke Gong, Xiaohui Shen, and Liang Lin. 2018. Look into person: Joint body parsing & pose estimation network and a new benchmark. IEEE Trans. Pattern Anal. Mach. Intell. (2018).","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"e_1_3_2_127_2","volume-title":"CVPRW","author":"Lim Bee","year":"2017","unstructured":"Bee Lim, Sanghyun Son, Heewon Kim, Seungjun Nah, and Kyoung Mu Lee. 2017. Enhanced deep residual networks for single image super-resolution. In CVPRW."},{"key":"e_1_3_2_128_2","doi-asserted-by":"crossref","unstructured":"Daoyu Lin Kun Fu Yang Wang Guangluan Xu and Xian Sun. 2017. MARTA GANs: Unsupervised representation learning for remote sensing image classification. IEEE Geoscience and Remote Sensing Letters 14 11 (2017) 2092\u20132096.","DOI":"10.1109\/LGRS.2017.2752750"},{"key":"e_1_3_2_129_2","article-title":"M6: A Chinese multimodal pretrainer","author":"Lin Junyang","year":"2021","unstructured":"Junyang Lin, Rui Men, An Yang, Chang Zhou, Ming Ding, Yichang Zhang, Peng Wang, Ang Wang, Le Jiang, Xianyan Jia, et\u00a0al. 2021. M6: A Chinese multimodal pretrainer. arXiv:2103.00823. Retrieved from https:\/\/arxiv.org\/abs\/2103.00823.","journal-title":"arXiv:2103.00823"},{"key":"e_1_3_2_130_2","article-title":"End-to-end human pose and mesh reconstruction with transformers","author":"Lin Kevin","year":"2020","unstructured":"Kevin Lin, Lijuan Wang, and Zicheng Liu. 2020. End-to-end human pose and mesh reconstruction with transformers. arXiv:2012.09760. Retrieved from https:\/\/arxiv.org\/abs\/2012.09760.","journal-title":"arXiv:2012.09760"},{"key":"e_1_3_2_131_2","volume-title":"CVPR","author":"Lin Tsung-Yi","year":"2017","unstructured":"Tsung-Yi Lin, Piotr Doll\u00e1r, Ross Girshick, Kaiming He, B. Hariharan, and S. Belongie. 2017. Feature pyramid networks for object detection. In CVPR."},{"key":"e_1_3_2_132_2","volume-title":"ECCV","author":"Lin Tsung-Yi","year":"2014","unstructured":"Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll\u00e1r, and C Lawrence Zitnick. 2014. Microsoft COCO: Common objects in context. In ECCV."},{"key":"e_1_3_2_133_2","doi-asserted-by":"publisher","DOI":"10.1145\/10.1109\/TPAMI.2019.2916873"},{"key":"e_1_3_2_134_2","volume-title":"ECCV","author":"Liu Wei","year":"2016","unstructured":"Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander C. Berg. 2016. SSD: Single shot multibox detector. In ECCV."},{"key":"e_1_3_2_135_2","article-title":"Self-supervised learning: Generative or contrastive","author":"Liu Xiao","year":"2020","unstructured":"Xiao Liu, Fanjin Zhang, Zhenyu Hou, Zhaoyu Wang, Li Mian, Jing Zhang, and Jie Tang. 2020. Self-supervised learning: Generative or contrastive. arXiv:2006.08218. Retrieved from https:\/\/arxiv.org\/abs\/2006.08218.","journal-title":"arXiv:2006.08218"},{"key":"e_1_3_2_136_2","article-title":"RoBERTa: A robustly optimized bert pretraining approach","author":"Liu Yinhan","year":"2019","unstructured":"Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A robustly optimized bert pretraining approach. arXiv:1907.11692. Retrieved from https:\/\/arxiv.org\/abs\/1907.11692.","journal-title":"arXiv:1907.11692"},{"key":"e_1_3_2_137_2","article-title":"Transformer in convolutional neural networks","author":"Liu Yun","year":"2021","unstructured":"Yun Liu, Guolei Sun, Yu Qiu, Le Zhang, Ajad Chhatkuli, and Luc Van Gool. 2021. Transformer in convolutional neural networks. arXiv:2106.03180. Retrieved from https:\/\/arxiv.org\/abs\/2106.03180.","journal-title":"arXiv:2106.03180"},{"key":"e_1_3_2_138_2","volume-title":"ICCV","author":"Liu Ze","year":"2021","unstructured":"Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. Swin transformer: Hierarchical vision transformer using shifted windows. In ICCV."},{"key":"e_1_3_2_139_2","volume-title":"NeurIPS","author":"Lu Jiasen","year":"2019","unstructured":"Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In NeurIPS."},{"key":"e_1_3_2_140_2","article-title":"Efficient transformer for single image super-resolution","author":"Lu Zhisheng","year":"2021","unstructured":"Zhisheng Lu, Hong Liu, Juncheng Li, and Linlin Zhang. 2021. Efficient transformer for single image super-resolution. arXiv:2108.11084. Retrieved from https:\/\/arxiv.org\/abs\/2108.11084.","journal-title":"arXiv:2108.11084"},{"key":"e_1_3_2_141_2","article-title":"Improving vision-and-language navigation with image-text pairs from the web","author":"Majumdar Arjun","year":"2020","unstructured":"Arjun Majumdar, Ayush Shrivastava, Stefan Lee, Peter Anderson, Devi Parikh, and Dhruv Batra. 2020. Improving vision-and-language navigation with image-text pairs from the web. arXiv:2004.14973. Retrieved from https:\/\/arxiv.org\/abs\/2004.14973.","journal-title":"arXiv:2004.14973"},{"key":"e_1_3_2_142_2","volume-title":"CVPR","author":"Mao Junhua","year":"2016","unstructured":"Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L. Yuille, and Kevin Murphy. 2016. Generation and comprehension of unambiguous object descriptions. In CVPR."},{"key":"e_1_3_2_143_2","article-title":"Do you even need attention? A stack of feed-forward layers does surprisingly well on imagenet","author":"Melas-Kyriazi Luke","year":"2021","unstructured":"Luke Melas-Kyriazi. 2021. Do you even need attention? A stack of feed-forward layers does surprisingly well on imagenet. arXiv:2105.02723. Retrieved from https:\/\/arxiv.org\/abs\/2105.02723.","journal-title":"arXiv:2105.02723"},{"key":"e_1_3_2_144_2","volume-title":"ECCV","author":"Misra Ishan","year":"2016","unstructured":"Ishan Misra, C. Lawrence Zitnick, and Martial Hebert. 2016. Shuffle and learn: Unsupervised learning using temporal order verification. In ECCV."},{"key":"e_1_3_2_145_2","article-title":"Intriguing properties of vision transformers","author":"Naseer Muzammal","year":"2021","unstructured":"Muzammal Naseer, Kanchana Ranasinghe, Salman Khan, Munawar Hayat, Fahad Shahbaz Khan, and Ming-Hsuan Yang. 2021. Intriguing properties of vision transformers. arXiv:2105.10497. Retrieved from https:\/\/arxiv.org\/abs\/2105.10497.","journal-title":"arXiv:2105.10497"},{"key":"e_1_3_2_146_2","volume-title":"NeurIPS","author":"Naseer Muhammad Muzammal","year":"2019","unstructured":"Muhammad Muzammal Naseer, Salman H. Khan, Muhammad Haris Khan, Fahad Shahbaz Khan, and Fatih Porikli. 2019. Cross-domain transferability of adversarial perturbations. In NeurIPS."},{"key":"e_1_3_2_147_2","article-title":"Video transformer network","author":"Neimark Daniel","year":"2021","unstructured":"Daniel Neimark, Omri Bar, Maya Zohar, and Dotan Asselmann. 2021. Video transformer network. arXiv:2102.00719. Retrieved from https:\/\/arxiv.org\/abs\/2102.00719.","journal-title":"arXiv:2102.00719"},{"key":"e_1_3_2_148_2","volume-title":"ICCV","author":"Neuhold Gerhard","year":"2017","unstructured":"Gerhard Neuhold, Tobias Ollmann, Samuel Rota Bulo, and Peter Kontschieder. 2017. The mapillary vistas dataset for semantic understanding of street scenes. In ICCV."},{"key":"e_1_3_2_149_2","volume-title":"ECCV","author":"Niu Ben","year":"2020","unstructured":"Ben Niu, Weilei Wen, Wenqi Ren, Xiangde Zhang, Lianping Yang, Shuzhen Wang, Kaihao Zhang, Xiaochun Cao, and Haifeng Shen. 2020. Single image super-resolution via a holistic attention network. In ECCV."},{"key":"e_1_3_2_150_2","volume-title":"ECCV","author":"Noroozi Mehdi","year":"2016","unstructured":"Mehdi Noroozi and Paolo Favaro. 2016. Unsupervised learning of visual representations by solving jigsaw puzzles. In ECCV."},{"key":"e_1_3_2_151_2","volume-title":"NeurIPS","author":"Ordonez Vicente","year":"2011","unstructured":"Vicente Ordonez, Girish Kulkarni, and Tamara L. Berg. 2011. Im2Text: Describing images using 1 million captioned photographs. In NeurIPS."},{"key":"e_1_3_2_152_2","volume-title":"WMT","author":"Ott Myle","year":"2018","unstructured":"Myle Ott, Sergey Edunov, David Grangier, and Michael Auli. 2018. Scaling neural machine translation. In WMT."},{"key":"e_1_3_2_153_2","volume-title":"ECCV","author":"Park Seong-Jin","year":"2018","unstructured":"Seong-Jin Park, Hyeongseok Son, Sunghyun Cho, Ki-Sang Hong, and Seungyong Lee. 2018. SRFEAT: Single image super-resolution with feature discrimination. In ECCV."},{"key":"e_1_3_2_154_2","volume-title":"NeurIPS","author":"Parmar Niki","year":"2019","unstructured":"Niki Parmar, Prajit Ramachandran, Ashish Vaswani, Irwan Bello, Anselm Levskaya, and Jon Shlens. 2019. Stand-alone self-attention in vision models. In NeurIPS."},{"key":"e_1_3_2_155_2","volume-title":"ICML","author":"Parmar Niki","year":"2018","unstructured":"Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, \u0141ukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran. 2018. Image transformer. In ICML."},{"key":"e_1_3_2_156_2","volume-title":"CVPR","author":"Pathak Deepak","year":"2016","unstructured":"Deepak Pathak, Philipp Krahenbuhl, Jeff Donahue, Trevor Darrell, and A. Efros. 2016. Context encoders: Feature learning by inpainting. In CVPR."},{"key":"e_1_3_2_157_2","volume-title":"ICLR","author":"Peng Hao","year":"2021","unstructured":"Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah A. Smith, and Lingpeng Kong. 2021. Random feature attention. In ICLR."},{"key":"e_1_3_2_158_2","volume-title":"ICLR","author":"P\u00e9rez Jorge","year":"2018","unstructured":"Jorge P\u00e9rez, Javier Marinkovi\u0107, and Pablo Barcel\u00f3. 2018. On the Turing completeness of modern neural network architectures. In ICLR."},{"key":"e_1_3_2_159_2","article-title":"Spatial temporal transformer network for skeleton-based action recognition","author":"Plizzari Chiara","year":"2020","unstructured":"Chiara Plizzari, Marco Cannici, and Matteo Matteucci. 2020. Spatial temporal transformer network for skeleton-based action recognition. arXiv:2008.07404. Retrieved from https:\/\/arxiv.org\/abs\/2008.07404.","journal-title":"arXiv:2008.07404"},{"key":"e_1_3_2_160_2","first-page":"700","volume-title":"BIBM","author":"Prangemeier Tim","year":"2020","unstructured":"Tim Prangemeier, Christoph Reich, and Heinz Koeppl. 2020. Attention-based transformers for instance segmentation of cells in microstructures. In BIBM. IEEE, 700\u2013707."},{"key":"e_1_3_2_161_2","first-page":"T2","article-title":"Learning transferable visual models from natural language supervision","volume":"2","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et\u00a0al. 2021. Learning transferable visual models from natural language supervision. Image 2 (2021), T2.","journal-title":"Image"},{"key":"e_1_3_2_162_2","article-title":"Unsupervised representation learning with deep convolutional generative adversarial networks","author":"Radford Alec","year":"2015","unstructured":"Alec Radford, Luke Metz, and Soumith Chintala. 2015. Unsupervised representation learning with deep convolutional generative adversarial networks. arXiv:1511.06434. Retrieved from https:\/\/arxiv.org\/abs\/1511.06434.","journal-title":"arXiv:1511.06434"},{"key":"e_1_3_2_163_2","volume-title":"Improving Language Understanding by Generative Pre-training","author":"Radford Alec","year":"2018","unstructured":"Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018. Improving Language Understanding by Generative Pre-training. Technical Report. OpenAI."},{"key":"e_1_3_2_164_2","volume-title":"Language Models Are Unsupervised Multitask Learners","author":"Radford Alec","year":"2019","unstructured":"Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language Models Are Unsupervised Multitask Learners. Technical Report. OpenAI."},{"key":"e_1_3_2_165_2","volume-title":"CVPR","author":"Radosavovic Ilija","year":"2020","unstructured":"Ilija Radosavovic, Raj Prateek Kosaraju, Ross Girshick, Kaiming He, and Piotr Doll\u00e1r. 2020. Designing network design spaces. In CVPR."},{"key":"e_1_3_2_166_2","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","author":"Raffel Colin","year":"2019","unstructured":"Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv:1910.10683. Retrieved from https:\/\/arxiv.org\/abs\/1910.10683.","journal-title":"arXiv:1910.10683"},{"key":"e_1_3_2_167_2","volume-title":"DALL  \\( {\\cdot } \\)  E: Creating Images from Text","author":"Ramesh Aditya","year":"2021","unstructured":"Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, and Scott Gray. 2021. DALL \\( {\\cdot } \\) E: Creating Images from Text. Technical Report. OpenAI."},{"key":"e_1_3_2_168_2","volume-title":"NeurISP","author":"Razavi Ali","year":"2019","unstructured":"Ali Razavi, Aaron van den Oord, and Oriol Vinyals. 2019. Generating diverse high-fidelity images with vq-vae-2. In NeurISP."},{"key":"e_1_3_2_169_2","volume-title":"CVPR","author":"Redmon Joseph","year":"2016","unstructured":"Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi. 2016. You only look once: Unified, real-time object detection. In CVPR."},{"key":"e_1_3_2_170_2","volume-title":"ICML","author":"Reed Scott","year":"2016","unstructured":"Scott Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, and Honglak Lee. 2016. Generative adversarial text to image synthesis. In ICML."},{"key":"e_1_3_2_171_2","article-title":"Faster R-CNN: Towards real-time object detection with region proposal networks","author":"Ren Shaoqing","year":"2016","unstructured":"Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2016. Faster R-CNN: Towards real-time object detection with region proposal networks. IEEE Trans. Pattern Anal. Mach. Intell. (2016).","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"e_1_3_2_172_2","volume-title":"ICCV","author":"Sajjadi Mehdi S. M.","year":"2017","unstructured":"Mehdi S. M. Sajjadi, Bernhard Scholkopf, and Michael Hirsch. 2017. EnhanceNet: Single image super-resolution through automated texture synthesis. In ICCV."},{"key":"e_1_3_2_173_2","volume-title":"GCPR","author":"Sayed Nawid","year":"2018","unstructured":"Nawid Sayed, Biagio Brattoli, and Bj\u00f6rn Ommer. 2018. Cross and learn: Cross-modal self-supervision. In GCPR."},{"key":"e_1_3_2_174_2","first-page":"0","volume-title":"ICCV Workshops","author":"Seong Hongje","year":"2019","unstructured":"Hongje Seong, Junhyuk Hyun, and Euntai Kim. 2019. Video multitask transformer network. In ICCV Workshops. 0\u20130."},{"key":"e_1_3_2_175_2","volume-title":"CVPR","author":"Shahroudy Amir","year":"2016","unstructured":"Amir Shahroudy, Jun Liu, Tian-Tsong Ng, and Gang Wang. 2016. NTU RGB+D: A large scale dataset for 3d human activity analysis. In CVPR."},{"key":"e_1_3_2_176_2","volume-title":"ACL","author":"Sharma Piyush","year":"2018","unstructured":"Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In ACL."},{"key":"e_1_3_2_177_2","volume-title":"NAACL","author":"Shaw Peter","year":"2018","unstructured":"Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. 2018. Self-attention with relative position representations. In NAACL."},{"key":"e_1_3_2_178_2","volume-title":"ECCV","author":"Sigurdsson Gunnar A.","year":"2016","unstructured":"Gunnar A. Sigurdsson, G\u00fcl Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta. 2016. Hollywood in homes: Crowdsourcing data collection for activity understanding. In ECCV."},{"key":"e_1_3_2_179_2","unstructured":"David R. So Chen Liang and Quoc V. Le. 2019. The evolved transformer. arxiv:cs.LG\/1901.11117. Retrieved from https:\/\/arxiv.org\/abs\/1901.11117."},{"key":"e_1_3_2_180_2","article-title":"UCF101: A dataset of 101 human actions classes from videos in the wild","author":"Soomro Khurram","year":"2012","unstructured":"Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. 2012. UCF101: A dataset of 101 human actions classes from videos in the wild. arXiv:1212.0402. Retrieved from https:\/\/arxiv.org\/abs\/1212.0402.","journal-title":"arXiv:1212.0402"},{"key":"e_1_3_2_181_2","doi-asserted-by":"crossref","unstructured":"Robin Strudel Ricardo Garcia Ivan Laptev and Cordelia Schmid. 2021. Segmenter: Transformer for semantic segmentation. arxiv:cs.CV\/2105.05633. Retrieved from https:\/\/arxiv.org\/abs\/2105.05633.","DOI":"10.1109\/ICCV48922.2021.00717"},{"key":"e_1_3_2_182_2","article-title":"VL-BERT: Pre-training of generic visual-linguistic representations","author":"Su Weijie","year":"2019","unstructured":"Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2019. VL-BERT: Pre-training of generic visual-linguistic representations. arXiv:1908.08530. Retrieved from https:\/\/arxiv.org\/abs\/1908.08530.","journal-title":"arXiv:1908.08530"},{"key":"e_1_3_2_183_2","article-title":"A corpus for reasoning about natural language grounded in photographs","author":"Suhr Alane","year":"2018","unstructured":"Alane Suhr, Stephanie Zhou, Ally Zhang, Iris Zhang, Huajun Bai, and Yoav Artzi. 2018. A corpus for reasoning about natural language grounded in photographs. arXiv:1811.00491. Retrieved from https:\/\/arxiv.org\/abs\/1811.00491.","journal-title":"arXiv:1811.00491"},{"key":"e_1_3_2_184_2","article-title":"Learning video representations using contrastive bidirectional transformer","author":"Sun Chen","year":"2019","unstructured":"Chen Sun, Fabien Baradel, Kevin Murphy, and Cordelia Schmid. 2019. Learning video representations using contrastive bidirectional transformer. arXiv:1906.05743. Retrieved from https:\/\/arxiv.org\/abs\/1906.05743.","journal-title":"arXiv:1906.05743"},{"key":"e_1_3_2_185_2","volume-title":"ICCV","author":"Sun Chen","year":"2019","unstructured":"Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid. 2019. VideoBERT: A joint model for video and language representation learning. In ICCV."},{"key":"e_1_3_2_186_2","first-page":"2818","volume-title":"CVPR","author":"Szegedy Christian","year":"2016","unstructured":"Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2016. Rethinking the inception architecture for computer vision. In CVPR. 2818\u20132826."},{"key":"e_1_3_2_187_2","article-title":"Intriguing properties of neural networks","author":"Szegedy Christian","year":"2013","unstructured":"Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. 2013. Intriguing properties of neural networks. arXiv:1312.6199. Retrieved from https:\/\/arxiv.org\/abs\/1312.6199.","journal-title":"arXiv:1312.6199"},{"key":"e_1_3_2_188_2","volume-title":"CVPR","author":"Tai Ying","year":"2017","unstructured":"Ying Tai, Jian Yang, and Xiaoming Liu. 2017. Image super-resolution via deep recursive residual network. In CVPR."},{"key":"e_1_3_2_189_2","volume-title":"EMNLP-IJCNLP","author":"Tan Hao","year":"2019","unstructured":"Hao Tan and Mohit Bansal. 2019. LXMERT: Learning cross-modality encoder representations from transformers. In EMNLP-IJCNLP."},{"key":"e_1_3_2_190_2","volume-title":"EMNLP","author":"Tan Hao","year":"2020","unstructured":"Hao Tan and Mohit Bansal. 2020. Vokenization: Improving language understanding with contextualized, visual-grounded supervision. In EMNLP."},{"key":"e_1_3_2_191_2","volume-title":"ICML","author":"Tan Mingxing","year":"2019","unstructured":"Mingxing Tan and Quoc V. Le. 2019. EfficientNet: Rethinking model scaling for convolutional neural networks. In ICML."},{"key":"e_1_3_2_192_2","volume-title":"ICML","author":"Tay Y.","year":"2021","unstructured":"Y. Tay, D. Bahri, D. Metzler, D. Juan, Z. Zhao, and C. Zheng. 2021. Synthesizer: Rethinking self-attention in transformer models. In ICML."},{"key":"e_1_3_2_193_2","volume-title":"ICML","author":"Tay Yi","year":"2020","unstructured":"Yi Tay, Dara Bahri, Liu Yang, Donald Metzler, and Da-Cheng Juan. 2020. Sparse sinkhorn attention. In ICML."},{"key":"e_1_3_2_194_2","article-title":"Efficient transformers: A survey","author":"Tay Yi","year":"2020","unstructured":"Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. 2020. Efficient transformers: A survey. arXiv:2009.06732. Retrieved from https:\/\/arxiv.org\/abs\/2009.06732.","journal-title":"arXiv:2009.06732"},{"key":"e_1_3_2_195_2","article-title":"Contrastive multiview coding","author":"Tian Yonglong","year":"2019","unstructured":"Yonglong Tian, Dilip Krishnan, and Phillip Isola. 2019. Contrastive multiview coding. arXiv:1906.05849. Retrieved from https:\/\/arxiv.org\/abs\/1906.05849.","journal-title":"arXiv:1906.05849"},{"key":"e_1_3_2_196_2","volume-title":"NIPS","author":"Tolstikhin Ilya","year":"2021","unstructured":"Ilya Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Peter Steiner, Daniel Keysers, Jakob Uszkoreit, et\u00a0al. 2021. Mlp-mixer: An all-mlp architecture for vision. In NIPS."},{"key":"e_1_3_2_197_2","article-title":"Resmlp: Feedforward networks for image classification with data-efficient training","author":"Touvron Hugo","year":"2021","unstructured":"Hugo Touvron, Piotr Bojanowski, Mathilde Caron, Matthieu Cord, Alaaeldin El-Nouby, Edouard Grave, Gautier Izacard, Armand Joulin, Gabriel Synnaeve, Jakob Verbeek, et\u00a0al. 2021. Resmlp: Feedforward networks for image classification with data-efficient training. arXiv:2105.03404. Retrieved from https:\/\/arxiv.org\/abs\/2105.03404.","journal-title":"arXiv:2105.03404"},{"key":"e_1_3_2_198_2","article-title":"Training data-efficient image transformers & distillation through attention","author":"Touvron Hugo","year":"2020","unstructured":"Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herv\u00e9 J\u00e9gou. 2020. Training data-efficient image transformers & distillation through attention. arXiv:2012.12877. Retrieved from https:\/\/arxiv.org\/abs\/2012.12877.","journal-title":"arXiv:2012.12877"},{"key":"e_1_3_2_199_2","article-title":"Going deeper with image transformers","author":"Touvron Hugo","year":"2021","unstructured":"Hugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve, and Herv\u00e9 J\u00e9gou. 2021. Going deeper with image transformers. arXiv:2103.17239. Retrieved from https:\/\/arxiv.org\/abs\/2103.17239.","journal-title":"arXiv:2103.17239"},{"key":"e_1_3_2_200_2","volume-title":"NeurIPS","author":"Oord Aaron Van den","year":"2016","unstructured":"Aaron Van den Oord, Nal Kalchbrenner, Lasse Espeholt, Oriol Vinyals, Alex Graves, et\u00a0al. 2016. Conditional image generation with pixelcnn decoders. In NeurIPS."},{"key":"e_1_3_2_201_2","first-page":"12894","volume-title":"CVPR","author":"Vaswani Ashish","year":"2021","unstructured":"Ashish Vaswani, Prajit Ramachandran, Aravind Srinivas, Niki Parmar, Blake Hechtman, and Jonathon Shlens. 2021. Scaling local self-attention for parameter efficient visual backbones. In CVPR. 12894\u201312904."},{"key":"e_1_3_2_202_2","volume-title":"NeurIPS","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In NeurIPS."},{"key":"e_1_3_2_203_1","volume-title":"CVPR","author":"Vinyals Oriol","year":"2015","unstructured":"Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015. Show and tell: A neural image caption generator. In CVPR."},{"key":"e_1_3_2_204_2","article-title":"Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned","author":"Voita Elena","year":"2019","unstructured":"Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019. Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned. arXiv:1905.09418. Retrieved from https:\/\/arxiv.org\/abs\/1905.09418.","journal-title":"arXiv:1905.09418"},{"key":"e_1_3_2_205_2","doi-asserted-by":"crossref","unstructured":"Hanrui Wang Zhanghao Wu Zhijian Liu Han Cai Ligeng Zhu Chuang Gan and Song Han. 2020. HAT: Hardware-aware transformers for efficient natural language processing. arxiv:cs.CL\/2005.14187. Retrieved from https:\/\/arxiv.org\/abs\/2005.14187.","DOI":"10.18653\/v1\/2020.acl-main.686"},{"key":"e_1_3_2_206_2","article-title":"Axial-DeepLab: Stand-alone axial-attention for panoptic segmentation","author":"Wang Huiyu","year":"2020","unstructured":"Huiyu Wang, Yukun Zhu, Bradley Green, Hartwig Adam, Alan Yuille, and Liang-Chieh Chen. 2020. Axial-DeepLab: Stand-alone axial-attention for panoptic segmentation. arXiv:2003.07853. Retrieved from https:\/\/arxiv.org\/abs\/2003.07853.","journal-title":"arXiv:2003.07853"},{"key":"e_1_3_2_207_2","article-title":"Long-short temporal contrastive learning of video transformers","author":"Wang Jue","year":"2021","unstructured":"Jue Wang, Gedas Bertasius, Du Tran, and Lorenzo Torresani. 2021. Long-short temporal contrastive learning of video transformers. arXiv:2106.09212. Retrieved from https:\/\/arxiv.org\/abs\/2106.09212.","journal-title":"arXiv:2106.09212"},{"key":"e_1_3_2_208_2","article-title":"Linformer: Self-attention with linear complexity","author":"Wang Sinong","year":"2020","unstructured":"Sinong Wang, Belinda Li, Madian Khabsa, Han Fang, and Hao Ma. 2020. Linformer: Self-attention with linear complexity. arXiv:2006.04768. Retrieved from https:\/\/arxiv.org\/abs\/2006.04768.","journal-title":"arXiv:2006.04768"},{"key":"e_1_3_2_209_2","unstructured":"Wenhai Wang Enze Xie Xiang Li Deng-Ping Fan Kaitao Song Ding Liang Tong Lu Ping Luo and Ling Shao. 2021. PVTv2: Improved baselines with pyramid vision transformer. arxiv:cs.CV\/2106.13797. Retrieved from https:\/\/arxiv.org\/abs\/2106.13797."},{"key":"e_1_3_2_210_2","article-title":"Pyramid vision transformer: A versatile backbone for dense prediction without convolutions","author":"Wang Wenhai","year":"2021","unstructured":"Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. 2021. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. arXiv:2102.12122. Retrieved from https:\/\/arxiv.org\/abs\/2102.12122.","journal-title":"arXiv:2102.12122"},{"key":"e_1_3_2_211_2","article-title":"CrossFormer: A versatile vision transformer based on cross-scale attention","author":"Wang Wenxiao","year":"2021","unstructured":"Wenxiao Wang, Lu Yao, Long Chen, Deng Cai, Xiaofei He, and Wei Liu. 2021. CrossFormer: A versatile vision transformer based on cross-scale attention. arXiv:2108.00154. Retrieved from https:\/\/arxiv.org\/abs\/2108.00154.","journal-title":"arXiv:2108.00154"},{"key":"e_1_3_2_212_2","volume-title":"CVPR","author":"Wang Xiaolong","year":"2018","unstructured":"Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. 2018. Non-local neural networks. In CVPR."},{"key":"e_1_3_2_213_2","article-title":"SceneFormer: Indoor scene generation with transformers","author":"Wang Xinpeng","year":"2020","unstructured":"Xinpeng Wang, Chandan Yeshwanth, and Matthias Nie\u00dfner. 2020. SceneFormer: Indoor scene generation with transformers. arXiv:2012.09793. Retreived from https:\/\/arxiv.org\/abs\/2012.09793.","journal-title":"arXiv:2012.09793"},{"key":"e_1_3_2_214_2","volume-title":"ECCVW","author":"Wang Xintao","year":"2018","unstructured":"Xintao Wang, Ke Yu, Shixiang Wu, Jinjin Gu, Yihao Liu, Chao Dong, Yu Qiao, and Chen Change Loy. 2018. ESRGAN: Enhanced super-resolution generative adversarial networks. In ECCVW."},{"key":"e_1_3_2_215_2","article-title":"End-to-end video instance segmentation with transformers","author":"Wang Yuqing","year":"2020","unstructured":"Yuqing Wang, Zhaoliang Xu, Xinlong Wang, Chunhua Shen, Baoshan Cheng, Hao Shen, and Huaxia Xia. 2020. End-to-end video instance segmentation with transformers. arXiv:2011.14503. Retrieved from https:\/\/arxiv.org\/abs\/2011.14503.","journal-title":"arXiv:2011.14503"},{"key":"e_1_3_2_216_2","unstructured":"Yingming Wang Xiangyu Zhang Tong Yang and Jian Sun. 2021. Anchor DETR: Query design for transformer-based detector. arxiv:cs.CV\/2109.07107. Retrieved from https:\/\/arxiv.org\/abs\/2109.07107."},{"key":"e_1_3_2_217_2","article-title":"Uformer: A general U-shaped transformer for image restoration","author":"Wang Zhendong","year":"2021","unstructured":"Zhendong Wang, Xiaodong Cun, Jianmin Bao, and Jianzhuang Liu. 2021. Uformer: A general U-shaped transformer for image restoration. arXiv:2106.03106. Retrieved from https:\/\/arxiv.org\/abs\/2106.03106.","journal-title":"arXiv:2106.03106"},{"key":"e_1_3_2_218_2","volume-title":"CVPR","author":"Wei Donglai","year":"2018","unstructured":"Donglai Wei, Joseph J. Lim, Andrew Zisserman, and William T. Freeman. 2018. Learning and using the arrow of time. In CVPR."},{"key":"e_1_3_2_219_2","article-title":"Cvt: Introducing convolutions to vision transformers","author":"Wu Haiping","year":"2021","unstructured":"Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyang Dai, Lu Yuan, and Lei Zhang. 2021. Cvt: Introducing convolutions to vision transformers. arXiv:2103.15808. Retrieved from https:\/\/arxiv.org\/abs\/2103.15808.","journal-title":"arXiv:2103.15808"},{"key":"e_1_3_2_220_2","first-page":"14618","volume-title":"ICCV","author":"Wu Xiaolei","year":"2021","unstructured":"Xiaolei Wu, Zhihao Hu, Lu Sheng, and Dong Xu. 2021. StyleFormer: Real-time arbitrary style transfer via parametric style composition. In ICCV. 14618\u201314627."},{"key":"e_1_3_2_221_2","article-title":"P2T: Pyramid pooling transformer for scene understanding","author":"Wu Yu-Huan","year":"2021","unstructured":"Yu-Huan Wu, Yun Liu, Xin Zhan, and Ming-Ming Cheng. 2021. P2T: Pyramid pooling transformer for scene understanding. arXiv:2106.12011. Retrieved from https:\/\/arxiv.org\/abs\/2106.12011.","journal-title":"arXiv:2106.12011"},{"key":"e_1_3_2_222_2","unstructured":"Enze Xie Wenhai Wang Zhiding Yu Anima Anandkumar Jose M. Alvarez and Ping Luo. 2021. SegFormer: Simple and efficient design for semantic segmentation with transformers. arxiv:cs.CV\/2105.15203. Retrieved from https:\/\/arxiv.org\/abs\/2105.15203."},{"key":"e_1_3_2_223_2","volume-title":"CVPR","author":"Xie Saining","year":"2017","unstructured":"Saining Xie, Ross Girshick, Piotr Doll\u00e1r, Zhuowen Tu, and Kaiming He. 2017. Aggregated residual transformations for deep neural networks. In CVPR."},{"key":"e_1_3_2_224_2","article-title":"Self-supervised learning with swin transformers","author":"Xie Zhenda","year":"2021","unstructured":"Zhenda Xie, Yutong Lin, Zhuliang Yao, Zheng Zhang, Qi Dai, Yue Cao, and Han Hu. 2021. Self-supervised learning with swin transformers. arXiv:2105.04553. Retrieved from https:\/\/arxiv.org\/abs\/2105.04553.","journal-title":"arXiv:2105.04553"},{"key":"e_1_3_2_225_2","volume-title":"AAAI","author":"Xiong Yunyang","year":"2021","unstructured":"Yunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan, Glenn Fung, Yin Li, and Vikas Singh. 2021. Nystr\u00f6mformer: A Nyst\u00f6m-based Algorithm for Approximating Self-Attention. In AAAI."},{"key":"e_1_3_2_226_2","volume-title":"CVPR","author":"Xu Tao","year":"2018","unstructured":"Tao Xu, Pengchuan Zhang, Qiuyuan Huang, Han Zhang, Zhe Gan, Xiaolei Huang, and Xiaodong He. 2018. AttnGAN: Fine-grained text to image generation with attentional generative adversarial networks. In CVPR."},{"key":"e_1_3_2_227_2","unstructured":"Weijian Xu Yifan Xu Tyler Chang and Zhuowen Tu. 2021. Co-Scale conv-attentional image transformers. arxiv:cs.CV\/2104.06399. Retrieved from https:\/\/arxiv.org\/abs\/2104.06399."},{"key":"e_1_3_2_228_2","volume-title":"CVPR","author":"Yang Fuzhi","year":"2020","unstructured":"Fuzhi Yang, Huan Yang, Jianlong Fu, Hongtao Lu, and B. Guo. 2020. Learning texture transformer network for image super-resolution. In CVPR."},{"key":"e_1_3_2_229_2","unstructured":"Jianwei Yang Chunyuan Li Pengchuan Zhang Xiyang Dai Bin Xiao Lu Yuan and Jianfeng Gao. 2021. Focal self-attention for local-global interactions in vision transformers. arxiv:cs.CV\/2107.00641. Retrieved from https:\/\/arxiv.org\/abs\/2107.00641."},{"key":"e_1_3_2_230_2","first-page":"5188","volume-title":"ICCV","author":"Yang Linjie","year":"2019","unstructured":"Linjie Yang, Yuchen Fan, and Ning Xu. 2019. Video instance segmentation. In ICCV. 5188\u20135197."},{"key":"e_1_3_2_231_2","volume-title":"CVPR","author":"Ye Han-Jia","year":"2020","unstructured":"Han-Jia Ye, Hexiang Hu, De-Chuan Zhan, and Fei Sha. 2020. Few-shot learning via embedding adaptation with set-to-set functions. In CVPR."},{"key":"e_1_3_2_232_2","volume-title":"CVPR","author":"Ye Linwei","year":"2019","unstructured":"Linwei Ye, Mrigank Rochan, Zhi Liu, and Yang Wang. 2019. Cross-modal self-attention network for referring image segmentation. In CVPR."},{"key":"e_1_3_2_233_2","volume-title":"ECCV","author":"Yu Licheng","year":"2016","unstructured":"Licheng Yu, Patrick Poirson, Shan Yang, Alexander C. Berg, and Tamara L. Berg. 2016. Modeling context in referring expressions. In ECCV."},{"key":"e_1_3_2_234_2","article-title":"Incorporating convolution designs into visual transformers","author":"Yuan Kun","year":"2021","unstructured":"Kun Yuan, Shaopeng Guo, Ziwei Liu, Aojun Zhou, Fengwei Yu, and Wei Wu. 2021. Incorporating convolution designs into visual transformers. arXiv:2103.11816. Retrieved from https:\/\/arxiv.org\/abs\/2103.11816.","journal-title":"arXiv:2103.11816"},{"key":"e_1_3_2_235_2","article-title":"Tokens-to-token ViT: Training vision transformers from scratch on imagenet","author":"Yuan Li","year":"2021","unstructured":"Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Francis E. H. Tay, Jiashi Feng, and Shuicheng Yan. 2021. Tokens-to-token ViT: Training vision transformers from scratch on imagenet. arXiv:2101.11986. Retrieved from https:\/\/arxiv.org\/abs\/2101.11986.","journal-title":"arXiv:2101.11986"},{"key":"e_1_3_2_236_2","first-page":"6023","volume-title":"ICCV","author":"Yun Sangdoo","year":"2019","unstructured":"Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo. 2019. Cutmix: Regularization strategy to train strong classifiers with localizable features. In ICCV. 6023\u20136032."},{"key":"e_1_3_2_237_2","article-title":"Restormer: Efficient transformer for high-resolution image restoration","author":"Zamir Syed Waqas","year":"2021","unstructured":"Syed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat, Fahad Shahbaz Khan, and Ming-Hsuan Yang. 2021. Restormer: Efficient transformer for high-resolution image restoration. arXiv:2111.09881. Retrieved from https:\/\/arxiv.org\/abs\/2111.09881.","journal-title":"arXiv:2111.09881"},{"key":"e_1_3_2_238_2","volume-title":"CVPR","author":"Zellers Rowan","year":"2019","unstructured":"Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. From recognition to cognition: Visual commonsense reasoning. In CVPR."},{"key":"e_1_3_2_239_2","unstructured":"Xiaohua Zhai Alexander Kolesnikov Neil Houlsby and Lucas Beyer. 2021. Scaling vision transformers. arxiv:cs.CV\/2106.04560. Retrieved from https:\/\/arxiv.org\/abs\/2106.04560."},{"key":"e_1_3_2_240_2","first-page":"7354","volume-title":"ICML","author":"Zhang Han","year":"2019","unstructured":"Han Zhang, Ian Goodfellow, Dimitris Metaxas, and Augustus Odena. 2019. Self-attention generative adversarial networks. In ICML. PMLR, 7354\u20137363."},{"key":"e_1_3_2_241_2","volume-title":"ICCV","author":"Zhang Han","year":"2017","unstructured":"Han Zhang, Tao Xu, Hongsheng Li, Shaoting Zhang, Xiaogang Wang, Xiaolei Huang, and Dimitris N. Metaxas. 2017. StackGAN: Text to photo-realistic image synthesis with stacked generative adversarial networks. In ICCV."},{"key":"e_1_3_2_242_2","article-title":"StackGAN++: Realistic image synthesis with stacked generative adversarial networks","author":"Zhang Han","year":"2018","unstructured":"Han Zhang, Tao Xu, Hongsheng Li, Shaoting Zhang, Xiaogang Wang, Xiaolei Huang, and Dimitris N Metaxas. 2018. StackGAN++: Realistic image synthesis with stacked generative adversarial networks. IEEE Trans. Pattern Anal. Mach. Intell. (2018).","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"e_1_3_2_243_2","article-title":"Multi-scale vision longformer: A new vision transformer for high-resolution image encoding","author":"Zhang Pengchuan","year":"2021","unstructured":"Pengchuan Zhang, Xiyang Dai, Jianwei Yang, Bin Xiao, Lu Yuan, Lei Zhang, and Jianfeng Gao. 2021. Multi-scale vision longformer: A new vision transformer for high-resolution image encoding. In ICCV (2021).","journal-title":"ICCV"},{"key":"e_1_3_2_244_2","first-page":"5579","volume-title":"CVPR","author":"Zhang Pengchuan","year":"2021","unstructured":"Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao. 2021. Vinvl: Revisiting visual representations in vision-language models. In CVPR. 5579\u20135588."},{"key":"e_1_3_2_245_2","article-title":"ResT: An efficient transformer for visual recognition","author":"Zhang Qinglong","year":"2021","unstructured":"Qinglong Zhang and Yubin Yang. 2021. ResT: An efficient transformer for visual recognition. arXiv:2105.13677. Retrieved from https:\/\/arxiv.org\/abs\/2105.13677.","journal-title":"arXiv:2105.13677"},{"key":"e_1_3_2_246_2","volume-title":"ECCV","author":"Zhang Richard","year":"2016","unstructured":"Richard Zhang, Phillip Isola, and Alexei A. Efros. 2016. Colorful image colorization. In ECCV."},{"key":"e_1_3_2_247_2","volume-title":"ECCV","author":"Zhang Yulun","year":"2018","unstructured":"Yulun Zhang, Kunpeng Li, Kai Li, Lichen Wang, Bineng Zhong, and Yun Fu. 2018. Image super-resolution using very deep residual channel attention networks. In ECCV."},{"key":"e_1_3_2_248_2","article-title":"Residual dense network for image restoration","author":"Zhang Yulun","year":"2020","unstructured":"Yulun Zhang, Yapeng Tian, Yu Kong, Bineng Zhong, and Yun Fu. 2020. Residual dense network for image restoration. IEEE Trans. Pattern Anal. Mach. Intell. (2020).","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"e_1_3_2_249_2","volume-title":"arXiv:2105.12723","author":"Zhang Zizhao","year":"2021","unstructured":"Zizhao Zhang, Han Zhang, Long Zhao, Ting Chen, and Tomas Pfister. 2021. Aggregating nested transformers. In arXiv:2105.12723. Retrieved from https:\/\/arxiv.org\/abs\/2105.12723."},{"key":"e_1_3_2_250_2","volume-title":"CVPR","author":"Zhao Hengshuang","year":"2020","unstructured":"Hengshuang Zhao, Jiaya Jia, and Vladlen Koltun. 2020. Exploring self-attention for image recognition. In CVPR."},{"key":"e_1_3_2_251_2","doi-asserted-by":"crossref","unstructured":"Sixiao Zheng Jiachen Lu Hengshuang Zhao Xiatian Zhu Zekun Luo Yabiao Wang Yanwei Fu Jianfeng Feng Tao Xiang Philip H. S. Torr and Li Zhang. 2021. Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. arxiv:cs.CV\/2012.15840. Retrieved from https:\/\/arxiv.org\/abs\/2012.15840.","DOI":"10.1109\/CVPR46437.2021.00681"},{"key":"e_1_3_2_252_2","volume-title":"CVPR","author":"Zhou Bolei","year":"2017","unstructured":"Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. 2017. Scene parsing through ade20k dataset. In CVPR."},{"key":"e_1_3_2_253_2","unstructured":"Daquan Zhou Bingyi Kang Xiaojie Jin Linjie Yang Xiaochen Lian Zihang Jiang Qibin Hou and Jiashi Feng. 2021. DeepViT: Towards deeper vision transformer. arxiv:cs.CV\/2103.11886. Retrieved from https:\/\/arxiv.org\/abs\/2103.11886."},{"key":"e_1_3_2_254_2","first-page":"13041","volume-title":"AAAI","author":"Zhou Luowei","year":"2020","unstructured":"Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason Corso, and Jianfeng Gao. 2020. Unified vision-language pre-training for image captioning and vqa. In AAAI, Vol. 34. 13041\u201313049."},{"key":"e_1_3_2_255_2","volume-title":"AAAI","author":"Zhou Luowei","year":"2018","unstructured":"Luowei Zhou, Chenliang Xu, and Jason Corso. 2018. Towards automatic learning of procedures from web instructional videos. In AAAI, Vol. 32."},{"key":"e_1_3_2_256_2","volume-title":"CVPR","author":"Zhou Luowei","year":"2018","unstructured":"Luowei Zhou, Yingbo Zhou, Jason Corso, R. Socher, and C. Xiong. 2018. End-to-end dense video captioning with masked transformer. In CVPR."},{"key":"e_1_3_2_257_2","article-title":"Deformable DETR: Deformable transformers for end-to-end object detection","author":"Zhu Xizhou","year":"2020","unstructured":"Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. 2020. Deformable DETR: Deformable transformers for end-to-end object detection. arXiv:2010.04159. Retrieved from https:\/\/arxiv.org\/abs\/2010.04159.","journal-title":"arXiv:2010.04159"}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3505244","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3505244","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:31:26Z","timestamp":1750188686000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3505244"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,1,31]]},"references-count":256,"journal-issue":{"issue":"10s","published-print":{"date-parts":[[2022,1,31]]}},"alternative-id":["10.1145\/3505244"],"URL":"https:\/\/doi.org\/10.1145\/3505244","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,1,31]]},"assertion":[{"value":"2021-03-02","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-12-07","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-09-13","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}