{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,4]],"date-time":"2026-07-04T16:56:02Z","timestamp":1783184162469,"version":"3.54.6"},"publisher-location":"New York, NY, USA","reference-count":28,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,3,9]],"date-time":"2022-03-09T00:00:00Z","timestamp":1646784000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,3,9]]},"DOI":"10.1145\/3508396.3512869","type":"proceedings-article","created":{"date-parts":[[2022,3,5]],"date-time":"2022-03-05T05:07:15Z","timestamp":1646456835000},"page":"1-7","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":48,"title":["Towards efficient vision transformer inference"],"prefix":"10.1145","author":[{"given":"Xudong","family":"Wang","sequence":"first","affiliation":[{"name":"Shanghai Jiao Tong University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Li Lyna","family":"Zhang","sequence":"additional","affiliation":[{"name":"Microsoft Research"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yang","family":"Wang","sequence":"additional","affiliation":[{"name":"Microsoft Research"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mao","family":"Yang","sequence":"additional","affiliation":[{"name":"Microsoft Research"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,3,9]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"AutoGTCO: Graph and Tensor Co-Optimize for Image Recognition with Transformers on GPU. In 2021 IEEE\/ACM International Conference On Computer Aided Design (ICCAD). IEEE, 1--9.","author":"Bai Yang","year":"2021","unstructured":"Yang Bai , Xufeng Yao , Qi Sun , and Bei Yu . 2021 . AutoGTCO: Graph and Tensor Co-Optimize for Image Recognition with Transformers on GPU. In 2021 IEEE\/ACM International Conference On Computer Aided Design (ICCAD). IEEE, 1--9. Yang Bai, Xufeng Yao, Qi Sun, and Bei Yu. 2021. AutoGTCO: Graph and Tensor Co-Optimize for Image Recognition with Transformers on GPU. In 2021 IEEE\/ACM International Conference On Computer Aided Design (ICCAD). IEEE, 1--9."},{"key":"e_1_3_2_1_2_1","volume-title":"Longformer: The Long-Document Transformer. arXiv:2004.05150","author":"Beltagy Iz","year":"2020","unstructured":"Iz Beltagy , Matthew E. Peters , and Arman Cohan . 2020 . Longformer: The Long-Document Transformer. arXiv:2004.05150 (2020). Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020. Longformer: The Long-Document Transformer. arXiv:2004.05150 (2020)."},{"key":"e_1_3_2_1_3_1","volume-title":"TVM: An Automated End-to-End Optimizing Compiler for Deep Learning. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18)","author":"Chen Tianqi","year":"2018","unstructured":"Tianqi Chen , Thierry Moreau , Ziheng Jiang , Lianmin Zheng , Eddie Yan , Haichen Shen , Meghan Cowan , Leyuan Wang , Yuwei Hu , Luis Ceze , Carlos Guestrin , and Arvind Krishnamurthy . 2018 . TVM: An Automated End-to-End Optimizing Compiler for Deep Learning. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18) . USENIX Association, Carlsbad, CA, 578--594. https:\/\/www.usenix.org\/conference\/osdi18\/presentation\/chen Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018. TVM: An Automated End-to-End Optimizing Compiler for Deep Learning. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18). USENIX Association, Carlsbad, CA, 578--594. https:\/\/www.usenix.org\/conference\/osdi18\/presentation\/chen"},{"key":"e_1_3_2_1_4_1","unstructured":"The TFLite developers. 2021. TFLite Model Benchmark Tool with C++ Binary. https:\/\/github.com\/tensorflow\/tensorflow\/tree\/r2.7\/tensorflow\/lite\/tools\/benchmark.  The TFLite developers. 2021. TFLite Model Benchmark Tool with C++ Binary. https:\/\/github.com\/tensorflow\/tensorflow\/tree\/r2.7\/tensorflow\/lite\/tools\/benchmark."},{"key":"e_1_3_2_1_5_1","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019","volume":"1","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2019 . BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding . In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019 , Minneapolis, MN, USA , June 2-7, 2019, Volume 1 (Long and Short Papers). 4171--4186. Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers). 4171--4186."},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01204"},{"key":"e_1_3_2_1_7_1","volume-title":"Trained Quantization and Huffman Coding. In International Conference on Learning Representations (ICLR).","author":"Han Song","unstructured":"Song Han , Huizi Mao , and William J. Dally . 2016. Deep Compression: Compressing Deep Neural Networks with Pruning , Trained Quantization and Huffman Coding. In International Conference on Learning Representations (ICLR). Song Han, Huizi Mao, and William J. Dally. 2016. Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding. In International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_1_8_1","volume-title":"AMC: AutoML for Model Compression and Acceleration on Mobile Devices. In European Conference on Computer Vision (ECCV).","author":"He Yihui","year":"2018","unstructured":"Yihui He , Ji Lin , Zhijian Liu , Hanrui Wang , Li-Jia Li , and Song Han . 2018 . AMC: AutoML for Model Compression and Acceleration on Mobile Devices. In European Conference on Computer Vision (ECCV). Yihui He, Ji Lin, Zhijian Liu, Hanrui Wang, Li-Jia Li, and Song Han. 2018. AMC: AutoML for Model Compression and Acceleration on Mobile Devices. In European Conference on Computer Vision (ECCV)."},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00140"},{"key":"e_1_3_2_1_10_1","unstructured":"Intel. 2021. Deploy High-Performance Deep Learning Inference OpenVINO. https:\/\/software.intel.com\/content\/www\/us\/en\/develop\/tools\/openvino-toolkit.html.  Intel. 2021. Deploy High-Performance Deep Learning Inference OpenVINO. https:\/\/software.intel.com\/content\/www\/us\/en\/develop\/tools\/openvino-toolkit.html."},{"key":"e_1_3_2_1_11_1","volume-title":"The International Conference on Learning Representations (ICLR).","author":"Kolesnikov Alexander","year":"2021","unstructured":"Alexander Kolesnikov , Alexey Dosovitskiy , Dirk Weissenborn , Georg Heigold , Jakob Uszkoreit , Lucas Beyer , Matthias Minderer , Mostafa Dehghani , Neil Houlsby , Sylvain Gelly , Thomas Unterthiner , and Xiaohua Zhai . 2021 . An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale . In The International Conference on Learning Representations (ICLR). Alexander Kolesnikov, Alexey Dosovitskiy, Dirk Weissenborn, Georg Heigold, Jakob Uszkoreit, Lucas Beyer, Matthias Minderer, Mostafa Dehghani, Neil Houlsby, Sylvain Gelly, Thomas Unterthiner, and Xiaohua Zhai. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In The International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.829"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/3446640"},{"key":"e_1_3_2_1_15_1","unstructured":"Sachin Mehta and Mohammad Rastegari. 2021. MobileViT: Light-weight General-purpose and Mobile-friendly Vision Transformer. [arxiv]2110.02178 [cs.CV]  Sachin Mehta and Mohammad Rastegari. 2021. MobileViT: Light-weight General-purpose and Mobile-friendly Vision Transformer. [arxiv]2110.02178 [cs.CV]"},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/3453483.3454083"},{"key":"e_1_3_2_1_17_1","volume-title":"Proceedings of the 38th International Conference on Machine Learning (ICML).","author":"Shi Han","unstructured":"Han Shi , Jiahui Gao , Xiaozhe Ren , Hang Xu , Xiaodan Liang , Zhenguo Li , and James T. Kwok . 2021. SparseBERT: Rethinking the Importance Analysis in Self-attention . In Proceedings of the 38th International Conference on Machine Learning (ICML). Han Shi, Jiahui Gao, Xiaozhe Ren, Hang Xu, Xiaodan Liang, Zhenguo Li, and James T. Kwok. 2021. SparseBERT: Rethinking the Importance Analysis in Self-attention. In Proceedings of the 38th International Conference on Machine Learning (ICML)."},{"key":"e_1_3_2_1_18_1","volume-title":"EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. In International Conference on Machine Learning (ICML). 6105--6114","author":"Tan Mingxing","year":"2019","unstructured":"Mingxing Tan and Quoc Le . 2019 . EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. In International Conference on Machine Learning (ICML). 6105--6114 . Mingxing Tan and Quoc Le. 2019. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. In International Conference on Machine Learning (ICML). 6105--6114."},{"key":"e_1_3_2_1_19_1","first-page":"21","article-title":"To Bridge Neural Network Design and Real-World Performance: A Behaviour Study for Neural Networks","volume":"3","author":"Tang Xiaohu","year":"2021","unstructured":"Xiaohu Tang , Shihao Han , Li Lyna Zhang , Ting Cao , and Yunxin Liu . 2021 . To Bridge Neural Network Design and Real-World Performance: A Behaviour Study for Neural Networks . In Proceedings of Machine Learning and Systems (MLSys) , Vol. 3. 21 -- 37 . Xiaohu Tang, Shihao Han, Li Lyna Zhang, Ting Cao, and Yunxin Liu. 2021. To Bridge Neural Network Design and Real-World Performance: A Behaviour Study for Neural Networks. In Proceedings of Machine Learning and Systems (MLSys), Vol. 3. 21--37.","journal-title":"Proceedings of Machine Learning and Systems (MLSys)"},{"key":"e_1_3_2_1_20_1","unstructured":"TFLite. 2021. Post-training quantization. https:\/\/www.tensorflow.org\/lite\/performance\/post_training_quant.  TFLite. 2021. Post-training quantization. https:\/\/www.tensorflow.org\/lite\/performance\/post_training_quant."},{"key":"e_1_3_2_1_21_1","volume-title":"International Conference on Machine Learning (ICML). 10347--10357","author":"Touvron Hugo","year":"2021","unstructured":"Hugo Touvron , Matthieu Cord , Matthijs Douze , Francisco Massa , Alexandre Sablayrolles , and Herve Jegou . 2021 . Training data-efficient image transformers and distillation through attention . In International Conference on Machine Learning (ICML). 10347--10357 . Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve Jegou. 2021. Training data-efficient image transformers and distillation through attention. In International Conference on Machine Learning (ICML). 10347--10357."},{"key":"e_1_3_2_1_22_1","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani , Noam Shazeer , Niki Parmar , Jakob Uszkoreit , Llion Jones , Aidan N Gomez , \u0141 ukasz Kaiser, and Illia Polosukhin . 2017 . Attention is All you Need . In Advances in Neural Information Processing Systems (NeurIPS) , Vol. 30 . Curran Associates, Inc. Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, \u0141 ukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. In Advances in Neural Information Processing Systems (NeurIPS), Vol. 30. Curran Associates, Inc."},{"key":"e_1_3_2_1_23_1","first-page":"30","article-title":"A Systematic Methodology for Analysis of Deep Learning Hardware and Software Platforms","volume":"2","author":"Wang Yu","year":"2020","unstructured":"Yu Wang , Gu-Yeon Wei , and David Brooks . 2020 . A Systematic Methodology for Analysis of Deep Learning Hardware and Software Platforms . In Proceedings of Machine Learning and Systems (MLSys) , Vol. 2. 30 -- 43 . Yu Wang, Gu-Yeon Wei, and David Brooks. 2020. A Systematic Methodology for Analysis of Deep Learning Hardware and Software Platforms. In Proceedings of Machine Learning and Systems (MLSys), Vol. 2. 30--43.","journal-title":"Proceedings of Machine Learning and Systems (MLSys)"},{"key":"e_1_3_2_1_24_1","unstructured":"Ziheng Wang. 2021. SparseDNN: Fast Sparse Deep Learning Inference on CPUs. [arxiv]2101.07948 [cs.LG]  Ziheng Wang. 2021. SparseDNN: Fast Sparse Deep Learning Inference on CPUs. [arxiv]2101.07948 [cs.LG]"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3308558.3313591"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00060"},{"key":"e_1_3_2_1_27_1","volume-title":"Big Bird: Transformers for Longer Sequences. In Advances in Neural Information Processing Systems (NeurIPS).","author":"Zaheer Manzil","year":"2020","unstructured":"Manzil Zaheer , Guru Guruganesh , Avinava Dubey , Joshua Ainslie , Chris Alberti , Santiago Ontanon , Philip Pham , Anirudh Ravula , Qifan Wang , Li Yang , and Amr Ahmed . 2020 . Big Bird: Transformers for Longer Sequences. In Advances in Neural Information Processing Systems (NeurIPS). Manzil Zaheer, Guru Guruganesh, Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, and Amr Ahmed. 2020. Big Bird: Transformers for Longer Sequences. In Advances in Neural Information Processing Systems (NeurIPS)."},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/3458864.3467882"}],"event":{"name":"HotMobile '22: The 23rd International Workshop on Mobile Computing Systems and Applications","location":"Tempe Arizona","acronym":"HotMobile '22","sponsor":["SIGMOBILE ACM Special Interest Group on Mobility of Systems, Users, Data and Computing"]},"container-title":["Proceedings of the 23rd Annual International Workshop on Mobile Computing Systems and Applications"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3508396.3512869","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3508396.3512869","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:49:36Z","timestamp":1750182576000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3508396.3512869"}},"subtitle":["a first study of transformers on mobile devices"],"short-title":[],"issued":{"date-parts":[[2022,3,9]]},"references-count":28,"alternative-id":["10.1145\/3508396.3512869","10.1145\/3508396"],"URL":"https:\/\/doi.org\/10.1145\/3508396.3512869","relation":{},"subject":[],"published":{"date-parts":[[2022,3,9]]},"assertion":[{"value":"2022-03-09","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}