{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,11]],"date-time":"2026-07-11T16:27:15Z","timestamp":1783787235313,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":108,"publisher":"ACM","license":[{"start":{"date-parts":[[2023,1,27]],"date-time":"2023-01-27T00:00:00Z","timestamp":1674777600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,1,27]]},"DOI":"10.1145\/3575693.3575747","type":"proceedings-article","created":{"date-parts":[[2023,1,30]],"date-time":"2023-01-30T22:56:55Z","timestamp":1675119415000},"page":"295-310","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":70,"title":["FLAT: An Optimized Dataflow for Mitigating Attention Bottlenecks"],"prefix":"10.1145","author":[{"given":"Sheng-Chun","family":"Kao","sequence":"first","affiliation":[{"name":"Georgia Institute of Technology, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Suvinay","family":"Subramanian","sequence":"additional","affiliation":[{"name":"Google, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Gaurav","family":"Agrawal","sequence":"additional","affiliation":[{"name":"Microsoft, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Amir","family":"Yazdanbakhsh","sequence":"additional","affiliation":[{"name":"Google, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tushar","family":"Krishna","sequence":"additional","affiliation":[{"name":"Georgia Institute of Technology, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,1,30]]},"reference":[{"key":"e_1_3_2_1_1_1","volume":"199","author":"Allen Randy","unstructured":"Randy Allen and Ken Kennedy. Vector Register Allocation. IEEE Computer Architecture Letters , 199 2. Randy Allen and Ken Kennedy. Vector Register Allocation. IEEE Computer Architecture Letters, 1992.","journal-title":"Ken Kennedy. Vector Register Allocation. IEEE Computer Architecture Letters"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2016.7783725"},{"key":"e_1_3_2_1_3_1","volume-title":"ISCA","author":"Baek Eunjin","year":"2020","unstructured":"Eunjin Baek , Dongup Kwon , and Jangwoo Kim . A Multi-Neural Network Acceleration Architecture . In ISCA , 2020 . Eunjin Baek, Dongup Kwon, and Jangwoo Kim. A Multi-Neural Network Acceleration Architecture. In ISCA, 2020."},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/CGO.2019.8661197"},{"key":"e_1_3_2_1_5_1","volume-title":"Longformer: The Long-document Transformer. arXiv preprint arXiv:2004.05150","author":"Beltagy Iz","year":"2020","unstructured":"Iz Beltagy , Matthew E Peters , and Arman Cohan . Longformer: The Long-document Transformer. arXiv preprint arXiv:2004.05150 , 2020 . Iz Beltagy, Matthew E Peters, and Arman Cohan. Longformer: The Long-document Transformer. arXiv preprint arXiv:2004.05150, 2020."},{"key":"e_1_3_2_1_6_1","volume-title":"Self-attention based context-aware 3d object detection. arXiv preprint arXiv:2101.02672","author":"Bhattacharyya Prarthana","year":"2021","unstructured":"Prarthana Bhattacharyya , Chengjie Huang , and Krzysztof Czarnecki . Self-attention based context-aware 3d object detection. arXiv preprint arXiv:2101.02672 , 2021 . Prarthana Bhattacharyya, Chengjie Huang, and Krzysztof Czarnecki. Self-attention based context-aware 3d object detection. arXiv preprint arXiv:2101.02672, 2021."},{"key":"e_1_3_2_1_7_1","volume-title":"ICML","author":"Chen Mark","year":"2020","unstructured":"Mark Chen , Alec Radford , Rewon Child , Jeffrey Wu , Heewoo Jun , David Luan , and Ilya Sutskever . Generative Pretraining from Pixels . In ICML , 2020 . Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever. Generative Pretraining from Pixels. In ICML, 2020."},{"key":"e_1_3_2_1_8_1","volume-title":"NeurIPS","author":"Chen Tianqi","year":"2018","unstructured":"Tianqi Chen , Lianmin Zheng , Eddie Yan , Ziheng Jiang , Thierry Moreau , Luis Ceze , Carlos Guestrin , and Arvind Krishnamurthy . Learning to Optimize Tensor Programs . In NeurIPS , 2018 . Tianqi Chen, Lianmin Zheng, Eddie Yan, Ziheng Jiang, Thierry Moreau, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. Learning to Optimize Tensor Programs. In NeurIPS, 2018."},{"key":"e_1_3_2_1_9_1","volume-title":"Eyeriss: An Energy-efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks. JSSC","author":"Yu-Hsin","year":"2016","unstructured":"Yu-Hsin Chen et al . Eyeriss: An Energy-efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks. JSSC , 2016 . Yu-Hsin Chen et al. Eyeriss: An Energy-efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks. JSSC, 2016."},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/JETCAS.2019.2910232"},{"key":"e_1_3_2_1_11_1","volume-title":"ISSCC","author":"Krishna Yu-Hsin","year":"2016","unstructured":"Chen, Yu-Hsin and Krishna , Tushar and Emer , Joel and Sze , Vivienne. Eyeriss : An Energy-Efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks . In ISSCC , 2016 . Chen, Yu-Hsin and Krishna, Tushar and Emer, Joel and Sze, Vivienne. Eyeriss: An Energy-Efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks. In ISSCC, 2016."},{"key":"e_1_3_2_1_12_1","volume-title":"Generating Long Sequences with Sparse Transformers. arXiv preprint arXiv:1904.10509","author":"Child Rewon","year":"2019","unstructured":"Rewon Child , Scott Gray , Alec Radford , and Ilya Sutskever . Generating Long Sequences with Sparse Transformers. arXiv preprint arXiv:1904.10509 , 2019 . Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. Generating Long Sequences with Sparse Transformers. arXiv preprint arXiv:1904.10509, 2019."},{"key":"e_1_3_2_1_13_1","volume-title":"Masked Language Modeling for Proteins via Linearly Scalable Long-context Transformers. arXiv preprint arXiv:2006.03555","author":"Choromanski Krzysztof","year":"2020","unstructured":"Krzysztof Choromanski , Valerii Likhosherstov , David Dohan , Xingyou Song , Jared Davis , Tamas Sarlos , David Belanger , Lucy Colwell , and Adrian Weller . Masked Language Modeling for Proteins via Linearly Scalable Long-context Transformers. arXiv preprint arXiv:2006.03555 , 2020 . Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Jared Davis, Tamas Sarlos, David Belanger, Lucy Colwell, and Adrian Weller. Masked Language Modeling for Proteins via Linearly Scalable Long-context Transformers. arXiv preprint arXiv:2006.03555, 2020."},{"key":"e_1_3_2_1_14_1","volume-title":"Rethinking Attention with Performers. arXiv preprint arXiv:2009.14794","author":"Choromanski Krzysztof","year":"2020","unstructured":"Krzysztof Choromanski , Valerii Likhosherstov , David Dohan , Xingyou Song , Andreea Gane , Tamas Sarlos , Peter Hawkins , Jared Davis , Afroz Mohiuddin , Lukasz Kaiser , David Belanger , Lucy Colwell , and Adrian Weller . Rethinking Attention with Performers. arXiv preprint arXiv:2009.14794 , 2020 . Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, David Belanger, Lucy Colwell, and Adrian Weller. Rethinking Attention with Performers. arXiv preprint arXiv:2009.14794, 2020."},{"key":"e_1_3_2_1_15_1","volume-title":"Nucleic Acids Research","author":"UniProt Consortium","year":"2019","unstructured":"UniProt Consortium . UniProt : A Worldwide Hub of Protein Knowledge . Nucleic Acids Research , 2019 . UniProt Consortium. UniProt: A Worldwide Hub of Protein Knowledge. Nucleic Acids Research, 2019."},{"key":"e_1_3_2_1_16_1","volume-title":"Adaptively Sparse Transformers. arXiv preprint arXiv:1909.00015","author":"Correia Gon\u00e7alo M","year":"2019","unstructured":"Gon\u00e7alo M Correia , Vlad Niculae , and Andr\u00e9 FT Martins . Adaptively Sparse Transformers. arXiv preprint arXiv:1909.00015 , 2019 . Gon\u00e7alo M Correia, Vlad Niculae, and Andr\u00e9 FT Martins. Adaptively Sparse Transformers. arXiv preprint arXiv:1909.00015, 2019."},{"key":"e_1_3_2_1_17_1","volume-title":"Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context. arXiv preprint arXiv:1901.02860","author":"Dai Zihang","year":"2019","unstructured":"Zihang Dai , Zhilin Yang , Yiming Yang , Jaime Carbonell , Quoc V Le , and Ruslan Salakhutdinov . Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context. arXiv preprint arXiv:1901.02860 , 2019 . Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context. arXiv preprint arXiv:1901.02860, 2019."},{"key":"e_1_3_2_1_18_1","volume-title":"TECS","author":"Dave Shail","year":"2019","unstructured":"Shail Dave , Youngbin Kim , Sasikanth Avancha , Kyoungwoo Lee , and Aviral Shrivastava . dMazeRunner : Executing Perfectly Nested Loops on Dataflow Accelerators . TECS , 2019 . Shail Dave, Youngbin Kim, Sasikanth Avancha, Kyoungwoo Lee, and Aviral Shrivastava. dMazeRunner: Executing Perfectly Nested Loops on Dataflow Accelerators. TECS, 2019."},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2003.09.005"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/2749469.2750389"},{"key":"e_1_3_2_1_21_1","volume-title":"Taming Transformers for High-Resolution Image Synthesis. arXiv preprint arXiv:2012.09841","author":"Esser Patrick","year":"2020","unstructured":"Patrick Esser , Robin Rombach , and Bj\u00f6rn Ommer . Taming Transformers for High-Resolution Image Synthesis. arXiv preprint arXiv:2012.09841 , 2020 . Patrick Esser, Robin Rombach, and Bj\u00f6rn Ommer. Taming Transformers for High-Resolution Image Synthesis. arXiv preprint arXiv:2012.09841, 2020."},{"key":"e_1_3_2_1_22_1","volume-title":"Radhika Thekkath. Collective Loop Fusion for Array Contraction. In International Workshop on Languages and Compilers for Parallel Computing","author":"Gao Guang","year":"1992","unstructured":"Guang Gao , Russ Olsen , Vivek Sarkar , and Radhika Thekkath. Collective Loop Fusion for Array Contraction. In International Workshop on Languages and Compilers for Parallel Computing , 1992 . Guang Gao, Russ Olsen, Vivek Sarkar, and Radhika Thekkath. Collective Loop Fusion for Array Contraction. In International Workshop on Languages and Compilers for Parallel Computing, 1992."},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/3037697.3037702"},{"key":"e_1_3_2_1_24_1","first-page":"820","volume-title":"ASPLOS","author":"Gao Mingyu","year":"2019","unstructured":"Mingyu Gao , Xuan Yang , Jing Pu , Mark Horowitz , and Christos Kozyrakis . TANGRAM : Optimized Coarse-Grained Dataflow for Scalable NN Accelerators . In ASPLOS , pages 807\u2013 820 , 2019 . Mingyu Gao, Xuan Yang, Jing Pu, Mark Horowitz, and Christos Kozyrakis. TANGRAM: Optimized Coarse-Grained Dataflow for Scalable NN Accelerators. In ASPLOS, pages 807\u2013820, 2019."},{"key":"e_1_3_2_1_25_1","volume-title":"https:\/\/coral.ai\/","author":"Coral","year":"2020","unstructured":"Google. Coral AI. https:\/\/coral.ai\/ , 2020 . Google. Coral AI. https:\/\/coral.ai\/, 2020."},{"key":"e_1_3_2_1_26_1","volume-title":"https:\/\/www.tensorflow.org\/xla","author":"TensorFlow","year":"2021","unstructured":"Google. TensorFlow XLA. https:\/\/www.tensorflow.org\/xla , 2021 . Google. TensorFlow XLA. https:\/\/www.tensorflow.org\/xla, 2021."},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01204"},{"key":"e_1_3_2_1_28_1","volume-title":"Reweighted Proximal Pruning for Large-scale Language Representation. arXiv preprint arXiv:1909.12486","author":"Guo Fu-Ming","year":"2019","unstructured":"Fu-Ming Guo , Sijia Liu , Finlay S Mungall , Xue Lin , and Yanzhi Wang . Reweighted Proximal Pruning for Large-scale Language Representation. arXiv preprint arXiv:1909.12486 , 2019 . Fu-Ming Guo, Sijia Liu, Finlay S Mungall, Xue Lin, and Yanzhi Wang. Reweighted Proximal Pruning for Large-scale Language Representation. arXiv preprint arXiv:1909.12486, 2019."},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA47549.2020.00035"},{"key":"e_1_3_2_1_30_1","volume-title":"Lightweight Self-Attention Mechanism in Neural Networks. In ISCA","author":"Ham Tae Jun","year":"2021","unstructured":"Tae Jun Ham , Yejin Lee , Seong Hoon Seo , Soosung Kim , Hyunji Choi , Sung Jun Jung , and Jae W Lee . ELSA : Hardware-Software Co-design for Efficient , Lightweight Self-Attention Mechanism in Neural Networks. In ISCA , 2021 . Tae Jun Ham, Yejin Lee, Seong Hoon Seo, Soosung Kim, Hyunji Choi, Sung Jun Jung, and Jae W Lee. ELSA: Hardware-Software Co-design for Efficient, Lightweight Self-Attention Mechanism in Neural Networks. In ISCA, 2021."},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_1_32_1","volume-title":"ASPLOS","author":"Hegde Kartik","year":"2021","unstructured":"Kartik Hegde , Po-An Tsai , Sitao Huang , Vikas Chandra , Angshuman Parashar , and Christopher W Fletcher . Mind Mappings : Enabling Efficient Algorithm-Accelerator Mapping Space Search Extended Abstract . In ASPLOS , 2021 . Kartik Hegde, Po-An Tsai, Sitao Huang, Vikas Chandra, Angshuman Parashar, and Christopher W Fletcher. Mind Mappings: Enabling Efficient Algorithm-Accelerator Mapping Space Search Extended Abstract. In ASPLOS, 2021."},{"key":"e_1_3_2_1_33_1","volume-title":"Compound Word Transformer: Learning to Compose Full-Song Music over Dynamic Directed Hypergraphs. arXiv preprint arXiv:2101.02402","author":"Hsiao Wen-Yi","year":"2021","unstructured":"Wen-Yi Hsiao , Jen-Yu Liu , Yin-Cheng Yeh , and Yi-Hsuan Yang . Compound Word Transformer: Learning to Compose Full-Song Music over Dynamic Directed Hypergraphs. arXiv preprint arXiv:2101.02402 , 2021 . Wen-Yi Hsiao, Jen-Yu Liu, Yin-Cheng Yeh, and Yi-Hsuan Yang. Compound Word Transformer: Learning to Compose Full-Song Music over Dynamic Directed Hypergraphs. arXiv preprint arXiv:2101.02402, 2021."},{"key":"e_1_3_2_1_34_1","volume-title":"ICLR","author":"Anna Huang Cheng-Zhi","year":"2018","unstructured":"Cheng-Zhi Anna Huang , Ashish Vaswani , Jakob Uszkoreit , Ian Simon , Curtis Hawthorne , Noam Shazeer , Andrew M Dai , Matthew D Hoffman , Monica Dinculescu , and Douglas Eck . Music Transformer : Generating Music with Long-Term Structure . In ICLR , 2018 . Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Ian Simon, Curtis Hawthorne, Noam Shazeer, Andrew M Dai, Matthew D Hoffman, Monica Dinculescu, and Douglas Eck. Music Transformer: Generating Music with Long-Term Structure. In ICLR, 2018."},{"key":"e_1_3_2_1_35_1","volume-title":"CoSA: Scheduling by Constrained Optimization for Spatial Accelerators. arXiv preprint arXiv:2105.01898","author":"Huang Qijing","year":"2021","unstructured":"Qijing Huang , Minwoo Kang , Grace Dinh , Thomas Norell , Aravind Kalaiah , James Demmel , John Wawrzynek , and Yakun Sophia Shao . CoSA: Scheduling by Constrained Optimization for Spatial Accelerators. arXiv preprint arXiv:2105.01898 , 2021 . Qijing Huang, Minwoo Kang, Grace Dinh, Thomas Norell, Aravind Kalaiah, James Demmel, John Wawrzynek, and Yakun Sophia Shao. CoSA: Scheduling by Constrained Optimization for Spatial Accelerators. arXiv preprint arXiv:2105.01898, 2021."},{"key":"e_1_3_2_1_36_1","volume-title":"MLSys","author":"Ivanov Andrei","year":"2021","unstructured":"Andrei Ivanov , Nikoli Dryden , Tal Ben-Nun , Shigang Li , and Torsten Hoefler . Data Movement Is All You Need: A Case Study on Optimizing Transformers . In MLSys , 2021 . Andrei Ivanov, Nikoli Dryden, Tal Ben-Nun, Shigang Li, and Torsten Hoefler. Data Movement Is All You Need: A Case Study on Optimizing Transformers. In MLSys, 2021."},{"key":"e_1_3_2_1_37_1","volume-title":"Beyond Data and Model Parallelism for Deep Neural Networks. arXiv preprint arXiv:1807.05358","author":"Jia Zhihao","year":"2018","unstructured":"Zhihao Jia , Matei Zaharia , and Alex Aiken . Beyond Data and Model Parallelism for Deep Neural Networks. arXiv preprint arXiv:1807.05358 , 2018 . Zhihao Jia, Matei Zaharia, and Alex Aiken. Beyond Data and Model Parallelism for Deep Neural Networks. arXiv preprint arXiv:1807.05358, 2018."},{"key":"e_1_3_2_1_38_1","volume-title":"TinyBERT: Distilling BERT for Natural Language Understanding. arXiv preprint arXiv:1909.10351","author":"Jiao Xiaoqi","year":"2019","unstructured":"Xiaoqi Jiao , Yichun Yin , Lifeng Shang , Xin Jiang , Xiao Chen , Linlin Li , Fang Wang , and Qun Liu . TinyBERT: Distilling BERT for Natural Language Understanding. arXiv preprint arXiv:1909.10351 , 2019 . Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. TinyBERT: Distilling BERT for Natural Language Understanding. arXiv preprint arXiv:1909.10351, 2019."},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/3360307"},{"key":"e_1_3_2_1_40_1","volume-title":"ISCA","author":"Jouppi Norman P","year":"2017","unstructured":"Norman P Jouppi , Cliff Young , Nishant Patil , David Patterson , Gaurav Agrawal , Raminder Bajwa , Sarah Bates , Suresh Bhatia , Nan Boden , Al Borchers , In-Datacenter Performance Analysis of a Tensor Processing Unit . In ISCA , 2017 . Norman P Jouppi, Cliff Young, Nishant Patil, David Patterson, Gaurav Agrawal, Raminder Bajwa, Sarah Bates, Suresh Bhatia, Nan Boden, Al Borchers, et al. In-Datacenter Performance Analysis of a Tensor Processing Unit. In ISCA, 2017."},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/3400302.3415639"},{"key":"e_1_3_2_1_42_1","volume-title":"ICML","author":"Katharopoulos Angelos","year":"2020","unstructured":"Angelos Katharopoulos , Apoorv Vyas , Nikolaos Pappas , and Fran\u00e7ois Fleuret . Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention . In ICML , 2020 . Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and Fran\u00e7ois Fleuret. Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention. In ICML, 2020."},{"key":"e_1_3_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.5555\/645671.665526"},{"key":"e_1_3_2_1_44_1","volume-title":"I-BERT: Integer-only BERT Quantization. arXiv preprint arXiv:2101.01321","author":"Kim Sehoon","year":"2021","unstructured":"Sehoon Kim , Amir Gholami , Zhewei Yao , Michael W Mahoney , and Kurt Keutzer . I-BERT: Integer-only BERT Quantization. arXiv preprint arXiv:2101.01321 , 2021 . Sehoon Kim, Amir Gholami, Zhewei Yao, Michael W Mahoney, and Kurt Keutzer. I-BERT: Integer-only BERT Quantization. arXiv preprint arXiv:2101.01321, 2021."},{"key":"e_1_3_2_1_45_1","volume-title":"GRIP: A Graph Neural Network Accelerator Architecture. arXiv preprint arXiv:2007.13828","author":"Kiningham Kevin","year":"2020","unstructured":"Kevin Kiningham , Christopher Re , and Philip Levis . GRIP: A Graph Neural Network Accelerator Architecture. arXiv preprint arXiv:2007.13828 , 2020 . Kevin Kiningham, Christopher Re, and Philip Levis. GRIP: A Graph Neural Network Accelerator Architecture. arXiv preprint arXiv:2007.13828, 2020."},{"key":"e_1_3_2_1_46_1","volume-title":"Reformer: The Efficient Transformer. arXiv preprint arXiv:2001.04451","author":"Kitaev Nikita","year":"2020","unstructured":"Nikita Kitaev , \u0141ukasz Kaiser , and Anselm Levskaya . Reformer: The Efficient Transformer. arXiv preprint arXiv:2001.04451 , 2020 . Nikita Kitaev, \u0141ukasz Kaiser, and Anselm Levskaya. Reformer: The Efficient Transformer. arXiv preprint arXiv:2001.04451, 2020."},{"key":"e_1_3_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/3133901"},{"key":"e_1_3_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-01767-4"},{"key":"e_1_3_2_1_49_1","volume-title":"ICLR","author":"Kumar Aviral","year":"2022","unstructured":"Aviral Kumar , Amir Yazdanbakhsh , Milad Hashemi , Kevin Swersky , and Sergey Levine . Data-Driven Offline Optimization for Architecting Hardware Accelerators . In ICLR , 2022 . Aviral Kumar, Amir Yazdanbakhsh, Milad Hashemi, Kevin Swersky, and Sergey Levine. Data-Driven Offline Optimization for Architecting Hardware Accelerators. In ICLR, 2022."},{"key":"e_1_3_2_1_50_1","volume-title":"MICRO","author":"Kwon Hyoukjun","year":"2019","unstructured":"Hyoukjun Kwon , Prasanth Chatarasi , Michael Pellauer , Angshuman Parashar , Vivek Sarkar , and Tushar Krishna . Understanding Reuse , Performance, and Hardware Cost of DNN Dataflow : A Data-Centric Approach . In MICRO , 2019 . Hyoukjun Kwon, Prasanth Chatarasi, Michael Pellauer, Angshuman Parashar, Vivek Sarkar, and Tushar Krishna. Understanding Reuse, Performance, and Hardware Cost of DNN Dataflow: A Data-Centric Approach. In MICRO, 2019."},{"key":"e_1_3_2_1_51_1","volume-title":"ASPLOS","author":"Kwon Hyoukjun","year":"2018","unstructured":"Hyoukjun Kwon , Ananda Samajdar , and Tushar Krishna . MAERI : Enabling Flexible Dataflow Mapping over DNN Accelerators via Reconfigurable Interconnects . In ASPLOS , 2018 . Hyoukjun Kwon, Ananda Samajdar, and Tushar Krishna. MAERI: Enabling Flexible Dataflow Mapping over DNN Accelerators via Reconfigurable Interconnects. In ASPLOS, 2018."},{"key":"e_1_3_2_1_52_1","volume-title":"Cross-Lingual Language Model Pretraining. arXiv preprint arXiv:1901.07291","author":"Lample Guillaume","year":"2019","unstructured":"Guillaume Lample and Alexis Conneau . Cross-Lingual Language Model Pretraining. arXiv preprint arXiv:1901.07291 , 2019 . Guillaume Lample and Alexis Conneau. Cross-Lingual Language Model Pretraining. arXiv preprint arXiv:1901.07291, 2019."},{"key":"e_1_3_2_1_53_1","volume-title":"Flaubert: Unsupervised Language Model Pre-training for French. arXiv preprint arXiv:1912.05372","author":"Le Hang","year":"2019","unstructured":"Hang Le , Lo\u00efc Vial , Jibril Frej , Vincent Segonne , Maximin Coavoux , Benjamin Lecouteux , Alexandre Allauzen , Beno\u00eet Crabb\u00e9 , Laurent Besacier , and Didier Schwab . Flaubert: Unsupervised Language Model Pre-training for French. arXiv preprint arXiv:1912.05372 , 2019 . Hang Le, Lo\u00efc Vial, Jibril Frej, Vincent Segonne, Maximin Coavoux, Benjamin Lecouteux, Alexandre Allauzen, Beno\u00eet Crabb\u00e9, Laurent Besacier, and Didier Schwab. Flaubert: Unsupervised Language Model Pre-training for French. arXiv preprint arXiv:1912.05372, 2019."},{"key":"e_1_3_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA51647.2021.00070"},{"key":"e_1_3_2_1_55_1","volume-title":"ISCA","author":"Li Zheng","year":"2022","unstructured":"Zheng Li , Soroush Ghodrati , Amir Yazdanbakhsh , Hadi Esmaeilzadeh , and Mingu Kang . Accelerating Attention through Gradient-Based Learned Runtime Pruning . In ISCA , 2022 . Zheng Li, Soroush Ghodrati, Amir Yazdanbakhsh, Hadi Esmaeilzadeh, and Mingu Kang. Accelerating Attention through Gradient-Based Learned Runtime Pruning. In ISCA, 2022."},{"key":"e_1_3_2_1_56_1","volume-title":"Generating Wikipedia by Summarizing Long Sequences. arXiv preprint arXiv:1801.10198","author":"Liu Peter J","year":"2018","unstructured":"Peter J Liu , Mohammad Saleh , Etienne Pot , Ben Goodrich , Ryan Sepassi , Lukasz Kaiser , and Noam Shazeer . Generating Wikipedia by Summarizing Long Sequences. arXiv preprint arXiv:1801.10198 , 2018 . Peter J Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, and Noam Shazeer. Generating Wikipedia by Summarizing Long Sequences. arXiv preprint arXiv:1801.10198, 2018."},{"key":"e_1_3_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/3466752.3480125"},{"key":"e_1_3_2_1_58_1","volume-title":"FlexFlow: A Flexible Dataflow Accelerator Architecture for Convolutional Neural Networks. In HPCA","author":"Wenyan","year":"2017","unstructured":"Wenyan Lu et al . FlexFlow: A Flexible Dataflow Accelerator Architecture for Convolutional Neural Networks. In HPCA , 2017 . Wenyan Lu et al. FlexFlow: A Flexible Dataflow Accelerator Architecture for Convolutional Neural Networks. In HPCA, 2017."},{"key":"e_1_3_2_1_59_1","volume-title":"ZigZag: A Memory-Centric Rapid DNN Accelerator Design Space Exploration Framework. arXiv preprint arXiv:2007.11360","author":"Mei Linyan","year":"2020","unstructured":"Linyan Mei , Pouya Houshmand , Vikram Jain , Sebastian Giraldo , and Marian Verhelst . ZigZag: A Memory-Centric Rapid DNN Accelerator Design Space Exploration Framework. arXiv preprint arXiv:2007.11360 , 2020 . Linyan Mei, Pouya Houshmand, Vikram Jain, Sebastian Giraldo, and Marian Verhelst. ZigZag: A Memory-Centric Rapid DNN Accelerator Design Space Exploration Framework. arXiv preprint arXiv:2007.11360, 2020."},{"key":"e_1_3_2_1_60_1","volume-title":"Karthik Mandakolathur, and Shar Narasimhan. Boosting NVIDIA MLPerf Training v1.1 Performance with Full Stack Optimization. https:\/\/tinyurl.com\/3dku474c","author":"Nguyen Vinh","year":"2021","unstructured":"Vinh Nguyen , Sukru Burc Eryilmax , Karthik Mandakolathur, and Shar Narasimhan. Boosting NVIDIA MLPerf Training v1.1 Performance with Full Stack Optimization. https:\/\/tinyurl.com\/3dku474c , 2021 . Vinh Nguyen, Sukru Burc Eryilmax, Karthik Mandakolathur, and Shar Narasimhan. Boosting NVIDIA MLPerf Training v1.1 Performance with Full Stack Optimization. https:\/\/tinyurl.com\/3dku474c, 2021."},{"key":"e_1_3_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1145\/3453483.3454083"},{"key":"e_1_3_2_1_62_1","unstructured":"Nvidia. NVDLA Deep Learning Accelerator. http:\/\/nvdla.org 2017. \t\t\t\t  Nvidia. NVDLA Deep Learning Accelerator. http:\/\/nvdla.org 2017."},{"key":"e_1_3_2_1_63_1","volume-title":"https:\/\/github.com\/NVIDIA\/FasterTransformer","year":"2021","unstructured":"Nvidia. FasterTransforemr. https:\/\/github.com\/NVIDIA\/FasterTransformer , 2021 . Nvidia. FasterTransforemr. https:\/\/github.com\/NVIDIA\/FasterTransformer, 2021."},{"key":"e_1_3_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2019.00042"},{"key":"e_1_3_2_1_65_1","volume-title":"ICML","author":"Parmar Niki","year":"2018","unstructured":"Niki Parmar , Ashish Vaswani , Jakob Uszkoreit , Lukasz Kaiser , Noam Shazeer , Alexander Ku , and Dustin Tran . Image Transformer . In ICML , 2018 . Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran. Image Transformer. In ICML, 2018."},{"key":"e_1_3_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA47549.2020.00015"},{"key":"e_1_3_2_1_67_1","volume-title":"Sinong Wang, and Jie Tang. Blockwise Self-Attention for Long Document Understanding. arXiv preprint arXiv:1911.02972","author":"Qiu Jiezhong","year":"2019","unstructured":"Jiezhong Qiu , Hao Ma , Omer Levy , Scott Wen-tau Yih , Sinong Wang, and Jie Tang. Blockwise Self-Attention for Long Document Understanding. arXiv preprint arXiv:1911.02972 , 2019 . Jiezhong Qiu, Hao Ma, Omer Levy, Scott Wen-tau Yih, Sinong Wang, and Jie Tang. Blockwise Self-Attention for Long Document Understanding. arXiv preprint arXiv:1911.02972, 2019."},{"key":"e_1_3_2_1_68_1","volume-title":"Rabe and Charles Staats. Self-attention Does Not Need O(n^2) Memory. arXiv preprint arXiv:2112.05682","author":"Markus","year":"2021","unstructured":"Markus N. Rabe and Charles Staats. Self-attention Does Not Need O(n^2) Memory. arXiv preprint arXiv:2112.05682 , 2021 . Markus N. Rabe and Charles Staats. Self-attention Does Not Need O(n^2) Memory. arXiv preprint arXiv:2112.05682, 2021."},{"key":"e_1_3_2_1_69_1","volume-title":"Compressive Transformers for Long-Range Sequence Modelling. arXiv preprint arXiv:1911.05507","author":"Rae Jack W","year":"2019","unstructured":"Jack W Rae , Anna Potapenko , Siddhant M Jayakumar , and Timothy P Lillicrap . Compressive Transformers for Long-Range Sequence Modelling. arXiv preprint arXiv:1911.05507 , 2019 . Jack W Rae, Anna Potapenko, Siddhant M Jayakumar, and Timothy P Lillicrap. Compressive Transformers for Long-Range Sequence Modelling. arXiv preprint arXiv:1911.05507, 2019."},{"key":"e_1_3_2_1_70_1","volume-title":"Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. arXiv preprint arXiv:1910.10683","author":"Raffel Colin","year":"2019","unstructured":"Colin Raffel , Noam Shazeer , Adam Roberts , Katherine Lee , Sharan Narang , Michael Matena , Yanqi Zhou , Wei Li , and Peter J Liu . Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. arXiv preprint arXiv:1910.10683 , 2019 . Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. arXiv preprint arXiv:1910.10683, 2019."},{"key":"e_1_3_2_1_71_1","doi-asserted-by":"publisher","DOI":"10.1145\/2491956.2462176"},{"key":"e_1_3_2_1_72_1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00353"},{"key":"e_1_3_2_1_73_1","volume-title":"Poor Man\u2019s BERT: Smaller and Faster Transformer Models. arXiv preprint arXiv:2004.03844","author":"Sajjad Hassan","year":"2020","unstructured":"Hassan Sajjad , Fahim Dalvi , Nadir Durrani , and Preslav Nakov . Poor Man\u2019s BERT: Smaller and Faster Transformer Models. arXiv preprint arXiv:2004.03844 , 2020 . Hassan Sajjad, Fahim Dalvi, Nadir Durrani, and Preslav Nakov. Poor Man\u2019s BERT: Smaller and Faster Transformer Models. arXiv preprint arXiv:2004.03844, 2020."},{"key":"e_1_3_2_1_74_1","volume-title":"SCALE-Sim: Systolic CNN Accelerator Simulator. arXiv preprint arXiv:1811.02883","author":"Samajdar Ananda","year":"2018","unstructured":"Ananda Samajdar , Yuhao Zhu , Paul Whatmough , Matthew Mattina , and Tushar Krishna . SCALE-Sim: Systolic CNN Accelerator Simulator. arXiv preprint arXiv:1811.02883 , 2018 . Ananda Samajdar, Yuhao Zhu, Paul Whatmough, Matthew Mattina, and Tushar Krishna. SCALE-Sim: Systolic CNN Accelerator Simulator. arXiv preprint arXiv:1811.02883, 2018."},{"key":"e_1_3_2_1_75_1","volume-title":"Faster, Cheaper and Lighter. arXiv preprint arXiv:1910.01108","author":"Sanh Victor","year":"2019","unstructured":"Victor Sanh , Lysandre Debut , Julien Chaumond , and Thomas Wolf . DistilBERT , A Distilled Version of BERT: Smaller , Faster, Cheaper and Lighter. arXiv preprint arXiv:1910.01108 , 2019 . Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. DistilBERT, A Distilled Version of BERT: Smaller, Faster, Cheaper and Lighter. arXiv preprint arXiv:1910.01108, 2019."},{"key":"e_1_3_2_1_76_1","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC55918.2022.00017"},{"key":"e_1_3_2_1_77_1","volume-title":"MICRO","author":"Shao Yakun Sophia","year":"2019","unstructured":"Yakun Sophia Shao , Jason Clemons , Rangharajan Venkatesan , Brian Zimmer , Matthew Fojtik , Nan Jiang , Ben Keller , Alicia Klinefelter , Nathaniel Pinckney , Priyanka Raina , : Scaling Deep-Learning Inference with Multi-Chip-Module-Based Architecture . In MICRO , 2019 . Yakun Sophia Shao, Jason Clemons, Rangharajan Venkatesan, Brian Zimmer, Matthew Fojtik, Nan Jiang, Ben Keller, Alicia Klinefelter, Nathaniel Pinckney, Priyanka Raina, et al. Simba: Scaling Deep-Learning Inference with Multi-Chip-Module-Based Architecture. In MICRO, 2019."},{"key":"e_1_3_2_1_78_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i05.6409"},{"key":"e_1_3_2_1_79_1","volume-title":"Bi-Directional Block Self-Attention for Fast and Memory-Efficient Sequence Modeling. arXiv preprint arXiv:1804.00857","author":"Shen Tao","year":"2018","unstructured":"Tao Shen , Tianyi Zhou , Guodong Long , Jing Jiang , and Chengqi Zhang . Bi-Directional Block Self-Attention for Fast and Memory-Efficient Sequence Modeling. arXiv preprint arXiv:1804.00857 , 2018 . Tao Shen, Tianyi Zhou, Guodong Long, Jing Jiang, and Chengqi Zhang. Bi-Directional Block Self-Attention for Fast and Memory-Efficient Sequence Modeling. arXiv preprint arXiv:1804.00857, 2018."},{"key":"e_1_3_2_1_80_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2019.00027"},{"key":"e_1_3_2_1_81_1","volume-title":"TransTrack: Multiple Object Tracking with Transformer. arXiv preprint arXiv:2012.15460","author":"Sun Peize","year":"2020","unstructured":"Peize Sun , Yi Jiang , Rufeng Zhang , Enze Xie , Jinkun Cao , Xinting Hu , Tao Kong , Zehuan Yuan , Changhu Wang , and Ping Luo . TransTrack: Multiple Object Tracking with Transformer. arXiv preprint arXiv:2012.15460 , 2020 . Peize Sun, Yi Jiang, Rufeng Zhang, Enze Xie, Jinkun Cao, Xinting Hu, Tao Kong, Zehuan Yuan, Changhu Wang, and Ping Luo. TransTrack: Multiple Object Tracking with Transformer. arXiv preprint arXiv:2012.15460, 2020."},{"key":"e_1_3_2_1_82_1","volume-title":"MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices. arXiv preprint arXiv:2004.02984","author":"Sun Zhiqing","year":"2020","unstructured":"Zhiqing Sun , Hongkun Yu , Xiaodan Song , Renjie Liu , Yiming Yang , and Denny Zhou . MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices. arXiv preprint arXiv:2004.02984 , 2020 . Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou. MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices. arXiv preprint arXiv:2004.02984, 2020."},{"key":"e_1_3_2_1_83_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-01766-7"},{"key":"e_1_3_2_1_84_1","volume-title":"HCS","author":"Tambe Thierry","year":"2021","unstructured":"Thierry Tambe , En-Yu Yang , Glenn G Ko , Yuji Chai , Coleman Hooper , Marco Donato , Paul N Whatmough , Alexander M Rush , David Brooks , and Gu-Yeon Wei . SM6 : A 16nm System-on-Chip for Accurate and Noise-Robust Attention-Based NLP Applications . In HCS , 2021 . Thierry Tambe, En-Yu Yang, Glenn G Ko, Yuji Chai, Coleman Hooper, Marco Donato, Paul N Whatmough, Alexander M Rush, David Brooks, and Gu-Yeon Wei. SM6: A 16nm System-on-Chip for Accurate and Noise-Robust Attention-Based NLP Applications. In HCS, 2021."},{"key":"e_1_3_2_1_85_1","volume-title":"Synthesizer: Rethinking self-attention in transformer models. arXiv preprint arXiv:2005.00743","author":"Tay Yi","year":"2020","unstructured":"Yi Tay , Dara Bahri , Donald Metzler , Da-Cheng Juan , Zhe Zhao , and Che Zheng . Synthesizer: Rethinking self-attention in transformer models. arXiv preprint arXiv:2005.00743 , 2020 . Yi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan, Zhe Zhao, and Che Zheng. Synthesizer: Rethinking self-attention in transformer models. arXiv preprint arXiv:2005.00743, 2020."},{"key":"e_1_3_2_1_86_1","volume-title":"ICLR","author":"Tay Yi","year":"2021","unstructured":"Yi Tay , Mostafa Dehghani , Samira Abnar , Yikang Shen , Dara Bahri , Philip Pham , Jinfeng Rao , Liu Yang , Sebastian Ruder , and Donald Metzler . Long Range Arena: A Benchmark for Efficient Transformers . In ICLR , 2021 . Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler. Long Range Arena: A Benchmark for Efficient Transformers. In ICLR, 2021."},{"key":"e_1_3_2_1_87_1","volume-title":"Nvidia T4 Tensor Core GPU. https:\/\/www.nvidia.com\/en-us\/data-center\/tesla-t4\/","year":"2018","unstructured":"Tesla, Nvidia. Nvidia T4 Tensor Core GPU. https:\/\/www.nvidia.com\/en-us\/data-center\/tesla-t4\/ , 2018 . Tesla, Nvidia. Nvidia T4 Tensor Core GPU. https:\/\/www.nvidia.com\/en-us\/data-center\/tesla-t4\/, 2018."},{"key":"e_1_3_2_1_88_1","volume-title":"V100 GPU Architecture. https:\/\/www.nvidia.com\/en-us\/data-center\/v100\/","year":"2018","unstructured":"Tesla, Nvidia. V100 GPU Architecture. https:\/\/www.nvidia.com\/en-us\/data-center\/v100\/ , 2018 . Tesla, Nvidia. V100 GPU Architecture. https:\/\/www.nvidia.com\/en-us\/data-center\/v100\/, 2018."},{"key":"e_1_3_2_1_89_1","volume-title":"NeurIPS","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani , Noam Shazeer , Niki Parmar , Jakob Uszkoreit , Llion Jones , Aidan N Gomez , \u0141ukasz Kaiser , and Illia Polosukhin . Attention Is All You Need . In NeurIPS , 2017 . Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. Attention Is All You Need. In NeurIPS, 2017."},{"key":"e_1_3_2_1_90_1","volume-title":"DeepTools: Compiler and Execution Runtime Extensions for RaPiD AI Accelerator","author":"Venkataramani Swagath","year":"2019","unstructured":"Swagath Venkataramani , Jungwook Choi , Vijayalakshmi Srinivasan , Wei Wang , Jintao Zhang , Marcel Schaal , Mauricio J. Serrano , Kazuaki Ishizaki , Hiroshi Inoue , Eri Ogawa , Moriyoshi Ohara , Leland Chang , and Kailash Gopalakrishnan . DeepTools: Compiler and Execution Runtime Extensions for RaPiD AI Accelerator . IEEE Micro , 2019 . Swagath Venkataramani, Jungwook Choi, Vijayalakshmi Srinivasan, Wei Wang, Jintao Zhang, Marcel Schaal, Mauricio J. Serrano, Kazuaki Ishizaki, Hiroshi Inoue, Eri Ogawa, Moriyoshi Ohara, Leland Chang, and Kailash Gopalakrishnan. DeepTools: Compiler and Execution Runtime Extensions for RaPiD AI Accelerator. IEEE Micro, 2019."},{"key":"e_1_3_2_1_91_1","volume-title":"MICRO","author":"Venkataramani Swagath","year":"2017","unstructured":"Swagath Venkataramani , Ashish Ranjan , Subarno Banerjee , Dipankar Das , Sasikanth Avancha , Ashok Jagannathan , Ajaya Durg , Dheemanth Nagaraj , Bharat Kaul , Pradeep Dubey , and Anand Raghunathan . SCALEDEEP : A Scalable Compute Architecture for Learning and Evaluating Deep Networks . In MICRO , 2017 . Swagath Venkataramani, Ashish Ranjan, Subarno Banerjee, Dipankar Das, Sasikanth Avancha, Ashok Jagannathan, Ajaya Durg, Dheemanth Nagaraj, Bharat Kaul, Pradeep Dubey, and Anand Raghunathan. SCALEDEEP: A Scalable Compute Architecture for Learning and Evaluating Deep Networks. In MICRO, 2017."},{"key":"e_1_3_2_1_92_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA51647.2021.00018"},{"key":"e_1_3_2_1_93_1","volume-title":"Linformer: Self-Attention with Linear Complexity. arXiv preprint arXiv:2006.04768","author":"Wang Sinong","year":"2020","unstructured":"Sinong Wang , Belinda Li , Madian Khabsa , Han Fang , and Hao Ma . Linformer: Self-Attention with Linear Complexity. arXiv preprint arXiv:2006.04768 , 2020 . Sinong Wang, Belinda Li, Madian Khabsa, Han Fang, and Hao Ma. Linformer: Self-Attention with Linear Complexity. arXiv preprint arXiv:2006.04768, 2020."},{"key":"e_1_3_2_1_94_1","volume-title":"MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers. arXiv preprint arXiv:2002.10957","author":"Wang Wenhui","year":"2020","unstructured":"Wenhui Wang , Furu Wei , Li Dong , Hangbo Bao , Nan Yang , and Ming Zhou . MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers. arXiv preprint arXiv:2002.10957 , 2020 . Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou. MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers. arXiv preprint arXiv:2002.10957, 2020."},{"key":"e_1_3_2_1_95_1","volume-title":"Structured Pruning of Large Language Models. arXiv preprint arXiv:1910.04732","author":"Wang Ziheng","year":"2019","unstructured":"Ziheng Wang , Jeremy Wohlwend , and Tao Lei . Structured Pruning of Large Language Models. arXiv preprint arXiv:1910.04732 , 2019 . Ziheng Wang, Jeremy Wohlwend, and Tao Lei. Structured Pruning of Large Language Models. arXiv preprint arXiv:1910.04732, 2019."},{"key":"e_1_3_2_1_96_1","doi-asserted-by":"publisher","DOI":"10.1145\/3240765.3240856"},{"key":"e_1_3_2_1_97_1","doi-asserted-by":"publisher","DOI":"10.1145\/3061639.3062207"},{"key":"e_1_3_2_1_99_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00009"},{"key":"e_1_3_2_1_100_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCAD45719.2019.8942149"},{"key":"e_1_3_2_1_101_1","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378514"},{"key":"e_1_3_2_1_102_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO56248.2022.00059"},{"key":"e_1_3_2_1_103_1","volume-title":"ISCA","author":"Yazdanbakhsh Amir","year":"2018","unstructured":"Amir Yazdanbakhsh , Kambiz Samadi , Nam Sung Kim , and Hadi Esmaeilzadeh . GANAX : A Unified MIMD-SIMD Acceleration for Generative Adversarial Networks . In ISCA , 2018 . Amir Yazdanbakhsh, Kambiz Samadi, Nam Sung Kim, and Hadi Esmaeilzadeh. GANAX: A Unified MIMD-SIMD Acceleration for Generative Adversarial Networks. In ISCA, 2018."},{"key":"e_1_3_2_1_104_1","volume-title":"Q8BERT: Quantized 8Bit BERT. arXiv preprint arXiv:1910.06188","author":"Zafrir Ofir","year":"2019","unstructured":"Ofir Zafrir , Guy Boudoukh , Peter Izsak , and Moshe Wasserblat . Q8BERT: Quantized 8Bit BERT. arXiv preprint arXiv:1910.06188 , 2019 . Ofir Zafrir, Guy Boudoukh, Peter Izsak, and Moshe Wasserblat. Q8BERT: Quantized 8Bit BERT. arXiv preprint arXiv:1910.06188, 2019."},{"key":"e_1_3_2_1_105_1","doi-asserted-by":"publisher","DOI":"10.1145\/2684746.2689060"},{"key":"e_1_3_2_1_106_1","volume-title":"Multi-Scale Vision Longformer: A New Vision Transformer for High-Resolution Image Encoding. arXiv preprint arXiv:2103.15358","author":"Zhang Pengchuan","year":"2021","unstructured":"Pengchuan Zhang , Xiyang Dai , Jianwei Yang , Bin Xiao , Lu Yuan , Lei Zhang , and Jianfeng Gao . Multi-Scale Vision Longformer: A New Vision Transformer for High-Resolution Image Encoding. arXiv preprint arXiv:2103.15358 , 2021 . Pengchuan Zhang, Xiyang Dai, Jianwei Yang, Bin Xiao, Lu Yuan, Lei Zhang, and Jianfeng Gao. Multi-Scale Vision Longformer: A New Vision Transformer for High-Resolution Image Encoding. arXiv preprint arXiv:2103.15358, 2021."},{"key":"e_1_3_2_1_107_1","volume-title":"TernaryBERT: Distillation-aware Ultra-low Bit BERT. arXiv preprint arXiv:2009.12812","author":"Zhang Wei","year":"2020","unstructured":"Wei Zhang , Lu Hou , Yichun Yin , Lifeng Shang , Xiao Chen , Xin Jiang , and Qun Liu . TernaryBERT: Distillation-aware Ultra-low Bit BERT. arXiv preprint arXiv:2009.12812 , 2020 . Wei Zhang, Lu Hou, Yichun Yin, Lifeng Shang, Xiao Chen, Xin Jiang, and Qun Liu. TernaryBERT: Distillation-aware Ultra-low Bit BERT. arXiv preprint arXiv:2009.12812, 2020."},{"key":"e_1_3_2_1_108_1","volume-title":"NeurIPS","author":"Zhu Chen","year":"2021","unstructured":"Chen Zhu , Wei Ping , Chaowei Xiao , Mohammad Shoeybi , Tom Goldstein , Anima Anandkumar , and Bryan Catanzaro . Long-Short Transformer : Efficient Transformers for Language and Vision . NeurIPS , 2021 . Chen Zhu, Wei Ping, Chaowei Xiao, Mohammad Shoeybi, Tom Goldstein, Anima Anandkumar, and Bryan Catanzaro. Long-Short Transformer: Efficient Transformers for Language and Vision. NeurIPS, 2021."},{"key":"e_1_3_2_1_109_1","volume-title":"ICLR","author":"Zhu Xizhou","year":"2021","unstructured":"Xizhou Zhu , Weijie Su , Lewei Lu , Bin Li , Xiaogang Wang , and Jifeng Dai . Deformable DETR : Deformable Transformers for End-to-End Object Detection . In ICLR , 2021 . Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable DETR: Deformable Transformers for End-to-End Object Detection. In ICLR, 2021."}],"event":{"name":"ASPLOS '23: 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2","location":"Vancouver BC Canada","acronym":"ASPLOS '23","sponsor":["SIGARCH ACM Special Interest Group on Computer Architecture","SIGOPS ACM Special Interest Group on Operating Systems","SIGPLAN ACM Special Interest Group on Programming Languages"]},"container-title":["Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3575693.3575747","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3575693.3575747","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:51:20Z","timestamp":1750182680000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3575693.3575747"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,1,27]]},"references-count":108,"alternative-id":["10.1145\/3575693.3575747","10.1145\/3575693"],"URL":"https:\/\/doi.org\/10.1145\/3575693.3575747","relation":{},"subject":[],"published":{"date-parts":[[2023,1,27]]},"assertion":[{"value":"2023-01-30","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}