{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,3]],"date-time":"2026-06-03T07:56:06Z","timestamp":1780473366438,"version":"3.54.1"},"publisher-location":"New York, NY, USA","reference-count":49,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,10,14]],"date-time":"2022-10-14T00:00:00Z","timestamp":1665705600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,10,14]]},"DOI":"10.1145\/3495243.3517020","type":"proceedings-article","created":{"date-parts":[[2022,10,14]],"date-time":"2022-10-14T15:38:33Z","timestamp":1665761913000},"page":"487-500","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":13,"title":["Romou"],"prefix":"10.1145","author":[{"given":"Rendong","family":"Liang","sequence":"first","affiliation":[{"name":"University of California"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ting","family":"Cao","sequence":"additional","affiliation":[{"name":"Microsoft Research"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jicheng","family":"Wen","sequence":"additional","affiliation":[{"name":"Microsoft STCA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Manni","family":"Wang","sequence":"additional","affiliation":[{"name":"Microsoft Research and Xi'an Jiao Tong University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yang","family":"Wang","sequence":"additional","affiliation":[{"name":"Microsoft Research"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jianhua","family":"Zou","sequence":"additional","affiliation":[{"name":"Xi'an Jiao Tong University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yunxin","family":"Liu","sequence":"additional","affiliation":[{"name":"Tsinghua University"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,10,14]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"Chameleon: Adaptive Code Optimization for Expedited Deep Neural Network Compilation.","author":"Ahn Byung Hoon","year":"2020","unstructured":"Byung Hoon Ahn , Prannoy Pilligundla , Amir Yazdanbakhsh , and Hadi Esmaeilzadeh . 2020 . Chameleon: Adaptive Code Optimization for Expedited Deep Neural Network Compilation. (2020). https:\/\/openreview.net\/forum?id=rygG4AVFvH Byung Hoon Ahn, Prannoy Pilligundla, Amir Yazdanbakhsh, and Hadi Esmaeilzadeh. 2020. Chameleon: Adaptive Code Optimization for Expedited Deep Neural Network Compilation. (2020). https:\/\/openreview.net\/forum?id=rygG4AVFvH"},{"key":"e_1_3_2_1_2_1","volume-title":"Performance characterization of mobile GP-GPUs","author":"Andargie Fitsum Assamnew","unstructured":"Fitsum Assamnew Andargie and Jonathan Rose . 2015. Performance characterization of mobile GP-GPUs . In AFRICON. IEEE , 1--6. Fitsum Assamnew Andargie and Jonathan Rose. 2015. Performance characterization of mobile GP-GPUs. In AFRICON. IEEE, 1--6."},{"key":"e_1_3_2_1_3_1","unstructured":"ARM. 2020. Arm Mali Bifrost and Valhall OpenCL. (2020).  ARM. 2020. Arm Mali Bifrost and Valhall OpenCL. (2020)."},{"key":"e_1_3_2_1_4_1","unstructured":"ARM. 2020. The Bifrost Shader Core. https:\/\/developer.arm.com\/solutions\/graphics-and-gaming\/developer-guides\/learn-the-basics\/the-bifrost-shader-core\/the-bifrost-shader-core  ARM. 2020. The Bifrost Shader Core. https:\/\/developer.arm.com\/solutions\/graphics-and-gaming\/developer-guides\/learn-the-basics\/the-bifrost-shader-core\/the-bifrost-shader-core"},{"key":"e_1_3_2_1_5_1","unstructured":"Krishnaraj Bhat. 2020. A tool which profiles OpenCL devices to find their peak capacities. https:\/\/github.com\/krrishnarraj\/clpeak  Krishnaraj Bhat. 2020. A tool which profiles OpenCL devices to find their peak capacities. https:\/\/github.com\/krrishnarraj\/clpeak"},{"key":"e_1_3_2_1_6_1","volume-title":"Strategy Analytics: Q2 2019 Smartphone and Tablet GPU Market Share: Apple Gains Share as Arm Falters. https:\/\/www.businesswire.com\/news\/home\/20191118005549\/en\/","year":"2019","unstructured":"Businesswire. 2019 . Strategy Analytics: Q2 2019 Smartphone and Tablet GPU Market Share: Apple Gains Share as Arm Falters. https:\/\/www.businesswire.com\/news\/home\/20191118005549\/en\/ Businesswire. 2019. Strategy Analytics: Q2 2019 Smartphone and Tablet GPU Market Share: Apple Gains Share as Arm Falters. https:\/\/www.businesswire.com\/news\/home\/20191118005549\/en\/"},{"key":"e_1_3_2_1_7_1","volume-title":"13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18)","author":"Chen Tianqi","year":"2018","unstructured":"Tianqi Chen , Thierry Moreau , Ziheng Jiang , Lianmin Zheng , Eddie Yan , Haichen Shen , Meghan Cowan , Leyuan Wang , Yuwei Hu , Luis Ceze , 2018 . TVM: An automated end-to-end optimizing compiler for deep learning . In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18) . 578--594. Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, et al. 2018. TVM: An automated end-to-end optimizing compiler for deep learning. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18). 578--594."},{"key":"e_1_3_2_1_8_1","volume-title":"Autotuning OpenCL workgroup size for stencil patterns. arXiv preprint arXiv:1511.02490","author":"Cummins Chris","year":"2015","unstructured":"Chris Cummins , Pavlos Petoumenos , Michel Steuwer , and Hugh Leather . 2015. Autotuning OpenCL workgroup size for stencil patterns. arXiv preprint arXiv:1511.02490 ( 2015 ). Chris Cummins, Pavlos Petoumenos, Michel Steuwer, and Hugh Leather. 2015. Autotuning OpenCL workgroup size for stencil patterns. arXiv preprint arXiv:1511.02490 (2015)."},{"key":"e_1_3_2_1_9_1","unstructured":"CUTLASS. 2020. CUTLASS Convolution. https:\/\/github.com\/NVIDIA\/cutlass\/blob\/master\/media\/docs\/implicit_gemm_convolution.md  CUTLASS. 2020. CUTLASS Convolution. https:\/\/github.com\/NVIDIA\/cutlass\/blob\/master\/media\/docs\/implicit_gemm_convolution.md"},{"key":"e_1_3_2_1_10_1","volume-title":"The most used smartphone GPU -","year":"2019","unstructured":"DeviceAtlas. 2019. The most used smartphone GPU - 2019 . https:\/\/deviceatlas.com\/blog\/most-used-smartphone-gpu DeviceAtlas. 2019. The most used smartphone GPU - 2019. https:\/\/deviceatlas.com\/blog\/most-used-smartphone-gpu"},{"key":"e_1_3_2_1_11_1","volume-title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv preprint arXiv:1810.04805","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2018 . BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv preprint arXiv:1810.04805 (2018). Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv preprint arXiv:1810.04805 (2018)."},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2012.44"},{"key":"e_1_3_2_1_13_1","volume-title":"Machine Learning Based Auto-Tuning for Enhanced OpenCL Performance Portability. In IPDPS Workshops. IEEE Computer Society, 1231--1240","author":"Thomas","unstructured":"Thomas L. Falch and Anne C. Elster. 2015 . Machine Learning Based Auto-Tuning for Enhanced OpenCL Performance Portability. In IPDPS Workshops. IEEE Computer Society, 1231--1240 . Thomas L. Falch and Anne C. Elster. 2015. Machine Learning Based Auto-Tuning for Enhanced OpenCL Performance Portability. In IPDPS Workshops. IEEE Computer Society, 1231--1240."},{"key":"e_1_3_2_1_14_1","unstructured":"Google. 2019. TensorFlow Lite: Deploy machine learning models on mobile and IoT devices. https:\/\/www.tensorflow.org\/lite  Google. 2019. TensorFlow Lite: Deploy machine learning models on mobile and IoT devices. https:\/\/www.tensorflow.org\/lite"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"crossref","unstructured":"Dominik Grewe and Anton Lokhmotov. 2011. Automatically generating and tuning GPU code for sparse matrix-vector multiplication from a high-level representation. In GPGPU. ACM 12.  Dominik Grewe and Anton Lokhmotov. 2011. Automatically generating and tuning GPU code for sparse matrix-vector multiplication from a high-level representation. In GPGPU. ACM 12.","DOI":"10.1145\/1964179.1964196"},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/2906388.2906396"},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_1_18_1","unstructured":"Andrew G. Howard Menglong Zhu Bo Chen Dmitry Kalenichenko Weijun Wang Tobias Weyand Marco Andreetto and Hartwig Adam. [n.d.]. MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. CoRR ([n. d.]). http:\/\/arxiv.org\/abs\/1704.04861  Andrew G. Howard Menglong Zhu Bo Chen Dmitry Kalenichenko Weijun Wang Tobias Weyand Marco Andreetto and Hartwig Adam. [n.d.]. MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. CoRR ([n. d.]). http:\/\/arxiv.org\/abs\/1704.04861"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/2935643.2935650"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3081333.3081360"},{"key":"e_1_3_2_1_21_1","volume-title":"SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and &lt;1MB model size. CoRR abs\/1602.07360","author":"Iandola Forrest N.","year":"2016","unstructured":"Forrest N. Iandola , Matthew W. Moskewicz , Khalid Ashraf , Song Han , William J. Dally , and Kurt Keutzer . 2016. SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and &lt;1MB model size. CoRR abs\/1602.07360 ( 2016 ). http:\/\/arxiv.org\/abs\/1602.07360 Forrest N. Iandola, Matthew W. Moskewicz, Khalid Ashraf, Song Han, William J. Dally, and Kurt Keutzer. 2016. SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and &lt;1MB model size. CoRR abs\/1602.07360 (2016). http:\/\/arxiv.org\/abs\/1602.07360"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"crossref","unstructured":"Shiqi Jiang Lihao Ran Ting Cao Yusen Xu and Yunxin Liu. 2020. Profiling and optimizing deep learning inference on mobile GPUs. In APSys. ACM 75--81.  Shiqi Jiang Lihao Ran Ting Cao Yusen Xu and Yunxin Liu. 2020. Profiling and optimizing deep learning inference on mobile GPUs. In APSys. ACM 75--81.","DOI":"10.1145\/3409963.3410493"},{"key":"e_1_3_2_1_23_1","volume-title":"MNN: A Universal and Efficient Inference Engine. In MLSys. mlsys.org.","author":"Jiang Xiaotang","year":"2020","unstructured":"Xiaotang Jiang , Huan Wang , Yiliu Chen , Ziqi Wu , Lichuan Wang , Bin Zou , Yafeng Yang , Zongyang Cui , Yu Cai , Tianhang Yu , Chengfei Lyu , and Zhihua Wu . 2020 . MNN: A Universal and Efficient Inference Engine. In MLSys. mlsys.org. Xiaotang Jiang, Huan Wang, Yiliu Chen, Ziqi Wu, Lichuan Wang, Bin Zou, Yafeng Yang, Zongyang Cui, Yu Cai, Tianhang Yu, Chengfei Lyu, and Zhihua Wu. 2020. MNN: A Universal and Efficient Inference Engine. In MLSys. mlsys.org."},{"key":"e_1_3_2_1_24_1","unstructured":"Khronos. 2020. OpenCL-SDK. https:\/\/github.com\/KhronosGroup\/OpenCL-SDK\/  Khronos. 2020. OpenCL-SDK. https:\/\/github.com\/KhronosGroup\/OpenCL-SDK\/"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3302424.3303950"},{"key":"e_1_3_2_1_26_1","volume-title":"The fifth international workshop on automatic performance tuning","author":"Komatsu Kazuhiko","unstructured":"Kazuhiko Komatsu , Katsuto Sato , Yusuke Arai , Kentaro Koyama , Hiroyuki Takizawa , and Hiroaki Kobayashi . 2010. Evaluating performance and portability of OpenCL programs . In The fifth international workshop on automatic performance tuning , Vol. 66 . 1. Kazuhiko Komatsu, Katsuto Sato, Yusuke Arai, Kentaro Koyama, Hiroyuki Takizawa, and Hiroaki Kobayashi. 2010. Evaluating performance and portability of OpenCL programs. In The fifth international workshop on automatic performance tuning, Vol. 66. 1."},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPSN.2016.7460664"},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.435"},{"key":"e_1_3_2_1_29_1","unstructured":"Juhyun Lee Nikolay Chirkov Ekaterina Ignasheva Yury Pisarchyk Mogan Shieh Fabio Riccardi Raman Sarokin Andrei Kulik and Matthias Grundmann. 2019. On-Device Neural Net Inference with Mobile GPUs. (2019). arXiv:1907.01989 http:\/\/arxiv.org\/abs\/1907.01989  Juhyun Lee Nikolay Chirkov Ekaterina Ignasheva Yury Pisarchyk Mogan Shieh Fabio Riccardi Raman Sarokin Andrei Kulik and Matthias Grundmann. 2019. On-Device Neural Net Inference with Mobile GPUs. (2019). arXiv:1907.01989 http:\/\/arxiv.org\/abs\/1907.01989"},{"key":"e_1_3_2_1_30_1","unstructured":"Juhyun Lee and Raman Sarokin. 2020. Even Faster Mobile GPU Inference with OpenCL. https:\/\/blog.tensorflow.org\/2020\/08\/faster-mobile-gpu-inference-with-opencl.html  Juhyun Lee and Raman Sarokin. 2020. Even Faster Mobile GPU Inference with OpenCL. https:\/\/blog.tensorflow.org\/2020\/08\/faster-mobile-gpu-inference-with-opencl.html"},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-01970-8_89"},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11390-013-1314-8"},{"key":"e_1_3_2_1_33_1","unstructured":"Eddie Yan Lianmin Zheng. 2021. Auto-tuning a Convolutional Network for Mobile GPU. https:\/\/tvm.apache.org\/docs\/tutorials\/autotvm\/tune_relay_mobile_gpu.html  Eddie Yan Lianmin Zheng. 2021. Auto-tuning a Convolutional Network for Mobile GPU. https:\/\/tvm.apache.org\/docs\/tutorials\/autotvm\/tune_relay_mobile_gpu.html"},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01264-9_8"},{"key":"e_1_3_2_1_35_1","unstructured":"MACE. 2020. https:\/\/github.com\/XiaoMi\/mace.  MACE. 2020. https:\/\/github.com\/XiaoMi\/mace."},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/3081333.3081359"},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2016.2549523"},{"key":"e_1_3_2_1_38_1","volume-title":"CLTune: A Generic Auto-Tuner for OpenCL Kernels","author":"Nugteren Cedric","unstructured":"Cedric Nugteren and Valeriu Codreanu . 2015. CLTune: A Generic Auto-Tuner for OpenCL Kernels . In MCSoC. IEEE Computer Society , 195--202. Cedric Nugteren and Valeriu Codreanu. 2015. CLTune: A Generic Auto-Tuner for OpenCL Kernels. In MCSoC. IEEE Computer Society, 195--202."},{"key":"e_1_3_2_1_39_1","unstructured":"Qualcomm. 2017. Qualcomm Snapdragon Mobile Platform OpenCL General Programming and Optimization. (2017).  Qualcomm. 2017. Qualcomm Snapdragon Mobile Platform OpenCL General Programming and Optimization. (2017)."},{"key":"e_1_3_2_1_40_1","unstructured":"Qualcomm. 2021. TensorFlow Lite: Deploy machine learning models on mobile and IoT devices. https:\/\/developer.qualcomm.com\/software\/snapdragon-profiler  Qualcomm. 2021. TensorFlow Lite: Deploy machine learning models on mobile and IoT devices. https:\/\/developer.qualcomm.com\/software\/snapdragon-profiler"},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/2491956.2462176"},{"key":"e_1_3_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2017.52"},{"key":"e_1_3_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.308"},{"key":"e_1_3_2_1_44_1","volume-title":"Ting Cao, and Yunxin Liu.","author":"Tang Xiaohu","year":"2021","unstructured":"Xiaohu Tang , Shihao Han , Li Lyna Zhang , Ting Cao, and Yunxin Liu. 2021 . To Bridge Neural Network Design and Real-World Performance: A Behaviour Study for Neural Networks. In MLSys . mlsys.org. Xiaohu Tang, Shihao Han, Li Lyna Zhang, Ting Cao, and Yunxin Liu. 2021. To Bridge Neural Network Design and Real-World Performance: A Behaviour Study for Neural Networks. In MLSys. mlsys.org."},{"key":"e_1_3_2_1_45_1","unstructured":"Tencent. 2018. Tencent ncnn deep learning framework. https:\/\/github.com\/Tencent\/ncnn  Tencent. 2018. Tencent ncnn deep learning framework. https:\/\/github.com\/Tencent\/ncnn"},{"key":"e_1_3_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/1498765.1498785"},{"key":"e_1_3_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/3372224.3419192"},{"key":"e_1_3_2_1_48_1","volume-title":"14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20)","author":"Zheng Lianmin","year":"2020","unstructured":"Lianmin Zheng , Chengfan Jia , Minmin Sun , Zhao Wu , Cody Hao Yu , Ameer Haj-Ali , Yida Wang , Jun Yang , Danyang Zhuo , Koushik Sen , 2020 . Ansor: Generating high-performance tensor programs for deep learning . In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20) . 863--879. Lianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu, Cody Hao Yu, Ameer Haj-Ali, Yida Wang, Jun Yang, Danyang Zhuo, Koushik Sen, et al. 2020. Ansor: Generating high-performance tensor programs for deep learning. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20). 863--879."},{"key":"e_1_3_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378508"}],"event":{"name":"ACM MobiCom '22: The 28th Annual International Conference on Mobile Computing and Networking","location":"Sydney NSW Australia","acronym":"ACM MobiCom '22","sponsor":["SIGMOBILE ACM Special Interest Group on Mobility of Systems, Users, Data and Computing"]},"container-title":["Proceedings of the 28th Annual International Conference on Mobile Computing And Networking"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3495243.3517020","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3495243.3517020","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:12:03Z","timestamp":1750191123000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3495243.3517020"}},"subtitle":["rapidly generate high-performance tensor kernels for mobile GPUs"],"short-title":[],"issued":{"date-parts":[[2022,10,14]]},"references-count":49,"alternative-id":["10.1145\/3495243.3517020","10.1145\/3495243"],"URL":"https:\/\/doi.org\/10.1145\/3495243.3517020","relation":{},"subject":[],"published":{"date-parts":[[2022,10,14]]},"assertion":[{"value":"2022-10-14","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}