{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,8]],"date-time":"2026-08-08T04:31:27Z","timestamp":1786163487525,"version":"3.56.0"},"publisher-location":"New York, NY, USA","reference-count":63,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,8,14]],"date-time":"2021-08-14T00:00:00Z","timestamp":1628899200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,8,14]]},"DOI":"10.1145\/3447548.3467078","type":"proceedings-article","created":{"date-parts":[[2021,8,12]],"date-time":"2021-08-12T06:12:05Z","timestamp":1628748725000},"page":"2543-2553","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":94,"title":["Auto-Split: A General Framework of Collaborative Edge-Cloud AI"],"prefix":"10.1145","author":[{"given":"Amin","family":"Banitalebi-Dehkordi","sequence":"first","affiliation":[{"name":"Huawei Technologies Canada Co., Ltd., Vancouver, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Naveen","family":"Vedula","sequence":"additional","affiliation":[{"name":"Huawei Technologies Canada Co., Ltd., Vancouver, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jian","family":"Pei","sequence":"additional","affiliation":[{"name":"Simon Fraser University, Vancouver, BC, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Fei","family":"Xia","sequence":"additional","affiliation":[{"name":"Huawei Technologies, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lanjun","family":"Wang","sequence":"additional","affiliation":[{"name":"Huawei Technologies Canada Co., Ltd., Vancouver, BC, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yong","family":"Zhang","sequence":"additional","affiliation":[{"name":"Huawei Technologies Canada Co., Ltd., Vancouver, BC, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2021,8,14]]},"reference":[{"key":"e_1_3_2_2_1_1","volume-title":"Deploying a Model as an Edge Service","year":"2021","unstructured":"ModelArts : Deploying a Model as an Edge Service . 2021 . (2021). https:\/\/support.huaweicloud.com\/en-us\/engineers-modelarts\/modelarts_23_0069.html ModelArts: Deploying a Model as an Edge Service. 2021. (2021). https:\/\/support.huaweicloud.com\/en-us\/engineers-modelarts\/modelarts_23_0069.html"},{"key":"e_1_3_2_2_2_1","volume-title":"https:\/\/cloud.google.com\/anthos","author":"Anthos Google","year":"2021","unstructured":"Google Anthos . 2021. ( 2021 ). https:\/\/cloud.google.com\/anthos Google Anthos. 2021. (2021). https:\/\/cloud.google.com\/anthos"},{"key":"e_1_3_2_2_3_1","volume-title":"https:\/\/www.alibabacloud.com\/blog\/327841","author":"Is Alibaba Edge","year":"2021","unstructured":"Alibaba Edge Computing AP Is . 2021. ( 2021 ). https:\/\/www.alibabacloud.com\/blog\/327841 Alibaba Edge Computing APIs. 2021. (2021). https:\/\/www.alibabacloud.com\/blog\/327841"},{"key":"e_1_3_2_2_4_1","volume-title":"ACIQ: Analytical Clipping for Integer Quantization of neural networks. ArXiv","author":"Banner R.","year":"2018","unstructured":"R. Banner , Y. Nahshan , E. Hoffer , and D. Soudry . 2018 . ACIQ: Analytical Clipping for Integer Quantization of neural networks. ArXiv , Vol. abs\/ 1810 .05723 (2018). R. Banner, Y. Nahshan, E. Hoffer, and D. Soudry. 2018. ACIQ: Analytical Clipping for Integer Quantization of neural networks. ArXiv , Vol. abs\/1810.05723 (2018)."},{"key":"e_1_3_2_2_5_1","unstructured":"R. Banner Y. Nahshan and D. Soudry. 2019. Post training 4-bit quantization of convolutional networks for rapid-deployment. In NeurIPS.  R. Banner Y. Nahshan and D. Soudry. 2019. Post training 4-bit quantization of convolutional networks for rapid-deployment. In NeurIPS."},{"key":"e_1_3_2_2_6_1","unstructured":"T. B. Brown etal 2020. Language models are few-shot learners. arXiv preprint arXiv:2005.14165 (2020).  T. B. Brown et al. 2020. Language models are few-shot learners. arXiv preprint arXiv:2005.14165 (2020)."},{"key":"e_1_3_2_2_7_1","doi-asserted-by":"crossref","unstructured":"Y. Cai Zh. Yao Zh. Dong etal 2020. ZeroQ: A Novel Zero Shot Quantization Framework. 2020 IEEE\/CVF CVPR (2020) 13166--13175.  Y. Cai Zh. Yao Zh. Dong et al. 2020. ZeroQ: A Novel Zero Shot Quantization Framework. 2020 IEEE\/CVF CVPR (2020) 13166--13175.","DOI":"10.1109\/CVPR42600.2020.01318"},{"key":"e_1_3_2_2_8_1","volume-title":"https:\/\/e.huawei.com\/en\/products\/intelligent-vision\/cameras\/software-defined-camera","author":"Camera Software-Defined","year":"2021","unstructured":"Software-Defined Camera . 2021. ( 2021 ). https:\/\/e.huawei.com\/en\/products\/intelligent-vision\/cameras\/software-defined-camera Software-Defined Camera. 2021. (2021). https:\/\/e.huawei.com\/en\/products\/intelligent-vision\/cameras\/software-defined-camera"},{"key":"e_1_3_2_2_9_1","volume-title":"Eyeriss: A Spatial Architecture for Energy-Efficient Dataflow for Convolutional Neural Networks","author":"Chen Y.H.","year":"2017","unstructured":"Y.H. Chen , J. Emer , and V. Sze . 2017 . Eyeriss: A Spatial Architecture for Energy-Efficient Dataflow for Convolutional Neural Networks . IEEE Micro ( 2017). Y.H. Chen, J. Emer, and V. Sze. 2017. Eyeriss: A Spatial Architecture for Energy-Efficient Dataflow for Convolutional Neural Networks. IEEE Micro (2017)."},{"key":"e_1_3_2_2_10_1","unstructured":"Y. Cheng D. Wang P. Zhou and T. Zhang. 2017. A survey of model compression and acceleration for deep neural networks. arXiv:1710.09282 (2017).  Y. Cheng D. Wang P. Zhou and T. Zhang. 2017. A survey of model compression and acceleration for deep neural networks. arXiv:1710.09282 (2017)."},{"key":"e_1_3_2_2_11_1","volume-title":"PACT: Parameterized Clipping Activation for Quantized Neural Networks. ArXiv","author":"Choi J.","year":"2018","unstructured":"J. Choi , Z. Wang , S. Venkataramani , 2018 . PACT: Parameterized Clipping Activation for Quantized Neural Networks. ArXiv , Vol. abs\/ 1805 .06085 (2018). J. Choi, Z. Wang, S. Venkataramani, et al. 2018. PACT: Parameterized Clipping Activation for Quantized Neural Networks. ArXiv , Vol. abs\/1805.06085 (2018)."},{"key":"e_1_3_2_2_12_1","volume-title":"Prema: A predictive multi-task scheduling algorithm for preemptible neural processing units","author":"Choi Yujeong","year":"2020","unstructured":"Yujeong Choi and Minsoo Rhu . 2020 . Prema: A predictive multi-task scheduling algorithm for preemptible neural processing units . In HPCA. IEEE , 220--233. Yujeong Choi and Minsoo Rhu. 2020. Prema: A predictive multi-task scheduling algorithm for preemptible neural processing units. In HPCA. IEEE, 220--233."},{"key":"e_1_3_2_2_13_1","doi-asserted-by":"crossref","unstructured":"Y. Choukroun E. Kravchik and P. Kisilev. 2019. Low-bit Quantization of Neural Networks for Efficient Inference. 2019 IEEE\/CVF ICCVW (2019) 3009--3018.  Y. Choukroun E. Kravchik and P. Kisilev. 2019. Low-bit Quantization of Neural Networks for Efficient Inference. 2019 IEEE\/CVF ICCVW (2019) 3009--3018.","DOI":"10.1109\/ICCVW.2019.00363"},{"key":"e_1_3_2_2_14_1","volume-title":"https:\/\/www.alibabacloud.com\/blog\/594214","author":"Computing Alibaba Edge","year":"2021","unstructured":"Alibaba Edge Computing . 2021. ( 2021 ). https:\/\/www.alibabacloud.com\/blog\/594214 Alibaba Edge Computing. 2021. (2021). https:\/\/www.alibabacloud.com\/blog\/594214"},{"key":"e_1_3_2_2_15_1","doi-asserted-by":"crossref","unstructured":"X. Dai P. Zhang B. Wu etal 2019. ChamNet: Towards Efficient Network Design Through Platform-Aware Model Adaptation. CVPR (2019) 11390--11399.  X. Dai P. Zhang B. Wu et al. 2019. ChamNet: Towards Efficient Network Design Through Platform-Aware Model Adaptation. CVPR (2019) 11390--11399.","DOI":"10.1109\/CVPR.2019.01166"},{"key":"e_1_3_2_2_16_1","volume-title":"https:\/\/support.huaweicloud.com\/en-us\/usermanual-modelartspro\/modelartspro_01_0078.html","author":"Pro Model Deployment ModelArts","year":"2021","unstructured":"ModelArts Pro Model Deployment . 2021. ( 2021 ). https:\/\/support.huaweicloud.com\/en-us\/usermanual-modelartspro\/modelartspro_01_0078.html ModelArts Pro Model Deployment. 2021. (2021). https:\/\/support.huaweicloud.com\/en-us\/usermanual-modelartspro\/modelartspro_01_0078.html"},{"key":"e_1_3_2_2_17_1","volume-title":"HAWQ: Hessian AWare Quantization of Neural Networks With Mixed-Precision. ICCV","author":"Dong Zh.","year":"2019","unstructured":"Zh. Dong , Zh. Yao , A. Gholami , 2019 . HAWQ: Hessian AWare Quantization of Neural Networks With Mixed-Precision. ICCV (2019), 293--302. Zh. Dong, Zh. Yao, A. Gholami, et al. 2019. HAWQ: Hessian AWare Quantization of Neural Networks With Mixed-Precision. ICCV (2019), 293--302."},{"key":"e_1_3_2_2_18_1","volume-title":"https:\/\/www.huaweicloud.com\/intl\/en-us\/ascend\/mindxedge","author":"Edge X","year":"2021","unstructured":"Mind X Edge . 2021. ( 2021 ). https:\/\/www.huaweicloud.com\/intl\/en-us\/ascend\/mindxedge MindX Edge. 2021. (2021). https:\/\/www.huaweicloud.com\/intl\/en-us\/ascend\/mindxedge"},{"key":"e_1_3_2_2_19_1","unstructured":"S. Esser J. McKinstry etal 2019. Learned step size quantization. arXiv preprint arXiv:1902.08153 (2019).  S. Esser J. McKinstry et al. 2019. Learned step size quantization. arXiv preprint arXiv:1902.08153 (2019)."},{"key":"e_1_3_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/EDGE50951.2020.00015"},{"key":"e_1_3_2_2_21_1","doi-asserted-by":"crossref","unstructured":"D. Gao X. He Z. Zhou etal 2020. Rethinking Pruning for Accelerating Deep Inference At the Edge. In 26th ACM SIGKDD. 155--164.  D. Gao X. He Z. Zhou et al. 2020. Rethinking Pruning for Accelerating Deep Inference At the Edge. In 26th ACM SIGKDD. 155--164.","DOI":"10.1145\/3394486.3403058"},{"key":"e_1_3_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1098\/rsta.2019.0155"},{"key":"e_1_3_2_2_23_1","volume-title":"Deep Compression: Compressing Deep Neural Network with Pruning, Trained Quantization and Huffman Coding. CoRR","author":"Han S.","year":"2016","unstructured":"S. Han , H. Mao , and W. Dally . 2016 . Deep Compression: Compressing Deep Neural Network with Pruning, Trained Quantization and Huffman Coding. CoRR , Vol. abs\/ 1510 .00149 (2016). S. Han, H. Mao, and W. Dally. 2016. Deep Compression: Compressing Deep Neural Network with Pruning, Trained Quantization and Huffman Coding. CoRR , Vol. abs\/1510.00149 (2016)."},{"key":"e_1_3_2_2_24_1","volume":"201","author":"He Y.","unstructured":"Y. He , X. Zhang , and J. Sun. 201 7. Channel Pruning for Accelerating Very Deep Neural Networks. ICCV (2017), 1398--1406. Y. He, X. Zhang, and J. Sun. 2017. Channel Pruning for Accelerating Very Deep Neural Networks. ICCV (2017), 1398--1406.","journal-title":"J. Sun."},{"key":"e_1_3_2_2_25_1","volume-title":"Benchmarking neural network robustness to common corruptions and perturbations. arXiv preprint arXiv:1903.12261","author":"Hendrycks D.","year":"2019","unstructured":"D. Hendrycks and Th. Dietterich . 2019. Benchmarking neural network robustness to common corruptions and perturbations. arXiv preprint arXiv:1903.12261 ( 2019 ). D. Hendrycks and Th. Dietterich. 2019. Benchmarking neural network robustness to common corruptions and perturbations. arXiv preprint arXiv:1903.12261 (2019)."},{"key":"e_1_3_2_2_26_1","volume":"201","author":"Hinton G. E.","unstructured":"G. E. Hinton , O. Vinyals , and J. Dean. 201 5. Distilling the Knowledge in a Neural Network. ArXiv , Vol. abs\/1503.02531 (2015). G. E. Hinton, O. Vinyals, and J. Dean. 2015. Distilling the Knowledge in a Neural Network. ArXiv , Vol. abs\/1503.02531 (2015).","journal-title":"J. Dean."},{"key":"e_1_3_2_2_27_1","doi-asserted-by":"crossref","unstructured":"Ch. Hu W. Bao D. Wang and F. Liu. 2019. Dynamic Adaptive DNN Surgery for Inference Acceleration on the Edge. IEEE INFOCOM (2019) 1423--1431.  Ch. Hu W. Bao D. Wang and F. Liu. 2019. Dynamic Adaptive DNN Surgery for Inference Acceleration on the Edge. IEEE INFOCOM (2019) 1423--1431.","DOI":"10.1109\/INFOCOM.2019.8737614"},{"key":"e_1_3_2_2_28_1","unstructured":"Gao Huang Danlu Chen T. Li etal 2018. Multi-Scale Dense Networks for Resource Efficient Image Classification. In ICLR.  Gao Huang Danlu Chen T. Li et al. 2018. Multi-Scale Dense Networks for Resource Efficient Image Classification. In ICLR."},{"key":"e_1_3_2_2_29_1","unstructured":"Intel. 2020. Intel Nervana. 2020. Nervana's Early Exit Inference. (2020). https:\/\/nervanasystems.github.io\/distiller\/algo_earlyexit.html  Intel. 2020. Intel Nervana. 2020. Nervana's Early Exit Inference. (2020). https:\/\/nervanasystems.github.io\/distiller\/algo_earlyexit.html"},{"key":"e_1_3_2_2_30_1","doi-asserted-by":"crossref","unstructured":"B. Jacob S. Kligys Bo Chen etal 2018. Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference. CVPR (2018).  B. Jacob S. Kligys Bo Chen et al. 2018. Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference. CVPR (2018).","DOI":"10.1109\/CVPR.2018.00286"},{"key":"e_1_3_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/3037697.3037698"},{"key":"e_1_3_2_2_32_1","volume-title":"Quantizing deep convolutional networks for efficient inference: A whitepaper. ArXiv","author":"Krishnamoorthi R.","year":"2018","unstructured":"R. Krishnamoorthi . 2018. Quantizing deep convolutional networks for efficient inference: A whitepaper. ArXiv , Vol. abs\/ 1806 .08342 ( 2018 ). R. Krishnamoorthi. 2018. Quantizing deep convolutional networks for efficient inference: A whitepaper. ArXiv , Vol. abs\/1806.08342 (2018)."},{"key":"e_1_3_2_2_33_1","doi-asserted-by":"crossref","unstructured":"Stefanos Laskaridis Stylianos I. Venieris etal 2020. SPINN: synergistic progressive inference of neural networks over device and cloud. MobiCom (2020).  Stefanos Laskaridis Stylianos I. Venieris et al. 2020. SPINN: synergistic progressive inference of neural networks over device and cloud. MobiCom (2020).","DOI":"10.1145\/3372224.3419194"},{"key":"e_1_3_2_2_34_1","volume-title":"Edge AI: On-Demand Accelerating Deep Neural Network Inference via Edge Computing. TWC","author":"Li En","year":"2020","unstructured":"En Li , Liekang Zeng , , 2020 . Edge AI: On-Demand Accelerating Deep Neural Network Inference via Edge Computing. TWC (2020). En Li, Liekang Zeng, , et al. 2020. Edge AI: On-Demand Accelerating Deep Neural Network Inference via Edge Computing. TWC (2020)."},{"key":"e_1_3_2_2_35_1","unstructured":"C. Michaelis B. Mitzkus R. Geirhos etal 2019. Benchmarking robustness in object detection: Autonomous driving when winter is coming. arXiv preprint arXiv:1907.07484 (2019).  C. Michaelis B. Mitzkus R. Geirhos et al. 2019. Benchmarking robustness in object detection: Autonomous driving when winter is coming. arXiv preprint arXiv:1907.07484 (2019)."},{"key":"e_1_3_2_2_36_1","volume-title":"Apprentice: Using Knowledge Distillation Techniques To Improve Low-Precision Network Accuracy. arXiv:1711.05852","author":"Mishra A.","year":"2018","unstructured":"A. Mishra and D. Marr . 2018 . Apprentice: Using Knowledge Distillation Techniques To Improve Low-Precision Network Accuracy. arXiv:1711.05852 (2018). A. Mishra and D. Marr. 2018. Apprentice: Using Knowledge Distillation Techniques To Improve Low-Precision Network Accuracy. arXiv:1711.05852 (2018)."},{"key":"e_1_3_2_2_37_1","unstructured":"Y. Nahshan B. Chmiel Ch. Baskin etal 2019. Loss Aware Post-training Quantization. ArXiv Vol. abs\/1911.07190 (2019).  Y. Nahshan B. Chmiel Ch. Baskin et al. 2019. Loss Aware Post-training Quantization. ArXiv Vol. abs\/1911.07190 (2019)."},{"key":"e_1_3_2_2_38_1","volume-title":"https:\/\/aws.amazon.com\/sagemaker\/neo\/","author":"SageMaker Neo Amazon","year":"2021","unstructured":"Amazon SageMaker Neo . 2021. ( 2021 ). https:\/\/aws.amazon.com\/sagemaker\/neo\/ Amazon SageMaker Neo. 2021. (2021). https:\/\/aws.amazon.com\/sagemaker\/neo\/"},{"key":"e_1_3_2_2_39_1","volume-title":"https:\/\/github.com\/PaddlePaddle","author":"PaddlePaddle Baidu's","year":"2021","unstructured":"Baidu's PaddlePaddle . 2021. ( 2021 ). https:\/\/github.com\/PaddlePaddle Baidu's PaddlePaddle. 2021. (2021). https:\/\/github.com\/PaddlePaddle"},{"key":"e_1_3_2_2_40_1","volume-title":"Timeloop: A systematic approach to dnn accelerator evaluation. In ISPASS.","author":"Angshuman Parashar","year":"2019","unstructured":"Angshuman Parashar et al. 2019 . Timeloop: A systematic approach to dnn accelerator evaluation. In ISPASS. Angshuman Parashar et al. 2019. Timeloop: A systematic approach to dnn accelerator evaluation. In ISPASS."},{"key":"e_1_3_2_2_41_1","volume-title":"https:\/\/e.huawei.com\/en\/products\/cloud-computing-dc\/atlas\/atlas-200\/","author":"Platform HiLens","year":"2021","unstructured":"HiLens Platform and Edge Device . 2021. ( 2021 ). https:\/\/e.huawei.com\/en\/products\/cloud-computing-dc\/atlas\/atlas-200\/ HiLens Platform and Edge Device. 2021. (2021). https:\/\/e.huawei.com\/en\/products\/cloud-computing-dc\/atlas\/atlas-200\/"},{"key":"e_1_3_2_2_42_1","volume-title":"https:\/\/github.com\/Tencent\/PocketFlow","author":"PocketFlow Tencent","year":"2021","unstructured":"Tencent PocketFlow . 2021. ( 2021 ). https:\/\/github.com\/Tencent\/PocketFlow Tencent PocketFlow. 2021. (2021). https:\/\/github.com\/Tencent\/PocketFlow"},{"key":"e_1_3_2_2_43_1","volume-title":"Model compression via distillation and quantization. ArXiv","author":"Polino A.","year":"2018","unstructured":"A. Polino , R. Pascanu , and Dan Alistarh . 2018. Model compression via distillation and quantization. ArXiv , Vol. abs\/ 1802 .05668 ( 2018 ). A. Polino, R. Pascanu, and Dan Alistarh. 2018. Model compression via distillation and quantization. ArXiv , Vol. abs\/1802.05668 (2018)."},{"key":"e_1_3_2_2_44_1","unstructured":"M. Rusci A. Capotondi and L. Benini. 2020. Memory-Driven Mixed Low Precision Quantization For Enabling Deep Network Inference On Microcontrollers. ArXiv Vol. abs\/1905.13082 (2020).  M. Rusci A. Capotondi and L. Benini. 2020. Memory-Driven Mixed Low Precision Quantization For Enabling Deep Network Inference On Microcontrollers. ArXiv Vol. abs\/1905.13082 (2020)."},{"key":"e_1_3_2_2_45_1","unstructured":"A. Samajdar Y. Zhu P. Whatmough etal 2018. SCALE-Sim: Systolic CNN Accelerator. ArXiv Vol. abs\/1811.02883 (2018).  A. Samajdar Y. Zhu P. Whatmough et al. 2018. SCALE-Sim: Systolic CNN Accelerator. ArXiv Vol. abs\/1811.02883 (2018)."},{"key":"e_1_3_2_2_46_1","doi-asserted-by":"crossref","unstructured":"Y. Shoham and A. Gersho. 1988. Efficient bit allocation for an arbitrary set of quantizers. IEEE Trans. Acoustics Speech and Signal Processing Vol. 36 (1988).  Y. Shoham and A. Gersho. 1988. Efficient bit allocation for an arbitrary set of quantizers. IEEE Trans. Acoustics Speech and Signal Processing Vol. 36 (1988).","DOI":"10.1109\/29.90373"},{"key":"e_1_3_2_2_47_1","volume-title":"Platform-Aware Neural Architecture Search for Mobile. CVPR","author":"Tan M.","year":"2018","unstructured":"M. Tan , Bo Chen , R. Pang , V. Vasudevan , and Quoc V. Le . 2018 . Platform-Aware Neural Architecture Search for Mobile. CVPR 2018 . M. Tan, Bo Chen, R. Pang, V. Vasudevan, and Quoc V. Le. 2018. Platform-Aware Neural Architecture Search for Mobile. CVPR 2018."},{"key":"e_1_3_2_2_48_1","volume-title":"Branchynet: Fast inference via early exiting from deep neural networks","author":"Teerapittayanon S.","year":"2016","unstructured":"S. Teerapittayanon , B. McDanel , and H.-T. Kung . 2016 . Branchynet: Fast inference via early exiting from deep neural networks . In ICPR. IEEE , 2464--2469. S. Teerapittayanon, B. McDanel, and H.-T. Kung. 2016. Branchynet: Fast inference via early exiting from deep neural networks. In ICPR. IEEE, 2464--2469."},{"key":"e_1_3_2_2_49_1","unstructured":"Tenstorrent. 2020. Tenstorrent's Grayskull AI Chip. (2020). https:\/\/www.tenstorrent.com\/technology\/  Tenstorrent. 2020. Tenstorrent's Grayskull AI Chip. (2020). https:\/\/www.tenstorrent.com\/technology\/"},{"key":"e_1_3_2_2_50_1","doi-asserted-by":"crossref","unstructured":"J. Wang J. Zhang W. Bao etal 2018. Not just privacy: Improving performance of private deep learning in mobile cloud. In 24th ACM SIGKDD. 2407--2416.  J. Wang J. Zhang W. Bao et al. 2018. Not just privacy: Improving performance of private deep learning in mobile cloud. In 24th ACM SIGKDD. 2407--2416.","DOI":"10.1145\/3219819.3220106"},{"key":"e_1_3_2_2_51_1","volume-title":"Haq: Hardware-aware automated quantization with mixed precision. In CVPR. 8612--8620.","author":"Wang K.","year":"2019","unstructured":"K. Wang , Zh. Liu , Y. Lin , J. Lin , and Song H . 2019 . Haq: Hardware-aware automated quantization with mixed precision. In CVPR. 8612--8620. K. Wang, Zh. Liu, Y. Lin, J. Lin, and Song H. 2019. Haq: Hardware-aware automated quantization with mixed precision. In CVPR. 8612--8620."},{"key":"e_1_3_2_2_52_1","unstructured":"B. Wu Y. Wang P. Zhang etal 2018. Mixed Precision Quantization of ConvNets via Differentiable Neural Architecture Search. ArXiv Vol. abs\/1812.00090 (2018).  B. Wu Y. Wang P. Zhang et al. 2018. Mixed Precision Quantization of ConvNets via Differentiable Neural Architecture Search. ArXiv Vol. abs\/1812.00090 (2018)."},{"key":"e_1_3_2_2_53_1","volume-title":"Accelergy: An architecture-level energy estimation methodology for accelerator designs. In ICCAD.","author":"Wu Y. N.","year":"2019","unstructured":"Y. N. Wu , J. S. Emer , and V. Sze . 2019 . Accelergy: An architecture-level energy estimation methodology for accelerator designs. In ICCAD. Y. N. Wu, J. S. Emer, and V. Sze. 2019. Accelergy: An architecture-level energy estimation methodology for accelerator designs. In ICCAD."},{"key":"e_1_3_2_2_54_1","unstructured":"Haichuan Yang et al. 2018a. Energy-constrained compression for deep neural networks via weighted sparse projection and layer input masking. arXiv preprint arXiv:1806.04321 (2018).  Haichuan Yang et al. 2018a. Energy-constrained compression for deep neural networks via weighted sparse projection and layer input masking. arXiv preprint arXiv:1806.04321 (2018)."},{"key":"e_1_3_2_2_55_1","doi-asserted-by":"crossref","unstructured":"T.-J. Yang A. Howard B. Chen etal 2018b. NetAdapt: Platform-Aware Neural Network Adaptation for Mobile Applications. ArXiv Vol. abs\/1804.03230 (2018).  T.-J. Yang A. Howard B. Chen et al. 2018b. NetAdapt: Platform-Aware Neural Network Adaptation for Mobile Applications. ArXiv Vol. abs\/1804.03230 (2018).","DOI":"10.1007\/978-3-030-01249-6_18"},{"key":"e_1_3_2_2_56_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01237-3_23"},{"key":"e_1_3_2_2_57_1","volume-title":"SCAN: A Scalable Neural Networks Framework Towards Compact and Efficient Models. In NeurIPS.","author":"Zhang Linfeng","year":"2019","unstructured":"Linfeng Zhang , Zhanhong Tan , 2019 . SCAN: A Scalable Neural Networks Framework Towards Compact and Efficient Models. In NeurIPS. Linfeng Zhang, Zhanhong Tan, et al. 2019. SCAN: A Scalable Neural Networks Framework Towards Compact and Efficient Models. In NeurIPS."},{"key":"e_1_3_2_2_58_1","unstructured":"Shigeng Zhang Yinggang Li etal 2020. Towards Real-time Cooperative Deep Inference over the Cloud and Edge End Devices. IMWUT (2020).  Shigeng Zhang Yinggang Li et al. 2020. Towards Real-time Cooperative Deep Inference over the Cloud and Edge End Devices. IMWUT (2020)."},{"key":"e_1_3_2_2_59_1","volume-title":"ICML","author":"Zhao R.","year":"2019","unstructured":"R. Zhao , Y. Hu , J. Dotzel , C. D. Sa , and Z. Zhang . 2019. Improving Neural Network Quantization without Retraining using Outlier Channel Splitting . In ICML 2019 . R. Zhao, Y. Hu, J. Dotzel, C. D. Sa, and Z. Zhang. 2019. Improving Neural Network Quantization without Retraining using Outlier Channel Splitting. In ICML 2019."},{"key":"e_1_3_2_2_60_1","volume-title":"Quantization: Pareto-optimal Bit Allocation for Deep CNNs Compression.","author":"Zhe W.","year":"2020","unstructured":"W. Zhe , J. Lin , M. M. Sabry Aly , S. Young , V. Chandrasekhar , and B. Girod . 2020 . Towards Effective 2-bit Quantization: Pareto-optimal Bit Allocation for Deep CNNs Compression. (2020). https:\/\/openreview.net\/forum?id=H1eKT1SFvH W. Zhe, J. Lin, M. M. Sabry Aly, S. Young, V. Chandrasekhar, and B. Girod. 2020. Towards Effective 2-bit Quantization: Pareto-optimal Bit Allocation for Deep CNNs Compression. (2020). https:\/\/openreview.net\/forum?id=H1eKT1SFvH"},{"key":"e_1_3_2_2_61_1","volume":"201","author":"Zhou H.Y.","unstructured":"H.Y. Zhou , B.B. Gao , and J. Wu. 201 7. Adaptive feeding: Achieving fast and accurate detections by adaptively combining object detectors. In ICCV. H.Y. Zhou, B.B. Gao, and J. Wu. 2017. Adaptive feeding: Achieving fast and accurate detections by adaptively combining object detectors. In ICCV.","journal-title":"J. Wu."},{"key":"e_1_3_2_2_62_1","unstructured":"Sh. Zhou Z. Ni etal 2016. DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients. ArXiv Vol. abs\/1606.06160 (2016).  Sh. Zhou Z. Ni et al. 2016. DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients. ArXiv Vol. abs\/1606.06160 (2016)."},{"key":"e_1_3_2_2_63_1","unstructured":"N. Zmora G. Jacob etal 2019. Neural Network Distiller: A Python Package For DNN Compression Research. (October 2019). https:\/\/arxiv.org\/abs\/1910.12232  N. Zmora G. Jacob et al. 2019. Neural Network Distiller: A Python Package For DNN Compression Research. (October 2019). https:\/\/arxiv.org\/abs\/1910.12232"}],"event":{"name":"KDD '21: The 27th ACM SIGKDD Conference on Knowledge Discovery and Data Mining","location":"Virtual Event Singapore","acronym":"KDD '21","sponsor":["SIGMOD ACM Special Interest Group on Management of Data","SIGKDD ACM Special Interest Group on Knowledge Discovery in Data"]},"container-title":["Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery &amp; Data Mining"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3447548.3467078","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3447548.3467078","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T21:25:11Z","timestamp":1750195511000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3447548.3467078"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,8,14]]},"references-count":63,"alternative-id":["10.1145\/3447548.3467078","10.1145\/3447548"],"URL":"https:\/\/doi.org\/10.1145\/3447548.3467078","relation":{},"subject":[],"published":{"date-parts":[[2021,8,14]]},"assertion":[{"value":"2021-08-14","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}