{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,9]],"date-time":"2026-01-09T20:20:05Z","timestamp":1767990005793,"version":"3.49.0"},"reference-count":62,"publisher":"Association for Computing Machinery (ACM)","issue":"5","license":[{"start":{"date-parts":[[2022,9,30]],"date-time":"2022-09-30T00:00:00Z","timestamp":1664496000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc-nd\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Embed. Comput. Syst."],"published-print":{"date-parts":[[2022,9,30]]},"abstract":"<jats:p>Machine learning has shown tremendous success in a large variety of applications. The evolution of machine-learning applications from cloud-based systems to mobile and embedded devices has shifted the focus from only quality-related aspects towards the resource demand of machine learning. For embedded systems, dedicated accelerator hardware promises the energy-efficient execution of neural network inferences. Their precise resource demand in terms of execution time and power demand, however, is undocumented. Developers, therefore, face the challenge to fine-tune their neural networks such that their resource demand matches the available budgets.<\/jats:p>\n          <jats:p>\n            This article presents\n            <jats:sc>Precious<\/jats:sc>\n            , a comprehensive approach to estimate the resource demand of an embedded neural network accelerator. We generate randomised neural networks, analyse them statically, execute them on an embedded accelerator while measuring their actual power draw and execution time, and train estimators that map the statically analysed neural network properties to the measured resource demand. In addition, this article provides an in-depth analysis of the neural networks\u2019 resource demands and the responsible network properties. We demonstrate that the estimation error of\n            <jats:sc>Precious<\/jats:sc>\n            can be below 1.5% for both power draw and execution time. Furthermore, we discuss what estimator accuracy is practically achievable and how much effort is required to achieve sufficient accuracy.\n          <\/jats:p>","DOI":"10.1145\/3520132","type":"journal-article","created":{"date-parts":[[2022,3,21]],"date-time":"2022-03-21T12:37:55Z","timestamp":1647866275000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":7,"title":["Resource-demand Estimation for Edge Tensor Processing Units"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8725-3454","authenticated-orcid":false,"given":"Benedict","family":"Herzog","sequence":"first","affiliation":[{"name":"Ruhr-Universit\u00e4t Bochum (RUB), Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Stefan","family":"Reif","sequence":"additional","affiliation":[{"name":"Friedrich-Alexander-Universit\u00e4t Erlangen-N\u00fcrnberg (FAU), Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Judith","family":"Hemp","sequence":"additional","affiliation":[{"name":"Friedrich-Alexander-Universit\u00e4t Erlangen-N\u00fcrnberg (FAU), Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Timo","family":"H\u00f6nig","sequence":"additional","affiliation":[{"name":"Ruhr-Universit\u00e4t Bochum (RUB), Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wolfgang","family":"Schr\u00f6der-Preikschat","sequence":"additional","affiliation":[{"name":"Friedrich-Alexander-Universit\u00e4t Erlangen-N\u00fcrnberg (FAU), Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,10,8]]},"reference":[{"key":"e_1_3_2_2_2","first-page":"265","volume-title":"Proceedings of the 12th Symposium on Operating Systems Design and Implementation (OSDI\u201916)","author":"Abadi Mart\u00edn","year":"2016","unstructured":"Mart\u00edn Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. 2016. TensorFlow: A system for large-scale machine learning. In Proceedings of the 12th Symposium on Operating Systems Design and Implementation (OSDI\u201916). USENIX, 265\u2013283."},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/JETCAS.2019.2911899"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/3005348"},{"key":"e_1_3_2_5_2","first-page":"281","article-title":"Random search for hyper-parameter optimization","volume":"13","author":"Bergstra James","year":"2012","unstructured":"James Bergstra and Yoshua Bengio. 2012. Random search for hyper-parameter optimization. J. Mach. Learn. Res. 13, Feb. (2012), 281\u2013305.","journal-title":"J. Mach. Learn. Res."},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/ARITH.2019.00022"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/MSPEC.2019.8701189"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/MSPEC.2020.9126102"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/ASPDAC.2018.8297294"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/2897937.2898092"},{"key":"e_1_3_2_11_2","unstructured":"NVIDIA Corp.2015. GPU-based Deep Learning Inference: A Performance and Power Analysis. Retrieved from https:\/\/www.nvidia.com\/content\/tegra\/embedded-systems\/pdf\/jetson_tx1_whitepaper.pdf."},{"key":"e_1_3_2_12_2","unstructured":"NVIDIA Corp.2021. Jetson Nano Developer Kit. Retrieved from https:\/\/developer.nvidia.com\/embedded\/jetson-nano-developer-kit."},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2018.112130030"},{"key":"e_1_3_2_14_2","unstructured":"Scikit-Learn Developers. 2021. sklearn.ensemble: Ensemble Methods. Retrieved from https:\/\/scikit-learn.org\/stable\/modules\/classes.html#module-sklearn.ensemble."},{"key":"e_1_3_2_15_2","unstructured":"Scikit-Learn Developers. 2021. sklearn.linear_model: Linear Models. Retrieved from https:\/\/scikit-learn.org\/stable\/modules\/classes.html#module-sklearn.linear_model."},{"key":"e_1_3_2_16_2","unstructured":"Analog Devices. 2020. LTC2991. Retrieved from https:\/\/www.analog.com\/en\/products\/ltc2991.html."},{"key":"e_1_3_2_17_2","first-page":"343","volume-title":"Proceedings of the 15th Symposium on Networked Systems Design and Implementation (NSDI\u201918)","author":"Dong Mo","year":"2018","unstructured":"Mo Dong, Tong Meng, Doron Zarchy, Engin Arslan, Yossi Gilad, Brighten Godfrey, and Michael Schapira. 2018. PCC vivace: Online-learning congestion control. In Proceedings of the 15th Symposium on Networked Systems Design and Implementation (NSDI\u201918). USENIX, 343\u2013356."},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC50251.2020.00025"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/3240765.3243484"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3409977"},{"key":"e_1_3_2_21_2","unstructured":"Suyog Gupta and Mingxing Tan. 2019. EfficientNet-EdgeTPU: Creating Accelerator-optimized Neural Networks with AutoML. Retrieved from https:\/\/ai.googleblog.com\/2019\/08\/efficientnet-edgetpu-creating.html."},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2018.00059"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3297280.3297338"},{"key":"e_1_3_2_24_2","first-page":"288","volume-title":"Proceedings of the Perceptual Image Restoration and Manipulation Workshop and Challenge (PIRM\u201918)","author":"Ignatov Andrey","year":"2018","unstructured":"Andrey Ignatov, Radu Timofte, William Chou, Ke Wang, Max Wu, Tim Hartley, and Luc Van Gool. 2018. AI benchmark: Running deep neural networks on Android smartphones. In Proceedings of the Perceptual Image Restoration and Manipulation Workshop and Challenge (PIRM\u201918). Springer International Publishing, 288\u2013314."},{"key":"e_1_3_2_25_2","unstructured":"Canaan Inc. 2021. Kendryte K210. Retrieved from https:\/\/canaan.io\/product\/kendryteai."},{"key":"e_1_3_2_26_2","unstructured":"Intel Inc.2020. Intel Neural Compute Stick 2. Retrieved from https:\/\/software.intel.com\/en-us\/neural-compute-stick."},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2018.032271057"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080246"},{"key":"e_1_3_2_29_2","unstructured":"Dhiraj Kalamkar Dheevatsa Mudigere Naveen Mellempudi Dipankar Das Kunal Banerjee Sasikanth Avancha Dharma Teja Vooturi Nataraj Jammalamadaka Jianyu Huang Hector Yuen Jiyan Yang Jongsoo Park Alexander Heinecke Evangelos Georganas Sudarshan Srinivasan Abhisek Kundu Misha Smelyanskiy Bharat Kaul and Pradeep Dubey. 2019. A study of BFLOAT16 for deep learning training. Retrieved from https:\/\/arxiv.org\/abs\/1905.12322."},{"key":"e_1_3_2_30_2","first-page":"1","volume-title":"Proceedings of the Workshop on ML for Systems at the 33rd Conference on Neural Information Processing Systems (NeurIPS\u201919)","author":"Kaufman Samuel","year":"2019","unstructured":"Samuel Kaufman, Phitchaya Phothilimtha, and Mike Burrows. 2019. Learned TPU cost model for XLA tensor programs. In Proceedings of the Workshop on ML for Systems at the 33rd Conference on Neural Information Processing Systems (NeurIPS\u201919). 1\u20136."},{"key":"e_1_3_2_31_2","unstructured":"Keras Special Interest Group. 2020. Core Layers. Retrieved from https:\/\/keras.io\/api\/layers\/core_layers\/."},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPEC43674.2020.9286209"},{"key":"e_1_3_2_33_2","unstructured":"Alexandre Lacoste Alexandra Luccioni Victor Schmidt and Thomas Dandres. 2019. Quantifying the carbon emissions of machine learning. Retrieved from https:\/\/arxiv.org\/abs\/1910.09700."},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/MPRV.2017.2940968"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/2699343.2699349"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/2964284.2973801"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.5555\/3408352.3408370"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/BDCloud-SocialCom-SustainCom.2016.76"},{"key":"e_1_3_2_39_2","volume-title":"Proceedings of the 1st Workshop on Accelerated Machine Learning (AccML\u201920) at the European Network on High-performance Embedded Architecture and Compilation (HiPEAC\u201920)","author":"Libutti Leandro Ariel","year":"2020","unstructured":"Leandro Ariel Libutti, Francisco Igual, Luis Pi\u00f1uel, Laura De Giusti, and Marcelo Naiouf. 2020. Benchmarking performance and power of USB accelerators for inference with MLPerf. In Proceedings of the 1st Workshop on Accelerated Machine Learning (AccML\u201920) at the European Network on High-performance Embedded Architecture and Compilation (HiPEAC\u201920)."},{"key":"e_1_3_2_40_2","unstructured":"Google LLC. 2020. Deploy Machine Learning Models on Mobile and IoT Devices. Retrieved from https:\/\/www.tensorflow.org\/lite\/."},{"key":"e_1_3_2_41_2","unstructured":"Google LLC. 2020. Edge TPU Compiler. Retrieved from https:\/\/coral.ai\/docs\/edgetpu\/compiler\/."},{"key":"e_1_3_2_42_2","unstructured":"Google LLC. 2020. Get Started with TensorFlow Lite. Retrieved from https:\/\/www.tensorflow.org\/lite\/guide\/get_started."},{"key":"e_1_3_2_43_2","unstructured":"Google LLC. 2020. Get Started with the USB Accelerator. Retrieved from https:\/\/coral.ai\/docs\/accelerator\/get-started\/."},{"key":"e_1_3_2_44_2","unstructured":"Google LLC. 2020. Tensorflow Keras. Retrieved from https:\/\/www.tensorflow.org\/versions\/r2.0\/api_docs\/python\/tf\/keras."},{"key":"e_1_3_2_45_2","unstructured":"Google LLC. 2020. USB Accelerator. Retrieved from https:\/\/www.coral.ai\/products\/accelerator."},{"key":"e_1_3_2_46_2","unstructured":"Google LLC. 2020. USB Accelerator Datasheet. Retrieved from https:\/\/coral.ai\/docs\/accelerator\/datasheet\/."},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3123266.3123389"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2020.2974843"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2015.7095813"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC50251.2020.00011"},{"key":"e_1_3_2_51_2","unstructured":"David Patterson Joseph Gonzalez Quoc Le Chen Liang Lluis-Miquel Munguia Daniel Rothchild David So Maud Texier and Jeff Dean. 2021. Carbon emissions and large neural network training. Retrieved from https:\/\/arxiv.org\/abs\/2104.10350."},{"key":"e_1_3_2_52_2","first-page":"1","volume-title":"Proceedings of the 1st International Workshop on Benchmarking Machine Learning Workloads on Emerging Hardware (Challenge\u201920)","author":"Reif Stefan","year":"2020","unstructured":"Stefan Reif, Benedict Herzog, Judith Hemp, Timo H\u00f6nig, and Wolfgang Schr\u00f6der-Preikschat. 2020. Precious: Resource-demand estimation for embedded neural network accelerators. In Proceedings of the 1st International Workshop on Benchmarking Machine Learning Workloads on Emerging Hardware (Challenge\u201920). 1\u20139."},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1145\/3447555.3466579"},{"key":"e_1_3_2_54_2","unstructured":"Crefeda Faviola Rodrigues Graham Riley and Mikel Luj\u00e1n. 2020. Energy predictive models for convolutional neural networks on mobile platforms. Retrieved from https:\/\/arxiv.org\/abs\/2004.05137."},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.1145\/3381831"},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.5555\/3195638.3195697"},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISORC.2017.10"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature16961"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1355"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.23919\/DATE.2018.8342167"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2019.00048"},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2015.7056063"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1145\/3308558.3313591"}],"container-title":["ACM Transactions on Embedded Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3520132","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3520132","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3520132","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T18:10:32Z","timestamp":1750183832000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3520132"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,9,30]]},"references-count":62,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2022,9,30]]}},"alternative-id":["10.1145\/3520132"],"URL":"https:\/\/doi.org\/10.1145\/3520132","relation":{},"ISSN":["1539-9087","1558-3465"],"issn-type":[{"value":"1539-9087","type":"print"},{"value":"1558-3465","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,9,30]]},"assertion":[{"value":"2021-07-15","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-02-20","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-10-08","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}