{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,27]],"date-time":"2026-08-27T15:22:04Z","timestamp":1787844124952,"version":"build-2784847793"},"reference-count":56,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2022,12,1]],"date-time":"2022-12-01T00:00:00Z","timestamp":1669852800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Ministry of Education, India\/PMRF"},{"name":"Department of Science and Technology, India\/ICPS"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Meas. Anal. Comput. Syst."],"published-print":{"date-parts":[[2022,12]]},"abstract":"<jats:p>Deep Neural Networks (DNNs) have had a significant impact on domains like autonomous vehicles and smart cities through low-latency inferencing on edge computing devices close to the data source. However, DNN training on the edge is poorly explored. Techniques like federated learning and the growing capacity of GPU-accelerated edge devices like NVIDIA Jetson motivate the need for a holistic characterization of DNN training on the edge. Training DNNs is resource-intensive and can stress an edge's GPU, CPU, memory and storage capacities. Edge devices also have different resources compared to workstations and servers, such as slower shared memory and diverse storage media. Here, we perform a principled study of DNN training on individual devices of three contemporary Jetson device types: AGX Xavier, Xavier NX and Nano for three diverse DNN model--dataset combinations. We vary device and training parameters such as I\/O pipelining and parallelism, storage media, mini-batch sizes and power modes, and examine their effect on CPU and GPU utilization, fetch stalls, training time, energy usage, and variability. Our analysis exposes several resource inter-dependencies and counter-intuitive insights, while also helping quantify known wisdom. Our rigorous study can help tune the training performance on the edge, trade-off time and energy usage on constrained devices, and even select an ideal edge hardware for a DNN workload, and, in future, extend to federated learning too. As an illustration, we use these results to build a simple model to predict the training time and energy per epoch for any given DNN across different power modes, with minimal additional profiling.<\/jats:p>","DOI":"10.1145\/3570604","type":"journal-article","created":{"date-parts":[[2022,12,8]],"date-time":"2022-12-08T20:20:10Z","timestamp":1670530810000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":23,"title":["Characterizing the Performance of Accelerated Jetson Edge Devices for Training Deep Learning Models"],"prefix":"10.1145","volume":"6","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7490-3128","authenticated-orcid":false,"given":"Prashanthi","family":"S.K","sequence":"first","affiliation":[{"name":"Indian Institute of Science, Bangalore, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1270-7460","authenticated-orcid":false,"given":"Sai Anuroop","family":"Kesanapalli","sequence":"additional","affiliation":[{"name":"Indian Institute of Science, Bangalore, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4140-7774","authenticated-orcid":false,"given":"Yogesh","family":"Simmhan","sequence":"additional","affiliation":[{"name":"Indian Institute of Science, Bangalore, India"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,12,8]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/3434770.3459729"},{"key":"e_1_2_1_2_1","volume-title":"IEEE Intl. Conf. on Fog and Mobile Edge Comp. (FMEC).","author":"Hazem","unstructured":"Hazem A. Abdelhafez and Matei Ripeanu. 2019. Studying the Impact of CPU and Memory Controller Frequencies on Power Consumption of the Jetson TX1 . In IEEE Intl. Conf. on Fog and Mobile Edge Comp. (FMEC). Hazem A. Abdelhafez and Matei Ripeanu. 2019. Studying the Impact of CPU and Memory Controller Frequencies on Power Consumption of the Jetson TX1. In IEEE Intl. Conf. on Fog and Mobile Edge Comp. (FMEC)."},{"key":"e_1_2_1_3_1","unstructured":"Assemblyai. 2022. TF v\/s Pytorch. https:\/\/www.assemblyai.com\/blog\/pytorch-vs-tensorflow-in-2022\/.  Assemblyai. 2022. TF v\/s Pytorch. https:\/\/www.assemblyai.com\/blog\/pytorch-vs-tensorflow-in-2022\/."},{"key":"e_1_2_1_4_1","volume-title":"DeepEdgeBench: Benchmarking Deep Neural Networks on Edge Devices. In IEEE International Conference on Cloud Engineering.","author":"Baller S.","unstructured":"S. Baller , A. Jindal , M. Chadha , and M. Gerndt . 2021 . DeepEdgeBench: Benchmarking Deep Neural Networks on Edge Devices. In IEEE International Conference on Cloud Engineering. S. Baller, A. Jindal, M. Chadha, and M. Gerndt. 2021. DeepEdgeBench: Benchmarking Deep Neural Networks on Edge Devices. In IEEE International Conference on Cloud Engineering."},{"key":"e_1_2_1_5_1","volume-title":"Neural networks: Tricks of the trade","author":"Bengio Yoshua","unstructured":"Yoshua Bengio . 2012. Practical recommendations for gradient-based training of deep architectures . In Neural networks: Tricks of the trade . Springer , 437--478. Yoshua Bengio. 2012. Practical recommendations for gradient-based training of deep architectures. In Neural networks: Tricks of the trade. Springer, 437--478."},{"key":"e_1_2_1_6_1","volume-title":"Titouan Parcollet, Pedro Porto Buarque de Gusm\u00e3o, and Nicholas D. Lane.","author":"Beutel Daniel J.","year":"2020","unstructured":"Daniel J. Beutel , Taner Topal , Akhil Mathur , Xinchi Qiu , Javier Fernandez-Marques , Yan Gao , Lorenzo Sani , Kwing Hei Li , Titouan Parcollet, Pedro Porto Buarque de Gusm\u00e3o, and Nicholas D. Lane. 2020 . Flower : A friendly federated learning research framework. arXiv preprint arXiv:2007.14390 (2020). Daniel J. Beutel, Taner Topal, Akhil Mathur, Xinchi Qiu, Javier Fernandez-Marques, Yan Gao, Lorenzo Sani, Kwing Hei Li, Titouan Parcollet, Pedro Porto Buarque de Gusm\u00e3o, and Nicholas D. Lane. 2020. Flower: A friendly federated learning research framework. arXiv preprint arXiv:2007.14390 (2020)."},{"key":"e_1_2_1_7_1","volume-title":"David Petrou, Daniel Ramage, and Jason Roselander.","author":"Bonawitz Keith","year":"2019","unstructured":"Keith Bonawitz , Hubert Eichner , Wolfgang Grieskamp , Dzmitry Huba , Alex Ingerman , Vladimir Ivanov , Chlo\u00e9 Kiddon , Jakub Konen\u00fd , Stefano Mazzocchi , Brendan McMahan , Timon Van Overveldt , David Petrou, Daniel Ramage, and Jason Roselander. 2019 . Towards Federated Learning at Scale : System Design. In Proceedings of Machine Learning and Systems, A. Talwalkar, V. Smith, and M. Zaharia (Eds .), Vol. 1 . 374--388. https:\/\/proceedings.mlsys.org\/paper\/2019\/file\/ bd686fd640be98efaae0091fa301e613-Paper.pdf Keith Bonawitz, Hubert Eichner, Wolfgang Grieskamp, Dzmitry Huba, Alex Ingerman, Vladimir Ivanov, Chlo\u00e9 Kiddon, Jakub Konen\u00fd, Stefano Mazzocchi, Brendan McMahan, Timon Van Overveldt, David Petrou, Daniel Ramage, and Jason Roselander. 2019. Towards Federated Learning at Scale: System Design. In Proceedings of Machine Learning and Systems, A. Talwalkar, V. Smith, and M. Zaharia (Eds.), Vol. 1. 374--388. https:\/\/proceedings.mlsys.org\/paper\/2019\/file\/ bd686fd640be98efaae0091fa301e613-Paper.pdf"},{"key":"e_1_2_1_8_1","unstructured":"Shubham Chandel. 2022. Pytorch Model Summary. https:\/\/github.com\/sksq96\/pytorch-summary.  Shubham Chandel. 2022. Pytorch Model Summary. https:\/\/github.com\/sksq96\/pytorch-summary."},{"key":"e_1_2_1_9_1","volume-title":"On large-cohort training for federated learning. Advances in Neural Information Processing Systems 34","author":"Charles Zachary","year":"2021","unstructured":"Zachary Charles , Zachary Garrett , Zhouyuan Huo , Sergei Shmulyian , and Virginia Smith . 2021. On large-cohort training for federated learning. Advances in Neural Information Processing Systems 34 ( 2021 ). Zachary Charles, Zachary Garrett, Zhouyuan Huo, Sergei Shmulyian, and Virginia Smith. 2021. On large-cohort training for federated learning. Advances in Neural Information Processing Systems 34 (2021)."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2019.2921977"},{"key":"e_1_2_1_11_1","volume-title":"Demon: Improved Neural Network Training with Momentum Decay. In ICASSP 2022--2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 3958--3962","author":"Chen John","year":"2022","unstructured":"John Chen , Cameron Wolfe , Zhao Li , and Anastasios Kyrillidis . 2022 . Demon: Improved Neural Network Training with Momentum Decay. In ICASSP 2022--2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 3958--3962 . John Chen, Cameron Wolfe, Zhao Li, and Anastasios Kyrillidis. 2022. Demon: Improved Neural Network Training with Momentum Decay. In ICASSP 2022--2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 3958--3962."},{"key":"e_1_2_1_12_1","volume-title":"A survey on an emerging area: Deep learning for smart city data","author":"Chen Qi","year":"2019","unstructured":"Qi Chen , Wei Wang , Fangyu Wu , Suparna De , Ruili Wang , Bailing Zhang , and Xin Huang . 2019. A survey on an emerging area: Deep learning for smart city data . IEEE Transactions on Emerging Topics in Computational Intelligence ( 2019 ). Qi Chen, Wei Wang, Fangyu Wu, Suparna De, Ruili Wang, Bailing Zhang, and Xin Huang. 2019. A survey on an emerging area: Deep learning for smart city data. IEEE Transactions on Emerging Topics in Computational Intelligence (2019)."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.12262"},{"key":"e_1_2_1_14_1","volume-title":"On the computational inefficiency of large batch sizes for stochastic gradient descent. arXiv preprint arXiv:1811.12941","author":"Golmant Noah","year":"2018","unstructured":"Noah Golmant , Nikita Vemuri , Zhewei Yao , Vladimir Feinberg , Amir Gholami , Kai Rothauge , Michael W Mahoney , and Joseph Gonzalez . 2018. On the computational inefficiency of large batch sizes for stochastic gradient descent. arXiv preprint arXiv:1811.12941 ( 2018 ). Noah Golmant, Nikita Vemuri, Zhewei Yao, Vladimir Feinberg, Amir Gholami, Kai Rothauge, Michael W Mahoney, and Joseph Gonzalez. 2018. On the computational inefficiency of large batch sizes for stochastic gradient descent. arXiv preprint arXiv:1811.12941 (2018)."},{"key":"e_1_2_1_15_1","unstructured":"Google. 2022. Dev Board datasheet. https:\/\/coral.ai\/docs\/dev-board\/datasheet\/.  Google. 2022. Dev Board datasheet. https:\/\/coral.ai\/docs\/dev-board\/datasheet\/."},{"key":"e_1_2_1_16_1","unstructured":"Google. 2022. Google Coral Products. https:\/\/coral.ai\/products\/.  Google. 2022. Google Coral Products. https:\/\/coral.ai\/products\/."},{"key":"e_1_2_1_17_1","volume-title":"Euro-Par 2017: Parallel Processing","author":"Halawa Hassan","unstructured":"Hassan Halawa , Hazem A. Abdelhafez , Andrew Boktor , and Matei Ripeanu . 2017. NVIDIA Jetson Platform Characterization . In Euro-Par 2017: Parallel Processing . Springer International Publishing , Cham , 92--105. Hassan Halawa, Hazem A. Abdelhafez, Andrew Boktor, and Matei Ripeanu. 2017. NVIDIA Jetson Platform Characterization. In Euro-Par 2017: Parallel Processing. Springer International Publishing, Cham, 92--105."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/IGSC51522.2020.9290876"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00140"},{"key":"e_1_2_1_21_1","unstructured":"Intel. 2022. Intel Movidius VPUs. https:\/\/www.intel.com\/content\/www\/us\/en\/products\/details\/processors\/movidiusvpu.html.  Intel. 2022. Intel Movidius VPUs. https:\/\/www.intel.com\/content\/www\/us\/en\/products\/details\/processors\/movidiusvpu.html."},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3446382.3448606"},{"key":"e_1_2_1_23_1","unstructured":"Dimitrios Kollias et al. 2018. Dimitrios Kollias and Athanasios Tagaris and Andreas Stafylopatis and Stefanos Kollias and Georgios Tagaris. Complex & Intelligent Systems (2018).  Dimitrios Kollias et al. 2018. Dimitrios Kollias and Athanasios Tagaris and Andreas Stafylopatis and Stefanos Kollias and Georgios Tagaris. Complex & Intelligent Systems (2018)."},{"key":"e_1_2_1_24_1","unstructured":"Alex Krizhevsky and Geoffrey Hinton. 2009. Learning multiple layers of features from tiny images. (2009).  Alex Krizhevsky and Geoffrey Hinton. 2009. Learning multiple layers of features from tiny images. (2009)."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPSW.2019.00148"},{"key":"e_1_2_1_26_1","volume-title":"Quiver: An informed storage cache for deep learning. In 18th {USENIX} Conference on File and Storage Technologies ({FAST} 20).","author":"Kumar Abhishek Vijaya","year":"2020","unstructured":"Abhishek Vijaya Kumar and Muthian Sivathanu . 2020 . Quiver: An informed storage cache for deep learning. In 18th {USENIX} Conference on File and Storage Technologies ({FAST} 20). Abhishek Vijaya Kumar and Muthian Sivathanu. 2020. Quiver: An informed storage cache for deep learning. In 18th {USENIX} Conference on File and Storage Technologies ({FAST} 20)."},{"key":"e_1_2_1_27_1","volume-title":"A survey of deep learning applications to autonomous vehicle control","author":"Kuutti Sampo","year":"2020","unstructured":"Sampo Kuutti , Richard Bowden , Yaochu Jin , Phil Barber , and Saber Fallah . 2020. A survey of deep learning applications to autonomous vehicle control . IEEE Transactions on Intelligent Transportation Systems ( 2020 ). Sampo Kuutti, Richard Bowden, Yaochu Jin, Phil Barber, and Saber Fallah. 2020. A survey of deep learning applications to autonomous vehicle control. IEEE Transactions on Intelligent Transportation Systems (2020)."},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/5.726791"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICPADS47876.2019.00077"},{"key":"e_1_2_1_30_1","unstructured":"man page. 2021. iostat. https:\/\/man7.org\/linux\/man-pages\/man1\/iostat.1.html.  man page. 2021. iostat. https:\/\/man7.org\/linux\/man-pages\/man1\/iostat.1.html."},{"key":"e_1_2_1_31_1","unstructured":"man pages. 2021. vmtouch. https:\/\/linux.die.net\/man\/8\/vmtouch.  man pages. 2021. vmtouch. https:\/\/linux.die.net\/man\/8\/vmtouch."},{"key":"e_1_2_1_32_1","volume-title":"Revisiting small batch training for deep neural networks. arXiv preprint arXiv:1804.07612","author":"Masters Dominic","year":"2018","unstructured":"Dominic Masters and Carlo Luschi . 2018. Revisiting small batch training for deep neural networks. arXiv preprint arXiv:1804.07612 ( 2018 ). Dominic Masters and Carlo Luschi. 2018. Revisiting small batch training for deep neural networks. arXiv preprint arXiv:1804.07612 (2018)."},{"key":"e_1_2_1_33_1","volume-title":"Taylor Robie, Tom St John, Carole-Jean Wu, Lingjie Xu, Cliff Young, and Matei Zaharia.","author":"Mattson Peter","year":"2020","unstructured":"Peter Mattson , Christine Cheng , Gregory Diamos , Cody Coleman , Paulius Micikevicius , David Patterson , Hanlin Tang , Gu-Yeon Wei , Peter Bailis , Victor Bittorf , David Brooks , Dehao Chen , Debo Dutta , Udit Gupta , Kim Hazelwood , Andy Hock , Xinyuan Huang , Daniel Kang , David Kanter , Naveen Kumar , Jeffery Liao , Deepak Narayanan , Tayo Oguntebi , Gennady Pekhimenko , Lillian Pentecost , Vijay Janapa Reddi , Taylor Robie, Tom St John, Carole-Jean Wu, Lingjie Xu, Cliff Young, and Matei Zaharia. 2020 . MLPerf Training Benchmark , Vol . 2. 336--349. https:\/\/proceedings.mlsys.org\/ paper\/2020\/file\/02522a2b2726fb0a03bb19f2d8d9524d-Paper.pdf Peter Mattson, Christine Cheng, Gregory Diamos, Cody Coleman, Paulius Micikevicius, David Patterson, Hanlin Tang, Gu-Yeon Wei, Peter Bailis, Victor Bittorf, David Brooks, Dehao Chen, Debo Dutta, Udit Gupta, Kim Hazelwood, Andy Hock, Xinyuan Huang, Daniel Kang, David Kanter, Naveen Kumar, Jeffery Liao, Deepak Narayanan, Tayo Oguntebi, Gennady Pekhimenko, Lillian Pentecost, Vijay Janapa Reddi, Taylor Robie, Tom St John, Carole-Jean Wu, Lingjie Xu, Cliff Young, and Matei Zaharia. 2020. MLPerf Training Benchmark, Vol. 2. 336--349. https:\/\/proceedings.mlsys.org\/ paper\/2020\/file\/02522a2b2726fb0a03bb19f2d8d9524d-Paper.pdf"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.14778\/3446095.3446100"},{"key":"e_1_2_1_35_1","unstructured":"Nvidia. 2021. Jetson AGX Xavier Developer Kit. https:\/\/developer.nvidia.com\/embedded\/jetson-agx-xavier-developerkit.  Nvidia. 2021. Jetson AGX Xavier Developer Kit. https:\/\/developer.nvidia.com\/embedded\/jetson-agx-xavier-developerkit."},{"key":"e_1_2_1_36_1","unstructured":"Nvidia. 2021. Jetson Nano Developer Kit. https:\/\/developer.nvidia.com\/embedded\/jetson-nano-developer-kit.  Nvidia. 2021. Jetson Nano Developer Kit. https:\/\/developer.nvidia.com\/embedded\/jetson-nano-developer-kit."},{"key":"e_1_2_1_37_1","unstructured":"Nvidia. 2021. Jetson NX Xavier Developer Kit. https:\/\/developer.nvidia.com\/embedded\/jetson-xavier-nx.  Nvidia. 2021. Jetson NX Xavier Developer Kit. https:\/\/developer.nvidia.com\/embedded\/jetson-xavier-nx."},{"key":"e_1_2_1_38_1","unstructured":"Nvidia. 2021. Power modes for Nano. https:\/\/docs.nvidia.com\/jetson\/archives\/l4t-archived\/l4t-3261\/index.html#page\/ Tegra%20Linux%20Driver%20Package%20Development%20Guide\/power_management_nano.html#.  Nvidia. 2021. Power modes for Nano. https:\/\/docs.nvidia.com\/jetson\/archives\/l4t-archived\/l4t-3261\/index.html#page\/ Tegra%20Linux%20Driver%20Package%20Development%20Guide\/power_management_nano.html#."},{"key":"e_1_2_1_39_1","unstructured":"Nvidia. 2021. Power modes for NX and AGX. https:\/\/docs.nvidia.com\/jetson\/archives\/l4t-archived\/l4t-3261\/index.html# page\/Tegra%20Linux%20Driver%20Package%20Development%20Guide\/power_management_jetson_xavier.html#.  Nvidia. 2021. Power modes for NX and AGX. https:\/\/docs.nvidia.com\/jetson\/archives\/l4t-archived\/l4t-3261\/index.html# page\/Tegra%20Linux%20Driver%20Package%20Development%20Guide\/power_management_jetson_xavier.html#."},{"key":"e_1_2_1_40_1","volume-title":"Technical Brief: Nvidia Jetson AGX Orin. https:\/\/www.nvidia.com\/content\/dam\/en-zz\/Solutions\/gtcf21\/ jetson-orin\/nvidia-jetson-agx-orin-technical-brief.pdf.","year":"2021","unstructured":"Nvidia. 2021 . Technical Brief: Nvidia Jetson AGX Orin. https:\/\/www.nvidia.com\/content\/dam\/en-zz\/Solutions\/gtcf21\/ jetson-orin\/nvidia-jetson-agx-orin-technical-brief.pdf. Nvidia. 2021. Technical Brief: Nvidia Jetson AGX Orin. https:\/\/www.nvidia.com\/content\/dam\/en-zz\/Solutions\/gtcf21\/ jetson-orin\/nvidia-jetson-agx-orin-technical-brief.pdf."},{"key":"e_1_2_1_41_1","unstructured":"Nvidia. 2021. tegrastats. https:\/\/docs.nvidia.com\/jetson\/archives\/l4t-archived\/l4t-3231\/index.html#page\/Tegra% 20Linux%20Driver%20Package%20Development%20Guide\/AppendixTegraStats.html.  Nvidia. 2021. tegrastats. https:\/\/docs.nvidia.com\/jetson\/archives\/l4t-archived\/l4t-3231\/index.html#page\/Tegra% 20Linux%20Driver%20Package%20Development%20Guide\/AppendixTegraStats.html."},{"key":"e_1_2_1_42_1","unstructured":"Nvidia. 2022. Jetson AGX Orin Developer Kit. https:\/\/www.nvidia.com\/en-us\/autonomous-machines\/embeddedsystems\/jetson-orin\/.  Nvidia. 2022. Jetson AGX Orin Developer Kit. https:\/\/www.nvidia.com\/en-us\/autonomous-machines\/embeddedsystems\/jetson-orin\/."},{"key":"e_1_2_1_43_1","unstructured":"papers with code. 2021. Mobilenet V3. https:\/\/paperswithcode.com\/lib\/torchvision\/mobilenet-v3.  papers with code. 2021. Mobilenet V3. https:\/\/paperswithcode.com\/lib\/torchvision\/mobilenet-v3."},{"key":"e_1_2_1_44_1","volume-title":"PyTorch: An Imperative Style","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke , Sam Gross , Francisco Massa , Adam Lerer , James Bradbury , Gregory Chanan , Trevor Killeen , Zeming Lin , Natalia Gimelshein , Luca Antiga , Alban Desmaison , Andreas K\u00f6pf , Edward Yang , Zach DeVito , Martin Raison , Alykhan Tejani , Sasank Chilamkurthy , Benoit Steiner , Lu Fang , Junjie Bai , and Soumith Chintala . 2019. PyTorch: An Imperative Style , High-Performance Deep Learning Library . In Advances in Neural Information Processing Systems, Vol. 32 . https:\/\/proceedings.neurips.cc\/paper\/ 2019 \/file\/bdbca288fee7f92f2bfa9f7012727740-Paper.pdf Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas K\u00f6pf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Advances in Neural Information Processing Systems, Vol. 32. https:\/\/proceedings.neurips.cc\/paper\/2019\/file\/bdbca288fee7f92f2bfa9f7012727740-Paper.pdf"},{"key":"e_1_2_1_45_1","volume-title":"International Workshop on Demand Response, co-located with the ACM e-Energy.","author":"Prabhakar T","year":"2014","unstructured":"T Prabhakar , Nisha Bhaskar , Tejas Pande , and Chaitanya Kulkarni . 2014 . Joule Jotter: An interactive energy meter for metering, monitoring and control . In International Workshop on Demand Response, co-located with the ACM e-Energy. T Prabhakar, Nisha Bhaskar, Tejas Pande, and Chaitanya Kulkarni. 2014. Joule Jotter: An interactive energy meter for metering, monitoring and control. In International Workshop on Demand Response, co-located with the ACM e-Energy."},{"key":"e_1_2_1_46_1","unstructured":"pytorch. 2021. TORCH.UTILS.DATA. https:\/\/pytorch.org\/docs\/stable\/data.html.  pytorch. 2021. TORCH.UTILS.DATA. https:\/\/pytorch.org\/docs\/stable\/data.html."},{"key":"e_1_2_1_47_1","unstructured":"PyTorch. 2022. Cuda event. https:\/\/pytorch.org\/docs\/stable\/generated\/torch.cuda.Event.html.  PyTorch. 2022. Cuda event. https:\/\/pytorch.org\/docs\/stable\/generated\/torch.cuda.Event.html."},{"key":"e_1_2_1_48_1","volume-title":"Workshop on Parallel AI and Systems for the Edge - PAISE. In 2022 IEEE International Symposium on Parallel and Distributed Processing, Workshops and Phd Forum (IPDPSW).","author":"Aakash Khochare Prashanthi S. K","year":"2022","unstructured":"Prashanthi S. K , Aakash Khochare , Sai Anuroop Kesanapalli , Rahul Bhope , and Yogesh Simmhan . 2022 . Workshop on Parallel AI and Systems for the Edge - PAISE. In 2022 IEEE International Symposium on Parallel and Distributed Processing, Workshops and Phd Forum (IPDPSW). Prashanthi S. K, Aakash Khochare, Sai Anuroop Kesanapalli, Rahul Bhope, and Yogesh Simmhan. 2022. Workshop on Parallel AI and Systems for the Edge - PAISE. In 2022 IEEE International Symposium on Parallel and Distributed Processing, Workshops and Phd Forum (IPDPSW)."},{"key":"e_1_2_1_49_1","volume-title":"Measuring the effects of data parallelism on neural network training. arXiv preprint arXiv:1811.03600","author":"Shallue Christopher J","year":"2018","unstructured":"Christopher J Shallue , Jaehoon Lee , Joseph Antognini , Jascha Sohl-Dickstein , Roy Frostig , and George E Dahl . 2018. Measuring the effects of data parallelism on neural network training. arXiv preprint arXiv:1811.03600 ( 2018 ). Christopher J Shallue, Jaehoon Lee, Joseph Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E Dahl. 2018. Measuring the effects of data parallelism on neural network training. arXiv preprint arXiv:1811.03600 (2018)."},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.1409.1556"},{"key":"e_1_2_1_51_1","unstructured":"Vladislav Sovrasov. 2021. Flops counter. https:\/\/pypi.org\/project\/ptflops\/.  Vladislav Sovrasov. 2021. Flops counter. https:\/\/pypi.org\/project\/ptflops\/."},{"key":"e_1_2_1_52_1","unstructured":"TensorFlow. 2022. TFF GLDv2. https:\/\/www.tensorflow.org\/federated\/api_docs\/python\/tff\/simulation\/datasets\/gldv2\/ load_data  TensorFlow. 2022. TFF GLDv2. https:\/\/www.tensorflow.org\/federated\/api_docs\/python\/tff\/simulation\/datasets\/gldv2\/ load_data"},{"key":"e_1_2_1_53_1","first-page":"24","volume-title":"2001 USENIX Annual Technical Conference (USENIX ATC 01)","author":"van Riel Rik","year":"2001","unstructured":"Rik van Riel . 2001 . Page Replacement in Linux 2.4 Memory Management . In 2001 USENIX Annual Technical Conference (USENIX ATC 01) . USENIX Association, Boston, MA. https:\/\/www.usenix.org\/conference\/ 2001-usenix-annualtechnical-conference\/page-replacement-linux- 24 -memory-management Rik van Riel. 2001. Page Replacement in Linux 2.4 Memory Management. In 2001 USENIX Annual Technical Conference (USENIX ATC 01). USENIX Association, Boston, MA. https:\/\/www.usenix.org\/conference\/2001-usenix-annualtechnical-conference\/page-replacement-linux-24-memory-management"},{"key":"e_1_2_1_54_1","volume-title":"Benchmarking tpu, gpu, and cpu platforms for deep learning. arXiv preprint arXiv:1907.10701","author":"Wang Yu","year":"2019","unstructured":"Yu Wang , Gu-Yeon Wei , and David Brooks . 2019. Benchmarking tpu, gpu, and cpu platforms for deep learning. arXiv preprint arXiv:1907.10701 ( 2019 ). Yu Wang, Gu-Yeon Wei, and David Brooks. 2019. Benchmarking tpu, gpu, and cpu platforms for deep learning. arXiv preprint arXiv:1907.10701 (2019)."},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00265"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2019.2918951"}],"container-title":["Proceedings of the ACM on Measurement and Analysis of Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3570604","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3570604","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:46:17Z","timestamp":1750178777000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3570604"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,12]]},"references-count":56,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2022,12]]}},"alternative-id":["10.1145\/3570604"],"URL":"https:\/\/doi.org\/10.1145\/3570604","relation":{},"ISSN":["2476-1249"],"issn-type":[{"value":"2476-1249","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,12]]},"assertion":[{"value":"2022-12-08","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}