{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,18]],"date-time":"2026-08-18T01:44:33Z","timestamp":1787017473176,"version":"build-2736575974"},"reference-count":40,"publisher":"Association for Computing Machinery (ACM)","issue":"2","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62202123"],"award-info":[{"award-number":["62202123"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"Joint Funds of the National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["U22A2036"],"award-info":[{"award-number":["U22A2036"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2025,6,30]]},"abstract":"<jats:p>\n                    To improve the performance and energy efficiency of deep learning (DL) applications, recent edge computing platforms have built-in heterogeneous accelerators, such as general-purpose graphics processing units (GPUs) and neural processing units (NPUs). For example, widely used NVIDIA Jetson platforms contain CPU, GPU, and deep learning accelerator (DLA), a type of NPU. It is non-trivial to map DL workloads to suitable accelerators to improve performance, energy efficiency, or even both. This article presents JDIMO,\n                    <jats:xref ref-type=\"fn\">\n                      <jats:sup>1<\/jats:sup>\n                    <\/jats:xref>\n                    a Jetson-aware deep-learning inference workload mapping optimization framework, to simultaneously improve energy efficiency and performance. JDIMO first measures energy-performance data of the fundamental nodes and the sub-networks with energy-efficiency improvement potential according to the topology structure of a DL network. Then, under the guidance of an analytical energy-performance model, the framework exploits an algorithm based on the variable-length sliding window to find the optimal mapping configuration and the optimal number of CUDA streams. We evaluate JDIMO by applying it to seven DL applications on a Jetson Orin NX (16GB) platform. JDIMO saves 47.5% EDP (energy delay product) and 22.6% energy and improves 138.3% QPS (queries per second) on average compared to the DLA-possible configuration. JDIMO saves 22.5% EDP and 12.6% energy and improves 13.5% QPS on average compared to JEDI, the most similar work to ours. Meanwhile, JDIMO also reduces 93.8% optimization time on average compared to JEDI.\n                  <\/jats:p>","DOI":"10.1145\/3736175","type":"journal-article","created":{"date-parts":[[2025,5,20]],"date-time":"2025-05-20T07:20:27Z","timestamp":1747725627000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["Deep Learning Workload Mapping Optimization on Jetson Platforms"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2478-2153","authenticated-orcid":false,"given":"Farui","family":"Wang","sequence":"first","affiliation":[{"name":"Harbin Institute of Technology Shenzhen","place":["Shenzhen, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0043-4370","authenticated-orcid":false,"given":"Meng","family":"Hao","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology","place":["Harbin, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-0071-6218","authenticated-orcid":false,"given":"Siyu","family":"Yang","sequence":"additional","affiliation":[{"name":"Computer Science and Technology, Harbin Institute of Technology","place":["Harbin, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4783-876X","authenticated-orcid":false,"given":"Weizhe","family":"Zhang","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology Shenzhen","place":["Shenzhen, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,7,2]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.3390\/electronics8030292"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/DSD60849.2023.00015"},{"key":"e_1_3_2_4_2","doi-asserted-by":"crossref","unstructured":"Kamil Roszyk Micha\u0142 R. Nowicki and Piotr Skrzypczy\u0144ski. 2022. Adopting the YOLOv4 architecture for low-latency multispectral pedestrian detection in autonomous driving. Sensors 22 3 (2022) 1082.","DOI":"10.3390\/s22031082"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/DAC56929.2023.10247722"},{"key":"e_1_3_2_6_2","first-page":"578","volume-title":"Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201918)","author":"Chen Tianqi","year":"2018","unstructured":"Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, et\u00a0al. 2018. \\(\\lbrace\\) TVM \\(\\rbrace\\) : An automated \\(\\lbrace\\) End-to-End \\(\\rbrace\\) optimizing compiler for deep learning. In Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201918). 578\u2013594."},{"key":"e_1_3_2_7_2","article-title":"Learning to optimize tensor programs","author":"Chen Tianqi","year":"2018","unstructured":"Tianqi Chen, Lianmin Zheng, Eddie Yan, Ziheng Jiang, Thierry Moreau, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018. Learning to optimize tensor programs. In Proceedings of the 32nd International Conference on Neural Information Processing Systems.","journal-title":"Proceedings of the 32nd International Conference on Neural Information Processing Systems"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/RSDHA54838.2021.00006"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/3489517.3530572"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP49357.2023.10095003"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/3613424.3614285"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3555802"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/3508391"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/LES.2021.3087707"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2020.2977496"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3400302.3415639"},{"key":"e_1_3_2_17_2","doi-asserted-by":"crossref","unstructured":"Jangryul Kim and Soonhoi Ha. 2022. Energy-aware scenario-based mapping of deep learning applications onto heterogeneous processors under real-time constraints. IEEE Trans. Comput. 72 6 (2022) 1666\u20131680.","DOI":"10.1109\/TC.2022.3218991"},{"key":"e_1_3_2_18_2","doi-asserted-by":"crossref","unstructured":"Jangryul Kim Jaewoo Son and Soonhoi Ha. 2023. A novel technique to support deep learning applications in a model-based embedded software design methodology. IEEE Access 11 (2023) 54869\u201354880.","DOI":"10.1109\/ACCESS.2023.3281913"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature14539"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.324"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3410463.3414671"},{"key":"e_1_3_2_22_2","unstructured":"NVIDIA Corporation. 2024. CUDA Toolkit Documentation. Retrieved March 27 2024 from https:\/\/docs.nvidia.com\/cuda\/index.html"},{"key":"e_1_3_2_23_2","unstructured":"NVIDIA Corporation. 2024. Deep Learning Accelerator (DLA). Retrieved April 6 2024 from https:\/\/developer.nvidia.com\/deep-learning-accelerator"},{"key":"e_1_3_2_24_2","unstructured":"NVIDIA Corporation. 2024. Jetson - Embedded AI Computing Platform. Retrieved April 6 2024 from https:\/\/developer.nvidia.com\/embedded-computing"},{"key":"e_1_3_2_25_2","unstructured":"NVIDIA Corporation. 2024. Jetson Orin Technical Specifications. Retrieved April 18 2024 from https:\/\/www.nvidia.com\/en-us\/autonomous-machines\/embedded-systems\/jetson-orin\/"},{"key":"e_1_3_2_26_2","unstructured":"NVIDIA Corporation. 2024. NVIDIA Deep Learning TensorRT Documentation. Retrieved March 27 2024 from https:\/\/docs.nvidia.com\/deeplearning\/tensorrt\/index.html"},{"key":"e_1_3_2_27_2","unstructured":"NVIDIA Corporation. 2025. High-Performance In-Vehicle Computing for Autonomous Vehicles. Retrieved March 23 2025 from https:\/\/www.nvidia.com\/en-us\/self-driving-cars\/in-vehicle-computing\/"},{"key":"e_1_3_2_28_2","unstructured":"NVIDIA Corporation. 2025. NVIDIA Jetson Linux Developer Guide. Retrieved March 10 2025 from https:\/\/docs.nvidia.com\/jetson\/archives\/r36.4.3\/{\\}DeveloperGuide\/SD\/PlatformPowerAndPerformance\/JetsonOrinNanoSeriesJetsonOrinNxSeriesAndJetsonAgxOrinSeries.html"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/3609386"},{"key":"e_1_3_2_30_2","unstructured":"ONNX community. 2024. ONNX Model Zoo. Retrieved March 27 2024 from https:\/\/github.com\/onnx\/models"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00474"},{"key":"e_1_3_2_32_2","doi-asserted-by":"crossref","unstructured":"Zhongqin Bi Ling Yu Honghao Gao Ping Zhou and Hongyang Yao. 2021. Improved VGG model-based efficient traffic sign recognition for safe driving in 5G scenarios. International Journal of Machine Learning and Cybernetics 12 (2021) 3069\u20133080.","DOI":"10.1007\/s13042-020-01185-5"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA47549.2020.00036"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"e_1_3_2_35_2","first-page":"6105","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Tan Mingxing","year":"2019","unstructured":"Mingxing Tan and Quoc Le. 2019. Efficientnet: Rethinking model scaling for convolutional neural networks. In Proceedings of the International Conference on Machine Learning. PMLR, 6105\u20136114."},{"key":"e_1_3_2_36_2","unstructured":"Volvo Group. 2025. From Car to Cloud: Volvo Cars Expands Collaboration with NVIDIA. Retrieved March 23 2025 from https:\/\/www.media.volvocars.com\/global\/en-gb\/media\/pressreleases\/331848\/from-car-to-cloud-volvo-cars-expands-collaboration-with-nvidia"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/COMST.2020.2970550"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2023.3242200"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1145\/3470496.3527440"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378508"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3318216.3363312"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3736175","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,2]],"date-time":"2025-07-02T08:20:57Z","timestamp":1751444457000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3736175"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,30]]},"references-count":40,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2025,6,30]]}},"alternative-id":["10.1145\/3736175"],"URL":"https:\/\/doi.org\/10.1145\/3736175","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,6,30]]},"assertion":[{"value":"2024-11-02","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-05-07","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-07-02","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}