{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:12:25Z","timestamp":1750219945045,"version":"3.41.0"},"reference-count":47,"publisher":"Association for Computing Machinery (ACM)","issue":"5s","license":[{"start":{"date-parts":[[2023,9,9]],"date-time":"2023-09-09T00:00:00Z","timestamp":1694217600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Embed. Comput. Syst."],"published-print":{"date-parts":[[2023,10,31]]},"abstract":"<jats:p>In this work, we present a computationally efficient Reinforcement Learning mapping search heuristic for finding high quality mappings for N-dimensional convolution loops that uses a computationally inexpensive reward function based on potential data reuse of operands to guide the search process. We also present a RL state representation generalizable to N-dimensional convolution loops, and a state representation parsing strategy ensuring that only valid mappings are evaluated for quality. Our RL search heuristic is applicable to multi-core systems with a memory hierarchy. We show that our RL based search heuristic for a range of 3D convolution layers, at significantly lower computational expense than random search, generally yields mappings with lower Energy-Delay Product (EDP) for an architecture with multiple processing elements with shared memory connected to DRAM. Our evaluation results demonstrated across 19 3D convolution layers, shows that our RL method performed only an average 11.24% of the operations of that of Timeloop\u2019s random search for assessing same number of valid mappings. The mappings found using Timeloop had an average 12.51% higher EDP compared to lowest EDP mapping found using our RL method. Further, the lowest EDP mappings found using our method had an average only 4.69\u00d7 higher EDP than the theoretical lower bound EDP, with the best case being only 1.29\u00d7 higher.<\/jats:p>","DOI":"10.1145\/3609110","type":"journal-article","created":{"date-parts":[[2023,9,9]],"date-time":"2023-09-09T13:33:18Z","timestamp":1694266398000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Computationally Efficient DNN Mapping Search Heuristic using Deep Reinforcement Learning"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0009-0006-3569-9589","authenticated-orcid":false,"given":"Suyash","family":"Bakshi","sequence":"first","affiliation":[{"name":"University of Houston, Main Campus, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0337-879X","authenticated-orcid":false,"given":"Lennart","family":"Johnsson","sequence":"additional","affiliation":[{"name":"University of Houston, Main Campus, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,9,9]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"2023. Cadence. https:\/\/www.cadence.com\/en_US\/home\/tools\/ip\/tensilica-ip.html"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.1603.04467"},{"key":"e_1_3_1_4_2","article-title":"Chameleon: Adaptive code optimization for expedited deep neural network compilation","volume":"2001","author":"Ahn Byung Hoon","year":"2020","unstructured":"Byung Hoon Ahn, Prannoy Pilligundla, Amir Yazdanbakhsh, and Hadi Esmaeilzadeh. 2020. Chameleon: Adaptive code optimization for expedited deep neural network compilation. CoRR abs\/2001.08743 (2020). arXiv:2001.08743https:\/\/arxiv.org\/abs\/2001.08743","journal-title":"CoRR"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/SBAC-PAD.2012.26"},{"key":"e_1_3_1_6_2","unstructured":"Anon.2021. F3 Netherlands\u201d Marine Seismic Dataset. https:\/\/terranubis.com\/datainfo\/F3-Demo-2020"},{"key":"e_1_3_1_7_2","unstructured":"Anon.2023. Intel oneAPI Math Kernel Library. https:\/\/www.intel.com\/content\/www\/us\/en\/develop\/documentation\/onemkl-developer-reference-c\/top.html"},{"key":"e_1_3_1_8_2","unstructured":"Anon.2023. Minimal Standard Minstd_rand0 Generator. https:\/\/cplusplus.com\/reference\/random\/minstd_rand0\/"},{"key":"e_1_3_1_9_2","unstructured":"Anon.2023. NVIDIA CUDA Basic Linear Algebra Subroutine Library. https:\/\/docs.nvidia.com\/cuda\/cublas\/"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/SBAC-PAD49847.2020.00051"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/3085572"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2015.10"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939785"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.1802.04799"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2016.40"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0167-8191(00)00087-9"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/3361682"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3358198"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-020-2649-2"},{"key":"e_1_3_1_20_2","article-title":"Morph: Flexible acceleration for 3D CNN-based video understanding","volume":"1810","author":"Hegde Kartik","year":"2018","unstructured":"Kartik Hegde, Rohit Agrawal, Yulun Yao, and Christopher W. Fletcher. 2018. Morph: Flexible acceleration for 3D CNN-based video understanding. CoRR abs\/1810.06807 (2018). arXiv:1810.06807http:\/\/arxiv.org\/abs\/1810.06807","journal-title":"CoRR"},{"key":"e_1_3_1_21_2","article-title":"Mind mappings: Enabling efficient algorithm-accelerator mapping space search","volume":"2103","author":"Hegde Kartik","year":"2021","unstructured":"Kartik Hegde, Po-An Tsai, Sitao Huang, Vikas Chandra, Angshuman Parashar, et\u00a0al. 2021. Mind mappings: Enabling efficient algorithm-accelerator mapping space search. CoRR abs\/2103.01489 (2021). arXiv:2103.01489https:\/\/arxiv.org\/abs\/2103.01489","journal-title":"CoRR"},{"key":"e_1_3_1_22_2","unstructured":"Marius Hobbhahn. 2021. How to Measure FLOP\/s for Neural Networks Empirically?"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2105.01898"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/800076.802486"},{"key":"e_1_3_1_25_2","article-title":"In-datacenter performance analysis of a tensor processing unit","volume":"1704","author":"Jouppi Norman P.","year":"2017","unstructured":"Norman P. Jouppi, Cliff Young, Nishant Patil, David A. Patterson, Gaurav Agrawal, et\u00a0al. 2017. In-datacenter performance analysis of a tensor processing unit. CoRR abs\/1704.04760 (2017). arXiv:1704.04760http:\/\/arxiv.org\/abs\/1704.04760","journal-title":"CoRR"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/3400302.3415639"},{"key":"e_1_3_1_27_2","doi-asserted-by":"crossref","unstructured":"Induprakas Kodukula Nawaaz Ahmed and d Keshav Pingali. 1997. Data-centric multi-level blocking. 346\u2013357.","DOI":"10.1145\/258916.258946"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/3065386"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.parco.2008.12.010"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.1805.02566"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2012.03837"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2021.3059962"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1023\/A:1011119519789"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.5555\/2670022"},{"key":"e_1_3_1_35_2","volume-title":"Efficient LU Factorization for Texas Instruments Keystone Architecture Digital Signal Processors","author":"Netzer Gilbert","year":"2015","unstructured":"Gilbert Netzer. 2015. Efficient LU Factorization for Texas Instruments Keystone Architecture Digital Signal Processors. Master\u2019s Thesis. Royal Institute of Technology (KTH). http:\/\/www.diva-portal.org\/smash\/get\/diva2:837145\/FULLTEXT01"},{"key":"e_1_3_1_36_2","unstructured":"Badreddine Noune Philip Jones Daniel Justus Dominic Masters and Carlo Luschi. 2022. 8-bit Numerical Formats for Deep Neural Networks. arxiv:cs.LG\/2206.02915"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2019.00042"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.23919\/DATE.2019.8715007"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.1707.06347"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.1409.1556"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.1412.0767"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICPPW.2010.38"},{"key":"e_1_3_1_43_2","unstructured":"Naigang Wang Jungwook Choi Daniel Brand Chia-Yu Chen and Kailash Gopalakrishnan. 2018. Training deep neural networks with 8-bit floating point numbers. arxiv:cs.LG\/1812.08011"},{"key":"e_1_3_1_44_2","first-page":"1","article-title":"Accelergy: An architecture-level energy estimation methodology for accelerator designs","author":"Wu Yannan Nellie","year":"2019","unstructured":"Yannan Nellie Wu, Joel S. Emer, and Vivienne Sze. 2019. Accelergy: An architecture-level energy estimation methodology for accelerator designs. 2019 IEEE\/ACM International Conference on Computer-Aided Design (ICCAD) (2019), 1\u20138.","journal-title":"2019 IEEE\/ACM International Conference on Computer-Aided Design (ICCAD)"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378514"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.1903.06498"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378508"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.1606.06650"}],"container-title":["ACM Transactions on Embedded Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3609110","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3609110","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:48:58Z","timestamp":1750182538000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3609110"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,9,9]]},"references-count":47,"journal-issue":{"issue":"5s","published-print":{"date-parts":[[2023,10,31]]}},"alternative-id":["10.1145\/3609110"],"URL":"https:\/\/doi.org\/10.1145\/3609110","relation":{},"ISSN":["1539-9087","1558-3465"],"issn-type":[{"type":"print","value":"1539-9087"},{"type":"electronic","value":"1558-3465"}],"subject":[],"published":{"date-parts":[[2023,9,9]]},"assertion":[{"value":"2023-03-16","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-07-13","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-09-09","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}