{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,11]],"date-time":"2026-03-11T17:10:35Z","timestamp":1773249035732,"version":"3.50.1"},"reference-count":45,"publisher":"Association for Computing Machinery (ACM)","issue":"5s","license":[{"start":{"date-parts":[[2019,10,8]],"date-time":"2019-10-08T00:00:00Z","timestamp":1570492800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["CCF 1723476"],"award-info":[{"award-number":["CCF 1723476"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100003621","name":"Ministry of Science, ICT and Future Planning","doi-asserted-by":"crossref","award":["2014-3-00035"],"award-info":[{"award-number":["2014-3-00035"]}],"id":[{"id":"10.13039\/501100003621","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100003725","name":"National Research Foundation of Korea","doi-asserted-by":"publisher","award":["2015M3C4A7065522"],"award-info":[{"award-number":["2015M3C4A7065522"]}],"id":[{"id":"10.13039\/501100003725","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Embed. Comput. Syst."],"published-print":{"date-parts":[[2019,10,31]]},"abstract":"<jats:p>\n            Dataflow accelerators feature simplicity, programmability, and energy-efficiency and are visualized as a promising architecture for accelerating perfectly nested loops that dominate several important applications, including image and media processing and deep learning. Although numerous accelerator designs are being proposed, how to discover the most efficient way to execute the perfectly nested loop of an application onto computational and memory resources of a given dataflow accelerator (\n            <jats:italic>execution method<\/jats:italic>\n            ) remains an essential and yet unsolved challenge. In this paper, we propose\n            <jats:italic>dMazeRunner<\/jats:italic>\n            -- to efficiently and accurately explore the vast space of the different ways to spatiotemporally execute a perfectly nested loop on dataflow accelerators (execution methods). The novelty of dMazeRunner framework is in: i) a holistic representation of the loop nests, that can succinctly capture the various execution methods, ii) accurate energy and performance models that explicitly capture the computation and communication patterns, data movement, and data buffering of the different execution methods, and iii) drastic pruning of the vast search space by discarding invalid solutions and the solutions that lead to the same cost. Our experiments on various convolution layers (perfectly nested loops) of popular deep learning applications demonstrate that the solutions discovered by dMazeRunner are on average 9.16\u00d7 better in Energy-Delay-Product (EDP) and 5.83\u00d7 better in execution time, as compared to prior approaches. With additional pruning heuristics, dMazeRunner reduces the search time from days to seconds with a mere 2.56% increase in EDP, as compared to the optimal solution.\n          <\/jats:p>","DOI":"10.1145\/3358198","type":"journal-article","created":{"date-parts":[[2019,10,10]],"date-time":"2019-10-10T13:13:05Z","timestamp":1570713185000},"page":"1-27","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":71,"title":["dMazeRunner"],"prefix":"10.1145","volume":"18","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4262-3938","authenticated-orcid":false,"given":"Shail","family":"Dave","sequence":"first","affiliation":[{"name":"Arizona State University, Tempe, AZ, US"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Youngbin","family":"Kim","sequence":"additional","affiliation":[{"name":"Yonsei University, Seoul, South Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sasikanth","family":"Avancha","sequence":"additional","affiliation":[{"name":"Parallel Computing Lab, Intel Labs, Bangalore, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kyoungwoo","family":"Lee","sequence":"additional","affiliation":[{"name":"Yonsei University, Seoul, South Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Aviral","family":"Shrivastava","sequence":"additional","affiliation":[{"name":"Arizona State University, Tempe, AZ, US"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2019,10,8]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.634"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080246"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/JSSC.2016.2616357"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/3297858.3304028"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2017.29"},{"key":"e_1_2_1_6_1","volume-title":"Programmatic control of a compiler for generating high-performance spatial hardware. arXiv preprint arXiv:1711.07606","author":"Rong Hongbo","year":"2017","unstructured":"Hongbo Rong . 2017. Programmatic control of a compiler for generating high-performance spatial hardware. arXiv preprint arXiv:1711.07606 ( 2017 ). Hongbo Rong. 2017. Programmatic control of a compiler for generating high-performance spatial hardware. arXiv preprint arXiv:1711.07606 (2017)."},{"key":"e_1_2_1_7_1","volume-title":"Scale-sim: Systolic cnn accelerator. arXiv preprint arXiv:1811.02883","author":"Samajdar Ananda","year":"2018","unstructured":"Ananda Samajdar , Yuhao Zhu , Paul Whatmough , Matthew Mattina , and Tushar Krishna . 2018 . Scale-sim: Systolic cnn accelerator. arXiv preprint arXiv:1811.02883 (2018). Ananda Samajdar, Yuhao Zhu, Paul Whatmough, Matthew Mattina, and Tushar Krishna. 2018. Scale-sim: Systolic cnn accelerator. arXiv preprint arXiv:1811.02883 (2018)."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/2754930"},{"key":"e_1_2_1_9_1","volume-title":"Karthikeyan Sankaralingam, Cristian Estan, and Behnam Robatmili.","author":"Nowatzki Tony","year":"2013","unstructured":"Tony Nowatzki , Michael Sartin-Tarm , Lorenzo De Carli , Karthikeyan Sankaralingam, Cristian Estan, and Behnam Robatmili. 2013 . A general constraint-centric scheduling framework for spatial architectures. In ACM SIGPLAN Notices, Vol. 48 . ACM , 495--506. Tony Nowatzki, Michael Sartin-Tarm, Lorenzo De Carli, Karthikeyan Sankaralingam, Cristian Estan, and Behnam Robatmili. 2013. A general constraint-centric scheduling framework for spatial architectures. In ACM SIGPLAN Notices, Vol. 48. ACM, 495--506."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2019.2913833"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/3007787.3001177"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/2749469.2750389"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/JSSC.2017.2778281"},{"key":"e_1_2_1_14_1","volume-title":"TVM: End-to-end optimization stack for deep learning. arXiv preprint arXiv:1802.04799","author":"Chen Tianqi","year":"2018","unstructured":"Tianqi Chen , Thierry Moreau , Ziheng Jiang , Haichen Shen , Eddie Q. Yan , Leyuan Wang , Yuwei Hu , Luis Ceze , Carlos Guestrin , and Arvind Krishnamurthy . 2018 . TVM: End-to-end optimization stack for deep learning. arXiv preprint arXiv:1802.04799 (2018), 1--15. Tianqi Chen, Thierry Moreau, Ziheng Jiang, Haichen Shen, Eddie Q. Yan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018. TVM: End-to-end optimization stack for deep learning. arXiv preprint arXiv:1802.04799 (2018), 1--15."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/2847263.2847276"},{"key":"e_1_2_1_16_1","volume-title":"https:\/\/github.com\/ARM-software\/SCALE-Sim. ([n. d.]). Accessed","year":"2018","unstructured":"SCALE-Sim. https:\/\/github.com\/ARM-software\/SCALE-Sim. ([n. d.]). Accessed : November 5, 2018 . SCALE-Sim. https:\/\/github.com\/ARM-software\/SCALE-Sim. ([n. d.]). Accessed: November 5, 2018."},{"key":"e_1_2_1_17_1","volume-title":"Software-defined design space exploration for an efficient AI accelerator architecture. arXiv preprint arXiv:1903.07676","author":"Yu Ye","year":"2019","unstructured":"Ye Yu , Yingmin Li , Shuai Che , Niraj K Jha , and Weifeng Zhang . 2019. Software-defined design space exploration for an efficient AI accelerator architecture. arXiv preprint arXiv:1903.07676 ( 2019 ). Ye Yu, Yingmin Li, Shuai Che, Niraj K Jha, and Weifeng Zhang. 2019. Software-defined design space exploration for an efficient AI accelerator architecture. arXiv preprint arXiv:1903.07676 (2019)."},{"key":"e_1_2_1_18_1","volume-title":"Jeff Ou Setter, Kaidi Cao, Heonjae Ha, Christos Kozyrakis, et al.","author":"Yang Xuan","year":"2018","unstructured":"Xuan Yang , Mingyu Gao , Jing Pu , Ankita Nayak , Qiaoyi Liu , Steven Emberton Bell , Jeff Ou Setter, Kaidi Cao, Heonjae Ha, Christos Kozyrakis, et al. 2018 . DNN dataflow choice is overrated. arXiv preprint arXiv:1809.04070 (2018). Xuan Yang, Mingyu Gao, Jing Pu, Ankita Nayak, Qiaoyi Liu, Steven Emberton Bell, Jeff Ou Setter, Kaidi Cao, Heonjae Ha, Christos Kozyrakis, et al. 2018. DNN dataflow choice is overrated. arXiv preprint arXiv:1809.04070 (2018)."},{"key":"e_1_2_1_19_1","volume-title":"Hinton","author":"Krizhevsky Alex","year":"2012","unstructured":"Alex Krizhevsky , Ilya Sutskever , and Geoffrey E . Hinton . 2012 . Imagenet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems . 1097--1105. Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. 2012. Imagenet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems. 1097--1105."},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_2_1_21_1","unstructured":"Jongsoo Park Maxim Naumov Protonu Basu Summer Deng Aravind Kalaiah Daya Khudia James Law Parth Malani Andrey Malevich Satish Nadathur etal 2018. Deep learning inference in facebook data centers: Characterization performance optimizations and hardware implications. arXiv preprint arXiv:1811.09886 (2018).  Jongsoo Park Maxim Naumov Protonu Basu Summer Deng Aravind Kalaiah Daya Khudia James Law Parth Malani Andrey Malevich Satish Nadathur et al. 2018. Deep learning inference in facebook data centers: Characterization performance optimizations and hardware implications. arXiv preprint arXiv:1811.09886 (2018)."},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISSCC.2019.8662396"},{"key":"e_1_2_1_23_1","volume-title":"2018 51st Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). IEEE, 933--946","author":"Hegde Kartik","unstructured":"Kartik Hegde , Rohit Agrawal , Yulun Yao , and Christopher W. Fletcher . 2018. Morph: Flexible acceleration for 3D CNN-based video understanding . In 2018 51st Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). IEEE, 933--946 . Kartik Hegde, Rohit Agrawal, Yulun Yao, and Christopher W. Fletcher. 2018. Morph: Flexible acceleration for 3D CNN-based video understanding. In 2018 51st Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). IEEE, 933--946."},{"key":"e_1_2_1_24_1","volume-title":"DNN Energy Model and Optimizer. https:\/\/github.com\/xuanyoya\/CNN-blocking\/tree\/dev. Accessed","author":"Xuan Yang","year":"2018","unstructured":"Xuan Yang et al. DNN Energy Model and Optimizer. https:\/\/github.com\/xuanyoya\/CNN-blocking\/tree\/dev. Accessed : November 5, 2018 . Xuan Yang et al. DNN Energy Model and Optimizer. https:\/\/github.com\/xuanyoya\/CNN-blocking\/tree\/dev. Accessed: November 5, 2018."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/VLSIC.2018.8502276"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3240765.3240838"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2019.00027"},{"key":"e_1_2_1_28_1","volume-title":"Aho et al","author":"Alfred","year":"2007","unstructured":"Alfred V. Aho et al . 2007 . Compilers : Principles , techniques and tools. (2007). Alfred V. Aho et al. 2007. Compilers: Principles, techniques and tools. (2007)."},{"key":"e_1_2_1_29_1","volume-title":"Compiler Optimizations for Improving Data Locality","author":"Carr Steve","unstructured":"Steve Carr , Kathryn S. McKinley , and Chau-Wen Tseng . 1994. Compiler Optimizations for Improving Data Locality . Vol. 29 . ACM. Steve Carr, Kathryn S. McKinley, and Chau-Wen Tseng. 1994. Compiler Optimizations for Improving Data Locality. Vol. 29. ACM."},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/1230800.1230807"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/1379022.1375595"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11265-008-0244-0"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3061639.3062262"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.5555\/3049832.3049854"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/DAC.2018.8465892"},{"key":"e_1_2_1_36_1","volume-title":"URECA: A compiler solution to manage unified register file for CGRAs. In 2018 Design, Automation 8 Test in Europe Conference 8 Exhibition (DATE)","author":"Dave Shail","year":"2018","unstructured":"Shail Dave , Mahesh Balasubramanian , and Aviral Shrivastava . 2018 . URECA: A compiler solution to manage unified register file for CGRAs. In 2018 Design, Automation 8 Test in Europe Conference 8 Exhibition (DATE) . IEEE , 1081--1086. Shail Dave, Mahesh Balasubramanian, and Aviral Shrivastava. 2018. URECA: A compiler solution to manage unified register file for CGRAs. In 2018 Design, Automation 8 Test in Europe Conference 8 Exhibition (DATE). IEEE, 1081--1086."},{"key":"e_1_2_1_37_1","volume-title":"Optimally scheduling CNN convolutions for efficient memory access. arXiv preprint arXiv:1902.01492","author":"Stoutchinin Arthur","year":"2019","unstructured":"Arthur Stoutchinin , Francesco Conti , and Luca Benini . 2019. Optimally scheduling CNN convolutions for efficient memory access. arXiv preprint arXiv:1902.01492 ( 2019 ). Arthur Stoutchinin, Francesco Conti, and Luca Benini. 2019. Optimally scheduling CNN convolutions for efficient memory access. arXiv preprint arXiv:1902.01492 (2019)."},{"key":"e_1_2_1_38_1","volume-title":"International Conference on Machine Learning. 1737--1746","author":"Gupta Suyog","year":"2015","unstructured":"Suyog Gupta , Ankur Agrawal , Kailash Gopalakrishnan , and Pritish Narayanan . 2015 . Deep learning with limited numerical precision . In International Conference on Machine Learning. 1737--1746 . Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish Narayanan. 2015. Deep learning with limited numerical precision. In International Conference on Machine Learning. 1737--1746."},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2019.00040"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/2684746.2689060"},{"key":"e_1_2_1_41_1","volume-title":"MAESTRO: An open-source infrastructure for modeling dataflows within deep learning accelerators. CoRR abs\/1805.02566","author":"Kwon Hyoukjun","year":"2018","unstructured":"Hyoukjun Kwon , Michael Pellauer , and Tushar Krishna . 2018 . MAESTRO: An open-source infrastructure for modeling dataflows within deep learning accelerators. CoRR abs\/1805.02566 (2018). Hyoukjun Kwon, Michael Pellauer, and Tushar Krishna. 2018. MAESTRO: An open-source infrastructure for modeling dataflows within deep learning accelerators. CoRR abs\/1805.02566 (2018)."},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/CGO.2004.1281665"},{"key":"e_1_2_1_43_1","volume-title":"https:\/\/www.mathworks.com\/help\/optim\/ug\/fmincon.html. Accessed","year":"2018","unstructured":"fmincon. https:\/\/www.mathworks.com\/help\/optim\/ug\/fmincon.html. Accessed : November 5, 2018 . fmincon. https:\/\/www.mathworks.com\/help\/optim\/ug\/fmincon.html. Accessed: November 5, 2018."},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2006.49"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2011.2161217"}],"container-title":["ACM Transactions on Embedded Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3358198","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3358198","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3358198","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T22:32:58Z","timestamp":1750199578000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3358198"}},"subtitle":["Executing Perfectly Nested Loops on Dataflow Accelerators"],"short-title":[],"issued":{"date-parts":[[2019,10,8]]},"references-count":45,"journal-issue":{"issue":"5s","published-print":{"date-parts":[[2019,10,31]]}},"alternative-id":["10.1145\/3358198"],"URL":"https:\/\/doi.org\/10.1145\/3358198","relation":{},"ISSN":["1539-9087","1558-3465"],"issn-type":[{"value":"1539-9087","type":"print"},{"value":"1558-3465","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,10,8]]},"assertion":[{"value":"2019-04-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-07-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-10-08","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}