{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:25:45Z","timestamp":1750220745572,"version":"3.41.0"},"reference-count":64,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2020,9,29]],"date-time":"2020-09-29T00:00:00Z","timestamp":1601337600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Cluster of Excellence \u2018Center for Advancing Electronics Dresden\u2019"},{"name":"German Research Council (DFG) through the TraceSymm","award":["CA 1602\/4-1"],"award-info":[{"award-number":["CA 1602\/4-1"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Embed. Comput. Syst."],"published-print":{"date-parts":[[2020,11,30]]},"abstract":"<jats:p>\n            <jats:italic>Tensor contraction<\/jats:italic>\n            is a fundamental operation in many algorithms with a plethora of applications ranging from quantum chemistry over fluid dynamics and image processing to machine learning. The performance of tensor computations critically depends on the efficient utilization of on-chip\/off-chip memories. In the context of low-power embedded devices, efficient management of the memory space becomes even more crucial, in order to meet energy constraints. This work aims at investigating strategies for performance- and energy-efficient tensor contractions on embedded systems, using\n            <jats:italic>racetrack memory<\/jats:italic>\n            (RTM)-based\n            <jats:italic>scratch-pad memory<\/jats:italic>\n            (SPM) and DRAM-based off-chip memory. Compiler optimizations such as the loop access order and data layout transformations paired with architectural optimizations such as prefetching and preshifting are employed to reduce the shifting overhead in RTMs. Optimizations for off-chip memory such as memory access order, data mapping and the choice of a suitable memory access granularity are employed to reduce the contention in the off-chip memory. Experimental results demonstrate that the proposed optimizations improve the SPM performance and energy consumption by 32% and 73%, respectively, compared to an iso-capacity SRAM. The overall DRAM dynamic energy consumption improvements due to memory optimizations amount to 80%.\n          <\/jats:p>","DOI":"10.1145\/3396235","type":"journal-article","created":{"date-parts":[[2020,7,7]],"date-time":"2020-07-07T12:39:02Z","timestamp":1594125542000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Optimizing Tensor Contractions for Embedded Devices with Racetrack and DRAM Memories"],"prefix":"10.1145","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5130-9855","authenticated-orcid":false,"given":"Asif Ali","family":"Khan","sequence":"first","affiliation":[{"name":"Technische Universit\u00e4t Dresden, Dresden, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Norman A.","family":"Rink","sequence":"additional","affiliation":[{"name":"Technische Universit\u00e4t Dresden, Dresden, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fazal","family":"Hameed","sequence":"additional","affiliation":[{"name":"Institute of Space Technology, Islamabad, Pakistan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jeronimo","family":"Castrillon","sequence":"additional","affiliation":[{"name":"Technische Universit\u00e4t Dresden, Dresden, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2020,9,29]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Mart\u00edn Abadi Ashish Agarwal Paul Barham Eugene Brevdo Zhifeng Chen Craig Citro Greg S. Corrado Andy Davis Jeffrey Dean Matthieu Devin Sanjay Ghemawat Ian Goodfellow Andrew Harp Geoffrey Irving Michael Isard Rafal Jozefowicz Yangqing Jia Lukasz Kaiser Manjunath Kudlur Josh Levenberg Dan Man\u00e9 Mike Schuster Rajat Monga Sherry Moore Derek Murray Chris Olah Jonathon Shlens Benoit Steiner Ilya Sutskever Kunal Talwar Paul Tucker Vincent Vanhoucke Vijay Vasudevan Fernanda Vi\u00e9gas Oriol Vinyals Pete Warden Martin Wattenberg Martin Wicke Yuan Yu and Xiaoqiang Zheng. 2015. TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems. http:\/\/download.tensorflow.org\/paper\/whitepaper2015.pdf.  Mart\u00edn Abadi Ashish Agarwal Paul Barham Eugene Brevdo Zhifeng Chen Craig Citro Greg S. Corrado Andy Davis Jeffrey Dean Matthieu Devin Sanjay Ghemawat Ian Goodfellow Andrew Harp Geoffrey Irving Michael Isard Rafal Jozefowicz Yangqing Jia Lukasz Kaiser Manjunath Kudlur Josh Levenberg Dan Man\u00e9 Mike Schuster Rajat Monga Sherry Moore Derek Murray Chris Olah Jonathon Shlens Benoit Steiner Ilya Sutskever Kunal Talwar Paul Tucker Vincent Vanhoucke Vijay Vasudevan Fernanda Vi\u00e9gas Oriol Vinyals Pete Warden Martin Wattenberg Martin Wicke Yuan Yu and Xiaoqiang Zheng. 2015. TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems. http:\/\/download.tensorflow.org\/paper\/whitepaper2015.pdf."},{"volume-title":"Ullman","year":"2014","author":"Aho Alfred V.","key":"e_1_2_1_2_1"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.5555\/2830689.2830711"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2004.840311"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.25080\/Majora-92bf1922-003"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2020.2975719"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMSCS.2017.2771750"},{"key":"e_1_2_1_8_1","unstructured":"K. Chandrasekar C. Weis Y. Li S. Goossens M. Jung O. Naji B. Akesson N. Wehn and K. Goossens. [n.d.]. DRAMPower: Open-source DRAM Power and Energy Estimation Tool. http:\/\/www.drampower.info.  K. Chandrasekar C. Weis Y. Li S. Goossens M. Jung O. Naji B. Akesson N. Wehn and K. Goossens. [n.d.]. DRAMPower: Open-source DRAM Power and Energy Estimation Tool. http:\/\/www.drampower.info."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2002.1058095"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/2541940.2541967"},{"volume-title":"Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18)","year":"2018","author":"Chen Tianqi","key":"e_1_2_1_11_1"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2016.2537400"},{"volume-title":"Automated empirical optimizations of software and the ATLAS project. Parallel Comput. 27 (01","year":"2001","author":"Whaley R. Clinton","key":"e_1_2_1_13_1"},{"volume-title":"Proceedings of the 19th Annual ACM Symposium on Theory of Computing (STOC\u201987)","author":"Coppersmith D.","key":"e_1_2_1_14_1"},{"key":"e_1_2_1_15_1","doi-asserted-by":"crossref","unstructured":"Paul Feautrier and Christian Lengauer. 2011. Polyhedron Model. Springer US Boston MA 1581--1592. DOI:https:\/\/doi.org\/10.1007\/978-0-387-09766-4_502  Paul Feautrier and Christian Lengauer. 2011. Polyhedron Model. Springer US Boston MA 1581--1592. DOI:https:\/\/doi.org\/10.1007\/978-0-387-09766-4_502","DOI":"10.1007\/978-0-387-09766-4_502"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/3235029"},{"volume-title":"Proceedings of the Design, Automation Test in Europe Conference Exhibition (DATE). 828--831","author":"Goossens S.","key":"e_1_2_1_17_1"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/1356052.1356053"},{"volume-title":"Proceedings of the International Conference on Computational Sciences - Part I (ICCS\u201901)","author":"Gunnels John A.","key":"e_1_2_1_19_1"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2018.2804938"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2012.2202700"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCAD.2004.1382555"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2005.859478"},{"volume-title":"Proceedings of the 38th Annual Design Automation Conference (DAC\u201901)","author":"Kandemir M.","key":"e_1_2_1_24_1"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.23919\/DATE48585.2020.9116245"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/LCA.2019.2899306"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3372489"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/3316482.3326351"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/CGO.2019.8661182"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3133901"},{"volume-title":"Optimizing matrix multiplication for a short-vector SIMD architecture\u2014CELL processor. Parallel Comput. 35 (03","year":"2009","author":"Kurzak Jakub","key":"e_1_2_1_31_1"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.7873\/DATE.2015.0182"},{"key":"e_1_2_1_33_1","first-page":"3","article-title":"Basic linear algebra subprograms for Fortran usage","volume":"5","author":"Lawson C. L.","year":"1979","journal-title":"ACM Trans. Math. Softw."},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2017.2690855"},{"volume-title":"High-performance tensor contraction without BLAS. CoRR abs\/1607.00291","year":"2016","author":"Matthews Devin","key":"e_1_2_1_35_1"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/305138.305230"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2014.2324563"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.3390\/jlpea7030023"},{"key":"e_1_2_1_39_1","unstructured":"Steven S. Muchnick. 1997. Advanced Compiler Design and Implementation. Morgan Kaufmann.  Steven S. Muchnick. 1997. Advanced Compiler Design and Implementation. Morgan Kaufmann."},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISLPED.2019.8824954"},{"volume-title":"Proceedings of the Conference on High Performance Computing for Computational Science - VECPAR","year":"2006","author":"Ohshima S.","key":"e_1_2_1_41_1"},{"volume-title":"Proceedings of HPCMO User Group Conference.","author":"Park N.","key":"e_1_2_1_42_1"},{"key":"e_1_2_1_43_1","doi-asserted-by":"crossref","unstructured":"S. Parkin M. Hayashi and L. Thomas. 2008. Magnetic domain-wall racetrack memory. Science 320 5873 (2008) 190--194.  S. Parkin M. Hayashi and L. Thomas. 2008. Magnetic domain-wall racetrack memory. Science 320 5873 (2008) 190--194.","DOI":"10.1126\/science.1145799"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1038\/nnano.2015.41"},{"key":"e_1_2_1_45_1","unstructured":"Adam Paszke Sam Gross Soumith Chintala Gregory Chanan Edward Yang Zachary DeVito Zeming Lin Alban Desmaison Luca Antiga and Adam Lerer. 2017. Automatic differentiation in PyTorch. In NIPS-W.  Adam Paszke Sam Gross Soumith Chintala Gregory Chanan Edward Yang Zachary DeVito Zeming Lin Alban Desmaison Luca Antiga and Adam Lerer. 2017. Automatic differentiation in PyTorch. In NIPS-W."},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2004.840306"},{"volume-title":"Proceedings of the Real World Domain Specific Languages Workshop 2018 (RWDSL2018)","year":"1838","author":"Rink N. A.","key":"e_1_2_1_47_1"},{"volume-title":"Proceedings of the 32nd International Symposium on Computer Architecture (ISCA). 128--138","author":"Rixner S.","key":"e_1_2_1_48_1"},{"volume-title":"Proceedings of the 2016 International Symposium on Code Generation and Optimization (CGO\u201916)","author":"Daniele","key":"e_1_2_1_49_1"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/3157733"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/2463209.2488799"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/3278122.3278131"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1109\/IEDM.2011.6131603"},{"volume-title":"Tensor comprehensions: Framework-agnostic high-performance machine learning abstractions. CoRR abs\/1802.04730","year":"2018","author":"Vasilache Nicolas","key":"e_1_2_1_54_1"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/2213977.2214056"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1145\/2333660.2333707"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/2678373.2665710"},{"key":"e_1_2_1_58_1","doi-asserted-by":"crossref","unstructured":"D. Wang L. Ma M. Zhang J. An H. Li and Y. Chen. 2017. Shift-optimized energy-efficient racetrack-based main memory. Journal of Circuits Systems and Computers 27 (09 2017) 1--16. DOI:https:\/\/doi.org\/10.1142\/S0218126618500810  D. Wang L. Ma M. Zhang J. An H. Li and Y. Chen. 2017. Shift-optimized energy-efficient racetrack-based main memory. Journal of Circuits Systems and Computers 27 (09 2017) 1--16. DOI:https:\/\/doi.org\/10.1142\/S0218126618500810","DOI":"10.1142\/S0218126618500810"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2015.2422846"},{"volume-title":"Proceedings of the 1998 ACM\/IEEE Conference on Supercomputing (SC\u201998)","author":"Clint Whaley R.","key":"e_1_2_1_60_1"},{"volume-title":"Phase change memory. 98 (12","year":"2010","author":"Philip Wong H.-S.","key":"e_1_2_1_61_1"},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMSCS.2016.2536020"},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1109\/ASPDAC.2015.7058988"},{"key":"e_1_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-20119-1_2"}],"container-title":["ACM Transactions on Embedded Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3396235","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3396235","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T22:38:45Z","timestamp":1750199925000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3396235"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,9,29]]},"references-count":64,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2020,11,30]]}},"alternative-id":["10.1145\/3396235"],"URL":"https:\/\/doi.org\/10.1145\/3396235","relation":{},"ISSN":["1539-9087","1558-3465"],"issn-type":[{"type":"print","value":"1539-9087"},{"type":"electronic","value":"1558-3465"}],"subject":[],"published":{"date-parts":[[2020,9,29]]},"assertion":[{"value":"2019-11-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-04-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-09-29","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}