{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,8]],"date-time":"2025-12-08T22:31:35Z","timestamp":1765233095422,"version":"3.41.0"},"reference-count":52,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2021,4,23]],"date-time":"2021-04-23T00:00:00Z","timestamp":1619136000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Campus for Research Excellence and Technological Enterprise"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Model. Comput. Simul."],"published-print":{"date-parts":[[2021,4,30]]},"abstract":"<jats:p>\n            Spiking neural networks (SNN) are among the most computationally intensive types of simulation models, with node counts on the order of up to 10\n            <jats:sup>11<\/jats:sup>\n            . Currently, there is intensive research into hardware platforms suitable to support large-scale SNN simulations, whereas several of the most widely used simulators still rely purely on the execution on CPUs. Enabling the execution of these established simulators on heterogeneous hardware allows new studies to exploit the many-core hardware prevalent in modern supercomputing environments, while still being able to reproduce and compare with results from a vast body of existing literature. In this article, we propose a transition approach for CPU-based SNN simulators to enable the execution on heterogeneous hardware (e.g., CPUs, GPUs, and FPGAs), with only limited modifications to an existing simulator code base and without changes to model code. Our approach relies on manual porting of a small number of core simulator functionalities as found in common SNN simulators, whereas the unmodified model code is analyzed and transformed automatically. We apply our approach to the well-known simulator NEST and make a version executable on heterogeneous hardware available to the community. Our measurements show that at full utilization, a single GPU achieves the performance of about 9 CPU cores. A CPU-GPU co-execution with load balancing is also demonstrated, which shows better performance compared to CPU-only or GPU-only execution. Finally, an analytical performance model is proposed to heuristically determine the optimal parameters to execute the heterogeneous NEST.\n          <\/jats:p>","DOI":"10.1145\/3422389","type":"journal-article","created":{"date-parts":[[2021,4,23]],"date-time":"2021-04-23T16:40:24Z","timestamp":1619196024000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Transitioning Spiking Neural Network Simulators to Heterogeneous Hardware"],"prefix":"10.1145","volume":"31","author":[{"given":"Quang Anh Pham","family":"Nguyen","sequence":"first","affiliation":[{"name":"TUM Create Ltd. and Nanyang Technological University, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Philipp","family":"Andelfinger","sequence":"additional","affiliation":[{"name":"TUM Create Ltd. and Nanyang Technological University, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2204-0639","authenticated-orcid":false,"given":"Wen Jun","family":"Tan","sequence":"additional","affiliation":[{"name":"TUM Create Ltd. and Nanyang Technological University, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wentong","family":"Cai","sequence":"additional","affiliation":[{"name":"Nanyang Technological University, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Alois","family":"Knoll","sequence":"additional","affiliation":[{"name":"Techn. Universit\u00e4t M\u00fcnchen, Germany and Nanyang Technological University, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,4,23]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/WSC.2014.7020179"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/MASCOTS.2011.40"},{"key":"e_1_2_1_3_1","volume-title":"Proceedings of the Workshop on Compilers for Parallel Computers (CPC\u201910)","author":"Baghdadi Soufiane","year":"2010","unstructured":"Soufiane Baghdadi , Armin Gr\u00f6\u00dflinger , and Albert Cohen . 2010 . Putting automatic polyhedral compilation for GPGPU to work . In Proceedings of the Workshop on Compilers for Parallel Computers (CPC\u201910) . Vienna, Austria. Soufiane Baghdadi, Armin Gr\u00f6\u00dflinger, and Albert Cohen. 2010. Putting automatic polyhedral compilation for GPGPU to work. In Proceedings of the Workshop on Compilers for Parallel Computers (CPC\u201910). Vienna, Austria."},{"volume-title":"Proceedings of the International Conference on Compiler Construction. Springer","author":"Baskaran Muthu Manikandan","key":"e_1_2_1_4_1","unstructured":"Muthu Manikandan Baskaran , J. Ramanujam , and P. Sadayappan . 2010. Automatic C-to-CUDA code generation for affine programs . In Proceedings of the International Conference on Compiler Construction. Springer , Berlin, 244--263. Muthu Manikandan Baskaran, J. Ramanujam, and P. Sadayappan. 2010. Automatic C-to-CUDA code generation for affine programs. In Proceedings of the International Conference on Compiler Construction. Springer, Berlin, 244--263."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.3389\/fninf.2013.00048"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPSW.2010.5470899"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.3389\/fnbot.2018.00035"},{"key":"e_1_2_1_8_1","volume-title":"Proceedings of the International Conference on Parallel Architecture and Compilation Techniques (PACT\u201914)","author":"Bondhugula Uday","year":"2014","unstructured":"Uday Bondhugula , Vinayaka Bandishti , Albert Cohen , Guillain Potron , and Nicolas Vasilache . 2014 . Tiling and optimizing time-iterated computations over periodic domains . In Proceedings of the International Conference on Parallel Architecture and Compilation Techniques (PACT\u201914) . IEEE, IEEE, 39--50. Uday Bondhugula, Vinayaka Bandishti, Albert Cohen, Guillain Potron, and Nicolas Vasilache. 2014. Tiling and optimizing time-iterated computations over periodic domains. In Proceedings of the International Conference on Parallel Architecture and Compilation Techniques (PACT\u201914). IEEE, IEEE, 39--50."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1008925309027"},{"volume-title":"Proceedings of the International Joint Conference on Neural Networks (IJCNN\u201918)","author":"Chou Ting-Shuo","key":"e_1_2_1_10_1","unstructured":"Ting-Shuo Chou , Hirak J. Kashyap , Jinwei Xing , Stanislav Listopad , Emily L. Rounds , Michael Beyeler , Nikil Dutt , and Jeffrey L. Krichmar . 2018. CARLsim 4: An open source library for large scale, biologically detailed spiking neural network simulation using heterogeneous clusters . In Proceedings of the International Joint Conference on Neural Networks (IJCNN\u201918) . IEEE, 1--8. Ting-Shuo Chou, Hirak J. Kashyap, Jinwei Xing, Stanislav Listopad, Emily L. Rounds, Michael Beyeler, Nikil Dutt, and Jeffrey L. Krichmar. 2018. CARLsim 4: An open source library for large scale, biologically detailed spiking neural network simulation using heterogeneous clusters. In Proceedings of the International Joint Conference on Neural Networks (IJCNN\u201918). IEEE, 1--8."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-96983-1_36"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/2063384.2063396"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.parco.2012.03.004"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.parco.2014.12.003"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/ASAP.2009.24"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/ASONAM.2012.102"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.4249\/scholarpedia.1430"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.3389\/neuro.01.026.2009"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.5555\/2891307.2891313"},{"volume-title":"Proceedings of the International Conference on Software Engineering Research and Practice (SERP\u201913)","author":"Hawick K. A.","key":"e_1_2_1_20_1","unstructured":"K. A. Hawick and D. P. Playne . 2013. Simulation software generation using a domain-specific language for partial differential field equations . In Proceedings of the International Conference on Software Engineering Research and Practice (SERP\u201913) . The Steering Committee of The World Congress in Computer Science, Computer Engineering and Applied Computing (WorldComp), 69--75. K. A. Hawick and D. P. Playne. 2013. Simulation software generation using a domain-specific language for partial differential field equations. In Proceedings of the International Conference on Software Engineering Research and Practice (SERP\u201913). The Steering Committee of The World Congress in Computer Science, Computer Engineering and Applied Computing (WorldComp), 69--75."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.3389\/fninf.2012.00026"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.1201895109"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.6.1179"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.3389\/fninf.2013.00019"},{"volume-title":"ACM SIGARCH Computer Architecture News","author":"Hong Sungpack","key":"e_1_2_1_25_1","unstructured":"Sungpack Hong , Hassan Chafi , Edic Sedlar , and Kunle Olukotun . 2012. Green-Marl: A DSL for easy and efficient graph analysis . In ACM SIGARCH Computer Architecture News , Vol. 40 . ACM , ACM , 349--362. Sungpack Hong, Hassan Chafi, Edic Sedlar, and Kunle Olukotun. 2012. Green-Marl: A DSL for easy and efficient graph analysis. In ACM SIGARCH Computer Architecture News, Vol. 40. ACM, ACM, 349--362."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2018.00037"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.3389\/fninf.2017.00030"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.3389\/fninf.2018.00002"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2015.2467449"},{"key":"e_1_2_1_30_1","volume-title":"Neurorobotics: A strategic pillar of the human brain project. Sci. Robot.","author":"Knoll Alois","year":"2016","unstructured":"Alois Knoll and Marc-Oliver Gewaltig . 2016 . Neurorobotics: A strategic pillar of the human brain project. Sci. Robot. (2016), 25--34. Alois Knoll and Marc-Oliver Gewaltig. 2016. Neurorobotics: A strategic pillar of the human brain project. Sci. Robot. (2016), 25--34."},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.3389\/fninf.2017.00040"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.3389\/fninf.2014.00078"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1007\/3-540-57208-2_28"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICPP.2015.107"},{"key":"e_1_2_1_35_1","volume-title":"Proceedings of the International Conference on Networking, Architecture and Storage (NAS\u201915)","author":"Li Peilong","year":"2015","unstructured":"Peilong Li , Yan Luo , Ning Zhang , and Yu Cao . 2015 . Heterospark: A heterogeneous CPU\/GPU spark platform for machine learning algorithms . In Proceedings of the International Conference on Networking, Architecture and Storage (NAS\u201915) . IEEE, 347--348. Peilong Li, Yan Luo, Ning Zhang, and Yu Cao. 2015. Heterospark: A heterogeneous CPU\/GPU spark platform for machine learning algorithms. In Proceedings of the International Conference on Networking, Architecture and Storage (NAS\u201915). IEEE, 347--348."},{"key":"e_1_2_1_36_1","volume-title":"Jessica Mitchell, Jari Pronold, Jochen Martin Eppler, Chrisitan Keup, Alexander Peyser, Susanne Kunkel, Philipp Weidel, Yannick Nodem, et al.","author":"Linssen Charl","year":"2018","unstructured":"Charl Linssen , Mikkel Elle Lepper\u00f8d , Jessica Mitchell, Jari Pronold, Jochen Martin Eppler, Chrisitan Keup, Alexander Peyser, Susanne Kunkel, Philipp Weidel, Yannick Nodem, et al. 2018 . NEST 2.16.0. (Aug. 2018). Charl Linssen, Mikkel Elle Lepper\u00f8d, Jessica Mitchell, Jari Pronold, Jochen Martin Eppler, Chrisitan Keup, Alexander Peyser, Susanne Kunkel, Philipp Weidel, Yannick Nodem, et al. 2018. NEST 2.16.0. (Aug. 2018)."},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0893-6080(97)00011-7"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2006.11.029"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2015.2394802"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2013.2276056"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1162\/neco.2007.19.6.1437"},{"key":"e_1_2_1_42_1","volume-title":"Proceedings of the ACM SIGSIM Conference on Principles of Advanced Discrete Simulation. 115--126","author":"Pham Nguyen Quang Anh","year":"2019","unstructured":"Quang Anh Pham Nguyen , Philipp Andelfinger , Wentong Cai , and Alois Knoll . 2019 . Transitioning spiking neural network simulators to heterogeneous hardware . In Proceedings of the ACM SIGSIM Conference on Principles of Advanced Discrete Simulation. 115--126 . Quang Anh Pham Nguyen, Philipp Andelfinger, Wentong Cai, and Alois Knoll. 2019. Transitioning spiking neural network simulators to heterogeneous hardware. In Proceedings of the ACM SIGSIM Conference on Principles of Advanced Discrete Simulation. 115--126."},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/151261.151266"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.3389\/neuro.11.011.2009"},{"volume-title":"Proceedings of the Conference for Supercomputing (SC\u201914)","author":"Schenck Wolfram","key":"e_1_2_1_45_1","unstructured":"Wolfram Schenck , A. V. Adinetz , Y. V. Zaytsev , Dirk Pleiter , and A. Morrison . 2014. Performance model for large\u2013scale neural simulations with NEST . In Proceedings of the Conference for Supercomputing (SC\u201914) . Wolfram Schenck, A. V. Adinetz, Y. V. Zaytsev, Dirk Pleiter, and A. Morrison. 2014. Performance model for large\u2013scale neural simulations with NEST. In Proceedings of the Conference for Supercomputing (SC\u201914)."},{"key":"e_1_2_1_46_1","first-page":"1","article-title":"Brian2GeNN: Accelerating spiking neural network simulations with graphics hardware. Sci","volume":"10","author":"Stimberg Marcel","year":"2020","unstructured":"Marcel Stimberg , Dan F. M. Goodman , and Thomas Nowotny . 2020 . Brian2GeNN: Accelerating spiking neural network simulations with graphics hardware. Sci . Rep. 10 , 1 (2020), 1 -- 12 . Marcel Stimberg, Dan F. M. Goodman, and Thomas Nowotny. 2020. Brian2GeNN: Accelerating spiking neural network simulations with graphics hardware. Sci. Rep. 10, 1 (2020), 1--12.","journal-title":"Rep."},{"volume-title":"Proceedings of the International Joint Conference on Neural Networks. IEEE, 3118--3123","author":"Tiesel Jan-Phillip","key":"e_1_2_1_47_1","unstructured":"Jan-Phillip Tiesel and Anthony S. Maida . 2009. Using parallel GPU architecture for simulation of planar I\/F networks . In Proceedings of the International Joint Conference on Neural Networks. IEEE, 3118--3123 . Jan-Phillip Tiesel and Anthony S. Maida. 2009. Using parallel GPU architecture for simulation of planar I\/F networks. In Proceedings of the International Joint Conference on Neural Networks. IEEE, 3118--3123."},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/2400682.2400713"},{"key":"e_1_2_1_49_1","doi-asserted-by":"crossref","unstructured":"Jiajian Xiao Philipp Andelfinger Wentong Cai Paul Richmond Alois Knoll and David Eckhoff. 2020. OpenABLext: An automatic code generation framework for agent-based simulations on CPU-GPU-FPGA heterogeneous platforms. Concurr. Comput.: Pract. Exper. (2020) 32:e5807.  Jiajian Xiao Philipp Andelfinger Wentong Cai Paul Richmond Alois Knoll and David Eckhoff. 2020. OpenABLext: An automatic code generation framework for agent-based simulations on CPU-GPU-FPGA heterogeneous platforms. Concurr. Comput.: Pract. Exper. (2020) 32:e5807.","DOI":"10.1002\/cpe.5807"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/DISTRA.2018.8601016"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/3186729"},{"key":"e_1_2_1_52_1","first-page":"18854","article-title":"GeNN: A code generation framework for accelerated brain simulations. Sci","volume":"6","author":"Yavuz Esin","year":"2016","unstructured":"Esin Yavuz , James Turner , and Thomas Nowotny . 2016 . GeNN: A code generation framework for accelerated brain simulations. Sci . Rep. 6 (2016), 18854 . Esin Yavuz, James Turner, and Thomas Nowotny. 2016. GeNN: A code generation framework for accelerated brain simulations. Sci. Rep. 6 (2016), 18854.","journal-title":"Rep."}],"container-title":["ACM Transactions on Modeling and Computer Simulation"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3422389","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3422389","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T22:03:21Z","timestamp":1750197801000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3422389"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,4,23]]},"references-count":52,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2021,4,30]]}},"alternative-id":["10.1145\/3422389"],"URL":"https:\/\/doi.org\/10.1145\/3422389","relation":{},"ISSN":["1049-3301","1558-1195"],"issn-type":[{"type":"print","value":"1049-3301"},{"type":"electronic","value":"1558-1195"}],"subject":[],"published":{"date-parts":[[2021,4,23]]},"assertion":[{"value":"2020-02-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-09-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-04-23","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}