{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T04:21:23Z","timestamp":1782361283009,"version":"3.54.5"},"reference-count":29,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2018,1,30]],"date-time":"2018-01-30T00:00:00Z","timestamp":1517270400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Laboratory for Telecommunication Sciences"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Embed. Comput. Syst."],"published-print":{"date-parts":[[2018,3,31]]},"abstract":"<jats:p>The increasing use of heterogeneous embedded systems with multi-core CPUs and Graphics Processing Units (GPUs) presents important challenges in effectively exploiting pipeline, task, and data-level parallelism to meet throughput requirements of digital signal processing applications. Moreover, in the presence of system-level memory constraints, hand optimization of code to satisfy these requirements is inefficient and error prone and can therefore, greatly slow down development time or result in highly underutilized processing resources. In this article, we present vectorization and scheduling methods to effectively exploit multiple forms of parallelism for throughput optimization on hybrid CPU-GPU platforms, while conforming to system-level memory constraints. The methods operate on synchronous dataflow representations, which are widely used in the design of embedded systems for signal and information processing. We show that our novel methods can significantly improve system throughput compared to previous vectorization and scheduling approaches under the same memory constraints. In addition, we present a practical case-study of applying our methods to significantly improve the throughput of an orthogonal frequency division multiplexing receiver system for wireless communications.<\/jats:p>","DOI":"10.1145\/3157669","type":"journal-article","created":{"date-parts":[[2018,1,31]],"date-time":"2018-01-31T13:25:40Z","timestamp":1517405140000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":8,"title":["Memory-Constrained Vectorization and Scheduling of Dataflow Graphs for Hybrid CPU-GPU Platforms"],"prefix":"10.1145","volume":"17","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5550-6311","authenticated-orcid":false,"given":"Shuoxin","family":"Lin","sequence":"first","affiliation":[{"name":"University of Maryland, MD, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jiahao","family":"Wu","sequence":"additional","affiliation":[{"name":"University of Maryland, MD, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shuvra S.","family":"Bhattacharyya","sequence":"additional","affiliation":[{"name":"University of Maryland and Tampere University of Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2018,1,30]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1002\/cpe.1631"},{"key":"e_1_2_1_2_1","doi-asserted-by":"crossref","unstructured":"S. S. Bhattacharyya E. Deprettere R. Leupers and J. Takala (Eds.). 2013. Handbook of Signal Processing Systems (second ed.). Springer.   S. S. Bhattacharyya E. Deprettere R. Leupers and J. Takala (Eds.). 2013. Handbook of Signal Processing Systems (second ed.). Springer.","DOI":"10.1007\/978-1-4614-6859-2"},{"key":"e_1_2_1_3_1","doi-asserted-by":"crossref","unstructured":"S. S. Bhattacharyya P. K. Murthy and E. A. Lee. 1996. Software Synthesis from Dataflow Graphs. Kluwer Academic.   S. S. Bhattacharyya P. K. Murthy and E. A. Lee. 1996. Software Synthesis from Dataflow Graphs. Kluwer Academic.","DOI":"10.1007\/978-1-4613-1389-2"},{"key":"e_1_2_1_4_1","volume-title":"Proceedings of the Asia South Pacific Design Automation Conference. 127--132","author":"Chen Y.","unstructured":"Y. Chen and H. Zhou . 2012. Buffer minimization in pipelined SDF scheduling on multi-core platforms . In Proceedings of the Asia South Pacific Design Automation Conference. 127--132 . Y. Chen and H. Zhou. 2012. Buffer minimization in pipelined SDF scheduling on multi-core platforms. In Proceedings of the Asia South Pacific Design Automation Conference. 127--132."},{"key":"e_1_2_1_5_1","volume-title":"Proceedings of the International Workshop on Model Based Architecting and Construction of Embedded Systems.","author":"Ciccozzi F.","year":"2013","unstructured":"F. Ciccozzi . 2013 . Automatic synthesis of heterogeneous CPU-GPU embedded applications from a UML profile . In Proceedings of the International Workshop on Model Based Architecting and Construction of Embedded Systems. F. Ciccozzi. 2013. Automatic synthesis of heterogeneous CPU-GPU embedded applications from a UML profile. In Proceedings of the International Workshop on Model Based Architecting and Construction of Embedded Systems."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2015.7178142"},{"key":"e_1_2_1_7_1","volume-title":"Proceedings of the International Workshop on Hardware\/Software Codesign. 97--101","author":"Dick R. P.","unstructured":"R. P. Dick , D. L. Rhodes , and W. Wolf . 1998. TGFF: Task graphs for free . In Proceedings of the International Workshop on Hardware\/Software Codesign. 97--101 . R. P. Dick, D. L. Rhodes, and W. Wolf. 1998. TGFF: Task graphs for free. In Proceedings of the International Workshop on Hardware\/Software Codesign. 97--101."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1142\/S0129626411000151"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACSD.2006.33"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCC.2012.67"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.5555\/2015039.2015535"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2011.52"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/1970353.1970358"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1007\/s 11265-007-0114-1"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/PROC.1987.13876"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/2906363.2906374"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/PDP.2015.29"},{"key":"e_1_2_1_18_1","volume-title":"Proceedings of the IEEE Asilomar Conference on Signals, Systems, and Computers. 104--108","author":"Massey J. W.","unstructured":"J. W. Massey , J. Starr , S. Lee , D. Lee , A. Gerstlauer , and R. W. Heath . 2012. Implementation of a real-time wireless interference alignment network . In Proceedings of the IEEE Asilomar Conference on Signals, Systems, and Computers. 104--108 . J. W. Massey, J. Starr, S. Lee, D. Lee, A. Gerstlauer, and R. W. Heath. 2012. Implementation of a real-time wireless interference alignment network. In Proceedings of the IEEE Asilomar Conference on Signals, Systems, and Computers. 104--108."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/1810479.1810481"},{"key":"e_1_2_1_20_1","volume-title":"Proceedings of the International Conference on Application Specific Array Processors.","author":"Ritz S.","unstructured":"S. Ritz , M. Pankert , and H. Meyr . 1993. Optimum vectorization of scalable synchronous dataflow graphs . In Proceedings of the International Conference on Application Specific Array Processors. S. Ritz, M. Pankert, and H. Meyr. 1993. Optimum vectorization of scalable synchronous dataflow graphs. In Proceedings of the International Conference on Application Specific Array Processors."},{"key":"e_1_2_1_21_1","volume-title":"Proceedings of the IEEE Workshop on Embedded Systems for Real-Time Multimedia. 41--50","author":"Schor L.","unstructured":"L. Schor , A. Tretter , T. Scherer , and L. Thiele . 2013. Exploiting the parallelism of heterogeneous systems using dataflow graphs on top of OpenCL . In Proceedings of the IEEE Workshop on Embedded Systems for Real-Time Multimedia. 41--50 . L. Schor, A. Tretter, T. Scherer, and L. Thiele. 2013. Exploiting the parallelism of heterogeneous systems using dataflow graphs on top of OpenCL. In Proceedings of the IEEE Workshop on Embedded Systems for Real-Time Multimedia. 41--50."},{"key":"e_1_2_1_22_1","volume-title":"Proceedings of the Wireless Innovation Conference and Product Exposition. 640--645","author":"Shen C.","unstructured":"C. Shen , W. Plishker , H. Wu , and S. S. Bhattacharyya . 2010. A lightweight dataflow approach for design and implementation of SDR systems . In Proceedings of the Wireless Innovation Conference and Product Exposition. 640--645 . C. Shen, W. Plishker, H. Wu, and S. S. Bhattacharyya. 2010. A lightweight dataflow approach for design and implementation of SDR systems. In Proceedings of the Wireless Innovation Conference and Product Exposition. 640--645."},{"key":"e_1_2_1_23_1","volume-title":"Technical Report UMIACS-TR-2011-17. Institute for Advanced Computer Studies","author":"Shen C.","year":"2011","unstructured":"C. Shen , L. Wang , I. Cho , S. Kim , S. Won , W. Plishker , and S. S. Bhattacharyya . 2011 . The DSPCAD Lightweight Dataflow Environment: Introduction to LIDE Version 0.1. Technical Report UMIACS-TR-2011-17. Institute for Advanced Computer Studies , University of Maryland at College Park. C. Shen, L. Wang, I. Cho, S. Kim, S. Won, W. Plishker, and S. S. Bhattacharyya. 2011. The DSPCAD Lightweight Dataflow Environment: Introduction to LIDE Version 0.1. Technical Report UMIACS-TR-2011-17. Institute for Advanced Computer Studies, University of Maryland at College Park."},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.5555\/1550904"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/1146909.1147138"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/71.993206"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/2442116.2442133"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/CGO.2009.20"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11265-012-0696-0"}],"container-title":["ACM Transactions on Embedded Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3157669","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3157669","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3157669","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T02:11:29Z","timestamp":1750212689000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3157669"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,1,30]]},"references-count":29,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2018,3,31]]}},"alternative-id":["10.1145\/3157669"],"URL":"https:\/\/doi.org\/10.1145\/3157669","relation":{},"ISSN":["1539-9087","1558-3465"],"issn-type":[{"value":"1539-9087","type":"print"},{"value":"1558-3465","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,1,30]]},"assertion":[{"value":"2017-01-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2017-10-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2018-01-30","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}