{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,5]],"date-time":"2026-06-05T16:01:58Z","timestamp":1780675318296,"version":"3.54.1"},"reference-count":38,"publisher":"Springer Science and Business Media LLC","issue":"12","license":[{"start":{"date-parts":[[2021,5,15]],"date-time":"2021-05-15T00:00:00Z","timestamp":1621036800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,5,15]],"date-time":"2021-05-15T00:00:00Z","timestamp":1621036800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100012164","name":"National High-tech Research and Development Program","doi-asserted-by":"publisher","award":["2014AA01A301"],"award-info":[{"award-number":["2014AA01A301"]}],"id":[{"id":"10.13039\/501100012164","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Supercomput"],"published-print":{"date-parts":[[2021,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>The heterogeneous many-core architecture plays an important role in the fields of high-performance computing and scientific computing. It uses accelerator cores with on-chip memories to improve performance and reduce energy consumption. Scratchpad memory (SPM) is a kind of fast on-chip memory with lower energy consumption compared with a hardware cache. However, data transfer between SPM and off-chip memory can be managed only by a programmer or compiler. In this paper, we propose a compiler-directed multithreaded SPM data transfer model (MSDTM) to optimize the process of data transfer in a heterogeneous many-core architecture. We use compile-time analysis to classify data accesses, check dependences and determine the allocation of data transfer operations. We further present the data transfer performance model to derive the optimal granularity of data transfer and select the most profitable data transfer strategy. We implement the proposed MSDTM on the GCC complier and evaluate it on Sunway TaihuLight with selected test cases from benchmarks and scientific computing applications. The experimental result shows that the proposed MSDTM improves the application execution time by 5.49<jats:inline-formula><jats:alternatives><jats:tex-math>$$\\times$$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                  <mml:mo>\u00d7<\/mml:mo>\n                <\/mml:math><\/jats:alternatives><\/jats:inline-formula> and achieves an energy saving of 5.16<jats:inline-formula><jats:alternatives><jats:tex-math>$$\\times$$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                  <mml:mo>\u00d7<\/mml:mo>\n                <\/mml:math><\/jats:alternatives><\/jats:inline-formula> on average.<\/jats:p>","DOI":"10.1007\/s11227-021-03853-x","type":"journal-article","created":{"date-parts":[[2021,5,15]],"date-time":"2021-05-15T20:02:22Z","timestamp":1621108942000},"page":"14502-14524","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["Compiler-directed scratchpad memory data transfer optimization for multithreaded applications on a heterogeneous many-core architecture"],"prefix":"10.1007","volume":"77","author":[{"given":"Xiaohan","family":"Tao","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1835-5419","authenticated-orcid":false,"given":"Jianmin","family":"Pang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jinlong","family":"Xu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yu","family":"Zhu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2021,5,15]]},"reference":[{"key":"3853_CR1","doi-asserted-by":"crossref","unstructured":"Ao Y, Yang C, Wang X, Xue W, Fu H, Liu F, Gan L, Xu P, Ma W (2017) 26 pflops stencil computations for atmospheric modeling on sunway taihulight. In: 2017 IEEE International Parallel and Distributed Processing Symposium (IPDPS). IEEE, pp 535\u2013544","DOI":"10.1109\/IPDPS.2017.9"},{"issue":"3","key":"3853_CR2","doi-asserted-by":"publisher","first-page":"63","DOI":"10.1177\/109434209100500306","volume":"5","author":"D Bailey","year":"1991","unstructured":"Bailey D, Barszcz E, Barton J, Browning D, Carter R, Dagum L, Fatoohi R, Frederickson P, Lasinski T, Schreiber R, Simon H, Venkatakrishnan V, Weeratunga S (1991) The NAS parallel benchmarks. Int J Supercomput Appl 5(3):63\u201373. https:\/\/doi.org\/10.1177\/109434209100500306","journal-title":"Int J Supercomput Appl"},{"key":"3853_CR3","doi-asserted-by":"crossref","unstructured":"Banakar R, Steinke S, Lee BS, Balakrishnan M, Marwedel P (2002) Scratchpad memory: a design alternative for cache on-chip memory in embedded systems. In: Proceedings of the Tenth International Symposium on Hardware\/Software Codesign. CODES 2002 (IEEE Cat. No. 02TH8627). IEEE, pp 73\u201378","DOI":"10.1145\/774789.774805"},{"key":"3853_CR4","unstructured":"Bandyopadhyay S (2006) Automated memory allocation of actor code and data buffer in heterochronous dataflow models to scratchpad memory. Master\u2019s thesis, EECS Department, University of California, Berkeley"},{"key":"3853_CR5","doi-asserted-by":"crossref","unstructured":"Borkar S (2007) Thousand core chips: a technology perspective. In: Proceedings of the 44th Annual Design Automation Conference, pp 746\u2013749","DOI":"10.1145\/1278480.1278667"},{"issue":"5","key":"3853_CR6","doi-asserted-by":"publisher","first-page":"559","DOI":"10.1147\/rd.515.0559","volume":"51","author":"T Chen","year":"2007","unstructured":"Chen T, Raghavan R, Dale JN, Iwata E (2007) Cell broadband engine architecture and its first implementation: a performance view. IBM J Res Dev 51(5):559\u2013572","journal-title":"IBM J Res Dev"},{"key":"3853_CR7","doi-asserted-by":"crossref","unstructured":"Chen T, Sura Z, O\u2019Brien K, O\u2019Brien JK (2006) Optimizing the use of static buffers for DMA on a cell chip. In: International Workshop on Languages and Compilers for Parallel Computing. Springer, pp 314\u2013329","DOI":"10.1007\/978-3-540-72521-3_23"},{"key":"3853_CR8","doi-asserted-by":"crossref","unstructured":"Cho D, Pasricha S, Issenin I, Dutt N, Paek Y, Ko S (2008) Compiler driven data layout optimization for regular\/irregular array access patterns. In: Proceedings of the 2008 ACM SIGPLAN-SIGBED Conference on Languages, Compilers, and Tools for Embedded Systems, pp 41\u201350","DOI":"10.1145\/1379023.1375664"},{"key":"3853_CR9","unstructured":"Dongarra J (2016) Report on the sunway taihulight system. Technical report, UT-EECS-16-742. http:\/\/www.netlib.org\/utk\/people\/JackDongarra\/PAPERS\/sunway-report-2016.pdf"},{"key":"3853_CR10","first-page":"1581","volume-title":"Polyhedron model","author":"P Feautrier","year":"2011","unstructured":"Feautrier P, Lengauer C (2011) Polyhedron model. Springer, Boston, pp 1581\u20131592"},{"key":"3853_CR11","doi-asserted-by":"crossref","unstructured":"Francesco P, Marchal P, Atienza D, Benini L, Catthoor F, Mendias JM (2004) An integrated hardware\/software approach for run-time scratchpad management. In: Proceedings of the 41st Annual Design Automation Conference, pp 238\u2013243","DOI":"10.1145\/996566.996634"},{"issue":"7","key":"3853_CR12","doi-asserted-by":"publisher","first-page":"072001","DOI":"10.1007\/s11432-016-5588-7","volume":"59","author":"H Fu","year":"2016","unstructured":"Fu H, Liao J, Yang J, Wang L, Song Z, Huang X, Yang C, Xue W, Liu F, Qiao F et al (2016) The sunway taihulight supercomputer: system and applications. Sci China Inf Sci 59(7):072001","journal-title":"Sci China Inf Sci"},{"key":"3853_CR13","doi-asserted-by":"crossref","unstructured":"Gao Y, Zhang P (2016) A survey of homogeneous and heterogeneous system architectures in high performance computing. In: 2016 IEEE International Conference on Smart Cloud (SmartCloud). IEEE, pp 170\u2013175","DOI":"10.1109\/SmartCloud.2016.36"},{"key":"3853_CR14","doi-asserted-by":"crossref","unstructured":"Grosser T, Cohen A, Kelly PH, Ramanujam J, Sadayappan P, Verdoolaege S (2013) Split tiling for gpus: automatic parallelization using trapezoidal tiles. In: Proceedings of the 6th Workshop on General Purpose Processor Using Graphics Processing Units, pp 24\u201331","DOI":"10.1145\/2458523.2458526"},{"issue":"13","key":"3853_CR15","first-page":"11-02","volume":"6","author":"L Gwennap","year":"2011","unstructured":"Gwennap L (2011) Adapteva: more flops, less watts. Microprocess Rep 6(13):11\u201302","journal-title":"Microprocess Rep"},{"issue":"4","key":"3853_CR16","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/1186736.1186737","volume":"34","author":"JL Henning","year":"2006","unstructured":"Henning JL (2006) Spec cpu2006 benchmark descriptions. ACM SIGARCH Comput Archit News 34(4):1\u201317","journal-title":"ACM SIGARCH Comput Archit News"},{"key":"3853_CR17","doi-asserted-by":"crossref","unstructured":"Janapsatya A, Parameswaran S, Ignjatovic A (2004) Hardware\/software managed scratchpad memory for embedded system. In: IEEE\/ACM International Conference on Computer Aided Design, 2004. ICCAD-2004. IEEE, pp 370\u2013377","DOI":"10.1109\/ICCAD.2004.1382603"},{"key":"3853_CR18","doi-asserted-by":"publisher","unstructured":"Kelly W, Pugh W (1995) A unifying framework for iteration reordering transformations. In: Proceedings 1st International Conference on Algorithms and Architectures for Parallel Processing, vol 1, pp 153\u2013162. https:\/\/doi.org\/10.1109\/ICAPP.1995.472180","DOI":"10.1109\/ICAPP.1995.472180"},{"key":"3853_CR19","volume-title":"Optimizing compilers for modern architectures: a dependence-based approach","author":"K Kennedy","year":"2001","unstructured":"Kennedy K, Allen JR (2001) Optimizing compilers for modern architectures: a dependence-based approach. Morgan Kaufmann Publishers Inc, Burlington"},{"issue":"3","key":"3853_CR20","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/1582710.1582711","volume":"6","author":"L Li","year":"2009","unstructured":"Li L, Feng H, Xue J (2009) Compiler-directed scratchpad memory management via graph coloring. ACM Trans Archit Code Optim 6(3):1\u201317","journal-title":"ACM Trans Archit Code Optim"},{"key":"3853_CR21","doi-asserted-by":"crossref","unstructured":"Li P, Brunet E, Namyst R (2013) High performance code generation for stencil computation on heterogeneous multi-device architectures. In: 2013 IEEE 10th International Conference on High Performance Computing and Communications & 2013 IEEE International Conference on Embedded and Ubiquitous Computing. IEEE, pp 1512\u20131518","DOI":"10.1109\/HPCC.and.EUC.2013.213"},{"key":"3853_CR22","doi-asserted-by":"crossref","unstructured":"Lim AW, Liao SW, Lam MS (2001) Blocking and array contraction across arbitrarily nested loops using affine partitioning. In: Proceedings of the Eighth ACM SIGPLAN Symposium on Principles and practices of Parallel Programming, pp 103\u2013112","DOI":"10.1145\/568014.379586"},{"key":"3853_CR23","doi-asserted-by":"crossref","unstructured":"Liu T, Lin H, Chen T, O\u2019Brien JK, Shao L (2009) Dbdb: optimizing dma transfer for the cell be architecture. In: Proceedings of the 23rd International Conference on Supercomputing, pp 36\u201345","DOI":"10.1145\/1542275.1542286"},{"issue":"2","key":"3853_CR24","doi-asserted-by":"publisher","first-page":"222","DOI":"10.1109\/TC.2010.199","volume":"61","author":"A Marongiu","year":"2010","unstructured":"Marongiu A, Benini L (2010) An openmp compiler for efficient use of distributed scratchpad memory in mpsocs. IEEE Trans Comput 61(2):222\u2013236","journal-title":"IEEE Trans Comput"},{"issue":"2","key":"3853_CR25","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/2739047","volume":"12","author":"I Pananilath","year":"2015","unstructured":"Pananilath I, Acharya A, Vasista V, Bondhugula U (2015) An optimizing code generator for a class of lattice-Boltzmann computations. ACM Trans Archit Code Optim 12(2):1\u201323","journal-title":"ACM Trans Archit Code Optim"},{"issue":"3","key":"3853_CR26","doi-asserted-by":"publisher","first-page":"682","DOI":"10.1145\/348019.348570","volume":"5","author":"PR Panda","year":"2000","unstructured":"Panda PR, Dutt ND, Nicolau A (2000) On-chip vs. off-chip memory: the data partitioning problem in embedded processor-based systems. ACM Trans Des Autom Electron Syst 5(3):682\u2013704","journal-title":"ACM Trans Des Autom Electron Syst"},{"key":"3853_CR27","doi-asserted-by":"crossref","unstructured":"Rahman SMF, Yi Q, Qasem A (2011) Understanding stencil code performance on multicore architectures. In: Proceedings of the 8th ACM International Conference on Computing Frontiers, pp 1\u201310","DOI":"10.1145\/2016604.2016641"},{"key":"3853_CR28","unstructured":"Ren J, Luo J, Wu K, Zhang M, Li D (2019) Sentinel: Runtime data management on heterogeneous main memorysystems for deep learning"},{"key":"3853_CR29","unstructured":"Riesbeck CK, Martin C (1986) Direct memory access parsing. Experience, memory and reasoning, pp 209\u2013226"},{"issue":"4","key":"3853_CR30","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/2086696.2086716","volume":"8","author":"S Saidi","year":"2012","unstructured":"Saidi S, Tendulkar P, Lepley T, Maler O (2012) Optimizing explicit data transfers for data parallel applications on the cell architecture. ACM Trans Archit Code Optim 8(4):1\u201320","journal-title":"ACM Trans Archit Code Optim"},{"key":"3853_CR31","doi-asserted-by":"crossref","unstructured":"Sancho JC, Kerbyson DJ (2008) Analysis of double buffering on two different multicore architectures: Quad-core opteron and the cell-be. In: 2008 IEEE International Symposium on Parallel and Distributed Processing. IEEE, pp 1\u201312","DOI":"10.1109\/IPDPS.2008.4536316"},{"key":"3853_CR32","doi-asserted-by":"crossref","unstructured":"Sandrieser M, Benkner S, Pllana S (2011) Explicit platform descriptions for heterogeneous many-core architectures. In: 2011 IEEE International Symposium on Parallel and Distributed Processing Workshops and Phd Forum. IEEE, pp 1292\u20131299","DOI":"10.1109\/IPDPS.2011.280"},{"key":"3853_CR33","doi-asserted-by":"publisher","unstructured":"Shao Z, Li R, Hu D, Liao X, Jin H (2019) Improving performance of graph processing on fpga-dram platform by two-level vertex caching. In: Proceedings of the 2019 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays, FPGA \u201919, pp 320\u2013329. Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/3289602.3293900","DOI":"10.1145\/3289602.3293900"},{"issue":"3","key":"3853_CR34","doi-asserted-by":"publisher","first-page":"4","DOI":"10.1145\/3391920","volume":"13","author":"Z Shao","year":"2020","unstructured":"Shao Z, Liu C, Li R, Liao X, Jin H (2020) Processing grid-format real-world graphs on dram-based fpga accelerators with application-specific caching mechanisms. ACM Trans. Reconfig. Technol. Syst. 13(3):4. https:\/\/doi.org\/10.1145\/3391920","journal-title":"ACM Trans. Reconfig. Technol. Syst."},{"key":"3853_CR35","doi-asserted-by":"publisher","DOI":"10.1137\/1.9781611970999","volume-title":"Computational frameworks for the fast Fourier transform","author":"C Van Loan","year":"1992","unstructured":"Van Loan C (1992) Computational frameworks for the fast Fourier transform, vol 10. Siam, Philadelphia"},{"issue":"1","key":"3853_CR36","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3301308","volume":"18","author":"V Venkataramani","year":"2019","unstructured":"Venkataramani V, Chan MC, Mitra T (2019) Scratchpad-memory management for multi-threaded applications on many-core architectures. ACM Trans Embed Comput Syst 18(1):1\u201328","journal-title":"ACM Trans Embed Comput Syst"},{"issue":"8","key":"3853_CR37","doi-asserted-by":"publisher","first-page":"802","DOI":"10.1109\/TVLSI.2006.878469","volume":"14","author":"M Verma","year":"2006","unstructured":"Verma M, Marwedel P (2006) Overlay techniques for scratchpad memories in low power embedded processors. IEEE Trans Very Large Scale Integr Syst 14(8):802\u2013815","journal-title":"IEEE Trans Very Large Scale Integr Syst"},{"key":"3853_CR38","doi-asserted-by":"publisher","first-page":"1878","DOI":"10.1109\/TPDS.2020.2978045","volume":"31","author":"P Zhang","year":"2020","unstructured":"Zhang P, Fang J, Yang C, Huang C, Tang T, Wang Z (2020) Optimizing streaming parallelism on heterogeneous many-core architectures. IEEE Trans Parallel Distrib Syst 31:1878\u20131896","journal-title":"IEEE Trans Parallel Distrib Syst"}],"container-title":["The Journal of Supercomputing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-021-03853-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11227-021-03853-x\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-021-03853-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,11,16]],"date-time":"2021-11-16T11:27:46Z","timestamp":1637062066000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11227-021-03853-x"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,5,15]]},"references-count":38,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2021,12]]}},"alternative-id":["3853"],"URL":"https:\/\/doi.org\/10.1007\/s11227-021-03853-x","relation":{},"ISSN":["0920-8542","1573-0484"],"issn-type":[{"value":"0920-8542","type":"print"},{"value":"1573-0484","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,5,15]]},"assertion":[{"value":"29 April 2021","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 May 2021","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}