{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:12:45Z","timestamp":1750306365465,"version":"3.41.0"},"reference-count":9,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2016,4,22]],"date-time":"2016-04-22T00:00:00Z","timestamp":1461283200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["SIGARCH Comput. Archit. News"],"published-print":{"date-parts":[[2016,4,22]]},"abstract":"<jats:p>This paper presents a novel resource management approach for efficiently managing the computation and the data movements between the host and its accelerators in a heterogeneous platform. Our approach is based on OmpSs, with support for multi-core CPUs, GPGPUs and Maxeler Data Flow Engines based on FPGA technology; it exploits data locality, data transfer costs and data dependencies. The proposed approach is supported by an offline learning process coupled with online monitoring, allowing performance to be estimated while learning from past observations during execution. Its performance is compared against the current OmpSs scheduler using five benchmarks: matrix multiplication, bitonic sort, N-body simulation, Cholesky decomposition and AdPredictor. The results show the proposed approach can achieve up to 4.25 times speed-up for Cholesky decomposition. Moreover, an evaluation with AdPredictor indicates that the FPGA version is up to 46 times faster than the CPU version for large task sizes.<\/jats:p>","DOI":"10.1145\/2927964.2927972","type":"journal-article","created":{"date-parts":[[2016,4,25]],"date-time":"2016-04-25T19:51:13Z","timestamp":1461613873000},"page":"40-45","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["A Transfer-Aware Runtime System for Heterogeneous Asynchronous Parallel Execution"],"prefix":"10.1145","volume":"43","author":[{"given":"Soukaina N.","family":"Hmid","sequence":"first","affiliation":[{"name":"Imperial College London, United Kingdom"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jose G.F.","family":"Coutinho","sequence":"additional","affiliation":[{"name":"Imperial College London, United Kingdom"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wayne","family":"Luk","sequence":"additional","affiliation":[{"name":"Imperial College London, United Kingdom"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2016,4,22]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1002\/cpe.1631"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1142\/S0129626411000151"},{"key":"e_1_2_1_3_1","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Graepel T.","year":"2010","unstructured":"T. Graepel and all. Web-Scale Bayesian Click-Through Rate Prediction for Sponsored Search Advertising in Microsoft's Bing Search Engine . In Proceedings of the International Conference on Machine Learning , 2010 . T. Graepel and all. Web-Scale Bayesian Click-Through Rate Prediction for Sponsored Search Advertising in Microsoft's Bing Search Engine. In Proceedings of the International Conference on Machine Learning, 2010."},{"key":"e_1_2_1_4_1","unstructured":"S. Gratton. Cholesky Factorization in CUDA. In http:\/\/www.ast.cam.ac.uk\/ stg20\/cuda\/cholesky\/.  S. Gratton. Cholesky Factorization in CUDA. In http:\/\/www.ast.cam.ac.uk\/ stg20\/cuda\/cholesky\/."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/1346281.1346318"},{"key":"e_1_2_1_6_1","unstructured":"Maxeler Technologies. http:\/\/www.maxeler.com\/.  Maxeler Technologies. http:\/\/www.maxeler.com\/."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISPA.2014.28"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2013.53"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/1755888.1755906"}],"container-title":["ACM SIGARCH Computer Architecture News"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2927964.2927972","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2927964.2927972","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:56:21Z","timestamp":1750222581000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2927964.2927972"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2016,4,22]]},"references-count":9,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2016,4,22]]}},"alternative-id":["10.1145\/2927964.2927972"],"URL":"https:\/\/doi.org\/10.1145\/2927964.2927972","relation":{},"ISSN":["0163-5964"],"issn-type":[{"type":"print","value":"0163-5964"}],"subject":[],"published":{"date-parts":[[2016,4,22]]},"assertion":[{"value":"2016-04-22","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}