{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,8]],"date-time":"2025-12-08T22:28:44Z","timestamp":1765232924422,"version":"3.41.0"},"reference-count":28,"publisher":"Association for Computing Machinery (ACM)","issue":"5","license":[{"start":{"date-parts":[[2020,9,26]],"date-time":"2020-09-26T00:00:00Z","timestamp":1601078400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"HiPEAC Network of Excellence and the European Research Council","award":["772773"],"award-info":[{"award-number":["772773"]}]},{"DOI":"10.13039\/501100004837","name":"Spanish Ministry of Science and Innovation","doi-asserted-by":"crossref","award":["TIN2015-65316-P"],"award-info":[{"award-number":["TIN2015-65316-P"]}],"id":[{"id":"10.13039\/501100004837","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Spanish Ministry of Economy and Competitiveness","award":["FJCI-2017-34095"],"award-info":[{"award-number":["FJCI-2017-34095"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Embed. Comput. Syst."],"published-print":{"date-parts":[[2020,9,30]]},"abstract":"<jats:p>Critical real-time systems require strict resource provisioning in terms of memory and timing. The constant need for higher performance in these systems has led industry to recently include GPUs. However, GPU software ecosystems are by their nature closed source, forcing system engineers to consider them as black boxes, complicating resource provisioning. In this work, we reverse engineer the internal operations of the GPU system software to increase the understanding of their observed behaviour and how resources are internally managed. We present our methodology that is incorporated in GMAI (GPU Memory Allocation Inspector), a tool that allows system engineers to accurately determine the exact amount of resources required by their critical systems, avoiding underprovisioning. We first apply our methodology on a wide range of GPU hardware from different vendors showing its generality in obtaining the properties of the GPU memory allocators. Next, we demonstrate the benefits of such knowledge in resource provisioning of two case studies from the automotive domain, where the actual memory consumption is up to 5.6\u00d7 more than the memory requested by the application.<\/jats:p>","DOI":"10.1145\/3391896","type":"journal-article","created":{"date-parts":[[2020,7,7]],"date-time":"2020-07-07T12:39:02Z","timestamp":1594125542000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":8,"title":["GMAI"],"prefix":"10.1145","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2426-306X","authenticated-orcid":false,"given":"Alejandro J.","family":"Calder\u00f3n","sequence":"first","affiliation":[{"name":"Universitat Polit\u00e8cnica de Catalunya and Ikerlan Technology Research Centre, Mondragon, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9751-1058","authenticated-orcid":false,"given":"Leonidas","family":"Kosmidis","sequence":"additional","affiliation":[{"name":"Barcelona Supercomputing Center (BSC), Barcelona, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Carlos F.","family":"Nicol\u00e1s","sequence":"additional","affiliation":[{"name":"Ikerlan Technology Research Centre, Mondragon, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Francisco J.","family":"Cazorla","sequence":"additional","affiliation":[{"name":"Barcelona Supercomputing Center (BSC), Barcelona, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Peio","family":"Onaindia","sequence":"additional","affiliation":[{"name":"Ikerlan Technology Research Centre, Mondragon, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2020,9,26]]},"reference":[{"volume-title":"Proceedings of the Real-time Systems Symposium","author":"Amert T.","unstructured":"T. Amert , N. Otterness , M. Yang , J. H. Anderson , and F. Donelson Smith . 2018. GPU scheduling on the NVIDIA TX2: Hidden details revealed . In Proceedings of the Real-time Systems Symposium , Vol. 2018-January. T. Amert, N. Otterness, M. Yang, J. H. Anderson, and F. Donelson Smith. 2018. GPU scheduling on the NVIDIA TX2: Hidden details revealed. In Proceedings of the Real-time Systems Symposium, Vol. 2018-January.","key":"e_1_2_1_1_1"},{"unstructured":"ARINC. 2010. Avionics Application Software Standard Interface: ARINC Specification 653P1-3. Aeronautical Radio. Retrieved from https:\/\/www.aviation-ia.com\/product-categories\/600-series.  ARINC. 2010. Avionics Application Software Standard Interface: ARINC Specification 653P1-3. Aeronautical Radio. Retrieved from https:\/\/www.aviation-ia.com\/product-categories\/600-series.","key":"e_1_2_1_2_1"},{"key":"e_1_2_1_3_1","volume-title":"Retrieved on","author":"AUTOSAR.","year":"2019","unstructured":"AUTOSAR. 2019. AUTOSAR. Retrieved on April 2019 from https:\/\/www.autosar.org. AUTOSAR. 2019. AUTOSAR. Retrieved on April 2019 from https:\/\/www.autosar.org."},{"volume-title":"Proceedings of the International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS\u201900)","author":"Berger E. D.","unstructured":"E. D. Berger , K. S. McKinley , R. D. Blumofe , and P. R. Wilson . 2000. Hoard: A scalable memory allocator for multithreaded applications . In Proceedings of the International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS\u201900) . 117--128. E. D. Berger, K. S. McKinley, R. D. Blumofe, and P. R. Wilson. 2000. Hoard: A scalable memory allocator for multithreaded applications. In Proceedings of the International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS\u201900). 117--128.","key":"e_1_2_1_4_1"},{"key":"e_1_2_1_5_1","volume-title":"GMAI: GPU Memory Allocation Inspector.","author":"Calder\u00f3n A. J.","year":"2019","unstructured":"A. J. Calder\u00f3n , L. Kosmidis , C. F. Nicol\u00e1s , F. J. Cazorla , and P. Onaindia . 2019 . GMAI: GPU Memory Allocation Inspector. Retrieved from https:\/\/github.com\/ajcalderont\/gmai. A. J. Calder\u00f3n, L. Kosmidis, C. F. Nicol\u00e1s, F. J. Cazorla, and P. Onaindia. 2019. GMAI: GPU Memory Allocation Inspector. Retrieved from https:\/\/github.com\/ajcalderont\/gmai."},{"volume-title":"Proceedings of the Working Conference on Reverse Engineering (WCRE\u201913)","author":"Chen X.","unstructured":"X. Chen , A. Slowinska , and H. Bos . 2013. Who allocated my memory? Detecting custom memory allocators in C binaries . In Proceedings of the Working Conference on Reverse Engineering (WCRE\u201913) . 22--31. X. Chen, A. Slowinska, and H. Bos. 2013. Who allocated my memory? Detecting custom memory allocators in C binaries. In Proceedings of the Working Conference on Reverse Engineering (WCRE\u201913). 22--31.","key":"e_1_2_1_6_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_7_1","DOI":"10.1109\/TPDS.2018.2812853"},{"unstructured":"Free Software Foundation. 2019. The GNU Allocator. Retrieved from https:\/\/www.gnu.org\/software\/libc\/manual\/html_node\/The-GNU-Allocator.html.  Free Software Foundation. 2019. The GNU Allocator. Retrieved from https:\/\/www.gnu.org\/software\/libc\/manual\/html_node\/The-GNU-Allocator.html.","key":"e_1_2_1_8_1"},{"unstructured":"Green Hills Software. 1996. Integrity RTOS. Retrieved from https:\/\/www.ghs.com\/products\/rtos\/integrity.html.  Green Hills Software. 1996. Integrity RTOS. Retrieved from https:\/\/www.ghs.com\/products\/rtos\/integrity.html.","key":"e_1_2_1_9_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_10_1","DOI":"10.1016\/j.jss.2005.09.003"},{"volume-title":"Proceedings of the 10th IEEE International Conference on Computer and Information Technology (CIT\u201910) and 7th IEEE International Conference on Embedded Software and Systems (ICESS\u201910)","author":"Huang X.","unstructured":"X. Huang , C. I. Rodrigues , S. Jones , I. Buck , and W. Hwu . 2010. XMalloc: A scalable lock-free dynamic memory allocator for many-core machines . In Proceedings of the 10th IEEE International Conference on Computer and Information Technology (CIT\u201910) and 7th IEEE International Conference on Embedded Software and Systems (ICESS\u201910) (ScalCom\u201910). 1134--1139. X. Huang, C. I. Rodrigues, S. Jones, I. Buck, and W. Hwu. 2010. XMalloc: A scalable lock-free dynamic memory allocator for many-core machines. In Proceedings of the 10th IEEE International Conference on Computer and Information Technology (CIT\u201910) and 7th IEEE International Conference on Embedded Software and Systems (ICESS\u201910) (ScalCom\u201910). 1134--1139.","key":"e_1_2_1_11_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_12_1","DOI":"10.1007\/s11227-011-0680-7"},{"key":"e_1_2_1_13_1","volume-title":"Getting the Most from OpenCL 1.2: How to Increase Performance by Minimizing Buffer Copies on Intel Processor Graphics. Retrieved on","author":"Intel Corporation","year":"2019","unstructured":"Intel Corporation . 2014. Getting the Most from OpenCL 1.2: How to Increase Performance by Minimizing Buffer Copies on Intel Processor Graphics. Retrieved on October 2019 from https:\/\/software.intel.com\/en-us\/articles\/getting-the-most-from-opencl-12-how-to-increase-performance-by-minimizing-buffer-copies-on-intel-processor-graphics. Intel Corporation. 2014. Getting the Most from OpenCL 1.2: How to Increase Performance by Minimizing Buffer Copies on Intel Processor Graphics. Retrieved on October 2019 from https:\/\/software.intel.com\/en-us\/articles\/getting-the-most-from-opencl-12-how-to-increase-performance-by-minimizing-buffer-copies-on-intel-processor-graphics."},{"volume-title":"Proceedings of the IEEE\/ACM International Conference on Computer-aided Design, Digest of Technical Papers (ICCAD\u201918)","author":"Kosmidis L.","unstructured":"L. Kosmidis , C. Maxim , V. Jegu , F. Vatrinet , and F. J. Cazorla . 2018. Industrial experiences with resource management under software randomization in ARINC653 avionics environments . In Proceedings of the IEEE\/ACM International Conference on Computer-aided Design, Digest of Technical Papers (ICCAD\u201918) . L. Kosmidis, C. Maxim, V. Jegu, F. Vatrinet, and F. J. Cazorla. 2018. Industrial experiences with resource management under software randomization in ARINC653 avionics environments. In Proceedings of the IEEE\/ACM International Conference on Computer-aided Design, Digest of Technical Papers (ICCAD\u201918).","key":"e_1_2_1_14_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_15_1","DOI":"10.1109\/TPDS.2016.2549523"},{"key":"e_1_2_1_16_1","volume-title":"Self Driving Cars. Retrieved on","author":"NVIDIA Corporation","year":"2019","unstructured":"NVIDIA Corporation . 2019. Self Driving Cars. Retrieved on April 2019 from https:\/\/www.nvidia.com\/en-us\/self-driving-cars. NVIDIA Corporation. 2019. Self Driving Cars. Retrieved on April 2019 from https:\/\/www.nvidia.com\/en-us\/self-driving-cars."},{"key":"e_1_2_1_17_1","volume-title":"Proceedings of the 9th International Conference on Computational Intelligence and Communication Networks (CICN\u201917)","volume":"28","author":"Ozgunalp U.","year":"2018","unstructured":"U. Ozgunalp . 2018 . Combination of the symmetrical local threshold and the sobel edge detector for lane feature extraction . In Proceedings of the 9th International Conference on Computational Intelligence and Communication Networks (CICN\u201917) , Vol. 2018-January. 24-- 28 . U. Ozgunalp. 2018. Combination of the symmetrical local threshold and the sobel edge detector for lane feature extraction. In Proceedings of the 9th International Conference on Computational Intelligence and Communication Networks (CICN\u201917), Vol. 2018-January. 24--28."},{"doi-asserted-by":"publisher","key":"e_1_2_1_18_1","DOI":"10.1007\/978-981-13-1501-5_1"},{"volume-title":"Proceedings of the 18th Annual Network and Distributed System Security Symposium (NDSS\u201911)","author":"Slowinska A.","unstructured":"A. Slowinska , T. Stancescu , and H. Bos . 2011. Howard: A dynamic excavator for reverse engineering data structures . In Proceedings of the 18th Annual Network and Distributed System Security Symposium (NDSS\u201911) . A. Slowinska, T. Stancescu, and H. Bos. 2011. Howard: A dynamic excavator for reverse engineering data structures. In Proceedings of the 18th Annual Network and Distributed System Security Symposium (NDSS\u201911).","key":"e_1_2_1_19_1"},{"volume-title":"Proceedings of the Innovative Parallel Computing Conference (InPar\u201912)","author":"Steinberger M.","unstructured":"M. Steinberger , M. Kenzel , B. Kainz , and D. Schmalstieg . 2012. ScatterAlloc: Massively parallel dynamic memory allocation for the GPU . In Proceedings of the Innovative Parallel Computing Conference (InPar\u201912) . M. Steinberger, M. Kenzel, B. Kainz, and D. Schmalstieg. 2012. ScatterAlloc: Massively parallel dynamic memory allocation for the GPU. In Proceedings of the Innovative Parallel Computing Conference (InPar\u201912).","key":"e_1_2_1_20_1"},{"key":"e_1_2_1_21_1","volume-title":"Proceedings of the IEEE\/ACM International Conference on Computer-Aided Design, Digest of Technical Papers (ICCAD\u201917)","volume":"312","author":"Trompouki M. M.","unstructured":"M. M. Trompouki , L. Kosmidis , and N. Navarro . 2017. An open benchmark implementation for multi-CPU multi-GPU pedestrian detection in automotive systems . In Proceedings of the IEEE\/ACM International Conference on Computer-Aided Design, Digest of Technical Papers (ICCAD\u201917) , Vol. 2017-November. 305-- 312 . M. M. Trompouki, L. Kosmidis, and N. Navarro. 2017. An open benchmark implementation for multi-CPU multi-GPU pedestrian detection in automotive systems. In Proceedings of the IEEE\/ACM International Conference on Computer-Aided Design, Digest of Technical Papers (ICCAD\u201917), Vol. 2017-November. 305--312."},{"key":"e_1_2_1_22_1","volume-title":"Proceedings of the ASME Dynamic Systems and Control Conference (DSCC\u201917)","volume":"1","author":"Vishwanathan H.","unstructured":"H. Vishwanathan , D. L. Peters , and J. Z. Zhang . 2017. Traffic sign recognition in autonomous vehicles using edge detection . In Proceedings of the ASME Dynamic Systems and Control Conference (DSCC\u201917) , Vol. 1 . H. Vishwanathan, D. L. Peters, and J. Z. Zhang. 2017. Traffic sign recognition in autonomous vehicles using edge detection. In Proceedings of the ASME Dynamic Systems and Control Conference (DSCC\u201917), Vol. 1."},{"volume-title":"Proceedings of the 6th ACM Workshop on General Purpose Processor Using Graphics Processing Units. 120--126","author":"Widmer S.","unstructured":"S. Widmer , D. Wodniok , N. Weber , and M. Goesele . 2013. Fast dynamic memory allocator for massively parallel architectures . In Proceedings of the 6th ACM Workshop on General Purpose Processor Using Graphics Processing Units. 120--126 . S. Widmer, D. Wodniok, N. Weber, and M. Goesele. 2013. Fast dynamic memory allocator for massively parallel architectures. In Proceedings of the 6th ACM Workshop on General Purpose Processor Using Graphics Processing Units. 120--126.","key":"e_1_2_1_23_1"},{"doi-asserted-by":"crossref","unstructured":"P. R. Wilson M. S. Johnstone M. Neely and D. Boles. 1995. Dynamic Storage Allocation: A Survey and Critical Review. Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) Vol. 986. 1\u2013116.  P. R. Wilson M. S. Johnstone M. Neely and D. Boles. 1995. Dynamic Storage Allocation: A Survey and Critical Review. Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) Vol. 986. 1\u2013116.","key":"e_1_2_1_24_1","DOI":"10.1007\/3-540-60368-9_19"},{"volume-title":"Proceedings of the IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS\u201910)","author":"Wong H.","unstructured":"H. Wong , M. Papadopoulou , M. Sadooghi-Alvandi , and A. Moshovos . 2010. Demystifying GPU microarchitecture through microbenchmarking . In Proceedings of the IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS\u201910) . 235--246. H. Wong, M. Papadopoulou, M. Sadooghi-Alvandi, and A. Moshovos. 2010. Demystifying GPU microarchitecture through microbenchmarking. In Proceedings of the IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS\u201910). 235--246.","key":"e_1_2_1_25_1"},{"key":"e_1_2_1_26_1","volume-title":"Leibniz International Proceedings in Informatics, LIPIcs","volume":"106","author":"Yang M.","unstructured":"M. Yang , N. Otterness , T. Amert , J. Bakita , J. H. Anderson , and F. D. Smith . 2018. Avoiding pitfalls when using NVIDIA GPUs for real-time tasks in autonomous systems . In Leibniz International Proceedings in Informatics, LIPIcs , Vol. 106 . M. Yang, N. Otterness, T. Amert, J. Bakita, J. H. Anderson, and F. D. Smith. 2018. Avoiding pitfalls when using NVIDIA GPUs for real-time tasks in autonomous systems. In Leibniz International Proceedings in Informatics, LIPIcs, Vol. 106."},{"volume-title":"Proceedings of the 9th IEEE-GCC Conference and Exhibition (GCCCE\u201917)","author":"Younis R.","unstructured":"R. Younis and N. Bastaki . 2018. Accelerated fog removal from real images for car detection . In Proceedings of the 9th IEEE-GCC Conference and Exhibition (GCCCE\u201917) . R. Younis and N. Bastaki. 2018. Accelerated fog removal from real images for car detection. In Proceedings of the 9th IEEE-GCC Conference and Exhibition (GCCCE\u201917).","key":"e_1_2_1_27_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_28_1","DOI":"10.1007\/s11265-018-1352-0"}],"container-title":["ACM Transactions on Embedded Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3391896","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3391896","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T22:41:42Z","timestamp":1750200102000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3391896"}},"subtitle":["Understanding and Exploiting the Internals of GPU Resource Allocation in Critical Systems"],"short-title":[],"issued":{"date-parts":[[2020,9,26]]},"references-count":28,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2020,9,30]]}},"alternative-id":["10.1145\/3391896"],"URL":"https:\/\/doi.org\/10.1145\/3391896","relation":{},"ISSN":["1539-9087","1558-3465"],"issn-type":[{"type":"print","value":"1539-9087"},{"type":"electronic","value":"1558-3465"}],"subject":[],"published":{"date-parts":[[2020,9,26]]},"assertion":[{"value":"2019-11-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-03-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-09-26","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}