{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,21]],"date-time":"2025-09-21T07:14:36Z","timestamp":1758438876587,"version":"3.44.0"},"reference-count":69,"publisher":"Association for Computing Machinery (ACM)","issue":"3","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2025,9,30]]},"abstract":"<jats:p>Recent chip multiprocessors incorporate several on-chip accelerators, marking the beginning of the Accelerated Chip Multi-Processor (XMP) era in datacenters. Despite the close proximity of accelerators and general-purpose cores, offloading functions to accelerators may not always be beneficial. Offloading to hardware accelerators can introduce several end-to-end overheads that can negate the speedup of the accelerable function. In this article, we design RACER, a hardware architecture and runtime system that evades the danger of end-to-end slowdowns when using hardware acceleration. RACER leverages a low-overhead interface between general-purpose cores and on-chip accelerators, fine-grained context switching, accelerator-initiated preemption, and seamless data motion between general-purpose cores and accelerators to improve the performance of workloads that use on-chip accelerators. We evaluate RACER on five representative request processing workloads featuring diverse memory access patterns, accelerable functions, and compute intensities. RACER improves the performance of hardware acceleration on a real XMP by an average of 1.31\u00d7 on a range of diverse workloads and guarantees that accelerator offloads never cause slowdowns.<\/jats:p>","DOI":"10.1145\/3750448","type":"journal-article","created":{"date-parts":[[2025,7,24]],"date-time":"2025-07-24T11:23:34Z","timestamp":1753356214000},"page":"1-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["RACER: Avoiding End-to-End Slowdowns in Accelerated Chip Multi-Processors"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8055-4243","authenticated-orcid":false,"given":"Neel","family":"Patel","sequence":"first","affiliation":[{"name":"Electrical and Computer Engineering, Cornell University","place":["Ithaca, United States"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-2937-5804","authenticated-orcid":false,"given":"Ren","family":"Wang","sequence":"additional","affiliation":[{"name":"Intel Labs","place":["Portland, United States"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4622-2181","authenticated-orcid":false,"given":"Mohammad","family":"Alian","sequence":"additional","affiliation":[{"name":"Electrical and Computer Engineering, Cornell University","place":["Ithaca, United States"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,9,19]]},"reference":[{"key":"e_1_3_4_2_2","unstructured":"[n.d.]. 22. Intel(R) QuickAssist (QAT) Crypto Poll Mode Driver - Data Plane Development Kit 25.03.0 documentation. Retrieved May 11th 2025 from https:\/\/doc.dpdk.org\/guides\/cryptodevs\/qat.html"},{"key":"e_1_3_4_3_2","unstructured":"[n.d.]. 4th Gen Intel Xeon Processor Scalable Family sapphire rapids. Retrieved May 11th 2025 from https:\/\/www.intel.com\/content\/www\/us\/en\/developer\/articles\/technical\/fourth-generation-xeon-scalable-family-overview.html#gs.3wg3ec"},{"key":"e_1_3_4_4_2","unstructured":"[n.d.]. 5. IDXD DMA Device Driver - Data Plane Development Kit 25.03.0 documentation. Retrieved May 11th 2025 from https:\/\/doc.dpdk.org\/guides\/dmadevs\/idxd.html"},{"key":"e_1_3_4_5_2","unstructured":"[n.d.]. Apache Flink\u00ae - Stateful Computations over Data Streams. Retrieved May 11th 2025 from https:\/\/flink.apache.org\/"},{"key":"e_1_3_4_6_2","unstructured":"[n.d.]. Apache Pulsar | Apache Pulsar. Retrieved May 11th 2025 from https:\/\/pulsar.apache.org\/"},{"key":"e_1_3_4_7_2","unstructured":"[n.d.]. ARM A-profile PRFM Instruction. Retrieved May 11th 2025 from https:\/\/developer.arm.com\/documentation\/ddi0602\/2022-06\/Base-Instructions\/PRFM--immediate---Prefetch-Memory--immediate--"},{"key":"e_1_3_4_8_2","unstructured":"[n.d.]. Build Clickhouse with DEFLATE_QPL | ClickHouse Docs. Retrieved May 11th 2025 from https:\/\/clickhouse.com\/docs\/en\/development\/building_and_benchmarking_deflate_qpl"},{"key":"e_1_3_4_9_2","unstructured":"[n.d.]. CLDEMOTE - Cache Line Demote. Retrieved May 11th 2025 from https:\/\/www.felixcloutier.com\/x86\/cldemote"},{"key":"e_1_3_4_10_2","unstructured":"[n.d.]. CLFLUSH - Flush Cache Line. Retrieved May 11th 2025 from https:\/\/www.felixcloutier.com\/x86\/clflush"},{"key":"e_1_3_4_11_2","unstructured":"[n.d.]. CoreLink Level 2 Cache Controller L2C-310 Technical Reference Manual. Retrieved May 11th 2025 from https:\/\/developer.arm.com\/documentation\/ddi0246\/h\/programmers-model\/register-descriptions\/cache-maintenance-operations"},{"key":"e_1_3_4_12_2","unstructured":"[n.d.]. Intel\u00ae Data Streaming Accelerator (Intel\u00ae DSA). Retrieved May 11th 2025 from https:\/\/www.intel.com\/content\/www\/us\/en\/products\/docs\/accelerator-engines\/data-streaming-accelerator.html"},{"key":"e_1_3_4_13_2","unstructured":"[n.d.]. Intel\u2019s future \u201cCLDEMOTE\u201d instruction. Retrieved May 11th 2025 from https:\/\/sites.utexas.edu\/jdm4372\/2019\/02\/18\/intels-future-cldemote-instruction\/"},{"key":"e_1_3_4_14_2","unstructured":"[n.d.]. PREFETCHh - Prefetch Data Into Caches. Retrieved May 11th 2025 from https:\/\/www.felixcloutier.com\/x86\/prefetchh"},{"key":"e_1_3_4_15_2","unstructured":"[n.d.]. Protocol Buffers. Retrieved from https:\/\/protobuf.dev\/"},{"key":"e_1_3_4_16_2","unstructured":"[n.d.]. smhasher\/src\/MurmurHash3.cpp at master aappleby\/smhasher. Retrieved May 11th 2025 from https:\/\/github.com\/aappleby\/smhasher\/blob\/master\/src\/MurmurHash3.cpp"},{"key":"e_1_3_4_17_2","unstructured":"[n.d.]. UMWAIT - User Level Monitor Wait. Retrieved May 11th 2025 from https:\/\/www.felixcloutier.com\/x86\/umwait"},{"key":"e_1_3_4_18_2","unstructured":"2023. Intel Labs\u2019 Contributions to Latest Intel\u00ae Xeon\u00ae Scalable Processor. Retrieved from https:\/\/community.intel.com\/t5\/Blogs\/Tech-Innovation\/Data-Center\/Intel-Labs-Contributions-to-Latest-Intel-Xeon-Scalable-Processor\/post\/1441731. Section: Data Center."},{"key":"e_1_3_4_19_2","unstructured":"2024. AIX 7.2 dcbt (Data Cache Block Touch) instruction. Retrieved from https:\/\/www.ibm.com\/docs\/en\/aix\/7.2?topic=set-dcbt-data-cache-block-touch-instruction"},{"key":"e_1_3_4_20_2","unstructured":"2024. Arm A-Profile Architecture Developments 2024 - Architectures and Processors blog - Arm Community blogs - Arm Community. Retrieved from https:\/\/community.arm.com\/arm-community-blogs\/b\/architectures-and-processors-blog\/posts\/arm-a-profile-architecture-developments-2024"},{"key":"e_1_3_4_21_2","unstructured":"2024. boostorg\/context. Retrieved from https:\/\/github.com\/boostorg\/context. original-date: 2013-02-07T20:00:34Z."},{"key":"e_1_3_4_22_2","unstructured":"2024. DCBFL IBM Power9. Retrieved from https:\/\/www.ibm.com\/docs\/en\/xl-c-and-cpp-linux\/16.1.1?topic=functions-dcbfl"},{"key":"e_1_3_4_23_2","unstructured":"2024. intel\/DTO. Retrieved from https:\/\/github.com\/intel\/DTO. original-date: 2023-06-30T02:00:49Z."},{"key":"e_1_3_4_24_2","unstructured":"2024. intel\/idxd-config. Retrieved from https:\/\/github.com\/intel\/idxd-configoriginal-date: 2019-11-16T00:14:21Z."},{"key":"e_1_3_4_25_2","unstructured":"2024. intel\/ipp-crypto. Retrieved from https:\/\/github.com\/intel\/ipp-cryptooriginal-date: 2018-07-06T22:16:28Z."},{"key":"e_1_3_4_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO56248.2022.00042"},{"key":"e_1_3_4_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/3530819"},{"key":"e_1_3_4_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2018.00031"},{"key":"e_1_3_4_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/3140659.3080216"},{"key":"e_1_3_4_30_2","doi-asserted-by":"crossref","unstructured":"C\u00e9dric Augonnet Samuel Thibault Raymond Namyst and Pierre-Andr\u00e9 Wacrenier. 2011. StarPU: A unified platform for task scheduling on heterogeneous multicore architectures. CCPE-Concurrency and Computation: Practice and Experience special issue: Euro-par 2009 23 2 (Feb. 2011) 187\u2013198. Retrieved from http:\/\/hal.inria.fr\/inria-00550877 (2011).","DOI":"10.1002\/cpe.1631"},{"key":"e_1_3_4_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3676641.3716028"},{"key":"e_1_3_4_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2012.71"},{"key":"e_1_3_4_33_2","doi-asserted-by":"publisher","DOI":"10.1145\/3370748.3406564"},{"key":"e_1_3_4_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/LCA.2019.2910521"},{"key":"e_1_3_4_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/MCSE.2013.98"},{"key":"e_1_3_4_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/PACT.2011.21"},{"key":"e_1_3_4_37_2","volume-title":"Advances in Databases and Information Systems","author":"Elmasri R.","year":"2015","unstructured":"R. Elmasri, Shamkant B. Navathe, R. Elmasri, and S. B. Navathe. 2015. Fundamentals of database systems. In Advances in Databases and Information Systems, Vol. 139. Springer."},{"key":"e_1_3_4_38_2","doi-asserted-by":"publisher","DOI":"10.5555\/1012889.1012894"},{"key":"e_1_3_4_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2018.2877288"},{"key":"e_1_3_4_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/NOCS.2018.8512153"},{"key":"e_1_3_4_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3287624.3288755"},{"key":"e_1_3_4_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2005.23"},{"volume-title":"Intel\u00ae Architecture Instruction Set Extensions and Future Features (319433-044 ed.)","year":"2021","key":"e_1_3_4_43_2","unstructured":"Intel. 2021. Intel\u00ae Architecture Instruction Set Extensions and Future Features (319433-044 ed.). Retrieved from https:\/\/www.intel.com\/content\/dam\/develop\/external\/us\/en\/documents\/architecture-instruction-set-extensions-programming-reference.pdf"},{"key":"e_1_3_4_44_2","doi-asserted-by":"crossref","unstructured":"Iyer Rishabh Unal Musa Kogias Marios and Candea George. 2023. Achieving microsecond-scale tail latency efficiently with approximate optimal scheduling. In Proceedings of the 29th Symposium on Operating Systems Principles. ACM. Retrieved from https:\/\/dslab.epfl.ch\/pubs\/concord.pdf","DOI":"10.1145\/3600006.3613136"},{"key":"e_1_3_4_45_2","first-page":"345","volume-title":"Proceedings of the 16th USENIX Conference on Networked Systems Design and Implementation (NSDI\u201919)","author":"Kaffes Kostis","year":"2019","unstructured":"Kostis Kaffes, Timothy Chong, Jack Tigar Humphries, Adam Belay, David Mazi\u00e8res, and Christos Kozyrakis. 2019. Shinjuku: Preemptive scheduling for \u00b5second-scale tail latency. In Proceedings of the 16th USENIX Conference on Networked Systems Design and Implementation (NSDI\u201919). USENIX Association, USA, 345\u2013359."},{"key":"e_1_3_4_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2007.22"},{"key":"e_1_3_4_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/2749469.2750421"},{"key":"e_1_3_4_48_2","doi-asserted-by":"crossref","unstructured":"Reese Kuper Ipoom Jeong Yifan Yuan Ren Wang Narayan Ranganathan Nikhil Rao Jiayu Hu Sanjay Kumar Philip Lantz and Nam Sung Kim. 2024. A quantitative analysis and guidelines of data streaming accelerator in modern intel xeon scalable processors. In Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS\u201924). Association for Computing Machinery New York NY USA 37\u201354.","DOI":"10.1145\/3620665.3640401"},{"key":"e_1_3_4_49_2","doi-asserted-by":"publisher","unstructured":"Y. Li et\u00a0al. 2024. LibPreemptible: Enabling fast adaptive and hardware-assisted user-space Scheduling. IEEE International Symposium on High-Performance Computer Architecture (HPCA). 922\u2013936. DOI:10.1109\/HPCA57654.2024.00075","DOI":"10.1109\/HPCA57654.2024.00075"},{"key":"e_1_3_4_50_2","doi-asserted-by":"publisher","DOI":"10.1145\/3470496.3533042"},{"key":"e_1_3_4_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/3620665.3640381"},{"key":"e_1_3_4_52_2","unstructured":"Mark Rabkin Austin Appleby. [n.d.]. FurcHash. Retrieved from https:\/\/github.com\/facebook\/mcrouter\/blob\/main\/mcrouter\/lib\/fbi\/hash.c#L151"},{"key":"e_1_3_4_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2001.903254"},{"key":"e_1_3_4_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISSCC42614.2022.9731107"},{"key":"e_1_3_4_55_2","doi-asserted-by":"crossref","unstructured":"P. Deutsch. 1996. DEFLATE Compressed Data Format Specification version 1.3. Retrieved from https:\/\/www.ietf.org\/rfc\/rfc1951.txt","DOI":"10.17487\/rfc1951"},{"key":"e_1_3_4_56_2","first-page":"1359","volume-title":"Proceedings of the 2025 USENIX Annual Technical Conference (USENIX ATC 25)","author":"Patel Neel","year":"2025","unstructured":"Neel Patel and Mohammad Alian. 2025. XRT: An accelerator-aware runtime for accelerated chip multiprocessors. In Proceedings of the 2025 USENIX Annual Technical Conference (USENIX ATC 25). USENIX Association, Boston, MA, 1359\u20131369. Retrieved from https:\/\/www.usenix.org\/conference\/atc25\/presentation\/patel"},{"key":"e_1_3_4_57_2","first-page":"145","volume-title":"Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18)","author":"Qin Henry","year":"2018","unstructured":"Henry Qin, Qian Li, Jacqueline Speiser, Peter Kraft, and John Ousterhout. 2018. Arachne:{Core-aware} thread management. In Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18). 145\u2013160."},{"key":"e_1_3_4_58_2","unstructured":"Ravindran Binuraj and Giacchino Luca and Jee Kalyan and Babu Nandita Narendra. [n.d.]. Intel\u00ae in-memory analytics accelerator plugin for RocksDB storage engine (Intel\u00ae IAA plugin for RocksDB storage engine). ([n.d.]). Retrieved from https:\/\/cdrdv2-public.intel.com\/788783\/357191EN-Intel-IAA-Plugin-RocksDB.pdf"},{"key":"e_1_3_4_59_2","unstructured":"Ravindra P. Saraf Rahul Pal and Ashok Jagannathan. 2014. Sub-numa clustering. Retrieved from https:\/\/patents.google.com\/patent\/US8862828B2\/en"},{"key":"e_1_3_4_60_2","doi-asserted-by":"publisher","DOI":"10.5555\/3195638.3195697"},{"key":"e_1_3_4_61_2","first-page":"127","volume-title":"Proceedings of the 29th International Conference on Computers and Their Applications, CATA 2014","author":"Shelor Charles","unstructured":"Charles Shelor, Jim Buchanan, Krishna Kavi, and Ron Cytron. 2014. Quantifying wasted write energy in the memory hierarchy. In Proceedings of the 29th International Conference on Computers and Their Applications, CATA 2014. 127\u2013134."},{"key":"e_1_3_4_62_2","unstructured":"Jonathon Shlens. 2005. A tutorial on principal component analysis. (Dec.2005). Retrieved from https:\/\/www.cs.cmu.edu\/elaw\/papers\/pca.pdf"},{"key":"e_1_3_4_63_2","doi-asserted-by":"publisher","DOI":"10.1145\/2442516.2442530"},{"key":"e_1_3_4_64_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISSCC49657.2024.10454441"},{"key":"e_1_3_4_65_2","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378450"},{"key":"e_1_3_4_66_2","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC.2018.8573515"},{"key":"e_1_3_4_67_2","doi-asserted-by":"publisher","DOI":"10.1145\/3656019.3676896"},{"key":"e_1_3_4_68_2","first-page":"9","volume-title":"Proceedings of the 2nd International Workshop on MapReduce and Its Applications","author":"Talbot Justin","year":"2011","unstructured":"Justin Talbot, Richard M. Yoo, and Christos Kozyrakis. 2011. Phoenix++ modular mapreduce for shared-memory systems. In Proceedings of the 2nd International Workshop on MapReduce and Its Applications. 9\u201316."},{"key":"e_1_3_4_69_2","unstructured":"Will Brian and Shemer Karen. 2023. Intel\u00ae QuickAssist technology (intel\u00ae QAT) - NGINX* performance. (Feb.2023). Retrieved from https:\/\/cdrdv2-public.intel.com\/767645\/Intel_QuickAssist_Technology_NGINX%20Performance_Whitepaper_767645v1.pdf"},{"key":"e_1_3_4_70_2","doi-asserted-by":"publisher","DOI":"10.1145\/3466752.3480065"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3750448","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,20]],"date-time":"2025-09-20T00:49:17Z","timestamp":1758329357000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3750448"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,19]]},"references-count":69,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2025,9,30]]}},"alternative-id":["10.1145\/3750448"],"URL":"https:\/\/doi.org\/10.1145\/3750448","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"type":"print","value":"1544-3566"},{"type":"electronic","value":"1544-3973"}],"subject":[],"published":{"date-parts":[[2025,9,19]]},"assertion":[{"value":"2025-02-03","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-07-12","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-09-19","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}