{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T05:00:28Z","timestamp":1750309228635,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":44,"publisher":"ACM","license":[{"start":{"date-parts":[[2025,1,20]],"date-time":"2025-01-20T00:00:00Z","timestamp":1737331200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2025,1,20]]},"DOI":"10.1145\/3658617.3697546","type":"proceedings-article","created":{"date-parts":[[2025,3,4]],"date-time":"2025-03-04T14:23:57Z","timestamp":1741098237000},"page":"567-574","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Zipper: Latency-Tolerant Optimizations for High-Performance Buses"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9522-8934","authenticated-orcid":false,"given":"Shibo","family":"Chen","sequence":"first","affiliation":[{"name":"Computer Science and Engineering, Univ. of Michigan, Ann Arbor, Michigan, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-5810-1347","authenticated-orcid":false,"given":"Hailun","family":"Zhang","sequence":"additional","affiliation":[{"name":"Univ. of Wisconsin, Madison, Wisconsin, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0181-0852","authenticated-orcid":false,"given":"Todd","family":"Austin","sequence":"additional","affiliation":[{"name":"Computer Science and Engineering, Univ. of Michigan, Ann Arbor, Michigan, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,3,4]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"Software Prefetching for Indirect Memory Accesses. In 2017 IEEE\/ACM International Symposium on Code Generation and Optimization (CGO). IEEE, 305--317","author":"Ainsworth Sam","year":"2017","unstructured":"Sam Ainsworth and Timothy M Jones. 2017. Software Prefetching for Indirect Memory Accesses. In 2017 IEEE\/ACM International Symposium on Code Generation and Optimization (CGO). IEEE, 305--317."},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/HOTCHIPS.2009.7478337"},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.5555\/290940.290988"},{"key":"e_1_3_2_1_4_1","volume-title":"Effectively Prefetching Remote Memory with Leap. In 2020 USENIX Annual Technical Conference (USENIX ATC 20)","author":"Maruf Hasan Al","year":"2020","unstructured":"Hasan Al Maruf and Mosharaf Chowdhury. 2020. Effectively Prefetching Remote Memory with Leap. In 2020 USENIX Annual Technical Conference (USENIX ATC 20). 843--857."},{"key":"e_1_3_2_1_5_1","unstructured":"AMD. 2019. AMD Infinity Architecture: The Foundation of the Modern Datacenter. https:\/\/www.amd.com\/system\/files\/documents\/LE-70001-SB-InfinityArchitecture.pdf."},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1147\/rd.111.0008"},{"key":"e_1_3_2_1_7_1","unstructured":"ARM. 2021. AMBA AXI and ACE Protocol Specification. Version H.c. https:\/\/developer.arm.com\/documentation\/ihi0022\/hc\/?lang=en."},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/3015146"},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/SEED55351.2022.00014"},{"key":"e_1_3_2_1_10_1","volume-title":"VIP-Bench: A Benchmark Suite for Evaluating Privacy-Enhanced Computation Frameworks. In 2021 International Symposium on Secure and Private Execution Environment Design (SEED). IEEE, 139--149","author":"Biernacki Lauren","year":"2021","unstructured":"Lauren Biernacki, Meron Zerihun Demissie, Kidus Birkayehu Workneh, Galane Basha Namomsa, Plato Gebremedhin, Fitsum Assamnew Andargie, Brandon Reagen, and Todd Austin. 2021. VIP-Bench: A Benchmark Suite for Evaluating Privacy-Enhanced Computation Frameworks. In 2021 International Symposium on Secure and Private Execution Environment Design (SEED). IEEE, 139--149."},{"volume-title":"Posit NPB: Assessing the Precision Improvement in HPC Scientific Applications","author":"Chien Steven W. D.","key":"e_1_3_2_1_11_1","unstructured":"Steven W. D. Chien, Ivy B. Peng, and Stefano Markidis. 2020. Posit NPB: Assessing the Precision Improvement in HPC Scientific Applications. In Parallel Processing and Applied Mathematics, Roman Wyrzykowski, Ewa Deelman, Jack Dongarra, and Konrad Karczewski (Eds.). Springer International Publishing, Cham, 301--310."},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"crossref","unstructured":"Jongsok Choi Kevin Nam Andrew Canis Jason Anderson Stephen Brown and Tomasz Czajkowski. 2012. Impact of Cache Architecture and Interface on Performance and Area of FPGA-based Processor\/Parallel-accelerator Systems. In 2012 IEEE 20th International Symposium on Field-Programmable Custom Computing Machines. IEEE 17--24.","DOI":"10.1109\/FCCM.2012.13"},{"key":"e_1_3_2_1_13_1","volume-title":"Shipping to Vendors.","author":"Cutress Ian","year":"2018","unstructured":"Ian Cutress. 2018. Intel Shows Xeon Scalable Gold 6138P with Integrated FPGA, Shipping to Vendors. (2018). https:\/\/www.anandtech.com\/show\/12773\/intel-shows-xeon-scalable-gold-6138p-with-integrated-fpga-shipping-to-vendors"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/40.621209"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/320831.320833"},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/2.375174"},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/800046.801647"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/MASSP.1984.1162257"},{"key":"e_1_3_2_1_19_1","volume-title":"In-memory Computing with Resistive Switching Devices. Nature electronics 1, 6","author":"Ielmini Daniele","year":"2018","unstructured":"Daniele Ielmini and H-S Philip Wong. 2018. In-memory Computing with Resistive Switching Devices. Nature electronics 1, 6 (2018), 333--343."},{"key":"e_1_3_2_1_20_1","unstructured":"Intel. 2019. Intel Acceleration Stack for Intel\u00ae Xeon\u00ae CPU with FPGAs Core Cache Interface (CCI-P) Reference Manual. https:\/\/www.intel.com\/content\/www\/us\/en\/docs\/programmable\/683193\/current\/acceleration-stack-for-cpu-with-fpgas.html."},{"key":"e_1_3_2_1_21_1","unstructured":"Intel. 2019. Intel\u00ae Xeon\u00ae Processor Scalable Family Technical Overview. https:\/\/www.intel.com\/content\/www\/us\/en\/developer\/articles\/technical\/xeon-processor-scalable-family-technical-overview.html."},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3492321.3519583"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/2485922.2485951"},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA45697.2020.00070"},{"volume-title":"Software Caching for Tree-based Algorithms on Accelerator Cards. Master's thesis","author":"Knoben PAH","key":"e_1_3_2_1_25_1","unstructured":"PAH Knoben. 2021. Software Caching for Tree-based Algorithms on Accelerator Cards. Master's thesis. University of Twente."},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/DAC18074.2021.9586197"},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3087556.3087582"},{"key":"e_1_3_2_1_28_1","unstructured":"Enno Luebbers Song Liu and Michael Chu. [n. d.]. Simplify Software Integration for FPGA Accelerators with OPAE. ([n. d.]). http:\/\/eulerproject.com\/assets\/files\/open-programmable-acceleration-engine-paper.pdf"},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/2907071"},{"key":"e_1_3_2_1_30_1","unstructured":"Patrick Patrick Kennedy. 2022. Compute Express Link CXL Latency-How Much is Added at HC34. https:\/\/www.servethehome.com\/compute-express-link-cxl-latency-how-much-is-added-at-hc34\/#:~:text=The%20CXL%20Consortium%20is%20using 170%2D250ns%20for%20CXL%20memory.&text=If%20CXL%20seems%20to%20be with%20Q2%202022%20Wind%2DDown.."},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/224056.224064"},{"key":"e_1_3_2_1_32_1","volume-title":"Hoard: A Distributed Data Caching System to Accelerate Deep Learning Training on the Cloud. arXiv preprint arXiv:1812.00669","author":"Pinto Christian","year":"2018","unstructured":"Christian Pinto, Yiannis Gkoufas, Andrea Reale, Seetharami Seelam, and Steven Eliuk. 2018. Hoard: A Distributed Data Caching System to Accelerate Deep Learning Training on the Cloud. arXiv preprint arXiv:1812.00669 (2018)."},{"key":"e_1_3_2_1_33_1","volume-title":"Wolfram HP Pernice, C David Wright, Abu Sebastian, and Harish Bhaskaran.","author":"R\u00edos Carlos","year":"2019","unstructured":"Carlos R\u00edos, Nathan Youngblood, Zengguang Cheng, Manuel Le Gallo, Wolfram HP Pernice, C David Wright, Abu Sebastian, and Harish Bhaskaran. 2019. In-memory Computing on a Photonic Platform. Science Advances 5, 2 (2019), eaau5759."},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2001.903250"},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2018.2876312"},{"key":"e_1_3_2_1_36_1","volume-title":"Toward Cache-Friendly Hardware Accelerators. In HPCA Sensors and Cloud Architectures Workshop (SCAW). 1--6.","author":"Shao Yakun Sophia","year":"2015","unstructured":"Yakun Sophia Shao, Sam Xi, Viji Srinivasan, Gu-Yeon Wei, and David Brooks. 2015. Toward Cache-Friendly Hardware Accelerators. In HPCA Sensors and Cloud Architectures Workshop (SCAW). 1--6."},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/DSD.2018.00106"},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/C-M.1978.218016"},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/356887.356892"},{"key":"e_1_3_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2005.10"},{"key":"e_1_3_2_1_41_1","volume-title":"Needham","author":"Wheeler David J.","year":"1995","unstructured":"David J. Wheeler and Roger M. Needham. 1995. TEA, a Tiny Encryption Algorithm. In Fast Software Encryption, Bart Preneel (Ed.). Springer Berlin Heidelberg, Berlin, Heidelberg, 363--366."},{"key":"e_1_3_2_1_42_1","doi-asserted-by":"crossref","unstructured":"R. Wienands and W. Joppich. 2005. Practical Fourier Analysis for Multigrid Methods. Taylor & Francis. https:\/\/books.google.com\/books?id=IOSux5GxacsC","DOI":"10.1201\/9781420034998"},{"key":"e_1_3_2_1_43_1","unstructured":"Xilinx. [n. d.]. Xilinx Runtime Library (XRT). ([n. d.]). https:\/\/xilinx.github.io\/XRT\/master\/html\/index.html"},{"key":"e_1_3_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/IEDM.2015.7409607"}],"event":{"name":"ASPDAC '25: 30th Asia and South Pacific Design Automation Conference","sponsor":["SIGDA ACM Special Interest Group on Design Automation","IEICE","IPSJ","IEEE CAS","IEEE CEDA"],"location":"Tokyo Japan","acronym":"ASPDAC '25"},"container-title":["Proceedings of the 30th Asia and South Pacific Design Automation Conference"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3658617.3697546","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3658617.3697546","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T23:44:18Z","timestamp":1750290258000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3658617.3697546"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,1,20]]},"references-count":44,"alternative-id":["10.1145\/3658617.3697546","10.1145\/3658617"],"URL":"https:\/\/doi.org\/10.1145\/3658617.3697546","relation":{},"subject":[],"published":{"date-parts":[[2025,1,20]]},"assertion":[{"value":"2025-03-04","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}