{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,17]],"date-time":"2026-02-17T12:12:25Z","timestamp":1771330345559,"version":"3.50.1"},"reference-count":52,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2024,4,30]],"date-time":"2024-04-30T00:00:00Z","timestamp":1714435200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Reconfigurable Technol. Syst."],"published-print":{"date-parts":[[2024,6,30]]},"abstract":"<jats:p>\n            The introduction of High Bandwidth Memory (HBM) to the FPGA chip makes it possible for an FPGA-based accelerator to leverage the huge memory bandwidth of HBM to improve its performance when implementing a specific algorithm, which is especially true for the Breadth-First Search (BFS) algorithm that demands a high bandwidth for accessing the graph data stored in memory. Different from traditional FPGA-DRAM platforms where memory bandwidth is the precious resource due to the limited DRAM channels, FPGA chips equipped with HBM have much higher memory bandwidths provided by the large quantities of HBM channels, but still a limited amount of logic (LUT, FF, and BRAM\/URAM) resources. Therefore, the key to design a high-performance BFS accelerator on an HBM-enhanced FPGA chip is to efficiently use the logic resources to build as many as possible Processing Elements (PEs) and configure them flexibly to obtain as high as possible\n            <jats:italic>effective memory bandwidth<\/jats:italic>\n            that is useful to the algorithm from the HBM, rather than partially emphasizing the absolute memory bandwidth. To exploit as high as possible effective bandwidth from the HBM, ScalaBFS2 conducts BFS in graphs in a vertex-centric manner and proposes designs, including the independent module (HBM Reader) for memory accessing, multi-layer crossbar, and PEs that implement hybrid mode (i.e., capable of working in both push and pull modes) algorithm processing, to utilize the FPGA logic resources efficiently. Consequently, ScalaBFS2 is able to build up to 128 PEs on the XCU280 FPGA chip (produced with the 16 nm process and configured with two HBM2 stacks) of a Xilinx Alveo U280 board and achieves performance of 56.92 Giga Traversed Edges Per Second (GTEPS) by fully using its 32 HBM memory channels. Compared with the state-of-the-art graph processing system (i.e., ReGraph) built on top of the same board, ScalaBFS2 achieves 2.52x~4.40x performance speedups. Moreover, when compared with Gunrock running on an Nvidia A100 GPU that is produced with the 7 nm process and configured with five HBM2e stacks, ScalaBFS2 achieves 1.34x~2.40x speedups on absolute performance, and 7.35x~13.18x speedups on power efficiency.\n          <\/jats:p>","DOI":"10.1145\/3650037","type":"journal-article","created":{"date-parts":[[2024,2,29]],"date-time":"2024-02-29T12:28:24Z","timestamp":1709209704000},"page":"1-39","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":8,"title":["ScalaBFS2: A High-performance BFS Accelerator on an HBM-enhanced FPGA Chip"],"prefix":"10.1145","volume":"17","author":[{"ORCID":"https:\/\/orcid.org\/0009-0007-8431-0864","authenticated-orcid":false,"given":"Kexin","family":"Li","sequence":"first","affiliation":[{"name":"Huazhong University of Science and Technology, Wuhan, China and Zhejiang Lab, Hangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7580-0668","authenticated-orcid":false,"given":"Shaoxian","family":"Xu","sequence":"additional","affiliation":[{"name":"Huazhong University of Science and Technology, Wuhan, China and Zhejiang Lab, Hangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2139-6465","authenticated-orcid":false,"given":"Zhiyuan","family":"Shao","sequence":"additional","affiliation":[{"name":"Huazhong University of Science and Technology, Wuhan, China and Zhejiang Lab, Hangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3058-7581","authenticated-orcid":false,"given":"Ran","family":"Zheng","sequence":"additional","affiliation":[{"name":"Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6302-813X","authenticated-orcid":false,"given":"Xiaofei","family":"Liao","sequence":"additional","affiliation":[{"name":"Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3934-7605","authenticated-orcid":false,"given":"Hai","family":"Jin","sequence":"additional","affiliation":[{"name":"Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,4,30]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/77600.77615"},{"key":"e_1_3_2_3_2","doi-asserted-by":"crossref","first-page":"310","DOI":"10.1145\/3289602.3293901","volume-title":"Proceedings of the 2019 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA\u201919)","author":"Asiatici Mikhail","year":"2019","unstructured":"Mikhail Asiatici and Paolo Ienne. 2019. Stop crying over your cache miss rate: Handling efficiently thousands of outstanding misses in FPGAs. In Proceedings of the 2019 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA\u201919). 310\u2013319. 10.1145\/3289602.3293901"},{"key":"e_1_3_2_4_2","first-page":"609","volume-title":"Proceedings of the 48th ACM\/IEEE Annual International Symposium on Computer Architecture (ISCA\u201921)","author":"Asiatici Mikhail","year":"2021","unstructured":"Mikhail Asiatici and Paolo Ienne. 2021. Large-scale graph processing on FPGAs with caches for thousands of simultaneous misses. In Proceedings of the 48th ACM\/IEEE Annual International Symposium on Computer Architecture (ISCA\u201921). 609\u2013622. 10.1109\/ISCA52012.2021.00054"},{"key":"e_1_3_2_5_2","first-page":"228","volume-title":"Proceedings of the 2014 IEEE International Parallel & Distributed Processing Symposium Workshops (IPDPSW\u201914)","author":"Attia Osama G.","year":"2014","unstructured":"Osama G. Attia, Tyler Johnson, Kevin Townsend, Philip Jones, and Joseph Zambreno. 2014. CyGraph: A reconfigurable architecture for parallel breadth-first search. In Proceedings of the 2014 IEEE International Parallel & Distributed Processing Symposium Workshops (IPDPSW\u201914). 228\u2013235. 10.1109\/IPDPSW.2014.30"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1002\/widm.1178"},{"key":"e_1_3_2_7_2","first-page":"8","volume-title":"Proceedings of the 23rd IEEE International Conference on Application-specific Systems, Architectures and Processors (ASAP\u201912)","author":"Betkaoui Brahim","year":"2012","unstructured":"Brahim Betkaoui, Yu Wang, David B. Thomas, and Wayne Luk. 2012. A reconfigurable computing approach for efficient and scalable parallel graph exploration. In Proceedings of the 23rd IEEE International Conference on Application-specific Systems, Architectures and Processors (ASAP\u201912). 8\u201315. 10.1109\/ASAP.2012.30"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0169-7552(98)00110-X"},{"key":"e_1_3_2_9_2","doi-asserted-by":"crossref","first-page":"1342","DOI":"10.1109\/MICRO56248.2022.00092","volume-title":"Proceedings of the 2022 55th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201922)","author":"Chen Xinyu","year":"2022","unstructured":"Xinyu Chen, Yao Chen, Feng Cheng, Hongshi Tan, Bingsheng He, and Weng-Fai Wong. 2022. ReGraph: Scaling graph processing on HBM-enabled FPGAs with heterogeneous pipelines. In Proceedings of the 2022 55th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201922). 1342\u20131358. 10.1109\/MICRO56248.2022.00092"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/3517141"},{"key":"e_1_3_2_11_2","doi-asserted-by":"crossref","first-page":"116","DOI":"10.1145\/3431920.3439301","volume-title":"Proceedings of the 2021 ACM\/SIGDA International Symposium on Field Programmable Gate Arrays (FPGA\u201921)","author":"Choi Young-kyu","year":"2021","unstructured":"Young-kyu Choi, Yuze Chi, Weikang Qiao, Nikola Samardzic, and Jason Cong. 2021. HBM connect: High-performance HLS interconnect for FPGA HBM. In Proceedings of the 2021 ACM\/SIGDA International Symposium on Field Programmable Gate Arrays (FPGA\u201921). 116\u2013126. 10.1145\/3431920.3439301"},{"key":"e_1_3_2_12_2","first-page":"217","volume-title":"Proceedings of the 2017 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA\u201917)","author":"Dai Guohao","year":"2017","unstructured":"Guohao Dai, Tianhao Huang, Yuze Chi, Ningyi Xu, Yu Wang, and Huazhong Yang. 2017. ForeGraph: Exploring large-scale graph processing on multi-FPGA architecture. In Proceedings of the 2017 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA\u201917). 217\u2013226. 10.1145\/3020078.3021739"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/2049662.2049663"},{"key":"e_1_3_2_14_2","first-page":"143","volume-title":"Proceedings of the 14th IEEE Symposium on Field-Programmable Custom Computing Machines (FCCM\u201906)","author":"deLorimier Michael","year":"2006","unstructured":"Michael deLorimier, Nachiket Kapre, Nikil Mehta, Dominic Rizzo, Ian Eslick, Raphael Rubin, Tomas E. Uribe, Thomas F. Jr. Knight, and Andre DeHon. 2006. GraphStep: A system architecture for sparse-graph algorithms. In Proceedings of the 14th IEEE Symposium on Field-Programmable Custom Computing Machines (FCCM\u201906). 143\u2013151. 10.1109\/FCCM.2006.45"},{"key":"e_1_3_2_15_2","doi-asserted-by":"crossref","first-page":"251","DOI":"10.1145\/316188.316229","volume-title":"Proceedings of the Conference on Applications, Technologies, Architectures, and Protocols for Computer Communication","author":"Faloutsos Michalis","year":"1999","unstructured":"Michalis Faloutsos, Petros Faloutsos, and Christos Faloutsos. 1999. On power-law relationships of the internet topology. In Proceedings of the Conference on Applications, Technologies, Architectures, and Protocols for Computer Communication. 251\u2013262. 10.1145\/316188.316229"},{"key":"e_1_3_2_16_2","first-page":"1","volume-title":"Proceedings of the 56th Annual Design Automation Conference (DAC ;19)","author":"Finnerty Eric","year":"2019","unstructured":"Eric Finnerty, Zachary Sherer, Hang Liu, and Yan Luo. 2019. Dr. BFS: Data centric breadth-first search on FPGAs. In Proceedings of the 56th Annual Design Automation Conference (DAC ;19). 1\u20136. 10.1145\/3316781.3317802"},{"key":"e_1_3_2_17_2","first-page":"17","volume-title":"Proceedings of the 10th USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201912)","author":"Gonzalez Joseph E.","year":"2012","unstructured":"Joseph E. Gonzalez, Yucheng Low, Haijie Gu, Danny Bickson, and Carlos Guestrin. 2012. PowerGraph: Distributed graph-parallel computation on natural graphs. In Proceedings of the 10th USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201912). 17\u201330. https:\/\/www.usenix.org\/conference\/osdi12\/technical-sessions\/presentation\/gonzalez"},{"key":"e_1_3_2_18_2","first-page":"81","volume-title":"Proceedings of the 2021 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA\u201921)","author":"Guo Licheng","year":"2021","unstructured":"Licheng Guo, Yuze Chi, Jie Wang, Jason Lau, Weikang Qiao, Ecenur Ustun, Zhiru Zhang, and Jason Cong. 2021. AutoBridge: Coupling coarse-grained floorplanning and pipelining for high-frequency HLS design on multi-die FPGAs. In Proceedings of the 2021 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA\u201921). 81\u201392. 10.1145\/3431920.3439289"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/1941553.1941590"},{"key":"e_1_3_2_20_2","first-page":"1","volume-title":"Proceedings of the IEEE\/ACM International Conference On Computer Aided Design (ICCAD\u201921)","author":"Hu Yuwei","year":"2021","unstructured":"Yuwei Hu, Yixiao Du, Ecenur Ustun, and Zhiru Zhang. 2021. GraphLily: Accelerating graph linear algebra on HBM-equipped FPGAs. In Proceedings of the IEEE\/ACM International Conference On Computer Aided Design (ICCAD\u201921). 1\u20139. 10.1109\/ICCAD51958.2021.9643582"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2021.3075765"},{"key":"e_1_3_2_22_2","article-title":"High Bandwidth Memory (HBM) DRAM","year":"2021","unstructured":"JEDEC. 2021. High Bandwidth Memory (HBM) DRAM. https:\/\/www.jedec.org\/standards-documents\/docs\/jesd235a.","journal-title":"https:\/\/www.jedec.org\/standards-documents\/docs\/jesd235a"},{"key":"e_1_3_2_23_2","first-page":"1","volume-title":"Proceedings of the 2017 IEEE International Memory Workshop (IMW\u201917)","author":"Jun Hongshin","year":"2017","unstructured":"Hongshin Jun, Jinhee Cho, Kangseol Lee, Ho-Young Son, Kwiwook Kim, Hanho Jin, and Keith Kim. 2017. HBM (high bandwidth memory) DRAM technology and architecture. In Proceedings of the 2017 IEEE International Memory Workshop (IMW\u201917). 1\u20134. 10.1109\/IMW.2017.7939084"},{"key":"e_1_3_2_24_2","doi-asserted-by":"crossref","first-page":"239","DOI":"10.1145\/3174243.3174260","volume-title":"Proceedings of the 2018 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA\u201918)","author":"Khoram Soroosh","year":"2018","unstructured":"Soroosh Khoram, Jialiang Zhang, Maxwell Strange, and Jing Li. 2018. Accelerating graph analytics by co-optimizing storage and access on an FPGA-HMC Platform. In Proceedings of the 2018 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA\u201918). 239\u2013248. 10.1145\/3174243.3174260"},{"key":"e_1_3_2_25_2","article-title":"Xilinx Stacked Silicon Interconnect Technology Delivers Breakthrough FPGA Capacity, Bandwidth, and Power Efficiency","author":"Kirk Saban","year":"2012","unstructured":"Saban Kirk. 2012. Xilinx Stacked Silicon Interconnect Technology Delivers Breakthrough FPGA Capacity, Bandwidth, and Power Efficiency. https:\/\/docs.xilinx.com\/v\/u\/en-US\/wp380_Stacked_Silicon_Interconnect_Technology.","journal-title":"https:\/\/docs.xilinx.com\/v\/u\/en-US\/wp380_Stacked_Silicon_Interconnect_Technology"},{"issue":"5","key":"e_1_3_2_26_2","first-page":"313","article-title":"TorusBFS: A novel message-passing parallel breadth-first search architecture on FPGAs","volume":"5","author":"Lei Guoqing","year":"2015","unstructured":"Guoqing Lei, Rongchun Li, and Song Guo. 2015. TorusBFS: A novel message-passing parallel breadth-first search architecture on FPGAs. IRACST International Journal 5, 5 (2015), 313\u2013318.","journal-title":"IRACST International Journal"},{"key":"e_1_3_2_27_2","article-title":"ScalaBFS: A scalable BFS accelerator on HBM-enhanced FPGAs","author":"Li Kexin","year":"2021","unstructured":"Kexin Li, Chenhao Liu, Zhiyuan Shao, Zeke Wang, Minkang Wu, Jiajie Chen, Xiaofei Liao, and Hai Jin. 2021. ScalaBFS: A scalable BFS accelerator on HBM-enhanced FPGAs. arXiv preprint arXiv:2105.11754 (2021).","journal-title":"arXiv preprint arXiv:2105.11754"},{"key":"e_1_3_2_28_2","first-page":"315","volume-title":"Proceedings of the 2019 International Conference on Field-Programmable Technology (ICFPT\u201919)","author":"Liu Cheng","year":"2019","unstructured":"Cheng Liu, Xinyu Chen, Bingsheng He, Xiaofei Liao, Ying Wang, and Lei Zhang. 2019. OBFS: OpenCL based BFS optimizations on software programmable FPGAs. In Proceedings of the 2019 International Conference on Field-Programmable Technology (ICFPT\u201919). 315\u2013318.10.1109\/ICFPT47387.2019.00056"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/3431920.3439463"},{"key":"e_1_3_2_30_2","first-page":"1","volume-title":"Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC\u201915)","author":"Liu Hang","year":"2015","unstructured":"Hang Liu and H. Howie Huang. 2015. Enterprise: Breadth-first graph traversal on GPUs. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC\u201915). 1\u201312. 10.1145\/2807591.2807594"},{"key":"e_1_3_2_31_2","article-title":"Virtex UltraScale+ HBM FPGA: A Revolutionary Increase in Memory Performance","author":"Mike Wissolik","year":"2019","unstructured":"Wissolik Mike, Zacher Darren, Torza Anthony, and Brandon Day. 2019. Virtex UltraScale+ HBM FPGA: A Revolutionary Increase in Memory Performance. https:\/\/docs.xilinx.com\/v\/u\/en-US\/wp485-hbm","journal-title":"https:\/\/docs.xilinx.com\/v\/u\/en-US\/wp485-hbm"},{"key":"e_1_3_2_32_2","article-title":"Introducing the Graph 500","author":"Murphy Richard C.","year":"2010","unstructured":"Richard C. Murphy, Kyle B. Wheeler, Brian W. Barrett, and James A. Ang. 2010. Introducing the Graph 500. Cray User Group (CUG).","journal-title":"Cray User Group (CUG)"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1587\/elex.11.20130987"},{"key":"e_1_3_2_34_2","article-title":"Nvidia A100 Tensor Core GPU","year":"2020","unstructured":"Nvidia. 2020. Nvidia A100 Tensor Core GPU. https:\/\/www.nvidia.com\/content\/dam\/en-zz\/Solutions\/Data-Center\/a100\/pdf\/nvidia-a100-datasheet-nvidia-us-2188504-web.pdf.","journal-title":"https:\/\/www.nvidia.com\/content\/dam\/en-zz\/Solutions\/Data-Center\/a100\/pdf\/nvidia-a100-datasheet-nvidia-us-2188504-web.pdf"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.proeng.2012.09.545"},{"key":"e_1_3_2_36_2","doi-asserted-by":"crossref","first-page":"194","DOI":"10.1145\/3452296.3472889","volume-title":"Proceedings of the 2021 ACM SIGCOMM 2021 Conference","author":"Pan Tian","year":"2021","unstructured":"Tian Pan, Nianbing Yu, Chenhao Jia, Jianwen Pi, Liang Xu, Yisong Qiao, Zhiguo Li, Kun Liu, Jie Lu, Jianyuan Lu, Enge Song, Jiao Zhang, Tao Huang, and Shunmin Zhu. 2021. Sailfish: Accelerating cloud-scale multi-tenant multi-service gateways with programmable switches. In Proceedings of the 2021 ACM SIGCOMM 2021 Conference. 194\u2013206. 10.1145\/3452296.3472889"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11432-020-3219-6"},{"key":"e_1_3_2_38_2","first-page":"622","volume-title":"Proceedings of the 23th International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS\u201918)","author":"Sabet Amir Hossein Nodehi","year":"2018","unstructured":"Amir Hossein Nodehi Sabet, Junqiao Qiu, and Zhijia Zhao. 2018. TIGR: Transforming irregular graphs for GPU-friendly graph processing. In Proceedings of the 23th International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS\u201918). 622\u2013636. 10.1145\/3173162.3173180"},{"key":"e_1_3_2_39_2","doi-asserted-by":"crossref","first-page":"320","DOI":"10.1145\/3289602.3293900","volume-title":"Proceedings of the 2019 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA\u201919)","author":"Shao Zhiyuan","year":"2019","unstructured":"Zhiyuan Shao, Ruoshi Li, Diqing Hu, Xiaofei Liao, and Hai Jin. 2019. Improving performance of graph processing on FPGA-DRAM platform by two-level vertex caching. In Proceedings of the 2019 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA\u201919). 320\u2013329. 10.1145\/3289602.3293900"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/2442516.2442530"},{"key":"e_1_3_2_41_2","first-page":"1","volume-title":"Proceedings of the 25th International Conference on Field Programmable Logic and Applications (FPL\u201915)","author":"Umuroglu Yaman","year":"2015","unstructured":"Yaman Umuroglu, Donn Morrison, and Magnus Jahre. 2015. Hybrid breadth-first search on a single-chip FPGA-CPU heterogeneous platform. In Proceedings of the 25th International Conference on Field Programmable Logic and Applications (FPL\u201915). 1\u20138. 10.1109\/FPL.2015.7293939"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/FPT.2010.5681757"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/3108140"},{"key":"e_1_3_2_44_2","article-title":"Alveo U280 Data Center Accelerator Card","year":"2019","unstructured":"Xilinx. 2019. Alveo U280 Data Center Accelerator Card. https:\/\/www.xilinx.com\/products\/boards-and-kits\/alveo\/u280.html","journal-title":"https:\/\/www.xilinx.com\/products\/boards-and-kits\/alveo\/u280.html"},{"key":"e_1_3_2_45_2","article-title":"Block Memory Generator v8.4","year":"2021","unstructured":"Xilinx. 2021. Block Memory Generator v8.4. https:\/\/www.xilinx.com\/ support\/documentation\/ip_documentation\/blk_mem_gen\/v8_4\/pg058-blk-mem-gen.pdf","journal-title":"https:\/\/www.xilinx.com\/ support\/documentation\/ip_documentation\/blk_mem_gen\/v8_4\/pg058-blk-mem-gen.pdf"},{"key":"e_1_3_2_46_2","article-title":"Alveo U280 Data Center Accelerator Card User Guide","year":"2023","unstructured":"Xilinx. 2023. Alveo U280 Data Center Accelerator Card User Guide. https:\/\/docs.xilinx.com\/r\/en-US\/ug1314-alveo-u280-reconfig-accel","journal-title":"https:\/\/docs.xilinx.com\/r\/en-US\/ug1314-alveo-u280-reconfig-accel"},{"key":"e_1_3_2_47_2","article-title":"UltraFast Design Methodology Guide for FPGAs and SoCs","year":"2023","unstructured":"Xilinx. 2023. UltraFast Design Methodology Guide for FPGAs and SoCs. https:\/\/docs.xilinx.com\/r\/en-US\/ug949-vivado-design-methodology","journal-title":"https:\/\/docs.xilinx.com\/r\/en-US\/ug949-vivado-design-methodology"},{"key":"e_1_3_2_48_2","article-title":"Vitis High-level Synthesis User Guide (UG1399)","year":"2023","unstructured":"Xilinx. 2023. Vitis High-level Synthesis User Guide (UG1399). https:\/\/docs.xilinx.com\/r\/en-US\/ug1399-vitis-hls","journal-title":"https:\/\/docs.xilinx.com\/r\/en-US\/ug1399-vitis-hls"},{"key":"e_1_3_2_49_2","article-title":"Vitis Unified Software Platform Documentation: Application Acceleration Development (UG1393)","year":"2023","unstructured":"Xilinx. 2023. Vitis Unified Software Platform Documentation: Application Acceleration Development (UG1393). https:\/\/docs.xilinx.com\/r\/en-US\/ug1393-vitis-application-acceleration","journal-title":"https:\/\/docs.xilinx.com\/r\/en-US\/ug1393-vitis-application-acceleration"},{"key":"e_1_3_2_50_2","doi-asserted-by":"crossref","first-page":"207","DOI":"10.1145\/3020078.3021737","volume-title":"Proceedings of the 2017 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA\u201917)","author":"Zhang Jialiang","year":"2017","unstructured":"Jialiang Zhang, Soroosh Khoram, and Jing Li. 2017. Boosting the performance of FPGA-based graph processor using hybrid memory cube: A case for breadth first search. In Proceedings of the 2017 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA\u201917). 207\u2013216. 10.1145\/3020078.3021737"},{"key":"e_1_3_2_51_2","doi-asserted-by":"crossref","first-page":"229","DOI":"10.1145\/3174243.3174245","volume-title":"Proceedings of the 2018 ACM\/SIGDA International Symposium on Field-Programmable Gate Array (FPGA\u201918)","author":"Zhang Jialiang","year":"2018","unstructured":"Jialiang Zhang and Jing Li. 2018. Degree-aware hybrid graph traversal on FPGA-HMC platform. In Proceedings of the 2018 ACM\/SIGDA International Symposium on Field-Programmable Gate Array (FPGA\u201918). 229\u2013238. 10.1145\/3174243.3174245"},{"key":"e_1_3_2_52_2","first-page":"183","volume-title":"Proceedings of the 20th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (PPoPP\u201915)","author":"Zhang Kaiyuan","year":"2015","unstructured":"Kaiyuan Zhang, Rong Chen, and Haibo Chen. 2015. NUMA-aware graph-structured analytics. In Proceedings of the 20th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (PPoPP\u201915). 183\u2013193. 10.1145\/2688500.2688507"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2019.2910068"}],"container-title":["ACM Transactions on Reconfigurable Technology and Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3650037","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3650037","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:03:43Z","timestamp":1750291423000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3650037"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,4,30]]},"references-count":52,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2024,6,30]]}},"alternative-id":["10.1145\/3650037"],"URL":"https:\/\/doi.org\/10.1145\/3650037","relation":{},"ISSN":["1936-7406","1936-7414"],"issn-type":[{"value":"1936-7406","type":"print"},{"value":"1936-7414","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,4,30]]},"assertion":[{"value":"2023-07-31","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-02-15","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-04-30","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}