{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T12:14:55Z","timestamp":1784204095108,"version":"3.55.0"},"reference-count":38,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2022,12,22]],"date-time":"2022-12-22T00:00:00Z","timestamp":1671667200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Reconfigurable Technol. Syst."],"published-print":{"date-parts":[[2023,3,31]]},"abstract":"<jats:p>Graph convolutional networks (GCNs) have been introduced to effectively process non-Euclidean graph data. However, GCNs incur large amounts of irregularity in computation and memory access, which prevents efficient use of traditional neural network accelerators. Moreover, existing dedicated GCN accelerators demand high memory volumes and are difficult to implement onto resource limited edge devices.<\/jats:p>\n          <jats:p>In this work, we propose LW-GCN, a lightweight FPGA-based accelerator with a software-hardware co-designed process to tackle irregularity in computation and memory access in GCN inference. LW-GCN decomposes the main GCN operations into Sparse Matrix-Matrix Multiplication (SpMM) and Matrix-Matrix Multiplication (MM). We propose a novel compression format to balance workload across PEs and prevent data hazards. Moreover, we apply data quantization and workload tiling, and map both SpMM and MM of GCN inference onto a uniform architecture on resource limited hardware. Evaluation on GCN and GraphSAGE are performed on Xilinx Kintex-7 FPGA with three popular datasets. Compared to existing CPU, GPU, and state-of-the-art FPGA-based accelerator, LW-GCN reduces latency by up to 60\u00d7, 12\u00d7, and 1.7\u00d7 and increases power efficiency by up to 912\u00d7, 511\u00d7, and 3.87\u00d7, respectively. Furthermore, compared with NVIDIA\u2019s latest edge GPU Jetson Xavier NX, LW-GCN achieves speedup and energy savings of 32\u00d7 and 84\u00d7, respectively.<\/jats:p>","DOI":"10.1145\/3550075","type":"journal-article","created":{"date-parts":[[2022,8,4]],"date-time":"2022-08-04T12:00:59Z","timestamp":1659614459000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":30,"title":["LW-GCN: A Lightweight FPGA-based Graph Convolutional Network Accelerator"],"prefix":"10.1145","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0951-1811","authenticated-orcid":false,"given":"Zhuofu","family":"Tao","sequence":"first","affiliation":[{"name":"Electrical and Computer Engineering, University of California, Los Angeles, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9939-8922","authenticated-orcid":false,"given":"Chen","family":"Wu","sequence":"additional","affiliation":[{"name":"Electrical and Computer Engineering, University of California, Los Angeles, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1861-4479","authenticated-orcid":false,"given":"Yuan","family":"Liang","sequence":"additional","affiliation":[{"name":"Electrical and Computer Engineering, University of California, Los Angeles, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7288-1789","authenticated-orcid":false,"given":"Kun","family":"Wang","sequence":"additional","affiliation":[{"name":"Electrical and Computer Engineering, University of California, Los Angeles, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9698-0319","authenticated-orcid":false,"given":"Lei","family":"He","sequence":"additional","affiliation":[{"name":"Electrical and Computer Engineering, University of California, Los Angeles, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,12,22]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2021.3084599"},{"key":"e_1_3_1_3_2","article-title":"Explainability techniques for graph convolutional networks","author":"Baldassarre Federico","year":"2019","unstructured":"Federico Baldassarre and Hossein Azizpour. 2019. Explainability techniques for graph convolutional networks. arXiv:1905.13686. Retrieved from https:\/\/arxiv.org\/abs\/1905.13686.","journal-title":"arXiv:1905.13686"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.5194\/os-9-1-2013"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-06028-6_26"},{"key":"e_1_3_1_6_2","unstructured":"Ken Chapman. 2014. Multiplexer design techniques for datapath performance with minimized routing resources. Application Note. http:\/\/www.xilinx.com."},{"key":"e_1_3_1_7_2","volume-title":"Proceedings of the International Symposium on Robotics Research","author":"Chen Fanfei","year":"2019","unstructured":"Fanfei Chen, Jinkun Wang, Tixiao Shan, and Brendan Englot. 2019. Autonomous exploration under uncertainty via graph convolutional networks. In Proceedings of the International Symposium on Robotics Research. Springer, Cham, 676\u2013691."},{"key":"e_1_3_1_8_2","article-title":"Rubik: A hierarchical architecture for efficient graph learning","author":"Chen Xiaobing","year":"2020","unstructured":"Xiaobing Chen, Yuke Wang, Xinfeng Xie, Xing Hu, Abanti Basak, Ling Liang, Mingyu Yan, Lei Deng, Yufei Ding, Zidong Du, and Y. Chen. 2020. Rubik: A hierarchical architecture for efficient graph learning. arXiv:2009.12495.","journal-title":"arXiv:2009.12495"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/3292500.3330925"},{"key":"e_1_3_1_10_2","article-title":"Pubmed 200k rct: A dataset for sequential sentence classification in medical abstracts","author":"Dernoncourt Franck","year":"2017","unstructured":"Franck Dernoncourt and Ji Young Lee. 2017. Pubmed 200k rct: A dataset for sequential sentence classification in medical abstracts. arXiv:1710.06071.","journal-title":"arXiv:1710.06071"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICTAI50040.2020.00198"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00097"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.polymertesting.2003.09.013"},{"key":"e_1_3_1_14_2","doi-asserted-by":"crossref","unstructured":"Tong Geng Ang Li Runbin Shi Chunshu Wu Tianqi Wang Yanfei Li Pouya Haghi Antonino Tumeo Shuai Che Steve Reinhardt and Martin Herbordt. 2019. AWB-GCN: A graph convolutional network accelerator with runtime workload rebalancing. In 2020 53rd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO) . IEEE 922-936.","DOI":"10.1109\/MICRO50266.2020.00079"},{"key":"e_1_3_1_15_2","article-title":"Inductive representation learning on large graphs","volume":"30","author":"Hamilton Will","year":"2017","unstructured":"Will Hamilton, Zhitao Ying, and Jure Leskovec. 2017. Inductive representation learning on large graphs. In Advances in Neural Information Processing Systems, Vol. 30.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3007787.3001163"},{"key":"e_1_3_1_17_2","article-title":"Semi-supervised classification with graph convolutional networks","author":"Kipf Thomas N.","year":"2016","unstructured":"Thomas N. Kipf and Max Welling. 2016. Semi-supervised classification with graph convolutional networks. arXiv:1609.02907.","journal-title":"arXiv:1609.02907"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2005.852293"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA51647.2021.00070"},{"key":"e_1_3_1_20_2","article-title":"EnGN: A high-throughput and energy-efficient accelerator for large graph neural networks","author":"Liang Shengwen","year":"2020","unstructured":"Shengwen Liang, Ying Wang, Cheng Liu, Lei He, LI Huawei, Dawen Xu, and Xiaowei Li. 2020. EnGN: A high-throughput and energy-efficient accelerator for large graph neural networks. IEEE Trans. Comput. 70, 9 (2020), 1511\u20131525.","journal-title":"IEEE Trans. Comput."},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3007787.3001179"},{"key":"e_1_3_1_22_2","article-title":"An introduction to convolutional neural networks","author":"O\u2019Shea Keiron","year":"2015","unstructured":"Keiron O\u2019Shea and Ryan Nash. 2015. An introduction to convolutional neural networks. arXiv:1511.08458. Retrieved from https:\/\/arxiv.org\/abs\/1511.08458.","journal-title":"arXiv:1511.08458"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNN.2008.2005605"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-93417-4_38"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2021.3052138"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1093\/bib\/bbz042"},{"key":"e_1_3_1_27_2","volume-title":"Proceedings of the International Conference on Machine Learning (ICML\u201911)","author":"Sutskever Ilya","year":"2011","unstructured":"Ilya Sutskever, James Martens, and Geoffrey E. Hinton. 2011. Generating text with recurrent neural networks. In Proceedings of the International Conference on Machine Learning (ICML\u201911)."},{"key":"e_1_3_1_28_2","unstructured":"Shyam A. Tailor Javier Fernandez-Marques and Nicholas D. Lane. 2020. Degree-quant: Quantization-aware training for graph neural networks. arXiv preprint arXiv:2008.05000."},{"key":"e_1_3_1_29_2","article-title":"Attention-based graph neural network for semi-supervised learning","author":"Thekumparampil Kiran K.","year":"2018","unstructured":"Kiran K. Thekumparampil, Chong Wang, Sewoong Oh, and Li-Jia Li. 2018. Attention-based graph neural network for semi-supervised learning. arXiv:1803.03735. Retrieved from https:\/\/arxiv.org\/abs\/1803.03735.","journal-title":"arXiv:1803.03735"},{"key":"e_1_3_1_30_2","unstructured":"Petar Veli\u010dkovi\u0107 Guillem Cucurull Arantxa Casanova Adriana Romero Pietro Lio and Yoshua Bengio. 2017. Graph attention networks. arXiv preprint arXiv:1710.10903."},{"key":"e_1_3_1_31_2","volume-title":"Advances in Neural Information Processing Systems","author":"Xie Cong","year":"2014","unstructured":"Cong Xie, Ling Yan, Wu-Jun Li, and Zhihua Zhang. 2014. Distributed power-law graph computing: Theoretical and empirical analysis. In Advances in Neural Information Processing Systems, Z. Ghahramani, M. Welling, C. Cortes, N. Lawrence, and K. Q. Weinberger (Eds.), Vol. 27."},{"key":"e_1_3_1_32_2","article-title":"LogiCORE IP distributed memory generator v8.0 product guide","year":"2015","unstructured":"Xilinx. 2015. LogiCORE IP distributed memory generator v8.0 product guide. Xilinx Product Guide.","journal-title":"Xilinx Product Guide"},{"key":"e_1_3_1_33_2","unstructured":"Keyulu Xu Weihua Hu Jure Leskovec and Stefanie Jegelka. 2018. How powerful are graph neural networks? arXiv preprint arXiv:1810.00826."},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA47549.2020.00012"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/3219819.3219890"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2019.2939726"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3373087.3375311"},{"key":"e_1_3_1_38_2","article-title":"Graphsaint: Graph sampling based inductive learning method","author":"Zeng Hanqing","year":"2019","unstructured":"Hanqing Zeng, Hongkuan Zhou, Ajitesh Srivastava, Rajgopal Kannan, and Viktor Prasanna. 2019. Graphsaint: Graph sampling based inductive learning method. arXiv:1907.04931.","journal-title":"arXiv:1907.04931"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2016.7783723"}],"container-title":["ACM Transactions on Reconfigurable Technology and Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3550075","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3550075","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T18:08:14Z","timestamp":1750183694000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3550075"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,12,22]]},"references-count":38,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2023,3,31]]}},"alternative-id":["10.1145\/3550075"],"URL":"https:\/\/doi.org\/10.1145\/3550075","relation":{},"ISSN":["1936-7406","1936-7414"],"issn-type":[{"value":"1936-7406","type":"print"},{"value":"1936-7414","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,12,22]]},"assertion":[{"value":"2021-11-02","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-07-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-12-22","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}