{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,14]],"date-time":"2026-05-14T23:21:11Z","timestamp":1778800871913,"version":"3.51.4"},"reference-count":67,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2026,5,13]],"date-time":"2026-05-13T00:00:00Z","timestamp":1778630400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc-sa\/4.0\/legalcode"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Embed. Comput. Syst."],"published-print":{"date-parts":[[2026,5,31]]},"abstract":"<jats:p>\n                    Semantic understanding of 3D scenes is fundamental to many applications like robotics, autonomous driving, AR\/VR. State-of-the-art methods for different 3D scene understanding tasks use 3D convolutional neural networks (CNNs) operating on point clouds. Convolution on spatially sparse data like point cloud involves irregular data accesses and compute patterns leading to poor utilization and energy efficiency in CPU\/GPU implementations. The existing CNN accelerators designed for weight\/activation sparsity cannot be efficiently repurposed for 3D spatially sparse CNNs given the fundamental differences in locating non-zero operands and granularity of work-dispatches. To address the dataflow challenges due to spatial sparsity and the need for specialized microarchitecture for spatially sparse convolution we present\n                    <jats:sc>Ace-of-Spade<\/jats:sc>\n                    s (AoS), an algorithm-dataflow-architecture co-designed system. AoS enables the data reuse among spatially proximate points using a locality-aware metadata structure along with a surface orientation aware point cloud reordering algorithm. AoS uses a novel technique for spatial sparsity aware selection of optimal data tiles by modelling the sparsity induced variations in the point cloud with a near-zero latency overheads. To accelerate computation on spatially sparse data, we propose a novel hardware accelerator\n                    <jats:sc>Ss<\/jats:sc>\n                    p\n                    <jats:sc>nna<\/jats:sc>\n                    with a front-end to convert varying number of operations per point into a stream of dense work dispatches to the backend compute engine. The compute engine further exploits weight and input feature data reuse through dynamic systolic grouping and multicast interconnects. The\n                    <jats:sc>Ss<\/jats:sc>\n                    p\n                    <jats:sc>nna<\/jats:sc>\n                    core together with the 64 KB of L1 memory requires\n                    <jats:bold>\n                      0.31 mm\n                      <jats:sup>2<\/jats:sup>\n                    <\/jats:bold>\n                    of area in 10nm process at 1 GHz. Overall, AoS achieves speedup\/energy savings of\n                    <jats:bold>19.9x<\/jats:bold>\n                    \/\n                    <jats:bold>49.9x<\/jats:bold>\n                    and\n                    <jats:bold>2.2x<\/jats:bold>\n                    \/\n                    <jats:bold>7.1x<\/jats:bold>\n                    over the state-of-the-art CPU and GPU implementations respectively.\n                  <\/jats:p>","DOI":"10.1145\/3759457","type":"journal-article","created":{"date-parts":[[2025,8,19]],"date-time":"2025-08-19T11:28:08Z","timestamp":1755602888000},"page":"1-28","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["<scp>Ace-of-Spade<\/scp>\n                    s: Accelerating Spatially Sparse Convolution for 3D Scene Understanding"],"prefix":"10.1145","volume":"25","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9149-5605","authenticated-orcid":false,"given":"Om Ji","family":"Omer","sequence":"first","affiliation":[{"name":"Processor Architecture Research Lab, Intel Labs","place":["Bangalore, India"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-2305-6164","authenticated-orcid":false,"given":"Prashant","family":"Laddha","sequence":"additional","affiliation":[{"name":"Processor Architecture Research Lab, Intel Labs","place":["Bangalore, India"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-7018-7585","authenticated-orcid":false,"given":"Gurpreet","family":"Singh Kalsi","sequence":"additional","affiliation":[{"name":"Processor Architecture Research Lab, Intel Labs","place":["Bangalore, India"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-9206-8618","authenticated-orcid":false,"given":"Kamlesh","family":"Pillai","sequence":"additional","affiliation":[{"name":"Processor Architecture Research Lab, Intel Labs","place":["Bangalore, India"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-4892-0203","authenticated-orcid":false,"given":"Anirudh","family":"Thyagharajan","sequence":"additional","affiliation":[{"name":"Processor Architecture Research Lab, Intel Labs","place":["Bangalore, India"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-2093-6201","authenticated-orcid":false,"given":"Ahimanyu","family":"Kulkarni","sequence":"additional","affiliation":[{"name":"Processor Architecture Research Lab, Intel Labs","place":["Bangalore, India"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3878-8679","authenticated-orcid":false,"given":"Anbang","family":"Yao","sequence":"additional","affiliation":[{"name":"AI Research Lab, Intel Labs","place":["Minhang, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9333-1746","authenticated-orcid":false,"given":"Yurong","family":"Chen","sequence":"additional","affiliation":[{"name":"AI Research Lab, Intel Labs","place":["Minhang, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5372-0173","authenticated-orcid":false,"given":"Sreenivas","family":"Subramoney","sequence":"additional","affiliation":[{"name":"Processor Architecture Research Lab, Intel Labs","place":["Bangalore, India"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,5,13]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCD50377.2020.00036"},{"key":"e_1_3_1_3_2","first-page":"234","volume-title":"Proceedings of the International Conference on the Sciences of Electronics, Technologies of Information and Telecommunications","author":"Ayachi Riadh","year":"2018","unstructured":"Riadh Ayachi, Mouna Afif, Yahia Said, and Mohamed Atri. 2018. Strided convolution instead of max pooling for memory efficiency of convolutional neural networks. In Proceedings of the International Conference on the Sciences of Electronics, Technologies of Information and Telecommunications. Springer, 234\u2013243."},{"key":"e_1_3_1_4_2","first-page":"9297","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Behley Jens","year":"2019","unstructured":"Jens Behley, Martin Garbade, Andres Milioto, Jan Quenzel, Sven Behnke, Cyrill Stachniss, and Jurgen Gall. 2019. SemanticKITTI: A dataset for semantic scene understanding of LiDAR sequences. In Proceedings of the IEEE International Conference on Computer Vision. 9297\u20139307."},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/SFCS.1985.48"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/3007787.3001177"},{"key":"e_1_3_1_7_2","doi-asserted-by":"crossref","unstructured":"Yu-Hsin Chen Tien-Ju Yang Joel Emer and Vivienne Sze. 2019. Eyeriss v2: A flexible accelerator for emerging deep neural networks on mobile devices. IEEE Journal on Emerging and Selected Topics in Circuits and Systems 9 2 (2019) 292\u2013308.","DOI":"10.1109\/JETCAS.2019.2910232"},{"key":"e_1_3_1_8_2","unstructured":"Christopher Choy JunYoung Gwak and Silvio Savarese. 2019. 4d spatio-temporal convnets: Minkowski convolutional neural networks. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 3075\u20133084."},{"key":"e_1_3_1_9_2","first-page":"424","volume-title":"Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention","author":"\u00c7i\u00e7ek \u00d6zg\u00fcn","year":"2016","unstructured":"\u00d6zg\u00fcn \u00c7i\u00e7ek, Ahmed Abdulkadir, Soeren S. Lienkamp, Thomas Brox, and Olaf Ronneberger. 2016. 3D U-Net: Learning dense volumetric segmentation from sparse annotation. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, 424\u2013432."},{"key":"e_1_3_1_10_2","article-title":"Spconv: Spatially Sparse Convolution Library","author":"Contributors Spconv","year":"2022","unstructured":"Spconv Contributors. 2022. Spconv: Spatially Sparse Convolution Library. Retrieved Dec 15, 2024 from https:\/\/github.com\/traveller59\/spconv","journal-title":"R"},{"key":"e_1_3_1_11_2","unstructured":"Nvidia Corporation. 2026. Minkowski Engine. (n.d.). Retrieved Mar 23 2023 from https:\/\/github.com\/NVIDIA\/MinkowskiEngine"},{"key":"e_1_3_1_12_2","volume-title":"Proceedings of the Computer Vision and Pattern Recognition","author":"Dai Angela","year":"2017","unstructured":"Angela Dai, Angel X. Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nie\u00dfner. 2017. ScanNet: Richly-annotated 3D reconstructions of indoor scenes. In Proceedings of the Computer Vision and Pattern Recognition."},{"key":"e_1_3_1_13_2","unstructured":"Angela Dai Christian Diller and Matthias Nie\u00dfner. 2020. SG-NN: Sparse generative neural networks for self-supervised scene completion of RGB-D scans. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 849\u2013858."},{"key":"e_1_3_1_14_2","first-page":"189","volume-title":"Proceedings of the 2010 ACM\/IEEE International Symposium on Low-Power Electronics and Design.","author":"David Howard","year":"2010","unstructured":"Howard David, Eugene Gorbatov, Ulf R. Hanebutte, Rahul Khanna, and Christian Le. 2010. RAPL: Memory power estimation and capping. In Proceedings of the 2010 ACM\/IEEE International Symposium on Low-Power Electronics and Design.IEEE, 189\u2013194."},{"key":"e_1_3_1_15_2","first-page":"178","volume-title":"Proceedings of the 2019 28th International Conference on Parallel Architectures and Compilation Techniques","author":"Dong X.","year":"2019","unstructured":"X. Dong, L. Liu, P. Zhao, G. Li, J. Li, X. Wang, and X. Feng. 2019. Acorns: A framework for accelerating deep neural networks with input sparsity. In Proceedings of the 2019 28th International Conference on Parallel Architectures and Compilation Techniques. 178\u2013191."},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3470496.3527395"},{"key":"e_1_3_1_17_2","volume-title":"Proceedings of the 53th International Symposium on Microarchitecture","author":"Feng Yu","year":"2020","unstructured":"Yu Feng, Boyuan Tian, Tiancheng Xu, Paul Whatmough, and Yuhao Zhu. 2020. Mesorasi: Architecture support for point cloud analytics via delayed-aggregation. In Proceedings of the 53th International Symposium on Microarchitecture. ACM."},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3297858.3304014"},{"key":"e_1_3_1_19_2","first-page":"1","volume-title":"Proceedings of the 2016 IEEE International Conference on Networking, Architecture and Storage","author":"Giardino Michael","year":"2016","unstructured":"Michael Giardino and Bonnie Ferri. 2016. Correlating hardware performance events to CPU and DRAM power consumption. In Proceedings of the 2016 IEEE International Conference on Networking, Architecture and Storage. IEEE, 1\u20132."},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3352460.3358291"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00961"},{"key":"e_1_3_1_22_2","unstructured":"Benjamin Graham and Laurens van der Maaten. 2017. Submanifold sparse convolutional networks. arXiv:1706.01307. Retrieved from https:\/\/arxiv.org\/abs\/1706.01307"},{"key":"e_1_3_1_23_2","doi-asserted-by":"crossref","unstructured":"Yulan Guo Hanyun Wang Qingyong Hu Hao Liu Li Liu and Mohammed Bennamoun. 2020. Deep learning for 3d point clouds: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence 43 12 (2020) 4338\u20134364.","DOI":"10.1109\/TPAMI.2020.3005434"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00301"},{"key":"e_1_3_1_25_2","first-page":"933","volume-title":"Proceedings of the 2018 51st Annual IEEE\/ACM International Symposium on Microarchitecture.","author":"Hegde Kartik","year":"2018","unstructured":"Kartik Hegde, Rohit Agrawal, Yulun Yao, and Christopher W. Fletcher. 2018. Morph: Flexible acceleration for 3D CNN-based video understanding. In Proceedings of the 2018 51st Annual IEEE\/ACM International Symposium on Microarchitecture.IEEE, 933\u2013946."},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/3352460.3358275"},{"issue":"6","key":"e_1_3_1_27_2","first-page":"1","article-title":"Taichi: A language for high-performance computation on spatially sparse data structures","volume":"38","author":"Hu Yuanming","year":"2019","unstructured":"Yuanming Hu, Tzu-Mao Li, Luke Anderson, Jonathan Ragan-Kelley, and Fr\u00e9do Durand. 2019. Taichi: A language for high-performance computation on spatially sparse data structures. ACM Transactions on Graphics 38, 6 (2019), 1\u201316.","journal-title":"ACM Transactions on Graphics"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/3352460.3358283"},{"key":"e_1_3_1_29_2","first-page":"1","volume-title":"Proceedings of the 2020 ACM\/IEEE 47th Annual International Symposium on Computer Architecture.","author":"Hwang Ranggi","year":"2020","unstructured":"Ranggi Hwang, Taehun Kim, Youngeun Kwon, and Minsoo Rhu. 2020. Centaur: A chiplet-based, hybrid sparse-dense accelerator for personalized recommendations. In Proceedings of the 2020 ACM\/IEEE 47th Annual International Symposium on Computer Architecture.IEEE, 1."},{"key":"e_1_3_1_30_2","unstructured":"Google Inc. 2026. Google Sparse Hash. (n.d.). Retrieved Aug 12 2020 from https:\/\/github.com\/sparsehash\/sparsehash"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","unstructured":"Norman P. Jouppi Cliff Young Nishant Patil David Patterson Gaurav Agrawal Raminder Bajwa Sarah Bates Suresh Bhatia Nan Boden Al Borchers et\u00a0al. 2017. In-datacenter performance analysis of a tensor processing unit.InProceedings of the 44th Annual International Symposium on Computer Architecture. Association for Computing Machinery New York NY USA 1\u201312. DOI:10.1145\/3079856.3080246","DOI":"10.1145\/3079856.3080246"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.media.2016.10.004"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1145\/3352460.3358286"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/ASE.2017.8115709"},{"key":"e_1_3_1_35_2","first-page":"461","volume-title":"Proceedings of the 23rd International Conference on Architectural Support for Programming Languages and Operating Systems.","author":"Kwon Hyoukjun","year":"2018","unstructured":"Hyoukjun Kwon, Ananda Samajdar, and Tushar Krishna. 2018. MAERI: Enabling flexible dataflow mapping over DNN accelerators via reconfigurable interconnects. In Proceedings of the 23rd International Conference on Architectural Support for Programming Languages and Operating Systems.ACM, New York, NY, USA, 461\u2013475."},{"key":"e_1_3_1_36_2","unstructured":"Xuesong Li Jose Guivant Ngaiming Kwok Yongzhi Xu Ruowei Li and Hongkun Wu. 2019. Three-dimensional Backbone Network for 3D Object Detection in Traffic Scenes. arXiv:1901.08373. Retrieved from https:\/\/arxiv.org\/abs\/1901.08373"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2020.3015992"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","unstructured":"Xu Yan Yingfeng Lu Zhen Li Qing Wei Xin Gao Sheng Wang Song Wu and Shuguang Cui. 2022. PointSite: A point cloud segmentation tool for identification of protein ligand binding atoms. Journal of Chemical Information and Modeling 62 11 (2022) 2835\u20132845. DOI:10.1021\/acs.jcim.1c01512","DOI":"10.1021\/acs.jcim.1c01512"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","unstructured":"Zhidong Liang Ming Yang Hao Li and Chunxiang Wang. 2020. 3D Instance embedding learning with a structure-aware loss function for point cloud segmentation. IEEE Robotics and Automation Letters 5 3 (2020) 4915\u20134922. DOI:10.1109\/LRA.2020.3004802","DOI":"10.1109\/LRA.2020.3004802"},{"key":"e_1_3_1_40_2","volume-title":"Proceedings of the 54th Annual IEEE\/ACM International Symposium on Microarchitecture","author":"Lin Yujun","year":"2021","unstructured":"Yujun Lin, Zhekai Zhang, Haotian Tang, Hanrui Wang, and Song Han. 2021. PointAcc: Efficient point cloud accelerator. In Proceedings of the 54th Annual IEEE\/ACM International Symposium on Microarchitecture."},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00534"},{"key":"e_1_3_1_42_2","first-page":"553","volume-title":"Proceedings of the 2017 IEEE International Symposium on High Performance Computer Architecture","author":"Lu Wenyan","year":"2017","unstructured":"Wenyan Lu, Guihai Yan, Jiajun Li, Shijun Gong, Yinhe Han, and Xiaowei Li. 2017. Flexflow: A flexible dataflow accelerator architecture for convolutional neural networks. In Proceedings of the 2017 IEEE International Symposium on High Performance Computer Architecture. IEEE, 553\u2013564."},{"key":"e_1_3_1_43_2","unstructured":"DDR Micron. 2020. Power Calculator. (2020). Retrieved Jul 22 2021 from https:\/\/www.micron.com\/support\/tools-and-utilities\/power-calc"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2018.2886133"},{"key":"e_1_3_1_45_2","unstructured":"Nvidia. 2016. NVIDIA visual profiler. (2016). Retrieved Jul 22 2021 from https:\/\/developer.nvidia.com\/nvidia-visual-profiler"},{"key":"e_1_3_1_46_2","first-page":"27","volume-title":"Proceedings of the 2017 ACM\/IEEE 44th Annual International Symposium on Computer Architecture .","author":"Parashar Angshuman","year":"2017","unstructured":"Angshuman Parashar, Minsoo Rhu, Anurag Mukkara, Antonio Puglielli, Rangharajan Venkatesan, Brucek Khailany, Joel Emer, Stephen W. Keckler, and William J. Dally. 2017. Scnn: An accelerator for compressed-sparse convolutional neural networks. In Proceedings of the 2017 ACM\/IEEE 44th Annual International Symposium on Computer Architecture .IEEE, 27\u201340."},{"key":"e_1_3_1_47_2","first-page":"652","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Qi Charles R.","year":"2017","unstructured":"Charles R. Qi, Hao Su, Kaichun Mo, and Leonidas J. Guibas. 2017. Pointnet: Deep learning on point sets for 3d classification and segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 652\u2013660."},{"key":"e_1_3_1_48_2","first-page":"5099","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Qi Charles Ruizhongtai","year":"2017","unstructured":"Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J. Guibas. 2017. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. In Proceedings of the Advances in Neural Information Processing Systems. 5099\u20135108."},{"key":"e_1_3_1_49_2","unstructured":"Facebook Research. 2026. Submanifold Sparse Convolutional. (n.d.). Retrieved Jul 22 2020 from https:\/\/github.com\/facebookresearch\/SparseConvNet\/tree\/master\/sparseconvnet\/SCN"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","unstructured":"S. Schmohl and U. S\u00f6rgel. 2019. Submanifold sparse convolutional networks for semantic segmentation of large-scaleals point clouds. ISPRS Annals of the Photogrammetry Remote Sensing and Spatial Information Sciences IV-2\/W5 (2019) 77\u201384. DOI:10.5194\/isprs-annals-IV-2-W5-77-2019","DOI":"10.5194\/isprs-annals-IV-2-W5-77-2019"},{"key":"e_1_3_1_51_2","unstructured":"Shaoshuai Shi Chaoxu Guo Li Jiang Zhe Wang Jianping Shi Xiaogang Wang and Hongsheng Li. 2020. PV-RCNN: Point-voxel feature set abstraction for 3D object detection. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 10529\u201310538."},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","unstructured":"Shaoshuai Shi Zhe Wang Jianping Shi Xiaogang Wang and Hongsheng Li. 2021. From points to parts: 3D object detection from point cloud with part-aware and part-aggregation network. IEEE Transactions on Pattern Analysis and Machine Intelligence 43 8 (2021) 2647\u20132664. DOI:10.1109\/TPAMI.2020.2977026","DOI":"10.1109\/TPAMI.2020.2977026"},{"key":"e_1_3_1_53_2","first-page":"arXiv\u20131912","article-title":"Scalability in perception for autonomous driving: Waymo open dataset","author":"Sun Pei","year":"2019","unstructured":"Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, et\u00a0al. 2019. Scalability in perception for autonomous driving: Waymo open dataset. arXiv (2019), arXiv\u20131912.","journal-title":"arXiv"},{"key":"e_1_3_1_54_2","volume-title":"Proceedings of the Conference on Machine Learning and Systems.","author":"Tang Haotian","year":"2022","unstructured":"Haotian Tang, Zhijian Liu, Xiuyu Li, Yujun Lin, and Song Han. 2022. TorchSparse: Efficient point cloud inference engine. In Proceedings of the Conference on Machine Learning and Systems."},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58604-1_41"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1145\/3613424.3614303"},{"key":"e_1_3_1_57_2","first-page":"47","volume-title":"Proceedings of the Vmv","author":"Teschner Matthias","year":"2003","unstructured":"Matthias Teschner, Bruno Heidelberger, Matthias M\u00fcller, Danat Pomerantes, and Markus H. Gross. 2003. Optimized spatial hashing for collision detection of deformable objects. In Proceedings of the Vmv. 47\u201354."},{"key":"e_1_3_1_58_2","unstructured":"Linley Gwennap. 2018. Graphcore makes big AI splash. Microprocessor Rep. The Linley Group Mountain View CA USA (2018)."},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1145\/3624062.3624084"},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.scs.2019.102002"},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.1145\/3326362"},{"key":"e_1_3_1_62_2","doi-asserted-by":"publisher","DOI":"10.1145\/3352460.3358259"},{"key":"e_1_3_1_63_2","unstructured":"Xuan Yang Jing Pu Blaine Burton Rister Nikhil Bhagdikar Stephen Richardson Shahar Kvatinsky Jonathan Ragan-Kelley Ardavan Pedram and Mark Horowitz. 2016. A systematic approach to blocking convolutional neural networks. arXiv:1606.04209. Retrieved from https:\/\/arxiv.org\/abs\/1606.04209"},{"key":"e_1_3_1_64_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2019.09.086"},{"issue":"12","key":"e_1_3_1_65_2","doi-asserted-by":"crossref","first-page":"220104","DOI":"10.1007\/s11432-019-2636-x","article-title":"ARPNET: Attention region proposal network for 3D object detection","volume":"62","author":"Ye Yangyang","year":"2019","unstructured":"Yangyang Ye, Chi Zhang, and Xiaoli Hao. 2019. ARPNET: Attention region proposal network for 3D object detection. Science China Information Sciences 62, 12 (2019), 220104.","journal-title":"Science China Information Sciences"},{"key":"e_1_3_1_66_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01258-8_45"},{"key":"e_1_3_1_67_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2019.2958671"},{"key":"e_1_3_1_68_2","first-page":"20","volume-title":"Proceedings of the 49th Annual IEEE\/ACM International Symposium on Microarchitecture","author":"Zhang Shijin","year":"2016","unstructured":"Shijin Zhang, Zidong Du, Lei Zhang, Huiying Lan, Shaoli Liu, Ling Li, Qi Guo, Tianshi Chen, and Yunji Chen. 2016. Cambricon-x: An accelerator for sparse neural networks. In Proceedings of the 49th Annual IEEE\/ACM International Symposium on Microarchitecture. IEEE Press, 20."}],"container-title":["ACM Transactions on Embedded Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3759457","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,13]],"date-time":"2026-05-13T13:59:14Z","timestamp":1778680754000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3759457"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,13]]},"references-count":67,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,5,31]]}},"alternative-id":["10.1145\/3759457"],"URL":"https:\/\/doi.org\/10.1145\/3759457","relation":{},"ISSN":["1539-9087","1558-3465"],"issn-type":[{"value":"1539-9087","type":"print"},{"value":"1558-3465","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,5,13]]},"assertion":[{"value":"2024-11-30","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-06-07","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-05-13","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}