{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,22]],"date-time":"2026-07-22T01:55:50Z","timestamp":1784685350482,"version":"3.55.0"},"reference-count":54,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2023,6,20]],"date-time":"2023-06-20T00:00:00Z","timestamp":1687219200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key R&D Program of China","doi-asserted-by":"crossref","award":["2020YFB2104000"],"award-info":[{"award-number":["2020YFB2104000"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Key Program of National Natural Science Foundation of China","award":["U21A20461"],"award-info":[{"award-number":["U21A20461"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61872127 and 61751204"],"award-info":[{"award-number":["61872127 and 61751204"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Research Innovation Project for Postgraduate Students of Hunan Province","award":["CX20220412"],"award-info":[{"award-number":["CX20220412"]}]},{"name":"GHFUND A","award":["ghfund202107013482"],"award-info":[{"award-number":["ghfund202107013482"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Parallel Comput."],"published-print":{"date-parts":[[2023,6,30]]},"abstract":"<jats:p>\n            <jats:bold>Sparse Tensor-Times-Matrix (SpTTM)<\/jats:bold>\n            is the core calculation in tensor analysis. The sparse distributions of different tensors vary greatly, which poses a big challenge to designing efficient and general SpTTM. In this paper, we describe SpTTM on CPU-GPU heterogeneous hybrid systems and give a parallel execution strategy for SpTTM in different sparse formats. We analyze the theoretical computer powers and estimate the number of tasks to achieve the load balancing between the CPU and the GPU of the heterogeneous systems. We discuss a method to describe tensor sparse structure by graph structure and design a new graph neural network SPT-GCN to select a suitable tensor sparse format. Furthermore, we perform extensive experiments using real datasets to demonstrate the advantages and efficiency of our proposed input-aware slice-wise SpTTM. The experimental results show that our input-aware slice-wise SpTTM can achieve an average speedup of 1.310 \u00d7 compared to ParTI! library on a CPU-GPU heterogeneous system.\n          <\/jats:p>","DOI":"10.1145\/3584373","type":"journal-article","created":{"date-parts":[[2023,2,17]],"date-time":"2023-02-17T11:58:34Z","timestamp":1676635114000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":7,"title":["A Heterogeneous Parallel Computing Approach Optimizing SpTTM on CPU-GPU via GCN"],"prefix":"10.1145","volume":"10","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-0086-6301","authenticated-orcid":false,"given":"Haotian","family":"Wang","sequence":"first","affiliation":[{"name":"College of Computer Science and Electronic Engineering, Hunan University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2681-7898","authenticated-orcid":false,"given":"Wangdong","family":"Yang","sequence":"additional","affiliation":[{"name":"College of Computer Science and Electronic Engineering, Hunan University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8712-4561","authenticated-orcid":false,"given":"Renqiu","family":"Ouyang","sequence":"additional","affiliation":[{"name":"College of Computer Science and Electronic Engineering, Hunan University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4174-1326","authenticated-orcid":false,"given":"Rong","family":"Hu","sequence":"additional","affiliation":[{"name":"College of Computer Science and Electronic Engineering, Hunan University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2635-7716","authenticated-orcid":false,"given":"Kenli","family":"Li","sequence":"additional","affiliation":[{"name":"College of Computer Science and Electronic Engineering, Hunan University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5224-4048","authenticated-orcid":false,"given":"Keqin","family":"Li","sequence":"additional","affiliation":[{"name":"Department of Computer Science, State University of New York, USA and College of Computer Science and Electronic Engineering, Hunan University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,6,20]]},"reference":[{"key":"e_1_3_2_2_2","article-title":"Tensor decompositions for learning latent variable models","author":"Anandkumar Animashree","year":"2014","unstructured":"Animashree Anandkumar, Rong Ge, Daniel Hsu, Sham M. Kakade, and Matus Telgarsky. 2014. Tensor decompositions for learning latent variable models. Journal of Machine Learning Research (2014).","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/3380930"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1137\/060676489"},{"key":"e_1_3_2_5_2","volume-title":"PPSC","author":"Ballard Grey","year":"2020","unstructured":"Grey Ballard and Kathryn Rouse. 2020. General memory-independent lower bound for MTTKRP. In PPSC."},{"key":"e_1_3_2_6_2","volume-title":"ICCS","author":"Bassoy Cem Savas","year":"2019","unstructured":"Cem Savas Bassoy. 2019. Design of a high-performance tensor-vector multiplication with BLAS. In ICCS."},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICPP.2016.64"},{"key":"e_1_3_2_8_2","article-title":"Multilayer perceptron","author":"Chang Hong","year":"2021","unstructured":"Hong Chang. 2021. Multilayer perceptron. Machine Learning \u2013 A Journey to Deep Learning (2021).","journal-title":"Machine Learning \u2013 A Journey to Deep Learning"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/3385414"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/3512770"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2020.2990429"},{"key":"e_1_3_2_12_2","article-title":"Era of big data processing: A new approach via tensor networks and tensor decompositions","volume":"1403","author":"Cichocki Andrzej","year":"2014","unstructured":"Andrzej Cichocki. 2014. Era of big data processing: A new approach via tensor networks and tensor decompositions. ArXiv abs\/1403.2048 (2014).","journal-title":"ArXiv"},{"key":"e_1_3_2_13_2","article-title":"Heuristic adaptability to input dynamics for SpMM on GPUs","volume":"2202","author":"Dai Guohao","year":"2022","unstructured":"Guohao Dai, Guyue Huang, Shang Yang, Zhongming Yu, Hengrui Zhang, Yufei Ding, Yuan Xie, Huazhong Yang, and Yu Wang. 2022. Heuristic adaptability to input dynamics for SpMM on GPUs. ArXiv abs\/2202.08556 (2022).","journal-title":"ArXiv"},{"key":"e_1_3_2_14_2","volume-title":"ICLR Workshop on Representation Learning on Graphs and Manifolds","author":"Fey Matthias","year":"2019","unstructured":"Matthias Fey and Jan E. Lenssen. 2019. Fast graph representation learning with PyTorch geometric. In ICLR Workshop on Representation Learning on Graphs and Manifolds."},{"key":"e_1_3_2_15_2","first-page":"1263","volume-title":"Proceedings of the 34th International Conference on Machine Learning (Proceedings of Machine Learning Research)","volume":"70","author":"Gilmer Justin","year":"2017","unstructured":"Justin Gilmer, Samuel S. Schoenholz, Patrick F. Riley, Oriol Vinyals, and George E. Dahl. 2017. Neural message passing for quantum chemistry. In Proceedings of the 34th International Conference on Machine Learning (Proceedings of Machine Learning Research), Doina Precup and Yee Whye Teh (Eds.), Vol. 70. PMLR, 1263\u20131272."},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1137\/S1064827598341475"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jbi.2014.07.001"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICPP.2016.19"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1137\/07070111X"},{"key":"e_1_3_2_20_2","article-title":"Challenges in multimodal data fusion","author":"Lahat Dana","year":"2014","unstructured":"Dana Lahat, Tulay Adali, and Christian Jutten. 2014. Challenges in multimodal data fusion. European Signal Processing Conference (2014).","journal-title":"European Signal Processing Conference"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jneumeth.2012.03.005"},{"key":"e_1_3_2_22_2","first-page":"1","article-title":"An input-adaptive and in-place approach to dense tensor-times-matrix multiply","author":"Li Jiajia","year":"2015","unstructured":"Jiajia Li, Casey Battaglino, I. Perros, Jimeng Sun, and R. Vuduc. 2015. An input-adaptive and in-place approach to dense tensor-times-matrix multiply. SC15: International Conference for High Performance Computing, Networking, Storage and Analysis (2015), 1\u201312.","journal-title":"SC15: International Conference for High Performance Computing, Networking, Storage and Analysis"},{"key":"e_1_3_2_23_2","unstructured":"Jiajia Li Yuchen Ma and Richard Vuduc. 2018. ParTI! : A Parallel Tensor Infrastructure for multicore CPUs and GPUs. (Oct.2018). http:\/\/parti-project.org.Last updated: Jan. 2020."},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/IA3.2016.010"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3330345.3330366"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2014.2308221"},{"key":"e_1_3_2_27_2","article-title":"a-Tucker: Input-adaptive and matricization-free Tucker decomposition for dense tensors on CPUs and GPUs","volume":"2010","author":"Li Min","year":"2020","unstructured":"Min Li, Chuanfu Xiao, and Chao Yang. 2020. a-Tucker: Input-adaptive and matricization-free Tucker decomposition for dense tensors on CPUs and GPUs. ArXiv abs\/2010.10131 (2020).","journal-title":"ArXiv"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/CLUSTER.2017.75"},{"key":"e_1_3_2_29_2","article-title":"Sparta: High-performance, element-wise sparse tensor contraction on heterogeneous memory","author":"Liu Jiawen","year":"2021","unstructured":"Jiawen Liu, Jie Ren, Roberto Gioiosa, Dong Li, and Jiajia Li. 2021. Sparta: High-performance, element-wise sparse tensor contraction on heterogeneous memory. ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (2021).","journal-title":"ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS49936.2021.00049"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2018.07.018"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1145\/2717511"},{"key":"e_1_3_2_33_2","volume-title":"AAAI","author":"Morris Christopher","year":"2019","unstructured":"Christopher Morris, Martin Ritzert, M. Fey, William L. Hamilton, J. E. Lenssen, Gaurav Rattan, and Martin Grohe. 2019. Weisfeiler and Leman go neural: Higher-order graph neural networks. In AAAI."},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/3295500.3356216"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2019.00023"},{"key":"e_1_3_2_36_2","article-title":"Effective machine learning based format selection and performance modeling for SpMV on GPUs","author":"Nisa Israt","year":"2018","unstructured":"Israt Nisa, Charles Siegel, Aravind Sukumaran Rajam, Abhinav Vishnu, and P. Sadayappan. 2018. Effective machine learning based format selection and performance modeling for SpMV on GPUs. International Parallel and Distributed Processing Symposium (2018).","journal-title":"International Parallel and Distributed Processing Symposium"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS49936.2021.00016"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jocs.2019.02.007"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1145\/2751205.2751244"},{"key":"e_1_3_2_40_2","article-title":"Efficient parallel sparse symmetric Tucker decomposition for high-order tensors.","author":"Shivakumar Shruti","year":"2021","unstructured":"Shruti Shivakumar, Jiajia Li, Ramakrishnan Kannan, and Srinivas Aluru. 2021. Efficient parallel sparse symmetric Tucker decomposition for high-order tensors. ACDA (2021).","journal-title":"ACDA"},{"key":"e_1_3_2_41_2","article-title":"Tensor-matrix products with a compressed sparse tensor","author":"Smith Shaden","year":"2015","unstructured":"Shaden Smith and George Karypis. 2015. Tensor-matrix products with a compressed sparse tensor. Proceedings of the 5th Workshop on Irregular Applications: Architectures and Algorithms (2015).","journal-title":"Proceedings of the 5th Workshop on Irregular Applications: Architectures and Algorithms"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2015.27"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2014.06.002"},{"key":"e_1_3_2_44_2","first-page":"1","article-title":"SpTFS: Sparse tensor format selection for MTTKRP via deep learning","author":"Sun Qingxiao","year":"2020","unstructured":"Qingxiao Sun, Yi Liu, Ming Dun, Hailong Yang, Zhongzhi Luan, Lin Gan, Guangwen Yang, and Depei Qian. 2020. SpTFS: Sparse tensor format selection for MTTKRP via deep learning. SC20: International Conference for High Performance Computing, Networking, Storage and Analysis (2020), 1\u201314.","journal-title":"SC20: International Conference for High Performance Computing, Networking, Storage and Analysis"},{"key":"e_1_3_2_45_2","first-page":"1968","article-title":"Input-aware sparse tensor storage format selection for optimizing MTTKRP","volume":"71","author":"Sun Qingxiao","year":"2022","unstructured":"Qingxiao Sun, Yi Liu, Hailong Yang, Ming Dun, Zhongzhi Luan, Lin Gan, Guangwen Yang, and Depei Qian. 2022. Input-aware sparse tensor storage format selection for optimizing MTTKRP. IEEE Trans. Comput. 71 (2022), 1968\u20131981.","journal-title":"IEEE Trans. Comput."},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3218823"},{"key":"e_1_3_2_47_2","doi-asserted-by":"crossref","unstructured":"Endong Wang Qing Zhang Bo Shen Guangyong Zhang Xiaowei Lu Qing Wu and Yajuan Wang. 2014. Intel Math Kernel Library.","DOI":"10.1007\/978-3-319-06486-4_7"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2021.3090759"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2022.3221821"},{"key":"e_1_3_2_50_2","article-title":"Adaptive sparse matrix-matrix multiplication on the GPU","author":"Winter Martin","year":"2019","unstructured":"Martin Winter, Daniel Mlakar, Rhaleb Zayer, Hans-Peter Seidel, and Markus Steinberger. 2019. Adaptive sparse matrix-matrix multiplication on the GPU. ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (2019).","journal-title":"ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/3330345.3330354"},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/3363575"},{"key":"e_1_3_2_53_2","volume-title":"NeurIPS","author":"Ying Rex","year":"2018","unstructured":"Rex Ying, Jiaxuan You, Christopher Morris, Xiang Ren, William L. Hamilton, and J. Leskovec. 2018. Hierarchical graph representation learning with differentiable pooling. In NeurIPS."},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE51399.2021.00169"},{"key":"e_1_3_2_55_2","article-title":"Bridging the gap between deep learning and sparse matrix format selection","author":"Zhao Yue","year":"2018","unstructured":"Yue Zhao, Jiajia Li, Chunhua Liao, and X. Shen. 2018. Bridging the gap between deep learning and sparse matrix format selection. Proceedings of the 23rd ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (2018).","journal-title":"Proceedings of the 23rd ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming"}],"container-title":["ACM Transactions on Parallel Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3584373","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3584373","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:46:29Z","timestamp":1750178789000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3584373"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,6,20]]},"references-count":54,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2023,6,30]]}},"alternative-id":["10.1145\/3584373"],"URL":"https:\/\/doi.org\/10.1145\/3584373","relation":{},"ISSN":["2329-4949","2329-4957"],"issn-type":[{"value":"2329-4949","type":"print"},{"value":"2329-4957","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,6,20]]},"assertion":[{"value":"2022-05-17","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-02-13","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-06-20","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}