{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T22:54:33Z","timestamp":1777676073167,"version":"3.51.4"},"reference-count":25,"publisher":"SAGE Publications","issue":"3","license":[{"start":{"date-parts":[[2024,10,23]],"date-time":"2024-10-23T00:00:00Z","timestamp":1729641600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key R&D Program of China","doi-asserted-by":"crossref","award":["2021YFB0300800"],"award-info":[{"award-number":["2021YFB0300800"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Key Program of National Natural Science Foundation of China","award":["U21A20461, 92055213"],"award-info":[{"award-number":["U21A20461, 92055213"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61872127"],"award-info":[{"award-number":["61872127"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100004761","name":"Natural Science Foundation of Hunan Province, China","doi-asserted-by":"crossref","award":["2021JJ50158"],"award-info":[{"award-number":["2021JJ50158"]}],"id":[{"id":"10.13039\/501100004761","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Key-Area Research and Development Program of Guangdong Province","award":["2021B0101190004"],"award-info":[{"award-number":["2021B0101190004"]}]}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["The International Journal of High Performance Computing Applications"],"published-print":{"date-parts":[[2025,5]]},"abstract":"<jats:p>In circuit simulators that resemble the Simulation Program with Integrated Circuit Emphasis (SPICE), one of the most crucial steps is the solution of numerous sparse linear equations generated by frequency domain analysis or time domain analysis. The sparse direct solvers based on lower-upper (LU) factorization are extremely time-consuming, so their performance has become a significant bottleneck. Despite the existence of some parallel sparse direct solvers for circuit simulation problems, they remain challenging to adapt in terms of performance and scalability in the face of rapidly evolving parallel computers with multiple NUMA hardware based on ARM architecture. In this paper, we introduce a parallel sparse direct solver named HLU, which re-examines the performance of the parallel algorithm from the viewpoint of parallelism in pipeline mode and the computing efficiency of each task. To maximize task-level parallelism and further minimize the thread waiting time, HLU devises a fine-grained scheduling method based on an elimination tree in pipeline mode, which employs depth-first search (DFS-like) to iteratively search for parent tasks and then place dependent tasks in the same task queue. HLU also suggests two NUMA node affinity strategies: thread affinity optimization based on NUMA nodes topology to guarantee computational load balancing and data affinity optimization to enable effective memory placement when threads access data. The rationality and effectiveness of the sparse solver HLU are validated by the SuiteSparse Matrix Collection. In comparison with KLU and NICSLU, the experimental results and analysis show that HLU attains a speedup of up to 9.14\u00d7 and 1.26x (geometric mean) on a Huawei Kunpeng 920 Server, respectively.<\/jats:p>","DOI":"10.1177\/10943420241241491","type":"journal-article","created":{"date-parts":[[2024,10,23]],"date-time":"2024-10-23T20:19:21Z","timestamp":1729714761000},"page":"405-423","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":0,"title":["NUMA-aware parallel sparse LU factorization for SPICE-based circuit simulators on ARM multi-core processors"],"prefix":"10.1177","volume":"39","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8346-1811","authenticated-orcid":false,"given":"Junsheng","family":"Zhou","sequence":"first","affiliation":[{"name":"College of Computer Science and Electronic Engineering, Hunan University, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wangdong","family":"Yang","sequence":"additional","affiliation":[{"name":"College of Computer Science and Electronic Engineering, Hunan University, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fengkun","family":"Dong","sequence":"additional","affiliation":[{"name":"College of Computer Science and Electronic Engineering, Hunan University, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shengle","family":"Lin","sequence":"additional","affiliation":[{"name":"College of Computer Science and Electronic Engineering, Hunan University, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qinyun","family":"Cai","sequence":"additional","affiliation":[{"name":"College of Computer Science and Electronic Engineering, Hunan University, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kenli","family":"Li","sequence":"additional","affiliation":[{"name":"College of Computer Science and Electronic Engineering, Hunan University, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[2024,10,23]]},"reference":[{"key":"e_1_3_3_2_1","first-page":"43","article-title":"Corey: an operating system for many cores","volume":"8","author":"Boyd-Wickizer S","year":"2008","unstructured":"Boyd-Wickizer S, Chen H, Chen R, et al. (2008) Corey: an operating system for many cores. OSDI 8: 43\u201357.","journal-title":"OSDI"},{"key":"e_1_3_3_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2012.2217964"},{"key":"e_1_3_3_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2014.2312199"},{"key":"e_1_3_3_5_1","doi-asserted-by":"publisher","DOI":"10.3850\/9783981537079_0839"},{"key":"e_1_3_3_6_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-53429-9"},{"key":"e_1_3_3_7_1","volume-title":"Accelerating Analog Simulation with HSPICE Precision Parallel Technology","author":"Daniel R","year":"2010","unstructured":"Daniel R, Sosen HV, Elhak H (2010) Accelerating Analog Simulation with HSPICE Precision Parallel Technology. Sunnyvale: Synopsys Corporation. Technical report."},{"key":"e_1_3_3_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/2499368.2451157"},{"key":"e_1_3_3_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/992200.992206"},{"key":"e_1_3_3_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/2049662.2049670"},{"key":"e_1_3_3_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/1824801.1824814"},{"key":"e_1_3_3_12_1","doi-asserted-by":"publisher","DOI":"10.1017\/S0962492916000076"},{"key":"e_1_3_3_13_1","doi-asserted-by":"publisher","DOI":"10.1137\/S0895479897317685"},{"key":"e_1_3_3_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2011.308"},{"key":"e_1_3_3_15_1","volume-title":"A NUMA-aware scheduler for a parallel sparse direct solver. Research Report RR-7498","author":"Faverge M","year":"2010","unstructured":"Faverge M, Lacoste X, Ramet P (2010) A NUMA-aware scheduler for a parallel sparse direct solver. Research Report RR-7498. Le Chesnay-Rocquencourt: INRIA."},{"key":"e_1_3_3_16_1","doi-asserted-by":"publisher","DOI":"10.1137\/0909058"},{"key":"e_1_3_3_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/355841.355847"},{"key":"e_1_3_3_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2018.2858014"},{"key":"e_1_3_3_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/1089014.1089017"},{"key":"e_1_3_3_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3458817.3476199"},{"key":"e_1_3_3_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/2692916.2555271"},{"key":"e_1_3_3_22_1","doi-asserted-by":"publisher","DOI":"10.1177\/1094342020981153"},{"key":"e_1_3_3_23_1","volume-title":"SPICE 2:A Computer Program to Stimulate Semiconductor Circuits","author":"Nagel LW","year":"1975","unstructured":"Nagel LW (1975) SPICE 2:A Computer Program to Stimulate Semiconductor Circuits. Berkeley: Ph.D dissertation, Dept Electric Eng Comput Sci, University California."},{"key":"e_1_3_3_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/MDAT.2020.2974910"},{"key":"e_1_3_3_25_1","doi-asserted-by":"publisher","DOI":"10.23919\/DATE54114.2022.9774499"},{"key":"e_1_3_3_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/2463209.2488905"}],"container-title":["The International Journal of High Performance Computing Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/10943420241241491","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/10943420241241491","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/10943420241241491","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T08:17:35Z","timestamp":1777450655000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/10943420241241491"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,10,23]]},"references-count":25,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2025,5]]}},"alternative-id":["10.1177\/10943420241241491"],"URL":"https:\/\/doi.org\/10.1177\/10943420241241491","relation":{},"ISSN":["1094-3420","1741-2846"],"issn-type":[{"value":"1094-3420","type":"print"},{"value":"1741-2846","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,10,23]]}}}