{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:29:33Z","timestamp":1750220973931,"version":"3.41.0"},"reference-count":57,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2018,6,30]],"date-time":"2018-06-30T00:00:00Z","timestamp":1530316800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["1521092, 91430218, 31327901, 61472395, 61432018"],"award-info":[{"award-number":["1521092, 91430218, 31327901, 61472395, 61432018"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"National Key Research and Development Program of China","award":["2016YFB0201305,2016YFB0200504, 2017YFB0202105, 2016YFB0200803,2016YFB0200300"],"award-info":[{"award-number":["2016YFB0201305,2016YFB0200504, 2017YFB0202105, 2016YFB0200803,2016YFB0200300"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Parallel Comput."],"published-print":{"date-parts":[[2018,6,30]]},"abstract":"<jats:p>Automatic performance tuning (Autotuning) is an increasingly critical tuning technique for the high portable performance of Exascale applications. However, constructing an autotuner from scratch remains a challenge, even for domain experts. In this work, we propose a performance tuning and knowledge management suite (PAK) to help rapidly build autotuners. In order to accommodate existing autotuning techniques, we present an autotuning protocol that is composed of an extractor, producer, optimizer, evaluator, and learner. To achieve modularity and reusability, we also define programming interfaces for each protocol component as the fundamental infrastructure, which provides a customizable mechanism to deploy knowledge mining in the performance database. PAK\u2019s usability is demonstrated by studying two important computational kernels: stencil computation and sparse matrix-vector multiplication (SpMV). Our proposed autotuner based on PAK shows comparable performance and higher productivity than traditional autotuners by writing just a few tens of code using our autotuning protocol.<\/jats:p>","DOI":"10.1145\/3291527","type":"journal-article","created":{"date-parts":[[2019,1,7]],"date-time":"2019-01-07T13:42:28Z","timestamp":1546868548000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["An Autotuning Protocol to Rapidly Build Autotuners"],"prefix":"10.1145","volume":"5","author":[{"given":"Junhong","family":"Liu","sequence":"first","affiliation":[{"name":"State Key Laboratory of Computer Architecture, Institute of Computing Technology, Chinese Academy of Sciences, University of Chinese Academy of Sciences"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Guangming","family":"Tan","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Computer Architecture, Institute of Computing Technology, Chinese Academy of Sciences, University of Chinese Academy of Sciences"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yulong","family":"Luo","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Computer Architecture, Institute of Computing Technology, Chinese Academy of Sciences, University of Chinese Academy of Sciences"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiajia","family":"Li","sequence":"additional","affiliation":[{"name":"Computational Science and Engineering, Georgia Institute of Technology"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zeyao","family":"Mo","sequence":"additional","affiliation":[{"name":"Institute of Applied Physics and Computational Mathematics"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ninghui","family":"Sun","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Computer Architecture, Institute of Computing Technology, Chinese Academy of Sciences, University of Chinese Academy of Sciences"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2019,1,4]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/1274971.1275011"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/1542476.1542481"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/2628071.2628092"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1177\/1094342013493644"},{"volume-title":"Proceedings of the 20th Annual International Conference on High Performance Computing. 452--461","author":"Basu P.","key":"e_1_2_1_5_1","unstructured":"P. Basu , A. Venkat , M. Hall , S. Williams , B. Van Straalen , and L. Oliker . 2013b. Compiler generation and autotuning of communication-avoiding operators for geometric multigrid . In Proceedings of the 20th Annual International Conference on High Performance Computing. 452--461 . P. Basu, A. Venkat, M. Hall, S. Williams, B. Van Straalen, and L. Oliker. 2013b. Compiler generation and autotuning of communication-avoiding operators for geometric multigrid. In Proceedings of the 20th Annual International Conference on High Performance Computing. 452--461."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.parco.2017.04.002"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/263580.263662"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/1654059.1654065"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/CGO.2005.10"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/1693453.1693471"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.5555\/2388996.2389011"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/315253.314414"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1137\/070693199"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.5555\/1413370.1413375"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2004.840301"},{"key":"e_1_2_1_16_1","volume-title":"Proceedings of the 1st USENIX Workshop on Hot Topics in Parallelism (HotPar\u201909)","author":"Ganapathi Archana","year":"2009","unstructured":"Archana Ganapathi , Kaushik Datta , Armando Fox , and David Patterson . 2009 . A case for machine learning to optimize multicore performance . In Proceedings of the 1st USENIX Workshop on Hot Topics in Parallelism (HotPar\u201909) . Archana Ganapathi, Kaushik Datta, Armando Fox, and David Patterson. 2009. A case for machine learning to optimize multicore performance. In Proceedings of the 1st USENIX Workshop on Hot Topics in Parallelism (HotPar\u201909)."},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2009.5161004"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/2464996.2467268"},{"key":"e_1_2_1_19_1","unstructured":"R. Himeno. 2011. Himeno benchmark. Retrieved from http:\/\/accc.riken.jp\/2444.htm.  R. Himeno. 2011. Himeno benchmark. Retrieved from http:\/\/accc.riken.jp\/2444.htm."},{"volume-title":"Proceedings of the 2016 IEEE International Parallel and Distributed Processing Symposium (IPDPS\u201916)","author":"Hou K.","key":"e_1_2_1_20_1","unstructured":"K. Hou , H. Wang , and W. C. Feng . 2016. AAlign: A SIMD framework for pairwise sequence alignment on x86-based multi-and many-core processors . In Proceedings of the 2016 IEEE International Parallel and Distributed Processing Symposium (IPDPS\u201916) . 780--789. K. Hou, H. Wang, and W. C. Feng. 2016. AAlign: A SIMD framework for pairwise sequence alignment on x86-based multi-and many-core processors. In Proceedings of the 2016 IEEE International Parallel and Distributed Processing Symposium (IPDPS\u201916). 780--789."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/3075564.3075583"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2018.2789903"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2010.5470421"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/2400682.2400690"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1137\/130930352"},{"key":"e_1_2_1_26_1","volume-title":"Luebbers","author":"Kunz Karl S.","year":"1993","unstructured":"Karl S. Kunz and Raymond J . Luebbers . 1993 . The Finite Difference Time Domain Method for Electromagnetics. CRC Press . Karl S. Kunz and Raymond J. Luebbers. 1993. The Finite Difference Time Domain Method for Electromagnetics. CRC Press."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3126908.3126931"},{"volume-title":"Proceedings of the 2017 IEEE International Parallel and Distributed Processing Symposium (IPDPS\u201917)","author":"Li J.","key":"e_1_2_1_28_1","unstructured":"J. Li , J. Choi , I. Perros , J. Sun , and R. Vuduc . 2017a. Model-driven sparse CP decomposition for higher-order tensors . In Proceedings of the 2017 IEEE International Parallel and Distributed Processing Symposium (IPDPS\u201917) . 1048--1057. J. Li, J. Choi, I. Perros, J. Sun, and R. Vuduc. 2017a. Model-driven sparse CP decomposition for higher-order tensors. In Proceedings of the 2017 IEEE International Parallel and Distributed Processing Symposium (IPDPS\u201917). 1048--1057."},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/2499370.2462181"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3178487.3178529"},{"key":"e_1_2_1_31_1","unstructured":"Junhong Liu Guangming Tan Yulong Luo Zeyao Mo and Ninghui Sun. 2015. PAK. Retrieved from https:\/\/github.com\/PAA-NCIC\/PAK.  Junhong Liu Guangming Tan Yulong Luo Zeyao Mo and Ninghui Sun. 2015. PAK. Retrieved from https:\/\/github.com\/PAA-NCIC\/PAK."},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1002\/cpe.4244"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/2751205.2751209"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2015.06.010"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/2751205.2751214"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/2400682.2400718"},{"key":"e_1_2_1_38_1","volume-title":"TAU: Tuning and Analysis Utilities. Technical Report. LALP-99--205. Los Alamos National Laboratory Publication.","author":"Malony Allen D.","year":"1999","unstructured":"Allen D. Malony , Jan Cuny , and Sameer Shende . 1999 . TAU: Tuning and Analysis Utilities. Technical Report. LALP-99--205. Los Alamos National Laboratory Publication. Allen D. Malony, Jan Cuny, and Sameer Shende. 1999. TAU: Tuning and Analysis Utilities. Technical Report. LALP-99--205. Los Alamos National Laboratory Publication."},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/CLUSTER.2012.46"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/2063384.2063398"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2011.70"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/1542275.1542313"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/1513895.1513905"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3012011"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2010.2"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/GREEN.2012.6200963"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/3218823"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/1989493.1989508"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.5555\/762761.762771"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2009.5161054"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/1995896.1995932"},{"key":"e_1_2_1_52_1","volume-title":"Yelick","author":"Vuduc Richard","year":"2005","unstructured":"Richard Vuduc , James W. Demmel , and Katherine A . Yelick . 2005 . OSKI : A library of automatically tuned sparse matrix kernels. In Journal of Physics: Conference Series, Vol. 16 . IOP Publishing , 521. Richard Vuduc, James W. Demmel, and Katherine A. Yelick. 2005. OSKI: A library of automatically tuned sparse matrix kernels. In Journal of Physics: Conference Series, Vol. 16. IOP Publishing, 521."},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1145\/2925426.2926291"},{"volume-title":"Proceedings of the 2017 IEEE International Parallel and Distributed Processing Symposium (IPDPS\u201917)","author":"Wang H.","key":"e_1_2_1_54_1","unstructured":"H. Wang , J. Zhang , D. Zhang , S. Pumma , and W. C. Feng . 2017. PaPar: A parallel data partitioning framework for big data applications . In Proceedings of the 2017 IEEE International Parallel and Distributed Processing Symposium (IPDPS\u201917) . 605--614. H. Wang, J. Zhang, D. Zhang, S. Pumma, and W. C. Feng. 2017. PaPar: A parallel data partitioning framework for big data applications. In Proceedings of the 2017 IEEE International Parallel and Distributed Processing Symposium (IPDPS\u201917). 605--614."},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/3178487.3178513"},{"volume-title":"Proceedings of the 1998 ACM\/IEEE Conference on Supercomputing. IEEE Computer Society, 1--27","author":"Clint Whaley R.","key":"e_1_2_1_56_1","unstructured":"R. Clint Whaley and Jack J. Dongarra . 1998. Automatically tuned linear algebra software . In Proceedings of the 1998 ACM\/IEEE Conference on Supercomputing. IEEE Computer Society, 1--27 . R. Clint Whaley and Jack J. Dongarra. 1998. Automatically tuned linear algebra software. In Proceedings of the 1998 ACM\/IEEE Conference on Supercomputing. IEEE Computer Society, 1--27."},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/1498765.1498785"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1145\/3178487.3178495"}],"container-title":["ACM Transactions on Parallel Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3291527","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3291527","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T23:54:33Z","timestamp":1750204473000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3291527"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,6,30]]},"references-count":57,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2018,6,30]]}},"alternative-id":["10.1145\/3291527"],"URL":"https:\/\/doi.org\/10.1145\/3291527","relation":{},"ISSN":["2329-4949","2329-4957"],"issn-type":[{"type":"print","value":"2329-4949"},{"type":"electronic","value":"2329-4957"}],"subject":[],"published":{"date-parts":[[2018,6,30]]},"assertion":[{"value":"2017-01-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2018-06-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-01-04","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}