{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,2,21]],"date-time":"2025-02-21T07:39:01Z","timestamp":1740123541888,"version":"3.37.3"},"reference-count":37,"publisher":"Springer Science and Business Media LLC","issue":"11","license":[{"start":{"date-parts":[[2023,3,10]],"date-time":"2023-03-10T00:00:00Z","timestamp":1678406400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,3,10]],"date-time":"2023-03-10T00:00:00Z","timestamp":1678406400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61202076"],"award-info":[{"award-number":["61202076"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100005089","name":"Beijing Municipal Natural Science Foundation","doi-asserted-by":"publisher","award":["4192007"],"award-info":[{"award-number":["4192007"]}],"id":[{"id":"10.13039\/501100005089","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Supercomput"],"published-print":{"date-parts":[[2023,7]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Normally, threads in a warp do not severely interfere with each other. However, the scheduler must wait until all the threads within complete before scheduling the next warp, resulting in memory divergence. The crux of the problem is scheduling the warp in a more reasonable order. Therefore, we propose a new warp scheduling strategy called WSMP, which is based on multi-level feedback queue (MFQ) and perceptron-based prefetch filtering (PPF). All the warps are sorted beforehand according to the latency tolerance of the warps and pushed into a certain queue in MFQ. We also remold PPF to enhance the modified underlying prefetcher. We are able to strike a balance between cache hit rate and prefetch coverage then. We verify its feasibility using GPGPU-Sim, along with exclusive GPGPU workload. The results show that compared to the baseline, WSMP improves IPC by 26.45% and reduces L2 cache miss rate by 9.54% on average.<\/jats:p>","DOI":"10.1007\/s11227-023-05127-0","type":"journal-article","created":{"date-parts":[[2023,3,26]],"date-time":"2023-03-26T21:21:38Z","timestamp":1679865698000},"page":"12317-12340","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["WSMP: a warp scheduling strategy based on MFQ and PPF"],"prefix":"10.1007","volume":"79","author":[{"given":"Juan","family":"Fang","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Li\u2019ang","family":"Zhao","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Min","family":"Cai","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Huijing","family":"Yang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2023,3,10]]},"reference":[{"key":"5127_CR1","doi-asserted-by":"publisher","first-page":"4056","DOI":"10.1007\/s11227-015-1504-y","volume":"71","author":"J Fang","year":"2015","unstructured":"Fang J, Yu L, Liu S, Lu J, Chen T (2015) Kl_ga: an application mapping algorithm for mesh-of-tree (mot) architecture in network-on-chip design. J Supercomput 71:4056\u20134071. https:\/\/doi.org\/10.1007\/s11227-015-1504-y","journal-title":"J Supercomput"},{"key":"5127_CR2","doi-asserted-by":"publisher","unstructured":"Jablin JA, Jablin TB, Mutlu O, Herlihy M (2014) Warp-aware trace scheduling for gpus. In: 2014 23rd International Conference on Parallel Architecture and Compilation Techniques (PACT), pp. 163\u2013174 https:\/\/doi.org\/10.1145\/2628071.2628101","DOI":"10.1145\/2628071.2628101"},{"key":"5127_CR3","doi-asserted-by":"publisher","unstructured":"Choi H, Ahn J, Sung W (2012) Reducing off-chip memory traffic by selective cache management scheme in gpgpus. In: Proceedings of the 5th Annual Workshop on General Purpose Processing with Graphics Processing Units. GPGPU-5, pp. 110\u2013119. Association for Computing Machinery, New York, NY, USA https:\/\/doi.org\/10.1145\/2159430.2159443","DOI":"10.1145\/2159430.2159443"},{"issue":"11","key":"5127_CR4","doi-asserted-by":"publisher","first-page":"3153","DOI":"10.1109\/TC.2015.2395427","volume":"64","author":"Z Yu","year":"2015","unstructured":"Yu Z, Eeckhout L, Goswami N, Li T, John LK, Jin H, Xu C, Wu J (2015) Gpgpu-minibench: accelerating gpgpu micro-architecture simulation. IEEE Trans Comput 64(11):3153\u20133166. https:\/\/doi.org\/10.1109\/TC.2015.2395427","journal-title":"IEEE Trans Comput"},{"key":"5127_CR5","doi-asserted-by":"publisher","unstructured":"Koo, G, Jeon, H, Liu, Z., Kim, NS, Annavaram M (2018) Cta-aware prefetching and scheduling for gpu. In: 2018 IEEE International Parallel and Distributed Processing Symposium (IPDPS), pp. 137\u2013148 https:\/\/doi.org\/10.1109\/IPDPS.2018.00024","DOI":"10.1109\/IPDPS.2018.00024"},{"issue":"9","key":"5127_CR6","doi-asserted-by":"publisher","first-page":"1478","DOI":"10.1109\/TC.2017.2690855","volume":"66","author":"M Mao","year":"2017","unstructured":"Mao M, Wen W, Zhang Y, Chen Y, Li H (2017) An energy-efficient gpgpu register file architecture using racetrack memory. IEEE Trans Comput 66(9):1478\u20131490. https:\/\/doi.org\/10.1109\/TC.2017.2690855","journal-title":"IEEE Trans Comput"},{"key":"5127_CR7","doi-asserted-by":"publisher","unstructured":"Jang, B., Choi, M., Kim, K.K (2013). Algorithmic gpgpu memory optimization. In: 2013 International SoC Design Conference (ISOCC), pp. 154\u2013157 https:\/\/doi.org\/10.1109\/ISOCC.2013.6863959","DOI":"10.1109\/ISOCC.2013.6863959"},{"key":"5127_CR8","doi-asserted-by":"publisher","unstructured":"Liu HW, Kuo HK, Chen KT, Lai BCC (2013) Memory capacity aware non-blocking data transfer on gpgpu. In: SiPS 2013 Proceedings, pp. 395\u2013400 https:\/\/doi.org\/10.1109\/SiPS.2013.6674539","DOI":"10.1109\/SiPS.2013.6674539"},{"key":"5127_CR9","doi-asserted-by":"publisher","unstructured":"Gadhikar LM, Rao, YS (2018) Analysis of programs for gpgpu architectures. In: 2018 2nd International Conference on Trends in Electronics and Informatics (ICOEI), pp. 1\u20134 https:\/\/doi.org\/10.1109\/ICOEI.2018.8553918","DOI":"10.1109\/ICOEI.2018.8553918"},{"issue":"1","key":"5127_CR10","doi-asserted-by":"publisher","first-page":"30","DOI":"10.1109\/TC.2020.2980541","volume":"70","author":"L Yang","year":"2021","unstructured":"Yang L, Nie B, Jog A, Smirni E (2021) Practical resilience analysis of gpgpu applications in the presence of single- and multi-bit faults. IEEE Trans Comput 70(1):30\u201344. https:\/\/doi.org\/10.1109\/TC.2020.2980541","journal-title":"IEEE Trans Comput"},{"issue":"16","key":"5127_CR11","doi-asserted-by":"publisher","first-page":"12771","DOI":"10.1109\/JIOT.2020.3007751","volume":"8","author":"J Fang","year":"2021","unstructured":"Fang J, Ma A (2021) Iot application modules placement and dynamic task processing in edge-cloud computing. IEEE Int Things J 8(16):12771\u201312781. https:\/\/doi.org\/10.1109\/JIOT.2020.3007751","journal-title":"IEEE Int Things J"},{"issue":"7","key":"5127_CR12","doi-asserted-by":"publisher","first-page":"1711","DOI":"10.1109\/TC.2021.3104749","volume":"71","author":"A Segura","year":"2022","unstructured":"Segura A, Arnau J-M, Gonzalez A (2022) Energy-efficient stream compaction through filtering and coalescing accesses in gpgpu memory partitions. IEEE Trans Comput 71(7):1711\u20131723. https:\/\/doi.org\/10.1109\/TC.2021.3104749","journal-title":"IEEE Trans Comput"},{"key":"5127_CR13","doi-asserted-by":"publisher","unstructured":"Wu X, Long X (2017) Implementation of a global gpu management plugin for slurm. In: 2017 3rd International Conference on Computational Intelligence & Communication Technology (CICT), pp. 1\u20135 https:\/\/doi.org\/10.1109\/CIACT.2017.7977294","DOI":"10.1109\/CIACT.2017.7977294"},{"issue":"3","key":"5127_CR14","doi-asserted-by":"publisher","first-page":"630","DOI":"10.1109\/TPDS.2018.2868658","volume":"30","author":"KY Kim","year":"2019","unstructured":"Kim KY, Park J, Baek W (2019) Improving the performance and energy efficiency of gpgpu computing through integrated adaptive cache management. IEEE Trans Parallel Distrib Syst 30(3):630\u2013645. https:\/\/doi.org\/10.1109\/TPDS.2018.2868658","journal-title":"IEEE Trans Parallel Distrib Syst"},{"key":"5127_CR15","doi-asserted-by":"publisher","unstructured":"Lee, J., Lakshminarayana, N.B., Kim, H., Vuduc, R.: Many-thread aware prefetching mechanisms for gpgpu applications. In: 2010 43rd Annual IEEE\/ACM International Symposium on Microarchitecture, pp. 213\u2013224 (2010). https:\/\/doi.org\/10.1109\/MICRO.2010.44","DOI":"10.1109\/MICRO.2010.44"},{"issue":"4","key":"5127_CR16","doi-asserted-by":"publisher","first-page":"609","DOI":"10.1109\/TC.2018.2878671","volume":"68","author":"Y Oh","year":"2019","unstructured":"Oh Y, Kim K, Yoon MK, Park JH, Park Y, Annavaram M, Ro WW (2019) Adaptive cooperation of prefetching and warp scheduling on gpus. IEEE Trans Comput 68(4):609\u2013616. https:\/\/doi.org\/10.1109\/TC.2018.2878671","journal-title":"IEEE Trans Comput"},{"key":"5127_CR17","doi-asserted-by":"publisher","unstructured":"Zhao, Y., Cheng, G., Zhang, W., Chen, X., Li, J.: Bctcp: a feedback-based congestion control method. China Communications 17(6), 13\u201325 (2020). https:\/\/doi.org\/10.23919\/JCC.2020.06.002","DOI":"10.23919\/JCC.2020.06.002"},{"key":"5127_CR18","doi-asserted-by":"publisher","unstructured":"Kim, J., Pugsley, S.H., Gratz, P.V., Reddy, A.L.N., Wilkerson, C., Chishti, Z.: Path confidence based lookahead prefetching. In: 2016 49th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO), pp. 1\u201312 (2016). https:\/\/doi.org\/10.1109\/MICRO.2016.7783763","DOI":"10.1109\/MICRO.2016.7783763"},{"key":"5127_CR19","doi-asserted-by":"publisher","unstructured":"Michaud, P.: best-offset hardware prefetching. In: 2016 IEEE International Symposium on High Performance Computer Architecture (HPCA), pp. 469\u2013480 (2016). https:\/\/doi.org\/10.1109\/HPCA.2016.7446087","DOI":"10.1109\/HPCA.2016.7446087"},{"key":"5127_CR20","doi-asserted-by":"crossref","unstructured":"Bhatia, E., Chacon, G., Pugsley, S., Teran, E., Gratz, P.V., Jim\u00e9nez, D.A.: Perceptron-based prefetch filtering. In: 2019 ACM\/IEEE 46th Annual International Symposium on Computer Architecture (ISCA), pp. 1\u201313 (2019)","DOI":"10.1145\/3307650.3322207"},{"key":"5127_CR21","doi-asserted-by":"publisher","unstructured":"Leng, J., Hetherington, T., ElTantawy, A., Gilani, S., Kim, N.S., Aamodt, T.M., Reddi, V.J.: Gpuwattch: enabling energy optimizations in gpgpus. In: Proceedings of the 40th Annual International Symposium on Computer Architecture. ISCA \u201913, pp. 487\u2013498. Association for Computing Machinery, New York, NY, USA (2013). https:\/\/doi.org\/10.1145\/2485922.2485964","DOI":"10.1145\/2485922.2485964"},{"key":"5127_CR22","doi-asserted-by":"publisher","unstructured":"Wu, M., Pei, Y., Yu, L., Chen, T., Lou, X., Zhang, T.: Wap: The warp feature aware prefetching method for llc on cpu-gpu heterogeneous architecture. In: 2016 IEEE 18th International Conference on High Performance Computing and Communications; IEEE 14th International Conference on Smart City; IEEE 2nd International Conference on Data Science and Systems (HPCC\/SmartCity\/DSS), pp. 414\u2013421 (2016). https:\/\/doi.org\/10.1109\/HPCC-SmartCity-DSS.2016.0066","DOI":"10.1109\/HPCC-SmartCity-DSS.2016.0066"},{"issue":"3","key":"5127_CR23","doi-asserted-by":"publisher","first-page":"235","DOI":"10.1145\/2024723.2000093","volume":"39","author":"M Gebhart","year":"2011","unstructured":"Gebhart M, Johnson DR, Tarjan D, Keckler SW, Dally WJ, Lindholm E, Skadron K (2011) Energy-efficient mechanisms for managing thread context in throughput processors. SIGARCH Comput Archit News 39(3):235\u2013246. https:\/\/doi.org\/10.1145\/2024723.2000093","journal-title":"SIGARCH Comput Archit News"},{"key":"5127_CR24","doi-asserted-by":"publisher","unstructured":"Narasiman, V., Shebanow, M., Lee, C.J., Miftakhutdinov, R., Mutlu, O., Patt, Y.N.: Improving gpu performance via large warps and two-level warp scheduling. In: Proceedings of the 44th Annual IEEE\/ACM International Symposium on Microarchitecture. MICRO-44, pp. 308\u2013317. Association for Computing Machinery, New York, NY, USA (2011). https:\/\/doi.org\/10.1145\/2155620.2155656","DOI":"10.1145\/2155620.2155656"},{"issue":"11","key":"5127_CR25","doi-asserted-by":"publisher","first-page":"3142","DOI":"10.1109\/TPDS.2017.2704080","volume":"28","author":"MK Yoon","year":"2017","unstructured":"Yoon MK, Oh Y, Kim SH, Lee S, Kim D, Ro WW (2017) Dynamic resizing on active warps scheduler to hide operation stalls on gpus. IEEE Transa Parallel Distrib Syst 28(11):3142\u20133156. https:\/\/doi.org\/10.1109\/TPDS.2017.2704080","journal-title":"IEEE Transa Parallel Distrib Syst"},{"issue":"2","key":"5127_CR26","doi-asserted-by":"publisher","first-page":"90","DOI":"10.1109\/LCA.2014.2359882","volume":"14","author":"Z Zheng","year":"2015","unstructured":"Zheng Z, Wang Z, Lipasti M (2015) Adaptive cache and concurrency allocation on gpgpus. IEEE Comput Archit Lett 14(2):90\u201393. https:\/\/doi.org\/10.1109\/LCA.2014.2359882","journal-title":"IEEE Comput Archit Lett"},{"key":"5127_CR27","doi-asserted-by":"publisher","unstructured":"Chen, X., Wu, S., Chang, L.-W., Huang, W.-S., Pearson, C., Wang, Z., Hwu, W.-M.W.: Adaptive cache bypass and insertion for many-core accelerators. In: Proceedings of International Workshop on Manycore Embedded Systems. MES \u201914, pp. 1\u20138. Association for Computing Machinery, New York, NY, USA (2014). https:\/\/doi.org\/10.1145\/2613908.2613909","DOI":"10.1145\/2613908.2613909"},{"key":"5127_CR28","doi-asserted-by":"publisher","unstructured":"Hu, B., Rossbach, C.J.: Altis: Modernizing gpgpu benchmarks. In: 2020 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS), pp. 1\u201311 (2020). https:\/\/doi.org\/10.1109\/ISPASS48437.2020.00011","DOI":"10.1109\/ISPASS48437.2020.00011"},{"issue":"12","key":"5127_CR29","doi-asserted-by":"publisher","first-page":"2865","DOI":"10.1109\/TPDS.2020.3004623","volume":"31","author":"Q Wang","year":"2020","unstructured":"Wang Q, Chu X (2020) Gpgpu performance estimation with core and memory frequency scaling. IEEE Trans Parallel Distrib Syst 31(12):2865\u20132881. https:\/\/doi.org\/10.1109\/TPDS.2020.3004623","journal-title":"IEEE Trans Parallel Distrib Syst"},{"issue":"1","key":"5127_CR30","doi-asserted-by":"publisher","first-page":"351","DOI":"10.1145\/2964791.2901468","volume":"44","author":"A Jog","year":"2016","unstructured":"Jog A, Kayiran O, Pattnaik A, Kandemir MT, Mutlu O, Iyer R, Das CR (2016) Exploiting core criticality for enhanced gpu performance. SIGMETRICS Perform Eval Rev 44(1):351\u2013363. https:\/\/doi.org\/10.1145\/2964791.2901468","journal-title":"SIGMETRICS Perform Eval Rev"},{"key":"5127_CR31","doi-asserted-by":"publisher","unstructured":"Che, S., Boyer, M., Meng, J., Tarjan, D., Sheaffer, J.W., Lee, S.-H., Skadron, K.: Rodinia: A benchmark suite for heterogeneous computing. In: 2009 IEEE International Symposium on Workload Characterization (IISWC), pp. 44\u201354 (2009). https:\/\/doi.org\/10.1109\/IISWC.2009.5306797","DOI":"10.1109\/IISWC.2009.5306797"},{"key":"5127_CR32","doi-asserted-by":"publisher","unstructured":"Che, S., Sheaffer, J.W., Boyer, M., Szafaryn, L.G., Wang, L., Skadron, K.: A characterization of the rodinia benchmark suite with comparison to contemporary cmp workloads. In: IEEE International Symposium on Workload Characterization (IISWC\u201910), pp. 1\u201311 (2010). https:\/\/doi.org\/10.1109\/IISWC.2010.5650274","DOI":"10.1109\/IISWC.2010.5650274"},{"issue":"4","key":"5127_CR33","doi-asserted-by":"publisher","first-page":"83","DOI":"10.1109\/MDAT.2020.2986738","volume":"37","author":"J Wang","year":"2020","unstructured":"Wang J, Gao J (2020) AParallelizing gpgpu-sim for faster simulation with high fidelity. IEEE Des Test 37(4):83\u201391. https:\/\/doi.org\/10.1109\/MDAT.2020.2986738","journal-title":"IEEE Des Test"},{"key":"5127_CR34","doi-asserted-by":"publisher","unstructured":"Anand, A., Thomas, W., Toraskar, S., Singh, V.: Predictive warp scheduling for efficient execution in gpgpu. In: Proceedings of the 2021 on Great Lakes Symposium on VLSI. GLSVLSI \u201921, pp. 295\u2013300. Association for Computing Machinery, New York, NY, USA (2021). https:\/\/doi.org\/10.1145\/3453688.3461525","DOI":"10.1145\/3453688.3461525"},{"key":"5127_CR35","doi-asserted-by":"publisher","unstructured":"Khairy, M., Zahran, M., Wassal, A.G. Efficient utilization of gpgpu cache hierarchy. In: Proceedings of the 8th Workshop on General Purpose Processing Using GPUs. GPGPU-8, pp. 36\u201347. Association for Computing Machinery, New York, NY, USA (2015). https:\/\/doi.org\/10.1145\/2716282.2716291","DOI":"10.1145\/2716282.2716291"},{"key":"5127_CR36","doi-asserted-by":"publisher","unstructured":"Picchi, J., Zhang, W.: Impact of l2 cache locking on gpu performance. In: SoutheastCon 2015, pp. 1\u20134 (2015). https:\/\/doi.org\/10.1109\/SECON.2015.7133036","DOI":"10.1109\/SECON.2015.7133036"},{"key":"5127_CR37","doi-asserted-by":"crossref","unstructured":"Zhao, X., Adileh, A., Yu, Z., Wang, Z., Jaleel, A., AEeckhout, AL.: Adaptive memory-side last-level gpu caching. In: 2019 ACM\/IEEE 46th Annual International Symposium on Computer Architecture (ISCA), pp. 411\u2013423 (2019)","DOI":"10.1145\/3307650.3322235"}],"container-title":["The Journal of Supercomputing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-023-05127-0.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11227-023-05127-0\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-023-05127-0.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,6,9]],"date-time":"2023-06-09T07:17:10Z","timestamp":1686295030000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11227-023-05127-0"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,3,10]]},"references-count":37,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2023,7]]}},"alternative-id":["5127"],"URL":"https:\/\/doi.org\/10.1007\/s11227-023-05127-0","relation":{},"ISSN":["0920-8542","1573-0484"],"issn-type":[{"type":"print","value":"0920-8542"},{"type":"electronic","value":"1573-0484"}],"subject":[],"published":{"date-parts":[[2023,3,10]]},"assertion":[{"value":"21 February 2023","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"10 March 2023","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"The authors readily consent to have this paper published.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"Not applicable.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent to participate"}}]}}