{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,27]],"date-time":"2026-03-27T02:36:13Z","timestamp":1774578973796,"version":"3.50.1"},"publisher-location":"New York, NY, USA","reference-count":39,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,6,3]],"date-time":"2021-06-03T00:00:00Z","timestamp":1622678400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"U.S. Department of Energy Office of Science and the National Nuclear Security Administration ? Exascale Computing Project","award":["DE-AC05-00OR22725"],"award-info":[{"award-number":["DE-AC05-00OR22725"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,6,3]]},"DOI":"10.1145\/3447818.3461616","type":"proceedings-article","created":{"date-parts":[[2021,6,4]],"date-time":"2021-06-04T15:09:36Z","timestamp":1622819376000},"page":"88-101","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["Task-graph scheduling extensions for efficient synchronization and communication"],"prefix":"10.1145","author":[{"given":"Seonmyeong","family":"Bak","sequence":"first","affiliation":[{"name":"Georgia Institute of Technology"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Oscar","family":"Hernandez","sequence":"additional","affiliation":[{"name":"Oak Ridge National Laboratory"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mark","family":"Gates","sequence":"additional","affiliation":[{"name":"University of Tennessee"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Piotr","family":"Luszczek","sequence":"additional","affiliation":[{"name":"University of Tennessee"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Vivek","family":"Sarkar","sequence":"additional","affiliation":[{"name":"Georgia Institute of Technology"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,6,4]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/341800.341801"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2014.58"},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2017.2766064"},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/146941.146944"},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/CCGRID.2018.00018"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/1639950.1639989"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.5555\/2388996.2389086"},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/324133.324234"},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/MCSE.2013.98"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2016.49"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1177\/1094342007078442"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/1094811.1094852"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/2597652.2597665"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"crossref","unstructured":"J. Choi J. Demmel I. Dhillon J. Dongarra S. Ostrouchov A. Petitet K. Stanley D. Walker and R. C. Whaley. 1996. ScaLAPACK: A portable linear algebra library for distributed memory computers --- Design issues and performance. In Applied Parallel Computing Computations in Physics Chemistry and Engineering Science Jack Dongarra Kaj Madsen and Jerzy Wa\u015bniewski (Eds.). Springer Berlin Heidelberg Berlin Heidelberg 95--106.  J. Choi J. Demmel I. Dhillon J. Dongarra S. Ostrouchov A. Petitet K. Stanley D. Walker and R. C. Whaley. 1996. ScaLAPACK: A portable linear algebra library for distributed memory computers --- Design issues and performance. In Applied Parallel Computing Computations in Physics Chemistry and Engineering Science Jack Dongarra Kaj Madsen and Jerzy Wa\u015bniewski (Eds.). Springer Berlin Heidelberg Berlin Heidelberg 95--106.","DOI":"10.1007\/3-540-60902-4_12"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/2751205.2751235"},{"key":"e_1_3_2_1_16_1","volume-title":"Proceedings of the Conference on High Performance Computing Networking, Storage and Analysis. 1--11","author":"Dinan J.","unstructured":"J. Dinan , D. B. Larkins , P. Sadayappan , S. Krishnamoorthy , and J. Nieplocha . 2009. Scalable work stealing . In Proceedings of the Conference on High Performance Computing Networking, Storage and Analysis. 1--11 . J. Dinan, D. B. Larkins, P. Sadayappan, S. Krishnamoorthy, and J. Nieplocha. 2009. Scalable work stealing. In Proceedings of the Conference on High Performance Computing Networking, Storage and Analysis. 1--11."},{"key":"e_1_3_2_1_17_1","volume-title":"Kernighan","author":"Donovan Alan A.A.","year":"2015","unstructured":"Alan A.A. Donovan and Brian W . Kernighan . 2015 . The Go Programming Language (1st ed.). Addison-Wesley Professional . Alan A.A. Donovan and Brian W. Kernighan. 2015. The Go Programming Language (1st ed.). Addison-Wesley Professional."},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1007\/3-540-39997-6"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1016\/0743-7315(92)90014-E"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3295500.3356223"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/1693453.1693504"},{"key":"e_1_3_2_1_22_1","volume-title":"2010 IEEE International Symposium on Parallel Distributed Processing (IPDPS). 1--11","author":"Iancu C.","unstructured":"C. Iancu , S. Hofmeyr , F. Blagojevi\u0107 , and Y. Zheng . 2010. Oversubscription on multicore processors . In 2010 IEEE International Symposium on Parallel Distributed Processing (IPDPS). 1--11 . C. Iancu, S. Hofmeyr, F. Blagojevi\u0107, and Y. Zheng. 2010. Oversubscription on multicore processors. In 2010 IEEE International Symposium on Parallel Distributed Processing (IPDPS). 1--11."},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/PACT.2019.00011"},{"key":"e_1_3_2_1_24_1","volume-title":"The C++ Standard Library: A Tutorial and Reference","author":"Josuttis Nicolai M.","unstructured":"Nicolai M. Josuttis . 2012. The C++ Standard Library: A Tutorial and Reference ( 2 nd ed.). Addison-Wesley Professional . Nicolai M. Josuttis. 2012. The C++ Standard Library: A Tutorial and Reference (2nd ed.). Addison-Wesley Professional.","edition":"2"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/2676870.2676883"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPPS.1996.508060"},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/IA3.2016.019"},{"key":"e_1_3_2_1_28_1","volume-title":"Euro-Par 2019: Parallel Processing","author":"Kurzak Jakub","unstructured":"Jakub Kurzak , Mark Gates , Ali Charara , Asim YarKhan , Ichitaro Yamazaki , and Jack Dongarra . 2019. Linear Systems Solvers for Distributed-Memory Machines with GPU Accelerators . In Euro-Par 2019: Parallel Processing , Ramin Yahyapour (Ed.). Springer International Publishing , Cham , 495--506. Jakub Kurzak, Mark Gates, Ali Charara, Asim YarKhan, Ichitaro Yamazaki, and Jack Dongarra. 2019. Linear Systems Solvers for Distributed-Memory Machines with GPU Accelerators. In Euro-Par 2019: Parallel Processing, Ramin Yahyapour (Ed.). Springer International Publishing, Cham, 495--506."},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/2287076.2287103"},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"crossref","unstructured":"J. Lifflander S. Krishnamoorthy and L. V. Kale. 2014. Optimizing Data Locality for Fork\/Join Programs Using Constrained Work Stealing. In SC '14: Proceedings of the International Conference for High Performance Computing Networking Storage and Analysis. 857--868.  J. Lifflander S. Krishnamoorthy and L. V. Kale. 2014. Optimizing Data Locality for Fork\/Join Programs Using Constrained Work Stealing. In SC '14: Proceedings of the International Conference for High Performance Computing Networking Storage and Analysis . 857--868.","DOI":"10.1109\/SC.2014.75"},{"key":"e_1_3_2_1_32_1","volume-title":"Shenango: Achieving High CPU Efficiency for Latency-sensitive Datacenter Workloads. In 16th USENIX Symposium on Networked Systems Design and Implementation (NSDI 19)","author":"Ousterhout Amy","year":"2019","unstructured":"Amy Ousterhout , Joshua Fried , Jonathan Behrens , Adam Belay , and Hari Balakrishnan . 2019 . Shenango: Achieving High CPU Efficiency for Latency-sensitive Datacenter Workloads. In 16th USENIX Symposium on Networked Systems Design and Implementation (NSDI 19) . USENIX Association, Boston, MA, 361--378. https:\/\/www.usenix.org\/conference\/nsdi19\/presentation\/ousterhout Amy Ousterhout, Joshua Fried, Jonathan Behrens, Adam Belay, and Hari Balakrishnan. 2019. Shenango: Achieving High CPU Efficiency for Latency-sensitive Datacenter Workloads. In 16th USENIX Symposium on Networked Systems Design and Implementation (NSDI 19). USENIX Association, Boston, MA, 361--378. https:\/\/www.usenix.org\/conference\/nsdi19\/presentation\/ousterhout"},{"key":"e_1_3_2_1_33_1","volume-title":"Proceedings of the 3rd International Conference on Distributed Computing Systems, Miami\/Ft","author":"Ousterhout John K.","year":"1982","unstructured":"John K. Ousterhout . 1982 . Scheduling Techniques for Concurrent Systems . In Proceedings of the 3rd International Conference on Distributed Computing Systems, Miami\/Ft . Lauderdale, Florida, USA , October 18-22, 1982. IEEE Computer Society, 22--30. John K. Ousterhout. 1982. Scheduling Techniques for Concurrent Systems. In Proceedings of the 3rd International Conference on Distributed Computing Systems, Miami\/Ft. Lauderdale, Florida, USA, October 18-22, 1982. IEEE Computer Society, 22--30."},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/1806596.1806639"},{"key":"e_1_3_2_1_35_1","volume-title":"Arachne: Core-Aware Thread Management. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18)","author":"Qin Henry","year":"2018","unstructured":"Henry Qin , Qian Li , Jacqueline Speiser , Peter Kraft , and John Ousterhout . 2018 . Arachne: Core-Aware Thread Management. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18) . USENIX Association, Carlsbad, CA, 145--160. https:\/\/www.usenix.org\/conference\/osdi18\/presentation\/qin Henry Qin, Qian Li, Jacqueline Speiser, Peter Kraft, and John Ousterhout. 2018. Arachne: Core-Aware Thread Management. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18). USENIX Association, Carlsbad, CA, 145--160. https:\/\/www.usenix.org\/conference\/osdi18\/presentation\/qin"},{"key":"e_1_3_2_1_36_1","volume-title":"Euro-Par 2010 - Parallel Processing, Pasqua D'Ambra","author":"Quintin Jean-No\u00ebl","unstructured":"Jean-No\u00ebl Quintin and Fr\u00e9d\u00e9ric Wagner . 2010. Hierarchical Work-Stealing . In Euro-Par 2010 - Parallel Processing, Pasqua D'Ambra , Mario Guarracino, and Domenico Talia (Eds.). Springer Berlin Heidelberg, Berlin , Heidelberg , 217--229. Jean-No\u00ebl Quintin and Fr\u00e9d\u00e9ric Wagner. 2010. Hierarchical Work-Stealing. In Euro-Par 2010 - Parallel Processing, Pasqua D'Ambra, Mario Guarracino, and Domenico Talia (Eds.). Springer Berlin Heidelberg, Berlin, Heidelberg, 217--229."},{"key":"e_1_3_2_1_37_1","volume-title":"Euro-Par 2019: Parallel Processing","author":"Richard J\u00e9r\u00f4me","unstructured":"J\u00e9r\u00f4me Richard , Guillaume Latu , Julien Bigot , and Thierry Gautier . 2019. Fine-Grained MPI+OpenMP Plasma Simulations : Communication Overlap with Dependent Tasks . In Euro-Par 2019: Parallel Processing , Ramin Yahyapour (Ed.). Springer International Publishing , Cham , 419--433. J\u00e9r\u00f4me Richard, Guillaume Latu, Julien Bigot, and Thierry Gautier. 2019. Fine-Grained MPI+OpenMP Plasma Simulations: Communication Overlap with Dependent Tasks. In Euro-Par 2019: Parallel Processing, Ramin Yahyapour (Ed.). Springer International Publishing, Cham, 419--433."},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2017.2766062"},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2008.4536359"},{"key":"e_1_3_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2003.1206505"}],"event":{"name":"ICS '21: 2021 International Conference on Supercomputing","location":"Virtual Event USA","acronym":"ICS '21","sponsor":["SIGARCH ACM Special Interest Group on Computer Architecture"]},"container-title":["Proceedings of the ACM International Conference on Supercomputing"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3447818.3461616","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3447818.3461616","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T17:49:27Z","timestamp":1750268967000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3447818.3461616"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,6,3]]},"references-count":39,"alternative-id":["10.1145\/3447818.3461616","10.1145\/3447818"],"URL":"https:\/\/doi.org\/10.1145\/3447818.3461616","relation":{},"subject":[],"published":{"date-parts":[[2021,6,3]]},"assertion":[{"value":"2021-06-04","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}