{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T23:01:43Z","timestamp":1777676503610,"version":"3.51.4"},"reference-count":22,"publisher":"SAGE Publications","issue":"2","license":[{"start":{"date-parts":[[2013,4,2]],"date-time":"2013-04-02T00:00:00Z","timestamp":1364860800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["The International Journal of High Performance Computing Applications"],"published-print":{"date-parts":[[2013,5]]},"abstract":"<jats:p>The execution of a single process multiple data (SPMD) application involves running multiple instances of a process with possibly varying arguments. With the widespread adoption of massively multicore processors, there has been a focus towards harnessing the abundant compute resources effectively in a power-efficient manner. Although much work has been done towards optimizing distributed process launch using hierarchical techniques, there has been a void in studying the performance of spawning processes within a single node. Reducing the latency to spawn a new process locally results in faster global job launch. Further, emerging dynamic and resilient execution models are designed on the premise of maintaining process pools for fault isolation and launching several processes in a relatively shorter period of time. Optimizing the latency and throughput for spawning processes would help improve the overall performance of runtime systems, allow adaptive process-replication reliability and motivate the design and implementation of process management interfaces in future manycore operating systems. In this paper, we study the several limiting factors for efficient spawning of processes on massively multicore architectures. We have developed a library to optimize launching multiple instances of the same executable. Our microbenchmarks show a 20\u201380% decrease in the process spawn time for multiple executables. We further discuss the effects of memory locality and propose NUMA-aware extensions to optimize launching processes with large memory-mapped segments including dynamic shared libraries. Finally, we describe vector operating system interfaces for spawning a batch of processes from a given executable on specific cores. Our results show a speedup of a factor of 40\u201350 over the traditional method of launching new processes using fork and exec system calls.<\/jats:p>","DOI":"10.1177\/1094342013481483","type":"journal-article","created":{"date-parts":[[2013,4,4]],"date-time":"2013-04-04T02:19:14Z","timestamp":1365041954000},"page":"147-161","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":0,"title":["Optimizing process creation and execution on multi-core architectures"],"prefix":"10.1177","volume":"27","author":[{"given":"Abhishek","family":"Kulkarni","sequence":"first","affiliation":[{"name":"Center for Research in Extreme Scale Technologies, Department of Computer Science, Indiana University, Bloomington, IN, USA"},{"name":"Ultrascale Systems Research Center, Los Alamos National Laboratory, Los Alamos, NM, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Latchesar","family":"Ionkov","sequence":"additional","affiliation":[{"name":"Ultrascale Systems Research Center, Los Alamos National Laboratory, Los Alamos, NM, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Michael","family":"Lang","sequence":"additional","affiliation":[{"name":"Ultrascale Systems Research Center, Los Alamos National Laboratory, Los Alamos, NM, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Andrew","family":"Lumsdaine","sequence":"additional","affiliation":[{"name":"Center for Research in Extreme Scale Technologies, Department of Computer Science, Indiana University, Bloomington, IN, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[2013,4,2]]},"reference":[{"key":"bibr1-1094342013481483","doi-asserted-by":"publisher","DOI":"10.1109\/ICPP.2008.63"},{"key":"bibr2-1094342013481483","first-page":"1","volume-title":"Proceedings of the 9th USENIX conference on Operating Systems Design and Implementation (OSDI\u201910)","author":"Boyd-Wickizer S","year":"2010"},{"key":"bibr3-1094342013481483","doi-asserted-by":"crossref","first-page":"214","DOI":"10.1145\/258612.258690","author":"Brown AB","year":"1997","journal-title":"ACM SIGMETRICS Conference on Measurement and Modeling of Computer Systems"},{"key":"bibr4-1094342013481483","first-page":"313","author":"Chen Y","year":"2009","journal-title":"2009 2nd IEEE International Conference on Computer Science and Information Technology"},{"key":"bibr5-1094342013481483","first-page":"117","author":"Cui Y","year":"2010","journal-title":"2010 IEEE International Symposium on Performance Analysis of Systems Software (ISPASS)"},{"key":"bibr6-1094342013481483","unstructured":"Ehringer D (2010) The Dalvik Virtual Machine architecture. http:\/\/davidehringer.com\/software\/android\/The_Dalvik_Virtual_Machine.pdf"},{"key":"bibr7-1094342013481483","doi-asserted-by":"publisher","DOI":"10.1145\/2063384.2063443"},{"key":"bibr8-1094342013481483","first-page":"131","author":"Gingell R","year":"1987","journal-title":"Proceedings of the Summer 1987 USENIX Technical Conference"},{"key":"bibr9-1094342013481483","author":"Goehner JD","year":"2012","journal-title":"Parallel Computing"},{"key":"bibr10-1094342013481483","first-page":"44","volume-title":"Proceedings of Job Scheduling Strategies for Parallel Processing (JSSPP) 2003 ( Lecture Notes in Computer Science","volume":"2862","author":"Jette MA","year":"2003"},{"key":"bibr11-1094342013481483","first-page":"37","volume-title":"Proceedings of the 2012 USENIX Conference on Annual Technical Conference (USENIX ATC\u201912)","author":"Kato S","year":"2012"},{"key":"bibr12-1094342013481483","volume-title":"Exascale Computing Study: Technology Challenges in Achieving Exascale Systems","author":"Kogge P","year":"2008"},{"key":"bibr13-1094342013481483","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC.2007.4362185"},{"key":"bibr14-1094342013481483","volume-title":"Measuring and Improving the Performance of Berkeley UNIX","author":"McKusick MK","year":"1991"},{"key":"bibr15-1094342013481483","doi-asserted-by":"publisher","DOI":"10.1145\/1629575.1629597"},{"key":"bibr16-1094342013481483","volume-title":"Proceedings of the 24th International Workshop on Languages and Compilers for Parallel Computing (LCPC 2011)","author":"Orozco D","year":"2011"},{"key":"bibr17-1094342013481483","unstructured":"Pike R, Presotto D, Thompson K, Trickey H (1990) Plan 9 from Bell Labs. EUUG Newsletter 10(3): 2\u201311. http:\/\/plan9.bell-labs.com\/cm\/cs\/cstr\/158b.ps.gz."},{"key":"bibr18-1094342013481483","doi-asserted-by":"publisher","DOI":"10.1145\/2043556.2043579"},{"key":"bibr19-1094342013481483","first-page":"255","volume":"1","author":"Smith JM","year":"1988","journal-title":"Computing Systems"},{"key":"bibr20-1094342013481483","first-page":"1","volume-title":"Proceedings of the 9th USENIX Conference on Operating Systems Design and Implementation (OSDI\u201910)","author":"Soares L","year":"2010"},{"key":"bibr21-1094342013481483","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-89894-8_30"},{"key":"bibr22-1094342013481483","first-page":"31","volume-title":"Proceedings of the 13th USENIX Conference on Hot Topics in Operating Systems (HotOS\u201913)","author":"Vasudevan V","year":"2011"}],"container-title":["The International Journal of High Performance Computing Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1094342013481483","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/1094342013481483","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1094342013481483","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T08:19:11Z","timestamp":1777450751000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/1094342013481483"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2013,4,2]]},"references-count":22,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2013,5]]}},"alternative-id":["10.1177\/1094342013481483"],"URL":"https:\/\/doi.org\/10.1177\/1094342013481483","relation":{},"ISSN":["1094-3420","1741-2846"],"issn-type":[{"value":"1094-3420","type":"print"},{"value":"1741-2846","type":"electronic"}],"subject":[],"published":{"date-parts":[[2013,4,2]]}}}