{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:26:54Z","timestamp":1750307214082,"version":"3.41.0"},"reference-count":49,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2012,1,1]],"date-time":"2012-01-01T00:00:00Z","timestamp":1325376000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000143","name":"Division of Computing and Communication Foundations","doi-asserted-by":"publisher","award":["CCF-0963996CNS-0810906CCF-0905509CSR-0912850"],"award-info":[{"award-number":["CCF-0963996CNS-0810906CCF-0905509CSR-0912850"]}],"id":[{"id":"10.13039\/100000143","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000144","name":"Division of Computer and Network Systems","doi-asserted-by":"publisher","award":["CCF-0963996CNS-0810906CCF-0905509CSR-0912850"],"award-info":[{"award-number":["CCF-0963996CNS-0810906CCF-0905509CSR-0912850"]}],"id":[{"id":"10.13039\/100000144","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2012,1]]},"abstract":"<jats:p>To realize the performance potential of multicore systems, we must effectively manage the interactions between memory reference behavior and the operating system policies for thread scheduling and migration decisions. We observe that these interactions lead to significant variations in the performance of a given application, from one execution to the next, even when the program input remains unchanged and no other applications are being run on the system. Our experiments with multithreaded programs, including the TATP database application, SPECjbb2005, and a subset of PARSEC and SPEC OMP programs, on a 24-core Dell PowerEdge R905 server running OpenSolaris confirms the above observation. In this work we develop Thread Tranquilizer, an automatic technique for simultaneously reducing performance variation and improving performance by dynamically choosing appropriate memory allocation and process scheduling policies. Thread Tranquilizer uses simple utilities available on modern Operating Systems for monitoring cache misses and thread context-switches and then utilizes the collected information to dynamically select appropriate memory allocation and scheduling policies. In our experiments, Thread Tranquilizer yields up to 98% (average 68%) reduction in performance variation and up to 43% (average 15%) improvement in performance over default policies of OpenSolaris. We also demonstrate that Thread Tranquilizer simultaneously reduces performance variation and improves performance of the programs on Linux. Thread Tranquilizer is easy to use as it does not require any changes to the application source code or the OS kernel.<\/jats:p>","DOI":"10.1145\/2086696.2086725","type":"journal-article","created":{"date-parts":[[2012,1,24]],"date-time":"2012-01-24T16:47:14Z","timestamp":1327423634000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":21,"title":["Thread Tranquilizer"],"prefix":"10.1145","volume":"8","author":[{"given":"Kishore Kumar","family":"Pusukuri","sequence":"first","affiliation":[{"name":"University of California, Riverside, CA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Rajiv","family":"Gupta","sequence":"additional","affiliation":[{"name":"University of California, Riverside, CA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Laxmi N.","family":"Bhuyan","sequence":"additional","affiliation":[{"name":"University of California, Riverside, CA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2012,1,26]]},"reference":[{"volume-title":"Proceedings of the 9th International Symposium on High-Performance Computer Architecture (HPCA'03)","author":"Alameldeen A. R.","key":"e_1_2_1_1_1","unstructured":"Alameldeen , A. R. and Wood , D. A . 2003. Variability in architectural simulations of multi-threaded workloads . In Proceedings of the 9th International Symposium on High-Performance Computer Architecture (HPCA'03) . IEEE Computer Society, Los Alamitos, CA, 7--22. Alameldeen, A. R. and Wood, D. A. 2003. Variability in architectural simulations of multi-threaded workloads. In Proceedings of the 9th International Symposium on High-Performance Computer Architecture (HPCA'03). IEEE Computer Society, Los Alamitos, CA, 7--22."},{"key":"e_1_2_1_2_1","unstructured":"Attardi J. and Nadgir N. 2003. A comparison of memory allocators in multiprocessors. http:\/\/developers. sun.com\/solaris\/articles\/multiproc\/multiproc.html.  Attardi J. and Nadgir N. 2003. A comparison of memory allocators in multiprocessors. http:\/\/developers. sun.com\/solaris\/articles\/multiproc\/multiproc.html."},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/1128022.1128029"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10586-007-0047-2"},{"key":"e_1_2_1_5_1","unstructured":"Benson R. 2003. Identifying memory management bugs within applications using the libumem library. http:\/\/developers.sun.com\/solaris\/articles\/libumem_library.html.  Benson R. 2003. Identifying memory management bugs within applications using the libumem library. http:\/\/developers.sun.com\/solaris\/articles\/libumem_library.html."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/378993.379232"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/1454115.1454128"},{"volume-title":"Proceedings of the USENIX Conference on USENIX Annual Technical Conference (USENIXATC'11)","author":"Blagodurov S.","key":"e_1_2_1_8_1","unstructured":"Blagodurov , S. , Zhuravlev , S. , Dashti , M. , and Fedorova , A . 2011. A case for numa-aware contention management on multicore systems . In Proceedings of the USENIX Conference on USENIX Annual Technical Conference (USENIXATC'11) . USENIX Association, Berkeley, CA, 1--1. Blagodurov, S., Zhuravlev, S., Dashti, M., and Fedorova, A. 2011. A case for numa-aware contention management on multicore systems. In Proceedings of the USENIX Conference on USENIX Annual Technical Conference (USENIXATC'11). USENIX Association, Berkeley, CA, 1--1."},{"volume-title":"Proceedings of the 9th USENIX Symposium on Operating Systems Design and Implementation (OSDI'10)","author":"Boyd-Wickizer S.","key":"e_1_2_1_9_1","unstructured":"Boyd-Wickizer , S. , Clements , A. , Mao , Y. , Pesterev , A. , Kaashoek , M. F. , Morris , R. , and Zeldovich , N . 2010. An analysis of linux scalability to many cores . In Proceedings of the 9th USENIX Symposium on Operating Systems Design and Implementation (OSDI'10) . Boyd-Wickizer, S., Clements, A., Mao, Y., Pesterev, A., Kaashoek, M. F., Morris, R., and Zeldovich, N. 2010. An analysis of linux scalability to many cores. In Proceedings of the 9th USENIX Symposium on Operating Systems Design and Implementation (OSDI'10)."},{"volume-title":"Proceedings of the Annual Conference on USENIX Annual Technical Conference (ATEC'04)","author":"Cantrill B. M.","key":"e_1_2_1_10_1","unstructured":"Cantrill , B. M. , Shapiro , M. W. , and Leventhal , A. H . 2004. Dynamic instrumentation of production systems . In Proceedings of the Annual Conference on USENIX Annual Technical Conference (ATEC'04) . USENIX Association, Berkeley, CA, 2--2. Cantrill, B. M., Shapiro, M. W., and Leventhal, A. H. 2004. Dynamic instrumentation of production systems. In Proceedings of the Annual Conference on USENIX Annual Technical Conference (ATEC'04). USENIX Association, Berkeley, CA, 2--2."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/1105734.1105745"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/CLUSTR.2007.4629247"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/1870109.1870115"},{"volume-title":"Proceedings of the Annual IEEE Conference on Local Computer Networks. 538","author":"Evans J. J.","key":"e_1_2_1_14_1","unstructured":"Evans , J. J. , Hood , C. S. , and Gropp , W. D . 2003. Exploring the relationship between parallel application run-time variability and network performance in clusters . In Proceedings of the Annual IEEE Conference on Local Computer Networks. 538 . Evans, J. J., Hood, C. S., and Gropp, W. D. 2003. Exploring the relationship between parallel application run-time variability and network performance in clusters. In Proceedings of the Annual IEEE Conference on Local Computer Networks. 538."},{"volume-title":"Proceedings of the ACM\/IEEE Conference on Supercomputing (SC'08)","author":"Ferreira K. B.","key":"e_1_2_1_15_1","unstructured":"Ferreira , K. B. , Bridges , P. , and Brightwell , R . 2008. Characterizing application sensitivity to os interference using kernel-level noise injection . In Proceedings of the ACM\/IEEE Conference on Supercomputing (SC'08) . IEEE Press, Los Alamitos, CA, 19:1--19:12. Ferreira, K. B., Bridges, P., and Brightwell, R. 2008. Characterizing application sensitivity to os interference using kernel-level noise injection. In Proceedings of the ACM\/IEEE Conference on Supercomputing (SC'08). IEEE Press, Los Alamitos, CA, 19:1--19:12."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/CLUSTER.2010.16"},{"volume-title":"Proceedings of the Component and Middleware Performance Workshop OOPSLA","author":"Gu D.","key":"e_1_2_1_17_1","unstructured":"Gu , D. , Verbrugge , C. , and Gagnon , E . 2004. Code layout as a source of noise in jvm performance . In Proceedings of the Component and Middleware Performance Workshop OOPSLA , Gu, D., Verbrugge, C., and Gagnon, E. 2004. Code layout as a source of noise in jvm performance. In Proceedings of the Component and Middleware Performance Workshop OOPSLA,"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/107972.107985"},{"key":"e_1_2_1_19_1","unstructured":"Hensley J. Alter R. Duffy D. Fahey M. Higbie L. Oppe T. Ward W. Bullock M. and Becklehimer J. 2001. Minimizing runtime performance variation with cpusets on the SGI origin 3800. Tech. rep. 01-32 ERDC MSRC.  Hensley J. Alter R. Duffy D. Fahey M. Higbie L. Oppe T. Ward W. Bullock M. and Becklehimer J. 2001. Minimizing runtime performance variation with cpusets on the SGI origin 3800. Tech. rep. 01-32 ERDC MSRC."},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/1712605.1712640"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1008165010732"},{"volume-title":"Proceedings of the International Conference on Computational Science Part III (ICCS'03)","author":"Kramer W. T. C.","key":"e_1_2_1_22_1","unstructured":"Kramer , W. T. C. and Ryan , C . 2003. Performance variability of highly parallel architectures . In Proceedings of the International Conference on Computational Science Part III (ICCS'03) . Springer-Verlag, Berlin, 560--569. Kramer, W. T. C. and Ryan, C. 2003. Performance variability of highly parallel architectures. In Proceedings of the International Conference on Computational Science Part III (ICCS'03). Springer-Verlag, Berlin, 560--569."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2010.32"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2009.5161046"},{"volume-title":"Proceedings of the International Symposium on Performance Analysis of Systems and Software. 87--96","author":"McCurdy C.","key":"e_1_2_1_25_1","unstructured":"McCurdy , C. and Vetter , J. S . 2010. Memphis: Finding and fixing numa-related performance problems on multi-core platforms . In Proceedings of the International Symposium on Performance Analysis of Systems and Software. 87--96 . McCurdy, C. and Vetter, J. S. 2010. Memphis: Finding and fixing numa-related performance problems on multi-core platforms. In Proceedings of the International Symposium on Performance Analysis of Systems and Software. 87--96."},{"key":"e_1_2_1_26_1","unstructured":"McDougall R. and Mauro J. 2006. Solaris Internals 2nd Ed. Prentice Hall.   McDougall R. and Mauro J. 2006. Solaris Internals 2nd Ed. Prentice Hall."},{"key":"e_1_2_1_27_1","unstructured":"McDougall R. Mauro J. and Gregg B. 2006. Solaris Performance and Tools: DTrace and MDB Techniques for Solaris 10 and OpenSolaris. Prentice Hall.   McDougall R. Mauro J. and Gregg B. 2006. Solaris Performance and Tools: DTrace and MDB Techniques for Solaris 10 and OpenSolaris. Prentice Hall."},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/1755913.1755930"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.5555\/602770.602873"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/1362622.1362662"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/1048935.1050204"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1007\/11823285_27"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC.2011.6114208"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/2016604.2016647"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/1555349.1555384"},{"volume-title":"Proceedings of the IEEE International Symposium on Parallel & Distributed Processing. 1--12","author":"Seelam S.","key":"e_1_2_1_36_1","unstructured":"Seelam , S. , Fong , L. , Tantawi , A. , Lewars , J. , and J. Divirgilio , K. G. 2010. Extreme scale computing: Modeling the impact of system noise in multicore clustered systems . In Proceedings of the IEEE International Symposium on Parallel & Distributed Processing. 1--12 . Seelam, S., Fong, L., Tantawi, A., Lewars, J., and J. Divirgilio, K. G. 2010. Extreme scale computing: Modeling the impact of system noise in multicore clustered systems. In Proceedings of the IEEE International Symposium on Parallel & Distributed Processing. 1--12."},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/1736020.1736034"},{"volume-title":"Proceedings of the IEEE Workload Characterization Symposium, 137--149","author":"Skinner D.","key":"e_1_2_1_38_1","unstructured":"Skinner , D. and Kramer , W . 2005. Understanding the causes of performance variability in HPC workloads . In Proceedings of the IEEE Workload Characterization Symposium, 137--149 . Skinner, D. and Kramer, W. 2005. Understanding the causes of performance variability in HPC workloads. In Proceedings of the IEEE Workload Characterization Symposium, 137--149."},{"key":"e_1_2_1_39_1","unstructured":"solidDB. 2010. IBM soliddb 6.5 (build 2010-10-04). http:\/\/www-01.ibm.com\/software\/data\/soliddb\/soliddb\/.  solidDB. 2010. IBM soliddb 6.5 (build 2010-10-04). http:\/\/www-01.ibm.com\/software\/data\/soliddb\/soliddb\/."},{"key":"e_1_2_1_40_1","unstructured":"SPECjbb. 2005. http:\/\/www.spec.org\/jbb2005.  SPECjbb. 2005. http:\/\/www.spec.org\/jbb2005."},{"key":"e_1_2_1_41_1","unstructured":"SPECOMP. 2001. http:\/\/www.spec.org\/omp.  SPECOMP. 2001. http:\/\/www.spec.org\/omp."},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2005.52"},{"key":"e_1_2_1_43_1","unstructured":"TATP. 2009. IBM telecom application transaction processing benchmark description. http:\/\/tatpbench-mark.sourceforge.net.  TATP. 2009. IBM telecom application transaction processing benchmark description. http:\/\/tatpbench-mark.sourceforge.net."},{"volume-title":"Proceedings of the Performance Analysis of Systems and Software (ISPASSS'09)","author":"Teng Q.","key":"e_1_2_1_44_1","unstructured":"Teng , Q. , Sweeney , P. F. , and Duesterwald , E . 2009. Understanding the cost of thread migration for multi-threaded java applications running on multicore platform . In Proceedings of the Performance Analysis of Systems and Software (ISPASSS'09) . 123--132. Teng, Q., Sweeney, P. F., and Duesterwald, E. 2009. Understanding the cost of thread migration for multi-threaded java applications running on multicore platform. In Proceedings of the Performance Analysis of Systems and Software (ISPASSS'09). 123--132."},{"key":"e_1_2_1_45_1","unstructured":"Touati S.-A.-A. and Worms J. S. B. 2010. The speed test. Tech. rep. http:\/\/hal.archives-ouvertes.fr\/inria-00443839.  Touati S.-A.-A. and Worms J. S. B. 2010. The speed test. Tech. rep. http:\/\/hal.archives-ouvertes.fr\/inria-00443839."},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/1088149.1088190"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/237090.237205"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCMP-UGC.2009.72"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/1736020.1736036"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2086696.2086725","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2086696.2086725","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T10:06:43Z","timestamp":1750241203000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2086696.2086725"}},"subtitle":["Dynamically reducing performance variation"],"short-title":[],"issued":{"date-parts":[[2012,1]]},"references-count":49,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2012,1]]}},"alternative-id":["10.1145\/2086696.2086725"],"URL":"https:\/\/doi.org\/10.1145\/2086696.2086725","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"type":"print","value":"1544-3566"},{"type":"electronic","value":"1544-3973"}],"subject":[],"published":{"date-parts":[[2012,1]]},"assertion":[{"value":"2011-07-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2011-11-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2012-01-26","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}