{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:10:35Z","timestamp":1750219835775,"version":"3.41.0"},"reference-count":83,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2023,2,27]],"date-time":"2023-02-27T00:00:00Z","timestamp":1677456000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Meas. Anal. Comput. Syst."],"published-print":{"date-parts":[[2023,2,27]]},"abstract":"<jats:p>A diverse set of scheduling objectives (e.g., resource contention, fairness, priority, etc.) breed a series of objective-specific schedulers for multi-core architectures. Existing designs incorporate thread-to-thread statistics at runtime, and schedule threads based on such an abstraction (we formalize thread-to-thread interaction as the Thread-Interaction Matrix). However, such an abstraction also reveals a consistently-overlooked issue: the Thread-Interaction Matrix (TIM) is highly sparse. Therefore, existing designs can only deliver sub-optimal decisions, since the sparsity issue limits the amount of thread permutations (and its statistics) to be exploited when performing scheduling decisions.<\/jats:p>\n          <jats:p>We introduce Sparsity-Lightened Intelligent Thread Scheduling (SLITS), a general scheduler design for mitigating the sparsity issue of TIM, with the customizability for different scheduling objectives. SLITS is designed upon the key insight that: the sparsity issue of the TIM can be effectively mitigated via advanced Machine Learning (ML) techniques. SLITS has three components. First, SLITS profiles Thread Interactions for only a small number of thread permutations, and form the TIM using the run-time statistics. Second, SLITS estimates the missing values in the TIM using Factorization Machine (FM), a novel ML technique that can fill in the missing values within a large-scale sparse matrix based on the limited information. Third, SLITS leverages Lazy Reschedule, a general mechanism as the building block for customizing different scheduling policies for different scheduling objectives. We show how SLITS can be (1) customized for different scheduling objectives, including resource contention and fairness; and (2) implemented with only negligible hardware costs. We also discuss how SLITS can be potentially applied to other contexts of thread scheduling.<\/jats:p>\n          <jats:p>We evaluate two SLITS variants against four state-of-the-art scheduler designs. We highlight that, averaged across 11 benchmarks, SLITS achieves an average speedup of 1.08X over the de facto standard for thread scheduler - the Completely Fair Scheduler, under the 16-core setting for a variety of number of threads (i.e., 32, 64 and 128). Our analysis reveals that the benefits of SLITS are credited to significant improvements of cache utilization. In addition, our experimental results confirm that SLITS is scalable and the benefits are robust when of the number of threads increases. We also perform extensive studies to (1) break down SLITS components to justify the synergy of our design choices, (2) examine the impacts of varying the estimation coverage of FM, (3) justify the necessity of Lazy Reschedule rather than periodic rescheduling, and (4) demonstrate the hardware overheads for SLITS implementations can be marginal (&lt;1% chip area and power).<\/jats:p>","DOI":"10.1145\/3579436","type":"journal-article","created":{"date-parts":[[2023,3,2]],"date-time":"2023-03-02T23:50:57Z","timestamp":1677801057000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["SLITS: Sparsity-Lightened Intelligent Thread Scheduling"],"prefix":"10.1145","volume":"7","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1607-2781","authenticated-orcid":false,"given":"Wangkai","family":"Jin","sequence":"first","affiliation":[{"name":"Succincter, Hangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4442-9744","authenticated-orcid":false,"given":"Xiangjun","family":"Peng","sequence":"additional","affiliation":[{"name":"Succincter, Hangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,3,2]]},"reference":[{"unstructured":"\"Intel Xeon Gold 6150 \"https:\/\/en.wikichip.org\/wiki\/intel\/xeon_gold\/6150.","key":"e_1_2_1_1_1"},{"unstructured":"\"Intel Xeon Platinum 8180M \"https:\/\/en.wikichip.org\/wiki\/intel\/xeon_platinum\/8180m.","key":"e_1_2_1_2_1"},{"unstructured":"Xlearn. [Online]. Available: https:\/\/github.com\/aksnzhy\/xlearn","key":"e_1_2_1_3_1"},{"unstructured":"\"Chapter 6 - Performance \"in Principles of Computer System Design 2009.","key":"e_1_2_1_4_1"},{"key":"e_1_2_1_5_1","author":"Akturk I.","year":"2019","unstructured":"I. Akturk and O. Ozturk,\"Adaptive Thread Scheduling in Chip Multiprocessors,\" International Journal of Parallel Programming, 2019.","journal-title":"\"Adaptive Thread Scheduling in Chip Multiprocessors,\" International Journal of Parallel Programming"},{"key":"e_1_2_1_6_1","volume-title":"Multitask Implementation of Synchronous Reactive Models with Earliest Deadline First Scheduling,\"in SIES","author":"Z.","year":"2013","unstructured":"Z. Al-bayati, H. Zeng, M. Di Natale, and Z. Gu, \"Multitask Implementation of Synchronous Reactive Models with Earliest Deadline First Scheduling,\"in SIES, 2013."},{"doi-asserted-by":"publisher","key":"e_1_2_1_7_1","DOI":"10.1609\/aimag.v32i3.2368"},{"unstructured":"AMD \"BIOS and Kernel Developer's Guide for AMD Family 15h processors \"2013. [Online]. Available: https:\/\/www.amd.com\/system\/files\/TechDocs\/42301_15h_Mod_00h-0Fh_BKDG.pdf","key":"e_1_2_1_8_1"},{"key":"e_1_2_1_9_1","volume-title":"Parallel Real-Time Task Scheduling on Multicore Platforms,\"in RTSS","author":"Anderson J. H.","year":"2006","unstructured":"J. H. Anderson and J. M. Calandrino, \"Parallel Real-Time Task Scheduling on Multicore Platforms,\"in RTSS, 2006."},{"key":"e_1_2_1_10_1","volume-title":"Pfair Scheduling: Beyond Periodic Task Systems,\"in Proceedings Seventh International Conference on Real-Time Computing Systems and Applications","author":"Anderson J. H.","year":"2000","unstructured":"J. H. Anderson and A. Srinivasan,\"Pfair Scheduling: Beyond Periodic Task Systems,\"in Proceedings Seventh International Conference on Real-Time Computing Systems and Applications, 2000."},{"key":"e_1_2_1_11_1","author":"Anderson J.","year":"2004","unstructured":"J. Anderson and A. Srinivasan, \"Mixed Pfair\/ERfair Scheduling of Asynchronous Periodic Tasks,\"Journal of Computer and System Sciences, 2004.","journal-title":"\"Mixed Pfair\/ERfair Scheduling of Asynchronous Periodic Tasks,\"Journal of Computer and System Sciences"},{"key":"e_1_2_1_12_1","volume-title":"Early-release Fair Scheduling,\"in Euromicro RTS","author":"Anderson J.","year":"2000","unstructured":"J. Anderson and A. Srinivasan, \"Early-release Fair Scheduling,\"in Euromicro RTS, 2000."},{"key":"e_1_2_1_13_1","volume-title":"Thread Scheduling for Multiprogrammed Multiprocessors,\"in SPAA","author":"Arora N. S.","year":"1998","unstructured":"N. S. Arora, R. D. Blumofe, and C. G. Plaxton, \"Thread Scheduling for Multiprogrammed Multiprocessors,\"in SPAA, 1998."},{"key":"e_1_2_1_14_1","volume-title":"Memory Hierarchy for Web Search,\"in HPCA","author":"Ayers G.","year":"2018","unstructured":"G. Ayers, J. H. Ahn, C. Kozyrakis, and P. Ranganathan,\"Memory Hierarchy for Web Search,\"in HPCA, 2018."},{"doi-asserted-by":"publisher","key":"e_1_2_1_15_1","DOI":"10.1145\/167088.167194"},{"doi-asserted-by":"publisher","key":"e_1_2_1_16_1","DOI":"10.5555\/899222"},{"doi-asserted-by":"publisher","key":"e_1_2_1_17_1","DOI":"10.1145\/1629575.1629579"},{"key":"e_1_2_1_18_1","volume-title":"Jigsaw: Scalable Software-Defined Caches,\"in PACT","author":"Beckmann N.","year":"2013","unstructured":"N. Beckmann and D. Sanchez, \"Jigsaw: Scalable Software-Defined Caches,\"in PACT, 2013."},{"doi-asserted-by":"publisher","key":"e_1_2_1_19_1","DOI":"10.1145\/1454115.1454128"},{"key":"e_1_2_1_20_1","volume-title":"Hyper-Threading Aware Process Scheduling Heuristics,\"in USENIX ATC","author":"Bulpin J. R.","year":"2005","unstructured":"J. R. Bulpin and I. A. Pratt, \"Hyper-Threading Aware Process Scheduling Heuristics,\"in USENIX ATC, 2005."},{"key":"e_1_2_1_21_1","volume-title":"On the Design and Implementation of a Cache-Aware Multicore Real-Time Scheduler,\" in Euromicro RTS","author":"Calandrino J. M.","year":"2009","unstructured":"J. M. Calandrino and J. H. Anderson,\"On the Design and Implementation of a Cache-Aware Multicore Real-Time Scheduler,\" in Euromicro RTS, 2009."},{"doi-asserted-by":"publisher","key":"e_1_2_1_22_1","DOI":"10.1145\/2063384.2063454"},{"doi-asserted-by":"publisher","key":"e_1_2_1_23_1","DOI":"10.1109\/EMRTS.1999.777450"},{"key":"e_1_2_1_24_1","volume-title":"Global EDF Schedulability Analysis for Parallel Tasks on Multi-Core Platforms,\" IEEE TPDS","author":"Chwa H. S.","year":"2017","unstructured":"H. S. Chwa, J. Lee, J. Lee, K.-M. Phan, A. Easwaran, and I. Shin, \"Global EDF Schedulability Analysis for Parallel Tasks on Multi-Core Platforms,\" IEEE TPDS, 2017."},{"key":"e_1_2_1_25_1","volume-title":"Responsive, High-Throughput Client OS for Many-core Architectures,\" in IEEE HotChips","author":"Colmenares J. A.","year":"2011","unstructured":"J. A. Colmenares, S. Bird, G. Eads, S. A. Hofmeyr, A. Kim, R. Poddar, H. Alkaff, K. Asanovic, and J. Kubiatowicz,\"Tessel- lation Operating System: Building a Real-Time, Responsive, High-Throughput Client OS for Many-core Architectures,\" in IEEE HotChips, 2011."},{"key":"e_1_2_1_26_1","volume-title":"Scheduling Heterogeneous Multi-cores through Performance Impact Estimation (PIE),\"in ISCA","author":"Craeynest K. V.","year":"2012","unstructured":"K. V. Craeynest, A. Jaleel, L. Eeckhout, P. Narv\u00e1ez, and J. S. Emer,\"Scheduling Heterogeneous Multi-cores through Performance Impact Estimation (PIE),\"in ISCA, 2012."},{"doi-asserted-by":"crossref","unstructured":"C. Delimitrou and C. Kozyrakis \"Paragon: QoS-Aware Scheduling for Heterogeneous Datacenters \"2013.","key":"e_1_2_1_27_1","DOI":"10.1145\/2451116.2451125"},{"doi-asserted-by":"publisher","key":"e_1_2_1_28_1","DOI":"10.1145\/2168836.2168873"},{"key":"e_1_2_1_29_1","volume-title":"Per-Thread Cycle Accounting in SMT Processors,\"in ASPLOS","author":"Eyerman S.","year":"2009","unstructured":"S. Eyerman and L. Eeckhout,\"Per-Thread Cycle Accounting in SMT Processors,\"in ASPLOS, 2009."},{"doi-asserted-by":"crossref","unstructured":"S. Eyerman and L. Eeckhout \"Probabilistic Job Symbiosis Modeling for SMT Processor Scheduling \"2010.","key":"e_1_2_1_30_1","DOI":"10.1145\/1736020.1736033"},{"key":"e_1_2_1_31_1","volume-title":"Improving Performance Isolation on Chip Multiprocessors via an Operating System Scheduler,\"in PACT","author":"Fedorova A.","year":"2007","unstructured":"A. Fedorova, M. Seltzer, and M. D. Smith,\"Improving Performance Isolation on Chip Multiprocessors via an Operating System Scheduler,\"in PACT, 2007."},{"key":"e_1_2_1_32_1","volume-title":"Parallel Distrib. Syst.","author":"Feliu J.","year":"2020","unstructured":"J. Feliu, J. Sahuquillo, S. Petit, and L. Eeckhout,\"Thread Isolation to Improve Symbiotic Scheduling on SMT Multicore Processors,\"IEEE Trans. Parallel Distrib. Syst., 2020."},{"key":"e_1_2_1_33_1","volume-title":"L1-bandwidth Aware Thread Allocation in Multicore SMT Processors,\"in PACT","author":"Feliu J.","year":"2013","unstructured":"J. Feliu, J. Sahuquillo, S. Petit, and J. Duato,\"L1-bandwidth Aware Thread Allocation in Multicore SMT Processors,\"in PACT, 2013."},{"key":"e_1_2_1_34_1","volume-title":"A Task Scheduling Algorithm based on Multi-core Processors,\"in MEC","author":"Geng X.","year":"2011","unstructured":"X. Geng, G. Xu, D. Wang, and Y. Shi,\"A Task Scheduling Algorithm based on Multi-core Processors,\"in MEC, 2011."},{"key":"e_1_2_1_35_1","volume-title":"Haswell: The Fourth-Generation Intel Core Processor,\"IEEE Micro","author":"Hammarlund P.","year":"2014","unstructured":"P. Hammarlund, A. J. Martinez, A. A. Bajwa, D. L. Hill, E. G. Hallnor, H. Jiang, M. G. Dixon, M. Derr, M. Hunsaker, R. Kumar, R. B. Osborne, R. Rajwar, R. Singhal, R. D'Sa, R. Chappell, S. Kaushik, S. Chennupaty, S. Jourdan, S. Gunther, T. Piazza, and T. Burton, \"Haswell: The Fourth-Generation Intel Core Processor,\"IEEE Micro, 2014."},{"key":"e_1_2_1_36_1","volume-title":"Easy and Expressive LLC Contention Model,\"in HPCS","author":"Hemani R.","year":"2016","unstructured":"R. Hemani, S. Banerjee, and A. Guha, \"Easy and Expressive LLC Contention Model,\"in HPCS, 2016."},{"key":"e_1_2_1_37_1","volume-title":"The Fair Share Scheduler,\"AT&T Bell Laboratories Technical Journal","author":"Henry G. J.","year":"1984","unstructured":"G. J. Henry,\"The UNIX system: The Fair Share Scheduler,\"AT&T Bell Laboratories Technical Journal, 1984."},{"doi-asserted-by":"publisher","key":"e_1_2_1_38_1","DOI":"10.1145\/2628071.2628089"},{"unstructured":"Intel \"Intel 64 and IA-32 Architecture Software Developer Manual \"2014.","key":"e_1_2_1_39_1"},{"key":"e_1_2_1_40_1","author":"Jain P. N.","year":"2020","unstructured":"P. N. Jain and S. K. Surve,\"A Review on Shared Resource Contention in Multicores and its Mitigating Techniques,\" IJHPSA, 2020.","journal-title":"\"A Review on Shared Resource Contention in Multicores and its Mitigating Techniques,\" IJHPSA"},{"key":"e_1_2_1_41_1","volume-title":"Unison Cache: A Scalable and Effective Die-Stacked DRAM Cache,\"in MICRO","author":"Jevdjic D.","year":"2014","unstructured":"D. Jevdjic, G. H. Loh, C. Kaynak, and B. Falsafi,\"Unison Cache: A Scalable and Effective Die-Stacked DRAM Cache,\"in MICRO, 2014."},{"key":"e_1_2_1_42_1","volume-title":"Latency, or Bandwidth? Have It All with Footprint Cache,\" in ISCA","author":"Jevdjic D.","year":"2013","unstructured":"D. Jevdjic, S. Volos, and B. Falsafi,\"Die-stacked DRAM Caches for Servers: Hit Ratio, Latency, or Bandwidth? Have It All with Footprint Cache,\" in ISCA, 2013."},{"doi-asserted-by":"publisher","key":"e_1_2_1_43_1","DOI":"10.1145\/3477132.3483548"},{"key":"e_1_2_1_44_1","volume-title":"Profiling a Warehouse- Scale Computer,\" IEEE Micro","author":"Kanev S.","year":"2016","unstructured":"S. Kanev, J. P. Darago, K. M. Hazelwood, P. Ranganathan, T. Moseley, G. Wei, and D. M. Brooks,\"Profiling a Warehouse- Scale Computer,\" IEEE Micro, 2016."},{"key":"e_1_2_1_45_1","volume-title":"A Fair Share Scheduler,\"CACM","author":"Kay J.","year":"1988","unstructured":"J. Kay and P. Lauder,\"A Fair Share Scheduler,\"CACM, 1988."},{"key":"e_1_2_1_46_1","volume-title":"Profile-assisted Compiler Support for Dynamic Predication in Diverge-Merge Processors,\"in CGO","author":"Kim H.","year":"2007","unstructured":"H. Kim, J. A. Joao, O. Mutlu, and Y. N. Patt,\"Profile-assisted Compiler Support for Dynamic Predication in Diverge-Merge Processors,\"in CGO, 2007."},{"key":"e_1_2_1_47_1","volume-title":"Fair Cache Sharing and Partitioning in a Chip Multiprocessor Architecture,\"in PACT","author":"Kim S.","year":"2004","unstructured":"S. Kim, D. Chandra, and Y. Solihin,\"Fair Cache Sharing and Partitioning in a Chip Multiprocessor Architecture,\"in PACT, 2004."},{"key":"e_1_2_1_48_1","volume-title":"Staccato: Shared-Memory Work-Stealing Task Scheduler with Cache-aware Memory Management,\"IJWGS","author":"Kuchumov R.","year":"2019","unstructured":"R. Kuchumov, A. S. Sokolov, and V. Korkhov,\"Staccato: Shared-Memory Work-Stealing Task Scheduler with Cache-aware Memory Management,\"IJWGS, 2019."},{"key":"e_1_2_1_49_1","volume-title":"CuttleSys: Data-Driven Resource Management for Interactive Services on Reconfigurable Multicores,\"in MICRO","author":"Kulkarni N.","year":"2020","unstructured":"N. Kulkarni, G. Gonzalez-Pumariega, A. Khurana, C. A. Shoemaker, C. Delimitrou, and D. H. Albonesi, \"CuttleSys: Data-Driven Resource Management for Interactive Services on Reconfigurable Multicores,\"in MICRO, 2020."},{"key":"e_1_2_1_50_1","volume-title":"Predicting Thread Profiles across Core Types via Machine Learning on Heterogeneous Multiprocessors,\"in SBESC","author":"Li C. V.","year":"2016","unstructured":"C. V. Li, V. Petrucci, and D. Moss\u00e9,\"Predicting Thread Profiles across Core Types via Machine Learning on Heterogeneous Multiprocessors,\"in SBESC, 2016."},{"key":"e_1_2_1_51_1","volume-title":"Exploring Machine Learning for Thread Characterization on Heterogeneous Multiprocessors,\"ACM OSR","author":"Li C. V.","year":"2017","unstructured":"C. V. Li, V. Petrucci, and D. Moss\u00e9, \"Exploring Machine Learning for Thread Characterization on Heterogeneous Multiprocessors,\"ACM OSR, 2017."},{"key":"e_1_2_1_52_1","volume-title":"Global EDF Scheduling for Parallel Real-Time Tasks,\"in Springer RTS","author":"Li J.","year":"2015","unstructured":"J. Li, Z. Luo, D. Ferry, K. Agrawal, C. Gill, and C. Lu,\"Global EDF Scheduling for Parallel Real-Time Tasks,\"in Springer RTS, 2015."},{"doi-asserted-by":"crossref","unstructured":"C. Lin T. Huang and M. D. F. Wong \"An Efficient Work-Stealing Scheduler for Task Dependency Graph \"in ICPADS 2020.","key":"e_1_2_1_53_1","DOI":"10.1109\/ICPADS51040.2020.00018"},{"key":"e_1_2_1_54_1","volume-title":"Pfair Scheduling of Periodic Tasks with Allocation Constraints on Multiple Processors,\"in IPDPS","author":"Liu D.","year":"2004","unstructured":"D. Liu and Y. Lee,\"Pfair Scheduling of Periodic Tasks with Allocation Constraints on Multiple Processors,\"in IPDPS, 2004."},{"key":"e_1_2_1_55_1","volume-title":"Efficiently Enabling Conventional Block Sizes for Very Large Die-Stacked DRAM Caches,\" in MICRO","author":"Loh G. H.","year":"2011","unstructured":"G. H. Loh and M. D. Hill, \"Efficiently Enabling Conventional Block Sizes for Very Large Die-Stacked DRAM Caches,\" in MICRO, 2011."},{"doi-asserted-by":"publisher","key":"e_1_2_1_56_1","DOI":"10.1145\/2901318.2901326"},{"doi-asserted-by":"publisher","key":"e_1_2_1_57_1","DOI":"10.1145\/1065010.1065034"},{"doi-asserted-by":"crossref","unstructured":"C. Mattihalli \"Designing and Implementing of Earliest Deadline First Scheduling Algorithm on Standard Linux \"in CPSCom P. Zhu L. Wang F. Xia H. Chen I. McLoughlin S. Tsao M. Sato S. Chai and I. King Eds. 2010.","key":"e_1_2_1_58_1","DOI":"10.1109\/GreenCom-CPSCom.2010.82"},{"key":"e_1_2_1_59_1","volume-title":"ESP: A Machine Learning Approach to Predicting Application Interference,\" in ICAC","author":"Mishra N.","year":"2017","unstructured":"N. Mishra, J. D. Lafferty, and H. Hoffmann,\"ESP: A Machine Learning Approach to Predicting Application Interference,\" in ICAC, 2017."},{"key":"e_1_2_1_60_1","volume-title":"Applying Machine Learning Techniques to Improve Linux Process Scheduling,\"in TENCON","author":"Negi A.","year":"2005","unstructured":"A. Negi and P. K. Kumar, \"Applying Machine Learning Techniques to Improve Linux Process Scheduling,\"in TENCON, 2005."},{"key":"e_1_2_1_61_1","volume-title":"A Machine Learning Approach for Performance Prediction and Scheduling on Heterogeneous CPUs,\"in SBAC-PAD","author":"Nemirovsky D.","year":"2017","unstructured":"D. Nemirovsky, T. Arkose, N. Markovic, M. Nemirovsky, O. S. Unsal, and A. Cristal,\"A Machine Learning Approach for Performance Prediction and Scheduling on Heterogeneous CPUs,\"in SBAC-PAD, 2017."},{"unstructured":"K. Pearson \"Note on Regression and Inheritance in the Case of Two Parents \"Royal Society of London 1895.","key":"e_1_2_1_62_1"},{"key":"e_1_2_1_63_1","volume-title":"Robinhood: Towards Efficient Work-Stealing in Virtualized Environments,\"IEEE TPDS","author":"Peng Y.","year":"2016","unstructured":"Y. Peng, S. Wu, and H. Jin, \"Robinhood: Towards Efficient Work-Stealing in Virtualized Environments,\"IEEE TPDS, 2016."},{"key":"e_1_2_1_64_1","volume-title":"Fundamental Latency Trade-off in Architecting DRAM Caches: Outperforming Impractical SRAM-Tags with a Simple and Practical Design,\"in MICRO","author":"Qureshi M. K.","year":"2012","unstructured":"M. K. Qureshi and G. H. Loh,\"Fundamental Latency Trade-off in Architecting DRAM Caches: Outperforming Impractical SRAM-Tags with a Simple and Practical Design,\"in MICRO, 2012."},{"doi-asserted-by":"publisher","key":"e_1_2_1_65_1","DOI":"10.1145\/2150976.2151002"},{"key":"e_1_2_1_66_1","volume-title":"Thread Assignment in Multicore\/Multithreaded Processors: A Statistical Approach,\"IEEE TC","author":"Radojkovic P.","year":"2016","unstructured":"P. Radojkovic, P. M. Carpenter, M. Moret\u00f3, V. Cakarevic, J. Verd\u00fa, A. Pajuelo, F. J. Cazorla, M. Nemirovsky, and M. Valero, \"Thread Assignment in Multicore\/Multithreaded Processors: A Statistical Approach,\"IEEE TC, 2016."},{"doi-asserted-by":"crossref","unstructured":"S. Rendle \"Factorization Machines \"in ICDM 2010.","key":"e_1_2_1_67_1","DOI":"10.1109\/ICDM.2010.127"},{"key":"e_1_2_1_68_1","volume-title":"Multi-core Real-Time Scheduling for Generalized Parallel Task Models,\"in RTSS","author":"Saifullah A.","year":"2011","unstructured":"A. Saifullah, K. Agrawal, C. Lu, and C. Gill,\"Multi-core Real-Time Scheduling for Generalized Parallel Task Models,\"in RTSS, 2011."},{"key":"e_1_2_1_69_1","volume-title":"A Mostly-Clean DRAM Cache for Effective Hit Speculation and Self-Balancing Dispatch,\"in MICRO","author":"Sim J.","year":"2012","unstructured":"J. Sim, G. H. Loh, H. Kim, M. O'Connor, and M. Thottethodi,\"A Mostly-Clean DRAM Cache for Effective Hit Speculation and Self-Balancing Dispatch,\"in MICRO, 2012."},{"doi-asserted-by":"publisher","key":"e_1_2_1_70_1","DOI":"10.14778\/3485450.3485454"},{"key":"e_1_2_1_71_1","volume-title":"Symbiotic Jobscheduling for a Simultaneous Multithreading Processor,\"in ASPLOS","author":"Snavely A.","year":"2000","unstructured":"A. Snavely and D. M. Tullsen, \"Symbiotic Jobscheduling for a Simultaneous Multithreading Processor,\"in ASPLOS, 2000."},{"key":"e_1_2_1_72_1","volume-title":"Data Sharing or Resource Contention: Toward Performance Transparency on Multicore Systems,\"in ATC","author":"Srikanthan S.","year":"2015","unstructured":"S. Srikanthan, S. Dwarkadas, and K. Shen,\"Data Sharing or Resource Contention: Toward Performance Transparency on Multicore Systems,\"in ATC, 2015."},{"key":"e_1_2_1_73_1","volume-title":"Optimal Rate-based Scheduling on Multiprocessors,\"JCSS","author":"Srinivasan A.","year":"2006","unstructured":"A. Srinivasan and J. H. Anderson, \"Optimal Rate-based Scheduling on Multiprocessors,\"JCSS, 2006."},{"key":"e_1_2_1_74_1","volume-title":"The Cache and Memory Subsystems of the IBM POWER8 Processor,\"IBM JRD","author":"Starke W. J.","year":"2015","unstructured":"W. J. Starke, J. Stuecheli, D. Daly, J. S. Dodson, F. Auernhammer, P. Sagmeister, G. L. Guthrie, C. F. Marino, M. S. Siegel, and B. Blaner,\"The Cache and Memory Subsystems of the IBM POWER8 Processor,\"IBM JRD, 2015."},{"doi-asserted-by":"publisher","key":"e_1_2_1_75_1","DOI":"10.1145\/1272996.1273004"},{"key":"e_1_2_1_76_1","volume-title":"A Comprehensive Memory Modeling Tool and Its Application to the Design and Analysis of Future Memory Hierarchies,\"in ISCA","author":"Thoziyoor S.","year":"2008","unstructured":"S. Thoziyoor, J. H. Ahn, M. Monchiero, J. B. Brockman, and N. P. Jouppi,\"A Comprehensive Memory Modeling Tool and Its Application to the Design and Analysis of Future Memory Hierarchies,\"in ISCA, 2008."},{"key":"e_1_2_1_77_1","volume-title":"Fat Caches for Scale-Out Servers,\"IEEE Micro","author":"Volos S.","year":"2017","unstructured":"S. Volos, D. Jevdjic, B. Falsafi, and B. Grot,\"Fat Caches for Scale-Out Servers,\"IEEE Micro, 2017."},{"doi-asserted-by":"publisher","key":"e_1_2_1_78_1","DOI":"10.1145\/1807128.1807132"},{"doi-asserted-by":"publisher","key":"e_1_2_1_79_1","DOI":"10.1145\/223982.223990"},{"key":"e_1_2_1_80_1","volume-title":"Cache Contention and Application Performance Prediction for Multi-core systems,\"in ISPASS","author":"Xu C.","year":"2010","unstructured":"C. Xu, x. Chen, R. Dick, and Z. Mao, \"Cache Contention and Application Performance Prediction for Multi-core systems,\"in ISPASS, 2010."},{"key":"e_1_2_1_81_1","author":"Xu D.","year":"2012","unstructured":"D. Xu, C. Wu, P.-C. Yew, J. Li, and Z. Wang,\"Providing Fairness on Shared-Memory Multiprocessors via Process Scheduling,\"SIGMETRICS Perform. Eval. Rev., 2012.","journal-title":"Eval. Rev."},{"doi-asserted-by":"crossref","unstructured":"A. Yasin \"A Top-Down Method for Performance Analysis and Counters Architecture \"in ISPASS 2014.","key":"e_1_2_1_82_1","DOI":"10.1109\/ISPASS.2014.6844459"},{"key":"e_1_2_1_83_1","volume-title":"Survey of Scheduling Techniques for Addressing Shared Resources in Multicore Processors,\" ACM CSUR","author":"Zhuravlev S.","year":"2012","unstructured":"S. Zhuravlev, J. C. Saez, S. Blagodurov, A. Fedorova, and M. Prieto,\"Survey of Scheduling Techniques for Addressing Shared Resources in Multicore Processors,\" ACM CSUR, 2012."}],"container-title":["Proceedings of the ACM on Measurement and Analysis of Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3579436","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3579436","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:46:41Z","timestamp":1750178801000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3579436"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,2,27]]},"references-count":83,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2023,2,27]]}},"alternative-id":["10.1145\/3579436"],"URL":"https:\/\/doi.org\/10.1145\/3579436","relation":{},"ISSN":["2476-1249"],"issn-type":[{"type":"electronic","value":"2476-1249"}],"subject":[],"published":{"date-parts":[[2023,2,27]]},"assertion":[{"value":"2023-03-02","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}