{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T23:46:08Z","timestamp":1783035968123,"version":"3.54.6"},"reference-count":69,"publisher":"Association for Computing Machinery (ACM)","issue":"5s","funder":[{"name":"U.S. NSF","award":["CNS-2344505"],"award-info":[{"award-number":["CNS-2344505"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Embed. Comput. Syst."],"published-print":{"date-parts":[[2025,9,30]]},"abstract":"<jats:p>GPUs have recently been adopted in many real-time embedded systems. However, existing GPU scheduling solutions are mostly open-loop and rely on the estimation of worst-case execution time (WCET). Although adaptive solutions, such as feedback control scheduling, have been previously proposed to handle this challenge for CPU-based real-time tasks, they cannot be directly applied to GPU, because GPUs have different and more complex architectures and so schedulable utilization bounds cannot apply to GPUs yet. In this article, we propose FC-GPU, the first Feedback Control GPU scheduling framework for real-time embedded systems. To model the GPU resource contention among tasks, we analytically derive a multi-input-multi-output (MIMO) system model that captures the impacts of task rate adaptation on the response times of different tasks. Building on this model, we design a MIMO controller that dynamically adjusts task rates based on measured response times. Our extensive hardware testbed results on an Nvidia RTX 3090 GPU and an AMD MI-100 GPU demonstrate that FC-GPU can provide better real-time performance even when the task execution times significantly increase at runtime.<\/jats:p>","DOI":"10.1145\/3761812","type":"journal-article","created":{"date-parts":[[2025,8,15]],"date-time":"2025-08-15T10:13:43Z","timestamp":1755252823000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["FC-GPU: Feedback Control GPU Scheduling for Real-time Embedded Systems"],"prefix":"10.1145","volume":"24","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5848-5667","authenticated-orcid":false,"given":"Srinivasan","family":"Subramaniyan","sequence":"first","affiliation":[{"name":"The Ohio State University","place":["Columbus, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9633-1418","authenticated-orcid":false,"given":"Xiaorui","family":"Wang","sequence":"additional","affiliation":[{"name":"Electrical and Computer Engineering, The Ohio State University","place":["Columbus, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,9,26]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"Apollo. 2007. Apollo 3.0 Software Architecture. GitHub. https:\/\/tinyurl.com\/mhd6dfka"},{"key":"e_1_3_1_3_2","unstructured":"Apollo. Apollo Traffic Light Perception. GitHub. https:\/\/github.com\/ApolloAuto\/apollo\/blob\/master\/docs\/specs\/traffic_light.md"},{"key":"e_1_3_1_4_2","doi-asserted-by":"crossref","unstructured":"Yidi Wang Mohsen Karimi Yecheng Xiang and Hyoseung Kim. 2021. Balancing energy efficiency and real-time performance in GPU scheduling. In 2021 IEEE Real-Time Systems Symposium (RTSS). IEEE 110\u2013122.","DOI":"10.1109\/RTSS52674.2021.00021"},{"key":"e_1_3_1_5_2","doi-asserted-by":"crossref","unstructured":"Chenyang Lu Xiaorui Wang and Christopher Gill. 2003. Feedback control real-time scheduling in ORB middleware. In The 9th IEEE Real-Time and Embedded Technology and Applications Symposium. IEEE 37\u201348.","DOI":"10.1109\/RTTAS.2003.1203035"},{"key":"e_1_3_1_6_2","unstructured":"NVIDIA Corporation. CUDA Samples: NVIDIA Benchmark Suite. GitHub. https:\/\/github.com\/NVIDIA\/cuda-samples"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/321738.321743"},{"key":"e_1_3_1_8_2","doi-asserted-by":"crossref","unstructured":"Lui Sha Ragunathan Rajkumar and Shirish S Sathaye. 1994. Generalized rate-monotonic scheduling theory: A framework for developing real-time systems. Proceedings of the IEEE 82 1 (1994) 68\u201382.","DOI":"10.1109\/5.259427"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.5555\/827270.829013"},{"key":"e_1_3_1_10_2","volume-title":"Proceedings of the 3rd Operating Systems Design and Implementation","author":"Steere David C.","year":"1999","unstructured":"David C. Steere, Ashvin Goel, Joshua Gruenberg, Dylan McNamee, Calton Pu, and Jonathan Walpole. 1999. A feedback-driven proportion allocator for real-rate scheduling. In Proceedings of the 3rd Operating Systems Design and Implementation."},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.5555\/827272.829115"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.5555\/874076.876476"},{"issue":"1","key":"e_1_3_1_13_2","first-page":"85","article-title":"Feedback control real-time scheduling: Framework, modeling, and algorithms","volume":"23","author":"Lu C.","year":"2002","unstructured":"C. Lu, J. A. Stankovic, G. Tao, and S. H. Son. 2002. Feedback control real-time scheduling: Framework, modeling, and algorithms. Journal of Real-Time Systems, Special Issue on Control-Theoretical Approaches to Real-Time Computing 23, 1\/2 (July2002), 85\u2013126.","journal-title":"Journal of Real-Time Systems, Special Issue on Control-Theoretical Approaches to Real-Time Computing"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDCS.2004.1281612"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.5555\/975344"},{"key":"e_1_3_1_16_2","volume-title":"Hard Real-time Computing Systems: Predictable Scheduling Algorithms and Applications (Real-Time Systems Series)","author":"Buttazzo Giorgio C.","year":"2004","unstructured":"Giorgio C. Buttazzo. 2004. Hard Real-time Computing Systems: Predictable Scheduling Algorithms and Applications (Real-Time Systems Series). Springer-Verlag TELOS, Santa Clara, CA, USA."},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/RTSS.2005.15"},{"key":"e_1_3_1_18_2","article-title":"Feedback utilization control in distributed real-time systems with end-to-end tasks","author":"Lu Chenyang","year":"2005","unstructured":"Chenyang Lu, Xiaorui Wang, and Xenofon Koutsoukos. 2005. Feedback utilization control in distributed real-time systems with end-to-end tasks. IEEE Transactions on Parallel and Distributed Systems 16, 6 (2005), 550\u2013561.","journal-title":"IEEE Transactions on Parallel and Distributed Systems"},{"key":"e_1_3_1_19_2","doi-asserted-by":"crossref","unstructured":"Xiaorui Wang Dong Jia Chenyang Lu and Xenofon Koutsoukos. 2007. DEUCON: Decentralized end-to-end utilization control for distributed real-time systems. IEEE Transactions on Parallel and Distributed Systems 18 7 (2007) 996\u20131009.","DOI":"10.1109\/TPDS.2007.1051"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/CRV.2007.54"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/RTCSA.2008.24"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/1450135.1450155"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/RTCSA.2009.49"},{"key":"e_1_3_1_24_2","volume-title":"Proceedings of the 17th IEEE International Workshop on Quality of Service","author":"Chen Ming","year":"2009","unstructured":"Ming Chen, Xiaorui Wang, and Ben Taylor. 2009. Integrated control of matching delay and CPU utilization in information dissemination systems. In Proceedings of the 17th IEEE International Workshop on Quality of Service."},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/RTAS.2009.12"},{"key":"e_1_3_1_26_2","volume-title":"IEEE 12th International Conference on Computer Vision Workshops, ICCV Workshops","author":"Kim Jun-Sik","year":"2009","unstructured":"Jun-Sik Kim, Myung Hwangbo, and Takeo Kanade. 2009. Realtime affine-photometric KLT feature tracker on GPU in CUDA framework. In IEEE 12th International Conference on Computer Vision Workshops, ICCV Workshops."},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC.2009.5306797"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/1555815.1555794"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/1815961.1815998"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jss.2010.01.031"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/ECRTS.2011.18"},{"key":"e_1_3_1_32_2","doi-asserted-by":"crossref","unstructured":"Po-Han Wang Chia-Lin Yang Yen-Ming Chen and Yu-Jung Cheng. 2011. Power gating strategies on GPUs. ACM Transactions on Architecture and Code Optimization (TACO) 8 3 (2011) 1\u201325.","DOI":"10.1145\/2019608.2019612"},{"key":"e_1_3_1_33_2","doi-asserted-by":"crossref","unstructured":"Kai Ma Xue Li Ming Chen and Xiaorui Wang. 2011. Scalable power control for many-core architectures running multi-threaded applications. In Proceedings of the 38th Annual International Symposium on Computer Architecture. 449\u2013460.","DOI":"10.1145\/2000064.2000117"},{"key":"e_1_3_1_34_2","volume-title":"USENIX Annual Technical Conference","author":"Kato Shinpei","year":"2011","unstructured":"Shinpei Kato, Karthik Lakshmanan, Ragunathan Rajkumar, Yutaka Ishikawa, et\u00a0al. 2011. TimeGraph: GPU scheduling for real-time multi-tasking environments. In USENIX Annual Technical Conference."},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/RTCSA.2011.65"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2012.6168946"},{"key":"e_1_3_1_37_2","volume-title":"Computers as Components: Principles of Embedded Computing System Design","author":"Wolf Marilyn","year":"2012","unstructured":"Marilyn Wolf. 2012. Computers as Components: Principles of Embedded Computing System Design. Elsevier."},{"key":"e_1_3_1_38_2","doi-asserted-by":"crossref","unstructured":"Glenn A Elliott and James H Anderson. 2012. Globally scheduled real-time multiprocessor systems with GPUs. Real-Time Systems 48 1 (2012) 34\u201374.","DOI":"10.1007\/s11241-011-9140-y"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11241-011-9141-x"},{"key":"e_1_3_1_40_2","doi-asserted-by":"crossref","unstructured":"Yu Wang and Jien Kato. 2012. Integrated pedestrian detection and localization using stereo cameras. In Digital Signal Processing for In-Vehicle Systems and Safety. New York NY: Springer New York 229\u2013238.","DOI":"10.1007\/978-1-4419-9607-7_16"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/RTSS.2013.12"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/2678373.2665702"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/ASPDAC.2014.6742976"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/RTAS.2015.7108420"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.5555\/2988385"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2016.2547916"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3018743.3018748"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/RTAS.2017.3"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/RTSS.2017.00017"},{"key":"e_1_3_1_50_2","volume-title":"Workshop on Operating Systems Platforms for Embedded Real Time Systems Applications (OSPERT)","author":"Otterness Nathan","year":"2017","unstructured":"Nathan Otterness, Ming Yang, Tanya Amert, James Anderson, and F. Donelson Smith. 2017. Inferring the scheduling policies of an embedded CUDA GPU. In Workshop on Operating Systems Platforms for Embedded Real Time Systems Applications (OSPERT)."},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/RTCSA.2017.8046309"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICIP.2017.8297110"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/RTSS.2018.00021"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/RTAS.2019.00011"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/RTCSA.2019.8864564"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDCS47774.2020.00092"},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/DSN48063.2020.00031"},{"key":"e_1_3_1_58_2","volume-title":"Proceedings of the 14th USENIX Conference on Operating Systems Design and Implementation","author":"Bai Zhihao","year":"2020","unstructured":"Zhihao Bai, Zhen Zhang, Yibo Zhu, and Xin Jin. 2020. PipeSwitch: Fast pipelined context switching for deep learning applications. In Proceedings of the 14th USENIX Conference on Operating Systems Design and Implementation."},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/RTAS48715.2020.00-17"},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/RTSS52674.2021.00038"},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/ETFA45728.2021.9613590"},{"key":"e_1_3_1_62_2","doi-asserted-by":"crossref","unstructured":"Guin Gilman and Robert J Walls. 2022. Characterizing concurrency mechanisms for NVIDIA GPUs under deep learning workloads. ACM SIGMETRICS Performance Evaluation Review 49 3 (2022) 32\u201334.","DOI":"10.1145\/3529113.3529124"},{"key":"e_1_3_1_63_2","doi-asserted-by":"publisher","DOI":"10.1145\/3492321.3519576"},{"key":"e_1_3_1_64_2","unstructured":"Wei Gao Qinghao Hu Zhisheng Ye Peng Sun Xiaolin Wang Yingwei Luo Tianwei Zhang and Yonggang Wen. 2022. Deep learning workload scheduling in GPU datacenters: Taxonomy challenges and vision. arXiv:205.11913. Retrieved from https:\/\/arxiv.org\/abs\/205.11913"},{"key":"e_1_3_1_65_2","volume-title":"16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22)","author":"Han Mingcong","year":"2022","unstructured":"Mingcong Han, Hanze Zhang, Rong Chen, and Haibo Chen. 2022. Microsecond-scale preemption for concurrent GPU-accelerated DNN inferences. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22)."},{"key":"e_1_3_1_66_2","doi-asserted-by":"crossref","unstructured":"Yunhao Bai Li Li Zejiang Wang Xiaorui Wang and Junmin Wang. 2022. Performance optimization of autonomous driving control under end-to-end deadlines. Real-Time Systems 58 4 (2022) 509\u2013547.","DOI":"10.1007\/s11241-022-09379-6"},{"key":"e_1_3_1_67_2","doi-asserted-by":"publisher","DOI":"10.1109\/RTSS55097.2022.00042"},{"key":"e_1_3_1_68_2","doi-asserted-by":"publisher","DOI":"10.1109\/RTAS58335.2023.00022"},{"key":"e_1_3_1_69_2","doi-asserted-by":"crossref","unstructured":"Guoyu Chen Srinivasan Subramaniyan and Xiaorui Wang. 2024. Latency-guaranteed co-location of inference and training for reducing data center expenses. 2024 IEEE 44th International Conference on Distributed Computing Systems (ICDCS). IEEE 2024.","DOI":"10.1109\/ICDCS60910.2024.00051"},{"key":"e_1_3_1_70_2","volume-title":"36th Euromicro Conference on Real-Time Systems","author":"Wang Yidi","year":"2024","unstructured":"Yidi Wang, Cong Liu, Daniel Wong, and Hyoseung Kim. 2024. GCAPS: GPU context-aware preemptive priority-based scheduling for real-time tasks. In 36th Euromicro Conference on Real-Time Systems. Schloss Dagstuhl\u2013Leibniz-Zentrum f\u00fcr Informatik."}],"container-title":["ACM Transactions on Embedded Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3761812","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,8]],"date-time":"2025-10-08T14:28:34Z","timestamp":1759933714000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3761812"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,26]]},"references-count":69,"journal-issue":{"issue":"5s","published-print":{"date-parts":[[2025,9,30]]}},"alternative-id":["10.1145\/3761812"],"URL":"https:\/\/doi.org\/10.1145\/3761812","relation":{},"ISSN":["1539-9087","1558-3465"],"issn-type":[{"value":"1539-9087","type":"print"},{"value":"1558-3465","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,9,26]]},"assertion":[{"value":"2025-08-09","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-08-11","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-09-26","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}