{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,11]],"date-time":"2025-09-11T19:19:35Z","timestamp":1757618375506,"version":"3.44.0"},"reference-count":73,"publisher":"Springer Science and Business Media LLC","issue":"8","license":[{"start":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T00:00:00Z","timestamp":1750204800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T00:00:00Z","timestamp":1750204800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["ECCS-1710859","ECCS-1710859","ECCS-1710859"],"award-info":[{"award-number":["ECCS-1710859","ECCS-1710859","ECCS-1710859"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100029699","name":"Penn State College of Engineering, Pennsylvania State University","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100029699","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Mach Learn"],"published-print":{"date-parts":[[2025,8]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>This paper studies the issues of ensuring all-time safety and sample-efficient meta update in online safe meta reinforcement learning (MRL) on physical agents (e.g., mobile robots). We propose a novel masked Follow-the-Last-Parameter-Policy (FTLPP) framework, which is composed of a policy masking framework and a sample-efficient online meta update method. The policy masking framework applies a masking function over the learned control policy and ensures all-time safety by suppressing the probability of executing unsafe actions to a sufficiently small value. To enhance sample efficiency, the problem of online update of the meta parameter is transformed into a policy optimization problem, where the tasks are the states and the meta parameters for the next task are the actions, and then is solved using an off-policy reinforcement learning algorithm. We evaluate our method on Frozen Lake, Acrobot, Half Cheetah and Hopper from OpenAI gym and compare it with baseline methods Meta SRL and the variants of FTML and SAILR.<\/jats:p>","DOI":"10.1007\/s10994-025-06810-4","type":"journal-article","created":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T15:31:18Z","timestamp":1750260678000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["All-time safety and sample-efficient meta update for online safe meta reinforcement learning under Markov task transition"],"prefix":"10.1007","volume":"114","author":[{"given":"Zhenyuan","family":"Yuan","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Siyuan","family":"Xu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Minghui","family":"Zhu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,6,18]]},"reference":[{"issue":"13","key":"6810_CR1","doi-asserted-by":"publisher","first-page":"1608","DOI":"10.1177\/0278364910371999","volume":"29","author":"P Abbeel","year":"2010","unstructured":"Abbeel, P., Coates, A., & Ng, A. Y. (2010). Autonomous helicopter aerobatics through apprenticeship learning. The International Journal of Robotics Research, 29(13), 1608\u20131639.","journal-title":"The International Journal of Robotics Research"},{"key":"6810_CR2","unstructured":"Acar, D., Zhu, R., & Saligrama, V. (2021). Memory efficient online meta learning. In Proc. Int. Conf Machine Learning (ICML) (pp. 32\u201342). PMLR."},{"key":"6810_CR3","unstructured":"Achiam, J., Held, D., Tamar, A., & Abbeel, P. (2017). Constrained policy optimization. In International Conference on Machine Learning (pp. 22\u201331). PMLR."},{"key":"6810_CR4","doi-asserted-by":"publisher","DOI":"10.1201\/9781315140223","volume-title":"Constrained Markov Decision Processes","author":"E Altman","year":"2021","unstructured":"Altman, E. (2021). Constrained Markov Decision Processes. Routledge."},{"key":"6810_CR5","volume-title":"Introduction to Stochastic Control Theory","author":"KJ \u00c5str\u00f6m","year":"2012","unstructured":"\u00c5str\u00f6m, K. J. (2012). Introduction to Stochastic Control Theory. Courier Corporation."},{"key":"6810_CR6","unstructured":"Balcan, M.-F., Khodak, M., & Talwalkar, A. (2019). Provable guarantees for gradient-based meta-learning. In: Proc. Int. Conf. Machine Learning (ICML) (pp. 424\u2013433)."},{"issue":"4","key":"6810_CR7","doi-asserted-by":"publisher","first-page":"880","DOI":"10.1287\/moor.1080.0324","volume":"33","author":"A Basu","year":"2008","unstructured":"Basu, A., Bhattacharyya, T., & Borkar, V. S. (2008). A learning algorithm for risk-sensitive cost. Mathematics of Operations Research, 33(4), 880\u2013898.","journal-title":"Mathematics of Operations Research"},{"issue":"6","key":"6810_CR8","doi-asserted-by":"publisher","first-page":"1864","DOI":"10.1109\/TIE.2009.2015748","volume":"56","author":"AG Beccuti","year":"2009","unstructured":"Beccuti, A. G., Mari\u00e9thoz, S., Cliquennois, S., Wang, S., & Morari, M. (2009). Explicit model predictive control of dc-dc switched-mode power supplies with extended Kalman filtering. IEEE Transactions on Industrial Electronics, 56(6), 1864\u20131874.","journal-title":"IEEE Transactions on Industrial Electronics"},{"key":"6810_CR9","unstructured":"Beck, J., Vuorio, R., Liu, E. Z., Xiong, Z., Zintgraf, L., Finn, C., & Whiteson, S. (2023). A survey of meta-reinforcement learning. arXiv preprint arXiv:2301.08028."},{"key":"6810_CR10","unstructured":"Berkenkamp, F., Turchetta, M., Schoellig, A., & Krause, A. (2017). Safe model-based reinforcement learning with stability guarantees. In Advances in Neural Information Processing Systems (pp. 908\u2013918)."},{"key":"6810_CR11","doi-asserted-by":"crossref","unstructured":"Bertoncelli, F., Ruggiero, F., & Sabattini, L. (2020). Linear time-varying mpc for nonprehensile object manipulation with a nonholonomic mobile robot. In Proc. Int. Conf. Robotics and Automation (ICRA) (pp. 11032\u201311038).","DOI":"10.1109\/ICRA40945.2020.9197173"},{"issue":"5","key":"6810_CR12","doi-asserted-by":"publisher","first-page":"339","DOI":"10.1016\/S0167-6911(01)00152-9","volume":"44","author":"VS Borkar","year":"2001","unstructured":"Borkar, V. S. (2001). A sensitivity formula for risk-sensitive cost and the actor-critic algorithm. Systems & Control Letters, 44(5), 339\u2013346.","journal-title":"Systems & Control Letters"},{"key":"6810_CR13","unstructured":"Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., & Zaremba, W. (2016). Openai gym. arXiv preprint arXiv:1606.01540."},{"key":"6810_CR14","doi-asserted-by":"crossref","unstructured":"Cardaliaguet, P., Quincampoix, M., & Saint-Pierre, P. (1999). Set-valued numerical analysis for optimal control and differential games. In Stochastic and Differential Games (pp. 177\u2013247). Springer.","DOI":"10.1007\/978-1-4612-1592-9_4"},{"issue":"4","key":"6810_CR15","doi-asserted-by":"publisher","first-page":"2563","DOI":"10.1287\/opre.2021.2151","volume":"70","author":"S Cen","year":"2022","unstructured":"Cen, S., Cheng, C., Chen, Y., Wei, Y., & Chi, Y. (2022). Fast global convergence of natural policy gradient methods with entropy regularization. Operations Research, 70(4), 2563\u20132578.","journal-title":"Operations Research"},{"key":"6810_CR16","doi-asserted-by":"crossref","unstructured":"Chaudhary, S., & Kalathil, D. (2022). Safe online convex optimization with unknown linear safety constraints. In Proceedings of the AAAI Conference on Artificial Intelligence, (Vol. 36, pp. 6175\u20136182).","DOI":"10.1609\/aaai.v36i6.20566"},{"key":"6810_CR17","doi-asserted-by":"crossref","unstructured":"Choi, J., Castaneda, F., Tomlin, C. J., & Sreenath, K. (2020). Reinforcement learning for safety-critical control under model uncertainty, using control Lyapunov functions and control barrier functions. arXiv preprint arXiv:2004.07584.","DOI":"10.15607\/RSS.2020.XVI.088"},{"issue":"3","key":"6810_CR18","doi-asserted-by":"publisher","first-page":"528","DOI":"10.1109\/9.847737","volume":"45","author":"SP Coraluppi","year":"2000","unstructured":"Coraluppi, S. P., & Marcus, S. I. (2000). Mixed risk-neutral\/minimax control of discrete-time, finite-state Markov decision processes. IEEE Transactions on Automatic Control, 45(3), 528\u2013532.","journal-title":"IEEE Transactions on Automatic Control"},{"key":"6810_CR19","unstructured":"Cosner, R., Tucker, M., Taylor, A., Li, K., Molnar, T., Ubelacker, W., Alan, A., Orosz, G., Yue, Y., & Ames, A. (2022). Safety-aware preference-based learning for safety-critical control. In Learning for Dynamics and Control Conference (pp. 1020\u20131033)."},{"key":"6810_CR20","doi-asserted-by":"crossref","unstructured":"Curi, S., Lederer, A., Hirche, S., & Krause, A. (2022). Safe reinforcement learning via confidence-based filters. In Proc. IEEE Conf. Decision and Control (CDC) (pp. 3409\u20133415).","DOI":"10.1109\/CDC51059.2022.9992470"},{"issue":"3","key":"6810_CR21","doi-asserted-by":"publisher","first-page":"1749","DOI":"10.1109\/TRO.2022.3232542","volume":"39","author":"C Dawson","year":"2023","unstructured":"Dawson, C., Gao, S., & Fan, C. (2023). Safe control with learned certificates: A survey of neural Lyapunov, barrier, and contraction methods for robotics and control. IEEE Transactions on Robotics, 39(3), 1749\u20131767.","journal-title":"IEEE Transactions on Robotics"},{"key":"6810_CR22","unstructured":"Finn, C., Abbeel, P., & Levine, S. (2017). Model-agnostic meta-learning for fast adaptation of deep networks. In Proc. Int. Conf Machine Learning (ICML) (pp. 1126\u20131135)."},{"key":"6810_CR23","unstructured":"Finn, C., Rajeswaran, A., Kakade, S., & Levine, S. (2019). Online meta-learning. In Proc. Int. Conf Machine Learning (ICML) (pp. 1920\u20131930)."},{"issue":"7","key":"6810_CR24","doi-asserted-by":"publisher","first-page":"2737","DOI":"10.1109\/TAC.2018.2876389","volume":"64","author":"JF Fisac","year":"2018","unstructured":"Fisac, J. F., Akametalu, A. K., Zeilinger, M. N., Kaynama, S., Gillula, J., & Tomlin, C. J. (2018). A general safety framework for learning-based control in uncertain robotic systems. IEEE Transactions on Automatic Control, 64(7), 2737\u20132752.","journal-title":"IEEE Transactions on Automatic Control"},{"issue":"2","key":"6810_CR25","doi-asserted-by":"publisher","first-page":"24","DOI":"10.3390\/machines7020024","volume":"7","author":"Y Fu","year":"2019","unstructured":"Fu, Y., Jha, D. K., Zhang, Z., Yuan, Z., & Ray, A. (2019). Neural network-based learning from demonstration of an autonomous ground robot. Machines, 7(2), 24.","journal-title":"Machines"},{"key":"6810_CR26","doi-asserted-by":"publisher","first-page":"515","DOI":"10.1613\/jair.3761","volume":"45","author":"J Garcia","year":"2012","unstructured":"Garcia, J., & Fern\u00e1ndez, F. (2012). Safe exploration of state and action spaces in reinforcement learning. Journal of Artificial Intelligence Research, 45, 515\u2013564.","journal-title":"Journal of Artificial Intelligence Research"},{"issue":"1","key":"6810_CR27","first-page":"1437","volume":"16","author":"J Garc\u0131a","year":"2015","unstructured":"Garc\u0131a, J., & Fern\u00e1ndez, F. (2015). A comprehensive survey on safe reinforcement learning. Journal of Machine Learning Research, 16(1), 1437\u20131480.","journal-title":"Journal of Machine Learning Research"},{"key":"6810_CR28","unstructured":"Gehring, C., & Precup, D. (2013). Smart exploration in reinforcement learning using absolute temporal difference errors. In Proceedings of the 2013 International Conference on Autonomous Agents and Multi-agent Systems (pp. 1037\u20131044)."},{"key":"6810_CR29","doi-asserted-by":"publisher","first-page":"81","DOI":"10.1613\/jair.1666","volume":"24","author":"P Geibel","year":"2005","unstructured":"Geibel, P., & Wysotzki, F. (2005). Risk-sensitive reinforcement learning applied to control under constraints. Journal of Artificial Intelligence Research, 24, 81\u2013108.","journal-title":"Journal of Artificial Intelligence Research"},{"key":"6810_CR30","doi-asserted-by":"publisher","first-page":"83","DOI":"10.1007\/s10846-013-9826-6","volume":"72","author":"A Geramifard","year":"2013","unstructured":"Geramifard, A., Redding, J., & How, J. P. (2013). Intelligent cooperative control architecture: A framework for performance improvement using safe learning. Journal of Intelligent & Robotic Systems, 72, 83\u2013103.","journal-title":"Journal of Intelligent & Robotic Systems"},{"key":"6810_CR31","volume-title":"Deep learning","author":"I Goodfellow","year":"2016","unstructured":"Goodfellow, I. (2016). Deep learning. MIT press."},{"key":"6810_CR32","unstructured":"Haarnoja, T., Zhou, A., Abbeel, P., & Levine, S. (2018). Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In Proc. Int. Conf. Machine Learning (ICML) (pp. 1861\u20131870)."},{"key":"6810_CR33","unstructured":"Ha, D., Dai, A. M., & Le, Q. V. (2017). Hypernetworks. In Proc. Int. Conf. Learning Representations (ICLR)."},{"key":"6810_CR34","unstructured":"Hallak, N., Mertikopoulos, P., & Cevher, V. (2021). Regret minimization in stochastic non-convex learning via a proximal-gradient approach. In International Conference on Machine Learning (pp. 4008\u20134017). PMLR."},{"issue":"4","key":"6810_CR35","doi-asserted-by":"publisher","first-page":"647","DOI":"10.1109\/JSTSP.2015.2404790","volume":"9","author":"EC Hall","year":"2015","unstructured":"Hall, E. C., & Willett, R. M. (2015). Online convex optimization in dynamic environments. IEEE Journal of Selected Topics in Signal Processing, 9(4), 647\u2013662.","journal-title":"IEEE Journal of Selected Topics in Signal Processing"},{"issue":"3\u20134","key":"6810_CR36","doi-asserted-by":"publisher","first-page":"157","DOI":"10.1561\/2400000013","volume":"2","author":"E Hazan","year":"2016","unstructured":"Hazan, E. (2016). Introduction to online convex optimization. Foundations and Trends\u00ae in Optimization, 2(3\u20134), 157\u2013325.","journal-title":"Foundations and Trends\u00ae in Optimization"},{"key":"6810_CR37","doi-asserted-by":"crossref","unstructured":"Hazan, E., & Seshadhri, C. (2009). Efficient learning algorithms for changing environments. In Proc. Int. Conf. Machine Learning (ICML) (pp. 393\u2013400).","DOI":"10.1145\/1553374.1553425"},{"key":"6810_CR38","doi-asserted-by":"crossref","unstructured":"Heger, M. (1994). Consideration of risk in reinforcement learning. In Proc. Int. Conf. Machine Learning (ICML) (pp. 105\u2013111).","DOI":"10.1016\/B978-1-55860-335-6.50021-0"},{"issue":"2","key":"6810_CR39","doi-asserted-by":"publisher","first-page":"151","DOI":"10.1023\/A:1007424614876","volume":"32","author":"M Herbster","year":"1998","unstructured":"Herbster, M., & Warmuth, M. K. (1998). Tracking the best expert. Machine Learning, 32(2), 151\u2013178.","journal-title":"Machine Learning"},{"issue":"3","key":"6810_CR40","doi-asserted-by":"publisher","first-page":"291","DOI":"10.1016\/j.jcss.2004.10.016","volume":"71","author":"A Kalai","year":"2005","unstructured":"Kalai, A., & Vempala, S. (2005). Efficient algorithms for online decision problems. Journal of Computer and System Sciences, 71(3), 291\u2013307.","journal-title":"Journal of Computer and System Sciences"},{"key":"6810_CR41","unstructured":"Khattar, V., Ding, Y., Sel, B., Lavaei, J., & Jin, M. (2023). A CMDP-within-online framework for meta-safe reinforcement learning. In Proc. Int. Conf. Learning Representations (ICLR)."},{"key":"6810_CR42","doi-asserted-by":"publisher","first-page":"79","DOI":"10.1109\/OJCSYS.2023.3256305","volume":"2","author":"N Kochdumper","year":"2023","unstructured":"Kochdumper, N., Krasowski, H., Wang, X., Bak, S., & Althoff, M. (2023). Provably safe reinforcement learning via action projection using reachability analysis and polynomial zonotopes. IEEE Open Journal of Control Systems, 2, 79\u201392.","journal-title":"IEEE Open Journal of Control Systems"},{"issue":"5","key":"6810_CR43","doi-asserted-by":"publisher","first-page":"1105","DOI":"10.1109\/TCST.2008.2012116","volume":"17","author":"Y Kuwata","year":"2009","unstructured":"Kuwata, Y., Teo, J., Fiore, G., Karaman, S., Frazzoli, E., & How, J. P. (2009). Real-time motion planning with applications to autonomous urban driving. IEEE Transactions on Control Systems Technology, 17(5), 1105\u20131118.","journal-title":"IEEE Transactions on Control Systems Technology"},{"key":"6810_CR44","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511546877","volume-title":"Planning Algorithms","author":"SM LaValle","year":"2006","unstructured":"LaValle, S. M. (2006). Planning Algorithms. Cambridge University Press."},{"key":"6810_CR45","unstructured":"Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., & Wierstra, D. (2015). Continuous control with deep reinforcement learning. In Proc. Int. Conf. Learning Representations (ICLR)."},{"key":"6810_CR46","unstructured":"Lin, Q., Tang, B., Wu, Z., Yu, C., Mao, S., Xie, Q., Wang, X., & Wang, D. (2023). Safe offline reinforcement learning with real-time budget constraints. In Proc. Int. Conf. Machine Learning (ICML) (pp. 21127\u201321152). PMLR."},{"key":"6810_CR47","unstructured":"Liu, Z., Guo, Z., Yao, Y., Cen, Z., Yu, W., Zhang, T., & Zhao, D. (2023). Constrained decision transformer for offline safe reinforcement learning. In Proc. Int. Conf. Machine Learning (ICML) (pp. 21611\u201321630). PMLR."},{"key":"6810_CR48","first-page":"25621","volume":"34","author":"Y Luo","year":"2021","unstructured":"Luo, Y., & Ma, T. (2021). Learning barrier certificates: Towards safe reinforcement learning with zero training-time violations. Advances in Neural Information Processing Systems, 34, 25621\u201325632.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"6810_CR49","doi-asserted-by":"crossref","unstructured":"Mart\u00edn, H., Antonio, J., & Lope, J. (2009). Learning autonomous helicopter flight with evolutionary reinforcement learning. In International Conference on Computer Aided Systems Theory (pp. 75\u201382). Springer.","DOI":"10.1007\/978-3-642-04772-5_11"},{"key":"6810_CR50","doi-asserted-by":"crossref","unstructured":"Mercy, T., Van\u00a0Loock, W., & Pipeleers, G. (2016). Real-time motion planning in the presence of moving obstacles. In Proc. European Control Conf. (ECC) (pp. 1586\u20131591).","DOI":"10.1109\/ECC.2016.7810517"},{"key":"6810_CR51","doi-asserted-by":"publisher","first-page":"267","DOI":"10.1023\/A:1017940631555","volume":"49","author":"O Mihatsch","year":"2002","unstructured":"Mihatsch, O., & Neuneier, R. (2002). Risk-sensitive reinforcement learning. Machine Learning, 49, 267\u2013290.","journal-title":"Machine Learning"},{"key":"6810_CR52","unstructured":"Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., & Kavukcuoglu, K. (2016). Asynchronous methods for deep reinforcement learning. In Proc. Int. Conf. Machine Learning (ICML) (pp. 1928\u20131937)."},{"key":"6810_CR53","unstructured":"Moldovan, T. M., & Abbeel, P. (2012). Safe exploration in Markov decision processes. In Proc. Int. Conf. Machine Learning (ICML) (pp. 1451\u20131458). Omnipress."},{"issue":"5","key":"6810_CR54","doi-asserted-by":"publisher","first-page":"780","DOI":"10.1287\/opre.1050.0216","volume":"53","author":"A Nilim","year":"2005","unstructured":"Nilim, A., & El Ghaoui, L. (2005). Robust control of Markov decision processes with uncertain transition matrices. Operations Research, 53(5), 780\u2013798.","journal-title":"Operations Research"},{"key":"6810_CR55","unstructured":"Pong, V. H., Nair, A. V., Smith, L. M., Huang, C., & Levine, S. (2022). Offline meta-reinforcement learning with online self-supervision. In Proc. Int. Conf. Machine Learning (ICML) (pp. 17811\u201317829)."},{"key":"6810_CR56","unstructured":"Rakelly, K., Zhou, A., Finn, C., Levine, S., & Quillen, D. (2019). Efficient off-policy meta-reinforcement learning via probabilistic context variables. In Proc. Int. Conf Machine Learning (ICML) (pp. 5331\u20135340)."},{"key":"6810_CR57","unstructured":"Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347."},{"key":"6810_CR58","unstructured":"Schulman, J., Levine, S., Abbeel, P., Jordan, M., & Moritz, P. (2015). Trust region policy optimization. In Proc. Int. Conf. Machine Learning (ICML) (pp. 1889\u20131897)."},{"issue":"2","key":"6810_CR59","doi-asserted-by":"publisher","first-page":"3663","DOI":"10.1109\/LRA.2021.3063989","volume":"6","author":"YS Shao","year":"2021","unstructured":"Shao, Y. S., Chen, C., Kousik, S., & Vasudevan, R. (2021). Reachability-based trajectory safeguard (rts): A safe and fast reinforcement learning safety layer for continuous control. IEEE Robotics and Automation Letters, 6(2), 3663\u20133670.","journal-title":"IEEE Robotics and Automation Letters"},{"issue":"1","key":"6810_CR60","doi-asserted-by":"publisher","first-page":"166","DOI":"10.1007\/s12555-012-0119-9","volume":"10","author":"Y Song","year":"2012","unstructured":"Song, Y., Li, Y.-B., Li, C.-H., & Zhang, G. (2012). An efficient initialization approach of q-learning for mobile robots. International Journal of Control, Automation and Systems, 10(1), 166\u2013172.","journal-title":"International Journal of Control, Automation and Systems"},{"key":"6810_CR61","unstructured":"Tamar, A., Di Castro, D., & Mannor, S. (2012). Policy gradients with variance related risk criteria. In Proc. Int. Conf. Machine Learning (ICML) (pp. 387\u2013396)."},{"issue":"6\u20137","key":"6810_CR62","doi-asserted-by":"publisher","first-page":"716","DOI":"10.1016\/j.artint.2007.09.009","volume":"172","author":"AL Thomaz","year":"2008","unstructured":"Thomaz, A. L., & Breazeal, C. (2008). Teachable robots: Understanding human teaching behavior to build more effective robot learners. Artificial Intelligence, 172(6\u20137), 716\u2013737.","journal-title":"Artificial Intelligence"},{"key":"6810_CR63","unstructured":"Torrey, L., & Taylor, M. E. (2012). Help an agent out: Student\/teacher learning in sequential decision tasks. In Proceedings of the Adaptive and Learning Agents Workshop (at AAMAS-12) (pp. 41\u201348). Citeseer."},{"issue":"3","key":"6810_CR64","doi-asserted-by":"publisher","DOI":"10.1115\/1.4037782","volume":"140","author":"N Virani","year":"2018","unstructured":"Virani, N., Jha, D. K., Yuan, Z., Shekhawat, I., & Ray, A. (2018). Imitation of demonstrations using Bayesian filtering with nonparametric data-driven models. Journal of Dynamic Systems, Measurement, and Control, 140(3), Article 030906.","journal-title":"Journal of Dynamic Systems, Measurement, and Control"},{"issue":"1","key":"6810_CR65","doi-asserted-by":"publisher","first-page":"176","DOI":"10.1109\/TAC.2021.3049335","volume":"67","author":"KP Wabersich","year":"2021","unstructured":"Wabersich, K. P., Hewing, L., Carron, A., & Zeilinger, M. N. (2021). Probabilistic model predictive safety certification for learning-based control. IEEE Trans. Automatic Control, 67(1), 176\u2013188.","journal-title":"IEEE Trans. Automatic Control"},{"issue":"5","key":"6810_CR66","doi-asserted-by":"publisher","first-page":"2638","DOI":"10.1109\/TAC.2022.3175628","volume":"68","author":"KP Wabersich","year":"2022","unstructured":"Wabersich, K. P., & Zeilinger, M. N. (2022). Predictive control barrier functions: Enhanced safety mechanisms for learning-based control. IEEE Transactions on Automatic Control, 68(5), 2638\u20132651.","journal-title":"IEEE Transactions on Automatic Control"},{"key":"6810_CR67","unstructured":"Wagener, N. C., Boots, B., & Cheng, C.-A. (2021). Safe reinforcement learning using advantage-based intervention. In Proc. Int. Conf. Machine Learning (ICML) (pp. 10630\u201310640)."},{"key":"6810_CR68","unstructured":"Xu, T., Liang, Y., & Lan, G. (2021). Crpo: A new approach for safe reinforcement learning with convergence guarantee. In Proc. Int. Conf. Machine Learning (ICML) (pp. 11480\u201311491)."},{"key":"6810_CR69","unstructured":"Xu, S., & Zhu, M. (2023). Online constrained meta-learning: Provable guarantees for generalization. In Advances in Neural Information Processing Systems (Vol. 36, pp. 15531\u201315544)."},{"key":"6810_CR70","unstructured":"Xu, S., & Zhu, M. (2024). Meta-reinforcement learning with universal policy adaptation: Provable near-optimality under all-task optimum comparator. Neural Information Processing Systems."},{"key":"6810_CR71","unstructured":"Yao, Y., Liu, Z., Cen, Z., Zhu, J., Yu, W., Zhang, T., & Zhao, D. (2024). Constraint-conditioned policy optimization for versatile safe reinforcement learning. In Proc. Advances in Neural Information Processing Systems (NeurIPS) (Vol. 36, pp. 12555\u201312568)."},{"key":"6810_CR72","doi-asserted-by":"crossref","unstructured":"Yuan, Z., & Zhu, M. (2022). dSLAP: Distributed safe learning and planning for multi-robot systems. In Proc. IEEE Conf. Decision and Control (CDC) (pp. 5864\u20135869).","DOI":"10.1109\/CDC51059.2022.9992938"},{"key":"6810_CR73","doi-asserted-by":"crossref","unstructured":"Zhou, Z., Oguz, O. S., Leibold, M., & Buss, M. (2020). A general framework to increase safety of learning algorithms for dynamical systems based on region of attraction estimation. IEEE Transactions on Robotics","DOI":"10.1109\/TRO.2020.2992981"}],"container-title":["Machine Learning"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-025-06810-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10994-025-06810-4\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-025-06810-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,6]],"date-time":"2025-09-06T20:37:58Z","timestamp":1757191078000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10994-025-06810-4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,18]]},"references-count":73,"journal-issue":{"issue":"8","published-print":{"date-parts":[[2025,8]]}},"alternative-id":["6810"],"URL":"https:\/\/doi.org\/10.1007\/s10994-025-06810-4","relation":{},"ISSN":["0885-6125","1573-0565"],"issn-type":[{"type":"print","value":"0885-6125"},{"type":"electronic","value":"1573-0565"}],"subject":[],"published":{"date-parts":[[2025,6,18]]},"assertion":[{"value":"30 September 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"21 April 2025","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"27 May 2025","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"18 June 2025","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"173"}}