{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,30]],"date-time":"2026-07-30T13:01:07Z","timestamp":1785416467759,"version":"3.56.0"},"reference-count":32,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2026,4,4]],"date-time":"2026-04-04T00:00:00Z","timestamp":1775260800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100002858","name":"China Postdoctoral Science Foundation","doi-asserted-by":"publisher","award":["2024M764088"],"award-info":[{"award-number":["2024M764088"]}],"id":[{"id":"10.13039\/501100002858","id-type":"DOI","asserted-by":"publisher"}]},{"award":["2024M764088"],"award-info":[{"award-number":["2024M764088"]}],"id":[{"id":"https:\/\/ror.org\/0426zh255","id-type":"ROR","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61971025"],"award-info":[{"award-number":["61971025"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62331002"],"award-info":[{"award-number":["62331002"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"award":["61971025"],"award-info":[{"award-number":["61971025"]}],"id":[{"id":"https:\/\/ror.org\/01h0zpd94","id-type":"ROR","asserted-by":"publisher"}]},{"award":["62331002"],"award-info":[{"award-number":["62331002"]}],"id":[{"id":"https:\/\/ror.org\/01h0zpd94","id-type":"ROR","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>With the rapid evolution toward 6G networks, ensuring robust physical layer security (PLS) in highly dynamic and heterogeneous wireless environments has become a key challenge. Traditional security methods often struggle to adapt to time-varying channels, especially in the absence of perfect channel state information. Furthermore, the dynamic nature of node selection and power allocation in heterogeneous networks creates a complex hybrid action space operating across multiple timescales, significantly complicating the design of efficient and adaptive security strategies. To address this, this paper proposes a novel constrained hierarchical reinforcement learning (CHRL) framework for secure cooperative communications in next-generation wireless systems. The framework is designed to optimize secrecy performance within a hybrid action space comprising both discrete node selection and continuous power allocation, operating at different timescales. By hierarchically decoupling the joint optimization problem, the upper layer performs risk-aware node selection to maximize long-term secrecy capacity (SC) while guaranteeing a stable and secure link. At the lower layer, we develop a constrained MiniMax Multi-objective Deep Deterministic Policy Gradient (M3DDPG) algorithm that optimizes power allocation considering worst-case conditions. Lagrange multipliers are integrated to enforce a strictly positive SC constraint throughout transmission, effectively preventing security outages. Simulation results under time-varying Rayleigh fading channels demonstrate that the proposed CHRL framework outperforms existing HRL methods, achieving up to 17% improvement in SC while strictly maintaining security constraints. These results validate the effectiveness of the proposed approach for enhancing PLS in next-generation cooperative wireless networks.<\/jats:p>","DOI":"10.3390\/e28040412","type":"journal-article","created":{"date-parts":[[2026,4,6]],"date-time":"2026-04-06T03:11:28Z","timestamp":1775445088000},"page":"412","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Secure Cooperative Communications in 6G Networks: A Constrained Hierarchical Reinforcement Learning Framework with Hybrid Action Space"],"prefix":"10.3390","volume":"28","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3627-7664","authenticated-orcid":false,"given":"Xiaosi","family":"Tian","sequence":"first","affiliation":[{"name":"School of Electronic and Information Engineering, Beihang University, Beijing 100191, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1328-7739","authenticated-orcid":false,"given":"Zulin","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Electronic and Information Engineering, Beihang University, Beijing 100191, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yuanhan","family":"Ni","sequence":"additional","affiliation":[{"name":"School of Electronic and Information Engineering, Beihang University, Beijing 100191, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2026,4,4]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"134","DOI":"10.1109\/MNET.001.1900287","article-title":"A Vision of 6G Wireless Systems: Applications, Trends, Technologies, and Open Research Problems","volume":"34","author":"Saad","year":"2020","journal-title":"IEEE Netw."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Bolgouras, V., Farao, A., and Xenakis, C. (2024, January 21\u201323). Roadmap to Secure 6G Networks. Proceedings of the International Workshop on Computer Aided Modeling and Design of Communication Links and Networks, Athens, Greece.","DOI":"10.1109\/CAMAD62243.2024.10942738"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1550","DOI":"10.1109\/SURV.2014.012314.00178","article-title":"Principles of Physical Layer Security in Multiuser Wireless Networks: A Survey","volume":"16","author":"Mukherjee","year":"2014","journal-title":"IEEE Commun. Surv. Tutor."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"4","DOI":"10.1109\/MWC.2019.8883122","article-title":"Safeguarding 5G-and-Beyond Networks with Physical Layer Security","volume":"26","author":"Wu","year":"2019","journal-title":"IEEE Wirel. Commun."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"2515","DOI":"10.1109\/TIT.2008.921908","article-title":"Wireless Information-Theoretic Security","volume":"54","author":"Bloch","year":"2008","journal-title":"IEEE Trans. Inf. Theory"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"9448","DOI":"10.1109\/TVT.2017.2703305","article-title":"Strategic Antieavesdropping Game for Physical Layer Security in Wireless Cooperative Networks","volume":"66","author":"Wang","year":"2017","journal-title":"IEEE Trans. Veh. Technol."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"9906","DOI":"10.1109\/TWC.2022.3180395","article-title":"Eavesdropping and Anti-Eavesdropping Game in UAV Wiretap System: A Differential Game Approach","volume":"21","author":"Wu","year":"2022","journal-title":"IEEE Trans. Wirel. Commun."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"3806","DOI":"10.1109\/TSP.2016.2551693","article-title":"Dynamic Potential Games With Constraints: Fundamentals and Applications in Communications","volume":"64","author":"Zazo","year":"2016","journal-title":"IEEE Trans. Signal Process"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"320","DOI":"10.1109\/LWC.2020.3029797","article-title":"Achievable Secure Degrees of Freedom of MIMO Interference Channel with a Cooperative Jammer","volume":"10","author":"Wang","year":"2021","journal-title":"IEEE Wirel. Commun. Lett."},{"key":"ref_10","unstructured":"Qu, J., Cai, Y., Lu, J., Wang, A., Zheng, J., Yang, W., and Weng, N. (2014, January 8\u201310). Power allocation based on Stackelberg game in a jammer-assisted secure network. Proceedings of the International Conference on Cyberspace Technology, Beijing, China."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"401","DOI":"10.1109\/TWC.2015.2474378","article-title":"Secure Communication With a Wireless-Powered Friendly Jammer","volume":"15","author":"Liu","year":"2016","journal-title":"IEEE Trans. Wirel. Commun."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"28","DOI":"10.1109\/TWC.2015.2466091","article-title":"Optimal Power Allocation for Physical Layer Security in Multi-Hop DF Relay Networks","volume":"15","author":"Lee","year":"2016","journal-title":"IEEE Trans. Wirel. Commun."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"1617","DOI":"10.1109\/JSYST.2024.3423012","article-title":"Performance Analysis and Secure Resource Allocation for Relay-Aided MISO-NOMA Systems","volume":"18","author":"Li","year":"2024","journal-title":"IEEE Syst. J."},{"key":"ref_14","unstructured":"Charalambous, C.D., and Menemenlis, N. (1999, January 7\u201310). Stochastic models for short-term multipath fading channels: Chi-square and Ornstein-Uhlenbeck processes. Proceedings of the 38th IEEE Conference on Decision and Control, Phoenix, AZ, USA."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"1478","DOI":"10.1109\/TCOMM.2007.902531","article-title":"Stochastic Differential Equation Theory Applied to Wireless Channels","volume":"55","author":"Feng","year":"2007","journal-title":"IEEE Trans. Commun."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"4040","DOI":"10.1109\/TWC.2025.3536615","article-title":"Stochastic Differential Equations for Performance Analysis of Wireless Communication Systems","volume":"24","author":"Tempone","year":"2025","journal-title":"IEEE Trans. Wirel. Commun."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Kurniawan, F., Agustian, H., Dermawan, D., Nurdin, R., Ahmadi, N., and Dinaryanto, O. (2025). Hybrid Rule-Based and Reinforcement Learning for Urban Signal Control in Developing Cities: A Systematic Literature Review and Practice Recommendations for Indonesia. Appl. Sci., 15.","DOI":"10.3390\/app151910761"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"82318","DOI":"10.1109\/ACCESS.2024.3412055","article-title":"Joint Design of Adaptive Modulation and Precoding for Physical Layer Security in Visible Light Communications Using Reinforcement Learning","volume":"12","author":"Hoang","year":"2024","journal-title":"IEEE Access"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Hoseini, S.A., Bouhafs, F., Aboutorab, N., Sadeghi, P., and Hartog, F.d. (2023, January 4\u20138). Cooperative Jamming for Physical Layer Security Enhancement Using Deep Reinforcement Learning. Proceedings of the IEEE Globecom Workshops, Kuala Lumpur, Malaysia.","DOI":"10.1109\/GCWkshps58843.2023.10465104"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"9966","DOI":"10.1109\/TITS.2023.3271642","article-title":"Safe-State Enhancement Method for Autonomous Driving via Direct Hierarchical Reinforcement Learning","volume":"24","author":"Gu","year":"2023","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"3170","DOI":"10.54517\/cte3170","article-title":"A survey on the applications of machine learning, deep learning, and reinforcement learning in wireless communications","volume":"3","author":"Tran","year":"2025","journal-title":"Comput. Telecommun. Eng."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"4101","DOI":"10.1109\/TIFS.2021.3103062","article-title":"Multi-Agent Reinforcement Learning-Based Buffer-Aided Relay Selection in IRS-Assisted Secure Cooperative Networks","volume":"16","author":"Huang","year":"2021","journal-title":"IEEE Trans. Inf. Forensics Secur."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"732","DOI":"10.1109\/TIFS.2022.3149396","article-title":"Safe Exploration in Wireless Security: A Safe Reinforcement Learning Algorithm With Hierarchical Structure","volume":"17","author":"Lu","year":"2022","journal-title":"IEEE Trans. Inf. Forensics Secur."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"1616","DOI":"10.1109\/LWC.2020.2999333","article-title":"Dynamic Spectrum Anti-Jamming in Broadband Communications: A Hierarchical Deep Reinforcement Learning Approach","volume":"9","author":"Li","year":"2020","journal-title":"IEEE Wirel. Commun. Lett."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"171","DOI":"10.1109\/TCOMM.2021.3119689","article-title":"Hierarchical Reinforcement Learning for Relay Selection and Power Optimization in Two-Hop Cooperative Relay Network","volume":"70","author":"Geng","year":"2022","journal-title":"IEEE Trans. Commun."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"75818","DOI":"10.1109\/ACCESS.2024.3406949","article-title":"Hierarchical Reinforcement Learning Based Resource Allocation for RAN Slicing","volume":"12","author":"Hokelek","year":"2024","journal-title":"IEEE Access"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Qian, J., Li, H., Zhu, P., Zhou, A., Liu, S., and Wang, F. (2025). Cooperative Jamming and Relay Selection for Covert Communications Based on Reinforcement Learning. Sensors, 25.","DOI":"10.3390\/s25196218"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"577","DOI":"10.1109\/LWC.2015.2466678","article-title":"Relay Selection for Security-Constrained Cooperative Communication in the Presence of Eavesdropper\u2019s Overhearing and Interference","volume":"4","author":"Vahidian","year":"2015","journal-title":"IEEE Wirel. Commun. Lett."},{"key":"ref_29","unstructured":"Li, S., Wu, Y., Cui, X., Dong, H., and Russell, S. (February, January 27). Robust Multi-Agent Reinforcement Learning via Minimax Deep Deterministic Policy Gradient. Proceedings of the AAAI Conference on Artificial Intelligence, Honolulu, HI, USA."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Wang, A., Tang, J., Luo, H., Wen, H., Han-Ho, P., and Chang, S.Y. (2023). Physical Layer Security Performance Prediction Based on Deep Learning. Proceedings of the 2023 2nd International Conference on Computing, Communication, Perception and Quantum Technology (CCPQT), IEEE.","DOI":"10.1109\/CCPQT60491.2023.00044"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Deng, D., Li, X., Zhao, M., Rabie, K.M., and Kharel, R. (2020). Deep Learning-Based Secure MIMO Communications with Imperfect CSI for Heterogeneous Networks. Sensors, 20.","DOI":"10.3390\/s20061730"},{"key":"ref_32","first-page":"62696","article-title":"Sample complexity of goal-conditioned hierarchical reinforcement learning","volume":"36","author":"Robert","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/28\/4\/412\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,8]],"date-time":"2026-04-08T04:19:53Z","timestamp":1775621993000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/28\/4\/412"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,4]]},"references-count":32,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2026,4]]}},"alternative-id":["e28040412"],"URL":"https:\/\/doi.org\/10.3390\/e28040412","relation":{},"ISSN":["1099-4300"],"issn-type":[{"value":"1099-4300","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,4]]}}}