{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T02:09:12Z","timestamp":1760234952071,"version":"build-2065373602"},"reference-count":25,"publisher":"MDPI AG","issue":"13","license":[{"start":{"date-parts":[[2021,6,27]],"date-time":"2021-06-27T00:00:00Z","timestamp":1624752000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>The demand for bandwidth-intensive and delay-sensitive services is surging daily with the development of 5G technology, resulting in fierce competition for scarce radio resources. Power domain Nonorthogonal Multiple Access (NOMA) technologies can dramatically improve system capacity and spectrum efficiency. Unlike existing NOMA scheduling that mainly focuses on fairness, this paper proposes a power control solution for uplink hybrid OMA and PD-NOMA in dual dynamic environments: dynamic and imperfect channel information together with the random user-specific hierarchical quality of service (QoS). This paper models the power control problem as a nonconvex stochastic, which aims to maximize system energy efficiency while guaranteeing hierarchical user QoS requirements. Then, the problem is formulated as a partially observable Markov decision process (POMDP). Owing to the difficulty of modeling time-varying scenes, the urgency of fast convergency, the adaptability in a dynamic environment, and the continuity of the variables, a Deep Reinforcement Learning (DRL)-based method is proposed. This paper also transforms the hierarchical QoS constraint under the NOMA serial interference cancellation (SIC) scene to fit DRL. The simulation results verify the effectiveness and robustness of the proposed algorithm under a dual uncertain environment. As compared with the baseline Particle Swarm Optimization algorithm (PSO), the proposed DRL-based method has demonstrated satisfying performance.<\/jats:p>","DOI":"10.3390\/s21134404","type":"journal-article","created":{"date-parts":[[2021,6,27]],"date-time":"2021-06-27T23:57:22Z","timestamp":1624838242000},"page":"4404","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Dual Dynamic Scheduling for Hierarchical QoS in Uplink-NOMA: A Reinforcement Learning Approach"],"prefix":"10.3390","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1015-4102","authenticated-orcid":false,"given":"Xiangjun","family":"Li","sequence":"first","affiliation":[{"name":"National Engineering Laboratory for Mobile Network Technologies, Beijing University of Posts and Telecommunications, Beijing 100876, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1720-220X","authenticated-orcid":false,"given":"Qimei","family":"Cui","sequence":"additional","affiliation":[{"name":"National Engineering Laboratory for Mobile Network Technologies, Beijing University of Posts and Telecommunications, Beijing 100876, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0533-2528","authenticated-orcid":false,"given":"Jinli","family":"Zhai","sequence":"additional","affiliation":[{"name":"National Engineering Laboratory for Mobile Network Technologies, Beijing University of Posts and Telecommunications, Beijing 100876, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xueqing","family":"Huang","sequence":"additional","affiliation":[{"name":"New York Institute of Technology, Old Westbury, NY 11568, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,6,27]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"2347","DOI":"10.1109\/JPROC.2017.2768666","article-title":"Nonorthogonal Multiple Access for 5G and Beyond","volume":"105","author":"Liu","year":"2017","journal-title":"Proc. IEEE"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"140302","DOI":"10.1007\/s11432-020-2986-8","article-title":"Application of NOMA for Cellular-Connected UAVs: Opportunities and Challenges","volume":"64","author":"New","year":"2021","journal-title":"Sci. China Inf. Sci."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"108057","DOI":"10.1016\/j.comnet.2021.108057","article-title":"A Deep Reinforcement Learning-Based Multi-Optimality Routing Scheme for Dynamic IoT Networks","volume":"192","author":"Cong","year":"2021","journal-title":"Comput. Netw."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"63","DOI":"10.1109\/MCOM.2019.1800644","article-title":"Stochastic Online Learning for Mobile Edge Computing: Learning from Changes","volume":"57","author":"Cui","year":"2019","journal-title":"IEEE Commun. Mag."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"78","DOI":"10.1109\/MWC.001.2000325","article-title":"Edge-Intelligence-Empowered, Unified Authentication and Trust Evaluation for Heterogeneous Beyond 5G Systems","volume":"28","author":"Cui","year":"2021","journal-title":"IEEE Wirel. Commun."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"3980","DOI":"10.1109\/TVT.2020.2972363","article-title":"Adaptive Bitrate Video Streaming in Non-Orthogonal Multiple Access Networks","volume":"69","author":"Zhang","year":"2020","journal-title":"IEEE Trans. Veh. Technol."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"140303","DOI":"10.1007\/s11432-020-2985-8","article-title":"Energy-Efficient Design for mmWave-Enabled NOMA-UAV Networks","volume":"64","author":"Pang","year":"2021","journal-title":"Sci. China Inf. Sci."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"2141","DOI":"10.1109\/TVT.2019.2960506","article-title":"Optimal Resource Allocation for Multicarrier NOMA in Short Packet Communications","volume":"69","author":"Chen","year":"2020","journal-title":"IEEE Trans. Veh. Technol."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"11388","DOI":"10.1109\/TVT.2017.2725641","article-title":"Cross-Layer Power Allocation in Nonorthogonal Multiple Access Systems for Statistical QoS Provisioning","volume":"66","author":"Liu","year":"2017","journal-title":"IEEE Trans. Veh. Technol."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"81783","DOI":"10.1109\/ACCESS.2019.2923713","article-title":"Joint User Clustering and Multi-Dimensional Resource Allocation in Downlink MIMO\u2013NOMA Networks","volume":"7","author":"Zhang","year":"2019","journal-title":"IEEE Access"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"169306","DOI":"10.1007\/s11432-020-3091-6","article-title":"Fairness-Improved and QoS-Guaranteed Resource Allocation for NOMA-Based S-IoT Network","volume":"64","author":"Jiao","year":"2021","journal-title":"Sci. China Inf. Sci."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"105800","DOI":"10.1109\/ACCESS.2019.2931657","article-title":"Energy-Efficient Resource Allocation With Hybrid TDMA\u2013NOMA for Cellular-Enabled Machine-to-Machine Communications","volume":"7","author":"Li","year":"2019","journal-title":"IEEE Access"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Adjif, M.A., Habachi, O., and Cances, J.P. (2019, January 15\u201318). Joint Channel Selection and Power Control for NOMA: A Multi-Armed Bandit Approach. Proceedings of the 2019 IEEE Wireless Communications and Networking Conference Workshop (WCNCW), Marrakech, Morocco.","DOI":"10.1109\/WCNCW.2019.8902878"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"1030","DOI":"10.1109\/LWC.2018.2845398","article-title":"Power Allocation and User Clustering for Uplink MC-NOMA in D2D Underlaid Cellular Networks","volume":"7","author":"Zheng","year":"2018","journal-title":"IEEE Wirel. Commun. Lett."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"1361","DOI":"10.1049\/el.2019.2095","article-title":"Game-theoretic Power Allocation Algorithm for Downlink NOMA Cellular System","volume":"55","author":"Aldebes","year":"2019","journal-title":"Electron. Lett."},{"key":"ref_16","unstructured":"Fujita, H., Fournier-Viger, P., Ali, M., and Sasaki, J. (2020). User Grouping and Power Allocation in NOMA Systems: A Reinforcement Learning-Based Solution. Trends in Artificial Intelligence Theory and Applications. Artificial Intelligence Practices, Springer International Publishing. Lecture Notes in Computer Science."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"1485","DOI":"10.1109\/OJCOMS.2020.3024778","article-title":"Unified User Association and Contract-Theoretic Resource Orchestration in NOMA Heterogeneous Wireless Networks","volume":"1","author":"Diamanti","year":"2020","journal-title":"IEEE Open J. Commun. Soc."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Giang, H.T.H., Hoan, T.N.K., Thanh, P.D., and Koo, I. (2020). Hybrid NOMA\/OMA-Based Dynamic Power Allocation Scheme Using Deep Reinforcement Learning in 5G Networks. Appl. Sci., 10.","DOI":"10.3390\/app10124236"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Wang, X., and Xu, Y. (2019, January 23\u201325). Energy-Efficient Resource Allocation in Uplink NOMA Systems with Deep Reinforcement Learning. Proceedings of the 2019 11th International Conference on Wireless Communications and Signal Processing (WCSP), Xi\u2019an, China.","DOI":"10.1109\/WCSP.2019.8927898"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Wang, S., Lv, T., Ni, W., Beaulieu, N.C., and Jay Guo, Y. (2021). Joint Resource Management for MC-NOMA: A Deep Reinforcement Learning Approach. IEEE Trans. Wirel. Commun.","DOI":"10.1109\/TWC.2021.3069240"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Huang, C., Chen, G., Gong, Y., Xu, P., Han, Z., and Chambers, J.A. (2021). Buffer-Aided Relay Selection for Cooperative Hybrid NOMA\/OMA Networks with Asynchronous Deep Reinforcement Learning. IEEE J. Sel. Areas Communi.","DOI":"10.1109\/JSAC.2021.3087225"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"710","DOI":"10.1109\/LWC.2020.3040402","article-title":"Deep Reinforcement Learning Based Dynamic User Access and Decode Order Selection for Uplink NOMA System With Imperfect SIC","volume":"10","author":"Shi","year":"2021","journal-title":"IEEE Wirel. Commun. Lett."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"7244","DOI":"10.1109\/TWC.2016.2599521","article-title":"A General Power Allocation Scheme to Guarantee Quality of Service in Downlink and Uplink NOMA Systems","volume":"15","author":"Yang","year":"2016","journal-title":"IEEE Trans. Wirel. Commun."},{"key":"ref_24","unstructured":"Lillicrap, T.P., Hunt, J.J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2019). Continuous Control with Deep Reinforcement Learning. arXiv."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Zhou, Y., Xu, X., Lu, N., and Xie, W. (2020, January 11\u201314). Research on Technical Scheme and Overhead Calculation of Dynamic Spectrum Sharing. Proceedings of the 2020 IEEE 6th International Conference on Computer and Communications (ICCC), Chengdu, China.","DOI":"10.1109\/ICCC51575.2020.9345099"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/13\/4404\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T06:25:15Z","timestamp":1760163915000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/13\/4404"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,6,27]]},"references-count":25,"journal-issue":{"issue":"13","published-online":{"date-parts":[[2021,7]]}},"alternative-id":["s21134404"],"URL":"https:\/\/doi.org\/10.3390\/s21134404","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2021,6,27]]}}}