{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,9]],"date-time":"2026-04-09T12:22:38Z","timestamp":1775737358575,"version":"3.50.1"},"reference-count":33,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2023,5,30]],"date-time":"2023-05-30T00:00:00Z","timestamp":1685404800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Basic Scientific Research Project of Heilongjiang Province","award":["2020-KYYWF-1003"],"award-info":[{"award-number":["2020-KYYWF-1003"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>To address the need for massive connections in Internet-of-Vehicle communications, local wireless networks utilize non-orthogonal multiple access (NOMA). Scholars have introduced deep reinforcement learning networks for user grouping and power allocation to reduce computational complexity. However, the traditional algorithm based on DQN (Deep Q-Network) still exhibits slow convergence speed and low training stability, while the uniform sampling method in the sample playback process suffers from low sampling efficiency. In order to address these issues, this paper proposes a user grouping and power allocation method for NOMA systems based on Prioritized Dueling DQN-DDPG joint optimization. Firstly, the paper introduces the user grouping network based on Dueling DQN, which considers both the state value and action value in the entire connection layer. The two values compete with each other, are summed up, and re-evaluated. The network significantly improves training stability and increases the convergence speed. Secondly, in this paper, a depth deterministic strategy gradient (DDPG) algorithm with symmetric properties is used. This algorithm works well for continuous action spaces and avoids the power quantization error because of the continuity of power value in the power allocation stage. Finally, the priority sampling based on TD-error (Temporal-difference error) is combined with the Dueling DQN network and DDPG network to ensure random sampling and improve the replay probability of important samples. Simulation results show that the proposed priority-based Dueling DQN-DDPG algorithm significantly improves the convergence speed of sample training. The research results of this paper provide a solid foundation for the following research content, which focuses on NOMA system resource allocation under the mobile user state.<\/jats:p>","DOI":"10.3390\/sym15061170","type":"journal-article","created":{"date-parts":[[2023,5,30]],"date-time":"2023-05-30T03:02:23Z","timestamp":1685415743000},"page":"1170","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["NOMA Resource Allocation Method Based on Prioritized Dueling DQN-DDPG Network"],"prefix":"10.3390","volume":"15","author":[{"given":"Yuan","family":"Liu","sequence":"first","affiliation":[{"name":"Electronic Engineering School, Heilongjiang University, Harbin 150001, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yue","family":"Li","sequence":"additional","affiliation":[{"name":"Electronic Engineering School, Heilongjiang University, Harbin 150001, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lin","family":"Li","sequence":"additional","affiliation":[{"name":"Electronic Engineering School, Heilongjiang University, Harbin 150001, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mengli","family":"He","sequence":"additional","affiliation":[{"name":"Electronic Engineering School, Heilongjiang University, Harbin 150001, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,5,30]]},"reference":[{"key":"ref_1","first-page":"551","article-title":"The 5G mobile communication: The development trends and its emerging key techniques","volume":"44","author":"You","year":"2014","journal-title":"Sci.-Sin. Inf."},{"key":"ref_2","first-page":"551","article-title":"Development trend and some key technologies of 5G mobile communication","volume":"44","author":"Yu","year":"2014","journal-title":"Sci. China Inf. Sci."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Goto, J., Nakamura, O., Yokomakura, K., Hamaguchi, Y., Ibi, S., and Sampei, S. (2014, January 14\u201317). A Frequency Domain Scheduling for Uplink Single Carrier Non-orthogonal Multiple Access with Iterative Interference Cancellation. Proceedings of the 2014 IEEE 80th Vehicular Technology Conference (VTC2014-Fall), Vancouver, BC, Canada.","DOI":"10.1109\/VTCFall.2014.6965813"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"40","DOI":"10.1109\/MWC.2018.1700099","article-title":"Resource allocation for downlink NOMA systems: Key techniques and open issues","volume":"25","author":"Islam","year":"2018","journal-title":"IEEE Wirel. Commun."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Zhang, H., Zhang, D.-K., Meng, W.-X., and Li, C. (2016, January 22\u201327). User pairing algorithm with SIC in non-orthogonal multiple access system. Proceedings of the International Conference on Communications, Kuala Lumpur, Malaysia.","DOI":"10.1109\/ICC.2016.7511620"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"1077","DOI":"10.1109\/TCOMM.2017.2650992","article-title":"Optimal Joint Power and Subcarrier Allocation for Full-Duplex Multicarrier Non-Orthogonal Multiple Access Systems","volume":"65","author":"Sun","year":"2017","journal-title":"IEEE Trans. Commun."},{"key":"ref_7","first-page":"1595","article-title":"Power allocation of NOMA system in Downlink","volume":"40","author":"Li","year":"2018","journal-title":"Syst. Eng. Electron."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"70","DOI":"10.1109\/TGCN.2022.3216209","article-title":"Energy-Efficient Backscatter-Assisted Coded Cooperative-NOMA for B5G Wireless Communications","volume":"7","author":"Asif","year":"2022","journal-title":"IEEE Trans. Green Commun. Netw."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"3496","DOI":"10.1109\/TCOMM.2019.2893304","article-title":"Energy Effcient Resource Allocation in Hybrid Non-Orthogonal Multiple Accrss Systems","volume":"67","author":"Shi","year":"2019","journal-title":"IEEE Trans. Commun."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"1351","DOI":"10.1109\/TVT.2018.2881314","article-title":"Joint energy effcient subchannel and power optimization for a downlink NOMA heterI ogeneous network","volume":"68","author":"Fang","year":"2019","journal-title":"IEEE Trans. Veh. Technol."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"876","DOI":"10.1109\/TVT.2019.2951822","article-title":"Deep Neural Network for Resource Management in NOMA Networks","volume":"69","author":"Yang","year":"2019","journal-title":"IEEE Trans. Veh. Technol."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"2200","DOI":"10.1109\/JSAC.2019.2933762","article-title":"Joint power allocation and channel assignment for NOMA with deep reinforcement learning","volume":"37","author":"He","year":"2019","journal-title":"IEEE J. Sel. Areas Commun."},{"key":"ref_13","unstructured":"Shamna, K.F., Siyad, C.I., Tamilselven, S., and Manoj, M.K. (2020, January 23\u201324). Deep Learning Aided NOMA for User Fairness in 5G. Proceedings of the 2020 7th International Conference on Smart Structures and Systems (ICSSS), Chennai, India."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Kumaresan, S.P., Tan, C.K., and Ng, Y.H. (2021). Deep Neural Network (DNN) for Efficient User Clustering and Power Allocation in Downlink Non-Orthogonal Multiple Access (NOMA) 5G Networks. Symmetry, 13.","DOI":"10.3390\/sym13081507"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"529","DOI":"10.1038\/nature14236","article-title":"Human-Level Control Through Deep Reinforcement Learning","volume":"518","author":"Mnih","year":"2015","journal-title":"Nature"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"5083","DOI":"10.1109\/TWC.2021.3065523","article-title":"Arumugam Nallanathan, Resource allocation in uplink NOMA-IoT networks: A reinforcement-learning approach","volume":"20","author":"Ahsan","year":"2021","journal-title":"IEEE Trans. Wirel. Commun."},{"key":"ref_17","first-page":"1995","article-title":"Dueling network architectures for deep reinforcement learning","volume":"48","author":"Wang","year":"2015","journal-title":"PMLR"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Zhang, S., Li, L., Yin, J., Liang, W., Li, X., Chen, W., and Han, Z. (2018, January 16\u201318). A dynamic power allocation scheme in power-domain NOMA using actor-critic reinforcement learning. Proceedings of the 2018 IEEE\/CIC International Conference on Communications in China (ICCC), Beijing, China.","DOI":"10.1109\/ICCChina.2018.8641248"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Liu, Y., Li, Y., Li, L., and He, M. (2022). NOMA Resource Allocation Method Based on Prioritized Dueling DQN-DDPG Network. Res. Sq., Preprint.","DOI":"10.21203\/rs.3.rs-2341741\/v1"},{"key":"ref_20","unstructured":"Schaul, T., Quan, J., Antonoglou, I., and Silver, D. (2015, January 7\u20139). Prioritized experience replay. Proceedings of the International Conference Learning, Representations, San Diego, CA, USA."},{"key":"ref_21","unstructured":"Lillicrap, T.P., Hunt, J.J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2015). Continuous control with deep reinforcement learning, in ICLR. arXiv."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"12872","DOI":"10.1109\/TVT.2021.3121217","article-title":"Learning-assisted user clustering in cell-free massive MIMO-NOMA networks","volume":"70","author":"Le","year":"2021","journal-title":"IEEE Trans. Veh. Technol."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"5379","DOI":"10.1109\/TII.2019.2947435","article-title":"NOMA-based resource allocation for cluster-based cognitive industrial internet of things","volume":"16","author":"Liu","year":"2019","journal-title":"IEEE Trans. Ind. Inform."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Wang, X., and Xu, Y. (2019, January 23\u201325). Energy-efcient resource allocation in uplink NOMA systems with deep reinforcement learning. Proceedings of the International Conference on Wireless Communications and Signal Processing (WCSP), Xi\u2019an, China.","DOI":"10.1109\/WCSP.2019.8927898"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"2215","DOI":"10.1109\/TSP.2020.2982786","article-title":"Joint subcarrier and power allocation in NOMA: Optimal and approximate algorithms","volume":"68","author":"Coupechoux","year":"2020","journal-title":"IEEE Trans. Signal Process."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"7279","DOI":"10.1109\/JIOT.2020.2982699","article-title":"DRL-based energy-efficient resource allocation frameworks for uplink NOMA systems","volume":"7","author":"Wang","year":"2020","journal-title":"IEEE Internet Things J."},{"key":"ref_27","first-page":"7842987","article-title":"Joint Time and Power Allocation Algorithm in NOMA Relaying Network","volume":"2019","author":"Cheng","year":"2019","journal-title":"Int. J. Antennas Propag."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"3377","DOI":"10.1109\/TVT.2017.2782726","article-title":"Reinforcement learning-based NOMA power allocation in the presence of smart jamming","volume":"67","author":"Xiao","year":"2017","journal-title":"IEEE Trans. Veh. Technol."},{"key":"ref_29","first-page":"5838186","article-title":"Power Allocation Intelligent Optimization for Mobile NOMA Communication System","volume":"2022","author":"Feng","year":"2022","journal-title":"Int. J. Antennas Propag."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"6255","DOI":"10.1109\/TWC.2020.3001736","article-title":"Power allocation in multi-user cellular networks: Deep reinforcement learning approaches","volume":"19","author":"Meng","year":"2020","journal-title":"IEEE Trans. Wirel. Commun."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"5734","DOI":"10.1109\/TVT.2021.3074892","article-title":"Uplink Power Control Framework Based on Reinforcement Learning for 5G Networks","volume":"70","author":"Neto","year":"2021","journal-title":"IEEE Trans. Veh. Technol."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"6070","DOI":"10.1109\/TCOMM.2020.3004524","article-title":"Deep Reinforcement Learning for Distributed Dynamic MISO Downlink-Beamforming Coordination","volume":"68","author":"Ge","year":"2020","journal-title":"IEEE Trans. Commun."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"1581","DOI":"10.1109\/TCOMM.2019.2961332","article-title":"Deep Reinforcement Learning for 5G Networks: Joint Beamforming, Power Control, and Interference Coordination","volume":"68","author":"Mismar","year":"2020","journal-title":"IEEE Trans. Commun."}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/15\/6\/1170\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T19:44:54Z","timestamp":1760125494000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/15\/6\/1170"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,5,30]]},"references-count":33,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2023,6]]}},"alternative-id":["sym15061170"],"URL":"https:\/\/doi.org\/10.3390\/sym15061170","relation":{},"ISSN":["2073-8994"],"issn-type":[{"value":"2073-8994","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,5,30]]}}}