{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,23]],"date-time":"2026-07-23T15:51:47Z","timestamp":1784821907852,"version":"3.55.0"},"reference-count":28,"publisher":"MDPI AG","issue":"24","license":[{"start":{"date-parts":[[2020,12,11]],"date-time":"2020-12-11T00:00:00Z","timestamp":1607644800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100003725","name":"National Research Foundation of Korea","doi-asserted-by":"publisher","award":["2019R1F1A1058716"],"award-info":[{"award-number":["2019R1F1A1058716"]}],"id":[{"id":"10.13039\/501100003725","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100003725","name":"National Research Foundation of Korea","doi-asserted-by":"publisher","award":["2020R1F1A1065109"],"award-info":[{"award-number":["2020R1F1A1065109"]}],"id":[{"id":"10.13039\/501100003725","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>In this paper, we consider a multiple-input multiple-output (MIMO)\u2014non-orthogonal multiple access (NOMA) system with reinforcement learning (RL). NOMA, which is a technique for increasing the spectrum efficiency, has been extensively studied in fifth-generation (5G) wireless communication systems. The application of MIMO to NOMA can result in an even higher spectral efficiency. Moreover, user pairing and power allocation problem are important techniques in NOMA. However, NOMA has a fundamental limitation of the high computational complexity due to rapidly changing radio channels. This limitation makes it difficult to utilize the characteristics of the channel and allocate radio resources efficiently. To reduce the computational complexity, we propose an RL-based joint user pairing and power allocation scheme. By applying Q-learning, we are able to perform user pairing and power allocation simultaneously, which reduces the computational complexity. The simulation results show that the proposed scheme achieves a sum rate similar to that achieved with the exhaustive search (ES).<\/jats:p>","DOI":"10.3390\/s20247094","type":"journal-article","created":{"date-parts":[[2020,12,13]],"date-time":"2020-12-13T23:39:36Z","timestamp":1607902776000},"page":"7094","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":16,"title":["Reinforcement Learning-Based Joint User Pairing and Power Allocation in MIMO-NOMA Systems"],"prefix":"10.3390","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7162-6911","authenticated-orcid":false,"given":"Jaehee","family":"Lee","sequence":"first","affiliation":[{"name":"Department of Electronic Engineering, Sogang University, Seoul 04107, Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7868-7197","authenticated-orcid":false,"given":"Jaewoo","family":"So","sequence":"additional","affiliation":[{"name":"Department of Electronic Engineering, Sogang University, Seoul 04107, Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2020,12,11]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Saito, Y., Kishiyama, Y., Benjebbour, A., Nakamura, T., Li, A., and Higuchi, K. (2013, January 2\u20135). Non-orthogonal multiple access (NOMA) for cellular future radio access. Proceedings of the 2013 IEEE 77th Vehicular Technology Conference (VTC Spring), Dresden, Germany.","DOI":"10.1109\/VTCSpring.2013.6692652"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"2294","DOI":"10.1109\/COMST.2018.2835558","article-title":"A Survey of Non-Orthogonal Multiple Access for 5G","volume":"20","author":"Dai","year":"2018","journal-title":"IEEE Commun. Surv. Tutor."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"660","DOI":"10.1109\/LCOMM.2018.2802488","article-title":"Improving the decoding threshold of tailbiting spatially coupled LDPC codes by energy shaping","volume":"22","author":"Jerkovits","year":"2018","journal-title":"IEEE Commun. Lett."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"85","DOI":"10.1109\/MVT.2019.2903343","article-title":"Outage-limit-approaching channel coding for future wireless communications:Root-protograph low-density parity-check codes","volume":"14","author":"Fang","year":"2019","journal-title":"IEEE Veh. Technol. Mag."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"537","DOI":"10.1109\/TWC.2015.2475746","article-title":"The application of MIMO to non-orthogonal multiple access","volume":"15","author":"Ding","year":"2016","journal-title":"IEEE Trans. Wirel. Commun."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"3697","DOI":"10.1109\/TWC.2018.2814048","article-title":"Joint user pairing and power allocation in virtual MIMO systems","volume":"17","author":"Jia","year":"2018","journal-title":"IEEE Trans. Wirel. Commun."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"788","DOI":"10.1109\/LCOMM.2017.2776206","article-title":"User pairing and pair scheduling in massive MIMO-NOMA systems","volume":"22","author":"Chen","year":"2018","journal-title":"IEEE Commun. Lett."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Sun, H., Xu, Y., and Hu, R.Q. (2016, January 15\u201318). A NOMA and MU-MIMO supported cellular network with underlaid D2D communications. Proceedings of the 2016 IEEE 83rd Vehicular Technology Conference (VTC Spring), Nanjing, China.","DOI":"10.1109\/VTCSpring.2016.7504086"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"405","DOI":"10.1109\/LWC.2015.2426709","article-title":"On the ergodic capacity of MIMO NOMA systems","volume":"4","author":"Sun","year":"2015","journal-title":"IEEE Wirel. Commun. Lett."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"1647","DOI":"10.1109\/LSP.2015.2417119","article-title":"Fairness for non-orthogonal multiple access in 5G systems","volume":"22","author":"Timotheou","year":"2015","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Guo, J., Wang, X., Yang, J., Zheng, J., and Zhao, B. (2016, January 4\u20138). User pairing and power allocation for downlink non-orthogonal multiple access. Proceedings of the IEEE Globecom Workshops (GC Wkshps), Washington, DC, USA.","DOI":"10.1109\/GLOCOMW.2016.7849074"},{"key":"ref_12","unstructured":"Liu, F., M\u00e4h\u00f6nen, P., and Petrova, M. (September, January 30). Proportional fairness-based user pairing and power allocation for non-orthogonal multiple access. Proceedings of the IEEE International Symposium on Personal, Indoor, and Mobile Radio Communication (PIMRC), Hong Kong, China."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"2224","DOI":"10.1109\/COMST.2019.2904897","article-title":"Deep learning in mobile and wireless networking: A survey","volume":"21","author":"Zhang","year":"2019","journal-title":"IEEE Commun. Surv. Tutor."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"3133","DOI":"10.1109\/COMST.2019.2916583","article-title":"Applications of deep reinforcement learning in communications and networking: A survey","volume":"21","author":"Luong","year":"2019","journal-title":"IEEE Commun. Surv. Tutor."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"720","DOI":"10.1109\/LCOMM.2018.2792019","article-title":"Deep learning-aided SCMA","volume":"22","author":"Kim","year":"2018","journal-title":"IEEE Commun. Lett."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"630","DOI":"10.1109\/TCOMM.2019.2947418","article-title":"Power allocation in cache-aided NOMA systems: Optimization and deep reinforcement learning approaches","volume":"68","author":"Doan","year":"2020","journal-title":"IEEE Trans. Commun."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"8440","DOI":"10.1109\/TVT.2018.2848294","article-title":"Deep learning for an effective nonorthogonal multiple access scheme","volume":"67","author":"Gui","year":"2018","journal-title":"IEEE Trans. Veh. Technol."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"3377","DOI":"10.1109\/TVT.2017.2782726","article-title":"Reinforcement learning-based NOMA power allocation in the presence of smart jamming","volume":"67","author":"Xiao","year":"2018","journal-title":"IEEE Trans. Veh. Technol."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Ye, P., Wang, Y., Li, J., and Xiao, L. (2020). Fast reinforcement learning for anti-jamming communications. arXiv.","DOI":"10.1109\/GLOBECOM42002.2020.9322486"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Wang, S., Lv, T., and Zhang, X. (2019, January 20\u201324). Multi-agent reinforcement learning-based user pairing in multi-carrier NOMA systems. Proceedings of the IEEE International Conference on Communications Workshops (ICC Workshops), Shanghai, China.","DOI":"10.1109\/ICCW.2019.8757016"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"2200","DOI":"10.1109\/JSAC.2019.2933762","article-title":"Joint power allocation and channel assignment for NOMA with deep reinforcement learning","volume":"37","author":"He","year":"2019","journal-title":"IEEE J. Sel. Areas Commun."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"257","DOI":"10.1109\/TCCN.2018.2809722","article-title":"Deep reinforcement learning for dynamic multichannel access in wireless networks","volume":"4","author":"Wang","year":"2018","journal-title":"IEEE Trans. Cognit. Commun. Netw."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"95032","DOI":"10.1109\/ACCESS.2020.2995456","article-title":"Multi-Agent Deep Learning for Multi-channel Access in Slotted Wireless Networks","volume":"8","author":"Mennes","year":"2020","journal-title":"IEEE Access"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Ahmed, K.I., and Hossain, E. (2019). A deep Q-learning methods for downlink power allocation in multi-cell networks. arXiv.","DOI":"10.1109\/MNET.2019.1900029"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"3414","DOI":"10.1109\/JSYST.2019.2937463","article-title":"Deep learning-based MIMO-NOMA with imperfect SIC decoding","volume":"14","author":"Kang","year":"2020","journal-title":"IEEE Syst. J."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"279","DOI":"10.1007\/BF00992698","article-title":"Technical note: Q-learning","volume":"8","author":"Watkins","year":"1992","journal-title":"Mach. Learn."},{"key":"ref_27","unstructured":"3GPP (2017). 3rd Generation Partnership Project; Technical Specification Group Radio Access Network; Study on New Radio Access Technology Physical Layer Aspects (Release 14), 3rd Generation Partnership Project (3GPP). Version 14.2.0; Technical Report (TR) 38.802."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"2181","DOI":"10.1109\/JSAC.2017.2725519","article-title":"A survey on non-orthogonal multiple access for 5G networks: Research challenges and future trends","volume":"35","author":"Ding","year":"2017","journal-title":"IEEE J. Sel. Areas Commun."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/24\/7094\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T10:43:42Z","timestamp":1760179422000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/24\/7094"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,12,11]]},"references-count":28,"journal-issue":{"issue":"24","published-online":{"date-parts":[[2020,12]]}},"alternative-id":["s20247094"],"URL":"https:\/\/doi.org\/10.3390\/s20247094","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,12,11]]}}}