{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,30]],"date-time":"2025-10-30T11:38:50Z","timestamp":1761824330448,"version":"build-2065373602"},"reference-count":29,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2021,3,4]],"date-time":"2021-03-04T00:00:00Z","timestamp":1614816000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Structureless communications such as Device-to-Device (D2D) relaying are undeniably of paramount importance to improving the performance of today\u2019s mobile networks. Such a communication paradigm requires a certain level of intelligence at the device level, thereby allowing it to interact with the environment and make proper decisions. However, decentralizing decision-making may induce paradoxical outcomes, resulting in a drop in performance, which sustains the design of self-organizing yet efficient systems. We propose that each device decides either to directly connect to the eNodeB or get access via another device through a D2D link. In the first part of this article, we describe a biform game framework to analyze the proposed self-organized system\u2019s performance, under pure and mixed strategies. We use two reinforcement learning (RL) algorithms, enabling devices to self-organize and learn their pure\/mixed equilibrium strategies in a fully distributed fashion. Decentralized RL algorithms are shown to play an important role in allowing devices to be self-organized and reach satisfactory performance with incomplete information or even under uncertainties. We point out through a simulation the importance of D2D relaying and assess how our learning schemes perform under slow\/fast channel fading.<\/jats:p>","DOI":"10.3390\/s21051755","type":"journal-article","created":{"date-parts":[[2021,3,5]],"date-time":"2021-03-05T00:39:07Z","timestamp":1614904747000},"page":"1755","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["D2D Mobile Relaying Meets NOMA\u2014Part II: A Reinforcement Learning Perspective"],"prefix":"10.3390","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1681-993X","authenticated-orcid":false,"given":"Safaa","family":"Driouech","sequence":"first","affiliation":[{"name":"NEST Research Group, LRI Lab., ENSEM, Hassan II University of Casablanca, Casablanca 20000, Morocco"},{"name":"Laoratoire de Reacherche en Informatique, Sorbonne Universit\u00e9, CNRS, LIP6, F-75005 Paris, France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9946-5761","authenticated-orcid":false,"given":"Essaid","family":"Sabir","sequence":"additional","affiliation":[{"name":"NEST Research Group, LRI Lab., ENSEM, Hassan II University of Casablanca, Casablanca 20000, Morocco"},{"name":"Department of Computer Science, University of Quebec at Montreal, Montreal, QC H2L 2C4, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0055-7867","authenticated-orcid":false,"given":"Mounir","family":"Ghogho","sequence":"additional","affiliation":[{"name":"TICLab, International University of Rabat, Rabat 11100, Morocco"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6630-5083","authenticated-orcid":false,"given":"El-Mehdi","family":"Amhoud","sequence":"additional","affiliation":[{"name":"School of Computer Science, Mohammed VI Polytechnic University, Ben Guerir 43150, Morocco"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,3,4]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"46317","DOI":"10.1109\/ACCESS.2019.2909490","article-title":"Quantum Machine Learning for 6G Communication Networks: State-of-the-Art and Vision for the Future","volume":"7","author":"Nawaz","year":"2019","journal-title":"IEEE Access"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Alsharif, M.H., Kelechi, A.H., Albreem, M.A., Chaudhry, S.A., Zia, M.S., and Kim, S. (2020). Sixth Generation (6G) Wireless Networks: Vision, Research Activities, Challenges and Potential Solutions. Symmetry, 12.","DOI":"10.3390\/sym12040676"},{"key":"ref_3","unstructured":"Ali, S., Saad, W., Rajatheva, N., Chang, K., Steinbach, D., Sliwa, B., Wietfeld, C., Mei, K., Shiri, H., and Zepernick, H.J. (2020). 6G White Paper on Machine Learning in Wireless Communication Networks. arXiv."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Dang, S., Amin, O., Shihada, B., and Alouini, M.S. (2019). What should 6G be?. arXiv.","DOI":"10.36227\/techrxiv.10247726.v2"},{"key":"ref_5","unstructured":"Aazhang, B., Ahokangas, P., Alves, H., Alouini, M.S., Beek, J., Benn, H., Bennis, M., Belfiore, J., Strinati, E., and Chen, F. (2019). Key Drivers and Research Challenges for 6G Ubiquitous Wireless Intelligence (White Paper), University of Oulu."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"134","DOI":"10.1109\/MNET.001.1900287","article-title":"A Vision of 6G Wireless Systems: Applications, Trends, Technologies, and Open Research Problems","volume":"34","author":"Saad","year":"2020","journal-title":"IEEE Netw."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"118","DOI":"10.1109\/MWC.001.1900488","article-title":"A Speculative Study on 6G","volume":"27","author":"Tariq","year":"2020","journal-title":"IEEE Wirel. Commun."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Piran, M., and Suh, D. (2019). Learning-Driven Wireless Communications, towards 6G. arXiv.","DOI":"10.1109\/iCCECE46942.2019.8941882"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Janbi, N., Katib, I., Albeshri, A., and Mehmood, R. (2020). Distributed Artificial Intelligence-as-a-Service (DAIaaS) for Smarter IoE and 6G Environments. Sensors, 20.","DOI":"10.3390\/s20205796"},{"key":"ref_10","unstructured":"Xiao, Y., Shi, G., and Krunz, M. (2020). Towards Ubiquitous AI in 6G with Federated Learning. arXiv."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"88","DOI":"10.1109\/MCOM.001.1900649","article-title":"Federated Learning for Edge Networks: Resource Optimization and Incentive Mechanism","volume":"58","author":"Khan","year":"2020","journal-title":"IEEE Commun. Mag."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Driouech, S., Sabir, E., Ghogho, M., and Amhoud, E.M. (2021). D2D Mobile Relaying Meets NOMA\u2014Part I: A Biform Game Analysis. Sensors, 21.","DOI":"10.3390\/s21030702"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Tembine, H. (2012). Distributed Strategic Learning for Wireless Engineers, CRC Press.","DOI":"10.1201\/b11896"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"2524","DOI":"10.3390\/e15072524","article-title":"A decentralized heuristic approach towards resource allocation in femtocell networks","volume":"15","author":"Shahid","year":"2013","journal-title":"Entropy"},{"key":"ref_15","unstructured":"Lu, X., and Schwartz, H.M. (July, January 29). Decentralized learning in two-player zero-sum games: A L-RI lagging anchor algorithm. Proceedings of the 2011 American Control Conference, San Francisco, CA, USA."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"41","DOI":"10.1109\/MCOM.2017.1600614","article-title":"Self-Organized Connected Objects: Rethinking QoS Provisioning for IoT Services","volume":"55","author":"Elhammouti","year":"2017","journal-title":"IEEE Commun. Mag."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Habachi, O., Meghdadi, V., Sabir, E., and Cances, J.P. (2020). Combined Beam Alignment and Power Allocation for NOMA-Empowered mmWave Communications, Springer International Publishing. Ubiquitous Networking.","DOI":"10.1007\/978-3-030-58008-7"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Al-Tous, H., and Barhumi, I. (2019, January 3\u20136). Distributed reinforcement learning algorithm for energy harvesting sensor networks. Proceedings of the 2019 IEEE International Black Sea Conference on Communications and Networking (BlackSeaCom), Sochi, Russia.","DOI":"10.1109\/BlackSeaCom.2019.8812862"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Park, H., and Lim, Y. (2020). Reinforcement Learning for Energy Optimization with 5G Communications in Vehicular Social Networks. Sensors, 20.","DOI":"10.3390\/s20082361"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Khan, M., Alam, M., Moullec, Y., and Yaacoub, E. (2017). Throughput-Aware Cooperative Reinforcement Learning for Adaptive Resource Allocation in Device-to-Device Communication. Future Internet, 9.","DOI":"10.3390\/fi9040072"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"6733","DOI":"10.1109\/ACCESS.2018.2890210","article-title":"A distributed multi-agent RL-based autonomous spectrum allocation scheme in D2D enabled multi-tier HetNets","volume":"7","author":"Zia","year":"2019","journal-title":"IEEE Access"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Li, Z., Guo, C., and Xuan, Y. (2019, January 9\u201313). A multi-agent deep reinforcement learning based spectrum allocation framework for D2D communications. Proceedings of the 2019 IEEE Global Communications Conference (GLOBECOM), Big Island, HI, USA.","DOI":"10.1109\/GLOBECOM38437.2019.9013763"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Handouf, S., Sabir, E., and Sadik, M. (2015). A pricing-based spectrum leasing framework with adaptive distributed learning for cognitive radio networks. International Symposium on Ubiquitous Networking, Springer.","DOI":"10.1007\/978-981-287-990-5_4"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Tathe, P.K., and Sharma, M. (2018, January 16\u201318). Dynamic actor-critic: Reinforcement learning based radio resource scheduling for LTE-advanced. Proceedings of the 2018 Fourth International Conference on Computing Communication Control and Automation (ICCUBEA), Pune, India.","DOI":"10.1109\/ICCUBEA.2018.8697808"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Sun, P., Li, J., Lan, J., Hu, Y., and Lu, X. (2018, January 7\u201310). RNN Deep Reinforcement Learning for Routing Optimization. Proceedings of the 2018 IEEE 4th International Conference on Computer and Communications (ICCC), Chengdu, China.","DOI":"10.1109\/CompComm.2018.8780950"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Khodayari, S., and Yazdanpanah, M.J. (2005, January 14\u201316). Network routing based on reinforcement learning in dynamically changing networks. Proceedings of the 17th IEEE International Conference on Tools with Artificial Intelligence (ICTAI\u201905), Hong Kong, China.","DOI":"10.1109\/ICTAI.2005.91"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"721","DOI":"10.1109\/COMST.2016.2621116","article-title":"Power-domain non-orthogonal multiple access (NOMA) in 5G systems: Potentials and challenges","volume":"19","author":"Islam","year":"2016","journal-title":"IEEE Commun. Surv. Tutor."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"2181","DOI":"10.1109\/JSAC.2017.2725519","article-title":"A survey on non-orthogonal multiple access for 5G networks: Research challenges and future trends","volume":"35","author":"Ding","year":"2017","journal-title":"IEEE J. Sel. Areas Commun."},{"key":"ref_29","unstructured":"Tabassum, H., Ali, M.S., Hossain, E., Hossain, M., and Kim, D.I. (2016). Non-orthogonal multiple access (NOMA) in cellular uplink and downlink: Challenges and enabling techniques. arXiv."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/5\/1755\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T05:32:26Z","timestamp":1760160746000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/5\/1755"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,3,4]]},"references-count":29,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2021,3]]}},"alternative-id":["s21051755"],"URL":"https:\/\/doi.org\/10.3390\/s21051755","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2021,3,4]]}}}