{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T03:25:07Z","timestamp":1760239507183,"version":"build-2065373602"},"reference-count":38,"publisher":"MDPI AG","issue":"22","license":[{"start":{"date-parts":[[2020,11,23]],"date-time":"2020-11-23T00:00:00Z","timestamp":1606089600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100010661","name":"Horizon 2020 Framework Programme","doi-asserted-by":"publisher","award":["H2020-WIDESPREAD-2014-2"],"award-info":[{"award-number":["H2020-WIDESPREAD-2014-2"]}],"id":[{"id":"10.13039\/100010661","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>In scenarios, like critical public safety communication networks, On-Scene Available (OSA) user equipment (UE) may be only partially connected with the network infrastructure, e.g., due to physical damages or on-purpose deactivation by the authorities. In this work, we consider multi-hop Device-to-Device (D2D) communication in a hybrid infrastructure where OSA UEs connect to each other in a seamless manner in order to disseminate critical information to a deployed command center. The challenge that we address is to simultaneously keep the OSA UEs alive as long as possible and send the critical information to a final destination (e.g., a command center) as rapidly as possible, while considering the heterogeneous characteristics of the OSA UEs. We propose a dynamic adaptation approach based on machine learning to improve a joint energy-spectral efficiency (ESE). We apply a Q-learning scheme in a hybrid fashion (partially distributed and centralized) in learner agents (distributed OSA UEs) and scheduler agents (remote radio heads or RRHs), for which the next hop selection and RRH selection algorithms are proposed. Our simulation results show that the proposed dynamic adaptation approach outperforms the baseline system by approximately 67% in terms of joint energy-spectral efficiency, wherein the energy efficiency of the OSA UEs benefit from a gain of approximately 30%. Finally, the results show also that our proposed framework with C-RAN reduces latency by approximately 50% w.r.t. the baseline.<\/jats:p>","DOI":"10.3390\/s20226692","type":"journal-article","created":{"date-parts":[[2020,11,23]],"date-time":"2020-11-23T08:18:23Z","timestamp":1606119503000},"page":"6692","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Q-Learning Based Joint Energy-Spectral Efficiency Optimization in Multi-Hop Device-to-Device Communication"],"prefix":"10.3390","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9035-9046","authenticated-orcid":false,"given":"Muhidul Islam","family":"Khan","sequence":"first","affiliation":[{"name":"Thomas Johann Seebeck Department of Electronics, School of Information Technology, Tallinn University of Technology, Ehitajate tee 5, 19086 Tallinn, Estonia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Luca","family":"Reggiani","sequence":"additional","affiliation":[{"name":"Dipartimento di Electtronica e Informazione, Politecnico di Milano, Via Ponzio 34\/5, 20133 Milano, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1055-7959","authenticated-orcid":false,"given":"Muhammad Mahtab","family":"Alam","sequence":"additional","affiliation":[{"name":"Thomas Johann Seebeck Department of Electronics, School of Information Technology, Tallinn University of Technology, Ehitajate tee 5, 19086 Tallinn, Estonia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4667-621X","authenticated-orcid":false,"given":"Yannick","family":"Le Moullec","sequence":"additional","affiliation":[{"name":"Thomas Johann Seebeck Department of Electronics, School of Information Technology, Tallinn University of Technology, Ehitajate tee 5, 19086 Tallinn, Estonia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7371-0888","authenticated-orcid":false,"given":"Navuday","family":"Sharma","sequence":"additional","affiliation":[{"name":"Thomas Johann Seebeck Department of Electronics, School of Information Technology, Tallinn University of Technology, Ehitajate tee 5, 19086 Tallinn, Estonia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Elias","family":"Yaacoub","sequence":"additional","affiliation":[{"name":"Faculty of Computer Studies, Arab Open University, Beirut 2058 4518, Lebanon"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9288-0452","authenticated-orcid":false,"given":"Maurizio","family":"Magarini","sequence":"additional","affiliation":[{"name":"Dipartimento di Electtronica e Informazione, Politecnico di Milano, Via Ponzio 34\/5, 20133 Milano, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,11,23]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Gohil, A., Modi, H., and Patel, S.K. (2013, January 1\u20132). 5G technology of mobile communication: A survey. Proceedings of the Intelligent Systems and Signal Processing (ISSP), Anand, India.","DOI":"10.1109\/ISSP.2013.6526920"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"3970","DOI":"10.1109\/JSYST.2017.2773633","article-title":"5G D2D Networks: Techniques, Challenges, and Future Prospects","volume":"12","author":"Ansari","year":"2018","journal-title":"IEEE Syst. J."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Holma, H., Toskala, A., and Reunanen, J. (2016). LTE Small Cell Optimization: 3GPP Evolution to Release 13, John Wiley & Sons.","DOI":"10.1002\/9781118912560"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"40","DOI":"10.1109\/MCOM.2014.6807945","article-title":"An overview of 3GPP device-to-device proximity services","volume":"52","author":"Lin","year":"2014","journal-title":"IEEE Commun. Mag."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Alnoman, A., and Anpalagan, A. (2017, January 20\u201321). On D2D communications for public safety applications. Proceedings of the the 3rd Humanitarian Technology Conference (IHTC), Toronto, ON, Canada.","DOI":"10.1109\/IHTC.2017.8058172"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Babun, L., Y\u00fcrekli, A.I., and G\u00fcven\u00e7, I. (2015, January 26\u201329). Multi-hop and D2D communications for extending coverage in public safety scenarios. Proceedings of the 40th IEEE Conference on Local Computer Networks (LCN), Clearwater Beach, FL, USA.","DOI":"10.1109\/LCNW.2015.7365946"},{"key":"ref_7","unstructured":"Babun, L. (2015). Extended coverage for public safety and critical communications using multi-hop and D2D communications. [Master\u2019s Thesis, Florida International University]."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"14643","DOI":"10.1109\/ACCESS.2018.2793532","article-title":"Disaster Management Using D2D Communication with Power Transfer and Clustering Techniques","volume":"6","author":"Ali","year":"2018","journal-title":"IEEE Access"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"7331","DOI":"10.1109\/TCOMM.2019.2924010","article-title":"Wireless Networks Design in the Era of Deep Learning: Model-Based, AI-Based, or Both?","volume":"67","author":"Zappone","year":"2019","journal-title":"IEEE Trans. Commun."},{"key":"ref_10","unstructured":"Rawat, P., Haddad, M., and Altman, E. (December, January 30). Towards efficient disaster management: 5G and Device to Device communication. Proceedings of the 2nd International Conference on Information and Communication Technologies for Disaster Management (ICT-DM), Rennes, France."},{"key":"ref_11","unstructured":"Krishnamoorthy, S., and Agrawala, A. (September, January 30). M-Urgency: A next generation, context-aware public safety application. Proceedings of the 13th International Conference on Human Computer Interaction with Mobile Devices and Services, Stockholm, Sweden."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Boumard, S., Harjula, I., Kanstren, T., and Rantala, S.J. (2018, January 26\u201329). Comparison of Spectral and Energy Efficiency Metrics Using Measurements in a LTE-A Network. Proceedings of the 2018 Network Traffic Measurement and Analysis Conference (TMA), Vienna, Austria.","DOI":"10.23919\/TMA.2018.8506529"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"3251","DOI":"10.1109\/TWC.2019.2912596","article-title":"Mobile-Traffic-Aware Offloading for Energy- and Spectral-Efficient Large-Scale D2D-Enabled Cellular Networks","volume":"18","author":"Zhao","year":"2019","journal-title":"IEEE Trans. Wirel. Commun."},{"key":"ref_14","unstructured":"Zia, K., Javed, N., Sial, M.N., Ahmed, S., Iram, H., and Pirzada, A.A. (2018). A Survey of Conventional and Artificial Intelligence \/ Learning based Resource Allocation and Interference Mitigation Schemes in D2D Enabled Networks. arXiv."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Khan, M.I., Alam, M.M., Le Moullec, Y., and Yaacoub, E. (2017). Throughput-Aware Cooperative Reinforcement Learning for Adaptive Resource Allocation in Device-to-Device Communication. Futur. Internet, 9.","DOI":"10.3390\/fi9040072"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Khan, M.I., Alam, M.M., Le Moullec, Y., and Yaacoub, E. (2018, January 5\u20138). Cooperative Reinforcement Learning for Adaptive Power Allocation in Device-to-Device Communication. Proceedings of the IEEE 4th World Forum on Internet of Things (WF-IoT), Singapore.","DOI":"10.1109\/WF-IoT.2018.8355169"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Nie, S., Fan, Z., Zhao, M., Gu, X., and Zhang, L. (2016, January 4\u20137). Q-learning based power control algorithm for D2D communication. Proceedings of the 27th IEEE International Symposium on Personal, Indoor and Mobile Radio Communications, Valencia, Spain.","DOI":"10.1109\/PIMRC.2016.7794793"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"2078","DOI":"10.1049\/iet-com.2018.6028","article-title":"Weighted cooperative reinforcement learning-based energy-efficient autonomous resource selection strategy for underlay D2D communication","volume":"13","author":"Sharma","year":"2019","journal-title":"IET Commun."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Lhazmir, S., Kobbane, A., and Ben-Othman, J. (2018, January 25\u201329). Channel assignment for D2D communication: A regret matching based approach. Proceedings of the 14th International Wireless Communications and Mobile Computing Conference (IWCMC 2018), Limassol, Cyprus.","DOI":"10.1109\/IWCMC.2018.8450520"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Bennis, M., and Niyato, D. (2010, January 6\u201310). A Q-learning based approach to interference avoidance in self-organized femtocell networks. Proceedings of the 2010 IEEE Globecom Workshops, Miami, FL, USA.","DOI":"10.1109\/GLOCOMW.2010.5700414"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Park, H., and Lim, Y. (2020). Reinforcement Learning for Energy Optimization with 5G Communications in Vehicular Social Networks. Sensors, 20.","DOI":"10.3390\/s20082361"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Najla, M., Gesbert, D., Becvar, Z., and Mach, P. (2019, January 9\u201313). Machine Learning for Power Control in D2D Communication Based on Cellular Channel Gains. Proceedings of the 2019 IEEE Globecom Workshops (GC Wkshps), Waikoloa, HI, USA.","DOI":"10.1109\/GCWkshps45667.2019.9024549"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Chen, W., and Zheng, J. (2019, January 25\u201326). A Multi-agent Reinforcement Learning Based Power Control Algorithm for D2D Communication Underlaying Cellular Networks. Proceedings of the Artificial Intelligence for Communications and Networks (AICON 2019), Harbin, China.","DOI":"10.1007\/978-3-030-22971-9_7"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"3996","DOI":"10.1109\/TCOMM.2016.2593468","article-title":"An autonomous learning-based algorithm for joint channel and power level selection by D2D pairs in heterogeneous cellular networks","volume":"64","author":"Asheralieva","year":"2016","journal-title":"IEEE Trans. Commun."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"AlQerm, I., and Shihada, B. (2017, January 8\u201313). Enhanced machine learning scheme for energy efficient resource allocation in 5G heterogeneous cloud radio access networks. Proceedings of the 2017 IEEE 28th Annual International Symposium on Personal, Indoor, and Mobile Radio Communications (PIMRC), Montreal, QC, Canada.","DOI":"10.1109\/PIMRC.2017.8292227"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"428","DOI":"10.1109\/JIOT.2015.2497712","article-title":"Energy-efficient resource allocation for D2D communications underlaying cloud-RAN-based LTE-A networks","volume":"3","author":"Zhou","year":"2016","journal-title":"IEEE Internet Things J."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Zhang, J., Sun, Y., and Ng, D.W.K. (2017, January 21\u201325). Energy-efficient transmission for wireless powered D2D communication networks. Proceedings of the 2017 IEEE International Conference on Communications (ICC), Paris, France.","DOI":"10.1109\/ICC.2017.7996666"},{"key":"ref_28","unstructured":"Koenig, S., and Simmons, R.G. (1992). Complexity Analysis of Real-Time Reinforcement Learning Applied to Finding Shortest Paths in Deterministic Domains, Carnegie Mellon School of Computer Science. Technical Report."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"2079","DOI":"10.1007\/s11276-013-0592-y","article-title":"Reinforcement learning based routing in wireless mesh networks","volume":"19","author":"Boushaba","year":"2013","journal-title":"Wirel. Netw."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"279","DOI":"10.1007\/BF00992698","article-title":"Q-learning","volume":"8","author":"Watkins","year":"1992","journal-title":"Mach. Learn."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Sutton, R.S., and Barto, A.G. (1998). Reinforcement Learning: An Introduction, MIT Press.","DOI":"10.1109\/TNN.1998.712192"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"201","DOI":"10.1109\/TSMCC.2011.2106494","article-title":"Experience replay for real-time reinforcement learning control","volume":"42","author":"Adam","year":"2012","journal-title":"IEEE Trans. Syst. Man, Cybern. Part C (Appl. Rev.)"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Moradi, M. (2016, January 20\u201321). A centralized reinforcement learning method for multi-agent job scheduling in Grid. Proceedings of the 2016 6th International Conference on Computer and Knowledge Engineering (ICCKE), Mashhad, Iran.","DOI":"10.1109\/ICCKE.2016.7802135"},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"9","DOI":"10.1007\/BF00115009","article-title":"Learning to predict by the methods of temporal differences","volume":"3","author":"Sutton","year":"1988","journal-title":"Mach. Learn."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"237","DOI":"10.1613\/jair.301","article-title":"Reinforcement learning: A survey","volume":"4","author":"Kaelbling","year":"1996","journal-title":"J. Artif. Intell. Res."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"73","DOI":"10.3233\/IFS-2009-0416","article-title":"Experimental analysis on Sarsa (\u03bb) and Q (\u03bb) with different eligibility traces strategies","volume":"20","author":"Leng","year":"2009","journal-title":"J. Intell. Fuzzy Syst."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Rouil, R., Cintr\u00f3n, F.J., Ben Mosbah, A., and Gamboa, S. (2017, January 13\u201314). Implementation and Validation of an LTE D2D Model for ns-3. Proceedings of the Workshop on ns-3 (WNS3 2017), Porto, Portugal.","DOI":"10.1145\/3067665.3067668"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"He, Y., Luan, X., Wang, J., Feng, M., and Wu, J. (2014, January 16\u201319). Power allocation for D2D communications in heterogeneous networks. Proceedings of the 16th International Conference on Advanced Communication Technology (ICACT 2014), PyeongChang, Korea.","DOI":"10.1109\/ICACT.2014.6779117"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/22\/6692\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T10:36:01Z","timestamp":1760178961000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/22\/6692"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,11,23]]},"references-count":38,"journal-issue":{"issue":"22","published-online":{"date-parts":[[2020,11]]}},"alternative-id":["s20226692"],"URL":"https:\/\/doi.org\/10.3390\/s20226692","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2020,11,23]]}}}