{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,29]],"date-time":"2025-10-29T13:30:28Z","timestamp":1761744628542,"version":"build-2065373602"},"reference-count":31,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2017,11,1]],"date-time":"2017-11-01T00:00:00Z","timestamp":1509494400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Future Internet"],"abstract":"<jats:p>Device-to-device (D2D) communication is an essential feature for the future cellular networks as it increases spectrum efficiency by reusing resources between cellular and D2D users. However, the performance of the overall system can degrade if there is no proper control over interferences produced by the D2D users. Efficient resource allocation among D2D User equipments (UE) in a cellular network is desirable since it helps to provide a suitable interference management system. In this paper, we propose a cooperative reinforcement learning algorithm for adaptive resource allocation, which contributes to improving system throughput. In order to avoid selfish devices, which try to increase the throughput independently, we consider cooperation between devices as promising approach to significantly improve the overall system throughput. We impose cooperation by sharing the value function\/learned policies between devices and incorporating a neighboring factor. We incorporate the set of states with the appropriate number of system-defined variables, which increases the observation space and consequently improves the accuracy of the learning algorithm. Finally, we compare our work with existing distributed reinforcement learning and random allocation of resources. Simulation results show that the proposed resource allocation algorithm outperforms both existing methods while varying the number of D2D users and transmission power in terms of overall system throughput, as well as D2D throughput by proper Resource block (RB)-power level combination with fairness measure and improving the Quality of service (QoS) by efficient controlling of the interference level.<\/jats:p>","DOI":"10.3390\/fi9040072","type":"journal-article","created":{"date-parts":[[2017,11,1]],"date-time":"2017-11-01T16:01:19Z","timestamp":1509552079000},"page":"72","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":25,"title":["Throughput-Aware Cooperative Reinforcement Learning for Adaptive Resource Allocation in Device-to-Device Communication"],"prefix":"10.3390","volume":"9","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9035-9046","authenticated-orcid":false,"given":"Muhidul","family":"Khan","sequence":"first","affiliation":[{"name":"Thomas Johann Seeback Department of Electronics, School of Information Technology, Tallinn University of Technology, Ehitajate tee 5, 19086 Tallinn, Estonia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1055-7959","authenticated-orcid":false,"given":"Muhammad","family":"Alam","sequence":"additional","affiliation":[{"name":"Thomas Johann Seeback Department of Electronics, School of Information Technology, Tallinn University of Technology, Ehitajate tee 5, 19086 Tallinn, Estonia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4667-621X","authenticated-orcid":false,"given":"Yannick","family":"Moullec","sequence":"additional","affiliation":[{"name":"Thomas Johann Seeback Department of Electronics, School of Information Technology, Tallinn University of Technology, Ehitajate tee 5, 19086 Tallinn, Estonia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Elias","family":"Yaacoub","sequence":"additional","affiliation":[{"name":"Faculty of Computer Studies, Arab Open University, Omar Bayhoum Str. - Park Sector, Beirut 2058 4518, Lebanon"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2017,11,1]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Doppler, K., Rinne, M., Wijting, C., Ribeiro, C.B., and Hugl, K. (2009). Device-to-device communication as an underlay to LTE-advanced networks. IEEE Commun. Mag., 47.","DOI":"10.1109\/MCOM.2009.5350367"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Fodor, G., Dahlman, E., Mildh, G., Parkvall, S., Reider, N., Mikl\u00f3s, G., and Tur\u00e1nyi, Z. (2012). Design aspects of network assisted device-to-device communications. IEEE Commun. Mag., 50.","DOI":"10.1109\/MCOM.2012.6163598"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Xiao, X., Tao, X., and Lu, J. (2011, January 5\u20138). A QoS-aware power optimization scheme in OFDMA systems with integrated device-to-device (D2D) communications. Proceedings of the 2011 IEEE Vehicular Technology Conference (VTC Fall), San Francisco, CA, USA.","DOI":"10.1109\/VETECF.2011.6093182"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Khan, M.I., and Rinner, B. (2012, January 19\u201323). Resource coordination in wireless sensor networks by cooperative reinforcement learning. Proceedings of the 2012 IEEE International Conference on Pervasive Computing and Communications Workshops (PERCOM Workshops), Lugano, Switzerland.","DOI":"10.1109\/PerComW.2012.6197639"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"1260","DOI":"10.1109\/TCOMM.2016.2616138","article-title":"Energy efficient D2D communications in dynamic TDD systems","volume":"65","author":"Fu","year":"2017","journal-title":"IEEE Trans. Commun."},{"key":"ref_6","first-page":"1821084","article-title":"Network-Assisted Distributed Fairness-Aware Interference Coordination for Device-to-Device Communication Underlaid Cellular Networks","volume":"2017","author":"Boabang","year":"2017","journal-title":"Mob. Inf. Syst."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Kai, Y., and Zhu, H. (2015, January 8\u201312). In Proceedings of the Resource allocation for multiple-pair D2D communications in cellular networks. Proceedings of the 2015 IEEE International Conference on Communications (ICC), London, UK.","DOI":"10.1109\/ICC.2015.7248776"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"3541","DOI":"10.1109\/TCOMM.2013.071013.120787","article-title":"Device-to-device communications underlaying cellular networks","volume":"61","author":"Feng","year":"2013","journal-title":"IEEE Trans. Commun."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Zulhasnine, M., Huang, C., and Srinivasan, A. (2010, January 11\u201313). Efficient resource allocation for device-to-device communication underlaying LTE network. Proceedings of the 2010 IEEE 6th International Conference on Wireless and Mobile Computing, Networking and Communications (WiMob), Niagara Falls, NU, Canada.","DOI":"10.1109\/WIMOB.2010.5645039"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Zhao, J., Chai, K.K., Chen, Y., Schormans, J., and Alonso-Zarate, J. (2015). Joint mode selection and resource allocation for machine-type D2D links. Trans. Emerg. Telecommun. Technol.","DOI":"10.1002\/ett.3000"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"3995","DOI":"10.1109\/TWC.2011.100611.101684","article-title":"Capacity enhancement using an interference limited area for device-to-device uplink underlaying cellular networks","volume":"10","author":"Min","year":"2011","journal-title":"IEEE Trans. Wirel. Commun."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"3814","DOI":"10.1109\/TCOMM.2014.2363092","article-title":"Joint mode selection and resource allocation for device-to-device communications","volume":"62","author":"Yu","year":"2014","journal-title":"IEEE Trans. Commun."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"An, R., Sun, J., Zhao, S., and Shao, S. (2012, January 25\u201327). Resource allocation scheme for device-to-device communication underlying lte downlink network. Proceedings of the 2012 International Conference on Wireless Communications & Signal Processing (WCSP), Huangshan, China.","DOI":"10.1109\/WCSP.2012.6542986"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"530","DOI":"10.1109\/LCOMM.2016.2517012","article-title":"Adaptive Resource Sharing Algorithm for Device-to-Device Communications Underlaying Cellular Networks","volume":"20","author":"Esmat","year":"2016","journal-title":"IEEE Commun. Lett."},{"key":"ref_15","unstructured":"Wang, F., Song, L., Han, Z., Zhao, Q., and Wang, X. (2013, January 7\u201310). Joint scheduling and resource allocation for device-to-device underlay communication. Proceedings of the 2013 IEEE Wireless Communications and Networking Conference (WCNC), Shanghai, China."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"1519","DOI":"10.1109\/TWC.2014.2368151","article-title":"Pricing-based interference coordination for D2D communications in cellular networks","volume":"14","author":"Yin","year":"2015","journal-title":"IEEE Trans. Wirel. Commun."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Luo, Y., Shi, Z., Zhou, X., Liu, Q., and Yi, Q. (2014, January 19\u201321). Dynamic resource allocations based on Q-learning for D2D communication in cellular networks. Proceedings of the 2014 11th International Computer Conference on Wavelet Active Media Technology and Information Processing (ICCWAMTIP), Chengdu, China.","DOI":"10.1109\/ICCWAMTIP.2014.7073432"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Nie, S., Fan, Z., Zhao, M., Gu, X., and Zhang, L. (2016, January 4\u20138). Q-learning based power control algorithm for D2D communication. Proceedings of the 2016 IEEE 27th Annual International Symposium on Personal, Indoor, and Mobile Radio Communications (PIMRC), Valencia, Spain.","DOI":"10.1109\/PIMRC.2016.7794793"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"67","DOI":"10.1016\/j.icte.2015.09.005","article-title":"On the throughput gain of device-to-device communications","volume":"1","author":"Hwang","year":"2015","journal-title":"ICT Express"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Mehta, M., Aliu, O.G., Karandikar, A., and Imran, M.A. (arXiv, 2011). A self-organized resource allocation using inter-cell interference coordination (ICIC) in relay-assisted cellular networks, arXiv.","DOI":"10.21917\/ijct.2011.0043"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"274","DOI":"10.1109\/TMC.2014.2318700","article-title":"An evolutionary game for distributed resource allocation in self-organizing small cells","volume":"14","author":"Semasinghe","year":"2015","journal-title":"IEEE Trans. Mob. Comput."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"102","DOI":"10.1109\/4234.991146","article-title":"A general correlation model for shadow fading in mobile radio systems","volume":"6","author":"Graziosi","year":"2002","journal-title":"IEEE Commun. Lett."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Zulhasnine, M., Huang, C., and Srinivasan, A. (2010, January 6\u20139). Penalty function method for peer selection over wireless mesh network. Proceedings of the 2010 IEEE 72nd Vehicular Technology Conference Fall (VTC 2010-Fall), Ottawa, ON, Canada.","DOI":"10.1109\/VETECF.2010.5594219"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Khan, M.I. (2016). Resource-aware task scheduling by an adversarial bandit solver method in wireless sensor networks. EURASIP J. Wirel. Commun. Netw., 2016.","DOI":"10.1186\/s13638-015-0515-y"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"237","DOI":"10.1613\/jair.301","article-title":"Reinforcement learning: A survey","volume":"4","author":"Kaelbling","year":"1996","journal-title":"J. Artif. Intell. Res."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Sutton, R.S., and Barto, A.G. (1998). Reinforcement Learning: An Introduction, MIT Press.","DOI":"10.1109\/TNN.1998.712192"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"765182","DOI":"10.1155\/2014\/765182","article-title":"Performance analysis of resource-aware task scheduling methods in Wireless sensor networks","volume":"10","author":"Khan","year":"2014","journal-title":"Int. J. Distrib. Sensor Netw."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Chen, M., Chen, J., Ma, Y., Yu, T., and Wu, Z. (2015, January 10\u201312). Base station assisted device-to-device communications for content update network. Proceedings of the 2015 First International Conference on Computational Intelligence Theory, Systems and Applications (CCITSA), Yilan, Taiwan.","DOI":"10.1109\/CCITSA.2015.30"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Shah, K., and Kumar, M. (2007, January 8\u201311). Distributed independent reinforcement learning (DIRL) approach to resource management in wireless sensor networks. Proceedings of the 2007 IEEE International Conference on Mobile Adhoc and Sensor Systems (MASS), Pisa, Italy.","DOI":"10.1109\/MOBHOC.2007.4428658"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Khan, M.I., and Rinner, B. (2014, January 10\u201314). Energy-aware task scheduling in wireless sensor networks based on cooperative reinforcement learning. Proceedings of the 2014 IEEE International Conference on Communications Workshops (ICC), Sydney, NSW, Australia.","DOI":"10.1109\/ICCW.2014.6881310"},{"key":"ref_31","unstructured":"Jain, R., Chiu, D.M., and Hawe, W.R. (1984). A Quantitative Measure of Fairness and Discrimination for Resource Allocation in Shared Computer System, Digital Equipment Corporation."}],"container-title":["Future Internet"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-5903\/9\/4\/72\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T18:49:09Z","timestamp":1760208549000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-5903\/9\/4\/72"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017,11,1]]},"references-count":31,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2017,12]]}},"alternative-id":["fi9040072"],"URL":"https:\/\/doi.org\/10.3390\/fi9040072","relation":{},"ISSN":["1999-5903"],"issn-type":[{"type":"electronic","value":"1999-5903"}],"subject":[],"published":{"date-parts":[[2017,11,1]]}}}