{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,7]],"date-time":"2025-11-07T09:23:57Z","timestamp":1762507437000,"version":"build-2065373602"},"reference-count":27,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2018,7,2]],"date-time":"2018-07-02T00:00:00Z","timestamp":1530489600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Special Funding Project of China Postdoctoral Science Foundation","award":["2014T70967"],"award-info":[{"award-number":["2014T70967"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Cognitive Radio (CR) is a promising technology to overcome spectrum scarcity, which currently faces lots of unsolved problems. One of the critical challenges for setting up such systems is how to coordinate multiple protocol layers such as routing and spectrum access in a partially observable environment. In this paper, a deep reinforcement learning approach is adopted for solving above problem. Firstly, for the purpose of compressing huge action space in the cross-layer design problem, a novel concept named responsibility rating is introduced to help decide the transmission power of every Secondary User (SU). In order to deal with problem of dimension curse while reducing replay memory, the Prioritized Memories Deep Q-Network (PM-DQN) is proposed. Furthermore, PM-DQN is applied to solve the joint routing and resource allocation problem in cognitive radio ad hoc network for minimizing the transmission delay and power consumption. Simulation results illustrates that our proposed algorithm can reduce the end-to-end delay, packet loss ratio and estimation error while achieving higher energy efficiency compared with traditional algorithm.<\/jats:p>","DOI":"10.3390\/s18072119","type":"journal-article","created":{"date-parts":[[2018,7,2]],"date-time":"2018-07-02T10:56:52Z","timestamp":1530529012000},"page":"2119","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":23,"title":["A Kind of Joint Routing and Resource Allocation Scheme Based on Prioritized Memories-Deep Q Network for Cognitive Radio Ad Hoc Networks"],"prefix":"10.3390","volume":"18","author":[{"given":"Yihang","family":"Du","sequence":"first","affiliation":[{"name":"Electronic Countermeasure Institute, National University of Defense Technology, Shushan District, Hefei 230000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fan","family":"Zhang","sequence":"additional","affiliation":[{"name":"Science and Technology Research Bureau of Anhui Xinhua University, Shushan District, Hefei 230000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lei","family":"Xue","sequence":"additional","affiliation":[{"name":"Electronic Countermeasure Institute, National University of Defense Technology, Shushan District, Hefei 230000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2018,7,2]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"2155","DOI":"10.1109\/TMC.2012.178","article-title":"Stochastic Power Adaptation with Multiagent Reinforcement Learning for Cognitive Wireless Mesh Networks","volume":"12","author":"Chen","year":"2013","journal-title":"IEEE Trans. Mob. Comput."},{"key":"ref_2","first-page":"96","article-title":"A survey on distributed channel selection technique using surf algorithm for information transfer in multi-hop cognitive radio networks","volume":"1","author":"Ruby","year":"2014","journal-title":"Int. Conf. Comput. Sci. Comput. Intell."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Zareei, M., Islam, A.K.M., Baharun, S., Vargas-Rosales, C., Azpilicueta, L., and Mansoor, N. (2017). Medium Access Control Protocols for Cognitive Radio Ad Hoc Networks: A Survey. Sensors, 17.","DOI":"10.3390\/s17092136"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Singh, K., and Moh, S. (2017). An Energy-Efficient and Robust Multipath Routing Protocol for Cognitive Radio Ad Hoc Networks. Sensors, 17.","DOI":"10.3390\/s17092027"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"155","DOI":"10.1109\/JCN.2014.000026","article-title":"Cognitive Routing for Multi-Hop Mobile Cognitive Radio Ad Hoc Networks","volume":"16","author":"Lee","year":"2014","journal-title":"J. Commun. Netw."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"228","DOI":"10.1016\/j.adhoc.2010.06.009","article-title":"Routing in Cognitive Radio Networks: Challenges and Solutions","volume":"9","author":"Cesana","year":"2011","journal-title":"Ad Hoc Netw."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"1969","DOI":"10.1109\/TVT.2010.2045403","article-title":"Cross-Layer Routing and Dynamic Spectrum Allocation in Cognitive Radio Ad Hoc Networks","volume":"59","author":"Ding","year":"2010","journal-title":"IEEE Trans. Veh. Technol."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"1958","DOI":"10.1109\/JSAC.2012.121111","article-title":"Spectrum-Aware Opportunistic Routing in Multi-Hop Cognitive Radio Networks","volume":"30","author":"Liu","year":"2012","journal-title":"IEEE J. Sel. Areas Commun."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"10024","DOI":"10.1109\/TVT.2017.2743058","article-title":"Efficient Resource Allocation in Device-to-Device Communication Using Cognitive Radio Technology","volume":"66","author":"Sultana","year":"2017","journal-title":"IEEE Trans. Veh. Technol."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"186","DOI":"10.1109\/TWC.2013.112513.122082","article-title":"Joint Routing and Resource Allocation for Delay Minimization in Cognitive Radio Based Mesh Networks","volume":"13","author":"Mohamed","year":"2014","journal-title":"IEEE Trans. Wirel. Commun."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"7110","DOI":"10.1109\/TWC.2015.2464801","article-title":"Cooperative Routing for Underlay Cognitive Radio Networks Using Mutual-Information Accumulation","volume":"14","author":"Chen","year":"2015","journal-title":"IEEE Trans. Wirel. Commun."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"46","DOI":"10.1109\/98.760423","article-title":"A Review of Current Routing Protocols for Ad Hoc Mobile Wireless Networks","volume":"6","author":"Royer","year":"2002","journal-title":"IEEE Pers. Commun."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"1490","DOI":"10.1109\/TMC.2016.2601926","article-title":"A Distributed Learning Automata Scheme for Spectrum Management in Self-Organized Cognitive Radio Network","volume":"16","author":"Fahimi","year":"2017","journal-title":"IEEE Trans. Mob. Comput."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"817","DOI":"10.1109\/TMC.2015.2442529","article-title":"Distributed Heuristically Accelerated Q-learning for Robust Cognitive Spectrum Management in LTE Cellular Systems","volume":"15","author":"Morozs","year":"2016","journal-title":"IEEE Trans. Mob. Comput."},{"key":"ref_15","first-page":"1","article-title":"A reinforcement learning-based routing scheme for cognitive radio ad hoc networks","volume":"7","author":"Yau","year":"2014","journal-title":"Wirel. Mob. Netw. Conf."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"146","DOI":"10.1109\/MWC.2017.1600117","article-title":"Clustering and Reinforcement-Learning-Based Routing for Cognitive Radio Networks","volume":"24","author":"Saleem","year":"2017","journal-title":"IEEE Wirel. Commun."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"720","DOI":"10.1007\/s11036-014-0551-6","article-title":"Multi-Agent Reinforcement Learning Based Opportunistic Routing and Channel Assignment for Mobile Cognitive Radio Ad Hoc Network","volume":"19","author":"Barve","year":"2014","journal-title":"Mob. Netw. Appl."},{"key":"ref_18","unstructured":"Sutton, R., and Barto, A. (2005). An Introduction to Reinforcement Learning, MIT Press."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"34","DOI":"10.1126\/science.153.3731.34","article-title":"Dynamic programming","volume":"153","author":"Bellman","year":"1966","journal-title":"Science"},{"key":"ref_20","unstructured":"Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. (arXiv, 2013). Playing Atari with Deep Reinforcement Learning, arXiv."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"529","DOI":"10.1038\/nature14236","article-title":"Human-Level Control Through Deep Reinforcement Learning","volume":"518","author":"Mnih","year":"2015","journal-title":"Nature"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"610","DOI":"10.1214\/aoms\/1177706645","article-title":"A Note on the Generation of Random Normal Deviates","volume":"29","author":"Box","year":"1958","journal-title":"Ann. Math. Stat."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Wellens, M., Riihijarvi, J., and Mahonen, P. (2010, January 6\u20139). Evaluation of adaptive MAC-layer sensing in realistic spectrum occupancy scenarios. Proceedings of the 2010 IEEE Symposium on New Frontiers in Dynamic Spectrum, Singapore.","DOI":"10.1109\/DYSPAN.2010.5457888"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"533","DOI":"10.1109\/TMC.2007.70751","article-title":"Efficient Discovery of Spectrum Opportunities with MAC-Layer Sensing in Cognitive Radio Networks","volume":"7","author":"Kim","year":"2008","journal-title":"IEEE Trans. Mob. Comput."},{"key":"ref_25","unstructured":"Schaul, T., Quan, J., Antonoglou, I., and Silver, D. (arXiv, 2015). Prioritized Experience Replay, arXiv."},{"key":"ref_26","unstructured":"Andre, D., Friedman, N., and Parr, R. (1997, January 2\u20134). Generalized prioritized sweeping. Proceedings of the Advances in Neural Information Processing Systems 10, Denver, CO, USA."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Du, Z., Wang, W., Yan, Z., Dong, W., and Wang, W. (2017). Variable Admittance Control Based on Fuzzy Reinforcement Learning for Minimally Invasive Surgery Manipulator. Sensors, 17.","DOI":"10.3390\/s17040844"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/18\/7\/2119\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T15:10:59Z","timestamp":1760195459000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/18\/7\/2119"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,7,2]]},"references-count":27,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2018,7]]}},"alternative-id":["s18072119"],"URL":"https:\/\/doi.org\/10.3390\/s18072119","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2018,7,2]]}}}