{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,29]],"date-time":"2026-05-29T22:47:03Z","timestamp":1780094823777,"version":"3.54.0"},"reference-count":17,"publisher":"MDPI AG","issue":"12","license":[{"start":{"date-parts":[[2018,12,7]],"date-time":"2018-12-07T00:00:00Z","timestamp":1544140800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61702413"],"award-info":[{"award-number":["61702413"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Technology Innovation Funds for the Ninth Academy of China Aerospace","award":["2016JY06"],"award-info":[{"award-number":["2016JY06"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>When a satellite performs complex tasks such as discarding a payload or capturing a non-cooperative target, it will encounter sudden changes in the attitude and mass parameters, causing unstable flying and rolling of the satellite. In such circumstances, the change of the movement and mass characteristics are unpredictable. Thus, the traditional attitude control methods are unable to stabilize the satellite since they are dependent on the mass parameters of the controlled object. In this paper, we proposed a reinforcement learning method to re-stabilize the attitude of a satellite under such circumstances. Specifically, we discretize the continuous control torque, and build a neural network model that can output the discretized control torque to control the satellite. A dynamics simulation environment of the satellite is built, and the deep Q Network algorithm is then performed to train the neural network in this simulation environment. The reward of the training is the stabilization of the satellite. Simulation experiments illustrate that, with the iteration of training progresses, the neural network model gradually learned to re-stabilize the attitude of a satellite after unknown disturbance. As a contrast, the traditional PD (Proportion Differential) controller was unable to re-stabilize the satellite due to its dependence on the mass parameters. The proposed method adopts self-learning to control satellite attitudes, shows considerable intelligence and certain universality, and has a strong application potential for future intelligent control of satellites performing complex space tasks.<\/jats:p>","DOI":"10.3390\/s18124331","type":"journal-article","created":{"date-parts":[[2018,12,10]],"date-time":"2018-12-10T03:36:41Z","timestamp":1544413001000},"page":"4331","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":30,"title":["Reinforcement Learning-Based Satellite Attitude Stabilization Method for Non-Cooperative Target Capturing"],"prefix":"10.3390","volume":"18","author":[{"given":"Zhong","family":"Ma","sequence":"first","affiliation":[{"name":"Xi\u2019an Microelectronics Technology Institute, Xi\u2019an 710065, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yuejiao","family":"Wang","sequence":"additional","affiliation":[{"name":"Xi\u2019an Microelectronics Technology Institute, Xi\u2019an 710065, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yidai","family":"Yang","sequence":"additional","affiliation":[{"name":"Xi\u2019an Microelectronics Technology Institute, Xi\u2019an 710065, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhuping","family":"Wang","sequence":"additional","affiliation":[{"name":"Xi\u2019an Microelectronics Technology Institute, Xi\u2019an 710065, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lei","family":"Tang","sequence":"additional","affiliation":[{"name":"Xi\u2019an Microelectronics Technology Institute, Xi\u2019an 710065, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Stephen","family":"Ackland","sequence":"additional","affiliation":[{"name":"Centre for Computational Intelligence, De Montfort University, Gateway House, Leicester LE1 9BH, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2018,12,7]]},"reference":[{"key":"ref_1","first-page":"639","article-title":"Acquisition-probability model of noncooperative maneuvering target detection in space","volume":"39","author":"Zhou","year":"2010","journal-title":"Infrared Laser Eng."},{"key":"ref_2","unstructured":"Wei, W.S. (2013). Research on the Parameters Identification and Attitude Tracking Coupling Control for the Mass Body Attached Spacecraft. [Ph.D. Thesis, Harbin Institute of Technology]."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Jiao, C.T., Liang, B., and Wang, X.Q. (2017, January 28\u201330). Adaptive reaction null-space control of dual-arm space robot for post-capture of non-cooperative target. Proceedings of the 29th Chinese Control and Decision Conference, Chongqing, China.","DOI":"10.1109\/CCDC.2017.7978151"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"59","DOI":"10.1016\/j.actaastro.2018.03.041","article-title":"Quasi-model free control for the post-capture operation of a non-cooperative target","volume":"147","author":"She","year":"2018","journal-title":"Acta Astronaut."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"900","DOI":"10.1109\/TCST.2008.2011888","article-title":"Adaptive control of uncertain hamiltonian multi-input multi-output systems: With application to spacecraft control","volume":"17","author":"Yoon","year":"2009","journal-title":"IEEE Trans. Control Syst. Technol."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"385","DOI":"10.2514\/2.4549","article-title":"Adaptive nonlinear control of multiple spacecraft formation flying","volume":"23","author":"Queiroz","year":"2015","journal-title":"J. Guidance Control Dyn."},{"key":"ref_7","first-page":"1906","article-title":"Adaptive sliding mode control of flexible spacecraft on input shaping","volume":"34","author":"Miao","year":"2013","journal-title":"Acta Aeronaut. Astronaut. Sin."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Hu, W.D. (2016). Fundamental Spacecraft Dynamics and Control, John Wiley Sons.","DOI":"10.1002\/9781119113034"},{"key":"ref_9","unstructured":"Meng, H.T. (2017). Research on Satellite Attitude Control Method Based on Backstepping. [Master Thesis, Bohai University]."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Huo, X., Zhang, A., Zhang, Z., and She, Z. (2016, January 1\u20134). A attitude control method for spacecraft considering actuator constraint and dynamics based backstepping. Proceedings of the 2016 Seventh International Conference on Intelligent Control and Information Processing, Siem Reap, Cambodia.","DOI":"10.1109\/ICICIP.2016.7885889"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"557","DOI":"10.1016\/j.actaastro.2018.05.046","article-title":"Adaptive prediction backstepping attitude control for liquid-filled micro-satellite with flexible appendages","volume":"152","author":"Huo","year":"2018","journal-title":"Acta Astronaut."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"156","DOI":"10.1109\/TSMCC.2007.913919","article-title":"A comprehensive survey of multiagent reinforcement learning","volume":"38","author":"Busoniu","year":"2008","journal-title":"IEEE Trans. Syst. Man Cybern. Part C Appl. Rev."},{"key":"ref_13","unstructured":"Zhu, Y., Mottaghi, R., Kolve, E., Lim, J.J., Gupta, A., Li, F.F., and Farhadi, A. (June, January 29). Target-driven visual navigation in indoor scenes using deep reinforcement learning. Proceedings of the IEEE International Conference on Robotics and Automation, Singapore."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Jin, B., Yang, J., Huang, X., and Khan, D. (2017, January 23\u201326). Deep deformable Q-Network: an extension of deep Q-Network. Proceedings of the International Conference on Web Intelligence, Leipzig, Germany.","DOI":"10.1145\/3106426.3109426"},{"key":"ref_15","unstructured":"Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. (2018, December 05). Playing Atari with Deep Reinforcement Learning, arXiv, Available online: https:\/\/arxiv.org\/abs\/1312.5602."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Wu, S., Liu, Y., Radice, G., and Tan, S. (2017). Autonomous pointing control of a large satellite antenna subject to parametric uncertainty. Sensors, 17.","DOI":"10.3390\/s17030560"},{"key":"ref_17","unstructured":"Kim, S.H., and Choi, H.L. (July, January 28). Convolutional neural network-based spacecraft attitude control for docking port alignment. Proceedings of the International Conference on Ubiquitous Robots and Ambient Intelligence, Jeju, Korea."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/18\/12\/4331\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T15:32:04Z","timestamp":1760196724000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/18\/12\/4331"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,12,7]]},"references-count":17,"journal-issue":{"issue":"12","published-online":{"date-parts":[[2018,12]]}},"alternative-id":["s18124331"],"URL":"https:\/\/doi.org\/10.3390\/s18124331","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,12,7]]}}}