{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,15]],"date-time":"2026-01-15T10:21:23Z","timestamp":1768472483365,"version":"3.49.0"},"reference-count":39,"publisher":"MDPI AG","issue":"19","license":[{"start":{"date-parts":[[2021,10,6]],"date-time":"2021-10-06T00:00:00Z","timestamp":1633478400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["U1913602"],"award-info":[{"award-number":["U1913602"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"National Key R\\&amp;D Program","award":["2018AAA0102804"],"award-info":[{"award-number":["2018AAA0102804"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>Most multi-view based human pose estimation techniques assume the cameras are fixed. While in dynamic scenes, the cameras should be able to move and seek the best views to avoid occlusions and extract 3D information of the target collaboratively. In this paper, we address the problem of online view selection for a fixed number of cameras to estimate multi-person 3D poses actively. The proposed method exploits a distributed multi-agent based deep reinforcement learning framework, where each camera is modeled as an agent, to optimize the action of all the cameras. An inter-agent communication protocol was developed to transfer the cameras\u2019 relative positions between agents for better collaboration. Experiments on the Panoptic dataset show that our method outperforms other view selection methods by a large margin given an identical number of cameras. To the best of our knowledge, our method is the first to address online active multi-view 3D pose estimation with multi-agent reinforcement learning.<\/jats:p>","DOI":"10.3390\/rs13193995","type":"journal-article","created":{"date-parts":[[2021,10,8]],"date-time":"2021-10-08T21:26:20Z","timestamp":1633728380000},"page":"3995","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":9,"title":["Multi-Agent Deep Reinforcement Learning for Online 3D Human Poses Estimation"],"prefix":"10.3390","volume":"13","author":[{"given":"Zhen","family":"Fan","sequence":"first","affiliation":[{"name":"Graduate School at Shenzhen, Tsinghua University, Shenzhen 518057, China"},{"name":"Department of Automation, Tsinghua University, Beijing 100084, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0403-1923","authenticated-orcid":false,"given":"Xiu","family":"Li","sequence":"additional","affiliation":[{"name":"Graduate School at Shenzhen, Tsinghua University, Shenzhen 518057, China"},{"name":"Department of Automation, Tsinghua University, Beijing 100084, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yipeng","family":"Li","sequence":"additional","affiliation":[{"name":"Department of Automation, Tsinghua University, Beijing 100084, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,10,6]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"172","DOI":"10.1109\/TPAMI.2019.2929257","article-title":"OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields","volume":"43","author":"Cao","year":"2019","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Habibie, I., Xu, W., Mehta, D., Pons-Moll, G., and Theobalt, C. (2019, January 16\u201320). In the wild human pose estimation using explicit 2D features and intermediate 3D representations. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01116"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Iskakov, K., Burkov, E., Lempitsky, V., and Malkov, Y. (2019, January 27\u201328). Learnable triangulation of human pose. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Korea.","DOI":"10.1109\/ICCV.2019.00781"},{"key":"ref_4","unstructured":"Doersch, C., and Zisserman, A. (2019, January 8\u201314). Sim2real transfer learning for 3D human pose estimation: Motion to the rescue. Proceedings of the Annual Conference on Neural Information Processing Systems 2019, Vancouver, BC, Canada."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Rhodin, H., Sp\u00f6rri, J., Katircioglu, I., Constantin, V., Meyer, F., M\u00fcller, E., Salzmann, M., and Fua, P. (2018, January 18\u201322). Learning monocular 3d human pose estimation from multi-view images. Proceedings of the 2018 Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00880"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Kocabas, M., Karagoz, S., and Akbas, E. (2019, January 16\u201320). Self-supervised learning of 3d human pose using multi-view geometry. Proceedings of the 2019 Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00117"},{"key":"ref_7","unstructured":"Pirinen, A., G\u00e4rtner, E., and Sminchisescu, C. (2019, January 8\u201314). Domes to Drones: Self-Supervised Active Triangulation for 3D Human Pose Reconstruction. Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"G\u00e4rtner, E., Pirinen, A., and Sminchisescu, C. (2020, January 7\u201312). Deep Reinforcement Learning for Active Human Pose Estimation. Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA.","DOI":"10.1609\/aaai.v34i07.6714"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Joo, H., Liu, H., Tan, L., Gui, L., Nabbe, B., Matthews, I., Kanade, T., Nobuhara, S., and Sheikh, Y. (2015, January 7\u201313). Panoptic studio: A massively multiview system for social motion capture. Proceedings of the 2015 IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.381"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Dong, J., Jiang, W., Huang, Q., Bao, H., and Zhou, X. (2019, January 16\u201320). Fast and Robust Multi-Person 3D Pose Estimation From Multiple Views. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00798"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Zhang, Y., An, L., Yu, T., Li, X., Li, K., and Liu, Y. (2020, January 14\u201319). 4D Association Graph for Realtime Multi-person Motion Capture Using Multiple Video Cameras. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00140"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Mehta, D., Sotnychenko, O., Mueller, F., Xu, W., Sridhar, S., Pons-Moll, G., and Theobalt, C. (2018, January 5\u20138). Single-Shot Multi-Person 3D Pose Estimation From Monocular RGB. Proceedings of the International Conference on 3D Vision (3DV), Verona, Italy.","DOI":"10.1109\/3DV.2018.00024"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Rogez, G., Weinzaepfel, P., and Schmid, C. (2017, January 21\u201326). Lcr-net: Localization-classification-regression for human pose. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.134"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Tome, D., Russell, C., and Agapito, L. (2017, January 21\u201326). Lifting from the deep: Convolutional 3d pose estimation from a single image. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.603"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Bogo, F., Kanazawa, A., Lassner, C., Gehler, P., Romero, J., and Black, M.J. (2016, January 11\u201314). Keep it SMPL: Automatic estimation of 3D human pose and shape from a single image. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46454-1_34"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Papandreou, G., Zhu, T., Kanazawa, N., Toshev, A., Tompson, J., Bregler, C., and Murphy, K. (2017, January 21\u201326). Towards accurate multi-person pose estimation in the wild. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.395"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Pishchulin, L., Insafutdinov, E., Tang, S., Andres, B., Andriluka, M., Gehler, P., and Schiele, B. (2016, January 27\u201330). DeepCut: Joint Subset Partition and Labeling for Multi Person Pose Estimation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.533"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Newell, A., Yang, K., and Deng, J. (2016, January 11\u201314). Stacked hourglass networks for human pose estimation. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46484-8_29"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Wei, S.E., Ramakrishna, V., Kanade, T., and Sheikh, Y. (2016, January 27\u201330). Convolutional pose machines. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.511"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Fabbri, M., Lanzi, F., Calderara, S., Palazzi, A., Vezzani, R., and Cucchiara, R. (2018, January 8\u201314). Learning to Detect and Track Visible and Occluded Body Joints in a Virtual World. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01225-0_27"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Kiciroglu, S., Rhodin, H., Sinha, S.N., Salzmann, M., and Fua, P. (2019, January 16\u201320). ActiveMoCap: Optimized Drone Flight for Active Human Motion Capture. Proceedings of the 2019 Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR42600.2020.00018"},{"key":"ref_22","first-page":"4491","article-title":"Active Perception Based Formation Control for Multiple Aerial Vehicles","volume":"4","author":"Tallamraju","year":"2019","journal-title":"ICRA"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Saini, N., Price, E., Tallamraju, R., Enficiaud, R., Ludwig, R., Martinovic, I., Ahmad, A., and Black, M.J. (2019, January 27\u201328). Markerless Outdoor Human Motion Capture Using Multiple Autonomous Micro Aerial Vehicles. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Korea.","DOI":"10.1109\/ICCV.2019.00091"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Zhou, X., Liu, S., Pavlakos, G., Kumar, V., and Daniilidis, K. (2018, January 21\u201325). Human Motion Capture Using a Drone. Proceedings of the 2018 IEEE International Conference on Robotics and Automation, Brisbane, Australia.","DOI":"10.1109\/ICRA.2018.8462830"},{"key":"ref_25","first-page":"1","article-title":"Real-time Environment-independent Multi-view Human Pose Estimation with Aerial Vehicles","volume":"37","author":"Oberholzer","year":"2018","journal-title":"ACM Trans. Graph. (TOG)"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Huang, C., Gao, F., Pan, J., Yang, Z., Qiu, W., Chen, P., Yang, X., Shen, S., and Cheng, K.T. (2018, January 21\u201325). ACT: An Autonomous Drone Cinematography System for Action Scenes. Proceedings of the 2018 IEEE International Conference on Robotics and Automation, Brisbane, Australia.","DOI":"10.1109\/ICRA.2018.8460703"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Kaba, M.D., Uzunbas, M.G., and Lim, S.N. (2017, January 21\u201327). A reinforcement learning approach to the view planning problem. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.541"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Peralta, D., Casimiro, J., Nilles, A.M., Aguilar, J.A., Atienza, R., and Cajote, R. (2020, January 23\u201328). Next-best view policy for 3D reconstruction. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-66823-5_33"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Chen, L., Lu, J., Song, Z., and Zhou, J. (2018, January 8\u201314). Part-Activated Deep Reinforcement Learning for Action Prediction. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01219-9_26"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Wang, B., Adeli, E., Chiu, H., Huang, D., and Niebles, J.C. (2019, January 27\u201328). Imitation Learning for Human Pose Prediction. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Korea.","DOI":"10.1109\/ICCV.2019.00722"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Yuan, Y., and Kitani, K. (2019, January 27\u201328). Ego-Pose Estimation and Forecasting as Real-Time PD Control. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Seoul, Korea.","DOI":"10.1109\/ICCV.2019.01018"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Long, P., Fanl, T., Liao, X., Liu, W., Zhang, H., and Pan, J. (2018, January 21\u201325). Towards Optimally Decentralized Multi-Robot Collision Avoidance via Deep Reinforcement Learning. Proceedings of the 2018 IEEE International Conference on Robotics and Automation, Brisbane, Australia.","DOI":"10.1109\/ICRA.2018.8461113"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Semnani, S.H., Liu, H.H.T., Everett, M., De Ruiter, A.H.J., and How, J.P. (August, January 31). Multi-agent Motion Planning for Dense and Dynamic Environments via Deep Reinforcement Learning. Proceedings of the 2020 IEEE International Conference on Robotics and Automation, Paris, France.","DOI":"10.1109\/LRA.2020.2974695"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Zanol, R., Chiariotti, F., and Zanella, A. (2019, January 15\u201318). Drone mapping through multi-agent reinforcement learning. Proceedings of the 2019 IEEE Wireless Communications and Networking Conference (WCNC), Marrakesh, Morocco.","DOI":"10.1109\/WCNC.2019.8885873"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Osokin, D. (2018). Real-time 2D Multi-Person Pose Estimation on CPU: Lightweight OpenPose. arXiv.","DOI":"10.5220\/0007555407440748"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"229","DOI":"10.1007\/BF00992696","article-title":"Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning","volume":"8","author":"Williams","year":"1992","journal-title":"Mach. Learn."},{"key":"ref_37","unstructured":"Guyon, I., Luxburg, U.V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., and Garnett, R. (2017). Value Prediction Network. NIPS, Curran Associates, Inc."},{"key":"ref_38","unstructured":"Clevert, D.A., Unterthiner, T., and Hochreiter, S. (2015). Fast and accurate deep network learning by exponential linear units (elus). arXiv."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Casiez, G., Roussel, N., and Vogel, D. (2012, January 5\u201310). 1\u20ac Filter: A Simple Speed-based Low-pass Filter for Noisy Input in Interactive Systems. Proceedings of the Conference on Human Factors in Computing Systems\u2014Proceedings, Austin, TX, USA.","DOI":"10.1145\/2207676.2208639"}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/13\/19\/3995\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T07:09:49Z","timestamp":1760166589000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/13\/19\/3995"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,10,6]]},"references-count":39,"journal-issue":{"issue":"19","published-online":{"date-parts":[[2021,10]]}},"alternative-id":["rs13193995"],"URL":"https:\/\/doi.org\/10.3390\/rs13193995","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,10,6]]}}}