{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,21]],"date-time":"2026-02-21T21:17:44Z","timestamp":1771708664014,"version":"3.50.1"},"reference-count":30,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2021,2,14]],"date-time":"2021-02-14T00:00:00Z","timestamp":1613260800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100003725","name":"National Research Foundation of Korea","doi-asserted-by":"publisher","award":["2018R1A2B2007934"],"award-info":[{"award-number":["2018R1A2B2007934"]}],"id":[{"id":"10.13039\/501100003725","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100002471","name":"Dongguk University","doi-asserted-by":"publisher","award":["Dongguk University Research Fund of 2020"],"award-info":[{"award-number":["Dongguk University Research Fund of 2020"]}],"id":[{"id":"10.13039\/501100002471","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Applications related to smart cities require virtual cities in the experimental development stage. To build a virtual city that are close to a real city, a large number of various types of human models need to be created. To reduce the cost of acquiring models, this paper proposes a method to reconstruct 3D human meshes from single images captured using a normal camera. It presents a method for reconstructing the complete mesh of the human body from a single RGB image and a generative adversarial network consisting of a newly designed shape\u2013pose-based generator (based on deep convolutional neural networks) and an enhanced multi-source discriminator. Using a machine learning approach, the reliance on multiple sensors is reduced and 3D human meshes can be recovered using a single camera, thereby reducing the cost of building smart cities. The proposed method achieves an accuracy of 92.1% in body shape recovery; it can also process 34 images per second. The method proposed in this paper approach significantly improves the performance compared with previous state-of-the-art approaches. Given a single view image of various humans, our results can be used to generate various 3D human models, which can facilitate 3D human modeling work to simulate virtual cities. Since our method can also restore the poses of the humans in the image, it is possible to create various human poses by given corresponding images with specific human poses.<\/jats:p>","DOI":"10.3390\/s21041350","type":"journal-article","created":{"date-parts":[[2021,2,14]],"date-time":"2021-02-14T10:01:05Z","timestamp":1613296865000},"page":"1350","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["Human Mesh Reconstruction with Generative Adversarial Networks from Single RGB Images"],"prefix":"10.3390","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6456-9496","authenticated-orcid":false,"given":"Rui","family":"Gao","sequence":"first","affiliation":[{"name":"Department of Multimedia Engineering, Dongguk University-Seoul, 30, Pildongro-1-gil, Jung-gu, Seoul 04620, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mingyun","family":"Wen","sequence":"additional","affiliation":[{"name":"Department of Multimedia Engineering, Dongguk University-Seoul, 30, Pildongro-1-gil, Jung-gu, Seoul 04620, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4304-1780","authenticated-orcid":false,"given":"Jisun","family":"Park","sequence":"additional","affiliation":[{"name":"Department of Multimedia Engineering, Dongguk University-Seoul, 30, Pildongro-1-gil, Jung-gu, Seoul 04620, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2219-0848","authenticated-orcid":false,"given":"Kyungeun","family":"Cho","sequence":"additional","affiliation":[{"name":"Department of Multimedia Engineering, Dongguk University-Seoul, 30, Pildongro-1-gil, Jung-gu, Seoul 04620, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,2,14]]},"reference":[{"key":"ref_1","first-page":"408","article-title":"SCAPE: Shape Completion and Animation of People","volume":"24","author":"Anguelov","year":"2005","journal-title":"ACM J."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Zhang, Q., Fu, B., Ye, M., and Yang, R. (2014, January 24\u201327). Quality dynamic human body modeling using a single low-cost depth camera. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.92"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"2720","DOI":"10.1109\/TPAMI.2013.47","article-title":"Markerless motion capture of multiple characters using multiview image segmentation","volume":"35","author":"Liu","year":"2013","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Yu, R., Russell, C., Campbell, N.D.F., and Agapito, L. (2015, January 7\u201313). Direct, dense, and deformable. Template-based non-rigid 3d reconstruction from rgb video. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.111"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"206","DOI":"10.1109\/MNET.2019.1800310","article-title":"Human behavior deep recognition architecture for smart city applications in the 5G environment","volume":"33","author":"Dai","year":"2019","journal-title":"IEEE Netw."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"1501","DOI":"10.1109\/LRA.2019.2895266","article-title":"Bio-lstm: A biomechanically inspired recurrent neural network for 3-d pedestrian pose and gait prediction","volume":"4","author":"Xiaoxiao","year":"2019","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Kanazawa, A., Black, M.J., Jacobs, D.W., and Malik, J. (2018, January 18\u201323). End-to-end recovery of human shape and pose. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00744"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Varol, G., Ceylan, D., Russell, B., Yang, J., Yumer, E., Laptev, I., and Schmid, C. (2018, January 8\u201314). Bodynet: Volumetric inference of 3d human body shapes. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_2"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Yang, W., Ouyang, W., Wang, X., Ren, J., Li, H., and Wang, X. (2018, January 18\u201323). 3d human pose estimation in the wild by adversarial learning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00551"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Jackson, A.S., Manafas, C., and Tzimiropoulos, G. (2018, January 8\u201314). 3d human body reconstruction from a single image via volumetric regression. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-11018-5_6"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Alp G\u00fcler, R., Neverova, N., and Kokkinos, I. (2018, January 18\u201323). Densepose: Dense human pose estimation in the wild. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00762"},{"key":"ref_12","unstructured":"Mir, R.I.H., and Little, J.J. (2018, January 23\u201328). Exploiting temporal information for 3d human pose estimation. Proceedings of the European Conference on Computer Vision (ECCV), Glasgow, UK."},{"key":"ref_13","unstructured":"Federica, B., Kanazawa, A., Lassner, C., Gehler, P., Romero, J., and Black, M.J. (2016). Keep it SMPL: Automatic estimation of 3D human pose and shape from a single image. European Conference on Computer Vision, Springer."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"337","DOI":"10.1111\/j.1467-8659.2009.01373.x","article-title":"A statistical model of human pose and body shape","volume":"Volume 28","author":"Hasler","year":"2009","journal-title":"Computer Graphics Forum"},{"key":"ref_15","unstructured":"Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014). Generative adversarial nets. Advances in Neural Information Processing Systems, ACM."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Pumarola, A., Agudo, A., Sanfeliu, A., and Moreno-Noguer, F. (2018, January 18\u201323). Unsupervised person image synthesis in arbitrary poses. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00899"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Pumarola, A., Agudo, A., Martinez, A.M., Sanfeliu, A., and Moreno-Noguert, F. (2018, January 8\u201314). Ganimation: Anatomically-aware facial animation from a single image. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01249-6_50"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3272127.3275075","article-title":"paGAN: Real-time avatars using dynamic textures","volume":"37","author":"Nagano","year":"2018","journal-title":"ACM Trans. Graph. (TOG)"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3072959.3073596","article-title":"Vnect: Real-time 3d human pose estimation with a single rgb camera","volume":"36","author":"Mehta","year":"2017","journal-title":"ACM Trans. Graph. (TOG)"},{"key":"ref_20","unstructured":"Alejandro, N., Yang, K., and Deng, J. (2016). Stacked hourglass networks for human pose estimation. European Conference on Computer Vision, Springer."},{"key":"ref_21","unstructured":"Alec, R., Metz, L., and Chintala, S. (2015). Unsupervised representation learning with deep convolutional generative adversarial networks. arXiv."},{"key":"ref_22","first-page":"1","article-title":"Clustered Pose and Nonlinear Appearance Models for Human Pose Estimation","volume":"2","author":"Sam","year":"2010","journal-title":"BMVC"},{"key":"ref_23","unstructured":"Sam, J., and Everingham, M. (2011, January 20\u201325). Learning effective human pose estimation from inaccurate annotation. Proceedings of the CVPR 2011, Providence, RI, USA."},{"key":"ref_24","unstructured":"Tsung-Yi, L., Maire, M., Belonge, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014). Microsoft coco: Common objects in context. European Conference on Computer Vision, Springer."},{"key":"ref_25","unstructured":"Dushyant, M., Rhodin, H., Casas, D., Fua, P., Sotnychenko, O., Xu, W., and Theobalt, C. (2017, January 10\u201312). Monocular 3d human pose estimation in the wild using improved cnn supervision. Proceedings of the 2017 International Conference on 3D Vision (3DV), Qingdao, China."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/2661229.2661273","article-title":"MoSh: Motion and shape capture from sparse markers","volume":"33","author":"Matthew","year":"2014","journal-title":"ACM Trans. Graph. (TOG)"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Varol, G., Romero, J., Martin, X., Mahmood, N., Black, M.J., Laptev, I., and Schmid, C. (2017, January 21\u201326). Learning from synthetic humans. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.492"},{"key":"ref_28","first-page":"1325","article-title":"Human3.6m: Large Scale Datasets and Predictive Methods for 3d Human Sensing in Natural Environments","volume":"36","author":"Catalin","year":"2013","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"4","DOI":"10.1007\/s11263-009-0273-6","article-title":"Humaneva: Synchronized video and motion capture dataset and baseline algorithm for evaluation of articulated human motion","volume":"87","author":"Leonid","year":"2010","journal-title":"Int. J. Comput. Vis."},{"key":"ref_30","unstructured":"Christoph, L., Romero, J., Kiefel, M., Bogo, F., Black, M.J., and Gehler, P.V. (2017, January 21\u201326). Unite the people: Closing the loop between 3d and 2d human representations. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/4\/1350\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T05:24:07Z","timestamp":1760160247000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/4\/1350"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,2,14]]},"references-count":30,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2021,2]]}},"alternative-id":["s21041350"],"URL":"https:\/\/doi.org\/10.3390\/s21041350","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,2,14]]}}}