{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T20:19:49Z","timestamp":1760300389466,"version":"3.37.3"},"reference-count":108,"publisher":"Springer Science and Business Media LLC","issue":"5","license":[{"start":{"date-parts":[[2020,4,24]],"date-time":"2020-04-24T00:00:00Z","timestamp":1587686400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2020,4,24]],"date-time":"2020-04-24T00:00:00Z","timestamp":1587686400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100000769","name":"University of Oxford","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100000769","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Comput Vis"],"published-print":{"date-parts":[[2020,5]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Deep generative modelling for human body analysis is an emerging problem with many interesting applications. However, the latent space learned by such approaches is typically not interpretable, resulting in less flexibility. In this work, we present deep generative models for human body analysis in which the body pose and the visual appearance are disentangled. Such a disentanglement allows independent manipulation of pose and appearance, and hence enables applications such as pose-transfer without specific training for such a task. Our proposed models, the Conditional-DGPose and the Semi-DGPose, have different characteristics. In the first, body pose labels are taken as conditioners, from a fully-supervised training set. In the second, our structured semi-supervised approach allows for pose estimation to be performed by the model itself and relaxes the need for labelled data. Therefore, the Semi-DGPose aims for the joint <jats:italic>understanding<\/jats:italic> and <jats:italic>generation<\/jats:italic> of people in images. It is not only capable of mapping images to interpretable latent representations but also able to map these representations back to the image space. We compare our models with relevant baselines, the ClothNet-Body and the Pose Guided Person Generation networks, demonstrating their merits on the Human3.6M, ChictopiaPlus and DeepFashion benchmarks.<\/jats:p>","DOI":"10.1007\/s11263-020-01306-1","type":"journal-article","created":{"date-parts":[[2020,4,24]],"date-time":"2020-04-24T06:26:06Z","timestamp":1587709566000},"page":"1537-1563","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["DGPose: Deep Generative Models for Human Body Analysis"],"prefix":"10.1007","volume":"128","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4860-0115","authenticated-orcid":false,"given":"Rodrigo","family":"de Bem","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Arnab","family":"Ghosh","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Thalaiyasingam","family":"Ajanthan","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ondrej","family":"Miksik","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Adnane","family":"Boukhayma","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"N.","family":"Siddharth","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Philip","family":"Torr","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2020,4,24]]},"reference":[{"key":"1306_CR1","unstructured":"3Lateral: 3Lateral. (2018). http:\/\/www.3lateral.com\/."},{"key":"1306_CR2","doi-asserted-by":"crossref","unstructured":"Achilles, F., Ichim, A. E., Coskun, H., Tombari, F., Noachtar, S., & Navab, N. (2016). Patient MoCap: Human pose estimation under blanket occlusion for hospital monitoring applications. In MICCAI.","DOI":"10.1007\/978-3-319-46720-7_57"},{"key":"1306_CR3","doi-asserted-by":"crossref","unstructured":"Andriluka, M., Pishchulin, L., Gehler, P., & Schiele, B. (2014). 2d human pose estimation: New benchmark and state of the art analysis. In CVPR.","DOI":"10.1109\/CVPR.2014.471"},{"key":"1306_CR4","doi-asserted-by":"crossref","unstructured":"Balakrishnan, G., Zhao, A., Dalca, A. V., Durand, F., & Guttag, J. (2018). Synthesizing images of humans in unseen poses. In CVPR","DOI":"10.1109\/CVPR.2018.00870"},{"key":"1306_CR5","unstructured":"de\u00a0Bem, R., Arnab, A., Sapienza, M., Golodetz, S., & Torr, P. (2018) Deep fully-connected part-based models for human pose estimation. In ACML."},{"key":"1306_CR6","unstructured":"de\u00a0Bem, R., Ghosh, A., Ajanthan, T., Miksik, O., Siddharth, N., & Torr, P. (2018). A semi-supervised deep generative model for human body analysis. In ECCV (HBUGEN)."},{"key":"1306_CR7","unstructured":"de\u00a0Bem, R., Ghosh, A., Ajanthan, T., Siddharth, N., & Torr, P. (2019). A conditional deep generative model of people in natural images. In WACV."},{"key":"1306_CR8","doi-asserted-by":"crossref","unstructured":"Blanz, V., Vetter, T., et\u00a0al. (1999). A morphable model for the synthesis of 3d faces. In SIGGRAPH.","DOI":"10.1145\/311535.311556"},{"key":"1306_CR9","unstructured":"Boeing: William Fetter\u2019s Boeing Man. (2018). https:\/\/secure.boeingimages.com\/archive\/William-Fetter-s-Boeing-Man-2F3XC5YCZNC.html."},{"key":"1306_CR10","doi-asserted-by":"crossref","unstructured":"Bogo, F., Romero, J., Loper, M., & Black, M. J. (2014). Faust: Dataset and evaluation for 3d mesh registration. In CVPR (pp. 3794\u20133801)","DOI":"10.1109\/CVPR.2014.491"},{"key":"1306_CR11","doi-asserted-by":"crossref","unstructured":"Borshukov, G., Piponi, D., Larsen, O., Lewis, J. P., & Tempelaar-lietz, C. (2005). Universal capture-image-based facial animation for \u201cThe Matrix Reloaded\u201d. In SIGGRAPH.","DOI":"10.1145\/1198555.1198596"},{"key":"1306_CR12","doi-asserted-by":"crossref","unstructured":"Bulat, A., & Tzimiropoulos, G. (2016). Human pose estimation via convolutional part heatmap regression. In ECCV.","DOI":"10.1007\/978-3-319-46478-7_44"},{"key":"1306_CR13","doi-asserted-by":"crossref","unstructured":"Cao, Z., Simon, T., Wei, S. E., & Sheikh, Y. (2017). Realtime multi-person 2d pose estimation using part affinity fields. In CVPR.","DOI":"10.1109\/CVPR.2017.143"},{"key":"1306_CR14","unstructured":"Chan, C., Ginosar, S., Zhou, T., & Efros, A. A. (2018). Everybody dance now. arXiv preprint arXiv:1808.07371."},{"key":"1306_CR15","doi-asserted-by":"crossref","unstructured":"Chu, X., Yang, W., Ouyang, W., Ma, C., Yuille, A. L., & Wang, X. (2017). Multi-context attention for human pose estimation. In CVPR.","DOI":"10.1109\/CVPR.2017.601"},{"issue":"9","key":"1306_CR16","doi-asserted-by":"publisher","first-page":"1793","DOI":"10.1109\/TPAMI.2011.33","volume":"33","author":"M de La Gorce","year":"2011","unstructured":"de La Gorce, M., Fleet, D. J., & Paragios, N. (2011). Model-based 3D hand pose estimation from monocular video. IEEE Transactions on Pattern Analysis and Machine Intelligence, 33(9), 1793\u20131805.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"1306_CR17","doi-asserted-by":"crossref","unstructured":"Elgammal, A., & Lee, C. S. (2004). Inferring 3d body pose from silhouettes using activity manifold learning. In CVPR.","DOI":"10.1109\/CVPR.2004.1315230"},{"key":"1306_CR18","doi-asserted-by":"crossref","unstructured":"Enzweiler, M., & Gavrila, D. M. (2008). A mixed generative-discriminative framework for pedestrian classification. In CVPR.","DOI":"10.1109\/CVPR.2008.4587592"},{"key":"1306_CR19","doi-asserted-by":"crossref","unstructured":"Esser, P., Sutter, E., & Ommer, B. (2018). A variational u-net for conditional appearance and shape generation. In CVPR.","DOI":"10.1109\/CVPR.2018.00923"},{"key":"1306_CR20","doi-asserted-by":"crossref","unstructured":"Ezzat, T., & Poggio, T. (1996). Facial analysis and synthesis using image-based models. In FG.","DOI":"10.1109\/AFGR.1996.557252"},{"issue":"9","key":"1306_CR21","doi-asserted-by":"publisher","first-page":"2180","DOI":"10.1109\/TPAMI.2017.2747150","volume":"40","author":"S Fan","year":"2018","unstructured":"Fan, S., Ng, T. T., Koenig, B. L., Herberg, J. S., Jiang, M., Shen, Z., et al. (2018). Image visual realism: From human perception to machine computation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(9), 2180\u20132193.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"1306_CR22","doi-asserted-by":"crossref","unstructured":"Fei-Fei, L., & Perona, P. (2005). A Bayesian hierarchical model for learning natural scene categories. In CVPR (Vol.\u00a02).","DOI":"10.1109\/CVPR.2005.16"},{"issue":"9","key":"1306_CR23","doi-asserted-by":"publisher","first-page":"9","DOI":"10.1109\/MCG.1982.1674468","volume":"2","author":"WA Fetter","year":"1982","unstructured":"Fetter, W. A. (1982). A progression of human figures simulated by computer graphics. IEEE Computer Graphics and Applications, 2(9), 9\u201313.","journal-title":"IEEE Computer Graphics and Applications"},{"issue":"2","key":"1306_CR24","doi-asserted-by":"publisher","first-page":"267","DOI":"10.1109\/TPAMI.2007.1174","volume":"30","author":"F Fleuret","year":"2007","unstructured":"Fleuret, F., Berclaz, J., Lengagne, R., & Fua, P. (2007). Multicamera people tracking with a probabilistic occupancy map. IEEE Transactions on Pattern Analysis and Machine Intelligence, 30(2), 267\u2013282.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"1306_CR25","doi-asserted-by":"crossref","unstructured":"Fossati, A., Dimitrijevic, M., Lepetit, V., & Fua, P. (2007). Bridging the gap between detection and tracking for 3d monocular video-based motion capture. In CVPR (pp. 1\u20138).","DOI":"10.1109\/CVPR.2007.383297"},{"key":"1306_CR26","unstructured":"Franco, J. S., & Boyer, E. (2005). Fusion of multi-view silhouette cues using a space occupancy grid. In ICCV (pp. 1747\u20131753)."},{"key":"1306_CR27","doi-asserted-by":"publisher","first-page":"167","DOI":"10.1146\/annurev.psych.58.110405.085632","volume":"59","author":"WS Geisler","year":"2008","unstructured":"Geisler, W. S. (2008). Visual perception and the statistical properties of natural scenes. Annual Review of Psychology, 59, 167\u2013192.","journal-title":"Annual Review of Psychology"},{"key":"1306_CR28","volume-title":"Deep learning","author":"I Goodfellow","year":"2016","unstructured":"Goodfellow, I., Bengio, Y., Courville, A., & Bengio, Y. (2016). Deep learning. Cambridge: MIT Press."},{"key":"1306_CR29","unstructured":"Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., & Bengio, Y. (2014). Generative adversarial nets. In NIPS."},{"issue":"9","key":"1306_CR30","doi-asserted-by":"publisher","first-page":"1027","DOI":"10.1007\/s11263-018-1077-3","volume":"126","author":"H Hattori","year":"2018","unstructured":"Hattori, H., Lee, N., Boddeti, V. N., Beainy, F., Kitani, K. M., & Kanade, T. (2018). Synthesizing a scene-specific pedestrian detector and pose estimator for static video surveillance. International Journal of Computer Vision, 126(9), 1027\u20131044.","journal-title":"International Journal of Computer Vision"},{"key":"1306_CR31","doi-asserted-by":"crossref","unstructured":"Hattori, H., Naresh\u00a0Boddeti, V., Kitani, K. M., & Kanade, T. (2015). Learning scene-specific pedestrian detectors without real data. In CVPR.","DOI":"10.1109\/CVPR.2015.7299006"},{"key":"1306_CR32","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., & Sun, J. (2015). Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In ICCV.","DOI":"10.1109\/ICCV.2015.123"},{"key":"1306_CR33","doi-asserted-by":"crossref","unstructured":"Hilton, A., Beresford, D., Gentils, T., Smith, R., & Sun, W. (1999). Virtual people: Capturing human models to populate virtual worlds. In Proceedings of computer animation (pp. 174\u2013185).","DOI":"10.1109\/CA.1999.781210"},{"key":"1306_CR34","unstructured":"Ian Spriggs. (2018). http:\/\/www.ianspriggs.com\/."},{"issue":"4","key":"1306_CR35","doi-asserted-by":"publisher","first-page":"45","DOI":"10.1145\/2766974","volume":"34","author":"AE Ichim","year":"2015","unstructured":"Ichim, A. E., Bouaziz, S., & Pauly, M. (2015). Dynamic 3d avatar creation from hand-held video input. ACM Transactions on Graphics (ToG), 34(4), 45.","journal-title":"ACM Transactions on Graphics (ToG)"},{"key":"1306_CR36","doi-asserted-by":"crossref","unstructured":"Insafutdinov, E., Pishchulin, L., Andres, B., Andriluka, M., & Schiele, B. (2016). Deepercut: A deeper, stronger, and faster multi-person pose estimation model. In ECCV.","DOI":"10.1007\/978-3-319-46466-4_3"},{"issue":"7","key":"1306_CR37","doi-asserted-by":"publisher","first-page":"1325","DOI":"10.1109\/TPAMI.2013.248","volume":"36","author":"C Ionescu","year":"2014","unstructured":"Ionescu, C., Papava, D., Olaru, V., & Sminchisescu, C. (2014). Human3.6m: Large scale datasets and predictive methods for 3d human sensing in natural environments. IEEE Transactions on Pattern Analysis and Machine Intelligence, 36(7), 1325\u20131339.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"1306_CR38","unstructured":"Isola, P., Zhu, J. Y., Zhou, T., & Efros, A. A. (2016). Image-to-image translation with conditional adversarial networks. arXiv preprint arXiv:1611.07004."},{"key":"1306_CR39","doi-asserted-by":"crossref","unstructured":"Jhuang, H., Gall, J., Zuffi, S., Schmid, C., & Black, M. J. (2013). Towards understanding action recognition. In ICCV.","DOI":"10.1109\/ICCV.2013.396"},{"key":"1306_CR40","unstructured":"Jia, Y., Shelhamer, E., Donahue, J., Karayev, S., Long, J., Girshick, R., Guadarrama, S., & Darrell, T. (2014). Caffe: Convolutional architecture for fast feature embedding. arXiv preprint arXiv:1408.5093."},{"key":"1306_CR41","doi-asserted-by":"crossref","unstructured":"Johnson, S., & Everingham, M. (2010). Clustered pose and nonlinear appearance models for human pose estimation. In BMVC.","DOI":"10.5244\/C.24.12"},{"issue":"1","key":"1306_CR42","doi-asserted-by":"publisher","first-page":"34","DOI":"10.1109\/93.580394","volume":"4","author":"T Kanade","year":"1997","unstructured":"Kanade, T., Rander, P., & Narayanan, P. J. (1997). Virtualized reality: Constructing virtual worlds from real scenes. IEEE Multimedia, 4(1), 34\u201347.","journal-title":"IEEE Multimedia"},{"key":"1306_CR43","unstructured":"Karras, T., Aila, T., Laine, S., & Lehtinen, J. (2018). Progressive growing of gans for improved quality, stability, and variation. In ICLR."},{"issue":"7185","key":"1306_CR44","doi-asserted-by":"publisher","first-page":"352","DOI":"10.1038\/nature06713","volume":"452","author":"KN Kay","year":"2008","unstructured":"Kay, K. N., Naselaris, T., Prenger, R. J., & Gallant, J. L. (2008). Identifying natural images from human brain activity. Nature, 452(7185), 352.","journal-title":"Nature"},{"key":"1306_CR45","unstructured":"Kingma, D., & Ba, J. (2015). Adam: A method for stochastic optimization. In ICLR."},{"key":"1306_CR46","unstructured":"Kingma, D. P., Mohamed, S., Rezende, D. J., & Welling, M. (2014). Semi-supervised learning with deep generative models. In NIPS."},{"key":"1306_CR47","unstructured":"Kingma, D. P., & Welling, M. (2013). Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114."},{"key":"1306_CR48","unstructured":"Krizhevsky, A., Hinton, G., et\u00a0al. (2010). Factored 3-way restricted Boltzmann machines for modeling natural images. In Proceedings of the thirteenth international conference on artificial intelligence and statistics."},{"key":"1306_CR49","unstructured":"Kulkarni, T. D., Whitney, W. F., Kohli, P., & Tenenbaum, J. (2015). Deep convolutional inverse graphics network. In NIPS."},{"key":"1306_CR50","unstructured":"Larsen, A. B. L., S\u00f8nderby, S. K., Larochelle, H., & Winther, O. (2016). Autoencoding beyond pixels using a learned similarity metric. In ICML."},{"key":"1306_CR51","doi-asserted-by":"crossref","unstructured":"Lassner, C., Pons-Moll, G., & Gehler, P. V. (2017). A generative model for people in clothing. In ICCV.","DOI":"10.1109\/ICCV.2017.98"},{"key":"1306_CR52","doi-asserted-by":"crossref","unstructured":"Lee, C. S., & Elgammal, A. (2005). Facial expression analysis using nonlinear decomposable generative models. In International workshop on analysis and modeling of faces and gestures.","DOI":"10.1007\/11564386_3"},{"key":"1306_CR53","doi-asserted-by":"crossref","unstructured":"Liang, X., Liu, S., Shen, X., Yang, J., Liu, L., Dong, J., Lin, L., & Yan, S. (2015). Deep human parsing with active template regression. In TPAMI.","DOI":"10.1109\/TPAMI.2015.2408360"},{"key":"1306_CR54","doi-asserted-by":"crossref","unstructured":"Liu, Z., Luo, P., Qiu, S., Wang, X., & Tang, X. (2016). Deepfashion: Powering robust clothes recognition and retrieval with rich annotations. In CVPR.","DOI":"10.1109\/CVPR.2016.124"},{"issue":"6","key":"1306_CR55","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/2816795.2818013","volume":"34","author":"M Loper","year":"2015","unstructured":"Loper, M., Mahmood, N., Romero, J., Pons-Moll, G., & Black, M. J. (2015). SMPL: A skinned multi-person linear model. ACM Transactions on Graphics (TOG), 34(6), 1\u201316.","journal-title":"ACM Transactions on Graphics (TOG)"},{"key":"1306_CR56","unstructured":"Ma, L., Jia, X., Sun, Q., Schiele, B., Tuytelaars, T., & Gool, L. V. (2017). Pose guided person image generation. In NIPS."},{"key":"1306_CR57","doi-asserted-by":"crossref","unstructured":"Ma, L., Sun, Q., Georgoulis, S., Van\u00a0Gool, L., Schiele, B., & Fritz, M. (2018). Disentangled person image generation. In CVPR.","DOI":"10.1109\/CVPR.2018.00018"},{"issue":"3","key":"1306_CR58","doi-asserted-by":"publisher","first-page":"695","DOI":"10.1016\/j.chb.2008.12.026","volume":"25","author":"KF MacDorman","year":"2009","unstructured":"MacDorman, K. F., Green, R. D., Ho, C. C., & Koch, C. T. (2009). Too real for comfort? Uncanny responses to computer generated faces. Computers in Human Behavior, 25(3), 695\u2013710.","journal-title":"Computers in Human Behavior"},{"issue":"12","key":"1306_CR59","doi-asserted-by":"publisher","first-page":"997","DOI":"10.1007\/s00371-005-0363-6","volume":"21","author":"N Magnenat-Thalmann","year":"2005","unstructured":"Magnenat-Thalmann, N., & Thalmann, D. (2005). Virtual humans: Thirty years of research, what next? Visual Computer, 21(12), 997\u20131015.","journal-title":"Visual Computer"},{"key":"1306_CR60","volume-title":"Handbook of virtual humans","author":"N Magnenat-Thalmann","year":"2006","unstructured":"Magnenat-Thalmann, N., & Thalmann, D. (2006). Handbook of virtual humans. Hoboken: Wiley."},{"key":"1306_CR61","unstructured":"MakeHuman. (2018). http:\/\/www.makehumancommunity.org\/."},{"key":"1306_CR62","unstructured":"Massiceti, D., Siddharth, N., Dokania, P., & Torr, P. H. (2018). FlipDial: A generative model for two-way visual dialogue. In CVPR."},{"key":"1306_CR63","unstructured":"Massive Software. (2018). http:\/\/www.massivesoftware.com\/."},{"key":"1306_CR64","doi-asserted-by":"publisher","DOI":"10.1007\/978-0-85729-997-0","volume-title":"Visual analysis of humans","author":"TB Moeslund","year":"2011","unstructured":"Moeslund, T. B., Hilton, A., Kr\u00fcger, V., & Sigal, L. (2011). Visual analysis of humans. Berlin: Springer."},{"issue":"9","key":"1306_CR65","doi-asserted-by":"publisher","first-page":"902","DOI":"10.1007\/s11263-018-1073-7","volume":"126","author":"M M\u00fcller","year":"2018","unstructured":"M\u00fcller, M., Casser, V., Lahoud, J., Smith, N., & Ghanem, B. (2018). Sim4cv: A photo-realistic simulator for computer vision applications. International Journal of Computer Vision, 126(9), 902\u2013919.","journal-title":"International Journal of Computer Vision"},{"key":"1306_CR66","unstructured":"NASA . (1995). Space flight human-system standard volume 1. Technical report NASA-STD-3001, National Aeronautics and Space Administration\u2014NASA"},{"key":"1306_CR67","doi-asserted-by":"crossref","unstructured":"Neverova, N., Alp\u00a0Guler, R., & Kokkinos, I. (2018). Dense pose transfer. In ECCV.","DOI":"10.1007\/978-3-030-01219-9_8"},{"key":"1306_CR68","doi-asserted-by":"crossref","unstructured":"Newell, A., Yang, K., & Deng, J. (2016). Stacked hourglass networks for human pose estimation. In ECCV.","DOI":"10.1007\/978-3-319-46484-8_29"},{"key":"1306_CR69","doi-asserted-by":"crossref","unstructured":"Pighin, F., Hecker, J., Lischinski, D., Szeliski, R., & Salesin, D. H. (2006). Synthesizing realistic facial expressions from photographs. In SIGGRAPH (p.\u00a019). ACM.","DOI":"10.1145\/1185657.1185859"},{"key":"1306_CR70","unstructured":"Poser. (2018). https:\/\/www.posersoftware.com\/."},{"key":"1306_CR71","doi-asserted-by":"crossref","unstructured":"Pumarola, A., Agudo, A., Sanfeliu, A., & Moreno-Noguer, F. (2018). Unsupervised person image synthesis in arbitrary poses. In CVPR.","DOI":"10.1109\/CVPR.2018.00899"},{"key":"1306_CR72","unstructured":"Rezende, D. J., Mohamed, S., & Wierstra, D. (2014). Stochastic backpropagation and approximate inference in deep generative models. In ICML."},{"key":"1306_CR73","doi-asserted-by":"crossref","unstructured":"Rhodin, H., Salzmann, M., & Fua, P. (2018). Unsupervised geometry-aware representation for 3d human pose estimation. In ECCV.","DOI":"10.1007\/978-3-030-01249-6_46"},{"key":"1306_CR74","unstructured":"Rogez, G., & Schmid, C. (2016). Mocap-guided data augmentation for 3d pose estimation in the wild. In NIPS."},{"issue":"9","key":"1306_CR75","doi-asserted-by":"publisher","first-page":"993","DOI":"10.1007\/s11263-018-1071-9","volume":"126","author":"G Rogez","year":"2018","unstructured":"Rogez, G., & Schmid, C. (2018). Image-based synthesis for deep 3d human pose estimation. International Journal of Computer Vision, 126(9), 993\u20131008.","journal-title":"International Journal of Computer Vision"},{"issue":"6","key":"1306_CR76","doi-asserted-by":"publisher","first-page":"245","DOI":"10.1145\/3130800.3130883","volume":"36","author":"J Romero","year":"2017","unstructured":"Romero, J., Tzionas, D., & Black, M. J. (2017). Embodied hands: Modeling and capturing hands and bodies together. ACM Transactions on Graphics (ToG), 36(6), 245.","journal-title":"ACM Transactions on Graphics (ToG)"},{"key":"1306_CR77","doi-asserted-by":"crossref","unstructured":"Ronneberger, O., Fischer, P., & Brox, T. (2015). U-net: Convolutional networks for biomedical image segmentation. In MICCAI.","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"1306_CR78","doi-asserted-by":"crossref","unstructured":"Rosales, R., Athitsos, V., Sigal, L., & Sclaroff, S. (2001). 3d hand pose reconstruction using specialized mappings. In ICCV.","DOI":"10.21236\/ADA451286"},{"key":"1306_CR79","unstructured":"Schulman, J., Heess, N., Weber, T., & Abbeel, P. (2015). Gradient estimation using stochastic computation graphs. In NIPS."},{"key":"1306_CR80","doi-asserted-by":"crossref","unstructured":"Seemann, E., Nickel, K., & Stiefelhagen, R. (2004). Head pose estimation using stereo vision for human\u2013robot interaction. In FG.","DOI":"10.1109\/AFGR.2004.1301603"},{"key":"1306_CR81","doi-asserted-by":"crossref","unstructured":"Shan, Q., Adams, R., Curless, B., Furukawa, Y., & Seitz, S. M. (2013). The visual turing test for scene reconstruction. In 3DV.","DOI":"10.1109\/3DV.2013.12"},{"key":"1306_CR82","doi-asserted-by":"crossref","unstructured":"Shotton, J., Fitzgibbon, A. W., Cook, M., Sharp, T., Finocchio, M., Moore, R., Kipman, A., & Blake, A. (2011). Real-time human pose recognition in parts from single depth images. In CVPR.","DOI":"10.1109\/CVPR.2011.5995316"},{"key":"1306_CR83","doi-asserted-by":"crossref","unstructured":"Siarohin, A., Sangineto, E., Lathuili\u00e8re, S., & Sebe, N. (2018). Deformable gans for pose-based human image generation. In CVPR.","DOI":"10.1109\/CVPR.2018.00359"},{"key":"1306_CR84","unstructured":"Siddharth, N., Paige, B., Desmaison, A., van de Meent, J. W., Wood, F., Goodman, N. D., et al. (2017). Learning disentangled representations with semi-supervised deep generative models. In,. NIPS."},{"issue":"1","key":"1306_CR85","doi-asserted-by":"publisher","first-page":"1193","DOI":"10.1146\/annurev.neuro.24.1.1193","volume":"24","author":"EP Simoncelli","year":"2001","unstructured":"Simoncelli, E. P., & Olshausen, B. A. (2001). Natural image statistics and neural representation. Annual Review of Neuroscience, 24(1), 1193\u20131216.","journal-title":"Annual Review of Neuroscience"},{"key":"1306_CR86","unstructured":"Sohn, K., Lee, H., & Yan, X.(2015). Learning structured output representation using deep conditional generative models. In NIPS."},{"issue":"3","key":"1306_CR87","doi-asserted-by":"publisher","first-page":"21","DOI":"10.1109\/MCG.2007.68","volume":"27","author":"J Starck","year":"2007","unstructured":"Starck, J., & Hilton, A. (2007). Surface capture for performance-based animation. IEEE Computer Graphics and Applications, 27(3), 21\u201331.","journal-title":"IEEE Computer Graphics and Applications"},{"key":"1306_CR88","doi-asserted-by":"crossref","unstructured":"Starck, J., Miller, G., & Hilton, A. (2005). Video-based character animation. In SIGGRAPH (pp. 49\u201358). ACM","DOI":"10.1145\/1073368.1073375"},{"key":"1306_CR89","unstructured":"Theis, L., van\u00a0den Oord, A., & Bethge, M. (2016). A note on the evaluation of generative models. In ICLR."},{"key":"1306_CR90","doi-asserted-by":"crossref","unstructured":"Thies, J., Zollh\u00f6fer, M., Stamminger, M., Theobalt, C., & Nie\u00dfner, M. (2016). Face2face: Real-time face capture and reenactment of rgb videos. In CVPR.","DOI":"10.1145\/2929464.2929475"},{"key":"1306_CR91","unstructured":"Tompson, J., Jain, A., LeCun, Y., & Bregler, C. (2014). Joint training of a convolutional network and a graphical model for human pose estimation. In NIPS."},{"key":"1306_CR92","doi-asserted-by":"crossref","unstructured":"Trumble, M., Gilbert, A., Hilton, A., & Collomosse, J. (2018). Deep autoencoder for combined human pose estimation and body model upscaling. In ECCV18.","DOI":"10.1007\/978-3-030-01249-6_48"},{"key":"1306_CR93","doi-asserted-by":"crossref","unstructured":"Tulyakov, S., Liu, M. Y., Yang, X., & Kautz, J. (2018). Mocogan: Decomposing motion and content for video generation. In CVPR.","DOI":"10.1109\/CVPR.2018.00165"},{"key":"1306_CR94","unstructured":"Unreal Engine. (2018). https:\/\/www.unrealengine.com."},{"issue":"4","key":"1306_CR95","first-page":"892","volume":"22","author":"E Valenza","year":"1996","unstructured":"Valenza, E., Simion, F., Cassia, V. M., & Umilt\u00e0, C. (1996). Face preference at birth. Journal of Experimental Psychology: Human perception and performance, 22(4), 892\u2013903.","journal-title":"Journal of Experimental Psychology: Human perception and performance"},{"key":"1306_CR96","doi-asserted-by":"crossref","unstructured":"Varol, G., Romero, J., Martin, X., Mahmood, N., Black, M. J., Laptev, I., & Schmid, C. (2017). Learning from synthetic humans. In CVPR.","DOI":"10.1109\/CVPR.2017.492"},{"key":"1306_CR97","doi-asserted-by":"crossref","unstructured":"von Marcard, T., Rosenhahn, B., Black, M., & Pons-Moll, G. (2017). Sparse inertial poser: Automatic 3d human pose estimation from sparse imus. Eurographics.","DOI":"10.1111\/cgf.13131"},{"key":"1306_CR98","doi-asserted-by":"crossref","unstructured":"Walker, J., Marino, K., Gupta, A., & Hebert, M. (2017). The pose knows: Video forecasting by generating pose futures. In ICCV.","DOI":"10.1109\/ICCV.2017.361"},{"key":"1306_CR99","unstructured":"Wang, T. C., Liu, M. Y., Zhu, J. Y., Yakovenko, N., Tao, A., Kautz, J., Catanzaro, B. (2018). Video-to-video synthesis. In NIPS."},{"issue":"3","key":"1306_CR100","doi-asserted-by":"publisher","first-page":"677","DOI":"10.1111\/j.1467-8659.2004.00800.x","volume":"23","author":"Y Wang","year":"2004","unstructured":"Wang, Y., Huang, X., Lee, C. S., Zhang, S., Li, Z., Samaras, D., et al. (2004). High resolution acquisition, learning and transfer of dynamic 3-d facial expressions. Computer Graphics Forum, 23(3), 677\u2013686.","journal-title":"Computer Graphics Forum"},{"issue":"4","key":"1306_CR101","doi-asserted-by":"publisher","first-page":"600","DOI":"10.1109\/TIP.2003.819861","volume":"13","author":"Z Wang","year":"2004","unstructured":"Wang, Z., Bovik, A. C., Sheikh, H. R., & Simoncelli, E. P. (2004). Image quality assessment: From error visibility to structural similarity. IEEE Transactions on Image Processing, 13(4), 600\u2013612.","journal-title":"IEEE Transactions on Image Processing"},{"key":"1306_CR102","unstructured":"Wang, Z., Merel, J. S., Reed, S. E., de\u00a0Freitas, N., Wayne, G., & Heess, N. (2017). Robust imitation of diverse behaviors. In NIPS."},{"key":"1306_CR103","doi-asserted-by":"crossref","unstructured":"Wei, S. E., Ramakrishna, V., Kanade, T., & Sheikh, Y. (2016). Convolutional pose machines. In CVPR.","DOI":"10.1109\/CVPR.2016.511"},{"key":"1306_CR104","doi-asserted-by":"crossref","unstructured":"Yang, W., Li, S., Ouyang, W., Li, H., & Wang, X. (2017). Learning feature pyramids for human pose estimation. In ICCV.","DOI":"10.1109\/ICCV.2017.144"},{"key":"1306_CR105","doi-asserted-by":"crossref","unstructured":"Yang, Y., & Ramanan, D. (2011). Articulated pose estimation with flexible mixtures-of-parts. In CVPR.","DOI":"10.1109\/CVPR.2011.5995741"},{"issue":"7","key":"1306_CR106","doi-asserted-by":"publisher","first-page":"301","DOI":"10.1016\/j.tics.2006.05.002","volume":"10","author":"A Yuille","year":"2006","unstructured":"Yuille, A., & Kersten, D. (2006). Vision as bayesian inference: Analysis by synthesis? Trends in Cognitive Sciences, 10(7), 301\u2013308.","journal-title":"Trends in Cognitive Sciences"},{"key":"1306_CR107","doi-asserted-by":"crossref","unstructured":"Zanfir, M., Popa, A. I., Zanfir, A., & Sminchisescu, C. (2018). Human appearance transfer. In CVPR.","DOI":"10.1109\/CVPR.2018.00565"},{"key":"1306_CR108","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Guo, Y., Jin, Y., Luo, Y., He, Z., & Lee, H. (2018). Unsupervised discovery of object landmarks as structural representations. In CVPR.","DOI":"10.1109\/CVPR.2018.00285"}],"container-title":["International Journal of Computer Vision"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-020-01306-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11263-020-01306-1\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-020-01306-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,4,25]],"date-time":"2021-04-25T04:24:10Z","timestamp":1619324650000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11263-020-01306-1"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,4,24]]},"references-count":108,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2020,5]]}},"alternative-id":["1306"],"URL":"https:\/\/doi.org\/10.1007\/s11263-020-01306-1","relation":{},"ISSN":["0920-5691","1573-1405"],"issn-type":[{"type":"print","value":"0920-5691"},{"type":"electronic","value":"1573-1405"}],"subject":[],"published":{"date-parts":[[2020,4,24]]},"assertion":[{"value":"13 November 2018","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"7 February 2020","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"24 April 2020","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}