{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,30]],"date-time":"2026-03-30T14:42:08Z","timestamp":1774881728620,"version":"3.50.1"},"reference-count":35,"publisher":"MDPI AG","issue":"22","license":[{"start":{"date-parts":[[2019,11,13]],"date-time":"2019-11-13T00:00:00Z","timestamp":1573603200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>Generating 3D point clouds from a single image has attracted full attention from researchers in the field of multimedia, remote sensing and computer vision. With the recent proliferation of deep learning, various deep models have been proposed for the 3D point cloud generation. However, they require objects to be captured with absolutely clean backgrounds and fixed viewpoints, which highly limits their application in the real environment. To guide 3D point cloud generation, we propose a novel network, RealPoint3D, to integrate prior 3D shape knowledge into the network. Taking additional 3D information, RealPoint3D can handle 3D object generation from a single real image captured from any viewpoint and complex background. Specifically, provided a query image, we retrieve the nearest shape model from a pre-prepared 3D model database. Then, the image, together with the retrieved shape model, is fed into RealPoint3D to generate a fine-grained 3D point cloud. We evaluated the proposed RealPoint3D on the ShapeNet dataset and ObjectNet3D dataset for the 3D point cloud generation. Experimental results and comparisons with state-of-the-art methods demonstrate that our framework achieves superior performance. Furthermore, our proposed framework works well for real images in complex backgrounds (the image has the remaining objects in addition to the reconstructed object, and the reconstructed object may be occluded or truncated) with various viewing angles.<\/jats:p>","DOI":"10.3390\/rs11222644","type":"journal-article","created":{"date-parts":[[2019,11,13]],"date-time":"2019-11-13T09:11:27Z","timestamp":1573636287000},"page":"2644","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":13,"title":["RealPoint3D: Generating 3D Point Clouds from a Single Image of Complex Scenarios"],"prefix":"10.3390","volume":"11","author":[{"given":"Yan","family":"Xia","sequence":"first","affiliation":[{"name":"Photogrammetry and Remote Sensing, Technical University of Munich, 80333 Munich, Germany"},{"name":"Fujian Key Laboratory of Sensing and Computing, School of Informatics, Xiamen University, Xiamen 361005, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6075-796X","authenticated-orcid":false,"given":"Cheng","family":"Wang","sequence":"additional","affiliation":[{"name":"Fujian Key Laboratory of Sensing and Computing, School of Informatics, Xiamen University, Xiamen 361005, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5571-7808","authenticated-orcid":false,"given":"Yusheng","family":"Xu","sequence":"additional","affiliation":[{"name":"Photogrammetry and Remote Sensing, Technical University of Munich, 80333 Munich, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yu","family":"Zang","sequence":"additional","affiliation":[{"name":"Fujian Key Laboratory of Sensing and Computing, School of Informatics, Xiamen University, Xiamen 361005, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Weiquan","family":"Liu","sequence":"additional","affiliation":[{"name":"Fujian Key Laboratory of Sensing and Computing, School of Informatics, Xiamen University, Xiamen 361005, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7899-0049","authenticated-orcid":false,"given":"Jonathan","family":"Li","sequence":"additional","affiliation":[{"name":"Department of Geography and Environmental Management, University of Waterloo, Waterloo, ON N2L 3G1, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1184-0924","authenticated-orcid":false,"given":"Uwe","family":"Stilla","sequence":"additional","affiliation":[{"name":"Photogrammetry and Remote Sensing, Technical University of Munich, 80333 Munich, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2019,11,13]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"42","DOI":"10.1111\/j.1748-7692.2002.tb01017.x","article-title":"How do sperm whales catch squids?","volume":"18","author":"Fristrup","year":"2002","journal-title":"Mar. Mammal Sci."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Fan, H., Su, H., and Guibas, L.J. (2017, January 21\u201326). A Point Set Generation Network for 3D Object Reconstruction from a Single Image. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.264"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Tatarchenko, M., Dosovitskiy, A., and Brox, T. (2017). Octree Generating Networks: Efficient Convolutional Architectures for High-resolution 3D Outputs. arXiv.","DOI":"10.1109\/ICCV.2017.230"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Wang, N., Zhang, Y., Li, Z., Fu, Y., Liu, W., and Jiang, Y.G. (2018, January 8\u201314). Pixel2Mesh: Generating 3D Mesh Models from Single RGB Images. Proceedings of the European Conference on Computer Vision (ECCV) 2018, Munich, Germany.","DOI":"10.1007\/978-3-030-01252-6_4"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"37","DOI":"10.1145\/2601097.2601159","article-title":"Estimating image depth using shape collections","volume":"33","author":"Su","year":"2014","journal-title":"ACM Trans. Graph. (TOG)"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"87","DOI":"10.1145\/2766890","article-title":"Single-view reconstruction via joint analysis of image and shape collections","volume":"34","author":"Huang","year":"2015","journal-title":"ACM Trans. Graph. (TOG)"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"1923","DOI":"10.1109\/TIP.2018.2878958","article-title":"An augmented linear mixing model to address spectral variability for hyperspectral unmixing","volume":"28","author":"Hong","year":"2019","journal-title":"IEEE Trans. Image. Process."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Cicek, O., Abdulkadir, A., Lienkamp, S.S., Brox, T., and Ronneberger, O. (2016). 3D U-Net: Learning Dense Volumetric Segmentation from Sparse Annotation. International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer.","DOI":"10.1007\/978-3-319-46723-8_49"},{"key":"ref_9","unstructured":"Wu, J., Zhang, C., Xue, T., Freeman, W.T., and Tenenbaum, J.B. (2016). Learning a Probabilistic Latent Space of Object Shapes via 3D Generative-Adversarial Modeling. Advances in Neural Information Processing Systems, MIT Press."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Choy, C.B., Xu, D., Gwak, J.Y., Chen, K., and Savarese, S. (2016). 3D-R2N2: A Unified Approach for Single and Multi-view 3D Object Reconstruction. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-46484-8_38"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Maturana, D., and Scherer, S. (October, January 28). Voxnet: A 3d convolutional neural network for real-time object recognition. Proceedings of the 2015 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Hamburg, Germany.","DOI":"10.1109\/IROS.2015.7353481"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"999","DOI":"10.1016\/j.neucom.2015.10.031","article-title":"Robust palmprint recognition based on the fast variation Vese\u2013Osher model","volume":"174","author":"Hong","year":"2016","journal-title":"Neurocomputing"},{"key":"ref_13","unstructured":"Chang, A.X., Funkhouser, T., Guibas, L., Hanrahan, P., Huang, Q., Li, Z., Savarese, S., Savva, M., Song, S., and Su, H. (2015). ShapeNet: An Information-Rich 3D Model Repository. arXiv."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Xiang, Y., Kim, W., Chen, W., Ji, J., Choy, C., Su, H., Mottaghi, R., Guibas, L., and Savarese, S. (2016). Objectnet3d: A large scale database for 3d object recognition. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-46484-8_10"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"55","DOI":"10.1007\/s10462-012-9365-8","article-title":"Visual simultaneous localization and mapping: A survey","volume":"43","year":"2015","journal-title":"Artif. Intell. Rev."},{"key":"ref_16","first-page":"926","article-title":"The structure-from-motion reconstruction pipeline\u2014A survey with focus on short image sequences","volume":"46","author":"Peters","year":"2010","journal-title":"Kybernetika"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"4349","DOI":"10.1109\/TGRS.2018.2890705","article-title":"Cospace: Common subspace learning from hyperspectral-multispectral correspondences","volume":"57","author":"Hong","year":"2019","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"1670","DOI":"10.1109\/TPAMI.2014.2377712","article-title":"Shape, Illumination, and Reflectance from Shading","volume":"37","author":"Barron","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"149","DOI":"10.1023\/A:1007958829620","article-title":"Computing Local Surface Orientation and Shape from Texture for Curved Surfaces","volume":"23","author":"Malik","year":"1997","journal-title":"Int. J. Comput. Vis."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"305","DOI":"10.1007\/s11263-006-8323-9","article-title":"3D Reconstruction by Shadow Carving: Theory and Practical Evaluation","volume":"71","author":"Savarese","year":"2007","journal-title":"Int. J. Comput. Vis."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Karimi Mahabadi, R., Hane, C., and Pollefeys, M. (2015, January 7\u201312). Segment based 3D object shape priors. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition 2015, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298901"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Qi, C.R., Su, H., Niebner, M., Dai, A., Yan, M., and Guibas, L.J. (July, January 26). Volumetric and Multi-view CNNs for Object Classification on 3D Data. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition 2016, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.609"},{"key":"ref_23","unstructured":"Qi, C.R., Su, H., Mo, K., and Guibas, L.J. (2017, January 21\u201326). Pointnet: Deep learning on point sets for 3d classification and segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition 2017, Honolulu, HI, USA."},{"key":"ref_24","unstructured":"Qi, C.R., Yi, L., Su, H., and Guibas, L.J. (2017). Pointnet++: Deep hierarchical feature learning on point sets in a metric space. Advances in Neural Information Processing Systems, MIT Press."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"63","DOI":"10.5194\/isprs-annals-IV-2-W7-63-2019","article-title":"Semantic Labeling and Refinement of LIDAR Point Clouds Using Deep Neural Network in Urban Areas","volume":"IV-2\/W7","author":"Huang","year":"2019","journal-title":"ISPRS Ann. Photogramm. Remote Sens. Spat. Inf. Sci."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Yang, Y., Feng, C., Shen, Y., and Tian, D. (2018, January 18\u201322). FoldingNet: Point Cloud Auto-Encoder via Deep Grid Deformation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2018), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00029"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Li, J., Chen, B.M., and Hee Lee, G. (2018, January 18\u201322). SO-Net: Self-Organizing Network for Point Cloud Analysis. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2018), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00979"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"511","DOI":"10.1016\/j.neucom.2014.09.013","article-title":"A novel hierarchical approach for multispectral palmprint recognition","volume":"151","author":"Hong","year":"2015","journal-title":"Neurocomputing"},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"18","DOI":"10.1109\/MSP.2017.2693418","article-title":"Geometric Deep Learning: Going beyond Euclidean data","volume":"34","author":"Bronstein","year":"2016","journal-title":"IEEE Signal Process. Mag."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Yi, L., Su, H., Guo, X., and Guibas, L. (2017, January 21\u201326). SyncSpecCNN: Synchronized Spectral CNN for 3D Shape Segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition 2017, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.697"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Masci, J., Boscaini, D., Bronstein, M.M., and Vandergheynst, P. (2015, January 11\u201318). Geodesic Convolutional Neural Networks on Riemannian Manifolds. Proceedings of the IEEE International Conference on Computer Vision Workshop 2015, Santiago, Chile.","DOI":"10.1109\/ICCVW.2015.112"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Su, H., Qi, C.R., Li, Y., and Guibas, L.J. (2015, January 11\u201318). Render for cnn: Viewpoint estimation in images using cnns trained with rendered 3d model views. Proceedings of the IEEE International Conference on Computer Vision 2015, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.308"},{"key":"ref_33","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., and Fei-Fei, L. (2009, January 20\u201325). Imagenet: A large-scale hierarchical image database. Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Wu, X., Hong, D., Ghamisi, P., Li, W., and Tao, R. (2018). Msri-ccf: Multi-scale and rotation-insensitive convolutional channel features for geospatial object detection. Remote Sens., 10.","DOI":"10.3390\/rs10121990"}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/11\/22\/2644\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T13:33:58Z","timestamp":1760189638000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/11\/22\/2644"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,11,13]]},"references-count":35,"journal-issue":{"issue":"22","published-online":{"date-parts":[[2019,11]]}},"alternative-id":["rs11222644"],"URL":"https:\/\/doi.org\/10.3390\/rs11222644","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,11,13]]}}}