{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T03:10:29Z","timestamp":1760238629633,"version":"build-2065373602"},"reference-count":39,"publisher":"MDPI AG","issue":"17","license":[{"start":{"date-parts":[[2020,8,28]],"date-time":"2020-08-28T00:00:00Z","timestamp":1598572800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Cooperation Projects of CAS &amp; ITRI","award":["CAS-ITRI201905"],"award-info":[{"award-number":["CAS-ITRI201905"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61873259"],"award-info":[{"award-number":["61873259"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Due to the limitation of less information in a single image, it is very difficult to generate a high-precision 3D model based on the image. There are some problems in the generation of 3D voxel models, e.g., the information loss at the upper level of a network. To solve these problems, we design a 3D model generation network based on multi-modal data constraints and multi-level feature fusion, named as 3DMGNet. Moreover, 3DMGNet is trained by self-supervised method to achieve 3D voxel model generation from an image. The image feature extraction network (2DNet) and 3D feature extraction network (3D auxiliary network) are used to extract the features of the image and 3D voxel model. Then, feature fusion is used to integrate the low-level features into the high-level features in the 3D auxiliary network. To extract more effective features, each layer of the feature map in feature extraction network is processed by an attention network. Finally, the extracted features generate 3D models by a 3D deconvolution network. The feature extraction of 3D model and the generation of voxelization play an auxiliary role in the training of the whole network for the 3D model generation based on an image. Additionally, a multi-view contour constraint method is proposed, to enhance the effect of the 3D model generation. In the experiment, the ShapeNet dataset is adapted to prove the effect of the 3DMGNet, which verifies the robust performance of the proposed method.<\/jats:p>","DOI":"10.3390\/s20174875","type":"journal-article","created":{"date-parts":[[2020,8,28]],"date-time":"2020-08-28T09:17:08Z","timestamp":1598606228000},"page":"4875","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["3DMGNet: 3D Model Generation Network Based on Multi-Modal Data Constraints and Multi-Level Feature Fusion"],"prefix":"10.3390","volume":"20","author":[{"given":"Ende","family":"Wang","sequence":"first","affiliation":[{"name":"Shenyang Institute of Automation, Chinese Academy of Sciences, Shenyang 110016, China"},{"name":"Key Laboratory of Opto-Electronic Information Processing, Chinese Academy of Sciences, Shenyang 110016, China"},{"name":"Key Laboratory of Image Understanding and Computer Vision, Liaoning Province, Shenyang 110016, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lei","family":"Xue","sequence":"additional","affiliation":[{"name":"School of Information Science and Engineering, Shenyang Ligong University, Shenyang 110159, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9430-0914","authenticated-orcid":false,"given":"Yong","family":"Li","sequence":"additional","affiliation":[{"name":"College of Information Science and Engineering, Northeastern University, Shenyang 110819, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0076-7311","authenticated-orcid":false,"given":"Zhenxin","family":"Zhang","sequence":"additional","affiliation":[{"name":"Key Lab of 3D Information Acquisition and Application, MOE, and College of Resource Environment and Tourism, Capital Normal University, Beijing 100048, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xukui","family":"Hou","sequence":"additional","affiliation":[{"name":"Shenyang Institute of Automation, Chinese Academy of Sciences, Shenyang 110016, China"},{"name":"Key Laboratory of Opto-Electronic Information Processing, Chinese Academy of Sciences, Shenyang 110016, China"},{"name":"Key Laboratory of Image Understanding and Computer Vision, Liaoning Province, Shenyang 110016, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,8,28]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1586","DOI":"10.1109\/TMC.2019.2913364","article-title":"Furion: Engineering High-Quality Immersive Virtual Reality on Today\u2019s Mobile Devices","volume":"19","author":"Lai","year":"2020","journal-title":"IEEE Trans. Mob. Comput."},{"key":"ref_2","unstructured":"Chang, A.X., Funkhouser, T., Guibas, L., Hanrahan, P., Huang, Q., Li, Z., Savarese, S., Savva, M., Song, S., and Su, H. (2015). ShapeNet: An Information-Rich 3D Model Repository. arXiv, Available online: https:\/\/arxiv.org\/abs\/1512.03012."},{"key":"ref_3","first-page":"1","article-title":"Inferring 3D Shapes from Image Collections using Adversarial","volume":"128","author":"Matheus","year":"2020","journal-title":"Netw. Int. J. Comput. Vison"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Wang, N., Zhang, Y., Li, Z., Fu, Y., Liu, W., and Jiang, Y.G. (2018, January 8\u201314). Pixel2Mesh: Generating 3D Mesh Models from Single RGB Images. Proceedings of the European Conference on Computer Vision (ECCV 2018), Munich, Germany.","DOI":"10.1007\/978-3-030-01252-6_4"},{"key":"ref_5","unstructured":"Brilakis, L., and Haas, C. (2020). Infrastructure Computer Vision, Butterworch-Heinemann."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"105574","DOI":"10.1016\/j.knosys.2020.105574","article-title":"PGNet: A Part-based Generative Network for 3D object reconstruction","volume":"194","author":"Zhang","year":"2020","journal-title":"Knowl. -Based Syst."},{"key":"ref_7","unstructured":"Wu, J., Zhang, C., Xue, T., Freeman, W.T., and Tenenbaum, J.B. (2016). Learning a Probabilistic Latent Space of Object Shapes via 3D Generative-Adversarial Modeling. Advances in Neural Information Processing Systems 29 (NIPS 2016), Proceedings of the Neural Information Processing Systems 2016, Barcelona, Spain, 5\u201310 December 2016, MIT Press."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"221","DOI":"10.1109\/TPAMI.2012.59","article-title":"3D Convolutional Neural Networks for Human Action Recognition","volume":"35","author":"Ji","year":"2012","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Girdhar, R., Fouhey, D.F., Rodriguez, M., and Gupta, A. (2016). Learning a Predictable and Generative Vector Representation for Objects. arXiv, Available online: https:\/\/arxiv.org\/abs\/1603.08637.","DOI":"10.1007\/978-3-319-46466-4_29"},{"key":"ref_10","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv, Available online: https:\/\/arxiv.org\/abs\/1409.1556."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep Residual Learning for Image Recognition. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A. (2015, January 7\u201312). Going deeper with convolutions. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Tchapmi, L., Choy, C., Armeni, I., Gwak, J., and Savarese, S. (2017, January 10\u201312). SEGCloud: Semantic Segmentation of 3D Point Clouds. Proceedings of the 2017 International Conference on 3D Vision (3DV), Qingdao, China.","DOI":"10.1109\/3DV.2017.00067"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"436","DOI":"10.1038\/nature14539","article-title":"Deep learning","volume":"521","author":"LeCun","year":"2015","journal-title":"Nature"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Hu, J., Shen, L., and Sun, G. (2018, January 18\u201322). Squeeze-and-Excitation Networks. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPRW), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"ref_16","unstructured":"Ahmed, E., Saint, A., Shabayek, A.E.R., Cherenkova, K., Das, R., Gusev, G., Aouada, D., and Ottersten, B. (2018). A survey on deep learning advances on different 3D data representations. arXiv, Available online: https:\/\/arxiv.org\/abs\/1808.01462."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Maturana, D., and Scherer, S. (October, January 28). VoxNet: A 3D Convolutional Neural Network for real-time object recognition. Proceedings of the 2015 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Hamburg, Germany.","DOI":"10.1109\/IROS.2015.7353481"},{"key":"ref_18","unstructured":"Abhishek, S., Oliver, G., and Mario, F. (2016, January 8\u201310). VConv-DAE: Deep Volumetric Shape Learning Without Object Labels. Proceedings of the Geometry Meets Deep Learning Workshop at European Conference on Computer Vision (ECCV-W), Amsterdam, The Netherlands."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"199","DOI":"10.1016\/j.cag.2017.10.007","article-title":"Toward real-time 3D object recognition: A lightweight volumetric CNN framework using multitask learning","volume":"71","author":"Zhi","year":"2018","journal-title":"Comput. Graph."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Jing, L., and Tian, Y. (2020). Self-supervised Visual Feature Learning with Deep Neural Networks: A Survey. IEEE Trans. Pattern Anal. Mach. Intell., 1.","DOI":"10.1109\/TPAMI.2020.2992393"},{"key":"ref_21","unstructured":"Xu, B., Zhang, X., Li, Z., Leotta, M., Chang, S., and Shan, J. (2020). Deep Learning Guided Building Reconstruction from Satellite Imagery-derived Point Clouds. arXiv, Available online: https:\/\/arxiv.org\/abs\/2005.09223."},{"key":"ref_22","unstructured":"Wu, Z., Song, S., Khosla, A., Yu, F., Zhang, L., Tang, X., Xiao, J., and Fisher, Y. (2015, January 7\u201312). 3D ShapeNets: A deep representation for volumetric shapes. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Choy, C.B., Xu, D., Gwak, J., Chen, K., and Savarese, S. (2016). 3D-R2N2: A Unified Approach for Single and Multi-view 3D Object Reconstruction. arXiv, Available online: https:\/\/arxiv.org\/abs\/1604.00449.","DOI":"10.1007\/978-3-319-46484-8_38"},{"key":"ref_24","unstructured":"Mirza, M., and Osindero, S. (2014). Conditional generative adversarial nets. arXiv, Available online: https:\/\/arxiv.org\/abs\/1411.1784."},{"key":"ref_25","first-page":"1","article-title":"Sagnet: Structure-aware generative network for 3d-shape modeling","volume":"38","author":"Wu","year":"2019","journal-title":"ACM Trans. Graphic."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Mandikal, P., Navaneet, K.L., Agarwal, M., and Babu, R.V. (2018, January 3\u20136). 3D-LMNet: Latent embedding matching for accurate and diverse 3D point cloud reconstruction from a single image. Proceedings of the British Machine Vision Conference 2018 (BVLC), Newcastle, UK.","DOI":"10.1007\/978-3-030-11015-4_50"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Gadelha, M., Wang, R., and Maji, S. (2018, January 8\u201314). Multiresolution tree networks for 3d point cloud processing. Proceedings of the 15th European Conference on Computer Vision (ECCV 2018), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_7"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Jiang, L., Shi, S., Qi, X., and Jia, J. (2018, January 8\u201314). Gal: Geometric adversarial loss for single-view 3d-object reconstruction. Proceedings of the 15th European Conference on Computer Vision (ECCV 2018), Munich, Germany.","DOI":"10.1007\/978-3-030-01237-3_49"},{"key":"ref_29","unstructured":"Li, C.L., Zaheer, M., Zhang, Y., Poczos, B., and Salakhutdinov, R. (2018). Point cloud Gan. arXiv, Available online: https:\/\/arxiv.org\/abs\/1810.05795."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Mandikal, P., and Radhakrishnan, V.B. (2019, January 7\u201311). Dense 3d point cloud reconstruction using a deep pyramid network. Proceedings of the 2019 IEEE Winter Conference on Applications of Computer Vision (WACV), Waikoloa Village, HI, USA.","DOI":"10.1109\/WACV.2019.00117"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Sinha, A., Unmesh, A., Huang, Q., and Ramani, K. (2017, January 21\u201326). Surfnet: Generating 3d shape surfaces using deep residual networks. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.91"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Pumarola, A., Agudo, A., Porzi, L., Sanfeliu, A., Lepetit, V., and Moreno-Noguer, F. (2018, January 18\u201322). Geometry-aware network for non-rigid shape prediction from a single view. Proceedings of the 2018 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00492"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Groueix, T., Fisher, M., Kim, V.G., Russell, B.C., and Aubry, M. (2018, January 18\u201322). A papier-m\u00e2ch\u00e9 approach to learning 3d surface generation. Proceedings of the 2018 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00030"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L., Li, K., and Li, F. (2009, January 20\u201325). Imagenet: A large-scale hierarchical image database. Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Zeiler, M.D., Krishnan, D., Taylor, G.W., and Fergus, R. (2010, January 13\u201318). Deconvolutional networks. Proceedings of the 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPRW), San Francisco, CA, USA.","DOI":"10.1109\/CVPR.2010.5539957"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Tatarchenko, M., Dosovitskiy, A., and Brox, T. (2017, January 22\u201329). Octree generating networks: Efficient convolutional architectures forhigh-resolution 3D outputs. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.230"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Tulsiani, S., Zhou, T., Efros, A.A., and Malik., J. (2017, January 21\u201326). Multi-view supervision for single-view reconstruction via differentiable ray consistency. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.30"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Xie, H., Yao, H., Sun, X., Zhou, S., and Zhang, S. (November, January 27). Pix2vox: Context-aware 3d reconstruction from single and multi-view images. Proceedings of the 2019 IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Korea.","DOI":"10.1109\/ICCV.2019.00278"},{"key":"ref_39","first-page":"1","article-title":"Pix2Vox++: Multi-scale Context-aware 3D Object Reconstruction from Single and Multiple Images","volume":"1","author":"Xie","year":"2020","journal-title":"Int. J. Comput. Vis."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/17\/4875\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T10:04:11Z","timestamp":1760177051000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/17\/4875"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,8,28]]},"references-count":39,"journal-issue":{"issue":"17","published-online":{"date-parts":[[2020,9]]}},"alternative-id":["s20174875"],"URL":"https:\/\/doi.org\/10.3390\/s20174875","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2020,8,28]]}}}