{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,14]],"date-time":"2026-01-14T15:16:13Z","timestamp":1768403773492,"version":"3.49.0"},"reference-count":57,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2020,11,30]],"date-time":"2020-11-30T00:00:00Z","timestamp":1606694400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Open Project Program of the State Key Lab of CAD 8 CG"},{"DOI":"10.13039\/501100004835","name":"Zhejiang University","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100004835","id-type":"DOI","asserted-by":"crossref"}]},{"name":"2019 Tianjin New Generation Artificial Intelligence Major Program"},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61772359, 61572356, and 61872267"],"award-info":[{"award-number":["61772359, 61572356, and 61872267"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"2018 Tianjin New Generation Artificial Intelligence Major Program","award":["18ZXZNGX00150"],"award-info":[{"award-number":["18ZXZNGX00150"]}]},{"name":"Elite Scholar Program of Tianjin University","award":["2019XRX-0035"],"award-info":[{"award-number":["2019XRX-0035"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2020,11,30]]},"abstract":"<jats:p>In recent years, research into 3D shape recognition in the field of multimedia and computer vision has attracted wide attention. With the rapid development of deep learning, various deep models have achieved state-of-the-art performance based on different representations. There are many modalities for representing a 3D model, such as point cloud, multiview, and panorama view. Deep learning models based on these different modalities have different concerns, and all of them have achieved high performance for 3D shape recognition. However, all of these methods ignore the multimodality information in conditions where the same 3D model is represented by different modalities. Thus, we can obtain a better descriptor by guiding the training to consider these multiple representations. In this article, we propose MMFN, a novel multimodal fusion network for 3D shape recognition that employs correlations between the different modalities to generate a fused descriptor, which is more robust. In particular, we design two novel loss functions to help the model learn the correlation information during training. The first is correlation loss, which focuses on the correlations among different descriptors generated from different structures. This approach reduces the training time and improves the robustness of the fused descriptor of the 3D model. The second is instance loss, which preserves the independence of each modality and utilizes feature differentiation to guide model learning during the training process. More specifically, we use the weighted fusion method, which applies statistical methods to obtain robust descriptors that maximize the advantages of the information from the different modalities. We evaluated the proposed method on the ModelNet40 and ShapeNetCore55 datasets for 3D shape classification and retrieval tasks. The experimental results and comparisons with state-of-the-art methods demonstrate the superiority of our approach.<\/jats:p>","DOI":"10.1145\/3410439","type":"journal-article","created":{"date-parts":[[2020,12,17]],"date-time":"2020-12-17T17:49:26Z","timestamp":1608227366000},"page":"1-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":21,"title":["MMFN"],"prefix":"10.1145","volume":"16","author":[{"given":"Weizhi","family":"Nie","sequence":"first","affiliation":[{"name":"Tianjin University, TianJin, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qi","family":"Liang","sequence":"additional","affiliation":[{"name":"Tianjin University, TianJin, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yixin","family":"Wang","sequence":"additional","affiliation":[{"name":"Tianjin University, TianJin, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xing","family":"Wei","sequence":"additional","affiliation":[{"name":"Guilin University of Aerospace Technology"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuting","family":"Su","sequence":"additional","affiliation":[{"name":"Tianjin University, TianJin, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2020,12,17]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2017.2652071"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1111\/1467-8659.00669"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2017.2757769"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.691"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2018.2817042"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/78956.78958"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33018279"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00035"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2012.2199502"},{"key":"e_1_2_1_10_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR'17)","author":"Qi Charles R."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCYB.2016.2593940"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2016.2609814"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2016.2582532"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCYB.2017.2778764"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2018.2816821"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2019.2904460"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33018376"},{"key":"e_1_2_1_18_1","first-page":"2","article-title":"SeqViews2SeqLabels: Learning 3D global features via aggregating sequential views by RNN with attention","volume":"28","author":"Han Z.","year":"2019","journal-title":"IEEE Transactions on Image Processing"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.01054"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00208"},{"key":"e_1_2_1_21_1","unstructured":"Geoffrey Hinton Jeff Dean and Oriol Vinyals. 2014. Distilling the knowledge in a neural network. arXiv:1503.02531.  Geoffrey Hinton Jeff Dean and Oriol Vinyals. 2014. Distilling the knowledge in a neural network. arXiv:1503.02531."},{"key":"e_1_2_1_22_1","volume-title":"Advances in Neural Information Processing Systems 26: The 27th Annual Conference on Neural Information Processing Systems","author":"Johnson Rie","year":"2013"},{"key":"e_1_2_1_23_1","volume-title":"First Eurographics Symposium on Geometry Processing","volume":"43","author":"Kazhdan Michael M.","year":"2003"},{"key":"e_1_2_1_24_1","volume-title":"Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV\u201917)","author":"Klokov Roman"},{"key":"e_1_2_1_25_1","first-page":"2579","article-title":"Visualizing data using t-SNE","volume":"9","author":"Maaten Laurens Van Der","year":"2008","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-011-0873-3"},{"key":"e_1_2_1_27_1","unstructured":"Yangyan Li Rui Bu Mingchao Sun Wei Wu Xinhan Di and Baoquan Chen. 2018. PointCNN: Convolution on X-transformed points. In Advances in Neural Information Processing Systems 31 (NIPS\u201918). 828--838.  Yangyan Li Rui Bu Mingchao Sun Wei Wu Xinhan Di and Baoquan Chen. 2018. PointCNN: Convolution on X-transformed points. In Advances in Neural Information Processing Systems 31 (NIPS\u201918). 828--838."},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2018.2810191"},{"key":"e_1_2_1_29_1","volume-title":"Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201917)","author":"Liu H."},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33018778"},{"key":"e_1_2_1_31_1","volume-title":"Proceedings of the 2015 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS\u201915)","author":"Maturana Daniel"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3351009"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2016.2616143"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2006.12.026"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2019.2929600"},{"key":"e_1_2_1_36_1","volume-title":"Proceedings of the 2019 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP\u201919)","author":"Peng J."},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2017.134"},{"key":"e_1_2_1_38_1","volume-title":"Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201916)","author":"Qi Charles Ruizhongtai"},{"key":"e_1_2_1_39_1","volume-title":"Proceedings of Advances in Neural Information Processing Systems 30 (NIPS\u201917)","author":"Qi Charles Ruizhongtai"},{"key":"e_1_2_1_40_1","volume-title":"Proceedings of the Eurographics Workshop on 3D Object Retrieval (3DOR\u201917)","author":"Savva Manolis","year":"2017"},{"key":"e_1_2_1_41_1","volume-title":"Ensemble of PANORAMA-based convolutional neural networks for 3D model classification and retrieval. Computers 8 Graphics 71","author":"Sfikas Konstantinos","year":"2018"},{"key":"e_1_2_1_42_1","volume-title":"Proceedings of the Eurographics Workshop on 3D Object Retrieval (3DOR\u201917)","author":"Sfikas Konstantinos","year":"2017"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2018.2869747"},{"key":"e_1_2_1_44_1","volume-title":"Proceedings of the 25th International Conference on Neural Information Processing Systems (NIPS\u201912)","author":"Socher Richard"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00268"},{"key":"e_1_2_1_46_1","volume-title":"Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV\u201915)","author":"Su Hang"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00371-008-0304-2"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/3326362"},{"key":"e_1_2_1_49_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201915)","author":"Wu Zhirong","year":"2015"},{"key":"e_1_2_1_50_1","volume-title":"Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201915)","author":"Wu Zhirong","year":"2015"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.385"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2596722"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2017.06.034"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/3240508.3240702"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33019119"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2018.2862625"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2018.2881469"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3410439","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3410439","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T21:31:51Z","timestamp":1750195911000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3410439"}},"subtitle":["Multimodal Information Fusion Networks for 3D Model Classification and Retrieval"],"short-title":[],"issued":{"date-parts":[[2020,11,30]]},"references-count":57,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2020,11,30]]}},"alternative-id":["10.1145\/3410439"],"URL":"https:\/\/doi.org\/10.1145\/3410439","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,11,30]]},"assertion":[{"value":"2020-01-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-07-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-12-17","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}