{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,29]],"date-time":"2026-05-29T11:38:55Z","timestamp":1780054735533,"version":"3.54.0"},"reference-count":77,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2023,9,18]],"date-time":"2023-09-18T00:00:00Z","timestamp":1694995200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key R&D Program of China","doi-asserted-by":"crossref","award":["2021YFF0901502"],"award-info":[{"award-number":["2021YFF0901502"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61972009"],"award-info":[{"award-number":["61972009"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2024,1,31]]},"abstract":"<jats:p>\n            3D object representation learning is a fundamental challenge in computer vision to infer about the 3D world. Recent advances in deep learning have shown their efficiency in 3D object recognition, among which view-based methods have performed best so far. However, feature learning of multiple views in existing methods is mostly performed in a supervised fashion, which often requires a large amount of data labels with high costs. In contrast, self-supervised learning aims to learn multi-view feature representations without involving labeled data. To this end, we propose a novel self-supervised framework to learn Multi-View Transformation Equivariant Representations (MV-TER), exploring the equivariant transformations of a 3D object and its projected multiple views that we derive. Specifically, we perform a 3D transformation on a 3D object and obtain multiple views before and after the transformation via projection. Then, we train a representation encoding module to capture the intrinsic 3D object representation by decoding 3D transformation parameters from the fused feature representations of multiple views before and after the transformation. Experimental results demonstrate that the proposed MV-TER significantly outperforms the state-of-the-art view-based approaches in 3D object classification and retrieval tasks and show the generalization to real-world datasets. The code is available at\n            <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"url\" xlink:href=\"https:\/\/github.com\/gyshgx868\/mvter\">https:\/\/github.com\/gyshgx868\/mvter<\/jats:ext-link>\n            .\n          <\/jats:p>","DOI":"10.1145\/3597613","type":"journal-article","created":{"date-parts":[[2023,5,22]],"date-time":"2023-05-22T12:01:21Z","timestamp":1684756881000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["Self-supervised Multi-view Learning via Auto-encoding 3D Transformations"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2679-4019","authenticated-orcid":false,"given":"Xiang","family":"Gao","sequence":"first","affiliation":[{"name":"Wangxuan Institute of Computer Technology, Peking University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9860-0922","authenticated-orcid":false,"given":"Wei","family":"Hu","sequence":"additional","affiliation":[{"name":"Wangxuan Institute of Computer Technology, Peking University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3508-1851","authenticated-orcid":false,"given":"Guo-Jun","family":"Qi","sequence":"additional","affiliation":[{"name":"OPPO US Research Center, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,9,18]]},"reference":[{"key":"e_1_3_2_2_2","first-page":"40","volume-title":"Proceedings of the International Conference on Machine Learning (ICML)","author":"Achlioptas Panos","year":"2018","unstructured":"Panos Achlioptas, Olga Diamanti, Ioannis Mitliagkas, and Leonidas Guibas. 2018. Learning representations and generative models for 3D point clouds. In Proceedings of the International Conference on Machine Learning (ICML). PMLR, 40\u201349."},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/3197517.3201301"},{"key":"e_1_3_2_4_2","first-page":"15509","volume-title":"Proceedings of the Conference on Advances in Neural Information Processing Systems (NeurIPS)","author":"Bachman Philip","year":"2019","unstructured":"Philip Bachman, R. Devon Hjelm, and William Buchwalter. 2019. Learning representations by maximizing mutual information across views. In Proceedings of the Conference on Advances in Neural Information Processing Systems (NeurIPS). 15509\u201315519."},{"key":"e_1_3_2_5_2","article-title":"ShapeNet: An information-rich 3D model repository","author":"Chang Angel X.","year":"2015","unstructured":"Angel X. Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su et\u00a0al. 2015. ShapeNet: An information-rich 3D model repository. arXiv preprint arXiv:1512.03012 (2015).","journal-title":"arXiv preprint arXiv:1512.03012"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.5244\/C.28.6"},{"key":"e_1_3_2_7_2","first-page":"1597","volume-title":"Proceedings of the International Conference on Machine Learning (ICML)","author":"Chen Ting","year":"2020","unstructured":"Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A simple framework for contrastive learning of visual representations. In Proceedings of the International Conference on Machine Learning (ICML). PMLR, 1597\u20131607."},{"key":"e_1_3_2_8_2","article-title":"Improved baselines with momentum contrastive learning","author":"Chen Xinlei","year":"2020","unstructured":"Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He. 2020. Improved baselines with momentum contrastive learning. arXiv preprint arXiv:2003.04297 (2020).","journal-title":"arXiv preprint arXiv:2003.04297"},{"key":"e_1_3_2_9_2","first-page":"2990","volume-title":"Proceedings of the International Conference on Machine Learning (ICML)","author":"Cohen Taco","year":"2016","unstructured":"Taco Cohen and Max Welling. 2016. Group equivariant convolutional networks. In Proceedings of the International Conference on Machine Learning (ICML). 2990\u20132999."},{"key":"e_1_3_2_10_2","first-page":"1889","volume-title":"Proceedings of the International Conference on Machine Learning (ICML)","author":"Dieleman Sander","year":"2016","unstructured":"Sander Dieleman, Jeffrey De Fauw, and Koray Kavukcuoglu. 2016. Exploiting cyclic symmetry in convolutional neural networks. In Proceedings of the International Conference on Machine Learning (ICML). 1889\u20131898."},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1093\/mnras\/stv632"},{"key":"e_1_3_2_12_2","first-page":"264","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Feng Yifan","year":"2018","unstructured":"Yifan Feng, Zizhao Zhang, Xibin Zhao, Rongrong Ji, and Yue Gao. 2018. GVCNN: Group-view convolutional neural networks for 3D shape recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 264\u2013272."},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01234-2_7"},{"key":"e_1_3_2_14_2","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Gao Xiang","year":"2020","unstructured":"Xiang Gao, Wei Hu, and Guo-Jun Qi. 2020. GraphTER: Unsupervised learning of graph transformation equivariant representations via auto-encoding node-wise transformations. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/3377876"},{"issue":"4","key":"e_1_3_2_16_2","first-page":"2264","article-title":"Multi-level view associative convolution network for view-based 3D model retrieval","volume":"32","author":"Gao Zan","year":"2021","unstructured":"Zan Gao, Yan Zhang, Hua Zhang, Weili Guan, Dong Feng, and Shengyong Chen. 2021. Multi-level view associative convolution network for view-based 3D model retrieval. IEEE Trans. Circ. Syst. Vid. Technol. 32, 4 (2021), 2264\u20132278.","journal-title":"IEEE Trans. Circ. Syst. Vid. Technol."},{"key":"e_1_3_2_17_2","first-page":"2537","volume-title":"Proceedings of the Conference on Advances in Neural Information Processing Systems (NIPS)","author":"Gens Robert","year":"2014","unstructured":"Robert Gens and Pedro M. Domingos. 2014. Deep symmetry networks. In Proceedings of the Conference on Advances in Neural Information Processing Systems (NIPS). 2537\u20132545."},{"key":"e_1_3_2_18_2","article-title":"Generative adversarial nets","volume":"27","author":"Goodfellow Ian","year":"2014","unstructured":"Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. Generative adversarial nets. Adv. Neural Inf. Process. Syst. 27 (2014).","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_3_2_19_2","article-title":"Deep learning for 3D point clouds: A survey","author":"Guo Yulan","year":"2020","unstructured":"Yulan Guo, Hanyun Wang, Qingyong Hu, Hao Liu, Li Liu, and Mohammed Bennamoun. 2020. Deep learning for 3D point clouds: A survey. IEEE Trans. Pattern Anal. Mach. Intell. (2020).","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2019.2904460"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2018.2868426"},{"key":"e_1_3_2_22_2","first-page":"10441","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV)","author":"Han Zhizhong","year":"2019","unstructured":"Zhizhong Han, Xiyang Wang, Yu-Shen Liu, and Matthias Zwicker. 2019. Multi-angle point cloud-VAE: Unsupervised feature learning for 3D point clouds from multiple angles by joint self-reconstruction and half-to-half prediction. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV). IEEE, 10441\u201310450."},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475172"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00975"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_26_2","first-page":"1945","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","author":"He Xinwei","year":"2018","unstructured":"Xinwei He, Yang Zhou, Zhichao Zhou, Song Bai, and Xiang Bai. 2018. Triplet-center loss for multi-view 3D object retrieval. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 1945\u20131954."},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-21735-7_6"},{"key":"e_1_3_2_28_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Hjelm R. Devon","year":"2019","unstructured":"R. Devon Hjelm, Alex Fedorov, Samuel Lavoie-Marchildon, Karan Grewal, Phil Bachman, Adam Trischler, and Yoshua Bengio. 2019. Learning deep representations by mutual information estimation and maximization. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_29_2","first-page":"1074","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Hu Qianjiang","year":"2021","unstructured":"Qianjiang Hu, Xiao Wang, Wei Hu, and Guo-Jun Qi. 2021. AdCo: Adversarial contrast for efficient learning of unsupervised representations from self-trained negative adversaries. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 1074\u20131083."},{"key":"e_1_3_2_30_2","first-page":"8513","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence (AAAI)","author":"Jiang Jianwen","year":"2019","unstructured":"Jianwen Jiang, Di Bao, Ziqiang Chen, Xibin Zhao, and Yue Gao. 2019. MLVCNN: Multi-loop-view convolutional neural network for 3D shape retrieval. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI). 8513\u20138520."},{"key":"e_1_3_2_31_2","first-page":"5010","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Kanezaki Asako","year":"2018","unstructured":"Asako Kanezaki, Yasuyuki Matsushita, and Yoshifumi Nishida. 2018. RotationNet: Joint object categorization and pose estimation using multiviews from unsupervised viewpoints. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 5010\u20135019."},{"key":"e_1_3_2_32_2","article-title":"Auto-encoding variational Bayes","author":"Kingma Diederik P.","year":"2013","unstructured":"Diederik P. Kingma and Max Welling. 2013. Auto-encoding variational Bayes. arXiv preprint arXiv:1312.6114 (2013).","journal-title":"arXiv preprint arXiv:1312.6114"},{"key":"e_1_3_2_33_2","first-page":"1","volume-title":"Proceedings of the International Conference on Artificial Neural Networks (ICANN)","author":"Kivinen Jyri J.","year":"2011","unstructured":"Jyri J. Kivinen and Christopher K. I. Williams. 2011. Transformation equivariant Boltzmann machines. In Proceedings of the International Conference on Artificial Neural Networks (ICANN). Springer, 1\u20139."},{"key":"e_1_3_2_34_2","first-page":"863","volume-title":"Proceedings of the IEEE International Conference on Computer Vision (CVPR)","author":"Klokov Roman","year":"2017","unstructured":"Roman Klokov and Victor Lempitsky. 2017. Escape from cells: Deep Kd-networks for the recognition of 3D point cloud models. In Proceedings of the IEEE International Conference on Computer Vision (CVPR). 863\u2013872."},{"key":"e_1_3_2_35_2","first-page":"1097","volume-title":"Proceedings of the Conference on Advances in Neural Information Processing Systems (NIPS)","author":"Krizhevsky Alex","year":"2012","unstructured":"Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. 2012. ImageNet classification with deep convolutional neural networks. In Proceedings of the Conference on Advances in Neural Information Processing Systems (NIPS). 1097\u20131105."},{"key":"e_1_3_2_36_2","first-page":"1817","volume-title":"Proceedings of the IEEE International Conference on Robotics and Automation","author":"Lai Kevin","year":"2011","unstructured":"Kevin Lai, Liefeng Bo, Xiaofeng Ren, and Dieter Fox. 2011. A large-scale hierarchical multi-view RGB-D object dataset. In Proceedings of the IEEE International Conference on Robotics and Automation. IEEE, 1817\u20131824."},{"key":"e_1_3_2_37_2","first-page":"9204","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Le Truc","year":"2018","unstructured":"Truc Le and Ye Duan. 2018. PointGrid: A deep network for 3D shape understanding. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 9204\u20139214."},{"key":"e_1_3_2_38_2","first-page":"II\u2013409","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Leibe Bastian","year":"2003","unstructured":"Bastian Leibe and Bernt Schiele. 2003. Analyzing appearance and contour based methods for object categorization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, II\u2013409."},{"key":"e_1_3_2_39_2","first-page":"991","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Lenc Karel","year":"2015","unstructured":"Karel Lenc and Andrea Vedaldi. 2015. Understanding image representations by measuring their equivariance and equivalence. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 991\u2013999."},{"key":"e_1_3_2_40_2","first-page":"8844","volume-title":"Proceedings of the Conference on Advances in Neural Information Processing Systems (NeurIPS)","author":"Lenssen Jan Eric","year":"2018","unstructured":"Jan Eric Lenssen, Matthias Fey, and Pascal Libuschewski. 2018. Group equivariant capsule networks. In Proceedings of the Conference on Advances in Neural Information Processing Systems (NeurIPS). 8844\u20138853."},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2021.3099496"},{"key":"e_1_3_2_42_2","first-page":"820","volume-title":"Proceedings of the Conference on Advances in Neural Information Processing Systems (NeurIPS)","author":"Li Yangyan","year":"2018","unstructured":"Yangyan Li, Rui Bu, Mingchao Sun, Wei Wu, Xinhan Di, and Baoquan Chen. 2018. PointCNN: Convolution on x-transformed points. In Proceedings of the Conference on Advances in Neural Information Processing Systems (NeurIPS). 820\u2013830."},{"key":"e_1_3_2_43_2","first-page":"542","volume-title":"Proceedings of the International Conference on 3D Vision (3DV)","author":"Liu Shikun","year":"2018","unstructured":"Shikun Liu, Lee Giles, and Alexander Ororbia. 2018. Learning a hierarchical latent-variable model of 3D shapes. In Proceedings of the International Conference on 3D Vision (3DV). IEEE, 542\u2013551."},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2021.3090866"},{"key":"e_1_3_2_45_2","first-page":"8895","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Liu Yongcheng","year":"2019","unstructured":"Yongcheng Liu, Bin Fan, Shiming Xiang, and Chunhong Pan. 2019. Relation-shape convolutional neural network for point cloud analysis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 8895\u20138904."},{"key":"e_1_3_2_46_2","first-page":"922","volume-title":"Proceedings of the IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS)","author":"Maturana Daniel","year":"2015","unstructured":"Daniel Maturana and Sebastian Scherer. 2015. VoxNet: A 3D convolutional neural network for real-time object recognition. In Proceedings of the IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 922\u2013928."},{"key":"e_1_3_2_47_2","first-page":"1018","volume-title":"Proceedings of the International Conference on 3D Vision (3DV)","author":"Poursaeed Omid","year":"2020","unstructured":"Omid Poursaeed, Tianxing Jiang, Han Qiao, Nayun Xu, and Vladimir G. Kim. 2020. Self-supervised learning of point clouds via orientation estimation. In Proceedings of the International Conference on 3D Vision (3DV). IEEE, 1018\u20131028."},{"key":"e_1_3_2_48_2","first-page":"652","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Qi Charles R.","year":"2017","unstructured":"Charles R. Qi, Hao Su, Kaichun Mo, and Leonidas J. Guibas. 2017. PointNet: Deep learning on point sets for 3D classification and segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 652\u2013660."},{"key":"e_1_3_2_49_2","first-page":"5648","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Qi Charles R.","year":"2016","unstructured":"Charles R. Qi, Hao Su, Matthias Nie\u00dfner, Angela Dai, Mengyuan Yan, and Leonidas J. Guibas. 2016. Volumetric and multi-view CNNs for object classification on 3D data. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 5648\u20135656."},{"key":"e_1_3_2_50_2","first-page":"5099","volume-title":"Proceedings of the Conference on Advances in Neural Information Processing Systems (NIPS)","author":"Qi Charles Ruizhongtai","year":"2017","unstructured":"Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J. Guibas. 2017. PointNet++: Deep hierarchical feature learning on point sets in a metric space. In Proceedings of the Conference on Advances in Neural Information Processing Systems (NIPS). 5099\u20135108."},{"key":"e_1_3_2_51_2","volume-title":"Proceedings of the IEEE International Conference on Computer Vision (ICCV)","author":"Qi Guo-Jun","year":"2019","unstructured":"Guo-Jun Qi, Liheng Zhang, Chang Wen Chen, and Qi Tian. 2019. AVT: Unsupervised learning of transformation equivariant representations by autoencoding variational transformations. In Proceedings of the IEEE International Conference on Computer Vision (ICCV)."},{"key":"e_1_3_2_52_2","article-title":"Learning generalized transformation equivariant representations via autoencoding transformations","author":"Qi Guo-Jun","year":"2020","unstructured":"Guo-Jun Qi, Liheng Zhang, Feng Lin, and Xiao Wang. 2020. Learning generalized transformation equivariant representations via autoencoding transformations. IEEE Tr. Pattern Anal. Mach. Intell. (2020).","journal-title":"IEEE Tr. Pattern Anal. Mach. Intell."},{"key":"e_1_3_2_53_2","first-page":"3577","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Riegler Gernot","year":"2017","unstructured":"Gernot Riegler, Ali Osman Ulusoy, and Andreas Geiger. 2017. OctNet: Learning deep 3D representations at high resolutions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 3577\u20133586."},{"key":"e_1_3_2_54_2","first-page":"2050","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Schmidt Uwe","year":"2012","unstructured":"Uwe Schmidt and Stefan Roth. 2012. Learning rotation-aware features: From invariant priors to equivariant descriptors. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2050\u20132057."},{"key":"e_1_3_2_55_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Simonyan Karen","year":"2015","unstructured":"Karen Simonyan and Andrew Zisserman. 2015. Very deep convolutional networks for large-scale image recognition. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_56_2","volume-title":"Spherical Tensor Algebra for Biomedical Image Analysis","author":"Skibbe Henrik","year":"2013","unstructured":"Henrik Skibbe. 2013. Spherical Tensor Algebra for Biomedical Image Analysis. Ph. D. Dissertation. Verlag nicht ermittelbar."},{"key":"e_1_3_2_57_2","first-page":"1339","volume-title":"Proceedings of the International Conference on Machine Learning (ICML)","author":"Sohn Kihyuk","year":"2012","unstructured":"Kihyuk Sohn and Honglak Lee. 2012. Learning invariant representations with local transformations. In Proceedings of the International Conference on Machine Learning (ICML). 1339\u20131346."},{"key":"e_1_3_2_58_2","first-page":"945","volume-title":"Proceedings of the IEEE International Conference on Computer Vision (CVPR)","author":"Su Hang","year":"2015","unstructured":"Hang Su, Subhransu Maji, Evangelos Kalogerakis, and Erik Learned-Miller. 2015. Multi-view convolutional neural networks for 3D shape recognition. In Proceedings of the IEEE International Conference on Computer Vision (CVPR). 945\u2013953."},{"key":"e_1_3_2_59_2","first-page":"61","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Sun Yongbin","year":"2020","unstructured":"Yongbin Sun, Yue Wang, Ziwei Liu, Joshua Siegel, and Sanjay Sarma. 2020. PointGrow: Autoregressively learned point cloud generation with self-attention. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. 61\u201370."},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1145\/3240508.3240621"},{"key":"e_1_3_2_62_2","first-page":"1747","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Oord Aaron Van","year":"2016","unstructured":"Aaron Van Oord, Nal Kalchbrenner, and Koray Kavukcuoglu. 2016. Pixel recurrent neural networks. In Proceedings of the International Conference on Machine Learning. PMLR, 1747\u20131756."},{"key":"e_1_3_2_63_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Velickovic Petar","year":"2019","unstructured":"Petar Velickovic, William Fedus, William L. Hamilton, Pietro Li\u00f2, Yoshua Bengio, and R. Devon Hjelm. 2019. Deep Graph Infomax. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_64_2","first-page":"52","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"Wang Chu","year":"2018","unstructured":"Chu Wang, Babak Samari, and Kaleem Siddiqi. 2018. Local spectral graph convolution for point set feature learning. In Proceedings of the European Conference on Computer Vision (ECCV). 52\u201366."},{"key":"e_1_3_2_65_2","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Wang Jiayu","year":"2020","unstructured":"Jiayu Wang, Wengang Zhou, Guo-Jun Qi, Zhongqian Fu, Qi Tian, and Houqiang Li. 2020. Transformation GAN for unsupervised image synthesis and representation learning. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_3_2_66_2","doi-asserted-by":"publisher","DOI":"10.1145\/3450626.3459787"},{"key":"e_1_3_2_67_2","doi-asserted-by":"publisher","DOI":"10.1145\/3326362"},{"key":"e_1_3_2_68_2","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Wei Xin","year":"2020","unstructured":"Xin Wei, Ruixuan Yu, and Jian Sun. 2020. View-GCN: View-based graph convolutional network for 3D shape analysis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_3_2_69_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00985"},{"key":"e_1_3_2_70_2","first-page":"1912","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Wu Zhirong","year":"2015","unstructured":"Zhirong Wu, Shuran Song, Aditya Khosla, Fisher Yu, Linguang Zhang, Xiaoou Tang, and Jianxiong Xiao. 2015. 3D ShapeNets: A deep representation for volumetric shapes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 1912\u20131920."},{"key":"e_1_3_2_71_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3082310"},{"key":"e_1_3_2_72_2","first-page":"206","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Yang Yaoqing","year":"2018","unstructured":"Yaoqing Yang, Chen Feng, Yiru Shen, and Dong Tian. 2018. FoldingNet: Point cloud auto-encoder via deep grid deformation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 206\u2013215."},{"key":"e_1_3_2_73_2","first-page":"7505","volume-title":"Proceedings of the IEEE International Conference on Computer Vision (ICCV)","author":"Yang Ze","year":"2019","unstructured":"Ze Yang and Liwei Wang. 2019. Learning relationships for multi-view 3D object recognition. In Proceedings of the IEEE International Conference on Computer Vision (ICCV). 7505\u20137514."},{"key":"e_1_3_2_74_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3049968"},{"key":"e_1_3_2_75_2","first-page":"186","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Yu Tan","year":"2018","unstructured":"Tan Yu, Jingjing Meng, and Junsong Yuan. 2018. Multi-view harmonized bilinear network for 3D object recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 186\u2013194."},{"key":"e_1_3_2_76_2","first-page":"2547","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Zhang Liheng","year":"2019","unstructured":"Liheng Zhang, Guo-Jun Qi, Liqiang Wang, and Jiebo Luo. 2019. AET vs. AED: Unsupervised representation learning by auto-encoding transformations rather than data. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2547\u20132555."},{"key":"e_1_3_2_77_2","first-page":"6279","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Zhang Yingxue","year":"2018","unstructured":"Yingxue Zhang and Michael Rabbat. 2018. A graph-CNN for 3D point cloud classification. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 6279\u20136283."},{"key":"e_1_3_2_78_2","first-page":"1009","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Zhao Yongheng","year":"2019","unstructured":"Yongheng Zhao, Tolga Birdal, Haowen Deng, and Federico Tombari. 2019. 3D point capsule networks. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 1009\u20131018."}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3597613","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3597613","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T22:50:13Z","timestamp":1750287013000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3597613"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,9,18]]},"references-count":77,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2024,1,31]]}},"alternative-id":["10.1145\/3597613"],"URL":"https:\/\/doi.org\/10.1145\/3597613","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,9,18]]},"assertion":[{"value":"2022-12-19","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-05-09","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-09-18","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}