{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,20]],"date-time":"2026-07-20T17:03:05Z","timestamp":1784566985377,"version":"3.55.0"},"reference-count":69,"publisher":"Springer Science and Business Media LLC","issue":"6","license":[{"start":{"date-parts":[[2026,5,15]],"date-time":"2026-05-15T00:00:00Z","timestamp":1778803200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,5,15]],"date-time":"2026-05-15T00:00:00Z","timestamp":1778803200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100002920","name":"Research Grants Council, University Grants Committee","doi-asserted-by":"publisher","award":["11219324"],"award-info":[{"award-number":["11219324"]}],"id":[{"id":"10.13039\/501100002920","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Comput Vis"],"published-print":{"date-parts":[[2026,6]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    Cross-modal contrastive distillation has recently been explored for learning effective 3D representations. However, existing methods focus primarily on modality-shared features, neglecting the modality-specific features during the pre-training process, which leads to suboptimal representations. In this paper, we theoretically analyze the limitations of current contrastive methods for 3D representation learning and propose a new framework, namely CMCR (Cross-Modal Comprehensive Representation Learning), to address these shortcomings. Our approach improves upon traditional methods by better integrating both modality-shared and modality-specific features. Specifically, we introduce masked image modeling and occupancy estimation tasks to guide the network in learning more comprehensive modality-specific features. Furthermore, we introduce a novel multi-modal unified codebook that learns an embedding space shared across different modalities. Besides, we propose geometry-enhanced masked image modeling to further boost 3D representation learning. Extensive experiments demonstrate that our method mitigates the challenges faced by traditional approaches and consistently outperforms existing image-to-LiDAR contrastive distillation methods in downstream tasks. Code will be available at\u00a0\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/github.com\/Eaphan\/CMCR\" ext-link-type=\"uri\">https:\/\/github.com\/Eaphan\/CMCR<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1007\/s11263-026-02879-z","type":"journal-article","created":{"date-parts":[[2026,5,15]],"date-time":"2026-05-15T14:27:32Z","timestamp":1778855252000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Is Contrastive Distillation Enough for Learning Comprehensive 3D Representations?"],"prefix":"10.1007","volume":"134","author":[{"given":"Yifan","family":"Zhang","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3431-2021","authenticated-orcid":false,"given":"Junhui","family":"Hou","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,5,15]]},"reference":[{"key":"2879_CR1","doi-asserted-by":"crossref","unstructured":"Behley, J., Garbade, M., Milioto, A., Quenzel, J., Behnke, S., Stachniss, C., & Gall. J. (2019). Semantickitti: A dataset for semantic scene understanding of lidar sequences. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, (pp. 9297\u20139307).","DOI":"10.1109\/ICCV.2019.00939"},{"key":"2879_CR2","doi-asserted-by":"crossref","unstructured":"Berman, M., Triki, A.R., & Blaschko, M.B. (2018). The lov\u00e1sz-softmax loss: A tractable surrogate for the optimization of the intersection-over-union measure in neural networks. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, (pp. 4413\u20134421).","DOI":"10.1109\/CVPR.2018.00464"},{"key":"2879_CR3","doi-asserted-by":"crossref","unstructured":"Boulch, A., Sautier, C., Michele, B., Puy, G., & Marlet, R. (2023). Also: Automotive lidar self-supervision by occupancy estimation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, (pp. 13455\u201313465).","DOI":"10.1109\/CVPR52729.2023.01293"},{"key":"2879_CR4","doi-asserted-by":"crossref","unstructured":"Caesar, H., Bankiti, V., Lang, A.H., Vora, S., Liong, V.E., Xu, Q., Krishnan, A., Pan, Y., Baldan, G., & Beijbom, O. (2020). nuscenes: A multimodal dataset for autonomous driving. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, (pp 11621\u201311631).","DOI":"10.1109\/CVPR42600.2020.01164"},{"key":"2879_CR5","doi-asserted-by":"crossref","unstructured":"Caron, M., Touvron, H., Misra, I., J\u00e9gou, H., Mairal, J., Bojanowski, P., & Joulin, A. (2021). Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, (pp. 9650\u20139660).","DOI":"10.1109\/ICCV48922.2021.00951"},{"key":"2879_CR6","doi-asserted-by":"publisher","first-page":"221","DOI":"10.1609\/aaai.v36i1.19897","volume":"36","author":"C Chen","year":"2022","unstructured":"Chen, C., Chen, Z., Zhang, J., & Tao, D. (2022). Sasa: Semantics-augmented set abstraction for point-based 3d object detection. Proceedings of the AAAI Conference on Artificial Intelligence, 36, 221\u2013229.","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"2879_CR7","doi-asserted-by":"crossref","unstructured":"Chen, H., Zhang, Z., Qu, Y., Zhang, R., Tan, X., & Xie, Y. (2024). Building a strong pre-training baseline for universal 3d large-scale perception. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, (pp. 19925\u201319935).","DOI":"10.1109\/CVPR52733.2024.01883"},{"key":"2879_CR8","unstructured":"Chen, X., Fan, H., Girshick, R., & He, K. (2020). Improved baselines with momentum contrastive learning. arXiv preprint arXiv:2003.04297."},{"key":"2879_CR9","doi-asserted-by":"crossref","unstructured":"Chen, Y., Nie\u00dfner, M., & Dai, A. (2022b). 4dcontrast: Contrastive learning with dynamic correspondences for 3d scene understanding. In European Conference on Computer Vision, (pp. 543\u2013560).","DOI":"10.1007\/978-3-031-19824-3_32"},{"key":"2879_CR10","doi-asserted-by":"crossref","unstructured":"Chen, Y., Yuan, J., Tian, Y., Geng, S., Li, X., Zhou, D., Metaxas, D.N., & Yang, H. (2023). Revisiting multimodal representation in contrastive learning: from patch and token embeddings to finite discrete tokens. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, (pp. 15095\u201315104).","DOI":"10.1109\/CVPR52729.2023.01449"},{"key":"2879_CR11","doi-asserted-by":"crossref","unstructured":"Chibane, J., Engelmann, F., Anh\u00a0Tran, T., & Pons-Moll, G. (2022). Box2mask: Weakly supervised 3d semantic instance segmentation using bounding boxes. In European Conference on Computer Vision, (pp. 681\u2013699).","DOI":"10.1007\/978-3-031-19821-2_39"},{"key":"2879_CR12","doi-asserted-by":"crossref","unstructured":"Choe, J., Park, C., Rameau, F., Park, J., & Kweon, I.S. (2022). Pointmixer: Mlp-mixer for point cloud understanding. In European Conference on Computer Vision, Springer, (pp. 620\u2013640).","DOI":"10.1007\/978-3-031-19812-0_36"},{"key":"2879_CR13","doi-asserted-by":"crossref","unstructured":"Choy, C., Gwak, J., & Savarese, S. (2019). 4d spatio-temporal convnets: Minkowski convolutional neural networks. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, (pp. 3075\u20133084).","DOI":"10.1109\/CVPR.2019.00319"},{"key":"2879_CR14","doi-asserted-by":"crossref","unstructured":"Fadadu, S., Pandey, S., Hegde, D., Shi, Y., Chou, F.C., Djuric, N., & Vallespi-Gonzalez, C. (2022). Multi-view fusion of sensor data for improved perception and prediction in autonomous driving. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, (pp. 2349\u20132357).","DOI":"10.1109\/WACV51458.2022.00335"},{"issue":"1","key":"2879_CR15","doi-asserted-by":"publisher","first-page":"259","DOI":"10.1109\/18.272494","volume":"40","author":"M Feder","year":"1994","unstructured":"Feder, M., & Merhav, N. (1994). Relations between entropy and error probability. IEEE Transactions on Information theory, 40(1), 259\u2013266.","journal-title":"IEEE Transactions on Information theory"},{"key":"2879_CR16","doi-asserted-by":"crossref","unstructured":"He, K., Chen, X., Xie, S., Li, Y., Doll\u00e1r, P., & Girshick, R. (2022). Masked autoencoders are scalable vision learners. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, (pp. 16000\u201316009).","DOI":"10.1109\/CVPR52688.2022.01553"},{"key":"2879_CR17","doi-asserted-by":"crossref","unstructured":"Hess, G., Jaxing, J., Svensson, E., Hagerman, D., Petersson, C., & Svensson, L. (2023). Masked autoencoder for self-supervised pre-training on lidar point clouds. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, (pp. 350\u2013359).","DOI":"10.1109\/WACVW58289.2023.00039"},{"key":"2879_CR18","doi-asserted-by":"publisher","first-page":"49100","DOI":"10.52202\/075280-2134","volume":"36","author":"CJ Ho","year":"2023","unstructured":"Ho, C. J., Tai, C. H., Lin, Y. Y., Yang, M. H., & Tsai, Y. H. (2023). Diffusion-ss3d: Diffusion model for semi-supervised 3d object detection. Advances in Neural Information Processing Systems, 36, 49100\u201349112.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2879_CR19","doi-asserted-by":"crossref","unstructured":"Huang, S., Xie, Y., Zhu, S.C., & Zhu, Y. (2021). Spatio-temporal self-supervised representation learning for 3d point clouds. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, (pp. 6535\u20136545).","DOI":"10.1109\/ICCV48922.2021.00647"},{"key":"2879_CR20","doi-asserted-by":"crossref","unstructured":"Jiang, P., Osteen, P., Wigness, M., & Saripallig, S. (2021). Rellis-3d dataset: Data, benchmarks and analysis. In IEEE International Conference on Robotics and Automation, (pp. 1110\u20131116).","DOI":"10.1109\/ICRA48506.2021.9561251"},{"key":"2879_CR21","doi-asserted-by":"publisher","first-page":"79341","DOI":"10.1109\/ACCESS.2023.3298706","volume":"11","author":"AA Klokov","year":"2023","unstructured":"Klokov, A. A., Pak, D. U., Khorin, A., Yudin, D. A., Kochiev, L., Luchinskiy, V. D., & Bezuglyj, V. D. (2023). Daps3d: Domain adaptive projective segmentation of 3d lidar point clouds. IEEE Access, 11, 79341\u201379356.","journal-title":"IEEE Access"},{"key":"2879_CR22","doi-asserted-by":"crossref","unstructured":"Kong, L., Liu, Y., Chen, R., Ma, Y., Zhu, X., Li, Y., Hou, Y., Qiao, Y., & Liu, Z. (2023a). Rethinking range view representation for lidar segmentation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, (pp. 228\u2013240).","DOI":"10.1109\/ICCV51070.2023.00028"},{"key":"2879_CR23","doi-asserted-by":"crossref","unstructured":"Kong, L., Ren, J., Pan, L., & Liu, Z. (2023b). Lasermix for semi-supervised lidar semantic segmentation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, (pp. 21705\u201321715).","DOI":"10.1109\/CVPR52729.2023.02079"},{"key":"2879_CR24","doi-asserted-by":"publisher","first-page":"32971","DOI":"10.52202\/075280-1430","volume":"36","author":"PP Liang","year":"2023","unstructured":"Liang, P. P., Deng, Z., Ma, M. Q., Zou, J. Y., Morency, L. P., & Salakhutdinov, R. (2023). Factorized contrastive learning: Going beyond multi-view redundancy. Advances in Neural Information Processing Systems, 36, 32971\u201332998.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2879_CR25","doi-asserted-by":"publisher","first-page":"3351","DOI":"10.1609\/aaai.v38i4.28121","volume":"38","author":"G Liao","year":"2024","unstructured":"Liao, G., Li, J., & Ye, X. (2024). Vlm2scene: Self-supervised image-text-lidar learning with foundation models for autonomous driving scene understanding. Proceedings of the AAAI Conference on Artificial Intelligence, 38, 3351\u20133359.","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"2879_CR26","doi-asserted-by":"crossref","unstructured":"Liu, A.H., Jin, S., Lai, C., Rouditchenko, A., Oliva, A., & Glass, J.R. (2022a). Cross-modal discrete representation learning. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, (pp. 3013\u20133035).","DOI":"10.18653\/v1\/2022.acl-long.215"},{"key":"2879_CR27","doi-asserted-by":"publisher","first-page":"53433","DOI":"10.52202\/075280-2325","volume":"36","author":"K Liu","year":"2023","unstructured":"Liu, K., Zhan, F., Zhang, J., Xu, M., Yu, Y., El Saddik, A., Theobalt, C., Xing, E., & Lu, S. (2023). Weakly supervised 3d open-vocabulary segmentation. Advances in Neural Information Processing Systems, 36, 53433\u201353456.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2879_CR28","doi-asserted-by":"crossref","unstructured":"Liu, M., Zhou, Y., Qi, C.R., Gong, B., Su, H., & Anguelov, D. (2022b). Less: Label-efficient semantic segmentation for lidar point clouds. In European Conference on Computer Vision, (pp. 70\u201389).","DOI":"10.1007\/978-3-031-19842-7_5"},{"key":"2879_CR29","doi-asserted-by":"publisher","first-page":"37193","DOI":"10.52202\/075280-1617","volume":"36","author":"Y Liu","year":"2023","unstructured":"Liu, Y., Kong, L., Cen, J., Chen, R., Zhang, W., Pan, L., Chen, K., & Liu, Z. (2023). Segment any point cloud sequences by distilling vision foundation models. Advances in Neural Information Processing Systems, 36, 37193\u201337229.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2879_CR30","unstructured":"Liu, Y.C., Huang, Y.K., Chiang, H.Y., Su, H.T., Liu, Z.Y., Chen, C.T., Tseng, C.Y., & Hsu, W.H. (2021). Learning from 2d: Contrastive pixel-to-point knowledge transfer for 3d pretraining. arXiv preprint arXiv:2104.04687."},{"key":"2879_CR31","unstructured":"Luo, Y., Chen, Z., Wang, Z., Yu, X., Huang, Z., & Baktashmotlagh, M. (2023). Exploring active 3d object detection from a generalization perspective. In The Eleventh International Conference on Learning Representations, (pp. 1\u201313)."},{"key":"2879_CR32","doi-asserted-by":"crossref","unstructured":"Mahmoud, A., Hu, J.S., Kuai, T., Harakeh, A., Paull, L., & Waslander, S.L. (2023). Self-supervised image-to-point distillation via semantically tolerant contrastive loss. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, (pp. 7102\u20137110).","DOI":"10.1109\/CVPR52729.2023.00686"},{"issue":"2","key":"2879_CR33","doi-asserted-by":"publisher","first-page":"2116","DOI":"10.1109\/LRA.2022.3142440","volume":"7","author":"L Nunes","year":"2022","unstructured":"Nunes, L., Marcuzzi, R., Chen, X., Behley, J., & Stachniss, C. (2022). Segcontrast: 3d point cloud feature representation learning through self-supervised segment discrimination. IEEE Robotics and Automation Letters, 7(2), 2116\u20132123.","journal-title":"IEEE Robotics and Automation Letters"},{"key":"2879_CR34","doi-asserted-by":"crossref","unstructured":"Nunes, L., Wiesmann, L., Marcuzzi, R., Chen, X., Behley, J., & Stachniss, C. (2023). Temporal consistent 3d lidar representation learning for semantic perception in autonomous driving. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, (pp. 5217\u20135228).","DOI":"10.1109\/CVPR52729.2023.00505"},{"key":"2879_CR35","unstructured":"Oord, A.v.d., Li, Y., & Vinyals, O. (2018). Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748."},{"key":"2879_CR36","first-page":"1","volume":"2024","author":"M Oquab","year":"2024","unstructured":"Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., et al. (2024). Dinov2: Learning robust visual features without supervision. Trans Mach Learn Res, 2024, 1\u201328.","journal-title":"Trans Mach Learn Res"},{"key":"2879_CR37","doi-asserted-by":"crossref","unstructured":"Pan, Y., Gao, B., Mei, J., Geng, S., Li, C., & Zhao, H. (2020). Semanticposs: A point cloud dataset with large quantity of dynamic instances. In IEEE Intelligent Vehicles Symposium, (pp. 687\u2013693).","DOI":"10.1109\/IV47402.2020.9304596"},{"key":"2879_CR38","doi-asserted-by":"crossref","unstructured":"Pang, B., Xia, H., & Lu, C. (2023). Unsupervised 3d point cloud representation learning by triangle constrained contrast for autonomous driving. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, (pp. 5229\u20135239).","DOI":"10.1109\/CVPR52729.2023.00506"},{"key":"2879_CR39","doi-asserted-by":"crossref","unstructured":"Pang, Y., Wang, W., Tay, F.E., Liu, W., Tian, Y., & Yuan, L. (2022). Masked autoencoders for point cloud self-supervised learning. In European Conference on Computer Vision, (pp. 604\u2013621).","DOI":"10.1007\/978-3-031-20086-1_35"},{"key":"2879_CR40","unstructured":"Peng, Z., Dong, L., Bao, H., Ye, Q., & Wei, F. (2022). Beit v2: Masked image modeling with vector-quantized visual tokenizers. arXiv preprint arXiv:2208.06366."},{"key":"2879_CR41","doi-asserted-by":"crossref","unstructured":"Poursaeed, O., Jiang, T., Qiao, H., Xu, N., & Kim, V.G. (2020). Self-supervised learning of point clouds via orientation estimation. In International Conference on 3D Vision, (pp. 1018\u20131028).","DOI":"10.1109\/3DV50981.2020.00112"},{"key":"2879_CR42","doi-asserted-by":"crossref","unstructured":"Puy, G., Boulch, A., & Marlet, R. (2023). Using a waffle iron for automotive point cloud semantic segmentation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, (pp. 3379\u20133389).","DOI":"10.1109\/ICCV51070.2023.00313"},{"key":"2879_CR43","doi-asserted-by":"crossref","unstructured":"Puy, G., Gidaris, S., Boulch, A., Sim\u00e9oni, O., Sautier, C., P\u00e9rez, P., Bursuc, A., & Marlet, R. (2024). Three pillars improving vision foundation model distillation for lidar. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, (pp. 21519\u201321529).","DOI":"10.1109\/CVPR52733.2024.02033"},{"key":"2879_CR44","doi-asserted-by":"crossref","unstructured":"Ronneberger, O., Fischer, P., & Brox, T. (2015). U-net: Convolutional networks for biomedical image segmentation. In Medical Image Computing and Computer Assisted Intervention, Springer, (pp. 234\u2013241).","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"2879_CR45","first-page":"12942","volume":"32","author":"J Sauder","year":"2019","unstructured":"Sauder, J., & Sievers, B. (2019). Self-supervised deep learning on point clouds by reconstructing space. Advances in Neural Information Processing Systems, 32, 12942\u201312952.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2879_CR46","doi-asserted-by":"crossref","unstructured":"Sautier, C., Puy, G., Gidaris, S., Boulch, A., Bursuc, A., & Marlet, R. (2022). Image-to-lidar self-supervised distillation for autonomous driving data. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, (pp. 9891\u20139901).","DOI":"10.1109\/CVPR52688.2022.00966"},{"key":"2879_CR47","doi-asserted-by":"crossref","unstructured":"Smith, L.N. (2017). Cyclical learning rates for training neural networks. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, IEEE, (pp. 464\u2013472).","DOI":"10.1109\/WACV.2017.58"},{"key":"2879_CR48","first-page":"403","volume":"114","author":"K Sridharan","year":"2008","unstructured":"Sridharan, K., & Kakade, S. M. (2008). An information theoretic framework for multi-view learning. Annual Conference on Computational Learning Theory, 114, 403\u2013414.","journal-title":"Annual Conference on Computational Learning Theory"},{"key":"2879_CR49","unstructured":"Team, O., et\u00a0al. (2020). Openpcdet: An open-source toolbox for 3d object detection from point clouds."},{"key":"2879_CR50","doi-asserted-by":"publisher","first-page":"34899","DOI":"10.52202\/068431-2529","volume":"35","author":"Z Tian","year":"2022","unstructured":"Tian, Z., Chu, X., Wang, X., Wei, X., & Shen, C. (2022). Fully convolutional one-stage 3d object detection on lidar range images. Advances in Neural Information Processing Systems, 35, 34899\u201334911.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2879_CR51","doi-asserted-by":"crossref","unstructured":"Unal, O., Dai, D., & Gool, L.V. (2022). Scribble-supervised lidar semantic segmentation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, (pp. 2697\u20132707).","DOI":"10.1109\/CVPR52688.2022.00272"},{"key":"2879_CR52","unstructured":"Van Den\u00a0Oord, A., & Vinyals, O., et\u00a0al. (2017). Neural discrete representation learning. In Advances in Neural Information Processing Systems, (pp. 6306\u20136315)."},{"key":"2879_CR53","doi-asserted-by":"publisher","first-page":"63529","DOI":"10.52202\/075280-2774","volume":"36","author":"Y Xia","year":"2023","unstructured":"Xia, Y., Huang, H., Zhu, J., & Zhao, Z. (2023). Achieving cross modal generalization with multimodal unified representation. Advances in Neural Information Processing Systems, 36, 63529\u201363541.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2879_CR54","doi-asserted-by":"crossref","unstructured":"Xiao, A., Huang, J., Guan, D., Zhan, F., & Lu, S. (2022). Transfer learning from synthetic to real lidar point cloud for semantic segmentation. In AAAI Conference on Artificial Intelligence, (pp. 2795\u20132803).","DOI":"10.1609\/aaai.v36i3.20183"},{"key":"2879_CR55","doi-asserted-by":"crossref","unstructured":"Xiao, A., Huang, J., Xuan, W., Ren, R., Liu, K., Guan, D., Saddik, A.E., Lu, S., & Xing, E. (2023). 3d semantic segmentation in the wild: Learning generalized models for adverse-condition point clouds. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, (pp. 9382\u20139392).","DOI":"10.1109\/CVPR52729.2023.00905"},{"key":"2879_CR56","doi-asserted-by":"publisher","first-page":"48444","DOI":"10.52202\/075280-2102","volume":"36","author":"B Xie","year":"2023","unstructured":"Xie, B., Li, S., Guo, Q., Liu, C., & Cheng, X. (2023). Annotator: A generic active learning baseline for lidar semantic segmentation. Advances in Neural Information Processing Systems, 36, 48444\u201348458.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2879_CR57","doi-asserted-by":"crossref","unstructured":"Xie, S., Gu, J., Guo, D., Qi, C.R., Guibas, L., & Litany, O. (2020). Pointcontrast: Unsupervised pre-training for 3d point cloud understanding. In European Conference on Computer Vision, (pp. 574\u2013591).","DOI":"10.1007\/978-3-030-58580-8_34"},{"key":"2879_CR58","unstructured":"Xu, C., Tao, D., Xu, C. (2013). A survey on multi-view learning. arXiv preprint arXiv:1304.5634."},{"key":"2879_CR59","doi-asserted-by":"crossref","unstructured":"Xu, J., Zhang, R., Dou, J., Zhu, Y., Sun, J., & Pu, S. (2021). Rpvnet: A deep and efficient range-point-voxel fusion network for lidar point cloud segmentation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, (pp. 16024\u201316033).","DOI":"10.1109\/ICCV48922.2021.01572"},{"key":"2879_CR60","doi-asserted-by":"crossref","unstructured":"Xu, X., Kong, L., Shuai, H., Zhang, W., Pan, L., Chen, K., Liu, Z., & Liu, Q. (2025). 4d contrastive superflows are dense 3d representation learners. In European Conference on Computer Vision, (pp. 58\u201380).","DOI":"10.1007\/978-3-031-73232-4_4"},{"issue":"10","key":"2879_CR61","doi-asserted-by":"publisher","first-page":"3337","DOI":"10.3390\/s18103337","volume":"18","author":"Y Yan","year":"2018","unstructured":"Yan, Y., Mao, Y., & Li, B. (2018). Second: Sparsely embedded convolutional detection. Sensors, 18(10), 3337.","journal-title":"Sensors"},{"key":"2879_CR62","doi-asserted-by":"crossref","unstructured":"Yin, J., Zhou, D., Zhang, L., Fang, J., Xu, C.Z., Shen, J., & Wang, W. (2022). Proposalcontrast: Unsupervised pre-training for lidar-based 3d object detection. In European Conference on Computer Vision, (pp 17\u201333).","DOI":"10.1007\/978-3-031-19842-7_2"},{"key":"2879_CR63","doi-asserted-by":"crossref","unstructured":"Yin, T., Zhou, X., & Krahenbuhl, P. (2021). Center-based 3d object detection and tracking. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, (pp 11784\u201311793).","DOI":"10.1109\/CVPR46437.2021.01161"},{"issue":"7","key":"2879_CR64","doi-asserted-by":"publisher","first-page":"2585","DOI":"10.1007\/s11263-023-01981-w","volume":"132","author":"S Zhang","year":"2024","unstructured":"Zhang, S., Deng, J., Bai, L., Li, H., Ouyang, W., & Zhang, Y. (2024). Hvdistill: Transferring knowledge from images to point clouds via unsupervised hybrid-view distillation. International Journal of Computer Vision, 132(7), 2585\u20132599.","journal-title":"International Journal of Computer Vision"},{"key":"2879_CR65","doi-asserted-by":"publisher","first-page":"128396","DOI":"10.52202\/079017-4078","volume":"37","author":"Y Zhang","year":"2024","unstructured":"Zhang, Y., & Hou, J. (2024). Fine-grained image-to-lidar contrastive distillation with visual foundation models. Advances in Neural Information Processing Systems, 37, 128396\u2013128429.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2879_CR66","doi-asserted-by":"crossref","unstructured":"Zhang, Z., Girdhar, R., Joulin, A., & Misra, I. (2021). Self-supervised pretraining of 3d features on any point-cloud. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, (pp. 10252\u201310263).","DOI":"10.1109\/ICCV48922.2021.01009"},{"key":"2879_CR67","doi-asserted-by":"crossref","unstructured":"Zhou, Y., & Tuzel, O. (2018). Voxelnet: End-to-end learning for point cloud based 3d object detection. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, (pp. 4490\u20134499).","DOI":"10.1109\/CVPR.2018.00472"},{"key":"2879_CR68","doi-asserted-by":"crossref","unstructured":"Zhou, Z., Zhang, Y., & Foroosh, H. (2021). Panoptic-polarnet: Proposal-free lidar point cloud panoptic segmentation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, (pp. 13194\u201313203).","DOI":"10.1109\/CVPR46437.2021.01299"},{"key":"2879_CR69","doi-asserted-by":"crossref","unstructured":"Zhu, X., Zhou, H., Wang, T., Hong, F., Ma, Y., Li, W., Li, H., & Lin, D. (2021). Cylindrical and asymmetrical 3d convolution networks for lidar segmentation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, (pp. 9939\u20139948).","DOI":"10.1109\/CVPR46437.2021.00981"}],"container-title":["International Journal of Computer Vision"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-026-02879-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11263-026-02879-z","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-026-02879-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,20]],"date-time":"2026-07-20T16:13:44Z","timestamp":1784564024000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11263-026-02879-z"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,15]]},"references-count":69,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2026,6]]}},"alternative-id":["2879"],"URL":"https:\/\/doi.org\/10.1007\/s11263-026-02879-z","relation":{},"ISSN":["0920-5691","1573-1405"],"issn-type":[{"value":"0920-5691","type":"print"},{"value":"1573-1405","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,5,15]]},"assertion":[{"value":"12 December 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"4 May 2026","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 May 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The authors affirm that there are no commercial or associative relationships that could be perceived as a conflict of interest related to the submitted work.","order":1,"name":"Ethics","label":"Conflicts of Interest","group":{"name":"EthicsHeading","label":"Declarations"}}],"article-number":"271"}}