{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,31]],"date-time":"2026-07-31T15:38:09Z","timestamp":1785512289724,"version":"3.56.0"},"reference-count":82,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2022,11,24]],"date-time":"2022-11-24T00:00:00Z","timestamp":1669248000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,11,24]],"date-time":"2022-11-24T00:00:00Z","timestamp":1669248000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Max Planck Institute for Informatics"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Comput Vis"],"published-print":{"date-parts":[[2023,2]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>3D object detection is receiving increasing attention from both industry and academia thanks to its wide applications in various fields. In this paper, we propose Point-Voxel Region-based Convolution Neural Networks (PV-RCNNs) for 3D object detection on point clouds. First, we propose a novel 3D detector, PV-RCNN, which boosts the 3D detection performance by deeply integrating the feature learning of both point-based set abstraction and voxel-based sparse convolution through two novel steps,<jats:italic>i.e.<\/jats:italic>, the voxel-to-keypoint scene encoding and the keypoint-to-grid RoI feature abstraction. Second, we propose an advanced framework, PV-RCNN++, for more efficient and accurate 3D object detection. It consists of two major improvements: sectorized proposal-centric sampling for efficiently producing more representative keypoints, and VectorPool aggregation for better aggregating local point features with much less resource consumption. With these two strategies, our PV-RCNN++ is about<jats:inline-formula><jats:alternatives><jats:tex-math>$$3\\times $$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><mml:mrow><mml:mn>3<\/mml:mn><mml:mo>\u00d7<\/mml:mo><\/mml:mrow><\/mml:math><\/jats:alternatives><\/jats:inline-formula>faster than PV-RCNN, while also achieving better performance. The experiments demonstrate that our proposed PV-RCNN++ framework achieves state-of-the-art 3D detection performance on the large-scale and highly-competitive Waymo Open Dataset with 10 FPS inference speed on the detection range of<jats:inline-formula><jats:alternatives><jats:tex-math>$$150m \\times 150m$$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><mml:mrow><mml:mn>150<\/mml:mn><mml:mi>m<\/mml:mi><mml:mo>\u00d7<\/mml:mo><mml:mn>150<\/mml:mn><mml:mi>m<\/mml:mi><\/mml:mrow><\/mml:math><\/jats:alternatives><\/jats:inline-formula>.<\/jats:p>","DOI":"10.1007\/s11263-022-01710-9","type":"journal-article","created":{"date-parts":[[2022,11,25]],"date-time":"2022-11-25T10:05:03Z","timestamp":1669370703000},"page":"531-551","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":444,"title":["PV-RCNN++: Point-Voxel Feature Set Abstraction With Local Vector Representation for 3D Object Detection"],"prefix":"10.1007","volume":"131","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2558-181X","authenticated-orcid":false,"given":"Shaoshuai","family":"Shi","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Li","family":"Jiang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jiajun","family":"Deng","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhe","family":"Wang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chaoxu","family":"Guo","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jianping","family":"Shi","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiaogang","family":"Wang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hongsheng","family":"Li","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2022,11,24]]},"reference":[{"key":"1710_CR1","doi-asserted-by":"crossref","unstructured":"Brazil, G., Liu, X., (2019) M3d-rpn: Monocular 3d region proposal network for object detection. In: ICCV.","DOI":"10.1109\/ICCV.2019.00938"},{"key":"1710_CR2","doi-asserted-by":"crossref","unstructured":"Chabot, F., Chaouch, M., Rabarisoa, J., Teuliere, C., Chateau, T. (2017) Deep manta: A coarse-to-fine many-task network for joint 2d and 3d vehicle analysis from monocular image. In: CVPR.","DOI":"10.1109\/CVPR.2017.198"},{"key":"1710_CR3","doi-asserted-by":"crossref","unstructured":"Chen, Q., Sun, L., Wang, Z., Jia, K., Yuille, A. (2019a) Object as hotspots: An anchor-free 3d object detection approach via firing of hotspots.","DOI":"10.1007\/978-3-030-58589-1_5"},{"key":"1710_CR4","doi-asserted-by":"crossref","unstructured":"Chen, X., Kundu, K., Zhang, Z., Ma, H., Fidler, S., Urtasun, R. (2016) Monocular 3d object detection for autonomous driving. In: CVPR.","DOI":"10.1109\/CVPR.2016.236"},{"key":"1710_CR5","doi-asserted-by":"crossref","unstructured":"Chen, X., Ma, H., Wan, J., Li, B., Xia, T. (2017) Multi-view 3d object detection network for autonomous driving. In: CVPR.","DOI":"10.1109\/CVPR.2017.691"},{"key":"1710_CR6","doi-asserted-by":"crossref","unstructured":"Chen, Y., Liu, S., Shen, X., Jia, J. (2019b) Fast point r-cnn. In: ICCV.","DOI":"10.1109\/ICCV.2019.00987"},{"key":"1710_CR7","doi-asserted-by":"crossref","unstructured":"Chen, Y., Liu, S., Shen, X., Jia, J. (2020) Dsgn: Deep stereo geometry network for 3d object detection. In: CVPR.","DOI":"10.1109\/CVPR42600.2020.01255"},{"key":"1710_CR8","doi-asserted-by":"crossref","unstructured":"Choy, C., Gwak, J., Savarese, S. (2019) 4d spatio-temporal convnets: Minkowski convolutional neural networks. In: CVPR.","DOI":"10.1109\/CVPR.2019.00319"},{"key":"1710_CR9","doi-asserted-by":"crossref","unstructured":"Geiger, A., Lenz, P., Urtasun, R. (2012) Are we ready for autonomous driving? the kitti vision benchmark suite. In: CVPR.","DOI":"10.1109\/CVPR.2012.6248074"},{"key":"1710_CR10","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015) Fast r-cnn. In: ICCV.","DOI":"10.1109\/ICCV.2015.169"},{"key":"1710_CR11","doi-asserted-by":"crossref","unstructured":"Graham, B., Engelcke, M., van\u00a0der Maaten, L. (2018) 3d semantic segmentation with submanifold sparse convolutional networks. CVPR.","DOI":"10.1109\/CVPR.2018.00961"},{"key":"1710_CR12","unstructured":"Huang, J., Huang, G. (2022) Bevdet4d: Exploit temporal cues in multi-camera 3d object detection. arXiv preprint arXiv:2203.17054."},{"key":"1710_CR13","unstructured":"Huang, J., Huang, G., Zhu, Z., Du, D. (2021) Bevdet: High-performance multi-camera 3d object detection in bird-eye-view. arXiv preprint arXiv:2112.11790."},{"key":"1710_CR14","doi-asserted-by":"crossref","unstructured":"Huang, Q., Wang, W., Neumann, U. (2018) Recurrent slice networks for 3d segmentation of point clouds. In: CVPR.","DOI":"10.1109\/CVPR.2018.00278"},{"key":"1710_CR15","doi-asserted-by":"crossref","unstructured":"Huang, T., Liu, Z., Chen, X., Bai, X. (2020) Epnet: Enhancing point features with image semantics for 3d object detection. In: ECCV.","DOI":"10.1007\/978-3-030-58555-6_3"},{"key":"1710_CR16","doi-asserted-by":"crossref","unstructured":"Jaritz, M., Gu, J., Su, H. (2019) Multi-view pointnet for 3d scene understanding. In: ICCV Workshops.","DOI":"10.1109\/ICCVW.2019.00494"},{"key":"1710_CR17","doi-asserted-by":"crossref","unstructured":"Jiang, L., Zhao, H., Liu, S., Shen, X., Fu, C. W., Jia, J. (2019) Hierarchical point-edge interaction network for point cloud semantic segmentation. In: ICCV.","DOI":"10.1109\/ICCV.2019.01053"},{"key":"1710_CR18","unstructured":"Jiang, Y., Zhang, L., Miao, Z., Zhu, X., Gao, J., Hu, W., Jiang, Y. G. (2022) Polarformer: Multi-camera 3d object detection with polar transformers. arXiv preprint arXiv:2206.15398."},{"key":"1710_CR19","doi-asserted-by":"crossref","unstructured":"Ku, J., Mozifian, M., Lee, J., Harakeh, A., Waslander, S. (2018) Joint 3d proposal generation and object detection from view aggregation. IROS.","DOI":"10.1109\/IROS.2018.8594049"},{"key":"1710_CR20","doi-asserted-by":"crossref","unstructured":"Kuang, H., Wang, B., An, J., Zhang, M., Zhang, Z. (2020) Voxel-fpn: Multi-scale voxel feature aggregation for 3d object detection from lidar point clouds. Sensors.","DOI":"10.3390\/s20030704"},{"key":"1710_CR21","doi-asserted-by":"crossref","unstructured":"Lang, A. H., Vora, S., Caesar, H., Zhou, L., Yang, J., Beijbom, O. (2019) Pointpillars: Fast encoders for object detection from point clouds. CVPR.","DOI":"10.1109\/CVPR.2019.01298"},{"key":"1710_CR22","doi-asserted-by":"crossref","unstructured":"Li, B., Ouyang, W., Sheng, L., Zeng, X., Wang, X. (2019a) Gs3d: An efficient 3d object detection framework for autonomous driving. In: CVPR.","DOI":"10.1109\/CVPR.2019.00111"},{"key":"1710_CR23","doi-asserted-by":"crossref","unstructured":"Li, P., Chen, X., Shen, S. (2019b) Stereo r-cnn based 3d object detection for autonomous driving. In: CVPR.","DOI":"10.1109\/CVPR.2019.00783"},{"key":"1710_CR24","doi-asserted-by":"crossref","unstructured":"Li, P., Zhao, H., Liu, P., Cao, F. (2020) Rtm3d: Real-time monocular 3d detection from object keypoints for autonomous driving. In: ECCV.","DOI":"10.1007\/978-3-030-58580-8_38"},{"key":"1710_CR25","unstructured":"Li, Y., Bu, R., Sun, M., Wu, W., Di, X., Chen, B. (2018) Pointcnn: Convolution on x-transformed points. In: NeurIPS."},{"key":"1710_CR26","doi-asserted-by":"crossref","unstructured":"Li, Y., Ge, Z., Yu, G., Yang, J., Wang, Z., Shi, Y., Sun, J., Li, Z. (2022a) Bevdepth: Acquisition of reliable depth for multi-view 3d object detection. arXiv preprint arXiv:2206.10092.","DOI":"10.1609\/aaai.v37i2.25233"},{"key":"1710_CR27","doi-asserted-by":"crossref","unstructured":"Li, Z., Wang, F., Wang, N. (2021) Lidar r-cnn: An efficient and universal 3d object detector. In: CVPR.","DOI":"10.1109\/CVPR46437.2021.00746"},{"key":"1710_CR28","doi-asserted-by":"crossref","unstructured":"Li, Z., Wang, W., Li, H., Xie, E., Sima, C., Lu, T., Yu, Q., Dai, J. (2022b) Bevformer: Learning bird\u2019s-eye-view representation from multi-camera images via spatiotemporal transformers. arXiv preprint arXiv:2203.17270.","DOI":"10.1007\/978-3-031-20077-9_1"},{"key":"1710_CR29","doi-asserted-by":"crossref","unstructured":"Liang, M., Yang, B., Wang, S., Urtasun, R. (2018) Deep continuous fusion for multi-sensor 3d object detection. In: ECCV.","DOI":"10.1007\/978-3-030-01270-0_39"},{"key":"1710_CR30","doi-asserted-by":"crossref","unstructured":"Liang, M., Yang, B., Chen, Y., Hu, R., Urtasun, R. (2019) Multi-task multi-sensor fusion for 3d object detection. In: CVPR.","DOI":"10.1109\/CVPR.2019.00752"},{"key":"1710_CR31","doi-asserted-by":"crossref","unstructured":"Lin, T. Y., Doll\u00e1r, P., Girshick, R., He, K., Hariharan, B., Belongie, S. (2017) Feature pyramid networks for object detection. In: CVPR.","DOI":"10.1109\/CVPR.2017.106"},{"key":"1710_CR32","doi-asserted-by":"crossref","unstructured":"Lin, T. Y., Goyal, P., Girshick, R., He, K., Doll\u00e1r, P. (2018) Focal loss for dense object detection. TPAMI.","DOI":"10.1109\/ICCV.2017.324"},{"key":"1710_CR33","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C. Y., Berg, A. C. (2016) Ssd: Single shot multibox detector. In: ECCV.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"1710_CR34","doi-asserted-by":"crossref","unstructured":"Liu, Y., Wang, T., Zhang, X., Sun, J. (2022a) Petr: Position embedding transformation for multi-view 3d object detection. arXiv preprint arXiv:2203.05625.","DOI":"10.1007\/978-3-031-19812-0_31"},{"key":"1710_CR35","doi-asserted-by":"crossref","unstructured":"Liu, Y., Yan, J., Jia, F., Li, S., Gao, Q., Wang, T., Zhang, X., Sun, J. (2022b) Petrv2: A unified framework for 3d perception from multi-camera images. arXiv preprint arXiv:2206.01256.","DOI":"10.1109\/ICCV51070.2023.00302"},{"key":"1710_CR36","unstructured":"Liu, Z., Tang, H., Lin, Y., Han, S. (2019) Point-voxel cnn for efficient 3d deep learning. In: NeurIPS."},{"key":"1710_CR37","doi-asserted-by":"crossref","unstructured":"Liu, Z., Hu, H., Cao, Y., Zhang, Z., Tong, X. (2020) A closer look at local aggregation operators in point cloud analysis. arXiv preprint arXiv:2007.01294.","DOI":"10.1007\/978-3-030-58592-1_20"},{"key":"1710_CR38","doi-asserted-by":"crossref","unstructured":"Manhardt, F., Kehl, W., Gaidon, A. (2019) Roi-10d: Monocular lifting of 2d detection to 6d pose and metric shape. In: CVPR.","DOI":"10.1109\/CVPR.2019.00217"},{"key":"1710_CR39","doi-asserted-by":"crossref","unstructured":"Mao, J., Xue, Y., Niu, M., Bai, H., Feng, J., Liang, X., Xu, H., Xu, C. (2021) Voxel transformer for 3d object detection. In: ICCV.","DOI":"10.1109\/ICCV48922.2021.00315"},{"key":"1710_CR40","doi-asserted-by":"crossref","unstructured":"Mousavian, A., Anguelov, D., Flynn, J., Kosecka, J. (2017) 3d bounding box estimation using deep learning and geometry. In: CVPR.","DOI":"10.1109\/CVPR.2017.597"},{"key":"1710_CR41","doi-asserted-by":"crossref","unstructured":"Murthy, J. K., Krishna, G. S., Chhaya, F, Krishna, K. M. (2017) Reconstructing vehicles from a single image: Shape priors for road scene understanding. In: ICRA.","DOI":"10.1109\/ICRA.2017.7989089"},{"key":"1710_CR42","unstructured":"Ngiam, J., Caine, B., Han, W., Yang, B., Chai, Y., Sun, P., Zhou, Y., Yi, X., Alsharif, O., Nguyen, P., et al. (2019) Starnet: Targeted computation for object detection in point clouds. arXiv preprint arXiv:1908.11069."},{"key":"1710_CR43","doi-asserted-by":"crossref","unstructured":"Philion, J., Fidler, S. (2020) Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d. In: ECCV.","DOI":"10.1007\/978-3-030-58568-6_12"},{"key":"1710_CR44","unstructured":"Qi, C. R., Su, H., Mo, K., Guibas, L. J. (2017a) Pointnet: Deep learning on point sets for 3d classification and segmentation. In: CVPR."},{"key":"1710_CR45","unstructured":"Qi, C. R., Yi, L., Su ,H., Guibas, L. J. (2017b) Pointnet++: Deep hierarchical feature learning on point sets in a metric space. In: NeurIPS."},{"key":"1710_CR46","doi-asserted-by":"crossref","unstructured":"Qi, C. R., Liu, W., Wu, C., Su, H., Guibas, L. J. (2018) Frustum pointnets for 3d object detection from rgb-d data. In: CVPR.","DOI":"10.1109\/CVPR.2018.00102"},{"key":"1710_CR47","doi-asserted-by":"crossref","unstructured":"Qi, C. R., Litany, O., He, K., Guibas, L. J. (2019) Deep hough voting for 3d object detection in point clouds. In: ICCV.","DOI":"10.1109\/ICCV.2019.00937"},{"key":"1710_CR48","doi-asserted-by":"crossref","unstructured":"Qian, R., Garg, D., Wang, Y., You, Y., Belongie, S., Hariharan, B., Campbell, M., Weinberger, K. Q., Chao, W. L. (2020) End-to-end pseudo-lidar for image-based 3d object detection. In: CVPR.","DOI":"10.1109\/CVPR42600.2020.00592"},{"key":"1710_CR49","doi-asserted-by":"crossref","unstructured":"Reading, C., Harakeh, A., Chae, J., Waslander, S. L. (2021) Categorical depth distribution network for monocular 3d object detection. In: CVPR.","DOI":"10.1109\/CVPR46437.2021.00845"},{"key":"1710_CR50","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., Farhadi, A. (2016) You only look once: Unified, real-time object detection. In: CVPR.","DOI":"10.1109\/CVPR.2016.91"},{"key":"1710_CR51","unstructured":"Ren, S., He, K., Girshick, R., Sun, J. (2015) Faster r-cnn: Towards real-time object detection with region proposal networks. In: NeurIPS."},{"key":"1710_CR52","doi-asserted-by":"crossref","unstructured":"Sheng, H., Cai, S., Liu, Y., Deng, B., Huang, J., Hua, X. S., Zhao, M. J. (2021) Improving 3d object detection with channel-wise transformer. In: ICCV.","DOI":"10.1109\/ICCV48922.2021.00274"},{"key":"1710_CR53","doi-asserted-by":"crossref","unstructured":"Shi, S., Wang, X., Li, H. (2019) Pointrcnn: 3d object proposal generation and detection from point cloud. In: CVPR.","DOI":"10.1109\/CVPR.2019.00086"},{"key":"1710_CR54","doi-asserted-by":"crossref","unstructured":"Shi, S., Guo, C., Jiang, L., Wang, Z., Shi, J., Wang, X., Li, H. (2020a) Pv-rcnn: Point-voxel feature set abstraction for 3d object detection. In: CVPR.","DOI":"10.1109\/CVPR42600.2020.01054"},{"key":"1710_CR55","doi-asserted-by":"crossref","unstructured":"Shi, S., Wang, Z., Shi, J., Wang, X., Li, H. (2020b) From points to parts: 3d object detection from point cloud with part-aware and part-aggregation network. TPAMI.","DOI":"10.1109\/TPAMI.2020.2977026"},{"key":"1710_CR56","doi-asserted-by":"crossref","unstructured":"Song, S., Xiao, J. (2016) Deep sliding shapes for amodal 3d object detection in rgb-d images. In: CVPR.","DOI":"10.1109\/CVPR.2016.94"},{"key":"1710_CR57","doi-asserted-by":"crossref","unstructured":"Su, H., Jampani, V., Sun, D., Maji, S., Kalogerakis, E., Yang, M. H., Kautz, J. (2018) Splatnet: Sparse lattice networks for point cloud processing. In: CVPR.","DOI":"10.1109\/CVPR.2018.00268"},{"key":"1710_CR58","doi-asserted-by":"crossref","unstructured":"Sun, P., Kretzschmar, H., Dotiwalla, X., Chouard, A., Patnaik, V., Tsui, P., Guo, J., Zhou, Y., Chai, Y., Caine, B., Vasudevan, V., Han, W., Ngiam, J., Zhao, H., Timofeev, A., Ettinger, S., Krivokon, M., Gao, A., Joshi, A., Zhang, Y., Shlens, J., Chen, Z., Anguelov, D. (2020) Scalability in perception for autonomous driving: Waymo open dataset. In: CVPR.","DOI":"10.1109\/CVPR42600.2020.00252"},{"key":"1710_CR59","doi-asserted-by":"crossref","unstructured":"Sun, P., Wang, W., Chai, Y., Elsayed, G., Bewley, A., Zhang, X., Sminchisescu, C., Anguelov, D. (2021) Rsn: Range sparse net for efficient, accurate lidar 3d object detection. In: CVPR.","DOI":"10.1109\/CVPR46437.2021.00567"},{"key":"1710_CR60","unstructured":"Sun, S., Pang, J., Shi, J., Yi, S., Ouyang, W. (2018) Fishnet: A versatile backbone for image, region, and pixel level prediction. In: NeurIPS."},{"key":"1710_CR61","doi-asserted-by":"crossref","unstructured":"Thomas, H., Qi, C. R., Deschaud, J. E., Marcotegui, B., Goulette, F., Guibas, L. J. (2019) Kpconv: Flexible and deformable convolution for point clouds. In: ICCV.","DOI":"10.1109\/ICCV.2019.00651"},{"key":"1710_CR62","doi-asserted-by":"crossref","unstructured":"Vora, S., Lang, A. H., Helou, B., Beijbom, O. (2020) Pointpainting: Sequential fusion for 3d object detection. In: CVPR.","DOI":"10.1109\/CVPR42600.2020.00466"},{"key":"1710_CR63","doi-asserted-by":"crossref","unstructured":"Wang, Y., Chao, W. L., Garg, D., Hariharan, B., Campbell, M., Weinberger, K. Q. (2019a) Pseudo-lidar from visual depth estimation: Bridging the gap in 3d object detection for autonomous driving. In: CVPR.","DOI":"10.1109\/CVPR.2019.00864"},{"key":"1710_CR64","doi-asserted-by":"crossref","unstructured":"Wang, Y., Sun, Y., Liu, Z., Sarma, S. E., Bronstein, M. M., Solomon, J. M. (2019b) Dynamic graph cnn for learning on point clouds. TOG.","DOI":"10.1145\/3326362"},{"key":"1710_CR65","doi-asserted-by":"crossref","unstructured":"Wang Y, Fathi A, Kundu A, Ross DA, Pantofaru C, Funkhouser T, Solomon J (2020) Pillar-based object detection for autonomous driving. In: ECCV","DOI":"10.1007\/978-3-030-58542-6_2"},{"key":"1710_CR66","unstructured":"Wang, Y., Guizilini, V. C., Zhang, T., Wang, Y., Zhao, H., Solomon, J. (2022) Detr3d: 3d object detection from multi-view images via 3d-to-2d queries. In: CoRL."},{"key":"1710_CR67","doi-asserted-by":"crossref","unstructured":"Wang, Z., Jia, K. (2019) Frustum convnet: Sliding frustums to aggregate local point-wise features for amodal 3d object detection. In: IROS.","DOI":"10.1109\/IROS40897.2019.8968513"},{"key":"1710_CR68","doi-asserted-by":"crossref","unstructured":"Wu, W., Qi, Z., Fuxin, L. (2019) Pointconv: Deep convolutional networks on 3d point clouds. In: CVPR.","DOI":"10.1109\/CVPR.2019.00985"},{"key":"1710_CR69","unstructured":"Xie, E., Yu, Z., Zhou, D., Philion, J., Anandkumar, A., Fidler, S., Luo, P., Alvarez, J. M. (2022) M2bev: Multi-camera joint 3d detection and segmentation with unified birds-eye view representation. arXiv preprint arXiv:2204.05088."},{"key":"1710_CR70","doi-asserted-by":"crossref","unstructured":"Yan, Y., Mao, Y., Li, B. (2018) Second: Sparsely embedded convolutional detection. Sensors.","DOI":"10.3390\/s18103337"},{"key":"1710_CR71","unstructured":"Yang, B., Liang, M., Urtasun, R. (2018a) Hdnet: Exploiting hd maps for 3d object detection. In: CoRL."},{"key":"1710_CR72","doi-asserted-by":"crossref","unstructured":"Yang, B., Luo, W., Urtasun, R. (2018b) Pixor: Real-time 3d object detection from point clouds. In: CVPR.","DOI":"10.1109\/CVPR.2018.00798"},{"key":"1710_CR73","doi-asserted-by":"crossref","unstructured":"Yang. Z., Sun, Y., Liu, S., Shen, X., Jia, J. (2019) STD: sparse-to-dense 3d object detector for point cloud. ICCV.","DOI":"10.1109\/ICCV.2019.00204"},{"key":"1710_CR74","doi-asserted-by":"crossref","unstructured":"Yang, Z., Sun, Y., Liu, S., Jia, J. (2020) 3dssd: Point-based 3d single stage object detector. In: CVPR.","DOI":"10.1109\/CVPR42600.2020.01105"},{"key":"1710_CR75","doi-asserted-by":"crossref","unstructured":"Yang, Z., Zhou, Y., Chen, Z., Ngiam, J. (2021) 3d-man: 3d multi-frame attention network for object detection. In: CVPR.","DOI":"10.1109\/CVPR46437.2021.00190"},{"key":"1710_CR76","doi-asserted-by":"crossref","unstructured":"Ye, M., Xu, S., Cao, T. (2020) Hvnet: Hybrid voxel network for lidar based 3d object detection. In: CVPR.","DOI":"10.1109\/CVPR42600.2020.00170"},{"key":"1710_CR77","doi-asserted-by":"crossref","unstructured":"Yin, T., Zhou, X., Krahenbuhl, P. (2021) Center-based 3d object detection and tracking. In: CVPR.","DOI":"10.1109\/CVPR46437.2021.01161"},{"key":"1710_CR78","doi-asserted-by":"crossref","unstructured":"Yoo, J. H., Kim, Y., Kim, J. S., Choi, J. W. (2020) 3d-cvf: Generating joint camera and lidar features using cross-view spatial feature fusion for 3d object detection. In: ECCV.","DOI":"10.1007\/978-3-030-58583-9_43"},{"key":"1710_CR79","unstructured":"You, Y., Wang, Y., Chao, W. L., Garg, D., Pleiss, G., Hariharan, B., Campbell, M., Weinberger, K. Q. (2020) Pseudo-lidar++: Accurate depth for 3d object detection in autonomous driving. In: ICLR."},{"key":"1710_CR80","doi-asserted-by":"crossref","unstructured":"Zhao, H., Jiang, L., Fu, C. W., Jia, J. (2019) Pointweb: Enhancing local neighborhood features for point cloud processing. In: CVPR.","DOI":"10.1109\/CVPR.2019.00571"},{"key":"1710_CR81","doi-asserted-by":"crossref","unstructured":"Zhou, Y., Tuzel, O. (2018) Voxelnet: End-to-end learning for point cloud based 3d object detection. In: CVPR.","DOI":"10.1109\/CVPR.2018.00472"},{"key":"1710_CR82","unstructured":"Zhou, Y., Sun, P., Zhang, Y., Anguelov, D., Gao, J., Ouyang, T., Guo, J., Ngiam, J., Vasudevan, V. (2020) End-to-end multi-view fusion for 3d object detection in lidar point clouds. In: CoRL."}],"container-title":["International Journal of Computer Vision"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-022-01710-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11263-022-01710-9\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-022-01710-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,10,9]],"date-time":"2024-10-09T14:56:09Z","timestamp":1728485769000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11263-022-01710-9"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,11,24]]},"references-count":82,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2023,2]]}},"alternative-id":["1710"],"URL":"https:\/\/doi.org\/10.1007\/s11263-022-01710-9","relation":{},"ISSN":["0920-5691","1573-1405"],"issn-type":[{"value":"0920-5691","type":"print"},{"value":"1573-1405","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,11,24]]},"assertion":[{"value":"3 April 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"31 October 2022","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"24 November 2022","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}