{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,11]],"date-time":"2026-03-11T11:43:35Z","timestamp":1773229415628,"version":"3.50.1"},"reference-count":37,"publisher":"MDPI AG","issue":"12","license":[{"start":{"date-parts":[[2023,6,20]],"date-time":"2023-06-20T00:00:00Z","timestamp":1687219200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Natural Science Foundation of China","award":["62173192"],"award-info":[{"award-number":["62173192"]}]},{"name":"National Natural Science Foundation of China","award":["JCYJ20220530162202005"],"award-info":[{"award-number":["JCYJ20220530162202005"]}]},{"name":"Shenzhen Natural Science Foundation","award":["62173192"],"award-info":[{"award-number":["62173192"]}]},{"name":"Shenzhen Natural Science Foundation","award":["JCYJ20220530162202005"],"award-info":[{"award-number":["JCYJ20220530162202005"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Point cloud registration plays a crucial role in 3D mapping and localization. Urban scene point clouds pose significant challenges for registration due to their large data volume, similar scenarios, and dynamic objects. Estimating the location by instances (bulidings, traffic lights, etc.) in urban scenes is a more humanized matter. In this paper, we propose PCRMLP (point cloud registration MLP), a novel model for urban scene point cloud registration that achieves comparable registration performance to prior learning-based methods. Compared to previous works that focused on extracting features and estimating correspondence, PCRMLP estimates transformation implicitly from concrete instances. The key innovation lies in the instance-level urban scene representation method, which leverages semantic segmentation and density-based spatial clustering of applications with noise (DBSCAN) to generate instance descriptors, enabling robust feature extraction, dynamic object filtering, and logical transformation estimation. Then, a lightweight network consisting of Multilayer Perceptrons (MLPs) is employed to obtain transformation in an encoder\u2013decoder manner. Experimental validation on the KITTI dataset demonstrates that PCRMLP achieves satisfactory coarse transformation estimates from instance descriptors within a remarkable time of 0.0028 s. With the incorporation of an ICP refinement module, our proposed method outperforms prior learning-based approaches, yielding a rotation error of 2.01\u00b0 and a translation error of 1.58 m. The experimental results highlight PCRMLP\u2019s potential for coarse registration of urban scene point clouds, thereby paving the way for its application in instance-level semantic mapping and localization.<\/jats:p>","DOI":"10.3390\/s23125758","type":"journal-article","created":{"date-parts":[[2023,6,21]],"date-time":"2023-06-21T02:30:51Z","timestamp":1687314651000},"page":"5758","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["PCRMLP: A Two-Stage Network for Point Cloud Registration in Urban Scenes"],"prefix":"10.3390","volume":"23","author":[{"given":"Jingyang","family":"Liu","sequence":"first","affiliation":[{"name":"College of Artificial Intelligence, Nankai University, Tianjin 300071, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yucheng","family":"Xu","sequence":"additional","affiliation":[{"name":"School of Informatics, University of Edinburgh, Edinburgh EH8 9YL, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lu","family":"Zhou","sequence":"additional","affiliation":[{"name":"College of Artificial Intelligence, Nankai University, Tianjin 300071, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8886-9012","authenticated-orcid":false,"given":"Lei","family":"Sun","sequence":"additional","affiliation":[{"name":"College of Artificial Intelligence, Nankai University, Tianjin 300071, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,6,20]]},"reference":[{"key":"ref_1","unstructured":"Yan, G., Luo, Z., Liu, Z., and Li, Y. (2023). SensorX2car: Sensors-to-car calibration for autonomous driving in road scenarios. arXiv."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"2074","DOI":"10.1109\/TRO.2022.3150683","article-title":"Lcdnet: Deep loop closure detection and point cloud registration for lidar slam","volume":"38","author":"Cattaneo","year":"2022","journal-title":"IEEE Trans. Robot."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Jiang, B., and Shen, S. (2023). Contour Context: Abstract Structural Distribution for 3D LiDAR Loop Detection and Metric Pose Estimation. arXiv.","DOI":"10.1109\/ICRA48891.2023.10160337"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"1238","DOI":"10.1109\/LRA.2021.3138779","article-title":"Optimal target shape for LiDAR pose estimation","volume":"7","author":"Huang","year":"2021","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Wu, C.Y., Johnson, J., Malik, J., Feichtenhofer, C., and Gkioxari, G. (2023). Multiview Compressive Coding for 3D Reconstruction. arXiv.","DOI":"10.1109\/CVPR52729.2023.00875"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Geiger, A., Ziegler, J., and Stiller, C. (2011, January 5\u20139). Stereoscan: Dense 3d reconstruction in real-time. Proceedings of the 2011 IEEE Intelligent Vehicles Symposium (IV), Baden-Baden, Germany.","DOI":"10.1109\/IVS.2011.5940405"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"99","DOI":"10.1109\/MRA.2006.1678144","article-title":"Simultaneous localization and mapping: Part I","volume":"13","author":"Bailey","year":"2006","journal-title":"IEEE Robot. Autom. Mag."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Shan, T., and Englot, B. (2018, January 1\u20135). Lego-loam: Lightweight and ground-optimized lidar odometry and mapping on variable terrain. Proceedings of the 2018 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Madrid, Spain.","DOI":"10.1109\/IROS.2018.8594299"},{"key":"ref_9","first-page":"586","article-title":"Method for registration of 3-D shapes. In Proceedings of the Sensor fusion IV: Control paradigms and data structures","volume":"1611","author":"Besl","year":"1992","journal-title":"SPIE"},{"key":"ref_10","first-page":"435","article-title":"Generalized","volume":"Volume 2","author":"Segal","year":"2009","journal-title":"Robotics: Science and Systems"},{"key":"ref_11","unstructured":"Biber, P., and Stra\u00dfer, W. (2003, January 27\u201331). The normal distributions transform: A new approach to laser scan matching. Proceedings of the 2003 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS 2003) (Cat. No. 03CH37453), Las Vegas, NV, USA."},{"key":"ref_12","unstructured":"Choy, C., Park, J., and Koltun, V. (November, January 27). Fully convolutional geometric features. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Bai, X., Luo, Z., Zhou, L., Chen, H., Li, L., Hu, Z., Fu, H., and Tai, C.L. (2021, January 20\u201325). Pointdsc: Robust point cloud registration using deep spatial consistency. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01560"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Choy, C., Dong, W., and Koltun, V. (2020, January 13\u201319). Deep global registration. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00259"},{"key":"ref_15","unstructured":"Sarode, V., Li, X., Goforth, H., Aoki, Y., Srivatsan, R.A., Lucey, S., and Choset, H. (2019). Pcrnet: Point cloud registration network using pointnet encoding. arXiv."},{"key":"ref_16","unstructured":"Lu, W., Wan, G., Zhou, Y., Fu, X., Yuan, P., and Song, S. (November, January 27). Deepvcp: An end-to-end deep neural network for point cloud registration. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Ao, S., Hu, Q., Yang, B., Markham, A., and Guo, Y. (2021, January 20\u201325). Spinnet: Learning a general surface descriptor for 3d point cloud registration. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01158"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Huang, S., Gojcic, Z., Usvyatsov, M., Wieser, A., and Schindler, K. (2021, January 20\u201325). Predator: Registration of 3d point clouds with low overlap. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00425"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"DeTone, D., Malisiewicz, T., and Rabinovich, A. (2018, January 18\u201323). Superpoint: Self-supervised interest point detection and description. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPRW.2018.00060"},{"key":"ref_20","unstructured":"Ester, M., Kriegel, H.P., Sander, J., and Xu, X. (1996, January 2\u20134). Density-based spatial clustering of applications with noise. Proceedings of the International Conference on Knowledge Discovery and Data Mining, Portland, OR, USA."},{"key":"ref_21","unstructured":"Qi, C.R., Su, H., Mo, K., and Guibas, L.J. (2017, January 21\u201326). Pointnet: Deep learning on point sets for 3d classification and segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Klokov, R., and Lempitsky, V. (2017, January 22\u201329). Escape from cells: Deep kd-networks for the recognition of 3d point cloud models. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.99"},{"key":"ref_23","unstructured":"Qi, C.R., Yi, L., Su, H., and Guibas, L.J. (2017, January 4\u20139). Pointnet++: Deep hierarchical feature learning on point sets in a metric space. Proceedings of the 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Yan, Y., Mao, Y., and Li, B. (2018). Second: Sparsely embedded convolutional detection. Sensors, 18.","DOI":"10.3390\/s18103337"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Maturana, D., and Scherer, S. (October, January 28). Voxnet: A 3d convolutional neural network for real-time object recognition. Proceedings of the 2015 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Hamburg, Germany.","DOI":"10.1109\/IROS.2015.7353481"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Zhou, Y., and Tuzel, O. (2018, January 18\u201323). VoxelNet: End-to-End Learning for Point Cloud Based 3D Object Detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00472"},{"key":"ref_27","first-page":"965","article-title":"Point-voxel cnn for efficient 3d deep learning","volume":"32","author":"Liu","year":"2019","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_28","first-page":"1","article-title":"Linear least-squares optimization for point-to-plane icp surface registration","volume":"4","author":"Low","year":"2004","journal-title":"Chapel Hill Univ. North Carol."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"381","DOI":"10.1145\/358669.358692","article-title":"Random sample consensus: A paradigm for model fitting with applications to image analysis and automated cartography","volume":"24","author":"Fischler","year":"1981","journal-title":"Commun. ACM"},{"key":"ref_30","unstructured":"Wang, Y., and Solomon, J.M. (November, January 27). Deep closest point: Learning representations for point cloud registration. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_31","unstructured":"Huang, X., Mei, G., Zhang, J., and Abbas, R. (2021). A comprehensive survey on point cloud registration. arXiv."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Zhao, X., Yang, S., Huang, T., Chen, J., Ma, T., Li, M., and Liu, Y. (2022, January 23\u201327). SuperLine3D: Self-supervised Line Segmentation and Description for LiDAR Point Cloud. Proceedings of the Computer Vision\u2013ECCV 2022: 17th European Conference, Tel Aviv, Israel. Proceedings, Part IX.","DOI":"10.1007\/978-3-031-20077-9_16"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Cao, A.Q., Puy, G., Boulch, A., and Marlet, R. (2021, January 10\u201317). PCAM: Product of cross-attention matrices for rigid registration of point clouds. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.01298"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Qin, Z., Yu, H., Wang, C., Guo, Y., Peng, Y., and Xu, K. (2022, January 18\u201324). Geometric transformer for fast and robust point cloud registration. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01086"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Liu, J., Xu, Y., Lin, W., and Sun, L. (2023, January 24\u201326). PointTrans: Rethinking 3D Object Detection from a Translation Perspective with Transformer. Proceedings of the Chinese Control Conference, Tianjin, China. accepted.","DOI":"10.23919\/CCC58697.2023.10240559"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"\u00c7i\u00e7ek, \u00d6., Abdulkadir, A., Lienkamp, S.S., Brox, T., and Ronneberger, O. (2016, January 18\u201322). 3D U-Net: Learning dense volumetric segmentation from sparse annotation. Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Singapore.","DOI":"10.1007\/978-3-319-46723-8_49"},{"key":"ref_37","unstructured":"Behley, J., Garbade, M., Milioto, A., Quenzel, J., Behnke, S., Stachniss, C., and Gall, J. (November, January 27). Semantickitti: A dataset for semantic scene understanding of lidar sequences. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/12\/5758\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T19:57:31Z","timestamp":1760126251000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/12\/5758"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,6,20]]},"references-count":37,"journal-issue":{"issue":"12","published-online":{"date-parts":[[2023,6]]}},"alternative-id":["s23125758"],"URL":"https:\/\/doi.org\/10.3390\/s23125758","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,6,20]]}}}