{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T03:34:11Z","timestamp":1760240051272,"version":"build-2065373602"},"reference-count":52,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2019,2,28]],"date-time":"2019-02-28T00:00:00Z","timestamp":1551312000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>In recent years, estimating the 6D pose of object instances with convolutional neural network (CNN) has received considerable attention. Depending on whether intermediate cues are used, the relevant literature can be roughly divided into two broad categories: direct methods and two-stage pipelines. For the latter, intermediate cues, such as 3D object coordinates, semantic keypoints, or virtual control points instead of pose parameters are regressed by CNN in the first stage. Object pose can then be solved by correspondence constraints constructed with these intermediate cues. In this paper, we focus on the postprocessing of a two-stage pipeline and propose to combine two learning concepts for estimating object pose under challenging scenes: projection grouping on one side, and correspondence learning on the other. We firstly employ a local-patch based method to predict projection heatmaps which denote the confidence distribution of projection of 3D bounding box\u2019s corners. A projection grouping module is then proposed to remove redundant local maxima from each layer of heatmaps. Instead of directly feeding 2D\u20133D correspondences to the perspective-n-point (PnP) algorithm, multiple correspondence hypotheses are sampled from local maxima and its corresponding neighborhood and ranked by a correspondence\u2013evaluation network. Finally, correspondences with higher confidence are selected to determine object pose. Extensive experiments on three public datasets demonstrate that the proposed framework outperforms several state of the art methods.<\/jats:p>","DOI":"10.3390\/s19051032","type":"journal-article","created":{"date-parts":[[2019,3,1]],"date-time":"2019-03-01T03:29:21Z","timestamp":1551410961000},"page":"1032","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["DeepHMap++: Combined Projection Grouping and Correspondence Learning for Full DoF Pose Estimation"],"prefix":"10.3390","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6006-6699","authenticated-orcid":false,"given":"Mingliang","family":"Fu","sequence":"first","affiliation":[{"name":"State Key Laboratory of Robotics, Shenyang Institute of Automation, Chinese Academy of Sciences, Shenyang 110016, China"},{"name":"Institutes for Robotics and Intelligent Manufacturing, Chinese Academy of Sciences, Shenyang 110016, China"},{"name":"University of Chinese Academy of Sciences, Beijing 100049, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Weijia","family":"Zhou","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Robotics, Shenyang Institute of Automation, Chinese Academy of Sciences, Shenyang 110016, China"},{"name":"Institutes for Robotics and Intelligent Manufacturing, Chinese Academy of Sciences, Shenyang 110016, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2019,2,28]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Lepetit, V., and Fua, P. (2005). Monocular Model-Based 3D Tracking of Rigid Objects, Now Publishers Inc.","DOI":"10.1561\/9781933019536"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"91","DOI":"10.1023\/B:VISI.0000029664.99615.94","article-title":"Distinctive image features from scale-invariant keypoints","volume":"60","author":"Lowe","year":"2004","journal-title":"Int. J. Comput. Vis."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"231","DOI":"10.1007\/s11263-005-3674-1","article-title":"3D object modeling and recognition using local affine-invariant image descriptors and multi-view spatial constraints","volume":"66","author":"Rothganger","year":"2006","journal-title":"Int. J. Comput. Vis."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"1465","DOI":"10.1109\/TPAMI.2017.2708711","article-title":"Robust 3D object tracking from monocular images using stable parts","volume":"40","author":"Crivellaro","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"5","DOI":"10.1023\/A:1010933404324","article-title":"Random forests","volume":"45","author":"Breiman","year":"2001","journal-title":"Mach. Learn."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Brachmann, E., Krull, A., Michel, F., Gumhold, S., Shotton, J., and Rother, C. (2014, January 6\u201312). Learning 6D Object Pose Estimation using 3D Object Coordinates. Proceedings of the European Conference on Computer Vision, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10605-2_35"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Rad, M., and Lepetit, V. (2017, January 22\u201329). BB8: A Scalable, Accurate, Robust to Partial Occlusion Method for Predicting the 3D Poses of Challenging Objects without Using Depth. Proceedings of the International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.413"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"3960","DOI":"10.1109\/LRA.2018.2858446","article-title":"Detect Globally, Label Locally: Learning Accurate 6-DOF Object Pose Estimation by Joint Segmentation and Coordinate Regression","volume":"3","author":"Nigam","year":"2018","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Hodan, T., Haluza, P., Obdr\u017e\u00e1lek, \u0160., Matas, J., Lourakis, M., and Zabulis, X. (2017, January 24\u201331). T-LESS: An RGB-D dataset for 6D Pose Estimation of Texture-less Objects. Proceedings of the IEEE Winter Conference on Applications of Computer Vision, Santa Rosa, CA, USA.","DOI":"10.1109\/WACV.2017.103"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Crivellaro, A., Rad, M., Verdie, Y., Moo Yi, K., Fua, P., and Lepetit, V. (2015, January 7\u201313). A novel representation of parts for accurate 3D object detection and tracking in monocular images. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.499"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Oberweger, M., Rad, M., and Lepetit, V. (2018, January 8\u201314). Making Deep Heatmaps Robust to Partial Occlusions for 3D Object Pose Estimation. Proceedings of the European Conference on Computer Vision, Munich, Germany.","DOI":"10.1007\/978-3-030-01267-0_8"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Brachmann, E., and Rother, C. (2018, January 18\u201322). Learning less is more-6d camera localization via 3D surface regression. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00489"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Krull, A., Brachmann, E., Nowozin, S., Michel, F., Shotton, J., and Rother, C. (2017, January 21\u201326). Poseagent: Budget-constrained 6D object pose estimation via reinforcement learning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.275"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Tekin, B., Sinha, S.N., and Fua, P. (2018, January 18\u201322). Real-time seamless single shot 6D object pose prediction. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00038"},{"key":"ref_15","unstructured":"Hartley, R.I., and Zisserman, A. (2000). Multiple View Geometry in Computer Vision, Cambridge University Press. [2nd ed.]."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Brachmann, E., Krull, A., Nowozin, S., Shotton, J., Michel, F., Gumhold, S., and Rother, C. (2017, January 21\u201326). DSAC\u2014Differentiable RANSAC for Camera Localization. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.267"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Yi, K.M., Trulls, E., Ono, Y., Lepetit, V., Salzmann, M., and Fua, P. (2018, January 18\u201322). Learning to Find Good Correspondences. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00282"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Hinterstoisser, S., Lepetit, V., Ilic, S., Holzer, S., Bradski, G., Konolige, K., and Navab, N. (2012, January 5\u20139). Model based training, detection and pose estimation of texture-less 3d objects in heavily cluttered scenes. Proceedings of the Asian Conference on Computer Vision, Daejeon, Korea.","DOI":"10.1007\/978-3-642-33885-4_60"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Xiang, Y., Schmidt, T., Narayanan, V., and Fox, D. (2018, January 26\u201330). PoseCNN: A Convolutional Neural Network for 6D Object Pose Estimation in Cluttered Scenes. Proceedings of the Robotics: Science and Systems, Pittsburgh, PA, USA.","DOI":"10.15607\/RSS.2018.XIV.019"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"119","DOI":"10.1109\/TPAMI.2017.2665623","article-title":"Latent-Class Hough Forests for 6 DoF Object Pose Estimation","volume":"40","author":"Tejani","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Wohlhart, P., and Lepetit, V. (2015, January 7\u201312). Learning descriptors for object recognition and 3D pose estimation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298930"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Balntas, V., Doumanoglou, A., Sahin, C., Sock, J., Kouskouridas, R., and Kim, T.K. (2017, January 22\u201329). Pose Guided RGBD Feature Learning for 3D Object Pose Estimation. Proceedings of the International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.416"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Doumanoglou, A., Kouskouridas, R., Malassiotis, S., and Kim, T.K. (2016, January 27\u201330). Recovering 6D object pose and predicting next-best-view in the crowd. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NE, USA.","DOI":"10.1109\/CVPR.2016.390"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Li, C., Bai, J., and Hager, G.D. (2018, January 8\u201314). A Unified Framework for Multi-View Multi-Class Object Pose Estimation. Proceedings of the European Conference on Computer Vision, Munich, Germany.","DOI":"10.1007\/978-3-030-01270-0_16"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Mitash, C., Boularias, A., and Bekris, K. (arXiv, 2018). Physics-based Scene-level Reasoning for Object Pose Estimation in Clutter, arXiv.","DOI":"10.1177\/0278364919846551"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Wu, J., Zhou, B., Russell, R., Kee, V., Wagner, S., Hebert, M., Torralba, A., and Johnson, D. (2018, January 1\u20135). Real-Time Object Pose Estimation with Pose Interpreter Networks. Proceedings of the IEEE\/RSJ International Conference on Intelligent Robots and Systems, Madrid, Spain.","DOI":"10.1109\/IROS.2018.8593662"},{"key":"ref_27","unstructured":"Do, T.T., Cai, M., Pham, T., and Reid, I. (arXiv, 2018). Deep-6DPose: Recovering 6D Object Pose from a Single RGB Image, arXiv."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Periyasamy, A.S., Schwarz, M., and Behnke, S. (arXiv, 2018). Robust 6D Object Pose Estimation in Cluttered Scenes using Semantic Segmentation and Pose Regression Networks, arXiv.","DOI":"10.1109\/IROS.2018.8594406"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Sundermeyer, M., Marton, Z.C., Durner, M., Brucker, M., and Triebel, R. (2018, January 8\u201314). Implicit 3d orientation learning for 6d object detection from rgb images. Proceedings of the European Conference on Computer Vision, Munich, Germany.","DOI":"10.1007\/978-3-030-01231-1_43"},{"key":"ref_30","unstructured":"Doumanoglou, A., Balntas, V., Kouskouridas, R., and Kim, T.K. (arXiv, 2016). Siamese regression networks with efficient mid-level feature extraction for 3d object pose estimation, arXiv."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Kehl, W., Milletari, F., Tombari, F., Ilic, S., and Navab, N. (2016, January 8\u201316). Deep learning of local RGB-D patches for 3D object detection and 6D pose estimation. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46487-9_13"},{"key":"ref_32","unstructured":"Zhang, H., and Cao, Q. (2017, January 22\u201329). Combined Holistic and Local Patches for Recovering 6D Object Pose. Proceedings of the International Conference on Computer Vision, Venice, Italy."},{"key":"ref_33","unstructured":"Zeng, A., Yu, K.T., Song, S., Suo, D., Walker, E., Rodriguez, A., and Xiao, J. (June, January 29). Multi-view self-supervised deep learning for 6d pose estimation in the amazon picking challenge. Proceedings of the IEEE International Conference on Robotics and Automation, Singapore."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Brachmann, E., Michel, F., Krull, A., Ying Yang, M., and Gumhold, S. (2016, January 27\u201330). Uncertainty-driven 6d pose estimation of objects and scenes from a single rgb image. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.366"},{"key":"ref_35","unstructured":"Jafari, O.H., Mustikovela, S.K., Pertsch, K., Brachmann, E., and Rother, C. (2018, January 4\u20136). iPose: Instance-Aware 6D Pose Estimation of Partly Occluded Objects. Proceedings of the Asian Conference on Computer Vision, Perth, Australia."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Kehl, W., Manhardt, F., Tombari, F., Illic, S., and Navab, N. (2017, January 22\u201329). SSD-6D: Making RGB-based 3D detection and 6D pose estimation great again. Proceedings of the International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.169"},{"key":"ref_37","unstructured":"Pavlakos, G., Zhou, X., Chan, A., Derpanis, K.G., and Daniilidis, K. (June, January 29). 6-dof object pose from semantic keypoints. Proceedings of the IEEE International Conference on Robotics and Automation, Singapore."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Michel, F., Kirillov, A., Brachmann, E., Krull, A., Gumhold, S., Savchynskyy, B., and Rother, C. (2017, January 21\u201326). Global hypothesis generation for 6D object pose estimation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.20"},{"key":"ref_39","unstructured":"Sock, J., Kim, K.I., Sahin, C., and Kim, T.K. (arXiv, 2018). Multi-Task Deep Networks for Depth-Based 6D Object Pose and Joint Registration in Crowd Scenarios, arXiv."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"He, K., Gkioxari, G., Doll\u00e1r, P., and Girshick, R. (2017, January 22\u201329). Mask r-cnn. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.322"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Krull, A., Brachmann, E., Michel, F., Yang, M.Y., Gumhold, S., and Rother, C. (2015, January 7\u201313). Learning Analysis-by-Synthesis for 6D Pose Estimation in RGB-D Images. Proceedings of the 2015 IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.115"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Sahin, C., and Kim, T.K. (2018, January 8\u201314). Recovering 6D Object Pose: A Review and Multi-modal Analysis. Proceedings of the European Conference on Computer Vision Workshops on Assistive Computer Vision and Robotics, Munich, Germany.","DOI":"10.1007\/978-3-030-11024-6_2"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Opitz, M., Waltner, G., Poier, G., Possegger, H., and Bischof, H. (2016, January 11\u201314). Grid loss: Detecting occluded faces. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46487-9_24"},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"1845","DOI":"10.1109\/TPAMI.2017.2738644","article-title":"Faceness-Net: Face Detection through Deep Facial Part Responses","volume":"40","author":"Yang","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_45","unstructured":"Jaderberg, M., Simonyan, K., Zisserman, A., and Kavukcuoglu, K. (2015, January 11\u201312). Spatial transformer networks. Proceedings of the Advances In Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep Residual Learning for Image Recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_47","unstructured":"Nair, V., and Hinton, G.E. (2010, January 21\u201324). Rectified linear units improve restricted boltzmann machines. Proceedings of the 27th International Conference on Machine Learning, Haifa, Israel."},{"key":"ref_48","first-page":"1929","article-title":"Dropout: A simple way to prevent neural networks from overfitting","volume":"15","author":"Srivastava","year":"2014","journal-title":"J. Mach. Learn. Res."},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Dang, Z., Yi, K.M., Hu, Y., Wang, F., Fua, P., and Salzmann, M. (2018, January 8\u201314). Eigendecomposition-free Training of Deep Networks with Zero Eigenvalue-based Losses. Proceedings of the European Conference on Computer Vision, Munich, Germany.","DOI":"10.1007\/978-3-030-01228-1_47"},{"key":"ref_50","unstructured":"Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., and Isard, M. (2016, January 2\u20134). TensorFlow: A System for Large-Scale Machine Learning. Proceedings of the 12th USENIX Symposium on Operating Systems Design and Implementation, Savannah, GA, USA."},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Ferraz, L., Binefa, X., and Moreno-Noguer, F. (2014, January 23\u201328). Very Fast Solution to the PnP Problem with Algebraic Outlier Rejection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.71"},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Rad, M., Oberweger, M., and Lepetit, V. (2018, January 18\u201322). Feature Mapping for Learning Fast and Accurate 3D Pose Inference from Synthetic Images. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00490"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/5\/1032\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T12:35:22Z","timestamp":1760186122000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/5\/1032"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,2,28]]},"references-count":52,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2019,3]]}},"alternative-id":["s19051032"],"URL":"https:\/\/doi.org\/10.3390\/s19051032","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2019,2,28]]}}}