{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,27]],"date-time":"2026-01-27T08:26:32Z","timestamp":1769502392361,"version":"3.49.0"},"reference-count":55,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2021,3,1]],"date-time":"2021-03-01T00:00:00Z","timestamp":1614556800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Nature Foundation","award":["No.62071056."],"award-info":[{"award-number":["No.62071056."]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>This paper focuses on 6Dof object pose estimation from a single RGB image. We tackle this challenging problem with a two-stage optimization framework. More specifically, we first introduce a translation estimation module to provide an initial translation based on an estimated depth map. Then, a pose regression module combines the ROI (Region of Interest) and the original image to predict the rotation and refine the translation. Compared with previous end-to-end methods that directly predict rotations and translations, our method can utilize depth information as weak guidance and significantly reduce the searching space for the subsequent module. Furthermore, we design a new loss function function for symmetric objects, an approach that has handled such exceptionally difficult cases in prior works. Experiments show that our model achieves state-of-the-art object pose estimation for the YCB- video dataset (Yale-CMU-Berkeley).<\/jats:p>","DOI":"10.3390\/s21051692","type":"journal-article","created":{"date-parts":[[2021,3,1]],"date-time":"2021-03-01T10:25:18Z","timestamp":1614594318000},"page":"1692","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":8,"title":["DRNet: A Depth-Based Regression Network for 6D Object Pose Estimation"],"prefix":"10.3390","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4855-2464","authenticated-orcid":false,"given":"Lei","family":"Jin","sequence":"first","affiliation":[{"name":"School of Computer Science, Beijing University of Posts and Telecommunications, Beijing 100876, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaojuan","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Electronic Engineering, Beijing University of Posts and Telecommunications, Beijing 100876, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2896-4595","authenticated-orcid":false,"given":"Mingshu","family":"He","sequence":"additional","affiliation":[{"name":"School of Electronic Engineering, Beijing University of Posts and Telecommunications, Beijing 100876, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jingyue","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Electronic Engineering, Beijing University of Posts and Telecommunications, Beijing 100876, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,3,1]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1542","DOI":"10.1109\/TMM.2016.2568743","article-title":"6-DOF Image Localization from Massive Geo-tagged Reference Images","volume":"18","author":"Song","year":"2016","journal-title":"IEEE Trans. Multimed."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Chen, X., Ma, H., Wan, J., Li, B., and Xia, T. (2017, January 21\u201326). Multi-View 3d Object Detection Network for Autonomous Driving. Proceedings of the CVPR, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.691"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Geiger, A., Lenz, P., and Urtasun, R. (2012, January 16\u201321). Are We Ready for Autonomous Driving? The Kitti Vision Benchmark Suite. Proceedings of the CVPR, Providence, RI, USA.","DOI":"10.1109\/CVPR.2012.6248074"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Xu, D., Anguelov, D., and Jain, A. (2017). Pointfusion: Deep sensor fusion for 3d bounding box estimation. arXiv.","DOI":"10.1109\/CVPR.2018.00033"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"514","DOI":"10.1109\/TMM.2005.846787","article-title":"Live three-dimensional content for augmented reality","volume":"7","author":"Farbiz","year":"2005","journal-title":"IEEE Trans. Multimed."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Marder-Eppstein, E. (2016). Project tango. ACM SIGGRAPH 2016-Real-Time Live, Association for Computing Machinery.","DOI":"10.1145\/2933540.2933550"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"1147","DOI":"10.1109\/TMM.2006.879873","article-title":"Scalable and Efficient Video Coding Using 3-D Modeling","volume":"8","author":"Patrick","year":"2006","journal-title":"IEEE Trans. Multimed."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"1284","DOI":"10.1177\/0278364911401765","article-title":"The moped framework: Object recognition and pose estimation for manipulation","volume":"30","author":"Collet","year":"2011","journal-title":"Int. J. Robot. Res."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Zhu, M., Derpanis, K.G., Yang, Y., Brahmbhatt, S., Zhang, M., Phillips, C., Lecce, M., and Daniilidis, K. (June, January 31). Singleimage 3d object detection and pose estimation for grasping. Proceedings of the 2014 IEEE International Conference on Robotics and Automation (ICRA), Hong Kong, China.","DOI":"10.1109\/ICRA.2014.6907430"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Wang, C., Xu, D., Zhu, Y., Mart\u00edn-Mart\u00edn, R., Lu, C., Fei-Fei, L., and Savarese, S. (2019). DenseFusion: 6D Object Pose Estimation by Iterative Dense Fusion. arXiv.","DOI":"10.1109\/CVPR.2019.00346"},{"key":"ref_11","unstructured":"Cheng, Y., Zhu, H., Acar, C., Jing, W., and Lim, J.H. (2019). 6D Pose Estimation with Correlation Fusion. arXiv."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Brachmann, E., Krull, A., Michel, F., Gumhold, S., Shotton, J., and Rother, C. (2014). Learning 6d object pose estimation using 3d object coordinates. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-10605-2_35"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Hinterstoisser, S., Holzer, S., Cagniart, C., Ilic, S., Konolige, K., Navab, N., and Lepetit, V. (2011, January 6\u201313). Multimodal Templates for Real-Time Detection of Texture-Less Objects in Heavily Cluttered Scenes. Proceedings of the ICCV, Barcelona, Spain.","DOI":"10.1109\/ICCV.2011.6126326"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Hinterstoisser, S., Lepetit, V., Ilic, S., Holzer, S., Bradski, G., Konolige, K., and Navab, N. (2012). Model based training, detection and pose estimation of texture-less 3d objects in heavily cluttered scenes. Asian Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-642-33885-4_60"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Kehl, W., Milletari, F., Tombari, F., Ilic, S., and Navab, N. (2016). Deep learning of local rgb-d patches for 3d object detection and 6d pose estimation. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-46487-9_13"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Rios-Cabrera, R., and Tuytelaars, T. (2013, January 1\u20138). Discriminatively Trained Templates for 3d Object Detection: A Real Time Scalable Approach. Proceedings of the ICCV, Sydney, Australia.","DOI":"10.1109\/ICCV.2013.256"},{"key":"ref_17","unstructured":"Tejani, A., Tang, D., Kouskouridas, R., and Kim, T.K. Latent-class hough forests for 3d object detection and pose estimation. Proceedings of the European Conference on Computer Vision."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"876","DOI":"10.1109\/TPAMI.2011.206","article-title":"Gradient response maps for real-time detection of textureless objects","volume":"34","author":"Hinterstoisser","year":"2012","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"381","DOI":"10.1145\/358669.358692","article-title":"Random sample consensus: A paradigm for model fitting with applications to image analysis and automated cartography","volume":"24","author":"Fischler","year":"1981","journal-title":"Commun. ACM"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Xiang, Y., Schmidt, T., Narayanan, V., and Fox, D. (2017). PoseCNN: A Convolutional Neural Network for 6D Object Pose Estimation in Cluttered Scenes. arXiv.","DOI":"10.15607\/RSS.2018.XIV.019"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"3727","DOI":"10.1109\/LRA.2019.2928776","article-title":"SilhoNet: An RGB Method for 6D Object Pose Estimation","volume":"4","author":"Billings","year":"2019","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Aubry, M., Maturana, D., Efros, A.A., Russell, B.C., and Sivic, J. (2014, January 23\u201328). Seeing 3d Chairs: Exemplar Part-Based 2d-3d Alignment Using a Large Dataset of Cad Models. Proceedings of the CVPR, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.487"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Gu, C., and Ren, X. (2010). Discriminative Mixture-of-Templates for Viewpoint Classification, Springer.","DOI":"10.1007\/978-3-642-15555-0_30"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"850","DOI":"10.1109\/34.232073","article-title":"Comparing images using the hausdorff distance","volume":"15","author":"Huttenlocher","year":"1993","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_25","unstructured":"Suwajanakorn, S., Snavely, N., Tompson, J., and Norouzi, M. (2018). Discovery of Latent 3D Keypoints via End-to-end Geometric Reasoning. arXiv."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Mousavian, A., Anguelov, D., Flynn, J., and Kosecka, J. (2017, January 21\u201326). 3D Bounding Box Estimation Using Deep Learning and Geometry. Proceedings of the CVPR, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.597"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Pavlakos, G., Zhou, X., Chan, A., Derpanis, K.G., and Daniilidis, K. (2017). 6-dof object pose from semantic keypoints. arXiv.","DOI":"10.1109\/ICRA.2017.7989233"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Peng, S., Liu, Y., Huang, Q., Bao, H., and Zhou, X. (2019, January 15\u201321). PVNet: Pixel-Wise Voting Network for 6DoF Pose Estimation. Proceedings of the CVPR, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00469"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Kendall, A., Grimes, M., and Cipolla, R. (2015). Posenet: Aconvolutional Network for Real-Time 6-Dof Camera Relocalization. arXiv.","DOI":"10.1109\/ICCV.2015.336"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Li, Y., Wang, G., Ji, X., Xiang, Y., and Fox, D. (2018). DeepIM: Deep Iterative Matching for 6D Pose Estimation, Springer.","DOI":"10.1007\/978-3-030-01231-1_42"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Flynn, J., Neulander, I., Philbin, J., and Snavely, N. (2016, January 27\u201330). Deepstereo: Learning to Predict New Views from the World\u2019s Imagery. Proceedings of the CVPR, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.595"},{"key":"ref_32","unstructured":"Forsyth, D., and Ponce, J. (2002). Computer Vision: A Modern Approach, Prentice Hall."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"7","DOI":"10.1023\/A:1014573219977","article-title":"A taxonomy and evaluation of dense two-frame stereo correspondence algorithms","volume":"47","author":"Scharstein","year":"2002","journal-title":"Int. J. Comput. Vis."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Guler, R.A., Trigeorgis, G., Antonakos, E., Snape, P., and Kokkinos, I. (2017, January 21\u201326). DenseReg: Fully Convolutional Dense Shape Regression In-the-Wild. Proceedings of the CVPR, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.280"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Laina, I., Rupprecht, C., Belagiannis, V., Tombari, F., and Navab, N. (2016, January 25\u201328). Deeper Depth Prediction with Fully Convolutional Residual Networks. Proceedings of the 3DV, Stanford, CA, USA.","DOI":"10.1109\/3DV.2016.32"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Roy, A., and Todorovic, S. (2016, January 27\u201330). Monocular Depth Estimation Using Neural Regression Forest. Proceedings of the CVPR, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.594"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Eigen, D., and Fergus, R. (2015, January 7\u201313). Predicting Depth, Surface Normals and Semantic Labels with a Common Multi-Scale Convolutional Architecture. Proceedings of the ICCV, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.304"},{"key":"ref_38","unstructured":"Eigen, D., Puhrsch, C., and Fergus, R. (2014, January 8\u201313). Depth Map Prediction from a Single Image Using a Multi-Scale Deep Network. Proceedings of the NIPS, Montreal, QC, Canada."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Xie, J., Girshick, R., and Farhadi, A. (2016). Deep3d: Fully Automatic 2d-to-3d Video Conversion with Deep Convolutional Neural Networks, Springer.","DOI":"10.1007\/978-3-319-46493-0_51"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Fu, H., Gong, M., Wang, C., Batmanghelich, K., and Tao, D. (2018, January 18\u201323). Deep Ordinal Regression Network for Monocular Depth Estimation. Proceedings of the CVPR, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00214"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Garg, R., Carneiro, G., and Reid, I. (2016). Unsupervised Cnn for Single View Depth Estimation: Geometry to the Rescue, Springer.","DOI":"10.1007\/978-3-319-46484-8_45"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Kuznietsov, Y.J.S., and Leibe, B. (2017, January 21\u201326). Semi-Supervised Deep Learning for Monocular Depth Map Prediction. Proceedings of the CVPR, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.238"},{"key":"ref_43","unstructured":"Chen, L.C., Papandreou, G., Kokkinos, I., Murphy, K., and Yuille, A.L. (2016). Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. arXiv."},{"key":"ref_44","unstructured":"Kaiming, H., Xiangyu, Z., Shaoqing, R., and Jian, S. (2016, January 27\u201330). Deep Residual Learning for Image Recognition. Proceedings of the CVPR, Las Vegas, NV, USA."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Xiao, J., Hays, J., Ehinger, K.A., Oliva, A., and Torralba, A. (2010, January 13\u201318). Sun Database: Large-Scale Scene Recognition from Abbey to Zoo. Proceedings of the CVPR, San Francisco, CA, USA.","DOI":"10.1109\/CVPR.2010.5539970"},{"key":"ref_46","unstructured":"Nathan, S., Derek Hoiem, P.K., and Fergus, R. (2012). Indoor Segmentation and Support Inference from RGBD Images, Springer."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Hodan, T., Haluza, P., Obdr\u017e\u00e1lek, \u0160., Matas, J., Lourakis, M., and Zabulis, X. (2017, January 24\u201331). T-LESS: An RGB-D Dataset for 6D Pose Estimation of Texture-Less Objects. Proceedings of the WACV, Santa Rosa, CA, USA.","DOI":"10.1109\/WACV.2017.103"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Tekin, B., Sinha, S.N., and Fua, P. (2018, January 18\u201323). Real-Time Seamless Single Shot 6d Object Pose Prediction. Proceedings of the CVPR, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00038"},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Sundermeyer, M., Marton, Z.C., Durner, M., Brucker, M., and Triebel, R. (2018). Implicit 3d Orientation Learning for 6d Object Detection from Rgb Images, Springer.","DOI":"10.1007\/978-3-030-01231-1_43"},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"239","DOI":"10.1109\/34.121791","article-title":"A method for registration of 3-d shapes","volume":"14","author":"Besl","year":"1992","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_51","unstructured":"ZhiGang, L., and Gu Wang, X.J. (November, January 27). CDPN:Coordinates-Based Disentangled Pose Network for Real-Time RGB-Based 6-DoF Object Pose Estimation. Proceedings of the ICCV, Seoul, Korea."},{"key":"ref_52","unstructured":"Fabian, M., and Diego Martin, A. (November, January 27). Explaining the Ambiguity of Object Detection and 6D Pose From Visual Data. Proceedings of the ICCV, Seoul, Korea."},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Rad, M., and Lepetit, V. (2017, January 22\u201329). Bb8: A Scalable, Accurate, Robust to PARTIAL Occlusion Method for Predicting the 3d Poses of Challenging Objects without Using Depth. Proceedings of the ICCV, Venice, Italy.","DOI":"10.1109\/ICCV.2017.413"},{"key":"ref_54","unstructured":"Kiru, P., and Timothy Patten, M.V. (2019, January 27\u201328). Pix2Pose: Pixel-Wise Coordinate Regression of Objects for 6D Pose Estimation. Proceedings of the ICCV, Seoul, Korea."},{"key":"ref_55","unstructured":"Tomas, H., and Daniel, B. (2020, January 13\u201319). EPOS: Estimating 6D Pose of Objects With Symmetries. Proceedings of the CVPR, Seattle, WA, USA."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/5\/1692\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T05:31:02Z","timestamp":1760160662000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/5\/1692"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,3,1]]},"references-count":55,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2021,3]]}},"alternative-id":["s21051692"],"URL":"https:\/\/doi.org\/10.3390\/s21051692","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,3,1]]}}}