{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,18]],"date-time":"2026-01-18T10:18:22Z","timestamp":1768731502774,"version":"3.49.0"},"reference-count":83,"publisher":"Association for Computing Machinery (ACM)","issue":"12","license":[{"start":{"date-parts":[[2024,11,20]],"date-time":"2024-11-20T00:00:00Z","timestamp":1732060800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62072020"],"award-info":[{"award-number":["62072020"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Taishan Scholar Program of Shandong Province","award":["No.tstp20221128"],"award-info":[{"award-number":["No.tstp20221128"]}]},{"name":"National Natural Science Fund of Shanxi Province","award":["202103021224192"],"award-info":[{"award-number":["202103021224192"]}]},{"name":"Talented Young Teachers Training Program of Shandong University of Science and Technology","award":["BJ20231201"],"award-info":[{"award-number":["BJ20231201"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2024,12,31]]},"abstract":"<jats:p>Category-level pose estimation is proposed to predict the 6D pose of objects under a specific category and has wide applications in fields such as robotics, virtual reality, and autonomous driving. With the development of VR\/AR technology, pose estimation has gradually become a research hotspot in 3D scene understanding. However, most methods fail to fully utilize geometric and color information to solve intra-class shape variations, which leads to inaccurate prediction results. To solve the above problems, we propose a novel pose estimation and iterative refinement network, use an attention mechanism to fuse multi-modal information to obtain color features after a coordinate transformation, and design iterative modules to ensure the accuracy of object geometric features. Specifically, we use an encoder-decoder architecture to implicitly generate a coarse-grained initial pose and refine it through an iterative refinement module. In addition, due to the differences between rotation and position estimation, we design a multi-head pose decoder that utilizes the local geometry and global features. Finally, we design a transformer-based coordinate transformation attention module to extract pose-sensitive features from RGB images and supervise color information by correlating point cloud features in different coordinate systems. We train and test our network on the synthetic dataset CAMERA25 and the real dataset REAL275. Experimental results show that our method achieves state-of-the-art performance on multiple evaluation metrics.<\/jats:p>","DOI":"10.1145\/3695877","type":"journal-article","created":{"date-parts":[[2024,9,11]],"date-time":"2024-09-11T14:26:56Z","timestamp":1726064816000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Category-Level Pose Estimation and Iterative Refinement for Monocular RGB-D Image"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1010-7229","authenticated-orcid":false,"given":"Yongtang","family":"Bao","sequence":"first","affiliation":[{"name":"Shandong University of Science and Technology, Qingdao, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-3609-7676","authenticated-orcid":false,"given":"Chunjian","family":"Su","sequence":"additional","affiliation":[{"name":"Shandong University of Science and Technology, Qingdao, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-8887-6854","authenticated-orcid":false,"given":"Yutong","family":"Qi","sequence":"additional","affiliation":[{"name":"University of Toronto, Scarborough, ON, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-6971-3276","authenticated-orcid":false,"given":"Yanbing","family":"Geng","sequence":"additional","affiliation":[{"name":"North University of China, Taiyuan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-0334-3681","authenticated-orcid":false,"given":"Haojie","family":"Li","sequence":"additional","affiliation":[{"name":"Shandong University of Science and Technology, Qingdao, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,11,20]]},"reference":[{"key":"e_1_3_1_2_2","first-page":"7163","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Aoki Yasuhiro","year":"2019","unstructured":"Yasuhiro Aoki, Hunter Goforth, Rangaprasad Arun Srivatsan, and Simon Lucey. 2019. Pointnetlk: Robust & efficient point cloud registration using pointnet. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 7163\u20137172."},{"key":"e_1_3_1_3_2","first-page":"113","volume-title":"Computer Graphics Forum","author":"Bouaziz Sofien","year":"2013","unstructured":"Sofien Bouaziz, Andrea Tagliasacchi, and Mark Pauly. 2013. Sparse iterative closest point. In Computer Graphics Forum, Vol. 32, Wiley Online Library, 113\u2013123."},{"issue":"4","key":"e_1_3_1_4_2","doi-asserted-by":"crossref","first-page":"9597","DOI":"10.1109\/LRA.2022.3189792","article-title":"SDFEst: Categorical pose and shape estimation of objects from RGB-D using signed distance fields","volume":"7","author":"Bruns Leonard","year":"2022","unstructured":"Leonard Bruns and Patric Jensfelt. 2022. SDFEst: Categorical pose and shape estimation of objects from RGB-D using signed distance fields. IEEE Robotics and Automation Letters 7, 4 (2022), 9597\u20139604.","journal-title":"IEEE Robotics and Automation Letters"},{"key":"e_1_3_1_5_2","first-page":"3991","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Cao Anh-Quan","year":"2022","unstructured":"Anh-Quan Cao and Raoul de Charette. 2022. Monoscene: Monocular 3d semantic scene completion. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 3991\u20134001."},{"key":"e_1_3_1_6_2","first-page":"9387","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Cao Anh-Quan","year":"2023","unstructured":"Anh-Quan Cao and Raoul de Charette. 2023. Scenerf: Self-supervised monocular 3d scene reconstruction with radiance fields. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 9387\u20139398."},{"issue":"9","key":"e_1_3_1_7_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3557999","article-title":"Mobile augmented reality: User interfaces, frameworks, and intelligence","volume":"55","author":"Cao Jacky","year":"2023","unstructured":"Jacky Cao, Kit-Yung Lam, Lik-Hang Lee, Xiaoli Liu, Pan Hui, and Xiang Su. 2023. Mobile augmented reality: User interfaces, frameworks, and intelligence. ACM Computing Surveys 55, 9 (2023), 1\u201336.","journal-title":"ACM Computing Surveys"},{"key":"e_1_3_1_8_2","first-page":"5746","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Castro Pedro","year":"2023","unstructured":"Pedro Castro and Tae-Kyun Kim. 2023. Crt-6d: Fast 6d object pose estimation with cascaded refinement transformers. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, 5746\u20135755."},{"key":"e_1_3_1_9_2","unstructured":"Angel X. Chang Thomas Funkhouser Leonidas Guibas Pat Hanrahan Qixing Huang Zimo Li Silvio Savarese Manolis Savva Shuran Song Hao Su and Jianxiong Xiao Li Yi Fisher Yu. 2015. Shapenet: An information-rich 3d model repository. arXiv:1512.03012."},{"key":"e_1_3_1_10_2","first-page":"2781","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Chen Hansheng","year":"2022","unstructured":"Hansheng Chen, Pichao Wang, Fan Wang, Wei Tian, Lu Xiong, and Hao Li. 2022. Epro-pnp: Generalized end-to-end probabilistic perspective-n-points for monocular object pose estimation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 2781\u20132790."},{"key":"e_1_3_1_11_2","first-page":"2773","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Chen Kai","year":"2021","unstructured":"Kai Chen and Qi Dou. 2021. Sgpa: Structure-guided prior adaptation for category-level 6d object pose estimation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 2773\u20132782."},{"key":"e_1_3_1_12_2","first-page":"1581","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Chen Wei","year":"2021","unstructured":"Wei Chen, Xi Jia, Hyung Jin Chang, Jinming Duan, Linlin Shen, and Ales Leonardis. 2021. Fs-net: Fast shape-based network for category-level 6d object pose estimation with decoupled rotation mechanism. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 1581\u20131590."},{"key":"e_1_3_1_13_2","first-page":"201\u2013206","volume-title":"Proceedings of the 2023 6th International Conference on Signal Processing and Machine Learning","author":"Chen Wei","year":"2023","unstructured":"Wei Chen, Quanwen Zhao, Jueting Liu, Zehua Wang, Yingchun Liu, and Minda Yao. 2023. Improved YOLO-pose crowd pose estimation. In Proceedings of the 2023 6th International Conference on Signal Processing and Machine Learning. ACM, New York, NY, 201\u2013206."},{"key":"e_1_3_1_14_2","first-page":"8448","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Deng Shengheng","year":"2022","unstructured":"Shengheng Deng, Zhihao Liang, Lin Sun, and Kui Jia. 2022. Vista: Boosting 3d object detection via dual cross-view spatial attention. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 8448\u20138457."},{"issue":"2","key":"e_1_3_1_15_2","doi-asserted-by":"crossref","first-page":"1784","DOI":"10.1109\/LRA.2022.3142441","article-title":"iCaps: Iterative category-level object pose and shape estimation","volume":"7","author":"Deng Xinke","year":"2022","unstructured":"Xinke Deng, Junyi Geng, Timothy Bretl, Yu Xiang, and Dieter Fox. 2022. iCaps: Iterative category-level object pose and shape estimation. IEEE Robotics and Automation Letters 7, 2 (2022), 1784\u20131791.","journal-title":"IEEE Robotics and Automation Letters"},{"key":"e_1_3_1_16_2","first-page":"12396","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Di Yan","year":"2021","unstructured":"Yan Di, Fabian Manhardt, Gu Wang, Xiangyang Ji, Nassir Navab, and Federico Tombari. 2021. So-pose: Exploiting self-occlusion for direct 6d pose estimation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 12396\u201312405."},{"key":"e_1_3_1_17_2","first-page":"6781","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Di Yan","year":"2022","unstructured":"Yan Di, Ruida Zhang, Zhiqiang Lou, Fabian Manhardt, Xiangyang Ji, Nassir Navab, and Federico Tombari. 2022. Gpv-pose: Category-level object pose estimation via geometry-guided point-wise voting. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 6781\u20136791."},{"key":"e_1_3_1_18_2","doi-asserted-by":"crossref","first-page":"267","DOI":"10.1145\/2377576.2377625","volume-title":"Proceedings of the 10th Performance Metrics for Intelligent Systems Workshop","author":"Edwards Shaun M.","year":"2010","unstructured":"Shaun M. Edwards, William C. Flannigan, and Paul T. Evans. 2010. 6-DOF pose estimation: The need for standardization in industrial applications. In Proceedings of the 10th Performance Metrics for Intelligent Systems Workshop. ACM, New York, NY, 267\u2013270."},{"key":"e_1_3_1_19_2","first-page":"52","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Esteves Carlos","year":"2018","unstructured":"Carlos Esteves, Christine Allen-Blanchette, Ameesh Makadia, and Kostas Daniilidis. 2018. Learning so (3) equivariant representations with spherical cnns. In Proceedings of the European Conference on Computer Vision, 52\u201368."},{"key":"e_1_3_1_20_2","first-page":"220","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Fan Zhaoxin","year":"2022","unstructured":"Zhaoxin Fan, Zhenbo Song, Jian Xu, Zhicheng Wang, Kejian Wu, Hongyan Liu, and Jun He. 2022. Object level depth reconstruction for category level 6d object pose estimation from monocular RGB image. In Proceedings of the European Conference on Computer Vision. Springer, 220\u2013236."},{"key":"e_1_3_1_21_2","first-page":"2961","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"He Kaiming","year":"2017","unstructured":"Kaiming He, Georgia Gkioxari, Piotr Doll\u00e1r, and Ross Girshick. 2017. Mask R-CNN. In Proceedings of the IEEE International Conference on Computer Vision, 2961\u20132969."},{"key":"e_1_3_1_22_2","first-page":"770","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"He Kaiming","year":"2016","unstructured":"Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 770\u2013778."},{"key":"e_1_3_1_23_2","first-page":"3003","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"He Yisheng","year":"2021","unstructured":"Yisheng He, Haibin Huang, Haoqiang Fan, Qifeng Chen, and Jian Sun. 2021. Ffb6d: A full flow bidirectional fusion network for 6d pose estimation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 3003\u20133013."},{"key":"e_1_3_1_24_2","first-page":"11632","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"He Yisheng","year":"2020","unstructured":"Yisheng He, Wei Sun, Haibin Huang, Jianran Liu, Haoqiang Fan, and Jian Sun. 2020. Pvn3d: A deep point-wise 3d keypoints voting network for 6dof pose estimation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 11632\u201311641."},{"key":"e_1_3_1_25_2","first-page":"17853","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Hu Yihan","year":"2023","unstructured":"Yihan Hu, Jiazhi Yang, Li Chen, Keyu Li, Chonghao Sima, Xizhou Zhu, Siqi Chai, Senyao Du, Tianwei Lin, Wenhai Wang, Lewei Lu, Xiaosong Jia, Qiang Liu, Jifeng Dai, Yu Qiao, and Hongyang Li. 2023. Planning-oriented autonomous driving. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 17853\u201317862."},{"key":"e_1_3_1_26_2","first-page":"1060","volume-title":"Proceedings of the International Conference on Robot Learning","author":"Jantos Thomas Georg","year":"2023","unstructured":"Thomas Georg Jantos, Mohamed Amin Hamdad, Wolfgang Granig, Stephan Weiss, and Jan Steinbrener. 2023. PoET: Pose estimation transformer for single-view, multi-object 6D pose estimation. In Proceedings of the International Conference on Robot Learning, 1060\u20131070."},{"issue":"2","key":"e_1_3_1_27_2","first-page":"16 pages.","article-title":"GLPose: Global-local representation learning for human pose estimation","volume":"18","author":"Jiao Yingying","year":"2022","unstructured":"Yingying Jiao, Haipeng Chen, Runyang Feng, Haoming Chen, Sifan Wu, Yifang Yin, and Zhenguang Liu. 2022. GLPose: Global-local representation learning for human pose estimation. ACM Transactions on Multimedia Computing, Communications, and Applications 18, 2s (Oct 2022), Article 128, 16 pages.","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"key":"e_1_3_1_28_2","first-page":"574","volume-title":"Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920)","author":"Labb\u00e9 Yann","year":"2020","unstructured":"Yann Labb\u00e9, Justin Carpentier, Mathieu Aubry, and Josef Sivic. 2020. Cosypose: Consistent multi-view multi-object 6d pose estimation. In Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920). Springer, 574\u2013591."},{"key":"e_1_3_1_29_2","unstructured":"Lik-Hang Lee Tristan Braud Pengyuan Zhou Lin Wang Dianlei Xu Zijun Lin Abhishek Kumar Carlos Bermejo and Pan Hui. 2021. All one needs to know about metaverse: A complete survey on technological singularity virtual ecosystem and research agenda. arXiv:2110.05352."},{"issue":"4","key":"e_1_3_1_30_2","doi-asserted-by":"crossref","first-page":"8575","DOI":"10.1109\/LRA.2021.3110538","article-title":"Category-level metric scale object shape and pose estimation","volume":"6","author":"Lee Taeyeop","year":"2021","unstructured":"Taeyeop Lee, Byeong-Uk Lee, Myungchul Kim, and In So Kweon. 2021. Category-level metric scale object shape and pose estimation. IEEE Robotics and Automation Letters 6, 4 (2021), 8575\u20138582.","journal-title":"IEEE Robotics and Automation Letters"},{"key":"e_1_3_1_31_2","first-page":"14891","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Lee Taeyeop","year":"2022","unstructured":"Taeyeop Lee, Byeong-Uk Lee, Inkyu Shin, Jaesung Choe, Ukcheol Shin, In So Kweon, and Kuk-Jin Yoon. 2022. UDA-COPE: Unsupervised domain adaptation for category-level object pose estimation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 14891\u201314900."},{"key":"e_1_3_1_32_2","first-page":"2123","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Li Fu","year":"2023","unstructured":"Fu Li, Shishir Reddy Vutukur, Hao Yu, Ivan Shugurov, Benjamin Busam, Shaowu Yang, and Slobodan Ilic. 2023. Nerf-pose: A first-reconstruct-then-regress approach for weakly-supervised 6d object pose estimation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 2123\u20132133."},{"issue":"3","key":"e_1_3_1_33_2","doi-asserted-by":"crossref","first-page":"3243","DOI":"10.32604\/iasc.2023.035812","article-title":"Dual branch PnP based network for monocular 6D pose estimation","volume":"36","author":"Liang Jia-Yu","year":"2023","unstructured":"Jia-Yu Liang, Hong-Bo Zhang, Qing Lei, Ji-Xiang Du, and Tian-Liang Lin. 2023. Dual branch PnP based network for monocular 6D pose estimation. Intelligent Automation & Soft Computing 36, 3 (2023), 3243\u20133256.","journal-title":"Intelligent Automation & Soft Computing"},{"key":"e_1_3_1_34_2","first-page":"19","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Lin Jiehong","year":"2022","unstructured":"Jiehong Lin, Zewei Wei, Changxing Ding, and Kui Jia. 2022. Category-level 6D object pose and size estimation using self-supervised deep prior deformation networks. In Proceedings of the European Conference on Computer Vision. Springer, 19\u201334."},{"key":"e_1_3_1_35_2","first-page":"3560","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Lin Jiehong","year":"2021","unstructured":"Jiehong Lin, Zewei Wei, Zhihao Li, Songcen Xu, Kui Jia, and Yuanqing Li. 2021. Dualposenet: Category-level 6d object pose and size estimation using dual pose network with refined learning of pose consistency. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 3560\u20133569."},{"key":"e_1_3_1_36_2","first-page":"14001","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Lin Jiehong","year":"2023","unstructured":"Jiehong Lin, Zewei Wei, Yabin Zhang, and Kui Jia. 2023. VI-Net: Boosting category-level 6D object pose estimation via learning decoupled rotations on the spherical representations. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 14001\u201314011."},{"key":"e_1_3_1_37_2","first-page":"1","article-title":"Clipose: Category-level object pose estimation with pre-trained vision-language knowledge","author":"Lin Xiao","year":"2024","unstructured":"Xiao Lin, Minghao Zhu, Ronghao Dang, Guangliang Zhou, Shaolong Shu, Feng Lin, Chengju Liu, and Qijun Chen. 2024. Clipose: Category-level object pose estimation with pre-trained vision-language knowledge. IEEE Transactions on Circuits and Systems for Video Technology (2024), 1\u20131.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_1_38_2","first-page":"1800","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Lin Zhi-Hao","year":"2020","unstructured":"Zhi-Hao Lin, Sheng-Yu Huang, and Yu-Chiang Frank Wang. 2020. Convolution in the cloud: Learning deformable kernels in 3d graph convolution networks for point cloud analysis. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 1800\u20131809."},{"key":"e_1_3_1_39_2","first-page":"1","volume-title":"Proceedings of the IEEE Symposium on Computers and Communications","author":"Liu Chao","year":"2021","unstructured":"Chao Liu, Shuai Yu, Min Yu, Baole Wei, Boquan Li, Gang Li, and Weiqing Huang. 2021. Adaptive smooth L1 loss: A better way to regress scene texts with extreme aspect ratios. In Proceedings of the IEEE Symposium on Computers and Communications. IEEE, 1\u20137."},{"key":"e_1_3_1_40_2","volume-title":"IEEE International Conference on Computer Vision 2023","author":"Liu Jianhui","year":"2023","unstructured":"Jianhui Liu, Yukang Chen, Xiaoqing Ye, and Xiaojuan Qi. 2023. Prior-free category-level pose estimation with implicit space transformation. IEEE International Conference on Computer Vision 2023."},{"key":"e_1_3_1_41_2","first-page":"6469","volume-title":"IEEE Internet of Things Journal","author":"Liu Liangkai","year":"2020","unstructured":"Liangkai Liu, Sidi Lu, Ren Zhong, Baofu Wu, Yongtao Yao, Qingyang Zhang, and Weisong Shi. 2020. Computing systems for autonomous driving: State of the art and challenges. IEEE Internet of Things Journal 8, 8 (2020), 6469\u20136486."},{"key":"e_1_3_1_42_2","first-page":"499","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Liu Xingyu","year":"2022","unstructured":"Xingyu Liu, Gu Wang, Yi Li, and Xiangyang Ji. 2022. Catre: Iterative point clouds alignment for category-level object pose refinement. In Proceedings of the European Conference on Computer Vision. Springer, 499\u2013516."},{"key":"e_1_3_1_43_2","first-page":"800","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Manhardt Fabian","year":"2018","unstructured":"Fabian Manhardt, Wadim Kehl, Nassir Navab, and Federico Tombari. 2018. Deep model-based 6d pose refinement in rgb. In Proceedings of the European Conference on Computer Vision, 800\u2013815."},{"key":"e_1_3_1_44_2","first-page":"2901","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Mousavian Arsalan","year":"2019","unstructured":"Arsalan Mousavian, Clemens Eppner, and Dieter Fox. 2019. 6-dof graspnet: Variational grasp generation for object manipulation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 2901\u20132910."},{"key":"e_1_3_1_45_2","unstructured":"Suraj Nair Aravind Rajeswaran Vikash Kumar Chelsea Finn and Abhinav Gupta. 2022. R3m: A universal visual representation for robot manipulation. arXiv:2203.12601."},{"key":"e_1_3_1_46_2","first-page":"2861","volume-title":"IEEE Robotics and Automation Letters","author":"Pan Liang","year":"2024","unstructured":"Liang Pan, Zhongang Cai, and Ziwei Liu. 2024. Robust partial-to-partial point cloud registration in a full range. IEEE Robotics and Automation Letters (2024), 2861\u20132868."},{"key":"e_1_3_1_47_2","first-page":"652","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Qi Charles R.","year":"2017","unstructured":"Charles R. Qi, Hao Su, Kaichun Mo, and Leonidas J. Guibas. 2017. Pointnet: Deep learning on point sets for 3d classification and segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 652\u2013660."},{"key":"e_1_3_1_48_2","volume-title":"Proceedings of the 31st International Conference on Neural Information Processing Systems","author":"Qi Charles Ruizhongtai","year":"2017","unstructured":"Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J. Guibas. 2017. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. In Proceedings of the 31st International Conference on Neural Information Processing Systems."},{"issue":"3","key":"e_1_3_1_49_2","doi-asserted-by":"crossref","first-page":"1515","DOI":"10.1109\/LRA.2023.3240362","article-title":"i2c-net: Using instance-level neural networks for monocular category-level 6D pose estimation","volume":"8","author":"Remus Alberto","year":"2023","unstructured":"Alberto Remus, Salvatore D\u2019Avella, Francesco Di Felice, Paolo Tripicchio, and Carlo Alberto Avizzano. 2023. i2c-net: Using instance-level neural networks for monocular category-level 6D pose estimation. IEEE Robotics and Automation Letters 8, 3 (2023), 1515\u20131522.","journal-title":"IEEE Robotics and Automation Letters"},{"key":"e_1_3_1_50_2","first-page":"785","volume-title":"Proceedings of the International Conference on Robot Learning","author":"Shridhar Mohit","year":"2023","unstructured":"Mohit Shridhar, Lucas Manuelli, and Dieter Fox. 2023. Perceiver-actor: A multi-task transformer for robotic manipulation. In Proceedings of the International Conference on Robot Learning, 785\u2013799."},{"key":"e_1_3_1_51_2","first-page":"431","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Song Chen","year":"2020","unstructured":"Chen Song, Jiaru Song, and Qixing Huang. 2020. Hybridpose: 6d object pose estimation under hybrid representations. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 431\u2013440."},{"key":"e_1_3_1_52_2","first-page":"222","volume-title":"Proceedings of the IEEE International Symposium on Mixed and Augmented Reality Adjunct","author":"Su Yongzhi","year":"2019","unstructured":"Yongzhi Su, Jason Rambach, Nareg Minaskan, Paul Lesur, Alain Pagani, and Didier Stricker. 2019. Deep multi-state object pose estimation for augmented reality assembly. In Proceedings of the IEEE International Symposium on Mixed and Augmented Reality Adjunct. IEEE, 222\u2013227."},{"key":"e_1_3_1_53_2","unstructured":"Jingtao Sun Yaonan Wang and Danwei Wang. 2024. Towards real-world aerial vision guidance with categorical 6D pose tracker. arXiv:2401.04377."},{"key":"e_1_3_1_54_2","first-page":"530","volume-title":"Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920)","author":"Tian Meng","year":"2020","unstructured":"Meng Tian, Marcelo H. Ang, and Gim Hee Lee. 2020. Shape prior deformation for categorical 6d object pose and size estimation. In Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920). Springer, 530\u2013546."},{"issue":"4","key":"e_1_3_1_55_2","doi-asserted-by":"crossref","first-page":"376","DOI":"10.1109\/34.88573","article-title":"Least-squares estimation of transformation parameters between two point patterns","volume":"13","author":"Umeyama Shinji","year":"1991","unstructured":"Shinji Umeyama. 1991. Least-squares estimation of transformation parameters between two point patterns. IEEE Transactions on Pattern Analysis & Machine Intelligence 13, 4 (1991), 376\u2013380.","journal-title":"IEEE Transactions on Pattern Analysis & Machine Intelligence"},{"key":"e_1_3_1_56_2","volume-title":"Proceedings of the 31st International Conference on Neural Information Processing Systems","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the 31st International Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_57_2","first-page":"10059","volume-title":"Proceedings of the IEEE International Conference on Robotics and Automation","author":"Wang Chen","year":"2020","unstructured":"Chen Wang, Roberto Mart\u00edn-Mart\u00edn, Danfei Xu, Jun Lv, Cewu Lu, Li Fei-Fei, Silvio Savarese, and Yuke Zhu. 2020. 6-pack: Category-level 6d pose tracker with anchor-based keypoints. In Proceedings of the IEEE International Conference on Robotics and Automation. IEEE, 10059\u201310066."},{"key":"e_1_3_1_58_2","first-page":"16611","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wang Gu","year":"2021","unstructured":"Gu Wang, Fabian Manhardt, Federico Tombari, and Xiangyang Ji. 2021. Gdr-net: Geometry-guided direct regression network for monocular 6d object pose estimation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 16611\u201316621."},{"key":"e_1_3_1_59_2","doi-asserted-by":"crossref","first-page":"3676","DOI":"10.1145\/3581783.3612142","volume-title":"Proceedings of the 31st ACM International Conference on Multimedia","author":"Wang Haowen","year":"2023","unstructured":"Haowen Wang, Zhipeng Fan, Zhen Zhao, Zhengping Che, Zhiyuan Xu, Dong Liu, Feifei Feng, Yakun Huang, Xiuquan Qiao, and Jian Tang. 2023. Dtf-net: Category-level pose estimation and shape reconstruction via deformable template field. In Proceedings of the 31st ACM International Conference on Multimedia, 3676\u20133685."},{"key":"e_1_3_1_60_2","first-page":"2642","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wang He","year":"2019","unstructured":"He Wang, Srinath Sridhar, Jingwei Huang, Julien Valentin, Shuran Song, and Leonidas J. Guibas. 2019. Normalized object coordinate space for category-level 6d object pose and size estimation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 2642\u20132651."},{"key":"e_1_3_1_61_2","first-page":"4807","volume-title":"Proceedings of the IEEE\/RSJ International Conference on Intelligent Robots and Systems","author":"Wang Jiaze","year":"2021","unstructured":"Jiaze Wang, Kai Chen, and Qi Dou. 2021a. Category-level 6D object pose estimation via cascaded relation and recurrent reconstruction networks. In Proceedings of the IEEE\/RSJ International Conference on Intelligent Robots and Systems. IEEE, 4807\u20134814."},{"key":"e_1_3_1_62_2","first-page":"1742","volume-title":"Proceedings of the IEEE\/RSJ International Conference on Intelligent Robots and Systems","author":"Wang Zhixin","year":"2019","unstructured":"Zhixin Wang and Kui Jia. 2019. Frustum convnet: Sliding frustums to aggregate local point-wise features for amodal 3d object detection. In Proceedings of the IEEE\/RSJ International Conference on Intelligent Robots and Systems. IEEE, 1742\u20131749."},{"key":"e_1_3_1_63_2","volume-title":"IEEE International Conference on Robotics and Automation (ICRA)","author":"Wei Jiaxin","year":"2023","unstructured":"Jiaxin Wei, Xibin Song, Weizhe Liu, Laurent Kneip, Hongdong Li, and Pan Ji. 2023. RGB-based category-level object pose estimation via decoupled metric scale recovery. IEEE International Conference on Robotics and Automation (ICRA)."},{"key":"e_1_3_1_64_2","first-page":"8067","volume-title":"Proceedings of the IEEE\/RSJ International Conference on Intelligent Robots and Systems","author":"Wen Bowen","year":"2021","unstructured":"Bowen Wen and Kostas Bekris. 2021. Bundletrack: 6d pose tracking for novel objects without instance or category-level 3d models. In Proceedings of the IEEE\/RSJ International Conference on Intelligent Robots and Systems. IEEE, 8067\u20138074."},{"key":"e_1_3_1_65_2","first-page":"10367","volume-title":"Proceedings of the IEEE\/RSJ International Conference on Intelligent Robots and Systems","author":"Wen Bowen","year":"2020","unstructured":"Bowen Wen, Chaitanya Mitash, Baozhang Ren, and Kostas E. Bekris. 2020. se (3)-tracknet: Data-driven 6d pose tracking by calibrating image residuals in synthetic domains. In Proceedings of the IEEE\/RSJ International Conference on Intelligent Robots and Systems. IEEE, 10367\u201310373."},{"key":"e_1_3_1_66_2","first-page":"13209","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Weng Yijia","year":"2021","unstructured":"Yijia Weng, He Wang, Qiang Zhou, Yuzhe Qin, Yueqi Duan, Qingnan Fan, Baoquan Chen, Hao Su, and Leonidas J Guibas. 2021. Captra: Category-level pose tracking for rigid and articulated objects from point clouds. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 13209\u201313218."},{"key":"e_1_3_1_67_2","first-page":"13174","volume-title":"Proceedings of the 34th International Conference on Neural Information Processing Systems","author":"Wu Chaozheng","year":"2020","unstructured":"Chaozheng Wu, Jian Chen, Qiaoyu Cao, Jianchi Zhang, Yunxin Tai, Lin Sun, and Kui Jia. 2020. Grasp proposal networks: An end-to-end solution for visual learning of robotic grasps. In Proceedings of the 34th International Conference on Neural Information Processing Systems, 13174\u201313184."},{"issue":"8","key":"e_1_3_1_68_2","doi-asserted-by":"crossref","first-page":"3585","DOI":"10.1109\/TCSVT.2023.3237328","article-title":"SACF-Net: Skip-attention based correspondence filtering network for point cloud registration","volume":"33","author":"Wu Yue","year":"2023","unstructured":"Yue Wu, Xidao Hu, Yue Zhang, Maoguo Gong, Wenping Ma, and Qiguang Miao. 2023. SACF-Net: Skip-attention based correspondence filtering network for point cloud registration. IEEE Transactions on Circuits and Systems for Video Technology (2023), 33 (8), 3585\u20133595.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_1_69_2","doi-asserted-by":"crossref","unstructured":"Yu Xiang Tanner Schmidt Venkatraman Narayanan and Dieter Fox. 2017. Posecnn: A convolutional neural network for 6d object pose estimation in cluttered scenes. arXiv:1711.00199.","DOI":"10.15607\/RSS.2018.XIV.019"},{"key":"e_1_3_1_70_2","first-page":"21","volume-title":"Proceedings of the 4th International Symposium on Signal Processing Systems","author":"Yen Chia Chen","year":"2022","unstructured":"Chia Chen Yen, Tao Pin, and Hongmin Xu. 2022. Bilateral pose transformer for human pose estimation. In Proceedings of the 4th International Symposium on Signal Processing Systems. ACM, New York, NY, 21\u201329."},{"key":"e_1_3_1_71_2","first-page":"1323","volume-title":"Proceedings of the IEEE\/RSJ International Conference on Intelligent Robots and Systems","author":"Yen-Chen Lin","year":"2021","unstructured":"Lin Yen-Chen, Pete Florence, Jonathan T. Barron, Alberto Rodriguez, Phillip Isola, and Tsung-Yi Lin. 2021. inerf: Inverting neural radiance fields for pose estimation. In Proceedings of the IEEE\/RSJ International Conference on Intelligent Robots and Systems. IEEE, 1323\u20131330."},{"key":"e_1_3_1_72_2","first-page":"3959","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Yi Hongwei","year":"2022","unstructured":"Hongwei Yi, Chun-Hao P. Huang, Dimitrios Tzionas, Muhammed Kocabas, Mohamed Hassan, Siyu Tang, Justus Thies, and Michael J Black. 2022. Human-aware object placement for visual environment reconstruction. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 3959\u20133970."},{"key":"e_1_3_1_73_2","unstructured":"Yuta Yoshitake Mai Nishimura Shohei Nobuhara and Ko Nishino. 2023. TransPoser: Transformer as an optimizer for joint object shape and pose estimation. arXiv:2303.13477."},{"key":"e_1_3_1_74_2","first-page":"27469","volume-title":"Proceedings of the 36th International Conference on Neural Information Processing Systems","author":"Ze Yanjie","year":"2022","unstructured":"Yanjie Ze and Xiaolong Wang. 2022. Category-level 6d object pose estimation in the wild: A semi-supervised learning approach and a new dataset. In Proceedings of the 36th International Conference on Neural Information Processing Systems, 27469\u201327483."},{"key":"e_1_3_1_75_2","first-page":"8833","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhang Cheng","year":"2021","unstructured":"Cheng Zhang, Zhaopeng Cui, Yinda Zhang, Bing Zeng, Marc Pollefeys, and Shuaicheng Liu. 2021. Holistic 3d scene understanding from a single image with implicit representation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 8833\u20138842."},{"key":"e_1_3_1_76_2","first-page":"148","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Zhang Huijie","year":"2022","unstructured":"Huijie Zhang, Anthony Opipari, Xiaotong Chen, Jiyue Zhu, Zeren Yu, and Odest Chadwicke Jenkins. 2022. TransNet: Category-level transparent object pose estimation. In Proceedings of the European Conference on Computer Vision. Springer, 148\u2013164."},{"key":"e_1_3_1_77_2","first-page":"7452","volume-title":"Proceedings of the IEEE\/RSJ International Conference on Intelligent Robots and Systems","author":"Zhang Ruida","year":"2022","unstructured":"Ruida Zhang, Yan Di, Fabian Manhardt, Federico Tombari, and Xiangyang Ji. 2022. SSP-Pose: Symmetry-aware shape prior deformation for direct category-level object pose estimation. In Proceedings of the IEEE\/RSJ International Conference on Intelligent Robots and Systems. IEEE, 7452\u20137459."},{"key":"e_1_3_1_78_2","first-page":"2881","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Zhao Hengshuang","year":"2017","unstructured":"Hengshuang Zhao, Jianping Shi, Xiaojuan Qi, Xiaogang Wang, and Jiaya Jia. 2017. Pyramid scene parsing network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2881\u20132890."},{"key":"e_1_3_1_79_2","first-page":"14045","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Zhao Heng","year":"2023","unstructured":"Heng Zhao, Shenxing Wei, Dahu Shi, Wenming Tan, Zheyang Li, Ye Ren, Xing Wei, Yi Yang, and Shiliang Pu. 2023. Learning symmetry-aware geometry correspondences for 6D object pose estimation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 14045\u201314054."},{"key":"e_1_3_1_80_2","first-page":"17163","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zheng Linfang","year":"2023","unstructured":"Linfang Zheng, Chen Wang, Yinghan Sun, Esha Dasgupta, Hua Chen, Ale\u0161 Leonardis, Wei Zhang, and Hyung Jin Chang. 2023. HS-Pose: Hybrid Scope Feature Extraction for Category-level Object Pose Estimation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 17163\u201317173."},{"key":"e_1_3_1_81_2","first-page":"13967","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Zhou Jun","year":"2023","unstructured":"Jun Zhou, Kai Chen, Linlin Xu, Qi Dou, and Jing Qin. 2023. Deep fusion transformer network with weighted vector-wise keypoints voting for robust 6D object pose estimation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 13967\u201313977."},{"key":"e_1_3_1_82_2","first-page":"5745","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhou Yi","year":"2019","unstructured":"Yi Zhou, Connelly Barnes, Jingwan Lu, Jimei Yang, and Hao Li. 2019. On the continuity of rotation representations in neural networks. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 5745\u20135753."},{"key":"e_1_3_1_83_2","first-page":"1558","volume-title":"IEEE Transactions on Circuits and Systems for Video Technology","author":"Zou Lu","year":"2023","unstructured":"Lu Zou, Zhangjin Huang, Naijie Gu, and Guoping Wang. 2023. Gpt-cope: A graph-guided point transformer for category-level object pose estimation. IEEE Transactions on Circuits and Systems for Video Technology (2023), 1558\u20132205."},{"key":"e_1_3_1_84_2","doi-asserted-by":"crossref","first-page":"109896","DOI":"10.1016\/j.patcog.2023.109896","article-title":"Learning geometric consistency and discrepancy for category-level 6D object pose estimation from point clouds","volume":"145","author":"Zou Lu","year":"2024","unstructured":"Lu Zou, Zhangjin Huang, Naijie Gu, and Guoping Wang. 2024. Learning geometric consistency and discrepancy for category-level 6D object pose estimation from point clouds. Pattern Recognition 145 (2024), Article 109896.","journal-title":"Pattern Recognition"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3695877","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3695877","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:04:29Z","timestamp":1750291469000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3695877"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,11,20]]},"references-count":83,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2024,12,31]]}},"alternative-id":["10.1145\/3695877"],"URL":"https:\/\/doi.org\/10.1145\/3695877","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,11,20]]},"assertion":[{"value":"2024-03-21","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-09-06","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-11-20","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}