{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,30]],"date-time":"2025-12-30T17:44:55Z","timestamp":1767116695179,"version":"build-2065373602"},"reference-count":59,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2023,3,20]],"date-time":"2023-03-20T00:00:00Z","timestamp":1679270400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100000780","name":"European Union\u2019s Horizon 2020 Research and Innovation program","doi-asserted-by":"publisher","award":["101017151"],"award-info":[{"award-number":["101017151"]}],"id":[{"id":"10.13039\/501100000780","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J. Imaging"],"abstract":"<jats:p>In this work, a visual object detection and localization workflow integrated into a robotic platform is presented for the 6D pose estimation of objects with challenging characteristics in terms of weak texture, surface properties and symmetries. The workflow is used as part of a module for object pose estimation deployed to a mobile robotic platform that exploits the Robot Operating System (ROS) as middleware. The objects of interest aim to support robot grasping in the context of human\u2013robot collaboration during car door assembly in industrial manufacturing environments. In addition to the special object properties, these environments are inherently characterised by cluttered background and unfavorable illumination conditions. For the purpose of this specific application, two different datasets were collected and annotated for training a learning-based method that extracts the object pose from a single frame. The first dataset was acquired in controlled laboratory conditions and the second in the actual indoor industrial environment. Different models were trained based on the individual datasets and a combination of them were further evaluated in a number of test sequences from the actual industrial environment. The qualitative and quantitative results demonstrate the potential of the presented method in relevant industrial applications.<\/jats:p>","DOI":"10.3390\/jimaging9030072","type":"journal-article","created":{"date-parts":[[2023,3,20]],"date-time":"2023-03-20T05:46:42Z","timestamp":1679291202000},"page":"72","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["6D Object Localization in Car-Assembly Industrial Environment"],"prefix":"10.3390","volume":"9","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4553-7077","authenticated-orcid":false,"given":"Alexandra","family":"Papadaki","sequence":"first","affiliation":[{"name":"School of Rural Surveying and Geoinformatics Engineering, National Technical University of Athens, GR-15780 Athens, Greece"},{"name":"Institute of Communication and Computer Systems (ICCS), National Technical University of Athens, GR-15773 Athens, Greece"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8943-4598","authenticated-orcid":false,"given":"Maria","family":"Pateraki","sequence":"additional","affiliation":[{"name":"School of Rural Surveying and Geoinformatics Engineering, National Technical University of Athens, GR-15780 Athens, Greece"},{"name":"Institute of Communication and Computer Systems (ICCS), National Technical University of Athens, GR-15773 Athens, Greece"},{"name":"Institute of Computer Science, Foundation for Research and Technology-Hellas, GR-70013 Heraklion, Greece"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,3,20]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Hoda\u0148, T., Bar\u00e1th, D., and Matas, J. (2020, January 14\u201319). EPOS: Estimating 6D Pose of Objects with Symmetries. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Virtual.","DOI":"10.1109\/CVPR42600.2020.01172"},{"key":"ref_2","unstructured":"Clement, F., Shah, K., and Pancholi, D. (2019). A Review of methods for Textureless Object Recognition. arXiv."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1677","DOI":"10.1007\/s10462-020-09888-5","article-title":"Vision-based robotic grasping from object localization, object pose estimation to grasp estimation for parallel grippers: A review","volume":"54","author":"Du","year":"2021","journal-title":"Artif. Intell. Rev."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Kim, S.H., and Hwang, Y. (2021). A Survey on Deep learning-based Methods and Datasets for Monocular 3D Object Detection. Electronics, 10.","DOI":"10.3390\/electronics10040517"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"He, Z., Feng, W., Zhao, X., and Lv, Y. (2021). 6D Pose Estimation of Objects: Recent Technologies and Challenges. Appl. Sci., 11.","DOI":"10.3390\/app11010228"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Sahin, C., and Kim, T.K. (2018, January 8\u201314). Recovering 6D object pose: A review and multi-modal analysis. Proceedings of the European Conference on Computer Vision (ECCV) Workshops, Munich, Germany.","DOI":"10.1007\/978-3-030-11024-6_2"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"2947","DOI":"10.1109\/TIP.2019.2955239","article-title":"Recent advances in 3D object detection in the era of deep neural networks: A survey","volume":"29","author":"Rahman","year":"2019","journal-title":"IEEE Trans. Image Process."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"012049","DOI":"10.1088\/1742-6596\/1518\/1\/012049","article-title":"A Survey on Monocular 3D Object Detection Algorithms Based on Deep Learning","volume":"1518","author":"Wu","year":"2020","journal-title":"J. Phys. Conf. Ser."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Shi, Y., Huang, J., Xu, X., Zhang, Y., and Xu, K. (2021). StablePose: Learning 6D Object Poses from Geometrically Stable Patches. arXiv.","DOI":"10.1109\/CVPR46437.2021.01497"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Labbe, Y., Carpentier, J., Aubry, M., and Sivic, J. (2020, January 23\u201328). CosyPose: Consistent multi-view multi-object 6D pose estimation. Proceedings of the European Conference on Computer Vision (ECCV), Glasgow, UK.","DOI":"10.1007\/978-3-030-58520-4_34"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Jiang, X., Li, D., Chen, H., Zheng, Y., Zhao, R., and Wu, L. (2022, January 18\u201322). Uni6D: A Unified CNN Framework without Projection Breakdown for 6D Pose Estimation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01089"},{"key":"ref_12","unstructured":"Various authors (2022, December 30). Papers with Code\u20146D Pose Estimation Using RGB. Available online: https:\/\/paperswithcode.com\/task\/6d-pose-estimation."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Hoda\u0148, T., Sundermeyer, M., Drost, B., Labb\u00e9, Y., Brachmann, E., Michel, F., Rother, C., and Matas, J. (2020, January 23\u201328). BOP challenge 2020 on 6D object localization. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-66096-3_39"},{"key":"ref_14","unstructured":"Park, K., Patten, T., and Vincze, M. (November, January 27). Pix2Pose: Pixel-wise coordinate regression of objects for 6D pose estimation. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Li, Y., Wang, G., Ji, X., Xiang, Y., and Fox, D. (2018, January 8\u201314). DeepIM: Deep iterative matching for 6D pose estimation. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01231-1_42"},{"key":"ref_16","unstructured":"Tan, M., and Le, Q. (2019, January 9\u201315). EfficientNet: Rethinking model scaling for convolutional neural networks. Proceedings of the International Conference on Machine Learning, PMLR, Long Beach, CA, USA."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Dosovitskiy, A., Fischer, P., Ilg, E., Hausser, P., Hazirbas, C., Golkov, V., Van Der Smagt, P., Cremers, D., and Brox, T. (2015, January 7\u201313). FlowNet: Learning optical flow with convolutional networks. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.316"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"155","DOI":"10.1007\/s11263-008-0152-6","article-title":"EPnP: An accurate O(n) solution to the PnP problem","volume":"81","author":"Lepetit","year":"2009","journal-title":"Int. J. Comput. Vis."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"381","DOI":"10.1145\/358669.358692","article-title":"Random sample consensus: A paradigm for model fitting with applications to image analysis and automated cartography","volume":"24","author":"Fischler","year":"1981","journal-title":"Commun. ACM"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Jin, M., Li, J., and Zhang, L. (2022). DOPE++: 6D pose estimation algorithm for weakly textured objects based on deep neural networks. PLoS ONE, 17.","DOI":"10.1371\/journal.pone.0269175"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"He, Y., Sun, W., Huang, H., Liu, J., Fan, H., and Sun, J. (2020, January 13\u201319). PVN3D: A Deep Point-wise 3D Keypoints Voting Network for 6DoF Pose Estimation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01165"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"He, Y., Huang, H., Fan, H., Chen, Q., and Sun, J. (2021, January 20\u201325). Ffb6d: A full flow bidirectional fusion network for 6D pose estimation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00302"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"He, Y., Wang, Y., Fan, H., Sun, J., and Chen, Q. (2022, January 18\u201324). FS6D: Few-Shot 6D Pose Estimation of Novel Objects. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00669"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Cao, T., Luo, F., Fu, Y., Zhang, W., Zheng, S., and Xiao, C. (2022, January 18\u201324). DGECN: A Depth-Guided Edge Convolutional Network for End-to-End 6D Pose Estimation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00376"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"He, Z., and Zhang, L. (2019, January 27\u201328). Multi-adversarial faster-RCNN for unrestricted object detection. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea.","DOI":"10.1109\/ICCV.2019.00677"},{"key":"ref_26","unstructured":"Li, F., Yu, H., Shugurov, I., Busam, B., Yang, S., and Ilic, S. (2022). NeRF-Pose: A First-Reconstruct-Then-Regress Approach for Weakly-supervised 6D Object Pose Estimation. arXiv."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Hinterstoisser, S., Lepetit, V., Ilic, S., Holzer, S., Bradski, G., Konolige, K., and Navab, N. (2012, January 5\u20139). Model based training, detection and pose estimation of texture-less 3D objects in heavily cluttered scenes. Proceedings of the Asian Conference on Computer Vision, Daejeon, Republic of Korea.","DOI":"10.1007\/978-3-642-33885-4_60"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Brachmann, E., Krull, A., Michel, F., Gumhold, S., Shotton, J., and Rother, C. (2014, January 6\u201312). Learning 6D object pose estimation using 3D object coordinates. Proceedings of the European Conference on Computer Vision, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10605-2_35"},{"key":"ref_29","unstructured":"Kaskman, R., Zakharov, S., Shugurov, I., and Ilic, S. (November, January 27). Homebreweddb: RGB-D dataset for 6D pose estimation of 3D objects. Proceedings of the IEEE\/CVF International Conference on Computer Vision Workshops, Seoul, Republic of Korea."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"1179","DOI":"10.1109\/LRA.2016.2532924","article-title":"A dataset for improved RGBD-based object detection and pose estimation for warehouse pick-and-place","volume":"1","author":"Rennie","year":"2016","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Tejani, A., Tang, D., Kouskouridas, R., and Kim, T.K. (2014, January 6\u201312). Latent-class Hough forests for 3D object detection and pose estimation. Proceedings of the European Conference on Computer Vision, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10599-4_30"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Xiang, Y., Schmidt, T., Narayanan, V., and Fox, D. (2017). PoseCNN: A convolutional neural network for 6D object pose estimation in cluttered scenes. arXiv.","DOI":"10.15607\/RSS.2018.XIV.019"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Hodan, T., Haluza, P., Obdr\u017e\u00e1lek, \u0160., Matas, J., Lourakis, M., and Zabulis, X. (2017, January 24\u201331). T-LESS: An RGB-D dataset for 6D pose estimation of texture-less objects. Proceedings of the 2017 IEEE Winter Conference on Applications of Computer Vision (WACV), Santa Rosa, CA, USA.","DOI":"10.1109\/WACV.2017.103"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Drost, B., Ulrich, M., Bergmann, P., Hartinger, P., and Steger, C. (2017, January 22\u201329). Introducing MVTec ITODD\u2014A dataset for 3D object recognition in industry. Proceedings of the IEEE International Conference on Computer Vision Workshops, Venice, Italy.","DOI":"10.1109\/ICCVW.2017.257"},{"key":"ref_35","unstructured":"Doumanoglou, A., Kouskouridas, R., Malassiotis, S., and Kim, T.K. (July, January 26). Recovering 6D object pose and predicting next-best-view in the crowd. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA."},{"key":"ref_36","unstructured":"Chang, A.X., Funkhouser, T., Guibas, L., Hanrahan, P., Huang, Q., Li, Z., Savarese, S., Savva, M., Song, S., and Su, H. (2015). Shapenet: An information-rich 3D model repository. arXiv."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Byambaa, M., Koutaki, G., and Choimaa, L. (2022, January 21\u201322). 6D Pose Estimation of Transparent Objects Using Synthetic Data. Proceedings of the International Workshop on Frontiers of Computer Vision, Virtual.","DOI":"10.1007\/978-3-031-06381-7_1"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Hodan, T., Michel, F., Brachmann, E., Kehl, W., GlentBuch, A., Kraft, D., Drost, B., Vidal, J., Ihrke, S., and Zabulis, X. (2018, January 8\u201314). BOP: Benchmark for 6D object pose estimation. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01249-6_2"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Chen, L.C., Zhu, Y., Papandreou, G., Schroff, F., and Adam, H. (2018, January 8\u201314). encoder\u2013decoder with atrous separable convolution for semantic image segmentation. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_49"},{"key":"ref_40","unstructured":"Huber, P.J. (1992). Breakthroughs in Statistics, Springer."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Barath, D., and Matas, J. (2019, January 27\u201328). Progressive-x: Efficient, anytime, multi-model fitting algorithm. Proceedings of the IEEE\/CVF international Conference on Computer Vision, Seoul, Republic of Korea.","DOI":"10.1109\/ICCV.2019.00388"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Barath, D., and Matas, J. (2018, January 18\u201323). Graph-cut RANSAC. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00704"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Kneip, L., Scaramuzza, D., and Siegwart, R. (2011, January 20\u201325). A novel parametrization of the perspective-three-point problem for a direct computation of absolute camera position and orientation. Proceedings of the CVPR 2011, Springs, CO, USA.","DOI":"10.1109\/CVPR.2011.5995464"},{"key":"ref_44","unstructured":"Mor\u00e9, J.J. (1978). Numerical Analysis, Springer."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Hoda\u0148, T., Matas, J., and Obdr\u017e\u00e1lek, \u0160. (2016, January 11\u201314). On evaluation of 6D object pose estimation. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-49409-8_52"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Ronneberger, O., Fischer, P., and Brox, T. (2015, January 5\u20139). U-net: Convolutional networks for biomedical image segmentation. Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Munich, Germany.","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"ref_47","unstructured":"Cignoni, P., Callieri, M., Corsini, M., Dellepiane, M., Ganovelli, F., and Ranzuglia, G. (2008, January 2\u20134). Meshlab: An open-source mesh processing tool. Proceedings of the Eurographics Italian Chapter Conference, Salerno, Italy."},{"key":"ref_48","unstructured":"(2022, July 21). Intel RealSense Depth Camera D455. Available online: https:\/\/www.intelrealsense.com\/depth-camera-d455\/."},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"2280","DOI":"10.1016\/j.patcog.2014.01.005","article-title":"Automatic Generation and Detection of Highly Reliable Fiducial Markers under Occlusion","volume":"47","year":"2014","journal-title":"Pattern Recognit."},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Long, J., Shelhamer, E., and Darrell, T. (2015, January 7\u201312). Fully convolutional networks for semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298965"},{"key":"ref_51","unstructured":"Chen, L.C., Papandreou, G., Schroff, F., and Adam, H. (2017). Rethinking atrous convolution for semantic image segmentation. arXiv."},{"key":"ref_52","doi-asserted-by":"crossref","first-page":"2481","DOI":"10.1109\/TPAMI.2016.2644615","article-title":"Segnet: A deep convolutional encoder\u2013decoder architecture for image segmentation","volume":"39","author":"Badrinarayanan","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_53","unstructured":"Lourakis, M. (2022, July 21). Posest: A C\/C++ Library for Robust 6DoF Pose Estimation from 3D-2D Correspondences. Available online: https:\/\/users.ics.forth.gr\/~lourakis\/posest\/."},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Lourakis, M., and Zabulis, X. (2013, January 16\u201318). Model-based pose estimation for rigid objects. Proceedings of the International Conference on Computer Vision Systems, St. Petersburg, Russia.","DOI":"10.1007\/978-3-642-39402-7_9"},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Chollet, F. (2017, January 21\u201326). Xception: Deep learning with depthwise separable convolutions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.195"},{"key":"ref_56","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., and Fei-Fei, L. (2009, January 20\u201325). Imagenet: A large-scale hierarchical image database. Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_57","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014, January 6\u201312). Microsoft COCO: Common objects in context. Proceedings of the European Conference on Computer Vision, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_58","doi-asserted-by":"crossref","unstructured":"Hinterstoisser, S., Lepetit, V., Wohlhart, P., and Konolige, K. (2018, January 8\u201314). On pre-trained image features and synthetic images for deep learning. Proceedings of the European Conference on Computer Vision (ECCV) Workshops, Munich, Germany.","DOI":"10.1007\/978-3-030-11009-3_42"},{"key":"ref_59","unstructured":"Blume, F. (2022, December 30). 6DPAT. Available online: https:\/\/github.com\/florianblume\/6d-pat."}],"container-title":["Journal of Imaging"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2313-433X\/9\/3\/72\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T18:58:59Z","timestamp":1760122739000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2313-433X\/9\/3\/72"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,3,20]]},"references-count":59,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2023,3]]}},"alternative-id":["jimaging9030072"],"URL":"https:\/\/doi.org\/10.3390\/jimaging9030072","relation":{},"ISSN":["2313-433X"],"issn-type":[{"type":"electronic","value":"2313-433X"}],"subject":[],"published":{"date-parts":[[2023,3,20]]}}}