{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,15]],"date-time":"2026-06-15T14:38:55Z","timestamp":1781534335136,"version":"3.54.5"},"reference-count":52,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2026,5,29]],"date-time":"2026-05-29T00:00:00Z","timestamp":1780012800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100017142","name":"Gruppo Nazionale per il Calcolo Scientifico","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100017142","id-type":"DOI","asserted-by":"crossref"}]},{"id":[{"id":"https:\/\/ror.org\/02rnys730","id-type":"ROR","asserted-by":"publisher"}]},{"name":"European Union-FSE-REACT-EU"},{"id":[{"id":"https:\/\/ror.org\/019w4f821","id-type":"ROR","asserted-by":"publisher"}]},{"name":"PON Research and Innovation"},{"id":[{"id":"https:\/\/ror.org\/001aqnf71","id-type":"ROR","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J. Imaging"],"abstract":"<jats:p>In recent years, data-driven approaches have become increasingly important in industrial computer vision applications, particularly for 6-Degrees-of-Freedom (6-DoF) object pose estimation. However, benchmark datasets may unintentionally introduce biases that affect the reliability of learned models. In this work, we investigate the shortcut bias induced by fiducial ArUco markers in the widely used Linemod dataset. Although such markers are typically absent in real industrial environments, they introduce unintended visual cues that neural networks tend to exploit. As a result, model selection based on state-of-the-art benchmarks can be biased, since the reported performance often reflects reliance on these shortcuts rather than on robust feature extraction. Using saliency map analysis, we show that often a large portion of the model\u2019s attention is concentrated on these markers, revealing the presence of a shortcut that artificially boosts pose estimation performance. To mitigate this issue, we propose a data augmentation pipeline based on generative AI techniques that removes the markers and replaces the background with more realistic synthesized scenes. Experimental results indicate a noticeable drop in performance when models trained on the original Linemod dataset are evaluated in ArUco-free environments, confirming the presence of background-induced biases. Training with the proposed generative-swapped dataset leads to improved robustness and better generalization to unseen scenarios, although it does not fully eliminate the problem. Overall, the results highlight the impact of background-related biases in pose estimation benchmarks and suggest that the proposed augmentation strategy represents a practical and scalable step toward developing more reliable 6-DoF pose estimation systems for industrial applications, while leaving room for further improvements.<\/jats:p>","DOI":"10.3390\/jimaging12060244","type":"journal-article","created":{"date-parts":[[2026,5,29]],"date-time":"2026-05-29T17:36:18Z","timestamp":1780076178000},"page":"244","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Generative Data Augmentation for ArUco-Free RGB-Based 6-DoF Object Pose Estimation"],"prefix":"10.3390","volume":"12","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1006-7826","authenticated-orcid":false,"given":"Carmelo","family":"Scribano","sequence":"first","affiliation":[{"name":"Department of Physics, Informatics and Mathematics, University of Modena and Reggio Emilia, Via Campi 213\/B, 41125 Modena, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Iacopo","family":"Ferrari","sequence":"additional","affiliation":[{"name":"Department of Physics, Informatics and Mathematics, University of Modena and Reggio Emilia, Via Campi 213\/B, 41125 Modena, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9082-8087","authenticated-orcid":false,"given":"Giorgia","family":"Franchini","sequence":"additional","affiliation":[{"name":"Department of Physics, Informatics and Mathematics, University of Modena and Reggio Emilia, Via Campi 213\/B, 41125 Modena, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4167-3741","authenticated-orcid":false,"given":"Elena","family":"Govi","sequence":"additional","affiliation":[{"name":"HIPERT s.r.l, Via Inventori, 37, 41121 Modena, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6602-7394","authenticated-orcid":false,"given":"Davide","family":"Sapienza","sequence":"additional","affiliation":[{"name":"Department of Physics, Informatics and Mathematics, University of Modena and Reggio Emilia, Via Campi 213\/B, 41125 Modena, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-0573-8209","authenticated-orcid":false,"given":"Tobia","family":"Poppi","sequence":"additional","affiliation":[{"name":"Department of Physics, Informatics and Mathematics, University of Modena and Reggio Emilia, Via Campi 213\/B, 41125 Modena, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3898-8571","authenticated-orcid":false,"given":"Micaela","family":"Verucchi","sequence":"additional","affiliation":[{"name":"HIPERT s.r.l, Via Inventori, 37, 41121 Modena, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2115-4853","authenticated-orcid":false,"given":"Marko","family":"Bertogna","sequence":"additional","affiliation":[{"name":"Department of Physics, Informatics and Mathematics, University of Modena and Reggio Emilia, Via Campi 213\/B, 41125 Modena, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2026,5,29]]},"reference":[{"key":"ref_1","first-page":"1","article-title":"Tracking of industrial objects by using cad models","volume":"4","author":"Wuest","year":"2007","journal-title":"JVRB-J. Virtual Real. Broadcast."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Zhu, M., Derpanis, K.G., Yang, Y., Brahmbhatt, S., Zhang, M., Phillips, C., Lecce, M., and Daniilidis, K. (2014). Single image 3D object detection and pose estimation for grasping. Proceedings of the 2014 IEEE International Conference on Robotics and Automation (ICRA), IEEE.","DOI":"10.1109\/ICRA.2014.6907430"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"7535","DOI":"10.1109\/LRA.2023.3320028","article-title":"Model-Based Underwater 6D Pose Estimation from RGB","volume":"8","author":"Sapienza","year":"2023","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Chen, X., Ma, H., Wan, J., Li, B., and Xia, T. (2017). Multi-view 3d object detection network for autonomous driving. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR.2017.691"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Geiger, A., Lenz, P., and Urtasun, R. (2012). Are we ready for autonomous driving? The kitti vision benchmark suite. Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR.2012.6248074"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"2633","DOI":"10.1109\/TVCG.2015.2513408","article-title":"Pose estimation for augmented reality: A hands-on survey","volume":"22","author":"Marchand","year":"2015","journal-title":"IEEE Trans. Vis. Comput. Graph."},{"key":"ref_7","unstructured":"Shapiro, L.G., and Stockman, G.C. (2001). Computer Vision, Prentice Hall."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Hodan, T., Haluza, P., Obdr\u017e\u00e1lek, \u0160., Matas, J., Lourakis, M., and Zabulis, X. (2017). T-LESS: An RGB-D dataset for 6D pose estimation of texture-less objects. Proceedings of the 2017 IEEE Winter Conference on Applications of Computer Vision (WACV), IEEE.","DOI":"10.1109\/WACV.2017.103"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Hinterstoisser, S., Lepetit, V., Ilic, S., Holzer, S., Bradski, G., Konolige, K., and Navab, N. (2012). Model based training, detection and pose estimation of texture-less 3d objects in heavily cluttered scenes. Proceedings of the Asian Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-642-33885-4_60"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Brachmann, E., Krull, A., Michel, F., Gumhold, S., Shotton, J., and Rother, C. (2014). Learning 6d object pose estimation using 3d object coordinates. Proceedings of the European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-10605-2_35"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"665","DOI":"10.1038\/s42256-020-00257-z","article-title":"Shortcut learning in deep neural networks","volume":"2","author":"Geirhos","year":"2020","journal-title":"Nat. Mach. Intell."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"2280","DOI":"10.1016\/j.patcog.2014.01.005","article-title":"Automatic generation and detection of highly reliable fiducial markers under occlusion","volume":"47","year":"2014","journal-title":"Pattern Recognit."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Kaskman, R., Zakharov, S., Shugurov, I., and Ilic, S. (2019). Homebreweddb: Rgb-d dataset for 6d pose estimation of 3d objects. Proceedings of the IEEE\/CVF International Conference on Computer Vision Workshops, IEEE.","DOI":"10.1109\/ICCVW.2019.00338"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Tyree, S., Tremblay, J., To, T., Cheng, J., Mosier, T., Smith, J., and Birchfield, S. (2022). 6-DoF pose estimation of household objects for robotic manipulation: An accessible dataset and benchmark. arXiv.","DOI":"10.1109\/IROS47612.2022.9981838"},{"key":"ref_15","unstructured":"Kingma, D.P., and Welling, M. (2013). Auto-encoding variational Bayes. arXiv."},{"key":"ref_16","unstructured":"Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014). Generative adversarial nets. Proceedings of the Advances in Neural Information Processing Systems, NeurIPS."},{"key":"ref_17","unstructured":"Antoniou, A., Storkey, A., and Edwards, H. (2017). Data augmentation generative adversarial networks. arXiv."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Shrivastava, A., Pfister, T., Tuzel, O., Susskind, J., Wang, W., and Webb, R. (2017). Learning from simulated and unsupervised images through adversarial training. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR.2017.241"},{"key":"ref_19","unstructured":"Ho, J., Jain, A., and Abbeel, P. (2020). Denoising diffusion probabilistic models. Proceedings of the Advances in Neural Information Processing Systems, NeurIPS."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. (2022). High-resolution image synthesis with latent diffusion models. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Wang, H., Sridhar, S., Huang, J., Valentin, J., Song, S., and Guibas, L.J. (2019). Normalized object coordinate space for category-level 6d object pose and size estimation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR.2019.00275"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Tian, M., Ang, M.H., and Lee, G.H. (2020). Shape prior deformation for categorical 6d object pose and size estimation. Proceedings of the Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK, 23\u201328 August 2020, Springer. Proceedings, Part XXI 16.","DOI":"10.1007\/978-3-030-58589-1_32"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Chen, W., Jia, X., Chang, H.J., Duan, J., Shen, L., and Leonardis, A. (2021). Fs-net: Fast shape-based network for category-level 6d object pose estimation with decoupled rotation mechanism. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR46437.2021.00163"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Wen, B., Yang, W., Kautz, J., and Birchfield, S. (2024). Foundationpose: Unified 6d pose estimation and tracking of novel objects. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR52733.2024.01692"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Lin, J., Liu, L., Lu, D., and Jia, K. (2024). Sam-6d: Segment anything model meets zero-shot 6d object pose estimation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR52733.2024.02636"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Kendall, A., Grimes, M., and Cipolla, R. (2015). Posenet: A convolutional network for real-time 6-dof camera relocalization. Proceedings of the IEEE International Conference on Computer Vision, IEEE.","DOI":"10.1109\/ICCV.2015.336"},{"key":"ref_27","unstructured":"Bukschat, Y., and Vetter, M. (2020). EfficientPose: An efficient, accurate and scalable end-to-end 6D multi object pose estimation approach. arXiv."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Thalhammer, S., Patten, T., and Vincze, M. (2023). Cope: End-to-end trainable constant runtime object pose estimation. Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, IEEE.","DOI":"10.1109\/WACV56688.2023.00288"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Xiang, Y., Schmidt, T., Narayanan, V., and Fox, D. (2017). Posecnn: A convolutional neural network for 6d object pose estimation in cluttered scenes. arXiv.","DOI":"10.15607\/RSS.2018.XIV.019"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Labb\u00e9, Y., Carpentier, J., Aubry, M., and Sivic, J. (2020). Cosypose: Consistent multi-view multi-object 6d pose estimation. Proceedings of the Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK, 23\u201328 August 2020, Springer. Proceedings, Part XVII 16.","DOI":"10.1007\/978-3-030-58520-4_34"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Tekin, B., Sinha, S.N., and Fua, P. (2018). Real-time seamless single shot 6d object pose prediction. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR.2018.00038"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Hu, Y., Hugonot, J., Fua, P., and Salzmann, M. (2019). Segmentation-driven 6d object pose estimation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR.2019.00350"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Park, K., Patten, T., and Vincze, M. (2019). Pix2pose: Pixel-wise coordinate regression of objects for 6d pose estimation. Proceedings of the IEEE\/CVF International Conference on Computer Vision, IEEE.","DOI":"10.1109\/ICCV.2019.00776"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Peng, S., Liu, Y., Huang, Q., Zhou, X., and Bao, H. (2019). Pvnet: Pixel-wise voting network for 6dof pose estimation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR.2019.00469"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Hodan, T., Barath, D., and Matas, J. (2020). Epos: Estimating 6d pose of objects with symmetries. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR42600.2020.01172"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Wang, G., Manhardt, F., Tombari, F., and Ji, X. (2021). GDR-Net: Geometry-Guided Direct Regression Network for Monocular 6D Object Pose Estimation. arXiv.","DOI":"10.1109\/CVPR46437.2021.01634"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Di, Y., Manhardt, F., Wang, G., Ji, X., Navab, N., and Tombari, F. (2021). So-pose: Exploiting self-occlusion for direct 6d pose estimation. Proceedings of the IEEE\/CVF International Conference on Computer Vision, IEEE.","DOI":"10.1109\/ICCV48922.2021.01217"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Su, Y., Saleh, M., Fetzer, T., Rambach, J., Navab, N., Busam, B., Stricker, D., and Tombari, F. (2022). Zebrapose: Coarse to fine surface encoding for 6dof object pose estimation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR52688.2022.00662"},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"876","DOI":"10.1109\/TPAMI.2011.206","article-title":"Gradient response maps for real-time detection of textureless objects","volume":"34","author":"Hinterstoisser","year":"2011","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"850","DOI":"10.1109\/34.232073","article-title":"Comparing images using the Hausdorff distance","volume":"15","author":"Huttenlocher","year":"1993","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Kehl, W., Manhardt, F., Tombari, F., Ilic, S., and Navab, N. (2017). Ssd-6d: Making rgb-based 3d detection and 6d pose estimation great again. Proceedings of the IEEE International Conference on Computer Vision, IEEE.","DOI":"10.1109\/ICCV.2017.169"},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"714","DOI":"10.1007\/s11263-019-01243-8","article-title":"Augmented autoencoders: Implicit 3d orientation learning for 6d object detection","volume":"128","author":"Sundermeyer","year":"2020","journal-title":"Int. J. Comput. Vis."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Tan, M., Pang, R., and Le, Q.V. (2020). Efficientdet: Scalable and efficient object detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR42600.2020.01079"},{"key":"ref_44","unstructured":"Tan, M., and Le, Q. (2019). Efficientnet: Rethinking model scaling for convolutional neural networks. Proceedings of the International Conference on Machine Learning, PMLR."},{"key":"ref_45","unstructured":"Molnar, C. (2026, April 06). A Guide for Making Black Box Models Explainable. Available online: https:\/\/christophm.github.io\/interpretable-ml-book."},{"key":"ref_46","unstructured":"Simonyan, K., Vedaldi, A., and Zisserman, A. (2013). Deep inside convolutional networks: Visualising image classification models and saliency maps. arXiv."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Selvaraju, R.R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., and Batra, D. (2017). Grad-cam: Visual explanations from deep networks via gradient-based localization. Proceedings of the IEEE International Conference on Computer Vision, IEEE.","DOI":"10.1109\/ICCV.2017.74"},{"key":"ref_48","unstructured":"Ozbulak, U. (2026, April 10). PyTorch CNN Visualizations. Available online: https:\/\/github.com\/utkuozbulak\/pytorch-cnn-visualizations."},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"211","DOI":"10.1007\/s11263-015-0816-y","article-title":"Imagenet large scale visual recognition challenge","volume":"115","author":"Russakovsky","year":"2015","journal-title":"Int. J. Comput. Vis."},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Govi, E., Sapienza, D., De Dominicis, L., Capodieci, N., and Bertogna, M. (2026). Advancing Industrial Vision Research with SyGRID: Synthetically Generated Realistic Industrial Dataset. Res. Sq., preprint.","DOI":"10.21203\/rs.3.rs-8017422\/v1"},{"key":"ref_51","unstructured":"Lundberg, S.M., and Lee, S.I. (2017). A Unified Approach to Interpreting Model Predictions, Curran Associates, Inc."},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Ribeiro, M.T., Singh, S., and Guestrin, C. (2016). \u201cWhy should i trust you?\u201d Explaining the predictions of any classifier. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, ACM Digital Library.","DOI":"10.1145\/2939672.2939778"}],"container-title":["Journal of Imaging"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2313-433X\/12\/6\/244\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,13]],"date-time":"2026-06-13T04:11:04Z","timestamp":1781323864000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2313-433X\/12\/6\/244"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,29]]},"references-count":52,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2026,6]]}},"alternative-id":["jimaging12060244"],"URL":"https:\/\/doi.org\/10.3390\/jimaging12060244","relation":{},"ISSN":["2313-433X"],"issn-type":[{"value":"2313-433X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,5,29]]}}}