{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,20]],"date-time":"2026-07-20T17:03:17Z","timestamp":1784566997771,"version":"3.55.0"},"reference-count":72,"publisher":"Springer Science and Business Media LLC","issue":"6","license":[{"start":{"date-parts":[[2026,5,20]],"date-time":"2026-05-20T00:00:00Z","timestamp":1779235200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,5,20]],"date-time":"2026-05-20T00:00:00Z","timestamp":1779235200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100021856","name":"Ministero dell\u2019Universit\u00e0 e della Ricerca","doi-asserted-by":"crossref","award":["Future Artificial Intelligence Research (FAIR) \u2013 PNRR MUR Cod. PE0000013 - CUP: E63C22001940006"],"award-info":[{"award-number":["Future Artificial Intelligence Research (FAIR) \u2013 PNRR MUR Cod. PE0000013 - CUP: E63C22001940006"]}],"id":[{"id":"10.13039\/501100021856","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100021856","name":"Ministero dell\u2019Universit\u00e0 e della Ricerca","doi-asserted-by":"crossref","award":["EXTRA-EYE - PRIN 2022 - CUP E53D23008280006"],"award-info":[{"award-number":["EXTRA-EYE - PRIN 2022 - CUP E53D23008280006"]}],"id":[{"id":"10.13039\/501100021856","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Comput Vis"],"published-print":{"date-parts":[[2026,6]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    In this work, we explore the role of synthetic data in improving the detection of Hand-Object Interactions from egocentric images. Through extensive experimentation and comparative analysis on\n                    <jats:italic>VISOR<\/jats:italic>\n                    ,\n                    <jats:italic>EgoHOS<\/jats:italic>\n                    , and\n                    <jats:italic>ENIGMA-51<\/jats:italic>\n                    datasets, our findings demonstrate the potential of synthetic data to significantly improve HOI detection, particularly when real labeled data are scarce or unavailable. By using synthetic data and only\n                    <jats:inline-formula>\n                      <jats:alternatives>\n                        <jats:tex-math>$$10\\%$$<\/jats:tex-math>\n                        <mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                          <mml:mrow>\n                            <mml:mn>10<\/mml:mn>\n                            <mml:mo>%<\/mml:mo>\n                          <\/mml:mrow>\n                        <\/mml:math>\n                      <\/jats:alternatives>\n                    <\/jats:inline-formula>\n                    of the real labeled data, we achieve improvements in\n                    <jats:italic>Overall AP<\/jats:italic>\n                    over models trained exclusively on real data, with gains of\n                    <jats:inline-formula>\n                      <jats:alternatives>\n                        <jats:tex-math>$$+5.67\\%$$<\/jats:tex-math>\n                        <mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                          <mml:mrow>\n                            <mml:mo>+<\/mml:mo>\n                            <mml:mn>5.67<\/mml:mn>\n                            <mml:mo>%<\/mml:mo>\n                          <\/mml:mrow>\n                        <\/mml:math>\n                      <\/jats:alternatives>\n                    <\/jats:inline-formula>\n                    on\n                    <jats:italic>VISOR<\/jats:italic>\n                    ,\n                    <jats:inline-formula>\n                      <jats:alternatives>\n                        <jats:tex-math>$$+8.24\\%$$<\/jats:tex-math>\n                        <mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                          <mml:mrow>\n                            <mml:mo>+<\/mml:mo>\n                            <mml:mn>8.24<\/mml:mn>\n                            <mml:mo>%<\/mml:mo>\n                          <\/mml:mrow>\n                        <\/mml:math>\n                      <\/jats:alternatives>\n                    <\/jats:inline-formula>\n                    on\n                    <jats:italic>EgoHOS<\/jats:italic>\n                    , and\n                    <jats:inline-formula>\n                      <jats:alternatives>\n                        <jats:tex-math>$$+11.69\\%$$<\/jats:tex-math>\n                        <mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                          <mml:mrow>\n                            <mml:mo>+<\/mml:mo>\n                            <mml:mn>11.69<\/mml:mn>\n                            <mml:mo>%<\/mml:mo>\n                          <\/mml:mrow>\n                        <\/mml:math>\n                      <\/jats:alternatives>\n                    <\/jats:inline-formula>\n                    on\n                    <jats:italic>ENIGMA-51<\/jats:italic>\n                    . Furthermore, we systematically study how aligning synthetic data to specific real-world benchmarks with respect to objects, grasps, and environments, showing that the effectiveness of synthetic data consistently improves with better synthetic-real alignment. As a result of this work, we release a new data generation pipeline and the new\n                    <jats:italic>HOI-Synth<\/jats:italic>\n                    benchmark, which augments existing datasets with synthetic images of hand-object interaction. These data are automatically annotated with hand-object contact states, bounding boxes, and pixel-wise segmentation masks. All data, code, and tools for synthetic data generation are available at:\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/fpv-iplab.github.io\/HOI-Synth\/\" ext-link-type=\"uri\">https:\/\/fpv-iplab.github.io\/HOI-Synth\/<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1007\/s11263-026-02838-8","type":"journal-article","created":{"date-parts":[[2026,5,20]],"date-time":"2026-05-20T14:12:27Z","timestamp":1779286347000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Leveraging Synthetic Data for Enhancing Egocentric Hand-Object Interaction Detection"],"prefix":"10.1007","volume":"134","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-8693-3826","authenticated-orcid":false,"given":"Rosario","family":"Leonardi","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Antonino","family":"Furnari","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Francesco","family":"Ragusa","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Giovanni Maria","family":"Farinella","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,5,20]]},"reference":[{"key":"2838_CR1","doi-asserted-by":"crossref","unstructured":"Avetisyan, A., Xie, C., Howard-Jenkins, H., Yang, T.-Y., Aroudj, S., Patra, S., Zhang, F., Frost, D., Holland, L., Orme, C., et al. (2024). Scenescript: Reconstructing scenes with an autoregressive structured language model. arXiv preprint arXiv:2403.13064","DOI":"10.1007\/978-3-031-73030-6_14"},{"key":"2838_CR2","doi-asserted-by":"crossref","unstructured":"Besari, A.R.A., Saputra, A.A., Chin, W.H., Kubota, N., et\u00a0al. (2023). Hand\u2013object interaction recognition based on visual attention using multiscopic cyber-physical-social system. International Journal of Advances in Intelligent Informatics 9(2)","DOI":"10.26555\/ijain.v9i2.901"},{"key":"2838_CR3","doi-asserted-by":"crossref","unstructured":"Bousmalis, K., Silberman, N., Dohan, D., Erhan, D., & Krishnan, D. (2017). Unsupervised pixel-level domain adaptation with generative adversarial networks. In: CVPR, pp. 3722\u20133731","DOI":"10.1109\/CVPR.2017.18"},{"key":"2838_CR4","doi-asserted-by":"crossref","unstructured":"Cai, Q., Pan, Y., Ngo, C.-W., Tian, X., Duan, L., & Yao, T. (2019). Exploring object relation in mean teacher for cross-domain detection. In: CVPR, pp. 11457\u201311466","DOI":"10.1109\/CVPR.2019.01172"},{"issue":"3","key":"2838_CR5","doi-asserted-by":"publisher","first-page":"261","DOI":"10.1177\/0278364917700714","volume":"36","author":"B Calli","year":"2017","unstructured":"Calli, B., Singh, A., Bruce, J., Walsman, A., Konolige, K., Srinivasa, S., Abbeel, P., & Dollar, A. M. (2017). Yale-cmu-berkeley dataset for robotic manipulation research. The International Journal of Robotics Research, 36(3), 261\u2013268.","journal-title":"The International Journal of Robotics Research"},{"key":"2838_CR6","doi-asserted-by":"crossref","unstructured":"Carf\u00ec, A., Patten, T., Kuang, Y., Hammoud, A., Alameh, M., Maiettini, E., Weinberg, A. I., Faria, D., Mastrogiovanni, F., Aleny\u00e0, G., et\u00a0al. (2021). Hand-object interaction: From human demonstrations to robot manipulation. Frontiers in Robotics and AI,8, Article 714023.","DOI":"10.3389\/frobt.2021.714023"},{"key":"2838_CR7","unstructured":"Chang, A.X., Funkhouser, T., Guibas, L., Hanrahan, P., Huang, Q., Li, Z., Savarese, S., Savva, M., Song, S., Su, H., et\u00a0al. (2015). Shapenet: An information-rich 3d model repository"},{"key":"2838_CR8","doi-asserted-by":"crossref","unstructured":"Cheng, T., Shan, D., Hassen, A.S., Higgins, R.E.L., & Fouhey, D. (2023). Towards a richer 2d understanding of hands at scale. In: Thirty-seventh Conference on Neural Information Processing Systems","DOI":"10.52202\/075280-1327"},{"key":"2838_CR9","doi-asserted-by":"crossref","unstructured":"Choudhary, A., Mishra, D., & Karmakar, A. (2021). Domain adaptive egocentric person re-identification. In: Computer Vision and Image Processing (CVIP), pp. 81\u201392","DOI":"10.1007\/978-981-16-1103-2_8"},{"key":"2838_CR10","unstructured":"Contributors, M. (2020). OpenMMLab Pose Estimation Toolbox and Benchmark. https:\/\/github.com\/open-mmlab\/mmpose"},{"key":"2838_CR11","doi-asserted-by":"crossref","unstructured":"Csurka, G. (2017). Domain adaptation for visual applications: A comprehensive survey. https:\/\/arxiv.org\/abs\/1702.05374","DOI":"10.1007\/978-3-319-58347-1_1"},{"key":"2838_CR12","doi-asserted-by":"crossref","unstructured":"Damen, D., Doughty, H., Farinella, G.M., , Furnari, A., Ma, J., Kazakos, E., Moltisanti, D., Munro, J., Perrett, T., Price, W., & Wray, M. (2021). Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100. IJCV, 1\u201323","DOI":"10.1007\/s11263-021-01531-2"},{"key":"2838_CR13","doi-asserted-by":"crossref","unstructured":"Damen, D., Doughty, H., Farinella, G.M., Fidler, S., Furnari, A., Kazakos, E., Moltisanti, D., Munro, J., Perrett, T., Price, W., et\u00a0al. (2018). Scaling egocentric vision: The epic-kitchens dataset. In: ECCV, pp. 720\u2013736","DOI":"10.1007\/978-3-030-01225-0_44"},{"key":"2838_CR14","doi-asserted-by":"crossref","unstructured":"Darkhalil, A., Shan, D., Zhu, B., Ma, J., Kar, A., Higgins, R., Fidler, S., Fouhey, D., & Damen, D. (2022). Epic-kitchens visor benchmark: Video segmentations and object relations. In: NeurIPS, pp. 13745\u201313758","DOI":"10.52202\/068431-0999"},{"key":"2838_CR15","doi-asserted-by":"crossref","unstructured":"Deng, J., Li, W., Chen, Y., & Duan, L. (2021). Unbiased mean teacher for cross-domain object detection. In: CVPR, pp. 4091\u20134101","DOI":"10.1109\/CVPR46437.2021.00408"},{"key":"2838_CR16","doi-asserted-by":"publisher","first-page":"23241","DOI":"10.1007\/s11042-020-09597-9","volume":"80","author":"M Di Benedetto","year":"2021","unstructured":"Di Benedetto, M., Carrara, F., Meloni, E., Amato, G., Falchi, F., & Gennaro, C. (2021). Learning accurate personal protective equipment detection from virtual worlds. Multimedia Tools and Applications, 80, 23241\u201323253.","journal-title":"Multimedia Tools and Applications"},{"key":"2838_CR17","unstructured":"Dosovitskiy, A., Ros, G., Codevilla, F., Lopez, A., & Koltun, V. (2017). CARLA: An open urban driving simulator. In: Proceedings of the 1st Annual Conference on Robot Learning, pp. 1\u201316"},{"key":"2838_CR18","doi-asserted-by":"crossref","unstructured":"Downs, L., Francis, A., Koenig, N., Kinman, B., Hickman, R., Reymann, K., McHugh, T.B., & Vanhoucke, V. (2022). Google scanned objects: A high-quality dataset of 3d scanned household items. In: 2022 International Conference on Robotics and Automation (ICRA), pp. 2553\u20132560. IEEE","DOI":"10.1109\/ICRA46639.2022.9811809"},{"key":"2838_CR19","doi-asserted-by":"crossref","unstructured":"Edsinger, A., & Kemp, C.C. (2007). Human-robot interaction for cooperative manipulation: Handing objects to one another. In: RO-MAN 2007-The 16th IEEE International Symposium on Robot and Human Interactive Communication, pp. 1167\u20131172. IEEE","DOI":"10.1109\/ROMAN.2007.4415256"},{"key":"2838_CR20","unstructured":"Ester, M., Kriegel, H.-P., Sander, J., & Xu, X. (1996). Density-based spatial clustering of applications with noise. In: Int. Conf. Knowledge Discovery and Data Mining, vol. 240"},{"key":"2838_CR21","doi-asserted-by":"crossref","unstructured":"Fabbri, M., Bras\u00f3, G., Maugeri, G., O\u0161ep, A., Gasparini, R., Cetintas, O., Calderara, S., Leal-Taix\u00e9, L., & Cucchiara, R. (2021). Motsynth: How can synthetic data help pedestrian detection and tracking? In: ICCV","DOI":"10.1109\/ICCV48922.2021.01067"},{"key":"2838_CR22","doi-asserted-by":"crossref","unstructured":"Fu, Q., Liu, X., & Kitani, K.M. (2022). Sequential voting with relational box fields for active object detection. In: CVPR, pp. 2374\u20132383","DOI":"10.1109\/CVPR52688.2022.00241"},{"key":"2838_CR23","unstructured":"Ganin, Y., & Lempitsky, V. (2015). Unsupervised domain adaptation by backpropagation. In: International Conference on Machine Learning, pp. 1180\u20131189. PMLR"},{"key":"2838_CR24","doi-asserted-by":"crossref","unstructured":"Grauman, K., Westbury, A., Byrne, E., Chavis, Z.Q., Furnari, A., Girdhar, R., Hamburger, J., Jiang, H., Liu, M., Liu, X., Martin, M., Nagarajan, T., Radosavovic, I., Ramakrishnan, S.K., Ryan, F., Sharma, J., Wray, M., Xu, M., Xu, E.Z., & Malik, J. (2021). Ego4d: Around the world in 3,000 hours of egocentric video. In: CVPR, pp. 18995\u201319012","DOI":"10.1109\/CVPR52688.2022.01842"},{"key":"2838_CR25","doi-asserted-by":"crossref","unstructured":"Kappler, D., Bohg, J., & Schaal, S. (2015). Leveraging big data for grasp planning. In: 2015 IEEE International Conference on Robotics and Automation (ICRA), pp. 4304\u20134311. IEEE","DOI":"10.1109\/ICRA.2015.7139793"},{"issue":"8","key":"2838_CR26","doi-asserted-by":"publisher","first-page":"927","DOI":"10.1177\/0278364912445831","volume":"31","author":"A Kasper","year":"2012","unstructured":"Kasper, A., Xue, Z., & Dillmann, R. (2012). The kit object models database: An object model database for object recognition, localization and manipulation in service robotics. The International Journal of Robotics Research, 31(8), 927\u2013934.","journal-title":"The International Journal of Robotics Research"},{"key":"2838_CR27","doi-asserted-by":"crossref","unstructured":"Kirillov, A., Wu, Y., He, K., & Girshick, R. (2020). Pointrend: Image segmentation as rendering. In: CVPR, pp. 9799\u20139808","DOI":"10.1109\/CVPR42600.2020.00982"},{"key":"2838_CR28","unstructured":"Kolve, E., Mottaghi, R., Han, W., VanderBilt, E., Weihs, L., Herrasti, A., Gordon, D., Zhu, Y., Gupta, A., & Farhadi, A. (2017). AI2-THOR: An Interactive 3D Environment for Visual AI. arXiv"},{"key":"2838_CR29","unstructured":"Kolve, E., Mottaghi, R., Han, W., VanderBilt, E., Weihs, L., Herrasti, A., Gordon, D., Zhu, Y., Gupta, A., & Farhadi, A. (2017). Ai2-thor: An interactive 3d environment for visual ai. https:\/\/arxiv.org\/abs\/1712.05474"},{"key":"2838_CR30","doi-asserted-by":"crossref","unstructured":"Leonardi, R., Furnari, A., Ragusa, F.,& Farinella, G.M. (2025). Are synthetic data useful for egocentric hand-object interaction detection? In: European Conference on Computer Vision, pp. 36\u201354. Springer","DOI":"10.1007\/978-3-031-73209-6_3"},{"key":"2838_CR31","doi-asserted-by":"crossref","unstructured":"Leonardi, R., Ragusa, F., Furnari, A., & Farinella, G.M. (2022). Egocentric human-object interaction detection exploiting synthetic data. In: International Conference on Image Analysis and Processing, pp. 237\u2013248. Springer","DOI":"10.1007\/978-3-031-06430-2_20"},{"key":"2838_CR32","doi-asserted-by":"crossref","unstructured":"Leonardi, R., Ragusa, F., Furnari, A., & Farinella, G. M. (2024). Exploiting multimodal synthetic data for egocentric human-object interaction detection in an industrial scenario. Computer Vision and Image Understanding,242, Article 103984.","DOI":"10.1016\/j.cviu.2024.103984"},{"key":"2838_CR33","doi-asserted-by":"crossref","unstructured":"Li, Y.-J., Dai, X., Ma, C.-Y., Liu, Y.-C., Chen, K., Wu, B., He, Z., Kitani, K., & Vajda, P. (2022). Cross-domain adaptive teacher for object detection. In: CVPR, pp. 7581\u20137590","DOI":"10.1109\/CVPR52688.2022.00743"},{"key":"2838_CR34","doi-asserted-by":"crossref","unstructured":"Li, Y., Nagarajan, T., Xiong, B., & Grauman, K. (2021). Ego-exo: Transferring visual representations from third-person to first-person videos. In: CVPR, pp. 6943\u20136953","DOI":"10.1109\/CVPR46437.2021.00687"},{"key":"2838_CR35","unstructured":"Li, C., Xia, F., Mart\u00edn-Mart\u00edn, R., Lingelbach, M., Srivastava, S., Shen, B., Vainio, K.E., Gokmen, C., Dharan, G., Jain, T., Kurenkov, A., Liu, K., Gweon, H., Wu, J., Fei-Fei, L., & Savarese, S. (2022). igibson 2.0: Object-centric simulation for robot learning of everyday household tasks. In: Faust, A., Hsu, D., Neumann, G. (eds.) Proceedings of the 5th Conference on Robot Learning. Proceedings of Machine Learning Research, vol. 164, pp. 455\u2013465. PMLR. https:\/\/proceedings.mlr.press\/v164\/li22b.html"},{"key":"2838_CR36","doi-asserted-by":"crossref","unstructured":"Li, G., Zhao, K., Zhang, S., Lyu, X., Dusmanu, M., Zhang, Y., Pollefeys, M., & Tang, S. (2024). Egogen: An egocentric synthetic data generator. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 14497\u201314509","DOI":"10.1109\/CVPR52733.2024.01374"},{"key":"2838_CR37","doi-asserted-by":"crossref","unstructured":"Lin, T.-Y., Maire, M., Belongie, S.J., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., & Zitnick, C.L. (2014). Microsoft coco: Common objects in context. In: ECCV, pp. 740\u2013755","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"2838_CR38","unstructured":"Liu, Y.-C., Ma, C.-Y., He, Z., Kuo, C.-W., Chen, K., Zhang, P., Wu, B., Kira, Z., & Vajda, P. (2021). Unbiased teacher for semi-supervised object detection. In: ICLR"},{"key":"2838_CR39","doi-asserted-by":"crossref","unstructured":"Liu, Z., Mao, H., Wu, C.-Y., Feichtenhofer, C., Darrell, T., & Xie, S. (2022). A convnet for the 2020s. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 11976\u201311986","DOI":"10.1109\/CVPR52688.2022.01167"},{"key":"2838_CR40","doi-asserted-by":"crossref","unstructured":"Liu, S., Tripathi, S., Majumdar, S., & Wang, X. (2022). Joint hand motion and interaction hotspots prediction from egocentric videos. In: CVPR, pp. 3282\u20133292","DOI":"10.1109\/CVPR52688.2022.00328"},{"issue":"22","key":"2838_CR41","doi-asserted-by":"publisher","first-page":"11457","DOI":"10.3390\/app122211457","volume":"12","author":"Z Lv","year":"2022","unstructured":"Lv, Z., Poiesi, F., Dong, Q., Lloret, J., & Song, H. (2022). Deep learning for intelligent human-computer interaction. Applied Sciences, 12(22), 11457.","journal-title":"Applied Sciences"},{"key":"2838_CR42","doi-asserted-by":"crossref","unstructured":"Manolis Savva*, Abhishek Kadian*, Oleksandr Maksymets*, Zhao, Y., Wijmans, E., Jain, B., Straub, J., Liu, J., Koltun, V., Malik, J., Parikh, D., & Batra, D. (2019). Habitat: A Platform for Embodied AI Research. In: ICCV","DOI":"10.1109\/ICCV.2019.00943"},{"key":"2838_CR43","doi-asserted-by":"crossref","unstructured":"Munro, J., & Damen, D. (2020). Multi-modal domain adaptation for fine-grained action recognition. In: CVPR, pp. 122\u2013132","DOI":"10.1109\/CVPR42600.2020.00020"},{"key":"2838_CR44","unstructured":"Munro, J., Wray, M., Larlus, D., Csurka, G., & Damen, D. (2021). Domain adaptation in multi-view embedding for cross-modal video retrieval. ArXiv abs\/2110.12812"},{"key":"2838_CR45","unstructured":"NVIDIA. (2021). NVIDIA Isaac Sim. https:\/\/developer.nvidia.com\/isaac-sim"},{"key":"2838_CR46","unstructured":"Oquab, M., Darcet, T., Moutakanni, T., Vo, H.V., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., Howes, R., Huang, P.-Y., Xu, H., Sharma, V., Li, S.-W., Galuba, W., Rabbat, M., Assran, M., Ballas, N., & Bojanowski, P. (2023). DINOv2: Learning Robust Visual Features without Supervision"},{"key":"2838_CR47","doi-asserted-by":"crossref","unstructured":"Pasqualino, G., Furnari, A., Signorello, G., & Farinella, G. M. (2021). An unsupervised domain adaptation scheme for single-stage artwork recognition in cultural sites. Image and Vision Computing,107, Article 104098.","DOI":"10.1016\/j.imavis.2021.104098"},{"key":"2838_CR48","doi-asserted-by":"crossref","unstructured":"Plizzari, C., Perrett, T., Caputo, B., & Damen, D. (2023). What can a cook in italy teach a mechanic in india? action recognition generalisation over scenarios and locations. In: ICCV2023","DOI":"10.1109\/ICCV51070.2023.01256"},{"key":"2838_CR49","doi-asserted-by":"crossref","unstructured":"Quattrocchi, C., Mauro, D.D., Furnari, A., Lopes, A., Moltisanti, M., & Farinella, G.M. (2023). Put your ppe on: A tool for synthetic data generation and related benchmark in construction site scenarios. In: International Conference on Computer Vision Theory and Applications, pp. 656\u2013663","DOI":"10.5220\/0011718000003417"},{"key":"2838_CR50","doi-asserted-by":"crossref","unstructured":"Ragusa, F., Furnari, A., Livatino, S., & Farinella, G.M. (2021). The meccano dataset: Understanding human-object interactions from egocentric videos in an industrial-like domain. In: Winter Conference on Applications of Computer Vision, pp. 1569\u20131578","DOI":"10.1109\/WACV48630.2021.00161"},{"key":"2838_CR51","doi-asserted-by":"crossref","unstructured":"Ragusa, F., Leonardi, R., Mazzamuto, M., Bonanno, C., Scavo, R., Furnari, A., & Farinella, G.M. (2024). Enigma-51: Towards a fine-grained understanding of human behavior in industrial scenarios. In: Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, pp. 4549\u20134559","DOI":"10.1109\/WACV57701.2024.00449"},{"key":"2838_CR52","unstructured":"Ramakrishnan, S.K., Gokaslan, A., Wijmans, E., Maksymets, O., Clegg, A., Turner, J.M., Undersander, E., Galuba, W., Westbury, A., Chang, A.X., et\u00a0al. (2021). Habitat-matterport 3d dataset (hm3d): 1000 large-scale 3d environments for embodied ai. In: NeurIPS"},{"key":"2838_CR53","unstructured":"Rockstar Games. (2024). Rockstar Advanced Game Engine (RAGE). https:\/\/www.rockstargames.com\/. Proprietary game engine developed by Rockstar Games (Accessed 2024)"},{"issue":"3","key":"2838_CR54","doi-asserted-by":"publisher","first-page":"211","DOI":"10.1007\/s11263-015-0816-y","volume":"115","author":"O Russakovsky","year":"2015","unstructured":"Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., & Fei-Fei, L. (2015). ImageNet Large Scale Visual Recognition Challenge. IJCV, 115(3), 211\u2013252.","journal-title":"IJCV"},{"key":"2838_CR55","doi-asserted-by":"crossref","unstructured":"Saito, K., Watanabe, K., Ushiku, Y., & Harada, T. (2018). Maximum classifier discrepancy for unsupervised domain adaptation. In: CVPR, pp. 3723\u20133732","DOI":"10.1109\/CVPR.2018.00392"},{"key":"2838_CR56","doi-asserted-by":"crossref","unstructured":"Savva, M., Chang, A.X., & Hanrahan, P. (2015). Semantically-enriched 3d models for common-sense knowledge. In: CVPRW, pp. 24\u201331","DOI":"10.1109\/CVPRW.2015.7301289"},{"key":"2838_CR57","doi-asserted-by":"crossref","unstructured":"Savva, M., Kadian, A., Maksymets, O., Zhao, Y., Wijmans, E., Jain, B., Straub, J., Liu, J., Koltun, V., Malik, J., et\u00a0al. (2019). Habitat: A platform for embodied ai research. In: ICCV, pp. 9339\u20139347","DOI":"10.1109\/ICCV.2019.00943"},{"key":"2838_CR58","doi-asserted-by":"crossref","unstructured":"Sener, F., Chatterjee, D., Shelepov, D., He, K., Singhania, D., Wang, R., & Yao, A. (2022). Assembly101: A large-scale multi-view video dataset for understanding procedural activities. In: CVPR, pp. 21096\u201321106","DOI":"10.1109\/CVPR52688.2022.02042"},{"key":"2838_CR59","unstructured":"shadowrobot. (2005). ShadowHand. https:\/\/www.shadowrobot.com\/dexterous-hand-series\/"},{"key":"2838_CR60","doi-asserted-by":"crossref","unstructured":"Shan, D., Geng, J., Shu, M., & Fouhey, D.F. (2020). Understanding human hands in contact at internet scale. In: CVPR, pp. 9869\u20139878","DOI":"10.1109\/CVPR42600.2020.00989"},{"key":"2838_CR61","doi-asserted-by":"crossref","unstructured":"Singh, A., Sha, J., Narayan, K.S., Achim, T., & Abbeel, P. (2014). Bigbird: A large-scale 3d database of object instances. In: 2014 IEEE International Conference on Robotics and Automation (ICRA), pp. 509\u2013516. IEEE","DOI":"10.1109\/ICRA.2014.6906903"},{"key":"2838_CR62","first-page":"251","volume":"34","author":"A Szot","year":"2021","unstructured":"Szot, A., Clegg, A., Undersander, E., Wijmans, E., Zhao, Y., Turner, J., Maestre, N., Mukadam, M., Chaplot, D. S., Maksymets, O., et al. (2021). Habitat 2.0: Training home assistants to rearrange their habitat. Advances in Neural Information Processing Systems, 34, 251\u2013266.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2838_CR63","doi-asserted-by":"crossref","unstructured":"Tang, Y., Tian, Y., Lu, J., Feng, J., & Zhou, J. (2017). Action recognition in rgb-d egocentric videos. In: 2017 IEEE International Conference on Image Processing (ICIP), pp. 3410\u20133414. IEEE","DOI":"10.1109\/ICIP.2017.8296915"},{"key":"2838_CR64","unstructured":"Tarvainen, A., & Valpola, H. (2017). Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. NeurIPS 30"},{"key":"2838_CR65","doi-asserted-by":"crossref","unstructured":"Tzeng, E., Hoffman, J., Saenko, K., & Darrell, T. (2017). Adversarial discriminative domain adaptation. In: CVPR, pp. 7167\u20137176","DOI":"10.1109\/CVPR.2017.316"},{"key":"2838_CR66","unstructured":"Unity. (2022). SyntheticHumans Package (Unity Computer Vision). https:\/\/github.com\/Unity-Technologies\/com.unity.cv.synthetichumans"},{"key":"2838_CR67","doi-asserted-by":"crossref","unstructured":"Wang, R., Zhang, J., Chen, J., Xu, Y., Li, P., Liu, T., & Wang, H. (2023). Dexgraspnet: A large-scale robotic dexterous grasp dataset for general objects based on simulation. In: CVPR, pp. 11359\u201311366","DOI":"10.1109\/ICRA48891.2023.10160982"},{"key":"2838_CR68","doi-asserted-by":"crossref","unstructured":"Xia, F., R.\u00a0Zamir, A., He, Z.-Y., Sax, A., Malik, J., & Savarese, S. (2018). Gibson env: real-world perception for embodied agents. In: CVPR","DOI":"10.1109\/CVPR.2018.00945"},{"issue":"2","key":"2838_CR69","doi-asserted-by":"publisher","first-page":"713","DOI":"10.1109\/LRA.2020.2965078","volume":"5","author":"F Xia","year":"2020","unstructured":"Xia, F., Shen, W. B., Li, C., Kasimbeg, P., Tchapmi, M. E., Toshev, A., Mart\u00edn-Mart\u00edn, R., & Savarese, S. (2020). Interactive gibson benchmark: A benchmark for interactive navigation in cluttered environments. IEEE Robotics and Automation Letters, 5(2), 713\u2013720.","journal-title":"IEEE Robotics and Automation Letters"},{"key":"2838_CR70","doi-asserted-by":"crossref","unstructured":"Zhang, L., Zhou, S., Stent, S., & Shi, J. (2022). Fine-grained egocentric hand-object segmentation: Dataset, model, and applications. In: ECCV, pp. 127\u2013145","DOI":"10.1007\/978-3-031-19818-2_8"},{"key":"2838_CR71","doi-asserted-by":"crossref","unstructured":"Zhu, C., Xiao, F., Alvarado, A., Babaei, Y., Hu, J., El-Mohri, H., Culatana, S., Sumbaly, R., & Yan, Z. (2023). Egoobjects: A large-scale egocentric dataset for fine-grained object understanding. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision, pp. 20110\u201320120","DOI":"10.1109\/ICCV51070.2023.01840"},{"key":"2838_CR72","doi-asserted-by":"crossref","unstructured":"Zhuang, F., Qi, Z., Duan, K., Xi, D., Zhu, Y., Zhu, H., Xiong, H., & He, Q. (2020). A comprehensive survey on transfer learning. Proceedings of the IEEE,109(1), 43\u201376.","DOI":"10.1109\/JPROC.2020.3004555"}],"container-title":["International Journal of Computer Vision"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-026-02838-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11263-026-02838-8","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-026-02838-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,20]],"date-time":"2026-07-20T16:15:47Z","timestamp":1784564147000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11263-026-02838-8"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,20]]},"references-count":72,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2026,6]]}},"alternative-id":["2838"],"URL":"https:\/\/doi.org\/10.1007\/s11263-026-02838-8","relation":{},"ISSN":["0920-5691","1573-1405"],"issn-type":[{"value":"0920-5691","type":"print"},{"value":"1573-1405","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,5,20]]},"assertion":[{"value":"5 December 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"31 March 2026","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"20 May 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"279"}}