{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,7,2]],"date-time":"2025-07-02T17:16:35Z","timestamp":1751476595010,"version":"3.37.3"},"reference-count":30,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2024,4,29]],"date-time":"2024-04-29T00:00:00Z","timestamp":1714348800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,4,29]],"date-time":"2024-04-29T00:00:00Z","timestamp":1714348800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100006280","name":"Ministerio de Ciencia y Tecnolog\u00eda","doi-asserted-by":"publisher","award":["PID2019-103871GB-I00","PID2019-103871GB-I00","PID2019-103871GB-I00","PID2019-103871GB-I00"],"award-info":[{"award-number":["PID2019-103871GB-I00","PID2019-103871GB-I00","PID2019-103871GB-I00","PID2019-103871GB-I00"]}],"id":[{"id":"10.13039\/501100006280","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Universidad de C\u00f3rdoba"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Pattern Anal Applic"],"published-print":{"date-parts":[[2024,6]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Human interaction recognition (HIR) is a significant challenge in computer vision that focuses on identifying human interactions in images and videos. HIR presents a great complexity due to factors such as pose diversity, varying scene conditions, or the presence of multiple individuals. Recent research has explored different approaches to address it, with an increasing emphasis on human pose estimation. In this work, we propose Proxemics-Net++, an extension of the Proxemics-Net model, capable of addressing the problem of recognizing human interactions in images through two different tasks: the identification of the types of \u201ctouch codes\u201d or proxemics and the identification of the type of social relationship between pairs. To achieve this, we use RGB and body pose information together with the state-of-the-art deep learning architecture, ConvNeXt, as the backbone. We performed an ablative analysis to understand how the combination of RGB and body pose information affects these two tasks. Experimental results show that body pose information contributes significantly to proxemic recognition (first task) as it allows to improve the existing state of the art, while its contribution in the classification of social relations (second task) is limited due to the ambiguity of labelling in this problem, resulting in RGB information being more influential in this task.\n<\/jats:p>","DOI":"10.1007\/s10044-024-01270-3","type":"journal-article","created":{"date-parts":[[2024,4,29]],"date-time":"2024-04-29T13:02:06Z","timestamp":1714395726000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Proxemics-net++: classification of human interactions in still images"],"prefix":"10.1007","volume":"27","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1120-2874","authenticated-orcid":false,"given":"Isabel","family":"Jim\u00e9nez-Velasco","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1735-8837","authenticated-orcid":false,"given":"Jorge","family":"Zafra-Palma","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8773-8571","authenticated-orcid":false,"given":"Rafael","family":"Mu\u00f1oz-Salinas","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9294-6714","authenticated-orcid":false,"given":"Manuel J.","family":"Mar\u00edn-Jim\u00e9nez","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2024,4,29]]},"reference":[{"key":"1270_CR1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2012.24","author":"A Patron","year":"2012","unstructured":"Patron A, Reid I, Marszalek M, Zisserman A (2012) Structured learning of human interactions in TV shows. IEEE Trans Pattern Anal Mach Intell. https:\/\/doi.org\/10.1109\/TPAMI.2012.24","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"1270_CR2","doi-asserted-by":"publisher","unstructured":"Yang Y, Baker S, Kannan A, Ramanan D (2012) Recognizing proxemics in personal photos. In: IEEE conference on CVPR. https:\/\/doi.org\/10.1109\/CVPR.2012.6248095","DOI":"10.1109\/CVPR.2012.6248095"},{"issue":"4","key":"1270_CR3","doi-asserted-by":"publisher","first-page":"361","DOI":"10.14201\/ADCAIJ2021104361379","volume":"10","author":"AW Muhamada","year":"2021","unstructured":"Muhamada AW, Mohammed AA (2021) Review on recent computer vision methods for human action recognition. ADCAIJ 10(4):361\u2013379. https:\/\/doi.org\/10.14201\/ADCAIJ2021104361379","journal-title":"ADCAIJ"},{"key":"1270_CR4","doi-asserted-by":"publisher","DOI":"10.1155\/2022\/8323962","author":"VT Le","year":"2022","unstructured":"Le VT, Tran K, Truong V (2022) A comprehensive review of recent deep learning techniques for human activity recognition. Comput Intell Neurosci. https:\/\/doi.org\/10.1155\/2022\/8323962","journal-title":"Comput Intell Neurosci"},{"key":"1270_CR5","doi-asserted-by":"publisher","first-page":"653","DOI":"10.1007\/s10044-021-00988-8","volume":"25","author":"CMA Ilyas","year":"2022","unstructured":"Ilyas CMA, Rehm M, Nasrollahi K (2022) Deep transfer learning in human-robot interaction for cognitive and physical rehabilitation purposes. Pattern Anal Appl 25:653\u2013677. https:\/\/doi.org\/10.1007\/s10044-021-00988-8","journal-title":"Pattern Anal Appl"},{"key":"1270_CR6","doi-asserted-by":"publisher","DOI":"10.1007\/s10044-023-01202-7","author":"M Gutoski","year":"2023","unstructured":"Gutoski M, Lazzaretti AE, Lopes HS (2023) Unsupervised open-world human action recognition. Pattern Anal Appl. https:\/\/doi.org\/10.1007\/s10044-023-01202-7","journal-title":"Pattern Anal Appl"},{"key":"1270_CR7","doi-asserted-by":"publisher","first-page":"1750","DOI":"10.1007\/s11263-020-01295-1","volume":"128","author":"J Li","year":"2020","unstructured":"Li J, Wong Y, Zhao Q (2020) Visual social relationship recognition. IJCV 128:1750\u20131764. https:\/\/doi.org\/10.1007\/s11263-020-01295-1","journal-title":"IJCV"},{"key":"1270_CR8","doi-asserted-by":"publisher","first-page":"116265","DOI":"10.1016\/j.image.2021.116265","volume":"95","author":"G Tanisik","year":"2021","unstructured":"Tanisik G, Zalluhoglu C, Ikizler N (2021) Multi-stream pose convolutional neural networks for human interaction recognition in images. Signal Process Image Commun 95:116265. https:\/\/doi.org\/10.1016\/j.image.2021.116265","journal-title":"Signal Process Image Commun"},{"key":"1270_CR9","doi-asserted-by":"publisher","first-page":"108645","DOI":"10.1016\/j.patcog.2022.108645","volume":"128","author":"DG Lee","year":"2022","unstructured":"Lee DG, Lee SW (2022) Human interaction recognition framework based on interacting body part attention. Pattern Recognit. 128:108645. https:\/\/doi.org\/10.1016\/j.patcog.2022.108645","journal-title":"Pattern Recognit."},{"key":"1270_CR10","doi-asserted-by":"publisher","first-page":"100062","DOI":"10.1016\/j.birob.2022.100062","volume":"2","author":"R Sun","year":"2022","unstructured":"Sun R, Zhang Q, Luo C et al (2022) Human action recognition using a convolutional neural network based on skeleton heatmaps from two-stage pose estimation. Biomim Intell Robot 2:100062. https:\/\/doi.org\/10.1016\/j.birob.2022.100062","journal-title":"Biomim Intell Robot"},{"key":"1270_CR11","unstructured":"Dosovitskiy A, Beyer L, Kolesnikov A, et al (2021) An image is worth 16x16 words: transformers for image recognition at scale. ICLR"},{"key":"1270_CR12","doi-asserted-by":"publisher","unstructured":"Liu Z, Mao H, Wu CY (2022) A convnet for the 2020s. In: IEEE\/CVF conference on CVPR. https:\/\/doi.org\/10.1109\/CVPR52688.2022.01167","DOI":"10.1109\/CVPR52688.2022.01167"},{"key":"1270_CR13","doi-asserted-by":"publisher","unstructured":"Jim\u00e9nez I, Mu\u00f1oz R, Mar\u00edn MJ (2023) Proxemics-net: automatic proxemics recognition in images. In: Iberian conference on pattern recognition and image analysis, pp 402\u2013413. https:\/\/doi.org\/10.1007\/978-3-031-36616-1_32. IbPRIA 2023","DOI":"10.1007\/978-3-031-36616-1_32"},{"key":"1270_CR14","doi-asserted-by":"publisher","unstructured":"Guler RA, Neverova N, Kokkinos L (2018) Densepose: dense human pose estimation in the wild. In: Proceedings of the IEEE conference on CVPR, pp 7297\u20137306.https:\/\/doi.org\/10.1109\/CVPR.2018.00762","DOI":"10.1109\/CVPR.2018.00762"},{"issue":"5","key":"1270_CR15","doi-asserted-by":"publisher","first-page":"1003","DOI":"10.1525\/aa.1963.65.5.02a00020","volume":"65","author":"TH Edward","year":"1963","unstructured":"Edward TH (1963) A system for the notation of proxemic behavior. Am Anthropol 65(5):1003\u20131026. https:\/\/doi.org\/10.1525\/aa.1963.65.5.02a00020","journal-title":"Am Anthropol"},{"key":"1270_CR16","doi-asserted-by":"publisher","unstructured":"Chu X, Ouyang W, Yang W (2015) Multi-task recurrent neural network for immediacy prediction. In: Proceedings of the IEEE international conference on computer vision, pp 3352\u20133360. https:\/\/doi.org\/10.1109\/ICCV.2015.383","DOI":"10.1109\/ICCV.2015.383"},{"key":"1270_CR17","doi-asserted-by":"publisher","unstructured":"Jiang H, Grauman K (2017) Detangling people: individuating multiple close people and their body parts via region assembly. In: IEEE conference on CVPR, pp 3435\u20133443. https:\/\/doi.org\/10.1109\/CVPR.2017.366","DOI":"10.1109\/CVPR.2017.366"},{"key":"1270_CR18","doi-asserted-by":"publisher","unstructured":"Zhang M, Liu X, Liu W (2019) Multi-granularity reasoning for social relation recognition from images. In: IEEE international conference on multimedia and expo (ICME), pp 1618\u20131623. https:\/\/doi.org\/10.1109\/ICME.2019.00279","DOI":"10.1109\/ICME.2019.00279"},{"key":"1270_CR19","doi-asserted-by":"publisher","unstructured":"Goel A, Ma K, Tan C (2019) An end-to-end network for generating social relationship graphs. In: IEEE\/CVF conference on CVPR, pp 11178\u201311187. https:\/\/doi.org\/10.1109\/CVPR.2019.01144","DOI":"10.1109\/CVPR.2019.01144"},{"key":"1270_CR20","doi-asserted-by":"publisher","unstructured":"Li W, Duan Y, Lu J (2020) Graph-based social relation reasoning. In: European conference on computer vision, pp 18\u201334. https:\/\/doi.org\/10.1007\/978-3-030-58555-6_2","DOI":"10.1007\/978-3-030-58555-6_2"},{"key":"1270_CR21","doi-asserted-by":"publisher","first-page":"3979","DOI":"10.1007\/s00371-021-02244-w","volume":"38","author":"L Li","year":"2022","unstructured":"Li L, Qing L, Wang Y (2022) HF-SRGR: a new hybrid feature-driven social relation graph reasoning model. Vis Comput 38:3979\u20133992. https:\/\/doi.org\/10.1007\/s00371-021-02244-w","journal-title":"Vis Comput"},{"key":"1270_CR22","doi-asserted-by":"publisher","first-page":"99398","DOI":"10.1109\/ACCESS.2021.3096553","volume":"9","author":"X Yang","year":"2021","unstructured":"Yang X, Xu F, Wu K (2021) Gaze-aware graph convolutional network for social relation recognition. IEEE Access 9:99398\u201399408. https:\/\/doi.org\/10.1109\/ACCESS.2021.3096553","journal-title":"IEEE Access"},{"key":"1270_CR23","doi-asserted-by":"publisher","first-page":"103785","DOI":"10.1016\/j.cviu.2023.103785","volume":"235","author":"EV Sousa","year":"2023","unstructured":"Sousa EV, Macharet DG (2023) Structural reasoning for image-based social relation recognition. Comput Vis Image Underst 235:103785. https:\/\/doi.org\/10.1016\/j.cviu.2023.103785","journal-title":"Comput Vis Image Underst"},{"key":"1270_CR24","doi-asserted-by":"publisher","first-page":"1307","DOI":"10.1007\/s10044-018-0727-y","volume":"22","author":"M Farrajota","year":"2019","unstructured":"Farrajota M, Rodrigues JMF, Du JMH (2019) Human action recognition in videos with articulated pose information by deep networks. Pattern Anal Appl 22:1307\u20131318. https:\/\/doi.org\/10.1007\/s10044-018-0727-y","journal-title":"Pattern Anal Appl"},{"key":"1270_CR25","doi-asserted-by":"publisher","DOI":"10.1109\/TITS.2021.3069376","author":"L Bertoni","year":"2021","unstructured":"Bertoni L, Kreiss S, Alahi A (2021) Perceiving humans: from monocular 3d localization to social distancing. IEEE Trans Intell Transp Syst. https:\/\/doi.org\/10.1109\/TITS.2021.3069376","journal-title":"IEEE Trans Intell Transp Syst"},{"issue":"3","key":"1270_CR26","doi-asserted-by":"publisher","first-page":"211","DOI":"10.1007\/s11263-015-0816-y","volume":"115","author":"O Russakovsky","year":"2015","unstructured":"Russakovsky O, Deng J, Su H (2015) ImageNet large scale visual recognition challenge. IJCV 115(3):211\u2013252. https:\/\/doi.org\/10.1007\/s11263-015-0816-y","journal-title":"IJCV"},{"key":"1270_CR27","doi-asserted-by":"crossref","unstructured":"Liu Z, Lin Y, Cao Y, Hu H, Wei Y, Zhang Z, Lin S, Guo B (2021) Swin transformer: hierarchical vision transformer using shifted windows. In: Proceedings of the IEEE\/CVF ICCV","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"1270_CR28","doi-asserted-by":"publisher","unstructured":"Liu Z, Courant R, Kalogeiton V (2023) Funnynet: audiovisual learning of funny moments in videos. In: Computer vision\u2014ACCV 2022, pp 433\u2013450. https:\/\/doi.org\/10.1007\/978-3-031-26316-3_26","DOI":"10.1007\/978-3-031-26316-3_26"},{"key":"1270_CR29","unstructured":"Yang Y, Baker S, Kannan A, Ramanan L (2012) PROXEMICS dataset. https:\/\/www.dropbox.com\/s\/5zarkyny7ywc2fv\/PROXEMICS.zip?dl=0. Last visited: 26-October-2023"},{"key":"1270_CR30","unstructured":"Wu Y, Kirillov A, Massa F, et al (2019) Detectron2. https:\/\/github.com\/facebookresearch\/detectron2. Last visited: 26-October-2023"}],"container-title":["Pattern Analysis and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10044-024-01270-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10044-024-01270-3\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10044-024-01270-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,6,17]],"date-time":"2024-06-17T14:12:20Z","timestamp":1718633540000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10044-024-01270-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,4,29]]},"references-count":30,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2024,6]]}},"alternative-id":["1270"],"URL":"https:\/\/doi.org\/10.1007\/s10044-024-01270-3","relation":{},"ISSN":["1433-7541","1433-755X"],"issn-type":[{"type":"print","value":"1433-7541"},{"type":"electronic","value":"1433-755X"}],"subject":[],"published":{"date-parts":[[2024,4,29]]},"assertion":[{"value":"27 October 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"11 April 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"29 April 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors have no relevant financial or non-financial interests to disclose.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"For this article, the authors did not undertake work that involved humans or animals.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical approval"}}],"article-number":"49"}}