{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T09:03:13Z","timestamp":1777626193379,"version":"3.51.4"},"reference-count":49,"publisher":"MDPI AG","issue":"15","license":[{"start":{"date-parts":[[2022,8,1]],"date-time":"2022-08-01T00:00:00Z","timestamp":1659312000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Nature Science Foundation of China","doi-asserted-by":"publisher","award":["61871278"],"award-info":[{"award-number":["61871278"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Social relationships refer to the connections that exist between people and indicate how people interact in society. The effective recognition of social relationships is conducive to further understanding human behavioral patterns and thus can be vital for more complex social intelligent systems, such as interactive robots and health self-management systems. The existing works about social relation recognition (SRR) focus on extracting features on different scales but lack a comprehensive mechanism to orchestrate various features which show different degrees of importance. In this paper, we propose a new SRR framework, namely Multi-level Transformer-Based Social Relation Recognition (MT-SRR), for better orchestrating features on different scales. Specifically, a vision transformer (ViT) is firstly employed as a feature extraction module for its advantage in exploiting global features. An intra-relation transformer (Intra-TRM) is then introduced to dynamically fuse the extracted features to generate more rational social relation representations. Next, an inter-relation transformer (Inter-TRM) is adopted to further enhance the social relation representations by attentionally utilizing the logical constraints among relationships. In addition, a new margin related to inter-class similarity and a sample number are added to alleviate the challenges of a data imbalance. Extensive experiments demonstrate that MT-SRR can better fuse features on different scales as well as ameliorate the bad effect caused by a data imbalance. The results on the benchmark datasets show that our proposed model outperforms the state-of-the-art methods with significant improvement.<\/jats:p>","DOI":"10.3390\/s22155749","type":"journal-article","created":{"date-parts":[[2022,8,1]],"date-time":"2022-08-01T23:49:27Z","timestamp":1659397767000},"page":"5749","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":11,"title":["Multi-Level Transformer-Based Social Relation Recognition"],"prefix":"10.3390","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8697-365X","authenticated-orcid":false,"given":"Yuchen","family":"Wang","sequence":"first","affiliation":[{"name":"College of Electronics and Information Engineering, Sichuan University, Chengdu 610065, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3555-0005","authenticated-orcid":false,"given":"Linbo","family":"Qing","sequence":"additional","affiliation":[{"name":"College of Electronics and Information Engineering, Sichuan University, Chengdu 610065, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhengyong","family":"Wang","sequence":"additional","affiliation":[{"name":"College of Electronics and Information Engineering, Sichuan University, Chengdu 610065, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7282-7638","authenticated-orcid":false,"given":"Yongqiang","family":"Cheng","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Technology, University of Hull, Hull HU6 7RX, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5508-1819","authenticated-orcid":false,"given":"Yonghong","family":"Peng","sequence":"additional","affiliation":[{"name":"Department of Computing and Mathematics, Manchester Metropolitan University, Manchester M1 5GD, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2022,8,1]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"S54","DOI":"10.1177\/0022146510383501","article-title":"Social Relationships and Health: A Flashpoint for Health Policy","volume":"51","author":"Umberson","year":"2010","journal-title":"J. Health Soc. Behav."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Ramanathan, V., Yao, B., and Li, F.F. (2013, January 23\u201328). Social Role Discovery in Human Events. Proceedings of the 2013 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Portland, OR, USA.","DOI":"10.1109\/CVPR.2013.320"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Quiroz, M., Pati\u00f1o, R., Diaz-Amado, J., and Cardinale, Y. (2022). Group Emotion Detection Based on Social Robot Perception. Sensors, 22.","DOI":"10.3390\/s22103749"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Sou, K., Shiokawa, H., Yoh, K., and Doi, K. (2021). Street Design for Hedonistic Sustainability through AI and Human Co-Operative Evaluation. Sustainability, 13.","DOI":"10.3390\/su13169066"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Rato, D., and Prada, R. (2021). Towards Social Identity in Socio-Cognitive Agents. Sustainability, 13.","DOI":"10.3390\/su132011390"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"259","DOI":"10.26599\/BDMA.2020.9020006","article-title":"Survey on data analysis in social media: A practical application aspect","volume":"3","author":"Hou","year":"2020","journal-title":"Big Data Min. Anal."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Li, W., and Zlatanova, S. (2021). Significant Geo-Social Group Discovery over Location-Based Social Network. Sensors, 21.","DOI":"10.3390\/s21134551"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Minetto, A., Nardin, A., and Dovis, F. (2021). Modelling and Experimental Assessment of Inter-Personal Distancing Based on Shared GNSS Observables. Sensors, 21.","DOI":"10.3390\/s21082588"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Liu, M., Quan, Z.W., Wu, J.M., Liu, Y., and Han, M. (2022). Embedding temporal networks inductively via mining neighborhood and community influences. Appl. Intell., 1\u201320.","DOI":"10.1007\/s10489-021-03102-x"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Guo, X., Xiang, Y., and Chen, Q. (2011, January 26\u201328). A vector space model approach to social relation extraction from text corpus. Proceedings of the 2011 Eighth International Conference on Fuzzy Systems and Knowledge Discovery (FSKD), Shanghai, China.","DOI":"10.1109\/FSKD.2011.6019806"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Cernian, A., Vasile, N., and Sacala, I.S. (2021). Fostering Cyber-Physical Social Systems through an Ontological Approach to Personality Classification Based on Social Media Posts. Sensors, 21.","DOI":"10.3390\/s21196611"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Li, J., Wong, Y., Zhao, Q., and Kankanhalli, M. (2017, January 22\u201329). Dual-Glance Model for Deciphering Social Relationships. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.289"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Dai, P., Lv, J., and Wu, B. (2019, January 8\u201312). Two-Stage Model for Social Relationship Understanding from Videos. Proceedings of the 2019 IEEE International Conference on Multimedia and Expo (ICME), Shanghai, China.","DOI":"10.1109\/ICME.2019.00198"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Qing, L., Li, L., Xu, S., Huang, Y., Liu, M., Jin, R., Liu, B., Niu, T., Wen, H., and Wang, Y. (2021, January 10\u201317). Public Life in Public Space (PLPS): A multi-task, multi-group video dataset for public life research. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV) Workshops, Montreal, QC, Canada.","DOI":"10.1109\/ICCVW54120.2021.00404"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Goel, A., Ma, K.T., and Tan, C. (2019, January 15\u201320). An End-To-End Network for Generating Social Relationship Graphs. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01144"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"410","DOI":"10.1016\/j.patrec.2020.08.005","article-title":"Deep supervised feature selection for social relationship recognition","volume":"138","author":"Wang","year":"2020","journal-title":"Pattern Recognit. Lett."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Qing, L., Li, L., Wang, Y., Cheng, Y., and Peng, Y. (2021). SRR-LGR: Local\u2013Global Information-Reasoned Social Relation Recognition for Human-Oriented Observation. Remote Sens., 13.","DOI":"10.3390\/rs13112038"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Li, L., Qing, L., Wang, Y., Su, J., Cheng, Y., and Peng, Y. (2021). HF-SRGR: A new hybrid feature-driven social relation graph reasoning model. Vis. Comput., 1\u201314.","DOI":"10.1007\/s00371-021-02244-w"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Li, W., Duan, Y., Lu, J., Feng, J., and Zhou, J. (2020, January 23\u201328). Graph-based social relation reasoning. Proceedings of the 16th European Conference on Computer Vision (ECCV), Glasgow, UK.","DOI":"10.1007\/978-3-030-58555-6_2"},{"key":"ref_20","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2021, January 3\u20137). An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. Proceedings of the Ninth International Conference on Learning Representations (lCLR), Vienna, Austria."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Sun, Q., Schiele, B., and Fritz, M. (2017, January 21\u201326). A Domain Based Approach to Social Relation Recognition. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.54"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Fang, R., Tang, K.D., Snavely, N., and Chen, T. (2010, January 26\u201329). Towards computational models of kinship verification. Proceedings of the 2010 IEEE International Conference on Image Processing (ICIP), Hong Kong, China.","DOI":"10.1109\/ICIP.2010.5652590"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Dibeklioglu, H., Salah, A.A., and Gevers, T. (2013, January 1\u20138). Like father, like son: Facial expression dynamics for kinship verification. Proceedings of the 2013 IEEE International Conference on Computer Vision (ICCV), Sydney, Australia.","DOI":"10.1109\/ICCV.2013.189"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"243","DOI":"10.1016\/j.neucom.2021.05.097","article-title":"Multi-scale features based interpersonal relation recognition using higher-order graph neural network","volume":"456","author":"Gao","year":"2021","journal-title":"Neurocomputing"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Zhang, M., Liu, X., Liu, W., Zhou, A., Ma, H., and Mei, T. (2019, January 8\u201312). Multi-Granularity Reasoning for Social Relation Recognition From Images. Proceedings of the 2019 IEEE International Conference on Multimedia and Expo (ICME), Shanghai, China.","DOI":"10.1109\/ICME.2019.00279"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Wang, G., Gallagher, A., Luo, J., and Forsyth, D. (2010, January 5\u201311). Seeing people in social context: Recognizing people and social relationships. Proceedings of the 11th European Conference on Computer Vision (ECCV), Heraklion, Greece.","DOI":"10.1007\/978-3-642-15555-0_13"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"1046","DOI":"10.1109\/TMM.2012.2187436","article-title":"Understanding kin relationships in a photo","volume":"14","author":"Xia","year":"2012","journal-title":"IEEE Trans. Multimed."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"331","DOI":"10.1109\/TPAMI.2013.134","article-title":"Neighborhood Repulsed Metric Learning for Kinship Verification","volume":"36","author":"Lu","year":"2014","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Wang, Z., Chen, T., Ren, J., Yu, W., Cheng, H., and Lin, L. (2018, January 13\u201319). Deep reasoning with knowledge graph for social relationship understanding. Proceedings of the 27th International Joint Conference on Artificial Intelligence (IJCAI), Stockholm, Sweden.","DOI":"10.24963\/ijcai.2018\/142"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Wu, H., Codella, N., Liu, M., Dai, X., Yuan, L., and Zhang, L. (2021, January 10\u201317). CvT: Introducing Convolutions to Vision Transformers. Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00009"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021, January 11\u201317). Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Wang, W., Xie, E., Li, X., Fan, D.P., Song, K., Liang, D., Lu, T., Luo, P., and Shao, L. (2021, January 11\u201317). Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions. Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00061"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Wang, L., Li, R., Wang, D., Duan, C., Wang, T., and Meng, X. (2021). Transformer Meets Convolution: A Bilateral Awareness Network for Semantic Segmentation of Very Fine Resolution Urban Scene Images. Remote Sens., 13.","DOI":"10.3390\/rs13163065"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Bazi, Y., Bashmal, L., Rahhal, M.M.A., Dayil, R.A., and Ajlan, N.A. (2021). Vision Transformers for Remote Sensing Image Classification. Remote Sens., 13.","DOI":"10.3390\/rs13030516"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Zhang, J., Zhao, H., and Li, J. (2021). TRS: Transformers for Remote Sensing Scene Classification. Remote Sens., 13.","DOI":"10.3390\/rs13204143"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"4408715","DOI":"10.1109\/TGRS.2022.3144165","article-title":"Swin Transformer Embedding UNet for Remote Sensing Image Semantic Segmentation","volume":"60","author":"He","year":"2022","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Qiu, H., Hou, B., Ren, B., and Zhang, X. (2022). Spatio-Temporal Tuples Transformer for Skeleton-Based Action Recognition. arXiv.","DOI":"10.1016\/j.neucom.2022.10.084"},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"246","DOI":"10.1109\/TCDS.2020.3048883","article-title":"Trear: Transformer-Based RGB-D Egocentric Action Recognition","volume":"14","author":"Li","year":"2022","journal-title":"IEEE Trans. Cogn. Dev. Syst."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Bai, R., Li, M., Meng, B., Li, F., Ren, J., Jiang, M., and Sun, D. (2022). GCsT: Graph Convolutional Skeleton Transformer for Action Recognition. arXiv.","DOI":"10.1109\/ICME52920.2022.9859781"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"1452","DOI":"10.1109\/TPAMI.2017.2723009","article-title":"Places: A 10 million image database for scene recognition","volume":"40","author":"Zhou","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L., Li, K., and Li, F. (2009, January 20\u201325). ImageNet: A large-scale hierarchical image database. Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_42","unstructured":"Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (2019, January 3\u20135). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Minneapolis, MN, USA."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Feng, C., Zhong, Y., and Huang, W. (2021, January 11\u201317). Exploring Classification Equilibrium in Long-Tailed Object Detection. Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00340"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Zhang, N., Paluri, M., Taigman, Y., Fergus, R., and Bourdev, L. (2015, January 7\u201312). Beyond frontal faces: Improving person recognition using multiple cues. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7299113"},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"187","DOI":"10.1037\/0033-2909.126.2.187","article-title":"Acquisition of the algorithms of social life: A domain-based approach","volume":"126","author":"Bugental","year":"2000","journal-title":"Psychol. Bull."},{"key":"ref_46","unstructured":"Kingma, D., and Ba, J. (2015, January 7\u20139). Adam: A method for stochastic optimization. Proceedings of the 3rd International Conference on Learning Representations (ICLR), San Diego, CA, USA."},{"key":"ref_47","unstructured":"Li, Y., Zemel, R., Brockschmidt, M., and Tarlow, D. (2016, January 2\u20134). Gated Graph Sequence Neural Networks. Proceedings of the 4th International Conference on Learning Representation (ICLR), San Juan, Puerto Rico."},{"key":"ref_48","unstructured":"Kipf, T.N., and Welling, M. (2017, January 24\u201326). Semi-supervised classification with graph convolutional networks. Proceedings of the 5th International Conference on Learning Representation (ICLR), Toulon, France."},{"key":"ref_49","unstructured":"Veli\u010dkovi\u0107, P., Preixens, G.C., Paga, A.C., Romero, A., Li\u00f2, P., and Bengio, Y. (May, January 30). Graph attention networks. Proceedings of the International Conference on Learning Representations (ICLR), Vancouver, BC, Canada."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/15\/5749\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T00:00:47Z","timestamp":1760140847000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/15\/5749"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,8,1]]},"references-count":49,"journal-issue":{"issue":"15","published-online":{"date-parts":[[2022,8]]}},"alternative-id":["s22155749"],"URL":"https:\/\/doi.org\/10.3390\/s22155749","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,8,1]]}}}