{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,17]],"date-time":"2026-03-17T22:42:02Z","timestamp":1773787322909,"version":"3.50.1"},"reference-count":26,"publisher":"MDPI AG","issue":"8","license":[{"start":{"date-parts":[[2022,8,11]],"date-time":"2022-08-11T00:00:00Z","timestamp":1660176000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"the National Natural Science Foundation of China","award":["62072350"],"award-info":[{"award-number":["62072350"]}]},{"name":"the National Natural Science Foundation of China","award":["62171328"],"award-info":[{"award-number":["62171328"]}]},{"name":"the National Natural Science Foundation of China","award":["2019AAA045"],"award-info":[{"award-number":["2019AAA045"]}]},{"name":"the National Natural Science Foundation of China","award":["2018ZYYD059"],"award-info":[{"award-number":["2018ZYYD059"]}]},{"name":"the National Natural Science Foundation of China","award":["202001602011971"],"award-info":[{"award-number":["202001602011971"]}]},{"name":"Hubei Technology Innovation Project","award":["62072350"],"award-info":[{"award-number":["62072350"]}]},{"name":"Hubei Technology Innovation Project","award":["62171328"],"award-info":[{"award-number":["62171328"]}]},{"name":"Hubei Technology Innovation Project","award":["2019AAA045"],"award-info":[{"award-number":["2019AAA045"]}]},{"name":"Hubei Technology Innovation Project","award":["2018ZYYD059"],"award-info":[{"award-number":["2018ZYYD059"]}]},{"name":"Hubei Technology Innovation Project","award":["202001602011971"],"award-info":[{"award-number":["202001602011971"]}]},{"name":"the Central Government Guides Local Science and Technology Development Special Projects","award":["62072350"],"award-info":[{"award-number":["62072350"]}]},{"name":"the Central Government Guides Local Science and Technology Development Special Projects","award":["62171328"],"award-info":[{"award-number":["62171328"]}]},{"name":"the Central Government Guides Local Science and Technology Development Special Projects","award":["2019AAA045"],"award-info":[{"award-number":["2019AAA045"]}]},{"name":"the Central Government Guides Local Science and Technology Development Special Projects","award":["2018ZYYD059"],"award-info":[{"award-number":["2018ZYYD059"]}]},{"name":"the Central Government Guides Local Science and Technology Development Special Projects","award":["202001602011971"],"award-info":[{"award-number":["202001602011971"]}]},{"name":"the High value Intellectual Property Cultivation Project of Hubei Province, the Enterprise Technology Innovation Project of Wuhan","award":["62072350"],"award-info":[{"award-number":["62072350"]}]},{"name":"the High value Intellectual Property Cultivation Project of Hubei Province, the Enterprise Technology Innovation Project of Wuhan","award":["62171328"],"award-info":[{"award-number":["62171328"]}]},{"name":"the High value Intellectual Property Cultivation Project of Hubei Province, the Enterprise Technology Innovation Project of Wuhan","award":["2019AAA045"],"award-info":[{"award-number":["2019AAA045"]}]},{"name":"the High value Intellectual Property Cultivation Project of Hubei Province, the Enterprise Technology Innovation Project of Wuhan","award":["2018ZYYD059"],"award-info":[{"award-number":["2018ZYYD059"]}]},{"name":"the High value Intellectual Property Cultivation Project of Hubei Province, the Enterprise Technology Innovation Project of Wuhan","award":["202001602011971"],"award-info":[{"award-number":["202001602011971"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>Human\u2013object interaction (HOI) is a human-centered object detection task that aims to identify the interactions between persons and objects in an image. Previous end-to-end methods have used the attention mechanism of a transformer to spontaneously identify the associations between persons and objects in an image, which effectively improved detection accuracy; however, a transformer can increase computational demands and slow down detection processes. In addition, the end-to-end method can result in asymmetry between foreground and background information. The foreground data may be significantly less than the background data, while the latter consumes more computational resources without significantly improving detection accuracy. Therefore, we proposed an input-controlled transformer, \u201cratio-transformer\u201d to solve an HOI task, which could not only limit the amount of information in the input transformer by setting a sampling ratio, but also significantly reduced the computational demands while ensuring detection accuracy. The ratio-transformer consisted of a sampling module and a transformer network. The sampling module divided the input feature map into foreground versus background features. The irrelevant background features were a pooling sampler, which were then fused with the foreground features as input data for the transformer. As a result, the valid data input into the Transformer network remained constant, while irrelevant information was significantly reduced, which maintained the foreground and background information symmetry. The proposed network was able to learn the feature information of the target itself and the association features between persons and objects so it could query to obtain the complete HOI interaction triplet. The experiments on the VCOCO dataset showed that the proposed method reduced the computational demand of the transformer by 57% without any loss of accuracy, as compared to other current HOI methods.<\/jats:p>","DOI":"10.3390\/sym14081666","type":"journal-article","created":{"date-parts":[[2022,8,11]],"date-time":"2022-08-11T23:05:49Z","timestamp":1660259149000},"page":"1666","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Human\u2013Object Interaction Detection with Ratio-Transformer"],"prefix":"10.3390","volume":"14","author":[{"given":"Tianlang","family":"Wang","sequence":"first","affiliation":[{"name":"Hubei Key Laboratory of Intelligent Robot, School of Computer Science and Engineering, Wuhan Institute of Technology, Wuhan 430000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tao","family":"Lu","sequence":"additional","affiliation":[{"name":"Hubei Key Laboratory of Intelligent Robot, School of Computer Science and Engineering, Wuhan Institute of Technology, Wuhan 430000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wenhua","family":"Fang","sequence":"additional","affiliation":[{"name":"Hubei Key Laboratory of Intelligent Robot, School of Computer Science and Engineering, Wuhan Institute of Technology, Wuhan 430000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yanduo","family":"Zhang","sequence":"additional","affiliation":[{"name":"Hubei Key Laboratory of Intelligent Robot, School of Computer Science and Engineering, Wuhan Institute of Technology, Wuhan 430000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2022,8,11]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Guan, Z., Liu, K., Ma, Y., Qian, X., and Ji, T. (2018). Sequential Dual Attention: Coarse-to-Fine-Grained Hierarchical Generation for Image Captioning. Symmetry, 10.","DOI":"10.3390\/sym10110626"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S. (2020, January 23\u201328). End-to-end object detection with transformers. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Yang, W., Zhang, J., Cai, J., and Xu, Z. (2021). Relation Selective Graph Convolutional Network for Skeleton-Based Action Recognition. Symmetry, 13.","DOI":"10.3390\/sym13122275"},{"key":"ref_4","unstructured":"Gao, C., Zou, Y., and Huang, J.B. (2018). ican: Instance-centric attention network for human-object interaction detection. arXiv."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Gao, C., Xu, J., Zou, Y., and Huang, J.B. (2020, January 23\u201328). Drg: Dual relation graph for human-object interaction detection. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-58610-2_41"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Liao, Y., Liu, S., Wang, F., Chen, Y., Qian, C., and Feng, J. (2020, January 14\u201319). Ppdm: Parallel point detection and matching for real-time human-object interaction detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00056"},{"key":"ref_7","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017). Attention is all you need. Adv. Neural Inf. Process. Syst., 6000\u20136010."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Zou, C., Wang, B., Hu, Y., Liu, J., Wu, Q., Zhao, Y., Li, B., Zhang, C., Zhang, C., and Wei, Y. (2021, January 20\u201325). End-to-end human object interaction detection with hoi transformer. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01165"},{"key":"ref_9","unstructured":"Gupta, S., and Malik, J. (2015). Visual semantic role labeling. arXiv."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Gkioxari, G., Girshick, R., Doll\u00e1r, P., and He, K. (2018, January 18\u201323). Detecting and recognizing human-object interactions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00872"},{"key":"ref_11","unstructured":"Chen, J., and Yanai, K. (2021). QAHOI: Query-Based Anchors for Human-Object Interaction Detection. arXiv."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Yuan, H., Wang, M., Ni, D., and Xu, L. (2022). Detecting Human-Object Interactions with Object-Guided Cross-Modal Calibrated Semantics. arXiv.","DOI":"10.1609\/aaai.v36i3.20229"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015, January 7\u201313). Fast r-cnn. Proceedings of the IEEE International Conference on Computer Vision, Washington, DC, USA.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_14","unstructured":"Wang, S., Li, B.Z., Khabsa, M., Fang, H., and Ma, H. (2020). Linformer: Self-attention with linear complexity. arXiv."},{"key":"ref_15","unstructured":"Wu, Z., Liu, Z., Lin, J., Lin, Y., and Han, S. (2020). Lite transformer with long-short range attention. arXiv."},{"key":"ref_16","unstructured":"Zhu, X., Su, W., Lu, L., Li, B., Wang, X., and Dai, J. (2020). Deformable detr: Deformable transformers for end-to-end object detection. arXiv."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Wang, T., Yuan, L., Chen, Y., Feng, J., and Yan, S. (2021, January 10\u201317). PnP-DETR: Towards efficient visual analysis with transformers. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00462"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_19","unstructured":"Roh, B., Shin, J., Shin, W., and Kim, S. (2021). Sparse DETR: Efficient End-to-End Object Detection with Learnable Sparsity. arXiv."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Meng, D., Chen, X., Fan, Z., Zeng, G., Li, H., Yuan, Y., Sun, L., and Wang, J. (2021, January 10\u201317). Conditional detr for fast training convergence. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00363"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021, January 10\u201317). Swin transformer: Hierarchical vision transformer using shifted windows. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Chao, Y.W., Liu, Y., Liu, X., Zeng, H., and Deng, J. (2018, January 12\u201315). Learning to detect human-object interactions. Proceedings of the 2018 IEEE Winter Conference on Applications of Computer Vision (WACV), Lake Tahoe, NV, USA.","DOI":"10.1109\/WACV.2018.00048"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014, January 6\u201312). Microsoft coco: Common objects in context. Proceedings of the European Conference on Computer Vision, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Li, Y.L., Zhou, S., Huang, X., Xu, L., Ma, Z., Fang, H.S., Wang, Y., and Lu, C. (2019, January 16\u201317). Transferable interactiveness knowledge for human-object interaction detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00370"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Kolesnikov, A., Kuznetsova, A., Lampert, C., and Ferrari, V. (2019, January 27\u201328). Detecting visual relationships using box attention. Proceedings of the IEEE\/CVF International Conference on Computer Vision Workshops, Seoul, Korea.","DOI":"10.1109\/ICCVW.2019.00217"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Qi, S., Wang, W., Jia, B., Shen, J., and Zhu, S.C. (2018, January 8\u201314). Learning human-object interactions by graph parsing neural networks. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01240-3_25"}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/14\/8\/1666\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T00:07:16Z","timestamp":1760141236000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/14\/8\/1666"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,8,11]]},"references-count":26,"journal-issue":{"issue":"8","published-online":{"date-parts":[[2022,8]]}},"alternative-id":["sym14081666"],"URL":"https:\/\/doi.org\/10.3390\/sym14081666","relation":{},"ISSN":["2073-8994"],"issn-type":[{"value":"2073-8994","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,8,11]]}}}