{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T02:27:41Z","timestamp":1760236061886,"version":"build-2065373602"},"reference-count":43,"publisher":"MDPI AG","issue":"21","license":[{"start":{"date-parts":[[2021,10,28]],"date-time":"2021-10-28T00:00:00Z","timestamp":1635379200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"the Innovation Project of Shanghai Institute of Technical Physics of The Chinese Acade-my of Sciences","award":["CX-70"],"award-info":[{"award-number":["CX-70"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>In this paper, we provide external image features and use the internal attention mechanism to solve the VQA problem given a dataset of textual questions and related images. Most previous models for VQA use a pair of images and questions as input. In addition, the model adopts a question-oriented attention mechanism to extract the features of the entire image and then perform feature fusion. However, the shortcoming of these models is that they cannot effectively eliminate the irrelevant features of the image. In addition, the problem-oriented attention mechanism lacks in the mining of image features, which will bring in redundant image features. In this paper, we propose a VQA model based on adversarial learning and bidirectional attention. We exploit external image features that are not related to the question to form an adversarial mechanism to boost the accuracy of the model. Target detection is performed on the image\u2014that is, the image-oriented attention mechanism. The bidirectional attention mechanism is conducive to promoting model attention and eliminating interference. Experimental results are evaluated on benchmark datasets, and our model performs better than other models based on attention methods. In addition, the qualitative results show the attention maps on the images and leads to predicting correct answers.<\/jats:p>","DOI":"10.3390\/s21217164","type":"journal-article","created":{"date-parts":[[2021,10,28]],"date-time":"2021-10-28T23:52:35Z","timestamp":1635465155000},"page":"7164","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["Adversarial Learning with Bidirectional Attention for Visual Question Answering"],"prefix":"10.3390","volume":"21","author":[{"given":"Qifeng","family":"Li","sequence":"first","affiliation":[{"name":"Shanghai Institute of Technical Physics of the Chinese Academy of Sciences, Shanghai 200083, China"},{"name":"School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences, Beijing 100049, China"},{"name":"Key Laboratory of Infrared System Detection and Imaging Technology, Chinese Academy of Sciences, Shanghai 200083, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xinyi","family":"Tang","sequence":"additional","affiliation":[{"name":"Shanghai Institute of Technical Physics of the Chinese Academy of Sciences, Shanghai 200083, China"},{"name":"Key Laboratory of Infrared System Detection and Imaging Technology, Chinese Academy of Sciences, Shanghai 200083, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yi","family":"Jian","sequence":"additional","affiliation":[{"name":"Shanghai Institute of Technical Physics of the Chinese Academy of Sciences, Shanghai 200083, China"},{"name":"Key Laboratory of Infrared System Detection and Imaging Technology, Chinese Academy of Sciences, Shanghai 200083, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,10,28]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"6730","DOI":"10.1109\/TIP.2021.3097180","article-title":"Re-Attention for Visual Question Answering","volume":"30","author":"Guo","year":"2021","journal-title":"IEEE Trans. Image Process."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Zheng, X., Wang, B., Du, X., and Lu, X. (2021). Mutual Attention Inception Network for Remote Sensing Visual Question Answering. IEEE Trans. Geosci. Remote. Sens., 1\u201314.","DOI":"10.1109\/TGRS.2021.3079918"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Rahman, T., Chou, S.H., Sigal, L., and Carenini, G. (2021, January 21\u201324). An Improved Attention for Visual Question Answering. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPRW53098.2021.00181"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.inffus.2021.02.022","article-title":"Multimodal feature-wise co-attention method for visual question answering","volume":"73","author":"Zhang","year":"2021","journal-title":"Inf. Fusion"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Teney, D., Anderson, P., He, X., and Van Den Hengel, A. (2018, January 18\u201322). Tips and tricks for visual question answering: Learnings from the 2017 challenge. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00444"},{"key":"ref_6","unstructured":"Liu, Y., Zhang, X., Zhao, Z., Zhang, B., Cheng, L., and Li, Z. (2021). ALSA: Adversarial Learning of Supervised Attentions for Visual Question Answering. IEEE Trans. Cybern., 1\u201314."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"3894","DOI":"10.1109\/TNNLS.2020.3016083","article-title":"Adversarial Learning With Multi-Modal Attention for Visual Question Answering","volume":"32","author":"Liu","year":"2021","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Tang, R., Ma, C., Zhang, W.E., Wu, Q., and Yang, X. (2020). Semantic equivalent adversarial data augmentation for visual question answering. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-030-58529-7_26"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Ilievski, I., and Feng, J. (2017, January 23\u201327). Generative attention model with adversarial self-learning for visual question answering. Proceedings of the on Thematic Workshops of ACM Multimedia, Mountain View, CA, USA.","DOI":"10.1145\/3126686.3126695"},{"key":"ref_10","first-page":"289","article-title":"Hierarchical question-image co-attention for visual question answering","volume":"29","author":"Lu","year":"2016","journal-title":"Adv. Neural Inf. Process.Syst."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Shih, K.J., Singh, S., and Hoiem, D. (2016, January 27\u201330). Where to look: Focus regions for visual question answering. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.499"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Yang, Z., He, X., Gao, J., Deng, L., and Smola, A. (2016, January 27\u201330). Stacked attention networks for image question answering. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.10"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Xu, H., and Saenko, K. (2016). Ask, Attend and Answer Exploring Question-Guided Spatial Attention for Visual Question Answering. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-46478-7_28"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Shi, Y., Furlanello, T., Zha, S., and Anandkumar, A. (2018, January 8\u201314). Question Type Guided Attention in Visual Question Answering. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01225-0_10"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Pennington, J., Socher, R., and Manning, C.D. (2014, January 25\u201329). Glove: Global vectors for word representation. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar.","DOI":"10.3115\/v1\/D14-1162"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"50203","DOI":"10.1007\/s11432-018-9606-6","article-title":"A glove-based system for object recognition via visual-tactile fusion","volume":"62","author":"Fang","year":"2019","journal-title":"Sci. China Inf. Sci."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"42","DOI":"10.1016\/j.neucom.2018.03.030","article-title":"Face detection using deep learning: An improved faster RCNN approach","volume":"299","author":"Sun","year":"2018","journal-title":"Neurocomputing"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Sabokrou, M., Khalooei, M., Fathy, M., and Adeli, E. (2018, January 18\u201323). Adversarially Learned One-Class Classifier for Novelty Detection. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00356"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Liang, J., Cao, Y., Zhang, C., Chang, S., Bai, K., and Xu, Z. (2019, January 18\u201323). Additive adversarial learning for unbiased authentication. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2019.01169"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Jaiswal, A., Wu, Y., AbdAlmageed, W., Masi, I., and Natarajan, P. (2019, January 18\u201323). Aird: Adversarial learning framework for image repurposing detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2019.01159"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Agresti, G., Schaefer, H., Sartor, P., and Zanuttigh, P. (2019, January 18\u201323). Unsupervised domain adaptation for tof data denoising with adversarial learning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2019.00573"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"85","DOI":"10.1016\/j.inffus.2019.07.005","article-title":"Infrared and visible image fusion via detail preserving adversarial learning","volume":"54","author":"Ma","year":"2020","journal-title":"Inf. Fusion"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Chen, B.-C., and Kae, A. (2019, January 16\u201318). Toward Realistic Image Compositing with Adversarial Learning. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR.2019.00861"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Wu, Q., Wang, P., Shen, C., Reid, I., and Hengel, A.V.D. (2018, January 18\u201323). Are You Talking to Me? Reasoned Visual Dialog Generation Through Adversarial Learning. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00639"},{"key":"ref_25","first-page":"2589","article-title":"Auto-classification of insect images based on color histogram and GLCM","volume":"6","author":"Zhu","year":"2010","journal-title":"Proceedings 2010 Seventh International Conference on Fuzzy Systems and Knowledge Discovery, Yantai, China, 10\u201312 August 2010"},{"key":"ref_26","unstructured":"Huang, J., Kumar, S.R., Mitra, M., Zhu, W.J., and Zabih, R. (1997, January 17\u201319). Image indexing using color correlograms. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, San Juan, Puerto Rico."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Chen, D., Hua, G., Wen, F., and Sun, J. (2016). Supervised transformer network for efficient face detection. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-46454-1_8"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Jiang, H., and Learned-Miller, E. (June, January 30). Face detection with the faster R-CNN. Proceedings of the 2017 12th IEEE International Conference on Automatic Face Gesture Recognition, FG 2017, Washington, DC, USA.","DOI":"10.1109\/FG.2017.82"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Hara, K., Saito, D., and Shouno, H. (2015, January 12\u201316). Analysis of function of rectified linear unit used in deep learning. Proceedings of the 2015 International Joint Conference on Neural Networks, IJCNN, Killarney, Ireland.","DOI":"10.1109\/IJCNN.2015.7280578"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Yu, Z., Yu, J., Fan, J., and Tao, D. (2017, January 22\u201329). Multi-modal factorized bilinear pooling with co-attention learning for visual question answering. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.202"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Yuan, B. (2016, January 6\u20139). Efficient hardware architecture of softmax layer in deep neural network. Proceedings of the 2016 29th IEEE International System-on-Chip Conference, SOCC, Seattle, WA, USA.","DOI":"10.1109\/SOCC.2016.7905501"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Ramos, D., Franco-Pedroso, J., Lozano-Diez, A., and Gonzalez-Rodriguez, J. (2018). Deconstructing cross-entropy for probabilistic binary classifiers. Entropy, 20.","DOI":"10.3390\/e20030208"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D. (2017, January 21\u201326). Making the v in vqa matter: Elevating the role of image understanding in visual question answering. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.670"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., and Zhang, L. (2018, January 18\u201322). Bottom-up and top-down attention for image captioning and visual question answering. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00636"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C.L., and Parikh, D. (2015, January 7\u201313). Vqa: Visual question answering. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.279"},{"key":"ref_36","unstructured":"Ilievski, I., Yan, S., and Feng, J. (2016). A focused dynamic attention model for visual question answering. arXiv."},{"key":"ref_37","unstructured":"Noh, H., and Han, B. (2016). Training recurrent answering units with joint loss minimization for vqa. arXiv, preprint."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Nam, H., Ha, J.W., and Kim, J. (2017, January 21\u201326). Dual attention networks for multimodal reasoning and matching. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.232"},{"key":"ref_39","unstructured":"Kim, J.H., On, K.W., Lim, W., Kim, J., Ha, J.W., and Zhang, B.T. (2016). Hadamard product for low-rank bilinear pooling. arXiv, preprint."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Nguyen, D.K., and Okatani, T. (2018, January 18\u201322). Improved fusion of visual and language representations by dense symmetric co-attention for visual question answering. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00637"},{"key":"ref_41","unstructured":"Zhang, Y., Hare, J., and Pr\u00fcgel-Bennett, A. (2018). Learning to count objects in natural images for visual question answering. arXiv, preprint."},{"key":"ref_42","unstructured":"Li, L.H., Yatskar, M., Yin, D., Hsieh, C.J., and Chang, K.W. (2019). Visualbert: A simple and performant baseline for vision and language. arXiv, preprint."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Wang, T., Huang, J., Zhang, H., and Sun, Q. (2020, January 14\u201319). Visual commonsense r-cnn. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01077"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/21\/7164\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T07:22:03Z","timestamp":1760167323000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/21\/7164"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,10,28]]},"references-count":43,"journal-issue":{"issue":"21","published-online":{"date-parts":[[2021,11]]}},"alternative-id":["s21217164"],"URL":"https:\/\/doi.org\/10.3390\/s21217164","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2021,10,28]]}}}