{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,9]],"date-time":"2026-01-09T17:15:11Z","timestamp":1767978911331,"version":"3.49.0"},"publisher-location":"New York, NY, USA","reference-count":33,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,9,24]],"date-time":"2021-09-24T00:00:00Z","timestamp":1632441600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,9,24]]},"DOI":"10.1145\/3488933.3488941","type":"proceedings-article","created":{"date-parts":[[2022,2,25]],"date-time":"2022-02-25T11:36:59Z","timestamp":1645789019000},"page":"374-379","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["An Efficient Image Captioning Method Based on Generative Adversarial Networks"],"prefix":"10.1145","author":[{"given":"Lin","family":"Zhenxian","sequence":"first","affiliation":[{"name":"School of Science,Xi'an University of Posts and Telecommunications, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Feng","family":"Feirong","sequence":"additional","affiliation":[{"name":"School of Telecommunication and Information Engineering,Xi'an University of Posts and Telecommunications, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yang","family":"Xiaobao","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology,Northwestern Polytechnical University,China., China and School of Computer Science and Technology,Xi'an University of Posts and Telecommunications, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ding","family":"Chen","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology,Xi'an University of Posts and Telecommunications,China, China and Shaanxi Key Laboratory of Network Data Analysis and Intelligent Processing,Xi'an University of Posts and Telecommunications, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,2,25]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"crossref","unstructured":"Cao D. Zhu M. & Gao L. 2019. An image caption method based on object detection.\u00a0Multimedia Tools and Applications 78 35329 - 35350. https:\/\/doi.org\/10.1007\/s11042-019-08116-9.  Cao D. Zhu M. & Gao L. 2019. An image caption method based on object detection.\u00a0Multimedia Tools and Applications 78 35329 - 35350. https:\/\/doi.org\/10.1007\/s11042-019-08116-9.","DOI":"10.1007\/s11042-019-08116-9"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00435"},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.3390\/app9102024"},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"crossref","unstructured":"LeCun Y. Boser B. Denker J. Henderson D. Howard R. Hubbard W. & Jackel L. 1989. Backpropagation Applied to Handwritten Zip Code Recognition.\u00a0Neural Computation 1 541-551. https:\/\/doi.org\/10.1162\/neco.1989.1.4.541  LeCun Y. Boser B. Denker J. Henderson D. Howard R. Hubbard W. & Jackel L. 1989. Backpropagation Applied to Handwritten Zip Code Recognition.\u00a0Neural Computation 1 541-551. https:\/\/doi.org\/10.1162\/neco.1989.1.4.541","DOI":"10.1162\/neco.1989.1.4.541"},{"key":"e_1_3_2_1_5_1","volume-title":"Recurrent Neural Network Based Language Modeling in Meeting Recognition.\u00a0INTERSPEECH. https:\/\/arxiv.org\/abs\/1409","author":"Kombrink S.","year":"2011","unstructured":"Kombrink , S. , Mikolov , T. , Karafi\u00e1t , M. , & Burget , L. 2011 . Recurrent Neural Network Based Language Modeling in Meeting Recognition.\u00a0INTERSPEECH. https:\/\/arxiv.org\/abs\/1409 .2329 Kombrink, S., Mikolov, T., Karafi\u00e1t, M., & Burget, L. 2011. Recurrent Neural Network Based Language Modeling in Meeting Recognition.\u00a0INTERSPEECH. https:\/\/arxiv.org\/abs\/1409.2329"},{"key":"e_1_3_2_1_6_1","unstructured":"Mirza M. & Osindero S. 2014. Conditional Generative Adversarial Nets.\u00a0ArXiv abs\/1411.1784. https:\/\/arxiv.org\/pdf\/1411.1784  Mirza M. & Osindero S. 2014. Conditional Generative Adversarial Nets.\u00a0ArXiv abs\/1411.1784. https:\/\/arxiv.org\/pdf\/1411.1784"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"crossref","unstructured":"Dai B. Fidler S. Urtasun R. & Lin D. 2017. Towards Diverse and Natural Image Descriptions via a Conditional GAN.\u00a02017 IEEE International Conference on Computer Vision (ICCV) 2989-2998. https:\/\/arxiv.org\/pdf\/1703.06029  Dai B. Fidler S. Urtasun R. & Lin D. 2017. Towards Diverse and Natural Image Descriptions via a Conditional GAN.\u00a02017 IEEE International Conference on Computer Vision (ICCV) 2989-2998. https:\/\/arxiv.org\/pdf\/1703.06029","DOI":"10.1109\/ICCV.2017.323"},{"key":"e_1_3_2_1_8_1","unstructured":"Mao J. Xu W. Yang Y. Wang J. & Yuille A. 2014. Explain Images with Multimodal Recurrent Neural Networks. ArXiv abs\/1410.1090. https:\/\/arxiv.org\/pdf\/1410.1090.pdf  Mao J. Xu W. Yang Y. Wang J. & Yuille A. 2014. Explain Images with Multimodal Recurrent Neural Networks. ArXiv abs\/1410.1090. https:\/\/arxiv.org\/pdf\/1410.1090.pdf"},{"key":"e_1_3_2_1_9_1","unstructured":"Xu K. Ba J. Kiros R. Cho K. Courville A.C. Salakhutdinov R. Zemel R. & Bengio Y. 2015. Show Attend and Tell: Neural Image Caption Generation with Visual Attention.\u00a0ICML. https:\/\/arxiv.org\/pdf\/1502.03044  Xu K. Ba J. Kiros R. Cho K. Courville A.C. Salakhutdinov R. Zemel R. & Bengio Y. 2015. Show Attend and Tell: Neural Image Caption Generation with Visual Attention.\u00a0ICML. https:\/\/arxiv.org\/pdf\/1502.03044"},{"key":"e_1_3_2_1_10_1","volume-title":"Generating Diverse and Accurate Captions Based on Generative Adversarial Network. 2019 IEEE 4th International Conference on Image, Vision and Computing (ICIVC), 331-335","author":"Sun Z.","year":"2019","unstructured":"Sun , Z. , Wang , Y. , & Zhou , W. 2019 . Generating Diverse and Accurate Captions Based on Generative Adversarial Network. 2019 IEEE 4th International Conference on Image, Vision and Computing (ICIVC), 331-335 . https:\/\/ieeexplore.ieee.org\/document\/8981009 Sun, Z., Wang, Y., & Zhou, W. 2019. Generating Diverse and Accurate Captions Based on Generative Adversarial Network. 2019 IEEE 4th International Conference on Image, Vision and Computing (ICIVC), 331-335. https:\/\/ieeexplore.ieee.org\/document\/8981009"},{"key":"e_1_3_2_1_11_1","volume-title":"Encoder-Decoder Architecture for Image Caption Generation. 2020 3rd International Conference on Communication System, Computing and IT Applications (CSCITA), 174-179","author":"Parikh H.","year":"2020","unstructured":"Parikh , H. , Sawant , H. , Parmar , B. , Shah , R. , Chapaneri , S.V. , & Jayaswal , D. 2020 . Encoder-Decoder Architecture for Image Caption Generation. 2020 3rd International Conference on Communication System, Computing and IT Applications (CSCITA), 174-179 .https:\/\/ieeexplore.ieee.org\/document\/9137802 Parikh, H., Sawant, H., Parmar, B., Shah, R., Chapaneri, S.V., & Jayaswal, D. 2020. Encoder-Decoder Architecture for Image Caption Generation. 2020 3rd International Conference on Communication System, Computing and IT Applications (CSCITA), 174-179.https:\/\/ieeexplore.ieee.org\/document\/9137802"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"crossref","unstructured":"Wang D. Hu H. & Chen D. 2020. Transformer with sparse self-attention mechanism for image captioning.\u00a0Electronics Letters 56 764-766. https:\/\/doi.org\/10.1049\/el.2020.0635  Wang D. Hu H. & Chen D. 2020. Transformer with sparse self-attention mechanism for image captioning.\u00a0Electronics Letters 56 764-766. https:\/\/doi.org\/10.1049\/el.2020.0635","DOI":"10.1049\/el.2020.0635"},{"key":"e_1_3_2_1_13_1","unstructured":"Goodfellow I.J. Pouget-Abadie J. Mirza M. Xu B. Warde-Farley D. Ozair S. Courville A.C. & Bengio Y. 2014. Generative Adversarial Nets.\u00a0NIPS. https:\/\/arxiv.org\/abs\/1701.00160  Goodfellow I.J. Pouget-Abadie J. Mirza M. Xu B. Warde-Farley D. Ozair S. Courville A.C. & Bengio Y. 2014. Generative Adversarial Nets.\u00a0NIPS. https:\/\/arxiv.org\/abs\/1701.00160"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"crossref","unstructured":"He K. Zhang X. Ren S. & Sun J. 2016. Deep Residual Learning for Image Recognition.\u00a02016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) 770-778. https:\/\/arxiv.org\/pdf\/1512.03385  He K. Zhang X. Ren S. & Sun J. 2016. Deep Residual Learning for Image Recognition.\u00a02016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) 770-778. https:\/\/arxiv.org\/pdf\/1512.03385","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_1_15_1","unstructured":"Tan M. & Le Q.V. 2019. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks.\u00a0ArXiv abs\/1905.11946. https:\/\/arxiv.org\/pdf\/1905.11946  Tan M. & Le Q.V. 2019. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks.\u00a0ArXiv abs\/1905.11946. https:\/\/arxiv.org\/pdf\/1905.11946"},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"crossref","unstructured":"Cho K. Merrienboer B.V. G\u00fcl\u00e7ehre \u00c7. Bahdanau D. Bougares F. Schwenk H. & Bengio Y. 2014. Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation.\u00a0ArXiv abs\/1406.1078. https:\/\/arxiv.org\/pdf\/1502.03044  Cho K. Merrienboer B.V. G\u00fcl\u00e7ehre \u00c7. Bahdanau D. Bougares F. Schwenk H. & Bengio Y. 2014. Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation.\u00a0ArXiv abs\/1406.1078. https:\/\/arxiv.org\/pdf\/1502.03044","DOI":"10.3115\/v1\/D14-1179"},{"key":"e_1_3_2_1_17_1","unstructured":"Vaswani A. Shazeer N. Parmar N. Uszkoreit J. Jones L. Gomez A.N. Kaiser L. & Polosukhin I. 2017. Attention is All you Need.\u00a0ArXiv abs\/1706.03762. https:\/\/arxiv.org\/pdf\/1706.03762  Vaswani A. Shazeer N. Parmar N. Uszkoreit J. Jones L. Gomez A.N. Kaiser L. & Polosukhin I. 2017. Attention is All you Need.\u00a0ArXiv abs\/1706.03762. https:\/\/arxiv.org\/pdf\/1706.03762"},{"key":"e_1_3_2_1_18_1","volume-title":"Gradient Centralization: A New Optimization Technique for Deep Neural Networks.\u00a0ECCV. https:\/\/arxiv.org\/pdf\/2004.01461","author":"Yong H.","year":"2020","unstructured":"Yong , H. , Huang , J. , Hua , X. , & Zhang , L. 2020 . Gradient Centralization: A New Optimization Technique for Deep Neural Networks.\u00a0ECCV. https:\/\/arxiv.org\/pdf\/2004.01461 Yong, H., Huang, J., Hua, X., & Zhang, L. 2020. Gradient Centralization: A New Optimization Technique for Deep Neural Networks.\u00a0ECCV. https:\/\/arxiv.org\/pdf\/2004.01461"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"crossref","unstructured":"Papineni K. Roukos S. Ward T. & Zhu W. 2002. Bleu: a Method for Automatic Evaluation of Machine Translation.\u00a0ACL. https:\/\/www.aclweb.org\/anthology\/P02-1040  Papineni K. Roukos S. Ward T. & Zhu W. 2002. Bleu: a Method for Automatic Evaluation of Machine Translation.\u00a0ACL. https:\/\/www.aclweb.org\/anthology\/P02-1040","DOI":"10.3115\/1073083.1073135"},{"key":"e_1_3_2_1_20_1","unstructured":"Simonyan K. & Zisserman A. 2015. Very Deep Convolutional Networks for Large-Scale Image Recognition.\u00a0CoRR abs\/1409.1556. https:\/\/arxiv.org\/abs\/1409.1556  Simonyan K. & Zisserman A. 2015. Very Deep Convolutional Networks for Large-Scale Image Recognition.\u00a0CoRR abs\/1409.1556. https:\/\/arxiv.org\/abs\/1409.1556"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"crossref","unstructured":"Ren S. He K. Girshick R.B. & Sun J. 2015. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks.\u00a0IEEE Transactions on Pattern Analysis and Machine Intelligence 39 1137-1149. https:\/\/arxiv.org\/abs\/1506.01497  Ren S. He K. Girshick R.B. & Sun J. 2015. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks.\u00a0IEEE Transactions on Pattern Analysis and Machine Intelligence 39 1137-1149. https:\/\/arxiv.org\/abs\/1506.01497","DOI":"10.1109\/TPAMI.2016.2577031"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"crossref","unstructured":"Hochreiter S. & Schmidhuber J. 1997. Long Short-Term Memory.\u00a0Neural Computation 9 1735-1780. https:\/\/doi.org\/10.1162\/neco.1997.9.8.1735  Hochreiter S. & Schmidhuber J. 1997. Long Short-Term Memory.\u00a0Neural Computation 9 1735-1780. https:\/\/doi.org\/10.1162\/neco.1997.9.8.1735","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"crossref","unstructured":"Li J. Yao P. Guo L. & Zhang W. 2019. Boosted Transformer for Image Captioning.\u00a0Applied Sciences 9 3260. https:\/\/arxiv.org\/pdf\/1207.07258  Li J. Yao P. Guo L. & Zhang W. 2019. Boosted Transformer for Image Captioning.\u00a0Applied Sciences 9 3260. https:\/\/arxiv.org\/pdf\/1207.07258","DOI":"10.3390\/app9163260"},{"key":"e_1_3_2_1_24_1","volume-title":"Image Captioning: Transforming Objects into Words.\u00a0NeurIPS. https:\/\/arxiv.org\/pdf\/1906.05963","author":"Herdade S.","year":"2019","unstructured":"Herdade , S. , Kappeler , A. , Boakye , K. , & Soares , J. 2019 . Image Captioning: Transforming Objects into Words.\u00a0NeurIPS. https:\/\/arxiv.org\/pdf\/1906.05963 Herdade, S., Kappeler, A., Boakye, K., & Soares, J. 2019. Image Captioning: Transforming Objects into Words.\u00a0NeurIPS. https:\/\/arxiv.org\/pdf\/1906.05963"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"crossref","unstructured":"Vinyals O. Toshev A. Bengio S. & Erhan D. 2015. Show and tell: A neural image caption generator.\u00a02015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) 3156-3164. https:\/\/arxiv.org\/pdf\/1411.4555  Vinyals O. Toshev A. Bengio S. & Erhan D. 2015. Show and tell: A neural image caption generator.\u00a02015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) 3156-3164. https:\/\/arxiv.org\/pdf\/1411.4555","DOI":"10.1109\/CVPR.2015.7298935"},{"key":"e_1_3_2_1_26_1","unstructured":"Schwartz I. Schwing A.G. & Hazan T. 2017. High-Order Attention Models for Visual Question Answering.\u00a0NIPS. https:\/\/arxiv.org\/pdf\/1711.04323  Schwartz I. Schwing A.G. & Hazan T. 2017. High-Order Attention Models for Visual Question Answering.\u00a0NIPS. https:\/\/arxiv.org\/pdf\/1711.04323"},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"crossref","unstructured":"Wei Y. Wang L. Cao H. Shao M. & Wu C. 2020. Multi-Attention Generative Adversarial Network for image captioning.\u00a0Neurocomputing 387 91-99. https:\/\/doi.org\/10.1016\/j.neucom.2019.12.073  Wei Y. Wang L. Cao H. Shao M. & Wu C. 2020. Multi-Attention Generative Adversarial Network for image captioning.\u00a0Neurocomputing 387 91-99. https:\/\/doi.org\/10.1016\/j.neucom.2019.12.073","DOI":"10.1016\/j.neucom.2019.12.073"},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.3390\/app8050739"},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"crossref","unstructured":"Yu L. Zhang W. Wang J. & Yu Y. 2017. SeqGAN: Sequence Generative Adversarial Nets with Policy Gradient. AAAI.https:\/\/arxiv.org\/pdf\/1609.05473  Yu L. Zhang W. Wang J. & Yu Y. 2017. SeqGAN: Sequence Generative Adversarial Nets with Policy Gradient. AAAI.https:\/\/arxiv.org\/pdf\/1609.05473","DOI":"10.1609\/aaai.v31i1.10804"},{"key":"e_1_3_2_1_30_1","volume-title":"MSCap: Multi-Style Image Captioning With Unpaired Stylized Text. 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 4199-4208","author":"Guo L.","year":"2019","unstructured":"Guo , L. , Liu , J. , Yao , P. , Li , J. , & Lu , H. 2019 . MSCap: Multi-Style Image Captioning With Unpaired Stylized Text. 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 4199-4208 . https:\/\/dblp.org\/rec\/conf\/cvpr\/GuoLYLL19 Guo, L., Liu, J., Yao, P., Li, J., & Lu, H. 2019. MSCap: Multi-Style Image Captioning With Unpaired Stylized Text. 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 4199-4208. https:\/\/dblp.org\/rec\/conf\/cvpr\/GuoLYLL19"},{"key":"e_1_3_2_1_31_1","volume-title":"Adapt and Tell: Adversarial Training of Cross-Domain Image Captioner. 2017 IEEE International Conference on Computer Vision (ICCV), 521-530","author":"Chen T.","year":"2017","unstructured":"Chen , T. , Liao , Y. , Chuang , C. , Hsu , W. , Fu , J. , & Sun , M. 2017 . Show , Adapt and Tell: Adversarial Training of Cross-Domain Image Captioner. 2017 IEEE International Conference on Computer Vision (ICCV), 521-530 . https:\/\/arxiv.org\/pdf\/1705.00930.pdf Chen, T., Liao, Y., Chuang, C., Hsu, W., Fu, J., & Sun, M. 2017. Show, Adapt and Tell: Adversarial Training of Cross-Domain Image Captioner. 2017 IEEE International Conference on Computer Vision (ICCV), 521-530. https:\/\/arxiv.org\/pdf\/1705.00930.pdf"},{"key":"e_1_3_2_1_32_1","unstructured":"Chen C. Mu S. Xiao W. Ye Z. Wu L. Ma F. & Ju Q. 2019. Improving Image Captioning with Conditional Generative Adversarial Nets.\u00a0AAAI.https:\/\/arxiv.org\/pdf\/1805.07112.pdf  Chen C. Mu S. Xiao W. Ye Z. Wu L. Ma F. & Ju Q. 2019. Improving Image Captioning with Conditional Generative Adversarial Nets.\u00a0AAAI.https:\/\/arxiv.org\/pdf\/1805.07112.pdf"},{"key":"e_1_3_2_1_33_1","unstructured":"Sutton R. McAllester D.A. Singh S. & Mansour Y. 1999. Policy Gradient Methods for Reinforcement Learning with Function Approximation. NIPS. https:\/\/dl.acm.org\/doi\/10.5555\/3009657.3009806  Sutton R. McAllester D.A. Singh S. & Mansour Y. 1999. Policy Gradient Methods for Reinforcement Learning with Function Approximation. NIPS. https:\/\/dl.acm.org\/doi\/10.5555\/3009657.3009806"}],"event":{"name":"AIPR 2021: 2021 4th International Conference on Artificial Intelligence and Pattern Recognition","location":"Xiamen China","acronym":"AIPR 2021"},"container-title":["2021 4th International Conference on Artificial Intelligence and Pattern Recognition"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3488933.3488941","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3488933.3488941","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:31:28Z","timestamp":1750188688000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3488933.3488941"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,9,24]]},"references-count":33,"alternative-id":["10.1145\/3488933.3488941","10.1145\/3488933"],"URL":"https:\/\/doi.org\/10.1145\/3488933.3488941","relation":{},"subject":[],"published":{"date-parts":[[2021,9,24]]},"assertion":[{"value":"2022-02-25","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}