{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,5]],"date-time":"2026-07-05T07:40:16Z","timestamp":1783237216664,"version":"3.54.6"},"reference-count":44,"publisher":"Wiley","issue":"6","license":[{"start":{"date-parts":[[2023,7,5]],"date-time":"2023-07-05T00:00:00Z","timestamp":1688515200000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/onlinelibrary.wiley.com\/termsAndConditions#vor"},{"start":{"date-parts":[[2023,7,5]],"date-time":"2023-07-05T00:00:00Z","timestamp":1688515200000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/doi.wiley.com\/10.1002\/tdm_license_1.1"}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Expert Systems"],"published-print":{"date-parts":[[2026,6]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Automatic Video captioning system is a method of describing the content in a video by analysing its visual aspects with regard to space and time and producing a meaningful caption that explains the video. A decade of research in this area has resulted in a steep growth in the quality and appropriateness of the generated caption compared with the expected result. The research has been driven from the very basic method to most advanced transformer method. Machine generated caption of a video must be adhering to many expected standards. For humans, this task may be a trivial one, however its not the same for a machine to analyse the content and generate a semantically coherent description for it. The caption which is generated in a natural language must also adhere to its lexical and syntactical structure. The video captioning process is a culmination of computer vision and natural language processing tasks. Commencing with template based conventional approach, it has surpassed statistical method, traditional deep learning approaches and is now in the trend of using transformers. This work made an extensive study of the literature and has proposed an improved transformer\u2010based architecture for video captioning process. The transformer architecture made use of an encoder and decoder model that has two and three sublayers respectively. Multi\u2010head self attention and cross attention are part of the model which bring about very beneficial results. The decoder is auto\u2010regressive and uses a masked layer to prevent the model from foreseeing future words in the caption. An enhanced encoder\u2010decoder Transformer model with CNN for feature extraction has been used in our work. This model captures the long\u2010range dependencies and temporal relationships more effectively. The model has been evaluated with benchmark datasets and compared with state\u2010of\u2010the\u2010art methods and found to be slightly better in the performance. The performance scores are slightly varying for BLEU, METEOR, ROUGE and CIDEr. Furthermore, we propose the idea of curriculum learning if incorporated can improve the results again.<\/jats:p>","DOI":"10.1111\/exsy.13392","type":"journal-article","created":{"date-parts":[[2023,7,5]],"date-time":"2023-07-05T05:13:38Z","timestamp":1688534018000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Enhanced transformer model for video caption generation"],"prefix":"10.1111","volume":"43","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3037-4225","authenticated-orcid":false,"given":"Soumya","family":"Varma","sequence":"first","affiliation":[{"name":"Department of CSE Karunya Institute of Technology and Sciences  Coimbatore India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"J. Dinesh","family":"Peter","sequence":"additional","affiliation":[{"name":"Department of CSE Karunya Institute of Technology and Sciences  Coimbatore India"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"311","published-online":{"date-parts":[[2023,7,5]]},"reference":[{"key":"e_1_2_7_2_1","unstructured":"Chen C.\u2010F. Panda R. &Fan Q.(2021).RegionViT: Regional\u2010to\u2010local attention for vision transformers. arXiv:cs.CV\/2106.02689. Retrieved fromhttps:\/\/arxiv.org\/abs\/2106.02689"},{"key":"e_1_2_7_3_1","first-page":"190","volume-title":"ACL: Human Language Technologies\u2010 Volume 1","author":"Chen D.","year":"2011"},{"key":"e_1_2_7_4_1","doi-asserted-by":"crossref","unstructured":"Chen Y. Wang S. Zhang W. &Huang Q.(2018).Less is more: Picking informative frames for video captioning. Retrieved from: arXiv Preprint arXiv:1803.01457.","DOI":"10.1007\/978-3-030-01261-8_22"},{"key":"e_1_2_7_5_1","doi-asserted-by":"crossref","unstructured":"Das P. Xu C. Doell R. F. &Corso J. J.(2013).A thousand frames in just a few words: Lingual description of videos through latent topics and sparse object stitching. IEEE CVPR.","DOI":"10.1109\/CVPR.2013.340"},{"key":"e_1_2_7_6_1","unstructured":"Devlin J. Chang M.\u2010W. Lee K. &Toutanova K.(2018).BERT: Pretraining of deep bidirectional transformers for language understanding arXiv Preprint arXiv:1810.04805."},{"key":"e_1_2_7_7_1","volume-title":"Long\u2010term RCNN for visual recognition and description","author":"Donahue J.","year":"2015"},{"key":"e_1_2_7_8_1","unstructured":"Dosovitskiy A. Beyer L. Kolesnikov A. Weissenborn D. Zhai X. Unterthiner T. Dehghani M. Minderer M. Heigold G. Gelly S. et al. (2020).An image is worth 16 \u00d7\u200916 words: Transformers for image recognition at scale. arXiv:2010.11929. Retrieved fromhttps:\/\/arxiv.org\/abs\/2010.11929"},{"key":"e_1_2_7_9_1","first-page":"1292","article-title":"Image description using visual dependency representations","volume":"13","author":"Elliott D.","year":"2013","journal-title":"Proceedings of Empirical Methods in Natural Language Processing"},{"key":"e_1_2_7_10_1","doi-asserted-by":"crossref","unstructured":"Guo Z. Gao L. Song J. Xu X. Shao J. &Shen H. T.(2016).Attentionbased lstm with semantic consistency for videos captioning. Proceedings of the 24th ACM International Conference on Multimedia. 357\u2013361.","DOI":"10.1145\/2964284.2967242"},{"key":"e_1_2_7_11_1","doi-asserted-by":"crossref","unstructured":"He K. Zhang X. Ren S. &Sun J.(2016a).Deep residual learning for image recognition. IEEE CVPR.","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_2_7_12_1","doi-asserted-by":"crossref","unstructured":"He K. Zhang X. Ren S. &Sun J.(2016b).Deep residual learning for image recognition. Proceedings of the CVPR.","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_2_7_13_1","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_2_7_14_1","doi-asserted-by":"crossref","unstructured":"Koehn P. Hoang H. Birch A. Callison\u2010Burch C. Federico M. Bertoldi N. Cowan B. WadeShen C. M. Zens R. et al. (2007).Moses: Open source toolkit for statistical machine translation. Proceedings of the 45th Annual Meeting of the ACL on Interactive Poster and Demonstration Sessions. ACL. 177\u2013180.","DOI":"10.3115\/1557769.1557821"},{"key":"e_1_2_7_15_1","first-page":"706","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Krishna R.","year":"2017"},{"key":"e_1_2_7_16_1","doi-asserted-by":"crossref","unstructured":"Krishna R. Hata K. Ren F. Fei\u2010Fei L. &Niebles J. C.(2017).Dense\u2010Captioning Events in Videos. arXiv:1705.00754.","DOI":"10.1109\/ICCV.2017.83"},{"key":"e_1_2_7_17_1","unstructured":"Lin C. Y.(2004).ROUGE: A package for automatic evaluation of summaries Proceedings of the workshop on text summarization branches out Barcelona Spain (WAS2004)."},{"key":"e_1_2_7_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2940007"},{"key":"e_1_2_7_19_1","doi-asserted-by":"publisher","DOI":"10.1063\/5.0107029"},{"key":"e_1_2_7_20_1","unstructured":"Mubashira N. &James A.(2020).Transformer network for video to text translation In the Proceedings of IEEE International Conference on Power Instrumentation Control and Computing (PICC)."},{"key":"e_1_2_7_21_1","doi-asserted-by":"crossref","unstructured":"Papineni K. Roukos S. Ward T. &Zhu W. J.(2002).BLEU: A method for automatic evaluation of machine translation Proceedings of the 40th annual meeting on Association for Computational Linguistics (ACL'02). Association for Computational Linguistics Stroudsburg PA USA. 311\u2013318.","DOI":"10.3115\/1073083.1073135"},{"key":"e_1_2_7_22_1","unstructured":"Peter A. Xiaodong H. Chris B. Damein T. Mark J. Stephen G. &Lei Z.(2018).Bottom\u2010up and top\u2010down attention for image captioning and visual question answering. IEEE\/CVF Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_2_7_23_1","doi-asserted-by":"crossref","unstructured":"Rohrbach A. Rohrbach M. Tandon N. &Schiele B.(2015).A dataset for movie description. IEEE CVPR.","DOI":"10.1109\/CVPR.2015.7298940"},{"key":"e_1_2_7_24_1","doi-asserted-by":"crossref","unstructured":"Rohrbach M. Qiu W. Titov I. Thater S. Pinkal M. &Schiele B.(2013).Translating video content to natural language descriptions. In IEEE ICCV.","DOI":"10.1109\/ICCV.2013.61"},{"key":"e_1_2_7_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/78.650093"},{"key":"e_1_2_7_26_1","unstructured":"Torabi A. Pal C. Larochelle H. &Courville A.(2015).Using descriptive video services to create a large data source for video annotation research. arXiv Preprint arXiv:1503.01070."},{"key":"e_1_2_7_27_1","unstructured":"Touvron H. Cord M. Douze M. Massa F. Sablayrolles A. &J\u00e9gou H.(2020).Training data\u2010efficient image transformers & distillation through attention. arXiv:2012.12877. Retrieved fromhttps:\/\/arxiv.org\/abs\/2012.12877"},{"key":"e_1_2_7_28_1","doi-asserted-by":"crossref","unstructured":"Varma S. &Dinesh Peter J.(2022 109\u2013114).Video captioning model using a curriculum learning approach 2022 international conference on innovations in science and Technology for Sustainable Development (ICISTSD) Kollam India.https:\/\/doi.org\/10.1109\/ICISTSD55159.2022.10010501","DOI":"10.1109\/ICISTSD55159.2022.10010501"},{"key":"e_1_2_7_29_1","doi-asserted-by":"publisher","DOI":"10.1111\/exsy.12920"},{"key":"e_1_2_7_30_1","doi-asserted-by":"crossref","unstructured":"Varma S. &Sreeraj M.(2013 299\u2013303).Object detection and classification in surveillance system 2013 IEEE Recent Advances in Intelligent Computational Systems (RAICS).https:\/\/doi.org\/10.1109\/RAICS.2013.6745491","DOI":"10.1109\/RAICS.2013.6745491"},{"key":"e_1_2_7_31_1","doi-asserted-by":"crossref","unstructured":"Varma S. &Peter J. D.(2022 847\u2013850).Deep learning\u2010based video captioning technique using transformer 2022 8th International Conference on Advanced Computing and Communication Systems (ICACCS).https:\/\/doi.org\/10.1109\/ICACCS54159.2022.9785074","DOI":"10.1109\/ICACCS54159.2022.9785074"},{"key":"e_1_2_7_32_1","unstructured":"Vaswani A. Shazeer N. Parmar N. Uszkoreit J. Jones L. Gomez A. N. Kaiser \u0141. &Polosukhin I.(2017).Attention is all you need. NIPS 6000\u20136010."},{"key":"e_1_2_7_33_1","doi-asserted-by":"crossref","unstructured":"Vedantam R. Zitnick C. L. &Parikh D.(2015).CIDER: Consensus\u2010based image description evaluation Proceedings of IEEE Conference on Computer Vision and Pattern Recognition. 4566\u20134575.","DOI":"10.1109\/CVPR.2015.7299087"},{"key":"e_1_2_7_34_1","doi-asserted-by":"crossref","unstructured":"Venugopalan S. Rohrbach M. Donahue J. Mooney R. Darrell T. &Saenko K.(2015).Sequence to sequence\u2010video to text. In IEEE ICCV.","DOI":"10.1109\/ICCV.2015.515"},{"key":"e_1_2_7_35_1","doi-asserted-by":"crossref","unstructured":"Venugopalan S. Xu H. Donahue J. Rohrbach M. Mooney R. &Saenko K.(2014).Translating videos to natural language using deep recurrent neural networks. arXiv Preprint arXiv:1412.4729.","DOI":"10.3115\/v1\/N15-1173"},{"key":"e_1_2_7_36_1","doi-asserted-by":"crossref","unstructured":"Wang W. Chen J. Wu Y. W. &Wang W. Y.(2017).Video captioning via hierarchical reinforcement learning. Retrieved from: arXiv Preprint arXiv:1711.11135.","DOI":"10.1109\/CVPR.2018.00443"},{"key":"e_1_2_7_37_1","doi-asserted-by":"crossref","unstructured":"Wang W. Xie E. Li X. Fan D.\u2010P. Song K. Liang D. Lu T. Luo P. &Shao L.(2021).Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. arXiv:2102.12122. Retrieved fromhttps:\/\/arxiv.org\/abs\/2102.12122","DOI":"10.1109\/ICCV48922.2021.00061"},{"key":"e_1_2_7_38_1","doi-asserted-by":"crossref","unstructured":"Wang X. Chen W. Wu J. Wang Y. &Wang W. Y.(2017).Video captioning via hierarchical reinforcement learning. arXiv Preprint arXiv:1711.11135.","DOI":"10.1109\/CVPR.2018.00443"},{"key":"e_1_2_7_39_1","doi-asserted-by":"crossref","unstructured":"Wang Y. Gan W. Yang J. Wu W. &Yan J.(2019 5016\u20135025. [Online]).Dynamic curriculum learning for imbalanced data classification 2019 IEEE\/CVF International Conference on Computer Vision ICCV 2019 Seoul Korea (South) October 27\u2013November 2 2019. IEEE.https:\/\/doi.org\/10.1109\/ICCV.2019.00512","DOI":"10.1109\/ICCV.2019.00512"},{"key":"e_1_2_7_40_1","doi-asserted-by":"crossref","unstructured":"Xu J. Mei T. Yao T. &Rui Y.(2016).MSR\u2010VTT: A large video description dataset for bridging video and language. IEEE CVPR.","DOI":"10.1109\/CVPR.2016.571"},{"key":"e_1_2_7_41_1","unstructured":"Yang J. Li C. Zhang P. Dai X. Xiao B. Lu Y. &Gao J.(2021).Focal self attention for local\u2010global interactions in vision transformers. arXiv:cs.CV\/2107.00641. Retrieved fromhttps:\/\/arxiv. org\/abs\/2107.00641"},{"key":"e_1_2_7_42_1","doi-asserted-by":"crossref","unstructured":"Yao L. Torabi A. Cho K. Ballas N. Pal C. Larochelle H. &Courville A.(2015).Describing videos by exploiting temporal structure. In IEEE ICCV.","DOI":"10.1109\/ICCV.2015.512"},{"key":"e_1_2_7_43_1","doi-asserted-by":"crossref","unstructured":"Yu H. Wang J. Huang Z. Yang Y. &Xu W.(2016).Video paragraph captioning using hierarchical recurrent neural networks. In IEEE CVPR.","DOI":"10.1109\/CVPR.2016.496"},{"key":"e_1_2_7_44_1","doi-asserted-by":"crossref","unstructured":"Yuan L. Chen Y. Wang T. Yu W. Shi Y. Francis E. H. Tay J. F. &Yan S.(2021).Tokens\u2010to\u2010token ViT: Training vision transformers from scratch on imagenet. arXiv:2101.11986. Retrieved fromhttps:\/\/arxiv.org\/abs\/2101.11986","DOI":"10.1109\/ICCV48922.2021.00060"},{"key":"e_1_2_7_45_1","doi-asserted-by":"publisher","DOI":"10.7717\/peerj-cs.916"}],"container-title":["Expert Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/exsy.13392","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/full-xml\/10.1111\/exsy.13392","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/exsy.13392","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,14]],"date-time":"2026-05-14T08:28:13Z","timestamp":1778747293000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1111\/exsy.13392"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,7,5]]},"references-count":44,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2026,6]]}},"alternative-id":["10.1111\/exsy.13392"],"URL":"https:\/\/doi.org\/10.1111\/exsy.13392","archive":["Portico"],"relation":{},"ISSN":["0266-4720","1468-0394"],"issn-type":[{"value":"0266-4720","type":"print"},{"value":"1468-0394","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,7,5]]},"assertion":[{"value":"2023-03-08","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-06-15","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-07-05","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e13392"}}