{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,2]],"date-time":"2026-05-02T07:09:46Z","timestamp":1777705786501,"version":"3.51.4"},"reference-count":27,"publisher":"SAGE Publications","issue":"5","license":[{"start":{"date-parts":[[2021,12,15]],"date-time":"2021-12-15T00:00:00Z","timestamp":1639526400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["Journal of Intelligent &amp; Fuzzy Systems"],"published-print":{"date-parts":[[2022,3,31]]},"abstract":"<jats:p>Current State-of-the-Art image captioning systems that can read and integrate read text into the generated descriptions need high processing power and memory usage, which limits the sustainability and usability of the models (as they require expensive and very specialized hardware). The present work introduces two alternative versions (L-M4C and L-CNMT) of top architectures (on the TextCaps challenge), which were mainly adapted to achieve near-State-of-The-Art performance while being memory-lighter when compared to the original architectures, this is mainly achieved by using distilled or smaller pre-trained models on the text-and-OCR embedding modules. On the one hand, a distilled version of BERT was used in order to reduce the size of the text-embedding module (the distilled model has 59% fewer parameters), on the other hand, the OCR context processor on both architectures was replaced by Global Vectors (GloVe), instead of using FastText pre-trained vectors, this can reduce the memory used by the OCR-embedding module up to a 94% . Two of the three models presented in this work surpassed the baseline (M4C-Captioner) of the challenge on the evaluation and test sets, also, our best lighter architecture reached a CIDEr score of 88.24 on the test set, which is 7.25 points above the baseline model.<\/jats:p>","DOI":"10.3233\/jifs-219230","type":"journal-article","created":{"date-parts":[[2021,12,21]],"date-time":"2021-12-21T11:59:39Z","timestamp":1640087979000},"page":"4399-4410","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":2,"title":["Searching for memory-lighter architectures for OCR-augmented image captioning"],"prefix":"10.1177","volume":"42","author":[{"given":"Rafael","family":"Gallardo-Garc\u00eda","sequence":"first","affiliation":[{"name":"Language and Knowledge Engineering Laboratory, Benem\u00e9rita Universidad Aut\u00f3noma de Puebla, Puebla, Mexico"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Beatriz","family":"Beltr\u00e1n-Mart\u00ednez","sequence":"additional","affiliation":[{"name":"Language and Knowledge Engineering Laboratory, Benem\u00e9rita Universidad Aut\u00f3noma de Puebla, Puebla, Mexico"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Carlos","family":"Hern\u00e1ndez-Gracidas","sequence":"additional","affiliation":[{"name":"Faculty of Physical and Mathematical Sciences, Benem\u00e9rita Universidad Aut\u00f3noma de Puebla, Puebla, Mexico"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Darnes","family":"Vilari\u00f1o-Ayala","sequence":"additional","affiliation":[{"name":"Language and Knowledge Engineering Laboratory, Benem\u00e9rita Universidad Aut\u00f3noma de Puebla, Puebla, Mexico"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[2021,12,15]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"AmirianS. RasheedK. TahaT.R. and ArabniaH.R. A short review on image caption generation with deep learning In Proceedings of the International Conference on Image Processing Computer Vision and Pattern Recognition (IPCV) pages 10\u201318 The Steering Committee of The World Congress in Computer Science 2019."},{"key":"e_1_3_1_3_2","doi-asserted-by":"crossref","unstructured":"AndersonP. FernandoB. JohnsonM. and GouldS. Spice: Semantic propositional image caption evaluation In European Conference on Computer Vision pages 382\u2013398 Springer 2016.","DOI":"10.1007\/978-3-319-46454-1_24"},{"key":"e_1_3_1_4_2","doi-asserted-by":"crossref","unstructured":"BaekY. LeeB. HanD. YunS. and LeeH. Character region awareness for text detection In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition pages 9365\u20139374 2019.","DOI":"10.1109\/CVPR.2019.00959"},{"key":"e_1_3_1_5_2","doi-asserted-by":"crossref","unstructured":"BorisyukF. GordoA. and SivakumarV. Rosetta: Large scale system for text detection and recognition in images In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining pages 71\u201379 2018.","DOI":"10.1145\/3219819.3219861"},{"key":"e_1_3_1_6_2","doi-asserted-by":"crossref","unstructured":"DenkowskiM. and LavieA. Meteor universal: Language specific translation evaluation for any target language In Proceedings of the ninth workshop on statistical machine translation pages 376\u2013380 2014.","DOI":"10.3115\/v1\/W14-3348"},{"key":"e_1_3_1_7_2","unstructured":"DevlinJ. ChangM.-W. LeeK. and ToutanovaK. Bert: Pre-training of deep bidirectional transformers for language understanding In NAACL-HLT 2019."},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3295748"},{"key":"e_1_3_1_9_2","doi-asserted-by":"crossref","unstructured":"HuR. SinghA. DarrellT. and RohrbachM. Iterative answer prediction with pointer-augmented multimodal transformers for textvqa In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition pages 9992\u201310002 2020.","DOI":"10.1109\/CVPR42600.2020.01001"},{"key":"e_1_3_1_10_2","unstructured":"LinC.-Y. Rouge: A package for automatic evaluation of summaries In Text summarization branches out pages 74\u201381 2004."},{"key":"e_1_3_1_11_2","doi-asserted-by":"crossref","unstructured":"LiuY. ChenH. ShenC. HeT. JinL. and WangL. Abcnet: Real-time scene text spotting with adaptive bezier-curve network In Proceedings of the IEEE\/CVFConference on Computer Vision and Pattern Recognition pages 9809\u20139818 2020.","DOI":"10.1109\/CVPR42600.2020.00983"},{"key":"e_1_3_1_12_2","unstructured":"MikolovT. ChenK. CorradoG. and DeanJ. Efficient estimation ofword representations in vector space In ICLR 2013."},{"key":"e_1_3_1_13_2","unstructured":"MikolovT. GraveE. BojanowskiP. PuhrschC. and JoulinA. Advances in pre-training distributed word representations In Proceedings of the International Conference on Language Resources and Evaluation (LREC 2018) 2018."},{"key":"e_1_3_1_14_2","doi-asserted-by":"crossref","unstructured":"PapineniK. RoukosS. WardT. and ZhuW.-J. Bleu: a method for automatic evaluation of machine translation In Proceedings of the 40th annual meeting of the Association for Computational Linguistics pages 311\u2013318 2002.","DOI":"10.3115\/1073083.1073135"},{"key":"e_1_3_1_15_2","unstructured":"ParcalabescuL. TrostN. and FrankA. What is multimodality? arXiv preprint arXiv:2103.06304 2021."},{"key":"e_1_3_1_16_2","doi-asserted-by":"crossref","unstructured":"PenningtonJ. SocherR. and ManningC.D Glove: Global vectors for word representation In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP) pages 1532\u20131543 2014.","DOI":"10.3115\/v1\/D14-1162"},{"key":"e_1_3_1_17_2","first-page":"91","article-title":"Faster r-cnn: Towardsreal-time object detection with region proposal networks","volume":"28","author":"Ren S.","year":"2015","unstructured":"RenS., HeK., GirshickR., SunJ., Faster r-cnn: Towardsreal-time object detection with region proposal networks, Advances in Neural Information Processing Systems28 (2015), 91\u201399.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_18_2","unstructured":"SanhV. DebutL. ChaumondJ. and WolfT. Distilbert a distilled version of bert: smaller faster cheaper and lighter arXiv preprint arXiv:1910.01108 2019."},{"key":"e_1_3_1_19_2","doi-asserted-by":"crossref","unstructured":"SidorovO. HuR. RohrbachM. and SinghA. Textcaps: a dataset for image captioning with reading comprehension In European Conference on Computer Vision pages 742\u2013758. Springer 2020.","DOI":"10.1007\/978-3-030-58536-5_44"},{"key":"e_1_3_1_20_2","doi-asserted-by":"crossref","unstructured":"SinghA. NatarajanV. ShahM. JiangY. ChenX. BatraD. ParikhD. and RohrbachM. Towards vqa models that can read In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition pages 8317\u20138326 2019.","DOI":"10.1109\/CVPR.2019.00851"},{"key":"e_1_3_1_21_2","doi-asserted-by":"crossref","unstructured":"SudholtS. and FinkG.A. Phocnet: A deep convolutional neural network forword spotting in handwritten documents In 2016 15th International Conference on Frontiers in Handwriting Recognition (ICFHR) pages 277\u2013282 IEEE 2016.","DOI":"10.1109\/ICFHR.2016.0060"},{"key":"e_1_3_1_22_2","doi-asserted-by":"crossref","unstructured":"VedantamR. ZitnickC.L. and ParikhD. Cider: Consensus-based image description evaluation In Proceedings of the IEEE conference on computer vision and pattern recognition pages 4566\u20134575 2015.","DOI":"10.1109\/CVPR.2015.7299087"},{"key":"e_1_3_1_23_2","doi-asserted-by":"crossref","unstructured":"WangJ. TangJ. and LuoJ. Multimodal attention with image text spatial relationship for ocr-based image captioning In Proceedings of the 28th ACM International Conference on Multimedia pages 4337\u20134345 2020.","DOI":"10.1145\/3394171.3413753"},{"key":"e_1_3_1_24_2","doi-asserted-by":"crossref","unstructured":"WangJ. TangJ. YangM. BaiX. and LuoJ. Improving ocr-based image captioning by incorporating geometrical relationship In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition pages 1306\u20131315 2021.","DOI":"10.1109\/CVPR46437.2021.00136"},{"key":"e_1_3_1_25_2","doi-asserted-by":"crossref","unstructured":"WangZ. BaoR. WuQ. LiuS. Confidence-aware non-repetitive multimodal transformers for textcaps In AAAI 2021.","DOI":"10.1609\/aaai.v35i4.16389"},{"key":"e_1_3_1_26_2","doi-asserted-by":"crossref","unstructured":"XuG. NiuS. TanM. LuoY. DuQ. and WuQ. Towards accurate text-based image captioning with content diversity exploration In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition pages 12637\u201312646 2021.","DOI":"10.1109\/CVPR46437.2021.01245"},{"key":"e_1_3_1_27_2","doi-asserted-by":"crossref","unstructured":"YangZ. LuY. WangJ. YinX. FlorencioD. WangL. ZhangC. ZhangL. and LuoJ. Tap:Text-aware pre-training for text-vqa and text-caption In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition pages 8751\u20138761 2021.","DOI":"10.1109\/CVPR46437.2021.00864"},{"key":"e_1_3_1_28_2","doi-asserted-by":"crossref","unstructured":"ZhuQ. GaoC. WangP. and WuQ. Simple is not easy: A simple strong baseline for textvqa and textcaps In AAAI 2021.","DOI":"10.1609\/aaai.v35i4.16476"}],"container-title":["Journal of Intelligent &amp; Fuzzy Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.3233\/JIFS-219230","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.3233\/JIFS-219230","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.3233\/JIFS-219230","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T09:45:12Z","timestamp":1777455912000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.3233\/JIFS-219230"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,12,15]]},"references-count":27,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2022,3,31]]}},"alternative-id":["10.3233\/JIFS-219230"],"URL":"https:\/\/doi.org\/10.3233\/jifs-219230","relation":{},"ISSN":["1064-1246","1875-8967"],"issn-type":[{"value":"1064-1246","type":"print"},{"value":"1875-8967","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,12,15]]}}}