{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,7]],"date-time":"2026-03-07T01:39:33Z","timestamp":1772847573867,"version":"3.50.1"},"reference-count":39,"publisher":"World Scientific Pub Co Pte Ltd","issue":"04","funder":[{"name":"Conselho Nacional de Desenvolvimento Cientifico e Tecnol\u00f3ogico CNPq","award":["PQ 310075\/2019-0"],"award-info":[{"award-number":["PQ 310075\/2019-0"]}]},{"name":"Funda\u00e7\u00e3o de Amparo \u00e0 Pesquisa do Estado de Minas Gerais FAPEMIG","award":["Grants PPM-00006-18"],"award-info":[{"award-number":["Grants PPM-00006-18"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Int. J. Semantic Computing"],"published-print":{"date-parts":[[2023,12]]},"abstract":"<jats:p> A coherent description is an ultimate goal regarding video captioning via a couple of sentences because it might also affect the consistency and intelligibility of the generated results. In this context, a paragraph describing a video is affected by the activities used to both produce its specific narrative and provide some clues that can also assist in decreasing textual repetition. This work proposes a model, named Hierarchical time-aware Summarization with an Adaptive Transformer (HSAT), that uses a strategy to enhance the frame selection reducing the amount of information that needed to be processed along with attention mechanisms to enhance a memory-augmented transformer. This new approach increases the coherence among the generated sentences, assessing data importance (about the video segments) contained in the self-attention results and uses that to improve readability using only a small fraction of time spent by the other methods. The test results show the potential of this new approach as it provides higher coherence among the various video segments, decreasing the repetition in the generated sentences and improving the description diversity in the ActivityNet Captions dataset. <\/jats:p>","DOI":"10.1142\/s1793351x23640031","type":"journal-article","created":{"date-parts":[[2023,6,10]],"date-time":"2023-06-10T05:51:35Z","timestamp":1686376295000},"page":"569-592","source":"Crossref","is-referenced-by-count":4,"title":["Hierarchical Time-Aware Summarization with an Adaptive Transformer for Video Captioning"],"prefix":"10.1142","volume":"17","author":[{"given":"Leonardo Vilela","family":"Cardoso","sequence":"first","affiliation":[{"name":"Image and Multimedia Data Science Laboratory (IMSCIENCE), Pontifcia Universidade Catlica de Minas Gerais (PUC Minas), Av. Dom Jos\u00e9 Gaspar, 500 - Pr\u00e9dio 20, 30535-901, Belo Horizonte, Brazil"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Silvio Jamil Ferzoli","family":"Guimar\u00e3es","sequence":"additional","affiliation":[{"name":"Image and Multimedia Data Science Laboratory (IMSCIENCE), Pontifcia Universidade Catlica de Minas Gerais (PUC Minas), Av. Dom Jos\u00e9 Gaspar, 500 - Pr\u00e9dio 20, 30535-901, Belo Horizonte, Brazil"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zenilton Kleber Gon\u00e7alves","family":"do Patroc\u00ednio J\u00fanior","sequence":"additional","affiliation":[{"name":"Image and Multimedia Data Science Laboratory (IMSCIENCE), Pontifcia Universidade Catlica de Minas Gerais (PUC Minas), Av. Dom Jos\u00e9 Gaspar, 500 - Pr\u00e9dio 20, 30535-901, Belo Horizonte, Brazil"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"219","published-online":{"date-parts":[[2023,7,25]]},"reference":[{"key":"S1793351X23640031BIB001","first-page":"2603","volume-title":"Proc. 58th Annu. Meeting of the ACL","author":"Lei J.","year":"2020"},{"key":"S1793351X23640031BIB002","first-page":"836","volume-title":"Proc. 2021 IEEE 33rd Int. Conf. Tools with Artificial Intelligence","author":"Cardoso L. V.","year":"2021"},{"key":"S1793351X23640031BIB003","first-page":"37","volume-title":"Proc. 2022 IEEE 8th Int. Conf. Multimedia Big Data","author":"Cardoso L. V.","year":"2022"},{"key":"S1793351X23640031BIB004","first-page":"706","volume-title":"Proc. IEEE Int. Conf. Computer Vision","author":"Krishna R.","year":"2017"},{"issue":"6","key":"S1793351X23640031BIB005","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3355390","volume":"52","author":"Aafaq N.","year":"2019","journal-title":"ACM Comput. Surv."},{"key":"S1793351X23640031BIB006","doi-asserted-by":"crossref","first-page":"105667","DOI":"10.1016\/j.engappai.2022.105667","volume":"118","author":"Meena P.","year":"2023","journal-title":"Eng. Appl. Artif. Intell."},{"key":"S1793351X23640031BIB007","doi-asserted-by":"crossref","first-page":"103670","DOI":"10.1016\/j.jvcir.2022.103670","volume":"89","author":"Narwal P.","year":"2022","journal-title":"J. Visual Commun. Image Represent."},{"issue":"11","key":"S1793351X23640031BIB008","doi-asserted-by":"crossref","first-page":"1838","DOI":"10.1109\/JPROC.2021.3117472","volume":"109","author":"Apostolidis E.","year":"2021","journal-title":"Proc. IEEE"},{"key":"S1793351X23640031BIB009","first-page":"5998","volume-title":"Proc. 30th Annu. Conf. Neural Information Processing Systems","author":"Vaswani A.","year":"2017"},{"key":"S1793351X23640031BIB010","first-page":"7513","volume-title":"Proc. IEEE Int. Conf. Acoustics, Speech and Signal Processing","author":"Vydana H. K.","year":"2021"},{"key":"S1793351X23640031BIB011","doi-asserted-by":"crossref","first-page":"1154","DOI":"10.1145\/3437963.3441667","volume-title":"Proc. 14th ACM Int. Conf. Web Search and Data Mining","author":"Yates A.","year":"2021"},{"issue":"05","key":"S1793351X23640031BIB012","first-page":"7847","volume":"34","author":"Guo Q.","year":"2020","journal-title":"Proc. AAAI Conf. Artif. Intell."},{"key":"S1793351X23640031BIB013","first-page":"5059","volume-title":"Proc. 57th Annu. Meeting of the ACL","author":"Zhang X.","year":"2019"},{"key":"S1793351X23640031BIB015","first-page":"10971","volume-title":"Proc. IEEE\/CVF Conf. Computer Vision and Pattern Recognition","author":"Pan Y.","year":"2020"},{"key":"S1793351X23640031BIB016","first-page":"4634","volume-title":"Proc. IEEE\/CVF Int. Conf. Computer Vision","author":"Huang L.","year":"2019"},{"key":"S1793351X23640031BIB017","first-page":"6578","volume-title":"Proc. 58th Annu. Meeting of the ACL","author":"Tang H.","year":"2020"},{"key":"S1793351X23640031BIB018","first-page":"2978","volume-title":"Proc. 57th Annu. Meeting of the ACL","author":"Dai Z.","year":"2019"},{"key":"S1793351X23640031BIB020","doi-asserted-by":"crossref","first-page":"1001","DOI":"10.1016\/j.neucom.2015.08.057","volume":"173","author":"dos Santos Belo L.","year":"2016","journal-title":"Neurocomputing"},{"key":"S1793351X23640031BIB021","first-page":"6598","volume-title":"Proc. IEEE\/CVF Conf. Computer Vision and Pattern Recognition","author":"Park J. S.","year":"2019"},{"key":"S1793351X23640031BIB022","first-page":"6578","volume-title":"Proc. IEEE\/CVF Conf. Computer Vision and Pattern Recognition","author":"Zhou L.","year":"2019"},{"key":"S1793351X23640031BIB023","volume-title":"Proc. Int. Conf. Learning Representations","author":"Bahdanau D.","year":"2015"},{"key":"S1793351X23640031BIB025","first-page":"2048","volume-title":"Proc. Int. Conf. Machine Learning","author":"Xu K.","year":"2015"},{"key":"S1793351X23640031BIB026","first-page":"7132","volume-title":"Proc. IEEE Conf. Computer Vision and Pattern Recognition","author":"Hu J.","year":"2018"},{"key":"S1793351X23640031BIB027","first-page":"7464","volume-title":"Proc. IEEE\/CVF Int. Conf. Computer Vision","author":"Sun C.","year":"2019"},{"key":"S1793351X23640031BIB028","first-page":"8739","volume-title":"IEEE Conf. Computer Vision and Pattern Recognition","author":"Zhou L.","year":"2018"},{"key":"S1793351X23640031BIB029","volume-title":"Proc. Int. Conf. Learning Representations","author":"Chen Y.-C.","year":"2019"},{"issue":"8","key":"S1793351X23640031BIB030","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","volume":"9","author":"Hochreiter S.","year":"1997","journal-title":"Neural Comput."},{"key":"S1793351X23640031BIB031","first-page":"1724","volume-title":"Proc. 2014 Conf. Empirical Methods in Natural Language Processing","author":"Cho K.","year":"2014"},{"key":"S1793351X23640031BIB032","doi-asserted-by":"crossref","first-page":"479","DOI":"10.1007\/s10851-017-0768-7","volume":"60","author":"Cousty J.","year":"2018","journal-title":"J. Math. Imaging Vision"},{"issue":"1","key":"S1793351X23640031BIB033","first-page":"55","volume":"2","author":"Guimar\u00e3es S.","year":"2017","journal-title":"Math. Morphol. Theory Appl."},{"key":"S1793351X23640031BIB034","first-page":"272","volume-title":"Proc. 10th Int. Symp. Mathematical Morphology and Its Applications to Signal and Image Processing","author":"Cousty J.","year":"2011"},{"key":"S1793351X23640031BIB035","first-page":"4566","volume-title":"Proc. IEEE Conf. Computer Vision and Pattern Recognition","author":"Vedantam R.","year":"2015"},{"key":"S1793351X23640031BIB036","first-page":"468","volume-title":"Proc. European Conf. Computer Vision","author":"Xiong Y.","year":"2018"},{"key":"S1793351X23640031BIB037","first-page":"374","volume-title":"Proc. European Conf. Computer Vision","author":"Zhang B.","year":"2018"},{"key":"S1793351X23640031BIB038","first-page":"961","volume-title":"Proc. IEEE Conf. Computer Vision and Pattern Recognition","author":"Caba Heilbron F.","year":"2015"},{"key":"S1793351X23640031BIB039","first-page":"135","volume-title":"Proc. 11th Int. Symp. Mathematical Morphology and Its Applications to Signal and Image Processing","author":"Najman L.","year":"2013"},{"key":"S1793351X23640031BIB040","first-page":"770","volume-title":"Proc. IEEE Conf. Computer Vision and Pattern Recognition","author":"He K.","year":"2016"},{"key":"S1793351X23640031BIB041","first-page":"448","volume-title":"Proc. Int. Conf. Machine Learning","author":"Ioffe S.","year":"2015"},{"key":"S1793351X23640031BIB043","first-page":"311","volume-title":"Proc. 40th Annu. Meeting of the ACL","author":"Papineni K.","year":"2002"}],"container-title":["International Journal of Semantic Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.worldscientific.com\/doi\/pdf\/10.1142\/S1793351X23640031","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,11,24]],"date-time":"2023-11-24T08:08:55Z","timestamp":1700813335000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.worldscientific.com\/doi\/10.1142\/S1793351X23640031"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,7,25]]},"references-count":39,"journal-issue":{"issue":"04","published-print":{"date-parts":[[2023,12]]}},"alternative-id":["10.1142\/S1793351X23640031"],"URL":"https:\/\/doi.org\/10.1142\/s1793351x23640031","relation":{},"ISSN":["1793-351X","1793-7108"],"issn-type":[{"value":"1793-351X","type":"print"},{"value":"1793-7108","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,7,25]]}}}