{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,10]],"date-time":"2026-02-10T16:36:24Z","timestamp":1770741384274,"version":"3.49.0"},"reference-count":25,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2026,2,9]],"date-time":"2026-02-09T00:00:00Z","timestamp":1770595200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,2,9]],"date-time":"2026-02-09T00:00:00Z","timestamp":1770595200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001631","name":"University College Dublin","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100001631","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Multimed Tools Appl"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Traditional multimedia classification techniques rely on analyzing either its features or the associated annotated textual information. In this paper, we introduce a technique that leverages deep learning methodologies to extract both low-level visual features and high-level semantic information from the billboard images. We propose a multi-modal hybrid fusion model that integrates text and image features to categorize billboards in video frames into food, sports, and miscellaneous categories. Achieving a 5% accuracy improvement over image-based models and 10% over text-based models, our model demonstrates robust generalization across diverse datasets, benefiting advertising, media, and content creation industries.<\/jats:p>","DOI":"10.1007\/s11042-026-21192-y","type":"journal-article","created":{"date-parts":[[2026,2,9]],"date-time":"2026-02-09T22:27:59Z","timestamp":1770676079000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Multi-modal fusion for billboard categorization in video frames: A text and image-based approach"],"prefix":"10.1007","volume":"85","author":[{"given":"Sukriti","family":"Dhang","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jason","family":"Lok","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mimi","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Soumyabrata","family":"Dev","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2026,2,9]]},"reference":[{"key":"21192_CR1","first-page":"408","volume":"2004","author":"F Aldershoff","year":"2003","unstructured":"Aldershoff F, Gevers T (2003) Visual tracking and localization of billboards in streamed soccer matches. Storage and Retrieval Methods and Applications for Multimedia 2004:408\u2013416","journal-title":"Storage and Retrieval Methods and Applications for Multimedia"},{"key":"21192_CR2","doi-asserted-by":"publisher","first-page":"338","DOI":"10.1016\/j.procs.2019.08.210","volume":"156","author":"K Bochkarev","year":"2019","unstructured":"Bochkarev K, Smirnov E (2019) Detecting advertising on building fa\u00e7ades with computer vision. Procedia Computer Science 156:338\u2013346","journal-title":"Procedia Computer Science"},{"key":"21192_CR3","doi-asserted-by":"crossref","unstructured":"Cai G, Chen L, Li J (2003) Billboard advertising detection in sport tv. In: Seventh International Symposium on Signal Processing and Its Applications, 2003. Proceedings., pp 537\u2013540","DOI":"10.1109\/ISSPA.2003.1224759"},{"key":"21192_CR4","first-page":"57","volume":"2021","author":"S Chavan","year":"2021","unstructured":"Chavan S, Kerr D, Coleman S et al (2021) Billboard detection in the wild. Irish Machine Vision and Image Processing Conference 2021:57\u201364","journal-title":"Irish Machine Vision and Image Processing Conference"},{"key":"21192_CR5","doi-asserted-by":"crossref","unstructured":"Covell M, Baluja S, Fink M (2006) Advertisement detection and replacement using acoustic and visual repetition. In: 2006 IEEE workshop on multimedia signal processing, pp 461\u2013466","DOI":"10.1109\/MMSP.2006.285351"},{"key":"21192_CR6","unstructured":"Dauphinee T, Patel N, Rashidi M (2019) Modular multimodal architecture for document classification. arXiv preprint arXiv:1912.04376"},{"key":"21192_CR7","doi-asserted-by":"crossref","unstructured":"Dhang S, Idogho A, Zhang M et al (2024) Towards accurate billboard detection: an ablation and benchmarking study of deep learning models. In: 26th Irish Machine Vision and Image Processing Conference, IET Conference Proceedings CP887, pp 178\u2013185","DOI":"10.1049\/icp.2024.3303"},{"key":"21192_CR8","doi-asserted-by":"crossref","unstructured":"Dhang S, Zhang M, Dev S (2025) LaBINet - An Approach for Seamlessly Integrating New Advertisement into an Existing Scene. IEEE Trans on Artificial Intelligence, vol. 6(8):2281\u20132290","DOI":"10.1109\/TAI.2025.3544595"},{"key":"21192_CR9","unstructured":"Feng Z, Neumann J (2013) Real time commercial detection in videos. Tech rep Tech Rep 2013.[Online]"},{"key":"21192_CR10","unstructured":"Hossari M, Dev S, Nicholson M et\u00a0al (2018) ADNet: A deep network for detecting adverts. arXiv preprint arXiv:1811.04115"},{"key":"21192_CR11","doi-asserted-by":"crossref","unstructured":"Hussain Z, Zhang M, Zhang X et\u00a0al (2017) Automatic understanding of image and video advertisements. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 1705\u20131715","DOI":"10.1109\/CVPR.2017.123"},{"key":"21192_CR12","doi-asserted-by":"crossref","unstructured":"Intasuwan T, Kaewthong J, Vittayakorn S (2018) Text and object detection on billboards. In: 2018 10th International conference on information technology and electrical engineering (ICITEE), pp 6\u201311","DOI":"10.1109\/ICITEED.2018.8534879"},{"key":"21192_CR13","first-page":"1","volume":"2018","author":"CN Kamath","year":"2018","unstructured":"Kamath CN, Bukhari SS, Dengel A (2018) Comparative study between traditional machine learning and deep learning approaches for text classification. Proceedings of the ACM Symposium on Document Engineering 2018:1\u201311","journal-title":"Proceedings of the ACM Symposium on Document Engineering"},{"key":"21192_CR14","doi-asserted-by":"crossref","unstructured":"Krishna O, Aizawa K (2018) Billboard saliency detection in street videos for adults and elderly. In: 2018 25th IEEE International conference on image processing (ICIP), pp 2326\u20132330","DOI":"10.1109\/ICIP.2018.8451835"},{"key":"21192_CR15","doi-asserted-by":"crossref","unstructured":"Li B, Drozd A, Liu T et\u00a0al (2018) Subword-level composition functions for learning word embeddings. In: Proceedings of the second workshop on subword\/character level models, pp 38\u201348","DOI":"10.18653\/v1\/W18-1205"},{"issue":"1","key":"21192_CR16","first-page":"72","volume":"2","author":"R Mithe","year":"2013","unstructured":"Mithe R, Indalkar S, Divekar N (2013) Optical character recognition. International journal of recent technology and engineering (IJRTE) 2(1):72\u201375","journal-title":"International journal of recent technology and engineering (IJRTE)"},{"key":"21192_CR17","doi-asserted-by":"crossref","unstructured":"Nautiyal A, McCabe K, Hossari M et\u00a0al (2019) An advert creation system for next-gen publicity. In: Machine Learning and Knowledge Discovery in Databases: European Conference, ECML PKDD 2018, Dublin, Ireland, September 10\u201314, 2018, Proceedings, Part III 18, pp 663\u2013667","DOI":"10.1007\/978-3-030-10997-4_47"},{"issue":"10","key":"21192_CR18","first-page":"50","volume":"55","author":"C Patel","year":"2012","unstructured":"Patel C, Patel A, Patel D (2012) Optical character recognition by open source OCR tool tesseract: A case study. Int J Comp Appl 55(10):50\u201356","journal-title":"Int J Comp Appl"},{"issue":"5","key":"21192_CR19","doi-asserted-by":"publisher","first-page":"2659","DOI":"10.12928\/telkomnika.v17i5.11276","volume":"17","author":"RF Rahmat","year":"2019","unstructured":"Rahmat RF, Dennis D, Sitompul OS et al (2019) Advertisement billboard detection and geotagging system with inductive transfer learning in deep convolutional neural network. TELKOMNIKA (Telecommunication Computing Electronics and Control) 17(5):2659\u20132666","journal-title":"TELKOMNIKA (Telecommunication Computing Electronics and Control)"},{"key":"21192_CR20","doi-asserted-by":"crossref","unstructured":"Smith R (2007) An Overview of the Tesseract OCR Engine. In: Ninth international conference on document analysis and recognition (ICDAR 2007), pp 629\u2013633","DOI":"10.1109\/ICDAR.2007.4376991"},{"issue":"1","key":"21192_CR21","doi-asserted-by":"publisher","first-page":"104","DOI":"10.1016\/j.ipm.2013.08.006","volume":"50","author":"AK Uysal","year":"2014","unstructured":"Uysal AK, Gunal S (2014) The impact of preprocessing on text classification. Information processing & management 50(1):104\u2013112","journal-title":"Information processing & management"},{"key":"21192_CR22","doi-asserted-by":"crossref","unstructured":"Wang D, Mao K, Ng GW (2017) Convolutional neural networks and multimodal fusion for text aided image classification. In: 2017 20th International conference on information fusion (Fusion), pp 1\u20137","DOI":"10.23919\/ICIF.2017.8009768"},{"issue":"7","key":"21192_CR23","doi-asserted-by":"publisher","first-page":"994","DOI":"10.1016\/j.patrec.2008.01.022","volume":"29","author":"A Watve","year":"2008","unstructured":"Watve A, Sural S (2008) Soccer video processing for the detection of advertisement billboards. Pattern Recogn Lett 29(7):994\u20131006","journal-title":"Pattern Recogn Lett"},{"key":"21192_CR24","doi-asserted-by":"crossref","unstructured":"Yao T, Zhai Z, Gao B (2020) Text classification model based on fasttext. In: 2020 IEEE International conference on artificial intelligence and information systems (ICAIIS), pp 154\u2013157","DOI":"10.1109\/ICAIIS49377.2020.9194939"},{"issue":"6","key":"21192_CR25","doi-asserted-by":"publisher","first-page":"9065","DOI":"10.1007\/s11042-021-11461-3","volume":"82","author":"L Yu","year":"2023","unstructured":"Yu L, Li G, Yuan L et al (2023) Time-bounded targeted influence spread in online social networks. Multim Tools Appl 82(6):9065\u20139081","journal-title":"Multim Tools Appl"}],"container-title":["Multimedia Tools and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11042-026-21192-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11042-026-21192-y","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11042-026-21192-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,9]],"date-time":"2026-02-09T22:28:02Z","timestamp":1770676082000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11042-026-21192-y"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,2,9]]},"references-count":25,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2026,2]]}},"alternative-id":["21192"],"URL":"https:\/\/doi.org\/10.1007\/s11042-026-21192-y","relation":{},"ISSN":["1573-7721"],"issn-type":[{"value":"1573-7721","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,2,9]]},"assertion":[{"value":"7 February 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"14 November 2025","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"27 November 2025","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"9 February 2026","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing Interests"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical Approval"}}],"article-number":"153"}}