{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,25]],"date-time":"2026-03-25T05:51:42Z","timestamp":1774417902391,"version":"3.50.1"},"reference-count":38,"publisher":"Springer Science and Business Media LLC","issue":"5","license":[{"start":{"date-parts":[[2026,3,1]],"date-time":"2026-03-01T00:00:00Z","timestamp":1772323200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,3,6]],"date-time":"2026-03-06T00:00:00Z","timestamp":1772755200000},"content-version":"vor","delay-in-days":5,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Social Science University of Ankara"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Neural Comput &amp; Applic"],"published-print":{"date-parts":[[2026,3]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>The primary aim is to develop a model that can achieve high accuracy in solving multi label classification problems by training, testing, and analyzing deep learning models that utilize both image and text data. In this paper, we propose a novel hybrid model that jointly considers textual and visual data for the task of movie genre prediction, which is a representative example of multi label classification problems. In the proposed model, textual features extracted from movie summaries are obtained using the DistilBERT model, while visual features derived from movie posters are extracted using the ConvNeXt deep learning model. These features from the two modalities are then combined using the XGBoost machine learning algorithm to perform genre prediction. This approach aims to achieve higher accuracy and better generalizability in movie genre classification by integrating information from different modalities through a late fusion method. The ConvNeXt architecture was adapted to the problem using transfer learning and fine-tuning techniques. To achieve the highest performance from the DistilBERT model, optimization was performed for the token length and threshold hyperparameters, and the model with a token length of 256 and a threshold of 0.5 was used in the hybrid model. Furthermore, to maximize the overall performance of the novel hybrid model, optimization was conducted using the Grid Search algorithm. All three proposed models were trained and tested on a dataset obtained from the IMDB website. The performances of the models were evaluated using hamming loss, precision and F1 score metrics. Experimental results revealed that, overall, the text-based model outperformed the image-based model, while the proposed hybrid model achieved higher performance than both individual models. It was demonstrated that textual and visual features complement each other and positively enhance the overall performance. This paper presents an original and effective study for multi label movie genre classification, combining the fields of computer vision, natural language processing, and machine learning.<\/jats:p>","DOI":"10.1007\/s00521-026-11852-3","type":"journal-article","created":{"date-parts":[[2026,3,6]],"date-time":"2026-03-06T07:42:39Z","timestamp":1772782959000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Hybrid deep learning approach for multi-label classification problem: genre prediction"],"prefix":"10.1007","volume":"38","author":[{"given":"Fat\u0131ma Zehra","family":"\u00dcnal","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mehmet Serdar","family":"G\u00fczel","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Metehan","family":"\u00dcnal","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fatih","family":"Ekinci","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tun\u00e7","family":"A\u015furo\u011flu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Koray","family":"A\u00e7\u0131c\u0131","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2026,3,6]]},"reference":[{"key":"11852_CR1","doi-asserted-by":"publisher","unstructured":"Paulino MAD, Costa YMG, Feltrim VD (2022) Evaluating multimodal strategies for multi-label movie genre classification. In: Proceedings of the 29th International Conference on Systems, Signals and Image Processing (IWSSIP 2022), pp. 1\u20136. IEEE. https:\/\/doi.org\/10.1109\/IWSSIP55020.2022.9854451","DOI":"10.1109\/IWSSIP55020.2022.9854451"},{"issue":"19","key":"11852_CR2","doi-asserted-by":"publisher","first-page":"36823","DOI":"10.1007\/s11042-023-16121-2","volume":"83","author":"Z Cai","year":"2024","unstructured":"Cai Z, Ding H, Wu J, Xi Y, Wu X, Cui X (2024) Multi-label movie genre classification based on multimodal fusion. Multimed Tools Appl 83(19):36823\u201336840. https:\/\/doi.org\/10.1007\/s11042-023-16121-2","journal-title":"Multimedia Tools Appl"},{"issue":"30","key":"11852_CR3","doi-asserted-by":"publisher","first-page":"2304","DOI":"10.17485\/IJST\/v16i30.1238","volume":"16","author":"P Khant","year":"2023","unstructured":"Khant P, Tidke B (2023) Multimodal approach to recommend movie genres based on multi datasets. Indian J Sci Technol 16(30):2304\u20132310. https:\/\/doi.org\/10.17485\/IJST\/v16i30.1238","journal-title":"Indian J Sci Technol"},{"key":"11852_CR4","doi-asserted-by":"publisher","unstructured":"Li J, Qi G, Zhang C, Chen Y, Tan Y, Xia C, Tian Y (2023) Incorporating domain knowledge graph into multimodal movie genre classification with self-supervised attention and contrastive learning. In: Proceedings of the 31st ACM International Conference on Multimedia (MM\u201923), pp. 2814\u20132823. ACM. https:\/\/doi.org\/10.1145\/3581783.3612085","DOI":"10.1145\/3581783.3612085"},{"key":"11852_CR5","doi-asserted-by":"publisher","first-page":"125209","DOI":"10.1016\/j.eswa.2024.125209","volume":"258","author":"S Sulun","year":"2024","unstructured":"Sulun S, Viana P, Davies MEP (2024) Movie trailer genre classification using multimodal pretrained features. Expert Syst Appl 258:125209. https:\/\/doi.org\/10.1016\/j.eswa.2024.125209","journal-title":"Expert Syst Appl"},{"key":"11852_CR6","doi-asserted-by":"publisher","first-page":"113459","DOI":"10.1016\/j.knosys.2025.113459","volume":"318","author":"IT Soyk\u00f6k","year":"2025","unstructured":"Soyk\u00f6k IT, G\u00fcvenir HA (2025) Multi-label multi-modal classification of movie scenes. Knowl Based Syst 318:113459. https:\/\/doi.org\/10.1016\/j.knosys.2025.113459","journal-title":"Knowl Based Syst"},{"key":"11852_CR7","doi-asserted-by":"crossref","unstructured":"Braz Leod\u00e9cio et al (2021) Image-text integration using a multimodal fusion network module for movie genre classification. In: 11th International Conference of Pattern Recognition Systems (ICPRS 2021). Vol. IET, 2021","DOI":"10.1049\/icp.2021.1456"},{"key":"11852_CR8","doi-asserted-by":"crossref","unstructured":"KIELA, Douwe et al (2018) Efficient large-scale multi-modal classification. In: Proceedings of the AAAI conference on artificial intelligence","DOI":"10.1609\/aaai.v32i1.11945"},{"issue":"14","key":"11852_CR9","doi-asserted-by":"publisher","first-page":"19071","DOI":"10.1007\/s11042-020-10086-2","volume":"81","author":"MANGOLIN Rafael","year":"2022","unstructured":"Rafael MANGOLIN (2022) A multimodal approach for multi-label movie genre classification. Multimedia Tools Appl 81(14):19071\u201319096","journal-title":"Multimedia Tools Appl"},{"key":"11852_CR10","doi-asserted-by":"crossref","unstructured":"Jha A, Kishore et al (2022) Multimodal and Multilabel Genre Classification of Movie Trailers. In: International Conference on Sustainable Computing and Data Communication Systems (ICSCDS). IEEE, 2022","DOI":"10.1109\/ICSCDS53736.2022.9760773"},{"key":"11852_CR11","unstructured":"Cascante-Bonilla P et al Moviescope: Large-scale analysis of movies using multiple modalities. arXiv preprint arXiv:(1908). 03180 (2019)"},{"issue":"1","key":"11852_CR12","doi-asserted-by":"publisher","first-page":"945","DOI":"10.1007\/s11042-022-13211-5","volume":"82","author":"S Kumar","year":"2023","unstructured":"Kumar S et al (2023) Movie genre classification using binary relevance, label powerset, and machine learning classifiers. Multimedia Tools Appl 82(1):945\u2013968","journal-title":"Multimedia Tools Appl"},{"key":"11852_CR13","doi-asserted-by":"crossref","unstructured":"Hasan M, Mehedi et al (2021) Multilabel Movie Genre Classification from Movie Subtitle: Parameter Optimized Hybrid Classifier. In: 2021 4th International Symposium on Advanced Electrical and Communication Technologies (ISAECT). IEEE","DOI":"10.1109\/ISAECT53699.2021.9668427"},{"key":"11852_CR14","unstructured":"Zhang Z et al (2022) Effectively leveraging Multi-modal Features for Movie Genre Classification. arXiv preprint arXiv:2203.13281"},{"key":"11852_CR15","doi-asserted-by":"crossref","unstructured":"Nambiar G, Roy P, Singh D (2020) Multi Modal Genre Classification of Movies. In: 2020 IEEE International Conference for Innovation in Technology (INOCON). IEEE","DOI":"10.1109\/INOCON50539.2020.9298385"},{"key":"11852_CR16","doi-asserted-by":"publisher","first-page":"123","DOI":"10.1016\/j.neucom.2018.09.042","volume":"322","author":"J Wehrmann","year":"2018","unstructured":"Wehrmann J, Cerri R, Barros RC (2018) Movie genre classification: a multi-label approach based on convolutions through time. Neurocomputing 322:123\u2013132. https:\/\/doi.org\/10.1016\/j.neucom.2018.09.042","journal-title":"Neurocomputing"},{"issue":"29","key":"11852_CR17","doi-asserted-by":"publisher","first-page":"32469","DOI":"10.1007\/s11042-022-12961-6","volume":"81","author":"NK Rajput","year":"2022","unstructured":"Rajput NK, Grover BA (2022) A multi-label movie genre classification scheme based on the movie\u2019s subtitles. Multimedia Tools Appl 81(29):32469\u201332490. https:\/\/doi.org\/10.1007\/s11042-022-12961-6","journal-title":"Multimedia Tools Appl"},{"key":"11852_CR18","doi-asserted-by":"publisher","first-page":"e2945","DOI":"10.7717\/peerj-cs.2945","volume":"11","author":"F Shaukat","year":"2025","unstructured":"Shaukat F, Ejaz N, Ashraf Z, Alnfiai MM, Alotaibi NN, Alnefaie SMM (2025) An interpretable multi-transformer ensemble for text-based movie genre classification. PeerJ Comput Sci 11:e2945. https:\/\/doi.org\/10.7717\/peerj-cs.2945","journal-title":"PeerJ Comput Sci"},{"issue":"4","key":"11852_CR19","doi-asserted-by":"publisher","first-page":"5763","DOI":"10.1007\/s11042-022-13418-6","volume":"82","author":"T Behrouzi","year":"2023","unstructured":"Behrouzi T, Toosi R, Mohammad Ali Akhaee (2023) Multimodal movie genre classification using recurrent neural network. Multimed Tools Appl 82(4):5763\u20135784","journal-title":"Multimedia Tools Appl"},{"key":"11852_CR20","doi-asserted-by":"publisher","unstructured":"Sirirattanajakarin S, Thusaranon P (2019) Movie genre in multi-label classification using semantic extraction from only movie poster. In: Proceedings of the International Conference on Computer and Communication Management (ICCCM 2019), pp. 1\u20136. ACM. https:\/\/doi.org\/10.1145\/3348445.3348475","DOI":"10.1145\/3348445.3348475"},{"key":"11852_CR21","doi-asserted-by":"publisher","unstructured":"Visutsak P, Pensiri F, Netisopakul P (2024) Genre classification of movie trailers using spectrogram analysis and machine learning. In: Proceedings of the IEEE Black Sea Conference on Communications and Networking (BlackSeaCom 2024), pp. 1\u20136. IEEE. https:\/\/doi.org\/10.1109\/BLACKSEACOM61746.2024.10646256","DOI":"10.1109\/BLACKSEACOM61746.2024.10646256"},{"issue":"1","key":"11852_CR22","doi-asserted-by":"publisher","first-page":"65","DOI":"10.1007\/s10032-022-00413-8","volume":"26","author":"RASHEED","year":"2023","unstructured":"RASHEED, Assad et al (2023) Cover-based multiple book genre recognition using an improved multimodal network. Int J Doc Anal Recognit (IJDAR) 26(1):65\u201388","journal-title":"Int J Doc Anal Recognit (IJDAR)"},{"key":"11852_CR23","doi-asserted-by":"publisher","DOI":"10.22541\/au.174351635.54917844","author":"S Singhal","year":"2025","unstructured":"Singhal S, Singh V (2025) Cross-domain learning framework for book\u2014movie recommendation with RoBERTa and distilbert in action. Expert Syst (Preprint). https:\/\/doi.org\/10.22541\/au.174351635.54917844","journal-title":"Expert Systems (Preprint)"},{"key":"11852_CR24","doi-asserted-by":"publisher","unstructured":"Jaisankar V, Jayagopi DB (2024) Spectrogrand: Computational creativity driven audiovisuals\u2019 generation from text prompts. In: Proceedings of the Indian Conference on Computer Vision, Graphics and Image Processing (ICVGIP 2024), pp. 1\u20139. ACM. https:\/\/doi.org\/10.1145\/3702250.3702280","DOI":"10.1145\/3702250.3702280"},{"key":"11852_CR25","unstructured":"Fallah H, Bellot P, Bruno E, Murisasco E (2022) Adapting transformers for multi-label text classification. In: Proceedings of CIRCLE 2022 \u2013 Joint Conference of the Information Retrieval Communities in Europe, pp. 1\u201312. HAL. https:\/\/hal.science\/hal-03727927v1"},{"key":"11852_CR26","unstructured":"Oramas S et al (2017) Multi-label music genre classification from audio, text, and images using deep features. arXiv preprint arXiv:1707.04916"},{"key":"11852_CR27","doi-asserted-by":"publisher","unstructured":"Akalp H, \u00c7i\u011fdem EF, Y\u0131lmaz \u015e, B\u00f6l\u00fcc\u00fc N, Can B (2021) Language representation models for music genre classification using lyrics. In: Proceedings of the International Symposium on Electrical, Electronics and Information Engineering (ISEEIE 2021). ACM. https:\/\/doi.org\/10.1145\/3459104.3459171","DOI":"10.1145\/3459104.3459171"},{"key":"11852_CR28","unstructured":"Pizarro S, Zimmermann M, Offermann MS, Reither F (2024) Exploring genre and success classification through song lyrics using DistilBERT: A fun NLP venture. arXiv preprint, arXiv:2407.21068. https:\/\/arxiv.org\/abs\/2407.21068"},{"key":"11852_CR29","doi-asserted-by":"crossref","unstructured":"Liu Z, Mao H, Wu C-Y, Feichtenhofer C, Darrell T, ve Xie S (2022) A convnet for the 2020s, Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, 11976\u201311986","DOI":"10.1109\/CVPR52688.2022.01167"},{"key":"11852_CR30","unstructured":"Dataset (2024) Home page, https:\/\/www.kaggle.com\/datasets\/neha1703\/movie-genre-from-its-poster, last accessed: 2024\/01\/01."},{"key":"11852_CR31","doi-asserted-by":"crossref","unstructured":"Sechidis K, Tsoumakas G, ve Vlahavas I (2011) On the stratification of multi-label data, Machine Learning and Knowledge Discovery in Databases: European Conference, ECML PKDD 2011, Athens, Greece, September 5\u20139, 2011, Proceedings, Part III 22, Springer, 145\u2013158","DOI":"10.1007\/978-3-642-23808-6_10"},{"key":"11852_CR32","unstructured":"Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser \u0141, Polosukhin I (2017) Attention is all you need. In: Proceedings of the 31st International Conference on Neural Information Processing Systems (NIPS\u201917). Curran Associates Inc., Red Hook, NY, pp 6000\u20136010"},{"key":"11852_CR33","doi-asserted-by":"crossref","unstructured":"Devlin J, Chang M-W, Lee K, ve Toutanova K (2019) Bert: Pre-training of deep bidirectional transformers for language understanding. In: Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers), 4171\u20134186","DOI":"10.18653\/v1\/N19-1423"},{"key":"11852_CR34","unstructured":"Sanh V (2019) DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108"},{"issue":"4","key":"11852_CR35","doi-asserted-by":"publisher","first-page":"214","DOI":"10.3390\/info15040214","volume":"15","author":"I Branescu","year":"2024","unstructured":"Branescu I, Grigorescu O, ve Dascalu M (2024) Automated mapping of common vulnerabilities and exposures to mitre att&ck tactics. Information 15(4):214","journal-title":"Information"},{"key":"11852_CR36","doi-asserted-by":"crossref","unstructured":"Chen T, ve Guestrin C (2016) Xgboost: A scalable tree boosting system. In:  Proceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining, 785\u2013794","DOI":"10.1145\/2939672.2939785"},{"key":"11852_CR37","doi-asserted-by":"crossref","unstructured":"Hastie T, Tibshirani R, ve Friedman J (2009) The elements of statistical learning. In: Citeseer","DOI":"10.1007\/978-0-387-84858-7"},{"key":"11852_CR38","unstructured":"IMDb-Movie-Summaries-Dataset (2025) Home page, https:\/\/github.com\/metehanunal\/IMDb-Movie-Summaries-Dataset\/, last accessed: 2025\/11\/07."}],"container-title":["Neural Computing and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-026-11852-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s00521-026-11852-3","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-026-11852-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,25]],"date-time":"2026-03-25T04:54:10Z","timestamp":1774414450000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s00521-026-11852-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3]]},"references-count":38,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2026,3]]}},"alternative-id":["11852"],"URL":"https:\/\/doi.org\/10.1007\/s00521-026-11852-3","relation":{},"ISSN":["0941-0643","1433-3058"],"issn-type":[{"value":"0941-0643","type":"print"},{"value":"1433-3058","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3]]},"assertion":[{"value":"11 July 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"12 January 2026","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"6 March 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors state no Conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"129"}}