{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T05:45:29Z","timestamp":1784180729061,"version":"3.55.0"},"reference-count":33,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2023,5,13]],"date-time":"2023-05-13T00:00:00Z","timestamp":1683936000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,5,13]],"date-time":"2023-05-13T00:00:00Z","timestamp":1683936000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Centre for Research & Technology Hellas"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Multimed Info Retr"],"published-print":{"date-parts":[[2023,6]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Image memes and specifically their widely known variation <jats:italic>image macros<\/jats:italic> are a special new media type that combines text with images and are used in social media to playfully or subtly express humor, irony, sarcasm and even hate. It is important to accurately retrieve image memes from social media to better capture the cultural and social aspects of online phenomena and detect potential issues (hate-speech, disinformation). Essentially, the background image of an image macro is a regular image easily recognized as such by humans but cumbersome for the machine to do so due to feature map similarity with the complete image macro. Hence, accumulating suitable feature maps in such cases can lead to deep understanding of the notion of image memes. To this end, we propose a methodology, called <jats:italic>visual part utilization<\/jats:italic>, that utilizes the visual part of image memes as instances of the <jats:italic>regular image class<\/jats:italic> and the initial image memes as instances of the <jats:italic>image meme class<\/jats:italic> to force the model to concentrate on the critical parts that characterize an image meme. Additionally, we employ a trainable attention mechanism on top of a standard ViT architecture to enhance the model\u2019s ability to focus on these critical parts and make the predictions interpretable. Several training and test scenarios involving web-scraped regular images of controlled text presence are considered for evaluating the model in terms of robustness and accuracy. The findings indicate that light visual part utilization combined with sufficient text presence during training provides the best and most robust model, surpassing state of the art. Source code and dataset are available at <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/mever-team\/memetector\">https:\/\/github.com\/mever-team\/memetector<\/jats:ext-link>.<\/jats:p>","DOI":"10.1007\/s13735-023-00277-6","type":"journal-article","created":{"date-parts":[[2023,5,15]],"date-time":"2023-05-15T11:56:36Z","timestamp":1684151796000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":11,"title":["MemeTector: enforcing deep focus for meme detection"],"prefix":"10.1007","volume":"12","author":[{"given":"Christos","family":"Koutlis","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Manos","family":"Schinas","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Symeon","family":"Papadopoulos","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2023,5,13]]},"reference":[{"key":"277_CR1","doi-asserted-by":"publisher","unstructured":"Afridi TH, Alam A, Khan MN, et\u00a0al (2021) A multimodal memes classification: A survey and open research issues. In: Ben\u00a0Ahmed M, Rakip Kara\u015f \u0130, Santos D, et\u00a0al (eds) Innovations in smart cities applications, Vol. 4. Springer International Publishing, Cham, pp 1451\u20131466, https:\/\/doi.org\/10.1007\/978-3-030-66840-2_109","DOI":"10.1007\/978-3-030-66840-2_109"},{"key":"277_CR2","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1155\/2021\/5510253","volume":"2021","author":"A Aggarwal","year":"2021","unstructured":"Aggarwal A, Sharma V, Trivedi A et al (2021) Two-way feature extraction using sequential and multimodal approach for hateful meme classification. Complexity 2021:1\u20137. https:\/\/doi.org\/10.1155\/2021\/5510253","journal-title":"Complexity"},{"key":"277_CR3","doi-asserted-by":"publisher","unstructured":"Amalia A, Sharif A, Haisar F, et\u00a0al (2018) Meme opinion categorization by using optical character recognition (OCR) and na\u00efve Bayes algorithm. In: 2018 third international conference on informatics and computing (ICIC), pp. 1\u20135, https:\/\/doi.org\/10.1109\/IAC.2018.8780410","DOI":"10.1109\/IAC.2018.8780410"},{"key":"277_CR4","unstructured":"Ba LJ, Kiros JR, Hinton GE (2016) Layer normalization. arXiv preprint arXiv:1607.06450"},{"issue":"3","key":"277_CR5","doi-asserted-by":"publisher","first-page":"565","DOI":"10.1515\/cog-2017-0074","volume":"28","author":"B Dancygier","year":"2017","unstructured":"Dancygier B, Vandelanotte L (2017) Internet memes as multimodal constructions. Cogn Linguist 28(3):565\u2013598. https:\/\/doi.org\/10.1515\/cog-2017-0074","journal-title":"Cogn Linguist"},{"key":"277_CR6","unstructured":"Dosovitskiy A, Beyer L, Kolesnikov A, et\u00a0al (2021) An image is worth 16x16 words: transformers for image recognition at scale. In: 9th international conference on learning representations, ICLR 2021, Virtual Event, Austria. OpenReview.net, https:\/\/openreview.net\/forum?id=YicbFdNTTy"},{"key":"277_CR7","doi-asserted-by":"publisher","unstructured":"Fiorucci S (2020) SNK @ DANKMEMES: leveraging pretrained embeddings for multimodal meme detection. In: Basile V, Croce D, Maro M, et\u00a0al (eds) Proceedings of the 7th evaluation campaign of natural language processing and speech tools for Italian, EVALITA 2020, Torino, Italy. Accademia University Press, https:\/\/doi.org\/10.4000\/books.aaccademia.7352","DOI":"10.4000\/books.aaccademia.7352"},{"key":"277_CR8","doi-asserted-by":"publisher","unstructured":"Gaurav D, Shandilya S, Tiwari S, et\u00a0al (2020) A machine learning method for recognizing invasive content in memes. In: Villaz\u00f3n-Terrazas B, Ortiz-Rodr\u00edguez F, Tiwari SM, et\u00a0al (eds) Knowledge graphs and semantic web, KGSWC 2020. Springer International Publishing, Cham, pp 195\u2013213, https:\/\/doi.org\/10.1007\/978-3-030-65384-2_15","DOI":"10.1007\/978-3-030-65384-2_15"},{"key":"277_CR9","doi-asserted-by":"publisher","unstructured":"He K, Zhang X, Ren S, et\u00a0al (2016) Deep residual learning for image recognition. In: IEEE Conference on computer vision and pattern recognition, CVPR 2016. IEEE, pp 770\u2013778, https:\/\/doi.org\/10.1109\/CVPR.2016.90","DOI":"10.1109\/CVPR.2016.90"},{"key":"277_CR10","doi-asserted-by":"publisher","unstructured":"He S, Yang H, Zheng X, et\u00a0al (2019) Massive meme identification and popularity analysis in geopolitics. In: IEEE international conference on intelligence and security informatics, ISI 2019. IEEE, pp 116\u2013121, https:\/\/doi.org\/10.1109\/ISI.2019.8823294","DOI":"10.1109\/ISI.2019.8823294"},{"key":"277_CR11","unstructured":"Hendrycks D, Gimpel K (2016) Gaussian error linear units (GELUs). arXiv preprint arXiv: 1606.08415"},{"key":"277_CR12","unstructured":"Jetley S, Lord NA, Lee N, et\u00a0al (2018) Learn to pay attention. In: 6th international conference on learning representations, ICLR 2018, Vancouver. OpenReview.net, https:\/\/openreview.net\/forum?id=HyzbhfWRW"},{"key":"277_CR13","doi-asserted-by":"publisher","unstructured":"Khedkar S, Karsi P, Ahuja D, et\u00a0al (2022) Hateful memes, offensive or non-offensive! In: Khanna A, Gupta D, Bhattacharyya S, et\u00a0al (eds) international conference on innovative computing and communications, ICICC 2022, Delhi, India. Springer Singapore, pp 609\u2013621, https:\/\/doi.org\/10.1007\/978-981-16-2597-8_52","DOI":"10.1007\/978-981-16-2597-8_52"},{"key":"277_CR14","unstructured":"Kiela D, Firooz H, Mohan A, et\u00a0al (2020) The hateful memes challenge: detecting hate speech in multimodal memes. In: Larochelle H, Ranzato M, Hadsell R, et\u00a0al (eds) Annual conference on neural information processing systems, NeurIPS 2020, Virtual Event. Curran Associates, Inc., pp 2611\u20132624, https:\/\/proceedings.neurips.cc\/paper\/2020\/file\/1b84c4cee2b8b3d823b30e2d604b1878-Paper.pdf"},{"key":"277_CR15","unstructured":"Kiela D, Firooz H, Mohan A, et\u00a0al (2021) The hateful memes challenge: competition report. In: Proceedings of the NeurIPS 2020 competition and demonstration track, Proceedings of machine learning research, vol 133. PMLR, pp 344\u2013360, https:\/\/proceedings.mlr.press\/v133\/kiela21a.html"},{"key":"277_CR16","unstructured":"Loshchilov I, Hutter F (2019) Decoupled weight decay regularization. In: 7th international conference on learning representations, ICLR 2019, New Orleans. OpenReview.net, https:\/\/openreview.net\/forum?id=Bkg6RiCqY7"},{"key":"277_CR17","doi-asserted-by":"publisher","unstructured":"Miliani M, Giorgi G, Rama I, et\u00a0al (2020) DANKMEMES @ evalita 2020: the memeing of life: Memes, multimodality and politics. In: Basile V, Croce D, Maro M, et\u00a0al (eds) Proceedings of the 7th evaluation campaign of natural language processing and speech tools for Italian, EVALITA 2020, Torino. Accademia University Press, https:\/\/doi.org\/10.4000\/books.aaccademia.7330","DOI":"10.4000\/books.aaccademia.7330"},{"key":"277_CR18","unstructured":"Olivieri A, Noris A, Theng A, et\u00a0al (2022) What is a meme, technically speaking? https:\/\/wiki.digitalmethods.net\/Dmi\/WinterSchool2022WhatIsAMeme"},{"issue":"4","key":"277_CR19","doi-asserted-by":"publisher","first-page":"417","DOI":"10.1111\/jcc4.12120","volume":"20","author":"E Segev","year":"2015","unstructured":"Segev E, Nissenbaum A, Stolero N et al (2015) Families and networks of internet memes: the relationship between cohesiveness, uniqueness, and quiddity concreteness. J Comput-Mediat Commun 20(4):417\u2013433. https:\/\/doi.org\/10.1111\/jcc4.12120","journal-title":"J Comput-Mediat Commun"},{"key":"277_CR20","doi-asserted-by":"publisher","unstructured":"Setpal J, Sarti G (2020) ArchiMeDe @ DANKMEMES: a new model architecture for meme detection. In: Basile V, Croce D, Maro M, et\u00a0al (eds) Proceedings of the 7th evaluation campaign of natural language processing and speech tools for Italian, EVALITA 2020, Torino. Accademia University Press https:\/\/doi.org\/10.4000\/books.aaccademia.7405","DOI":"10.4000\/books.aaccademia.7405"},{"key":"277_CR21","doi-asserted-by":"publisher","unstructured":"Sharma P, Ding N, Goodman S, et\u00a0al (2018) Conceptual captions: a cleaned, hypernymed, image alt-text dataset for automatic image captioning. In: Proceedings of the 56th annual meeting of the association for computational linguistics, ACL 2018, Melbourne. Association for Computational Linguistics, pp 2556\u20132565, https:\/\/doi.org\/10.18653\/v1\/P18-1238","DOI":"10.18653\/v1\/P18-1238"},{"key":"277_CR22","doi-asserted-by":"publisher","unstructured":"Shrestha I, Rusert J (2020) NLP_UIOWA at SemEval-2020 Task 8: You\u2019re not the only one cursed with knowledge - multi branch model memotion analysis. In: Proceedings of the 14th workshop on semantic evaluation, SemEval 2020, Barcelona. International Committee for Computational Linguistics, pp 891\u2013900, https:\/\/doi.org\/10.18653\/v1\/2020.semeval-1.113","DOI":"10.18653\/v1\/2020.semeval-1.113"},{"key":"277_CR23","unstructured":"Simonyan K, Zisserman A (2015) Very deep convolutional networks for large-scale image recognition. In: Bengio Y, LeCun Y (eds) 3rd international conference on learning representations, ICLR 2015, San Diego. http:\/\/arxiv.org\/abs\/1409.1556"},{"key":"277_CR24","doi-asserted-by":"publisher","unstructured":"Sinha A, Patekar P, Mamidi R (2019) Unsupervised approach for monitoring satire on social media. In: Proceedings of the 11th forum for information retrieval evaluation, FIRE 2019. Association for Computing Machinery, pp 36\u201341, https:\/\/doi.org\/10.1145\/3368567.3368582","DOI":"10.1145\/3368567.3368582"},{"key":"277_CR25","doi-asserted-by":"publisher","unstructured":"Smitha ES, Sendhilkumar S, Mahalaksmi GS (2018) Meme classification using textual and visual features. In: Hemanth D, Smys S (eds) Computational vision and bio inspired computing, Lecture notes in computational vision and biomechanics, vol\u00a028. Springer International Publishing, pp 1015\u20131031, https:\/\/doi.org\/10.1007\/978-3-319-71767-8_87","DOI":"10.1007\/978-3-319-71767-8_87"},{"key":"277_CR26","unstructured":"Suryawanshi S, Chakravarthi BR, Arcan M, et\u00a0al (2020) Multimodal meme dataset (MultiOFF) for identifying offensive content in image and text. In: Proceedings of the 2nd workshop on trolling, aggression and cyberbullying, TRAC 2020, Marseille. European Language Resources Association, pp 32\u201341, https:\/\/aclanthology.org\/2020.trac-1.6"},{"key":"277_CR27","unstructured":"Tan M, Le Q (2019) Efficientnet: Rethinking model scaling for convolutional neural networks. In: Chaudhuri K, Salakhutdinov R (eds) 36th international conference on machine learning, ICML 2019, Long Beach, vol\u00a097. PMLR, pp 6105\u20136114, https:\/\/proceedings.mlr.press\/v97\/tan19a.html"},{"key":"277_CR28","unstructured":"Vaswani A, Shazeer N, Parmar N, et\u00a0al (2017) Attention is all you need. In: Guyon I, Luxburg UV, Bengio S, et\u00a0al (eds) Annual conference on neural information processing systems, NeurIPS 2017, Long Beach, vol\u00a030. Curran Associates, Inc., https:\/\/proceedings.neurips.cc\/paper\/2017\/file\/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf"},{"key":"277_CR29","doi-asserted-by":"publisher","unstructured":"Vlad GA, Zaharia GE, Cercel DC, et\u00a0al (2020) UPB @ DANKMEMES: Italian memes analysis - employing visual models and graph convolutional networks for meme identification and hate speech detection. In: Basile V, Croce D, Maro M, et\u00a0al (eds) Proceedings of the 7th evaluation campaign of natural language processing and speech tools for Italian, EVALITA 2020, Torino. Accademia University Press, https:\/\/doi.org\/10.4000\/books.aaccademia.7360","DOI":"10.4000\/books.aaccademia.7360"},{"key":"277_CR30","doi-asserted-by":"publisher","unstructured":"Xie L, Natsev A, Kender JR, et\u00a0al (2011) Visual memes in social media: tracking real-world news in youtube videos. In: Proceedings of the 19th ACM international conference on multimedia, MM 2011, Scottsdale. Association for Computing Machinery, pp 53\u201362, https:\/\/doi.org\/10.1145\/2072298.2072307","DOI":"10.1145\/2072298.2072307"},{"key":"277_CR31","doi-asserted-by":"publisher","unstructured":"Ye J, Chen Z, Liu J, et\u00a0al (2020) TextFuseNet: scene text detection with richer fused features. In: Bessiere C (ed) Proceedings of the 29th international joint conference on artificial intelligence, IJCAI 2020, Yokohama, Japan. International joint conferences on artificial intelligence organization, pp 516\u2013522, https:\/\/doi.org\/10.24963\/ijcai.2020\/72","DOI":"10.24963\/ijcai.2020\/72"},{"key":"277_CR32","doi-asserted-by":"publisher","unstructured":"Zannettou S, Caulfield T, Blackburn J, et\u00a0al (2018) On the origins of memes by means of fringe web communities. In: Proceedings of the Internet Measurement Conference, IMC 2018. Association for Computing Machinery, pp 188\u2013202, https:\/\/doi.org\/10.1145\/3278532.3278550","DOI":"10.1145\/3278532.3278550"},{"key":"277_CR33","doi-asserted-by":"publisher","unstructured":"Zhou Y, Chen Z, Yang H (2021) Multimodal learning for hateful memes detection. In: IEEE international conference on multimedia expo workshops, ICMEW 2021, Virtual Event. IEEE, pp 1\u20136, https:\/\/doi.org\/10.1109\/ICMEW53276.2021.9455994","DOI":"10.1109\/ICMEW53276.2021.9455994"}],"container-title":["International Journal of Multimedia Information Retrieval"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s13735-023-00277-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s13735-023-00277-6\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s13735-023-00277-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,6,14]],"date-time":"2023-06-14T15:29:02Z","timestamp":1686756542000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s13735-023-00277-6"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,5,13]]},"references-count":33,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2023,6]]}},"alternative-id":["277"],"URL":"https:\/\/doi.org\/10.1007\/s13735-023-00277-6","relation":{},"ISSN":["2192-6611","2192-662X"],"issn-type":[{"value":"2192-6611","type":"print"},{"value":"2192-662X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,5,13]]},"assertion":[{"value":"4 May 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"30 September 2022","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"20 January 2023","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 May 2023","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors have no competing interests to declare that are relevant to the content of this article.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}],"article-number":"11"}}