{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,3]],"date-time":"2026-04-03T09:32:46Z","timestamp":1775208766336,"version":"3.50.1"},"reference-count":52,"publisher":"PeerJ","license":[{"start":{"date-parts":[[2026,4,3]],"date-time":"2026-04-03T00:00:00Z","timestamp":1775174400000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Ministry of Higher Education, Science, and Technology of the Republic of Indonesia","award":["127\/C3\/DT.05.00\/PL\/2025"],"award-info":[{"award-number":["127\/C3\/DT.05.00\/PL\/2025"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"abstract":"<jats:p>Memes have become a dominant medium of online expression, blending humor, satire, and cultural commentary through visual and textual elements. While often used for entertainment and community building, memes can propagate hate speech in subtle and implicit ways, making automatic detection particularly challenging. This study introduces the Indonesian Multimodal Meme Dataset (INDOMEME), the first expert-annotated multimodal dataset for hateful meme detection in the Indonesian language. The dataset contains 5,023\u00a0memes collected from Facebook and annotated under three complementary schemes: hatefulness, appropriateness, and topical focus. Each meme is further enriched with optical character recognition (OCR) text and machine-generated captions, providing a comprehensive resource for multimodal analysis. Using this dataset, the study conducts extensive experiments addressing four research questions. First, unimodal models (text-only and image-only) are benchmarked against multimodal fusion models, showing that multimodal approaches outperform unimodal baselines; the best multimodal model (IndoBERTweet + Visual Transformers (ViT)) achieves a macro-F1 of 0.820 on hate speech detection and 0.809 on appropriateness classification. Second, several state-of-the-art multimodal large language models (MLLMs), including GPT-4o, Gemini 2.5 Flash, and Gemma3 27B, are evaluated in zero-shot settings, with GPT-4o reaching a macro-F1 of 0.772 for appropriateness detection, although MLLMs remain less effective for hatefulness classification compared to supervised approaches. Finally, multitask learning is explored by jointly modeling appropriateness and hatefulness using a dual-head architecture, demonstrating consistent performance gains across text-only models. These findings underscore the benefit of multimodal resources and multitask architectures in advancing Indonesian meme hate speech detection.<\/jats:p>","DOI":"10.7717\/peerj-cs.3736","type":"journal-article","created":{"date-parts":[[2026,4,3]],"date-time":"2026-04-03T08:36:21Z","timestamp":1775205381000},"page":"e3736","source":"Crossref","is-referenced-by-count":0,"title":["Decoding hate in memes: multimodal and multitask approaches for low-resource Indonesian social media"],"prefix":"10.7717","volume":"12","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0156-6754","authenticated-orcid":true,"given":"Endang Wahyu","family":"Pamungkas","sequence":"first","affiliation":[{"name":"Department of Informatics Engineering, Muhammadiyah University of Surakarta, Surakarta, Central Java, Indonesia"},{"name":"Social Informatics Research Center, Muhammadiyah University of Surakarta, Surakarta, Central Java, Indonesia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Cahyaningtyas Sekar","family":"Wahyuni","sequence":"additional","affiliation":[{"name":"Department of Information Systems, Muhammadiyah University of Surakarta, Surakarta, Central Java, Indonesia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ikhlasul","family":"Amal","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Electronics, Gadjah Mada University, Yogyakarta, Special Region of Yogyakarta, Indonesia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1632-5582","authenticated-orcid":true,"given":"Dian","family":"Purworini","sequence":"additional","affiliation":[{"name":"Social Informatics Research Center, Muhammadiyah University of Surakarta, Surakarta, Central Java, Indonesia"},{"name":"Department of Communication Science, Muhammadiyah University of Surakarta, Surakarta, Central Java, Indonesia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bagus Setya","family":"Rintyarna","sequence":"additional","affiliation":[{"name":"Department of Informatics, Muhammadiyah University of Jember, Jember, East Java, Indonesia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"4443","published-online":{"date-parts":[[2026,4,3]]},"reference":[{"key":"10.7717\/peerj-cs.3736\/ref-1","doi-asserted-by":"crossref","DOI":"10.1109\/CCGE50943.2021.9776440","article-title":"Hateful meme prediction model using multimodal deep learning","author":"Ahmed","year":"2021"},{"issue":"1","key":"10.7717\/peerj-cs.3736\/ref-2","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1057\/s41599-022-01381-4","article-title":"Mental health memes: beneficial or aversive in relation to psychiatric symptoms?","volume":"9","author":"Akram","year":"2022","journal-title":"Humanities and Social Sciences Communications"},{"issue":"17","key":"10.7717\/peerj-cs.3736\/ref-3","doi-asserted-by":"publisher","first-page":"5456","DOI":"10.3390\/s24175456","article-title":"PolyMeme: fine-grained internet meme sensing","volume":"24","author":"Arailopoulos","year":"2024","journal-title":"Sensors"},{"key":"10.7717\/peerj-cs.3736\/ref-4","doi-asserted-by":"publisher","first-page":"e2372","DOI":"10.7717\/peerj-cs.2372","article-title":"Understanding hate speech: the hateinsights dataset and model interpretability","volume":"10","author":"Arshad","year":"2024","journal-title":"PeerJ Computer Science"},{"key":"10.7717\/peerj-cs.3736\/ref-5","doi-asserted-by":"publisher","first-page":"22359","DOI":"10.1109\/access.2024.3361322","article-title":"Multimodal hate speech detection in memes using contrastive language-image pre-training","volume":"12","author":"Arya","year":"2024","journal-title":"IEEE Access"},{"issue":"5","key":"10.7717\/peerj-cs.3736\/ref-6","doi-asserted-by":"publisher","first-page":"2473","DOI":"10.1177\/14614448221088274","article-title":"Memetic persuasion and whatsappification in Indonesia\u2019s 2019 presidential election","volume":"26","author":"Baulch","year":"2024","journal-title":"New Media & Society"},{"issue":"9","key":"10.7717\/peerj-cs.3736\/ref-7","doi-asserted-by":"publisher","first-page":"e0274300","DOI":"10.1371\/journal.pone.0274300","article-title":"Multimodal detection of hateful memes by applying a vision-language pre-training model","volume":"17","author":"Chen","year":"2022","journal-title":"PLOS ONE"},{"key":"10.7717\/peerj-cs.3736\/ref-8","doi-asserted-by":"crossref","first-page":"8440","DOI":"10.18653\/v1\/2020.acl-main.747","article-title":"Unsupervised cross-lingual representation learning at scale","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Conneau","year":"2020"},{"key":"10.7717\/peerj-cs.3736\/ref-9","first-page":"17","article-title":"Detecting hate speech on memes using FixEfficientnet-l2","author":"Dandi","year":"2021"},{"key":"10.7717\/peerj-cs.3736\/ref-10","first-page":"15498","article-title":"BanglaAbuseMeme: a dataset for Bengali abusive meme classification","author":"Das","year":"2023"},{"key":"10.7717\/peerj-cs.3736\/ref-11","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2012.14891","article-title":"Detecting hate speech in multi-modal memes","author":"Das","year":"2020"},{"key":"10.7717\/peerj-cs.3736\/ref-12","first-page":"11","article-title":"Automated hate speech detection and the problem of offensive language","author":"Davidson","year":"2017"},{"key":"10.7717\/peerj-cs.3736\/ref-13","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2010.11929","article-title":"An image is worth 16x16 words: transformers for image recognition at scale","author":"Dosovitskiy","year":"2020"},{"key":"10.7717\/peerj-cs.3736\/ref-14","doi-asserted-by":"publisher","first-page":"114509","DOI":"10.1016\/j.dss.2025.114509","article-title":"Sentiment-aware cross-modal semantic interaction model for harmful meme detection","volume":"196","author":"Duan","year":"2025","journal-title":"Decision Support Systems"},{"issue":"4","key":"10.7717\/peerj-cs.3736\/ref-15","doi-asserted-by":"publisher","first-page":"847","DOI":"10.1080\/1369118x.2021.1993958","article-title":"Fidelis ad mortem: multimodal discourses and ideologies in black lives matter and blue lives matter (NON) humorous memes","volume":"26","author":"Dynel","year":"2023","journal-title":"Information, Communication & Society"},{"key":"10.7717\/peerj-cs.3736\/ref-16","doi-asserted-by":"publisher","first-page":"552","DOI":"10.37394\/23209.2025.22.46","article-title":"Hate speech typology of selected controversial figures on social media: a discourse-analytic perspective","volume":"22","author":"Fauziati","year":"2025","journal-title":"WSEAS Transactions on Information Science and Applications"},{"key":"10.7717\/peerj-cs.3736\/ref-17","doi-asserted-by":"publisher","first-page":"102269","DOI":"10.1016\/j.inffus.2024.102269","article-title":"KERMIT: knowledge-empowered model in harmful meme detection","volume":"106","author":"Grasso","year":"2024","journal-title":"Information Fusion"},{"issue":"11","key":"10.7717\/peerj-cs.3736\/ref-18","doi-asserted-by":"publisher","first-page":"12833","DOI":"10.1007\/s10462-023-10459-7","article-title":"Detecting hate speech in memes: a review","volume":"56","author":"Hermida","year":"2023","journal-title":"Artificial Intelligence Review"},{"key":"10.7717\/peerj-cs.3736\/ref-19","doi-asserted-by":"crossref","first-page":"46","DOI":"10.18653\/v1\/W19-3506","article-title":"Multi-label hate speech and abusive language detection in Indonesian Twitter","volume-title":"Proceedings of the Third Workshop on Abusive Language Online","author":"Ibrohim","year":"2019"},{"issue":"8","key":"10.7717\/peerj-cs.3736\/ref-20","doi-asserted-by":"publisher","first-page":"e18647","DOI":"10.1016\/j.heliyon.2023.e18647","article-title":"Hate speech and abusive language detection in Indonesian social media: progress and challenges","volume":"9","author":"Ibrohim","year":"2023","journal-title":"Heliyon"},{"key":"10.7717\/peerj-cs.3736\/ref-21","doi-asserted-by":"publisher","first-page":"1107","DOI":"10.18280\/isi.280430","article-title":"Application of LSTM and glove word embedding for hate speech detection in Indonesian Twitter data","volume":"28","author":"Imaduddin","year":"2023","journal-title":"Ing\u00e9nierie des Syst\u00e8mes d\u2019Information"},{"key":"10.7717\/peerj-cs.3736\/ref-22","first-page":"14","article-title":"The hateful memes challenge: detecting hate speech in multimodal memes","volume-title":"Proceedings of the 34th International Conference on Neural Information Processing Systems","author":"Kiela","year":"2020"},{"key":"10.7717\/peerj-cs.3736\/ref-23","doi-asserted-by":"crossref","first-page":"10660","DOI":"10.18653\/v1\/2021.emnlp-main.833","article-title":"IndoBERTweet: a pretrained language model for Indonesian Twitter with effective domain-specific vocabulary initialization","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Koto","year":"2021"},{"key":"10.7717\/peerj-cs.3736\/ref-24","doi-asserted-by":"crossref","first-page":"773","DOI":"10.30630\/joiv.7.3.1035","article-title":"Indonesian hate speech detection using IndoBERTweet and BILSTM on Twitter","volume":"7","author":"Kusuma","year":"2023","journal-title":"JOIV: International Journal on Informatics Visualization"},{"key":"10.7717\/peerj-cs.3736\/ref-25","doi-asserted-by":"publisher","first-page":"104848","DOI":"10.1016\/j.actpsy.2025.104848","article-title":"Psycho-physiological impact of virtual non-verbal communication on Gen z workforce: a study of memes","volume":"254","author":"Lamba","year":"2025","journal-title":"Acta Psychologica"},{"key":"10.7717\/peerj-cs.3736\/ref-26","doi-asserted-by":"publisher","first-page":"159","DOI":"10.2307\/2529310","article-title":"The measurement of observer agreement for categorical data","volume":"33","author":"Landis","year":"1977","journal-title":"Biometrics"},{"key":"10.7717\/peerj-cs.3736\/ref-27","article-title":"BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models","author":"Li","year":"2023"},{"key":"10.7717\/peerj-cs.3736\/ref-28","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1908.03557","article-title":"VisualBERT: a simple and performant baseline for vision and language","author":"Li","year":"2019"},{"key":"10.7717\/peerj-cs.3736\/ref-29","doi-asserted-by":"publisher","first-page":"101247","DOI":"10.1016\/j.ssaho.2024.101247","article-title":"Memes, freedom, and resilience to information disorders: information warfare between democracies and autocracies","volume":"11","author":"Liagusha","year":"2025","journal-title":"Social Sciences & Humanities Open"},{"key":"10.7717\/peerj-cs.3736\/ref-30","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2012.12871","article-title":"A multimodal framework for the detection of hateful memes","author":"Lippe","year":"2020"},{"key":"10.7717\/peerj-cs.3736\/ref-31","first-page":"12009","article-title":"Swin transformer v2: scaling up capacity and resolution","author":"Liu","year":"2022"},{"key":"10.7717\/peerj-cs.3736\/ref-32","first-page":"25","article-title":"Visual instruction tuning","volume-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems","author":"Liu","year":"2023"},{"key":"10.7717\/peerj-cs.3736\/ref-33","doi-asserted-by":"publisher","first-page":"100317","DOI":"10.1016\/j.osnem.2025.100317","article-title":"\u2018toxic\u2019 memes: a survey of computational perspectives on the detection and explanation of meme toxicities","volume":"47","author":"Martinez Pandiani","year":"2025","journal-title":"Online Social Networks and Media"},{"key":"10.7717\/peerj-cs.3736\/ref-34","volume-title":"The world made meme: public conversations and participatory media","author":"Milner","year":"2018"},{"issue":"6","key":"10.7717\/peerj-cs.3736\/ref-35","doi-asserted-by":"publisher","first-page":"102360","DOI":"10.1016\/j.ipm.2020.102360","article-title":"Misogyny detection in Twitter: a multilingual and cross-domain study","volume":"57","author":"Pamungkas","year":"2020","journal-title":"Information Processing & Management"},{"issue":"6","key":"10.7717\/peerj-cs.3736\/ref-36","doi-asserted-by":"publisher","first-page":"1175","DOI":"10.14569\/ijacsa.2023.01406125","article-title":"Hate speech detection in Bahasa Indonesia: challenges and opportunities","volume":"14","author":"Pamungkas","year":"2023","journal-title":"International Journal of Advanced Computer Science and Applications"},{"issue":"1","key":"10.7717\/peerj-cs.3736\/ref-37","doi-asserted-by":"publisher","first-page":"33","DOI":"10.1177\/09732586251335714","article-title":"Public opinion towards organisational crisis: insights from the cognitive appraisal theory","volume":"21","author":"Purworini","year":"2026","journal-title":"Journal of Creative Communications"},{"key":"10.7717\/peerj-cs.3736\/ref-38","first-page":"8748","article-title":"Learning transferable visual models from natural language supervision","author":"Radford","year":"2021"},{"issue":"2","key":"10.7717\/peerj-cs.3736\/ref-39","doi-asserted-by":"publisher","first-page":"239","DOI":"10.24815\/siele.v6i2.14020","article-title":"Meme as political criticism towards 2019 Indonesian general election: a critical discourse analysis","volume":"6","author":"Rahardi","year":"2019","journal-title":"Studies in English Language and Education"},{"issue":"1","key":"10.7717\/peerj-cs.3736\/ref-40","doi-asserted-by":"publisher","first-page":"179","DOI":"10.1108\/SAMPJ-03-2024-0303","article-title":"Digitalizing corporate whistleblowing systems: the role of country peers in the adoption of substantive CSR initiatives","volume":"17","author":"Rahmatdi","year":"2024","journal-title":"Sustainability Accounting, Management and Policy Journal"},{"issue":"5","key":"10.7717\/peerj-cs.3736\/ref-41","doi-asserted-by":"publisher","first-page":"103474","DOI":"10.1016\/j.ipm.2023.103474","article-title":"Recognizing misogynous memes: biased models and tricky archetypes","volume":"60","author":"Rizzi","year":"2023","journal-title":"Information Processing & Management"},{"key":"10.7717\/peerj-cs.3736\/ref-42","doi-asserted-by":"publisher","first-page":"e2801","DOI":"10.7717\/peerj-cs.2801","article-title":"Multimodal hate speech detection: a novel deep learning framework for multilingual text and images","volume":"11","author":"Saddozai","year":"2025","journal-title":"PeerJ Computer Science"},{"key":"10.7717\/peerj-cs.3736\/ref-43","doi-asserted-by":"crossref","DOI":"10.7551\/mitpress\/9429.001.0001","volume-title":"Memes in digital culture","author":"Shifman","year":"2013"},{"key":"10.7717\/peerj-cs.3736\/ref-44","first-page":"7","article-title":"A dataset for troll classification of Tamil memes","author":"Suryawanshi","year":"2020"},{"key":"10.7717\/peerj-cs.3736\/ref-45","first-page":"73","article-title":"Cltl@ multimodal hate speech event detection 2024: the winning approach to detecting multimodal hate speech and its targets","author":"Wang","year":"2024"},{"key":"10.7717\/peerj-cs.3736\/ref-46","first-page":"88","article-title":"Hateful symbols or hateful people? Predictive features for hate speech detection on Twitter","volume-title":"Proceedings of the NAACL Student Research Workshop","author":"Waseem","year":"2016"},{"key":"10.7717\/peerj-cs.3736\/ref-47","doi-asserted-by":"crossref","first-page":"843","DOI":"10.18653\/v1\/2020.aacl-main.85","article-title":"IndoNLU: enchmark and resources for evaluating Indonesian natural language understanding","volume-title":"Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing","author":"Wilie","year":"2020"},{"key":"10.7717\/peerj-cs.3736\/ref-48","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2309.17421","article-title":"The dawn of LMMS: preliminary explorations with GPT-4V(ision)","author":"Yang","year":"2023"},{"key":"10.7717\/peerj-cs.3736\/ref-49","doi-asserted-by":"publisher","first-page":"103204","DOI":"10.1016\/j.inffus.2025.103204","article-title":"MM-InstructEval: zero-shot evaluation of (multimodal) large language models on multimodal reasoning tasks","volume":"122","author":"Yang","year":"2025","journal-title":"Information Fusion"},{"issue":"12","key":"10.7717\/peerj-cs.3736\/ref-50","doi-asserted-by":"publisher","first-page":"nwae403","DOI":"10.1093\/nsr\/nwae403","article-title":"A survey on multimodal large language models","volume":"11","author":"Yin","year":"2024","journal-title":"National Science Review"},{"issue":"5","key":"10.7717\/peerj-cs.3736\/ref-51","doi-asserted-by":"publisher","first-page":"207","DOI":"10.3390\/info12050207","article-title":"Multi-task learning for sentiment analysis with hard-sharing and task recognition mechanisms","volume":"12","author":"Zhang","year":"2021","journal-title":"Information"},{"key":"10.7717\/peerj-cs.3736\/ref-52","doi-asserted-by":"crossref","DOI":"10.1109\/ICMEW53276.2021.9455994","article-title":"Multimodal learning for hateful memes detection","author":"Zhou","year":"2021"}],"container-title":["PeerJ Computer Science"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/peerj.com\/articles\/cs-3736.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/peerj.com\/articles\/cs-3736.xml","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/peerj.com\/articles\/cs-3736.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/peerj.com\/articles\/cs-3736.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,3]],"date-time":"2026-04-03T08:36:30Z","timestamp":1775205390000},"score":1,"resource":{"primary":{"URL":"https:\/\/peerj.com\/articles\/cs-3736"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,3]]},"references-count":52,"alternative-id":["10.7717\/peerj-cs.3736"],"URL":"https:\/\/doi.org\/10.7717\/peerj-cs.3736","archive":["CLOCKSS","LOCKSS","Portico"],"relation":{},"ISSN":["2376-5992"],"issn-type":[{"value":"2376-5992","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,3]]},"article-number":"e3736"}}