{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,16]],"date-time":"2026-02-16T17:22:11Z","timestamp":1771262531194,"version":"3.50.1"},"reference-count":0,"publisher":"TechForum Publishing Group","issue":"1","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Bull. Comput. Data Sci."],"published-print":{"date-parts":[[2025,3,30]]},"abstract":"<jats:p>Emoji usage and code-mixed language have become central to informal online communication, especially in multilingual communities. Predicting suitable emojis and inferring sentiment and emotion from noisy, code-mixed social media text is challenging due to non-standard spelling, cross-script mixing, and strong pragmatic effects. Prior work has proposed specialized encoders and multi-task frameworks for joint prediction of multi-label emojis, sentiment, and emotion on English\u2013Hindi code-mixed tweets. However, these approaches are built on relatively shallow architectures and do not fully exploit recent advances in self-supervised and large language model (LLM) based representation learning. In this paper, we introduce CodeMixLM, a self-supervised transformer model specialized for Hinglish code-mixed text. Starting from a multilingual transformer backbone, we continue pretraining on large-scale unlabeled code-mixed social media data with three auxiliary objectives: masked span denoising, script-aware language identification, and emoji-aware contrastive learning. We then fine-tune CodeMixLM in a multi-task setting for (i) multi-label emoji prediction, (ii) three-way sentiment classification, and (iii) seven-way emotion classification on the SENTIMOJI dataset of English\u2013Hindi code-mixed tweets. Across all tasks, CodeMixLM substantially outperforms prior task-specific architectures and strong multilingual transformer baselines, improving macro-F1 for emoji prediction by up to several points while also yielding better calibration and label efficiency under reduced supervision. Detailed analyses show that self-supervised code-mixed pretraining (1) improves robustness to spelling variants and code-switching patterns, and (2) better captures the interaction between emojis, sentiment, and emotion. Our results highlight the importance of domain-specialized self-supervised learning for code-mixed NLP and offer a stronger baseline for future work on emoji-aware affective computing.<\/jats:p>","DOI":"10.71448\/bcds2561-3","type":"journal-article","created":{"date-parts":[[2026,2,16]],"date-time":"2026-02-16T16:33:31Z","timestamp":1771259611000},"page":"39-60","source":"Crossref","is-referenced-by-count":0,"title":["Self-Supervised Code-Mixed Representation Learning for Multi-Label Emoji, Sentiment, and Emotion Prediction"],"prefix":"10.71448","volume":"6","author":[{"name":"Department of Computer Science and Engineering, Indian Institute of Technology Bombay","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Asif","family":"Shehzad","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Rahul","family":"Sharma","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"name":"Department of Computer Science and Engineering, Indian Institute of Technology Bombay","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Pushpak","family":"Bhattacharyya","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"name":"Department of Computer Science and Engineering, Indian Institute of Technology Bombay","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"52394","published-online":{"date-parts":[[2025,3,30]]},"container-title":["Bulletin of Computer and Data Sciences"],"original-title":[],"deposited":{"date-parts":[[2026,2,16]],"date-time":"2026-02-16T16:33:31Z","timestamp":1771259611000},"score":1,"resource":{"primary":{"URL":"https:\/\/bcds.ch\/self-supervised-code-mixed-representation-learning-for-multi-label-emoji-sentiment-and-emotion-prediction\/"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,3,30]]},"references-count":0,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2025,3,30]]},"published-print":{"date-parts":[[2025,3,30]]}},"URL":"https:\/\/doi.org\/10.71448\/bcds2561-3","relation":{},"ISSN":["3072-2926"],"issn-type":[{"value":"3072-2926","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,3,30]]}}}