{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,30]],"date-time":"2026-07-30T14:28:28Z","timestamp":1785421708961,"version":"3.56.0"},"reference-count":0,"publisher":"University of Minho","issue":"1","license":[{"start":{"date-parts":[[2026,6,1]],"date-time":"2026-06-01T00:00:00Z","timestamp":1780272000000},"content-version":"unspecified","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Linguam\u00e1tica"],"abstract":"<jats:p>Messages containing toxic language are a recurring problem on social media, highlighting the urgent need for effective automatic methods to mitigate their impact. Most existing approaches rely on large volumes of annotated data, which are costly, time-consuming, and highly labor-intensive. To address this challenge, this work proposes an ensemble of classifiers for the automatic annotation of toxic language in Portuguese, designed to operate under limited labeled data. The ensemble integrates three complementary strategies: a semi-supervised method based on heterogeneous graphs, a few-shot learning approach, and a Retrieval-Augmented Generation method, both grounded in large language models. The proposal is evaluated across multiple corpora, considering both their original versions and subsets filtered by total inter-annotator agreement. The results indicate that the ensemble exhibits competitive performance, surpassing the best individual method by up to 2% in scenarios of greater balance among the constituent classifiers and maintaining comparable performance in the remaining ones, while preserving moderate to substantial agreement with the original labels, demonstrating its potential for constructing annotated linguistic resources under data scarcity.<\/jats:p>","DOI":"10.21814\/lm.18.1.506","type":"journal-article","created":{"date-parts":[[2026,7,20]],"date-time":"2026-07-20T13:21:15Z","timestamp":1784553675000},"page":"27-53","source":"Crossref","is-referenced-by-count":0,"title":["Um Comit\u00ea de Classificadores para Anota\u00e7\u00e3o Autom\u00e1tica de Linguagem T\u00f3xica em Portugu\u00eas sob Escassez de Dados","An Ensemble of Classifiers for Automatic Annotation of Toxic Language in Portuguese under Data Scarcity"],"prefix":"10.21814","volume":"18","author":[{"given":"Francisco Assis","family":"Ricarte Neto","sequence":"first","affiliation":[{"id":[{"id":"https:\/\/ror.org\/033qmpy33","id-type":"ROR","asserted-by":"publisher"}],"name":"Instituto Federal do Piau\u00ed"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rafael Torres","family":"Anchi\u00eata","sequence":"additional","affiliation":[{"id":[{"id":"https:\/\/ror.org\/05rm1p636","id-type":"ROR","asserted-by":"publisher"}],"name":"Instituto Federal do Maranh\u00e3o"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Raimundo Santos","family":"Moura","sequence":"additional","affiliation":[{"id":[{"id":"https:\/\/ror.org\/00kwnx126","id-type":"ROR","asserted-by":"publisher"}],"name":"Universidade Federal do Piau\u00ed"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Pedro de Alc\u00e2ntara dos Santos","family":"Neto","sequence":"additional","affiliation":[{"id":[{"id":"https:\/\/ror.org\/00kwnx126","id-type":"ROR","asserted-by":"publisher"}],"name":"Universidade Federal do Piau\u00ed"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Andr\u00e9 Macedo","family":"Santana","sequence":"additional","affiliation":[{"id":[{"id":"https:\/\/ror.org\/00kwnx126","id-type":"ROR","asserted-by":"publisher"}],"name":"Universidade Federal do Piau\u00ed"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"9283","published-online":{"date-parts":[[2026,6,1]]},"container-title":["Linguam\u00e1tica"],"original-title":[],"link":[{"URL":"https:\/\/linguamatica.com\/index.php\/linguamatica\/article\/download\/506\/569","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/linguamatica.com\/index.php\/linguamatica\/article\/download\/506\/569","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,20]],"date-time":"2026-07-20T13:21:15Z","timestamp":1784553675000},"score":1,"resource":{"primary":{"URL":"https:\/\/linguamatica.com\/index.php\/linguamatica\/article\/view\/506"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,1]]},"references-count":0,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2026,1,30]]}},"URL":"https:\/\/doi.org\/10.21814\/lm.18.1.506","relation":{},"ISSN":["1647-0818"],"issn-type":[{"value":"1647-0818","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,1]]}}}