{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,4]],"date-time":"2026-05-04T10:20:45Z","timestamp":1777890045513,"version":"3.51.4"},"reference-count":31,"publisher":"SAGE Publications","issue":"1","license":[{"start":{"date-parts":[[2020,2,25]],"date-time":"2020-02-25T00:00:00Z","timestamp":1582588800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["Web Intelligence"],"published-print":{"date-parts":[[2020,3,9]]},"abstract":"<jats:p>With the advent of social media, our online feeds increasingly consist of short, informal, and unstructured text. Instagram is one of the largest social media platforms, containing both text and images. However, most of the prior research on text processing in social media is focused on analyzing Twitter data, and little attention has been paid to text mining of Instagram data. Moreover, many text mining methods rely on training data annotated manually by humans, which in practice is both difficult and expensive to obtain. In this paper, we present methods for weakly supervised text classification of Instagram text. We analyze a corpora of Instagram posts from the fashion domain and train a deep clothing classifier with weak supervision to classify Instagram posts based on the associated text.<\/jats:p>\n                  <jats:p>With our experiments, we demonstrate that in absence of annotated training data, using weak supervision to train models is a viable approach. With weak supervision we were able to label a large dataset in hours, something that would have taken months to do with human annotators. Using the dataset labeled with weak supervision in combination with generative modeling, an [Formula: see text] score of 0.61 is achieved on the task of classifying the image contents of Instagram posts based solely on the associated text, which is on level with human performance.<\/jats:p>","DOI":"10.3233\/web-200428","type":"journal-article","created":{"date-parts":[[2020,2,25]],"date-time":"2020-02-25T11:24:36Z","timestamp":1582629876000},"page":"53-67","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":3,"title":["Deep text classification of Instagram data using word embeddings and weak supervision"],"prefix":"10.1177","volume":"18","author":[{"given":"Kim","family":"Hammar","sequence":"first","affiliation":[{"name":"Department of Software and Computer Systems, KTH Royal Institute of Technology, Stockholm, Sweden. E-mails:\u00a0,\u00a0,\u00a0,\u00a0"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shatha","family":"Jaradat","sequence":"additional","affiliation":[{"name":"Department of Software and Computer Systems, KTH Royal Institute of Technology, Stockholm, Sweden. E-mails:\u00a0,\u00a0,\u00a0,\u00a0"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nima","family":"Dokoohaki","sequence":"additional","affiliation":[{"name":"Department of Software and Computer Systems, KTH Royal Institute of Technology, Stockholm, Sweden. E-mails:\u00a0,\u00a0,\u00a0,\u00a0"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mihhail","family":"Matskin","sequence":"additional","affiliation":[{"name":"Department of Software and Computer Systems, KTH Royal Institute of Technology, Stockholm, Sweden. E-mails:\u00a0,\u00a0,\u00a0,\u00a0"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[2020,2,25]]},"reference":[{"key":"ref001","unstructured":"N.\u00a0Alter, Four ways Instagram is redefining the fashion industry, 2016, Accessed: 2018-04-09."},{"key":"ref002","unstructured":"S.H.\u00a0Bach, D.\u00a0Rodriguez, Y.\u00a0Liu, C.\u00a0Luo, H.\u00a0Shao, C.\u00a0Xia, S.\u00a0Sen, A.\u00a0Ratner, B.\u00a0Hancock, H.\u00a0Alborzi, R.\u00a0Kuchhal, C.\u00a0R\u00e9 and R.\u00a0Malkin, Snorkel DryBell: A case study in deploying weak supervision at industrial scale, CoRR, 2018, http:\/\/arxiv.org\/abs\/1812.00417, abs\/1812.00417."},{"key":"ref003","unstructured":"T.\u00a0Baldwin, P.\u00a0Cook, M.\u00a0Lui, A.\u00a0MacKinlay and L.\u00a0Wang, How noisy social media text, how diffrnt social media sources? in: IJCNLP, Asian Federation of Natural Language Processing \/ ACL, 2013, pp.\u00a0356\u2013364."},{"key":"ref004","doi-asserted-by":"publisher","DOI":"10.1016\/j.bushor.2012.01.007"},{"key":"ref005","doi-asserted-by":"crossref","unstructured":"P.\u00a0Bojanowski, E.\u00a0Grave, A.\u00a0Joulin and T.\u00a0Mikolov, Enriching word vectors with subword information, CoRR, 2016. http:\/\/arxiv.org\/abs\/1607.04606.","DOI":"10.1162\/tacl_a_00051"},{"key":"ref006","doi-asserted-by":"publisher","DOI":"10.3233\/WEB-180389"},{"key":"ref007","doi-asserted-by":"crossref","unstructured":"B.\u00a0Chiu, G.K.O.\u00a0Crichton, A.\u00a0Korhonen and S.\u00a0Pyysalo, How to train good word embeddings for biomedical NLP, in: BioNLP@ACL, Association for Computational Linguistics, 2016, pp.\u00a0166\u2013174.","DOI":"10.18653\/v1\/W16-2922"},{"key":"ref008","first-page":"2493","volume":"12","author":"Collobert R.","year":"2011","journal-title":"J. Mach. Learn. Res."},{"key":"ref009","doi-asserted-by":"crossref","unstructured":"A.\u00a0Conneau, H.\u00a0Schwenk, L.\u00a0Barrault and Y.\u00a0LeCun, Very deep convolutional networks for natural language processing, CoRR, 2016.","DOI":"10.18653\/v1\/E17-1104"},{"key":"ref010","doi-asserted-by":"crossref","unstructured":"B.\u00a0Dhingra, Z.\u00a0Zhou, D.\u00a0Fitzpatrick, M.\u00a0Muehl and W.W.\u00a0Cohen, Tweet2Vec: Character-based distributed representations for social media, CoRR, 2016.","DOI":"10.18653\/v1\/P16-2044"},{"key":"ref011","doi-asserted-by":"crossref","unstructured":"B.\u00a0Eisner, T.\u00a0Rockt\u00e4schel, I.\u00a0Augenstein, M.\u00a0Bosnjak and S.\u00a0Riedel, emoji2vec: Learning emoji representations from their description, CoRR, 2016, http:\/\/arxiv.org\/abs\/1609.08359, abs\/1609.08359.","DOI":"10.18653\/v1\/W16-6208"},{"key":"ref012","doi-asserted-by":"publisher","DOI":"10.1145\/503104.503110"},{"key":"ref013","unstructured":"K.\u00a0Gimpel, N.\u00a0Schneider, B.\u00a0O\u2019Connor, D.\u00a0Das, D.\u00a0Mills, J.\u00a0Eisenstein, M.\u00a0Heilman, D.\u00a0Yogatama, J.\u00a0Flanigan and N.A.\u00a0Smith, Part-of-speech tagging for Twitter: Annotation, features, and experiments, in: Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies: Short Papers \u2013 Volume\u00a02, HLT \u201911, Association for Computational Linguistics, Stroudsburg, PA, USA, 2011, pp.\u00a042\u201347, http:\/\/dl.acm.org\/citation.cfm?id=2002736.2002747. ISBN 978-1-932432-88-6."},{"key":"ref014","unstructured":"Google, word2vec, 2013. https:\/\/code.google.com\/archive\/p\/word2vec\/."},{"key":"ref015","doi-asserted-by":"publisher","DOI":"10.1109\/WI.2018.00-94"},{"key":"ref016","doi-asserted-by":"publisher","DOI":"10.1162\/COLI_a_00237"},{"key":"ref017","doi-asserted-by":"publisher","DOI":"10.1145\/3109859.3109861"},{"key":"ref018","doi-asserted-by":"crossref","unstructured":"Y.\u00a0Kim, Convolutional neural networks for sentence classification, in: EMNLP, ACL, 2014, pp.\u00a01746\u20131751.","DOI":"10.3115\/v1\/D14-1181"},{"key":"ref019","doi-asserted-by":"publisher","DOI":"10.1109\/ICDMW.2011.171"},{"issue":"8","key":"ref020","first-page":"707","volume":"10","author":"Levenshtein V.I.","year":"1966","journal-title":"Soviet Physics Doklady"},{"key":"ref021","doi-asserted-by":"publisher","DOI":"10.3115\/1118108.1118117"},{"key":"ref022","unstructured":"M.\u00a0Lui and T.\u00a0Baldwin, Langid.Py: An off-the-shelf language identification tool, in: Proceedings of the ACL 2012 System Demonstrations, ACL \u201912, Association for Computational Linguistics, Stroudsburg, PA, USA, 2012, pp.\u00a025\u201330, http:\/\/dl.acm.org\/citation.cfm?id=2390470.2390475."},{"key":"ref023","unstructured":"V.\u00a0Major, A.\u00a0Surkis and Y.\u00a0Aphinyanaphongs, Utility of general and specific word embeddings for classifying translational stages of research, CoRR, 2017, http:\/\/arxiv.org\/abs\/1705.06262 abs\/1705.06262."},{"key":"ref024","doi-asserted-by":"crossref","unstructured":"J.\u00a0Pennington, R.\u00a0Socher and C.D.\u00a0Manning, Glove: Global vectors for word representation, in: EMNLP, ACL, 2014, pp.\u00a01532\u20131543.","DOI":"10.3115\/v1\/D14-1162"},{"key":"ref025","doi-asserted-by":"crossref","unstructured":"A.\u00a0Ratner, S.H.\u00a0Bach, H.R.\u00a0Ehrenberg, J.A.\u00a0Fries, S.\u00a0Wu and C.\u00a0R\u00e9, Snorkel: Rapid training data creation with weak supervision, CoRR, 2017, http:\/\/arxiv.org\/abs\/1711.10160.","DOI":"10.14778\/3157794.3157797"},{"key":"ref026","unstructured":"A.J.\u00a0Ratner, C.M.\u00a0De Sa, S.\u00a0Wu, D.\u00a0Selsam and C.\u00a0R\u00e9, Data programming: Creating large training sets, quickly, in: Advances in Neural Information Processing Systems 29, D.D.\u00a0Lee, M.\u00a0Sugiyama, U.V.\u00a0Luxburg, I.\u00a0Guyon and R.\u00a0Garnett, eds, Curran Associates, Inc., 2016, pp.\u00a03567\u20133575."},{"key":"ref027","unstructured":"A.\u00a0Ritter, C.\u00a0Cherry and B.\u00a0Dolan, Unsupervised modeling of Twitter conversations, in: Human Language Technologies: The 2010 Annual Conference of the North American Chapter of the Association for Computational Linguistics, HLT \u201910, Association for Computational Linguistics, Stroudsburg, PA, USA, 2010, pp.\u00a0172\u2013180. ISBN 1-932432-65-5."},{"key":"ref028","unstructured":"A.\u00a0Ritter, S.\u00a0Clark, Mausam and O.\u00a0Etzioni, Named entity recognition in tweets: An experimental study, in: Proceedings of the Conference on Empirical Methods in Natural Language Processing, EMNLP \u201911, Association for Computational Linguistics, Stroudsburg, PA, USA, 2011, pp.\u00a01524\u20131534, http:\/\/dl.acm.org\/citation.cfm?id=2145432.2145595. ISBN 978-1-937284-11-4."},{"key":"ref029","doi-asserted-by":"publisher","DOI":"10.1145\/2339530.2339704"},{"key":"ref030","unstructured":"A.J.\u00a0Tixier, M.\u00a0Vazirgiannis and M.R.\u00a0Hallowell, Word embeddings for the construction domain, CoRR, 2016, http:\/\/arxiv.org\/abs\/1610.09333, abs\/1610.09333."},{"key":"ref031","doi-asserted-by":"crossref","unstructured":"R.\u00a0Zsuzsanna Albert and A.L.\u00a0Barabasi, Statistical mechanics of complex networks\n                      74\n                      (2001).","DOI":"10.1103\/RevModPhys.74.47"}],"container-title":["Web Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.3233\/WEB-200428","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.3233\/WEB-200428","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.3233\/WEB-200428","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T05:27:17Z","timestamp":1777613237000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/full\/10.3233\/WEB-200428"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,2,25]]},"references-count":31,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2020,3,9]]}},"alternative-id":["10.3233\/WEB-200428"],"URL":"https:\/\/doi.org\/10.3233\/web-200428","relation":{},"ISSN":["2405-6456","2405-6464"],"issn-type":[{"value":"2405-6456","type":"print"},{"value":"2405-6464","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,2,25]]}}}