{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T23:55:34Z","timestamp":1784246134674,"version":"3.55.0"},"reference-count":69,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2022,1,8]],"date-time":"2022-01-08T00:00:00Z","timestamp":1641600000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Knowl. Discov. Data"],"published-print":{"date-parts":[[2022,8,31]]},"abstract":"<jats:p>\n            Domain-specific keyword extraction is a vital task in the field of text mining. There are various research tasks, such as spam e-mail classification, abusive language detection, sentiment analysis, and emotion mining, where a set of domain-specific keywords (aka lexicon) is highly effective. Existing works for keyword extraction list all keywords rather than\n            <jats:italic>domain-specific<\/jats:italic>\n            keywords from a document corpus. Moreover, most of the existing approaches perform well on formal document corpuses but fail on noisy and informal user-generated content in online social media. In this article, we present a hybrid approach by jointly modeling the local and global contextual semantics of words, utilizing the strength of distributional word representation and contrasting-domain corpus for domain-specific keyword extraction. Starting with a seed set of a few domain-specific keywords, we model the text corpus as a weighted word-graph. In this graph, the initial weight of a node (word) represents its semantic association with the target domain calculated as a linear combination of three semantic association metrics, and the weight of an edge connecting a pair of nodes represents the co-occurrence count of the respective words. Thereafter, a modified PageRank method is applied to the word-graph to identify the most relevant words for expanding the initial set of domain-specific keywords. We evaluate our method over both formal and informal text corpuses (comprising six datasets), and show that it performs significantly better in comparison to state-of-the-art methods. Furthermore, we generalize our approach to handle the language-agnostic case, and show that it outperforms existing language-agnostic approaches.\n          <\/jats:p>","DOI":"10.1145\/3494560","type":"journal-article","created":{"date-parts":[[2022,1,8]],"date-time":"2022-01-08T20:51:00Z","timestamp":1641675060000},"page":"1-30","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":20,"title":["Domain-Specific Keyword Extraction Using Joint Modeling of Local and Global Contextual Semantics"],"prefix":"10.1145","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3387-4743","authenticated-orcid":false,"given":"Muhammad","family":"Abulaish","sequence":"first","affiliation":[{"name":"South Asian University, Chanakyapuri, New Delhi, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mohd","family":"Fazil","sequence":"additional","affiliation":[{"name":"South Asian University, Chanakyapuri, New Delhi, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mohammed J.","family":"Zaki","sequence":"additional","affiliation":[{"name":"Rensselaer Polytechnic Institute, Troy, NY"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,1,8]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/COMSNETS.2019.8711451"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/3282373.3282421"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jksuci.2014.03.014"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.5555\/3270332.3270375"},{"key":"e_1_3_2_6_2","first-page":"180","volume-title":"Proceedings of the Italian Research Conference on Digital Libraries","author":"Basaldella Marco","year":"2018","unstructured":"Marco Basaldella, Elisa Antolli, Giuseppe Serra, and Carlo Tasso. 2018. Bidirectional LSTM recurrent neural networkfor keyphrase extraction. In Proceedings of the Italian Research Conference on Digital Libraries. Springer, Cham, 180\u2013187."},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.5555\/2457524.2457617"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.5555\/944919.944966"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2017.12.025"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.5555\/944919.944937"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1002\/asi.4630250505"},{"key":"e_1_3_2_12_2","first-page":"1","volume-title":"Proceedings of the 10th International Conference on Terminology and Artificial Intelligence","author":"Bordea Georgeta","year":"2013","unstructured":"Georgeta Bordea, Paul Buitelaar, and Tamara Polajnar. 2013. Domain-independent term extraction through domain modelling. In Proceedings of the 10th International Conference on Terminology and Artificial Intelligence. ICSA, 1\u20138."},{"key":"e_1_3_2_13_2","first-page":"1","volume-title":"Proceedings of the 6th International Conference on Recent Advances in Natural Language Processing","author":"Brewster Christopher","year":"2007","unstructured":"Christopher Brewster, Jose Iria, Ziqi Zhang, Fabio Ciravegna, Louise Guthrie, and Yorick Wilks. 2007. Dynamic iterative ontology learning. In Proceedings of the 6th International Conference on Recent Advances in Natural Language Processing. ACL, 1\u20135."},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0169-7552(98)00110-X"},{"issue":"11","key":"e_1_3_2_15_2","first-page":"1","article-title":"Us and them: Identifying cyber hate onTwitter across multiple protectedcharacteristics","volume":"5","author":"Burnap Pete","year":"2016","unstructured":"Pete Burnap and Matthew L. Williams. 2016. Us and them: Identifying cyber hate onTwitter across multiple protectedcharacteristics. EPJ Data Science 5, 11 (2016), 1\u201315.","journal-title":"EPJ Data Science"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1075\/term.9.2.05chu"},{"issue":"1","key":"e_1_3_2_17_2","first-page":"115","article-title":"Using statistics in lexical analysis","volume":"1","author":"Church Kenneth","year":"1991","unstructured":"Kenneth Church, William Gale, Patrick Hanks, and Donald Hindle. 1991. Using statistics in lexical analysis. Lexical Acquisition: Exploiting On-Line Resources to Build a Lexicon 1, 1 (1991), 115\u2013164.","journal-title":"Lexical Acquisition: Exploiting On-Line Resources to Build a Lexicon"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3489088.3489095"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1108\/eb026683"},{"key":"e_1_3_2_20_2","first-page":"61","volume-title":"Proceedings of the Symposium on Statistical Association Methods For Mechanized Documentation","author":"Dennis Sally F.","year":"1964","unstructured":"Sally F. Dennis. 1964. The construction of a thesaurus automatically from a sample of text. In Proceedings of the Symposium on Statistical Association Methods For Mechanized Documentation. ACL, 61\u2013148."},{"issue":"1","key":"e_1_3_2_21_2","first-page":"100","article-title":"sCAKE: Semantic connectivity aware keyword extraction","volume":"477","author":"Duari Swagata","year":"2018","unstructured":"Swagata Duari and Vasudha Bhatnagar. 2018. sCAKE: Semantic connectivity aware keyword extraction. Information Sciences 477, 1 (2018), 100\u2013117.","journal-title":"Information Sciences"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2007.01.015"},{"key":"e_1_3_2_23_2","first-page":"417","volume-title":"Proceedings of the 5th International Conference on Language Resources and Evaluation","author":"Esuli Andrea","year":"2006","unstructured":"Andrea Esuli and Fabrizio Sebastiani. 2006. SENTIWORDNET: A publicly available lexical resourcefor opinion mining. In Proceedings of the 5th International Conference on Language Resources and Evaluation. European Language Resources Association, 417\u2013422."},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.5555\/3297863.3297932"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17-1102"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-15712-8_13"},{"key":"e_1_3_2_27_2","first-page":"491","volume-title":"Proceedings of the 12th International Conference on Web and Social Media","author":"Founta Antigoni-Maria","year":"2018","unstructured":"Antigoni-Maria Founta, Constantinos Djouvas, Despoina Chatzakou, Ilias Leontiadis, Jeremy Blackburn, Gianluca Stringhini, Athena Vakali, Michael Sirivianos, and Nicolas Kourtellis. 2018. Large scale crowdsourcing and characterization of Twitter abusive behavior. In Proceedings of the 12th International Conference on Web and Social Media. Association for the Advancement of Artificial Intelligence, 491\u2013500."},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/3307339.3342147"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.5555\/3298023.3298031"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1002\/asi.4630260402"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P14-1119"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSC.2007.71"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1140\/epjb\/e2008-00206-x"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.3115\/1119355.1119383"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1075\/term.3.2.03kag"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1108\/eb026526"},{"key":"e_1_3_2_37_2","first-page":"94","volume-title":"Proceedings of the Australasian Language Technology Association Workshop","author":"Kim Su N.","year":"2009","unstructured":"Su N. Kim, Timothy Baldwin, and Min Y. Kan. 2009. Extracting domain-specific words - a statistical approach. In Proceedings of the Australasian Language Technology Association Workshop. ACL, 94\u201398."},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1075\/term.14.2.05kit"},{"key":"e_1_3_2_39_2","first-page":"67","volume-title":"Predictive Analysis on Twitter: Techniques and Applications","author":"Kursuncu Ugur","year":"2018","unstructured":"Ugur Kursuncu, Manas Gaur, Usha Lokala, Krishna prasad Thirunarayan, Amit Sheth, and Budak Arpinar. 2018. Predictive Analysis on Twitter: Techniques and Applications. Springer, Cham, Chapter Emerging Research Challenges and Opportunities in Computational Social Network Analysis and Mining, 67\u2013104."},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-18029-3_13"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N18-2100"},{"key":"e_1_3_2_42_2","first-page":"392","volume-title":"Proceedings of the 16th International Florida Artificial Intelligence Research Society Conference","author":"Matsuo Yutaka","year":"2003","unstructured":"Yutaka Matsuo and Mitsuru Ishizuka. 2003. Keyword extraction from a single documentusing word co-occurrence statistical information. In Proceedings of the 16th International Florida Artificial Intelligence Research Society Conference. AAAI, 392\u2013396."},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17-1054"},{"key":"e_1_3_2_44_2","first-page":"404","volume-title":"Proceedings of the International Conferences Empirical Methods in Natural Language Processing","author":"Mihalcea Rada","year":"2004","unstructured":"Rada Mihalcea and Paul Tarau. 2004. TextRank: Bringing order into text. In Proceedings of the International Conferences Empirical Methods in Natural Language Processing. ACL, 404\u2013411."},{"key":"e_1_3_2_45_2","unstructured":"Tomas Mikolov Kai Chen Greg Corrado and Jeffrey Dean. 2013. Efficient estimation of word representations inVector space. Computing Research Repository (CoRR) . 1\u201312. https:\/\/arxiv.org\/abs\/1301.3781."},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/219717.219748"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1111\/j.1467-8640.2012.00460.x"},{"key":"e_1_3_2_48_2","first-page":"1","volume-title":"The PageRank Citation Ranking: Bringing Order to the Web","author":"Page Lawrence","year":"1999","unstructured":"Lawrence Page, Sergey Brin, Rajeev Motwani, and Terry Winograd. 1999. The PageRank Citation Ranking: Bringing Order to the Web. Technical Report. Stanford InfoLab, 1\u201317."},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2018.06.004"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2008-537"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1162"},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1162\/coli_a_00034"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1145\/278459.258529"},{"key":"e_1_3_2_54_2","volume-title":"Automatic Keyword Extractionfrom Individual Documents","author":"Rose Stuart","year":"2010","unstructured":"Stuart Rose, Dave Engel, Nick Cramer, and Wendy Cowley. 2010. Automatic Keyword Extractionfrom Individual Documents. John Wiley & Sons."},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1163"},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.5555\/3192424.3192560"},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2013.131"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P14-2070"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D16-1198"},{"key":"e_1_3_2_60_2","first-page":"1083","volume-title":"Proceedings of the 4th International Conference on Language Resources and Evaluation","author":"Strapparava Carlo","year":"2004","unstructured":"Carlo Strapparava and Alessandro Valitutti. 2004. WordNet-Affect: An affective extension of WordNet. In Proceedings of the 4th International Conference on Language Resources and Evaluation. European Language Resources Association, 1083\u20131086."},{"key":"e_1_3_2_61_2","unstructured":"Kabir Taneja and Kriti M. Shah. 2019. The conflict in Jammu and Kashmir and the convergence of technology and terrorism. Royal United Services Institute for Defence and Security Studies Paper No. 11 (2019) 1\u201314."},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1145\/2948072"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.5555\/1620163.1620205"},{"key":"e_1_3_2_64_2","first-page":"1","volume-title":"Proceedings of the Software Engineering Research Conference","author":"Wang Rui","year":"2015","unstructured":"Rui Wang, Wei Liu, and Chris McDonald. 2015. Corpus-independent generic keyphrase extraction using word embedding vectors. In Proceedings of the Software Engineering Research Conference. 1\u20138."},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1145\/3301019.3323905"},{"issue":"5","key":"e_1_3_2_66_2","first-page":"1","article-title":"Automatic keyphrase extraction using word embeddings","volume":"23","author":"Zhang Yuxiang","year":"2019","unstructured":"Yuxiang Zhang, Huan Liu, Suge Wang, W. H. Ip., Wei Fan, and Chunjing Xiao. 2019. Automatic keyphrase extraction using word embeddings. Soft Computing 23, 5 (2019), 1\u201316.","journal-title":"Soft Computing"},{"key":"e_1_3_2_67_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2019.11.083"},{"key":"e_1_3_2_68_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2018.2865589"},{"key":"e_1_3_2_69_2","doi-asserted-by":"publisher","DOI":"10.1145\/3201408"},{"key":"e_1_3_2_70_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0378-4371(03)00625-3"}],"container-title":["ACM Transactions on Knowledge Discovery from Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3494560","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3494560","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:31:16Z","timestamp":1750188676000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3494560"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,1,8]]},"references-count":69,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2022,8,31]]}},"alternative-id":["10.1145\/3494560"],"URL":"https:\/\/doi.org\/10.1145\/3494560","relation":{},"ISSN":["1556-4681","1556-472X"],"issn-type":[{"value":"1556-4681","type":"print"},{"value":"1556-472X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,1,8]]},"assertion":[{"value":"2021-03-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-10-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-01-08","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}