{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:25:05Z","timestamp":1750220705709,"version":"3.41.0"},"reference-count":51,"publisher":"Association for Computing Machinery (ACM)","issue":"5","license":[{"start":{"date-parts":[[2020,6,21]],"date-time":"2020-06-21T00:00:00Z","timestamp":1592697600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Inner Mongolia Autonomous Region Key Laboratory of Big Data Research and Application of Agriculture and Animal Husbandry"},{"DOI":"10.13039\/501100004763","name":"Inner Mongolia Natural Science Foundation of China","doi-asserted-by":"crossref","award":["2015MS0628 and 2018MS06005"],"award-info":[{"award-number":["2015MS0628 and 2018MS06005"]}],"id":[{"id":"10.13039\/501100004763","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Research and Application of Big Data Key Technologies in Discipline Inspection and Supervision"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2020,9,30]]},"abstract":"<jats:p>Keywords are considered to be important words in the text and can provide a concise representation of the text. With the surge of unlabeled short text on the Internet, automatic keyword extraction task has proven useful in other information processing applications. Graph-based approaches are prevalent unsupervised models for this task. However, most of these methods emphasize the importance of the relation between words without considering other importance factors. Furthermore, when measuring the importance of a word in a text, the damping factor is set to 0.85 following PageRank. To the best of our knowledge, there is no existing work investigating the impact of the damping factor on the keyword extraction task. In addition, there are few publicly available labeled Chinese short text datasets for this task. In this article, we investigate the importance parts of words in a given document and propose an improved graph-based method for keyword extraction from short documents. Moreover, we analyze the impact of importance factors on performance. We also provide annotated long and short Chinese datasets for this task. The model is performed on Chinese and English datasets, and results show that our model obtains improvements in performance over the previous unsupervised models on short documents. Comparative experiments show that the damping factor is related to the text length, which is neglected in traditional methods.<\/jats:p>","DOI":"10.1145\/3388971","type":"journal-article","created":{"date-parts":[[2020,6,22]],"date-time":"2020-06-22T02:44:42Z","timestamp":1592793882000},"page":"1-15","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["Inside Importance Factors of Graph-Based Keyword Extraction on Chinese Short Text"],"prefix":"10.1145","volume":"19","author":[{"given":"Junjie","family":"Chen","sequence":"first","affiliation":[{"name":"College of Computer Science, Inner Mongolia University, China and College of Computer Science and Information Engineering, Inner Mongolia Agricultural University, Hohhot, Inner Mongolia, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hongxu","family":"Hou","sequence":"additional","affiliation":[{"name":"College of Computer Science, Inner Mongolia University, Hohhot, Inner Mongolia, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jing","family":"Gao","sequence":"additional","affiliation":[{"name":"College of Computer Science and Information Engineering, Inner Mongolia Agricultural University, Hohhot, Inner Mongolia, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2020,6,21]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1007\/3-540-45486-1_4"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/WI-IAT.2012.82"},{"key":"e_1_2_1_3_1","article-title":"Latent Dirichlet allocation","author":"Blei David M.","year":"2003","unstructured":"David M. Blei , Andrew Y. Ng , and Michael I. Jordan . 2003 . Latent Dirichlet allocation . Journal of Machine Learning Research 3 ( Jan. 2003), 993--1022. David M. Blei, Andrew Y. Ng, and Michael I. Jordan. 2003. Latent Dirichlet allocation. Journal of Machine Learning Research 3 (Jan. 2003), 993--1022.","journal-title":"Journal of Machine Learning Research 3"},{"key":"e_1_2_1_4_1","volume-title":"Proceedings of the 26th International Conference on Computational Linguistics: System Demonstrations (COLING\u201916)","author":"Boudin Florian","year":"2016","unstructured":"Florian Boudin . 2016 . pke: An open source Python-based keyphrase extraction toolkit . In Proceedings of the 26th International Conference on Computational Linguistics: System Demonstrations (COLING\u201916) . 69--73. Florian Boudin. 2016. pke: An open source Python-based keyphrase extraction toolkit. In Proceedings of the 26th International Conference on Computational Linguistics: System Demonstrations (COLING\u201916). 69--73."},{"key":"e_1_2_1_5_1","volume-title":"Proceedings of the International Joint Conference on Natural Language Processing (IJCNLP\u201913)","author":"Bougouin Adrien","year":"2013","unstructured":"Adrien Bougouin , Florian Boudin , and B\u00e9atrice Daille . 2013 . TopicRank: Graph-based topic ranking for keyphrase extraction . In Proceedings of the International Joint Conference on Natural Language Processing (IJCNLP\u201913) . 543--551. Adrien Bougouin, Florian Boudin, and B\u00e9atrice Daille. 2013. TopicRank: Graph-based topic ranking for keyphrase extraction. In Proceedings of the International Joint Conference on Natural Language Processing (IJCNLP\u201913). 543--551."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17-1102"},{"key":"e_1_2_1_7_1","volume-title":"Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics","volume":"1","author":"Gu Jiatao","unstructured":"Jiatao Gu , Zhengdong Lu , Hang Li , and Victor O. K. Li . 2016. Incorporating copying mechanism in sequence-to-sequence learning . In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics : Volume 1 , Long Papers. 1631--1640. Jiatao Gu, Zhengdong Lu, Hang Li, and Victor O. K. Li. 2016. Incorporating copying mechanism in sequence-to-sequence learning. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics: Volume 1, Long Papers. 1631--1640."},{"key":"e_1_2_1_8_1","volume-title":"Proceedings of the 23rd International Conference on Computational Linguistics: Posters. 365--373","author":"Hasan Kazi Saidul","year":"2010","unstructured":"Kazi Saidul Hasan and Vincent Ng . 2010 . Conundrums in unsupervised keyphrase extraction: Making sense of the state-of-the-art . In Proceedings of the 23rd International Conference on Computational Linguistics: Posters. 365--373 . Kazi Saidul Hasan and Vincent Ng. 2010. Conundrums in unsupervised keyphrase extraction: Making sense of the state-of-the-art. In Proceedings of the 23rd International Conference on Computational Linguistics: Posters. 365--373."},{"key":"e_1_2_1_9_1","volume-title":"Topic-sensitive PageRank. In Proceedings of the 11th International Conference on World Wide Web. ACM","author":"Haveliwala Taher H.","year":"2002","unstructured":"Taher H. Haveliwala . 2002 . Topic-sensitive PageRank. In Proceedings of the 11th International Conference on World Wide Web. ACM , New York, NY, 517--526. Taher H. Haveliwala. 2002. Topic-sensitive PageRank. In Proceedings of the 11th International Conference on World Wide Web. ACM, New York, NY, 517--526."},{"key":"e_1_2_1_10_1","volume-title":"Kamvar","author":"Haveliwala Taher H.","year":"2003","unstructured":"Taher H. Haveliwala and Ar D . Kamvar . 2003 . The Second Eigenvalue of the Google Matrix. Technical Report. Stanford . Taher H. Haveliwala and Ar D. Kamvar. 2003. The Second Eigenvalue of the Google Matrix. Technical Report. Stanford."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.3115\/1119355.1119383"},{"volume-title":"Proceedings of the 21st International Conference on Computational Linguistics and the 44th Annual Meeting of the Association for Computational Linguistics. 537--544","author":"Hulth Anette","key":"e_1_2_1_12_1","unstructured":"Anette Hulth and Be\u00e1ta B. Megyesi . 2006. A study on automatically extracted keywords in text categorization . In Proceedings of the 21st International Conference on Computational Linguistics and the 44th Annual Meeting of the Association for Computational Linguistics. 537--544 . Anette Hulth and Be\u00e1ta B. Megyesi. 2006. A study on automatically extracted keywords in text categorization. In Proceedings of the 21st International Conference on Computational Linguistics and the 44th Annual Meeting of the Association for Computational Linguistics. 537--544."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/775152.775191"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/E17-2068"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N16-1175"},{"key":"e_1_2_1_16_1","volume-title":"Proceedings of the International Conference on Asian Digital Libraries. 102--111","author":"Krapivin Mikalai","year":"2010","unstructured":"Mikalai Krapivin , Aliaksandr Autayeu , Maurizio Marchese , Enrico Blanzieri , and Nicola Segata . 2010 . n4: Improving machine learning approaches with natural language processing . In Proceedings of the International Conference on Asian Digital Libraries. 102--111 . Mikalai Krapivin, Aliaksandr Autayeu, Maurizio Marchese, Enrico Blanzieri, and Nicola Segata. 2010. n4: Improving machine learning approaches with natural language processing. In Proceedings of the International Conference on Asian Digital Libraries. 102--111."},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1080\/15427951.2004.10129091"},{"volume-title":"Natural Language Processing and Chinese Computing. Communications in Computer and Information Science, Vol.\u00a0496","author":"Li Guangyi","key":"e_1_2_1_18_1","unstructured":"Guangyi Li and Houfeng Wang . 2014. Improved automatic keyword extraction based on TextRank using domain knowledge . In Natural Language Processing and Chinese Computing. Communications in Computer and Information Science, Vol.\u00a0496 . Springer , 403--413. Guangyi Li and Houfeng Wang. 2014. Improved automatic keyword extraction based on TextRank using domain knowledge. In Natural Language Processing and Chinese Computing. Communications in Computer and Information Science, Vol.\u00a0496. Springer, 403--413."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/2699940"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/1772690.1772845"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.5555\/1870658.1870694"},{"key":"e_1_2_1_22_1","volume-title":"Anatole Gershman, and Jaime Carbonell.","author":"Marujo Lu\u00eds","year":"2013","unstructured":"Lu\u00eds Marujo , Miguel Bugalho , Jo\u00e3o Paulo da Silva Neto , Anatole Gershman, and Jaime Carbonell. 2013 . Hourly traffic prediction of news stories. arXiv:1306.4608. Lu\u00eds Marujo, Miguel Bugalho, Jo\u00e3o Paulo da Silva Neto, Anatole Gershman, and Jaime Carbonell. 2013. Hourly traffic prediction of news stories. arXiv:1306.4608."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P15-2105"},{"key":"e_1_2_1_24_1","volume-title":"Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing","volume":"3","author":"Medelyan Olena","unstructured":"Olena Medelyan , Eibe Frank , and Ian H. Witten . 2009. Human-competitive tagging using automatic keyphrase extraction . In Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing : Volume 3 . 1318--1327. Olena Medelyan, Eibe Frank, and Ian H. Witten. 2009. Human-competitive tagging using automatic keyphrase extraction. In Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing: Volume 3. 1318--1327."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17-1054"},{"volume-title":"TextRank: Bringing Order into Texts","author":"Mihalcea Rada","key":"e_1_2_1_26_1","unstructured":"Rada Mihalcea and Paul Tarau . 2004. TextRank: Bringing Order into Texts . Association for Computational Linguistics , Stroudsburg, PA . Rada Mihalcea and Paul Tarau. 2004. TextRank: Bringing Order into Texts. Association for Computational Linguistics, Stroudsburg, PA."},{"key":"e_1_2_1_27_1","unstructured":"Tomas Mikolov Kai Chen Greg Corrado and Jeffrey Dean. 2013a. Efficient estimation of word representations in vector space. arXiv:1301.3781.  Tomas Mikolov Kai Chen Greg Corrado and Jeffrey Dean. 2013a. Efficient estimation of word representations in vector space. arXiv:1301.3781."},{"key":"e_1_2_1_28_1","unstructured":"Tomas Mikolov Ilya Sutskever Kai Chen Greg S. Corrado and Jeff Dean. 2013b. Distributed representations of words and phrases and their compositionality. In Advances in Neural Information Processing Systems. 3111--3119.  Tomas Mikolov Ilya Sutskever Kai Chen Greg S. Corrado and Jeff Dean. 2013b. Distributed representations of words and phrases and their compositionality. In Advances in Neural Information Processing Systems. 3111--3119."},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.artint.2012.06.007"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-77094-7_41"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/ADL.1998.670375"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1162"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1007\/3-540-45735-6_13"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-15850-6_8"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/2396761.2398680"},{"key":"e_1_2_1_37_1","volume-title":"Proceedings of the 5th Italian Information Retrieval Workshop (IIR\u201914)","author":"Qureshi Muhammad Atif","year":"2014","unstructured":"Muhammad Atif Qureshi , Colm O\u2019Riordan , and Gabriella Pasi . 2014 . Exploiting Wikipedia to identify domain-specific key terms\/phrases from a short-text collection . In Proceedings of the 5th Italian Information Retrieval Workshop (IIR\u201914) . 63--74. Muhammad Atif Qureshi, Colm O\u2019Riordan, and Gabriella Pasi. 2014. Exploiting Wikipedia to identify domain-specific key terms\/phrases from a short-text collection. In Proceedings of the 5th Italian Information Retrieval Workshop (IIR\u201914). 63--74."},{"key":"e_1_2_1_38_1","volume-title":"Manning","author":"Toutanova Kristina","year":"2000","unstructured":"Kristina Toutanova and Christopher D . Manning . 2000 . Enriching the knowledge sources used in a maximum entropy part-of-speech tagger. In Proceedings of the 2000 Joint SIGDAT Conference on Empirical Methods in Natural Language Processing and Very Large Corpora: Held in Conjunction with the 38th Annual Meeting of the Association for Computational Linguistics , Volume 13 . 63--70. Kristina Toutanova and Christopher D. Manning. 2000. Enriching the knowledge sources used in a maximum entropy part-of-speech tagger. In Proceedings of the 2000 Joint SIGDAT Conference on Empirical Methods in Natural Language Processing and Very Large Corpora: Held in Conjunction with the 38th Annual Meeting of the Association for Computational Linguistics, Volume 13. 63--70."},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1009976227802"},{"key":"e_1_2_1_40_1","unstructured":"Peter D. Turney. 2003. Coherent keyphrase extraction via web mining. arXiv:cs\/0308033.  Peter D. Turney. 2003. Coherent keyphrase extraction via web mining. arXiv:cs\/0308033."},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.3115\/1599081.1599203"},{"key":"e_1_2_1_42_1","volume-title":"Proceedings of the 23rd National Conference on Artificial Intelligence (AAAI\u201908)","author":"Wan Xiaojun","year":"2008","unstructured":"Xiaojun Wan and Jianguo Xiao . 2008 b. Single document keyphrase extraction using neighborhood knowledge . In Proceedings of the 23rd National Conference on Artificial Intelligence (AAAI\u201908) , Vol.\u00a08. 855--860. Xiaojun Wan and Jianguo Xiao. 2008b. Single document keyphrase extraction using neighborhood knowledge. In Proceedings of the 23rd National Conference on Artificial Intelligence (AAAI\u201908), Vol.\u00a08. 855--860."},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-13560-1_11"},{"key":"e_1_2_1_44_1","volume-title":"Proceedings of the Software Engineering Research Conference. 39","author":"Wang Rui","year":"2014","unstructured":"Rui Wang , Wei Liu , and Chris McDonald . 2014 a. Corpus-independent generic keyphrase extraction using word embedding vectors . In Proceedings of the Software Engineering Research Conference. 39 . Rui Wang, Wei Liu, and Chris McDonald. 2014a. Corpus-independent generic keyphrase extraction using word embedding vectors. In Proceedings of the Software Engineering Research Conference. 39."},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/2396761.2398706"},{"volume-title":"Proceedings of the 4th ACM Conference on Digital Libraries. ACM","author":"Witten Ian H.","key":"e_1_2_1_46_1","unstructured":"Ian H. Witten , Gordon W. Paynter , Eibe Frank , Carl Gutwin , and Craig G . Nevill-Manning. 1999. KEA: Practical automatic keyphrase extraction . In Proceedings of the 4th ACM Conference on Digital Libraries. ACM , New York, NY, 254--255. Ian H. Witten, Gordon W. Paynter, Eibe Frank, Carl Gutwin, and Craig G. Nevill-Manning. 1999. KEA: Practical automatic keyphrase extraction. In Proceedings of the 4th ACM Conference on Digital Libraries. ACM, New York, NY, 254--255."},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-45817-5_49"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/2766462.2767744"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D16-1080"},{"key":"e_1_2_1_50_1","volume-title":"Proceedings of the 23rd International Joint Conference on Artificial Intelligence (IJCAI\u201913)","author":"Zhang Wei","year":"2013","unstructured":"Wei Zhang , Wei Feng , and Jianyong Wang . 2013 . Integrating semantic relatedness and words\u2019 intrinsic features for keyword extraction . In Proceedings of the 23rd International Joint Conference on Artificial Intelligence (IJCAI\u201913) . 2225--2231. Wei Zhang, Wei Feng, and Jianyong Wang. 2013. Integrating semantic relatedness and words\u2019 intrinsic features for keyword extraction. In Proceedings of the 23rd International Joint Conference on Artificial Intelligence (IJCAI\u201913). 2225--2231."},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N18-1151"},{"key":"e_1_2_1_52_1","volume-title":"Proceedings of the 30th AAAI Conference on Artificial Intelligence (AAAI\u201916)","author":"Zheng Hao","year":"2016","unstructured":"Hao Zheng , Zhoujun Li , Senzhang Wang , Zhao Yan , and Jianshe Zhou . 2016 . Aggregating inter-sentence information to enhance relation extraction . In Proceedings of the 30th AAAI Conference on Artificial Intelligence (AAAI\u201916) . 3108--3115. Hao Zheng, Zhoujun Li, Senzhang Wang, Zhao Yan, and Jianshe Zhou. 2016. Aggregating inter-sentence information to enhance relation extraction. In Proceedings of the 30th AAAI Conference on Artificial Intelligence (AAAI\u201916). 3108--3115."}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3388971","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3388971","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T22:33:02Z","timestamp":1750199582000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3388971"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,6,21]]},"references-count":51,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2020,9,30]]}},"alternative-id":["10.1145\/3388971"],"URL":"https:\/\/doi.org\/10.1145\/3388971","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"type":"print","value":"2375-4699"},{"type":"electronic","value":"2375-4702"}],"subject":[],"published":{"date-parts":[[2020,6,21]]},"assertion":[{"value":"2017-10-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-03-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-06-21","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}