{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,17]],"date-time":"2025-10-17T14:05:05Z","timestamp":1760709905151,"version":"3.37.3"},"reference-count":85,"publisher":"World Scientific Pub Co Pte Ltd","issue":"04","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Int. J. Patt. Recogn. Artif. Intell."],"published-print":{"date-parts":[[2020,4]]},"abstract":"<jats:p>Paraphrase identification is a natural language processing (NLP) problem that involves the determination of whether two text segments have the same meaning. Various NLP applications rely on a solution to this problem, including automatic plagiarism detection, text summarization, machine translation (MT), and question answering. The methods for identifying paraphrases found in the literature fall into two main classes: similarity-based methods and classification methods. This paper presents a critical study and an evaluation of existing methods for paraphrase identification and its application to automatic plagiarism detection. It presents the classes of paraphrase phenomena, the main methods, and the sets of features used by each particular method. All the methods and features used are discussed and enumerated in a table for easy comparison. Their performances on benchmark corpora are also discussed and compared via tables. Automatic plagiarism detection is presented as an application of paraphrase identification. The performances on benchmark corpora of existing plagiarism detection systems able to detect paraphrases are compared and discussed. The main outcome of this study is the identification of word overlap, structural representations, and MT measures as feature subsets that lead to the best performance results for support vector machines in both paraphrase identification and plagiarism detection on corpora. The performance results achieved by deep learning techniques highlight that these techniques are the most promising research direction in this field.<\/jats:p>","DOI":"10.1142\/s0218001420530043","type":"journal-article","created":{"date-parts":[[2019,5,16]],"date-time":"2019-05-16T02:33:31Z","timestamp":1557974011000},"page":"2053004","source":"Crossref","is-referenced-by-count":14,"title":["Evaluation of State-of-the-Art Paraphrase Identification and Its Application to Automatic Plagiarism Detection"],"prefix":"10.1142","volume":"34","author":[{"given":"Alaa","family":"Altheneyan","sequence":"first","affiliation":[{"name":"Department of Computer Science, College of Computer and Information Sciences, King Saud University, Riyadh, P. O. Box 89638, Saudi Arabia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5981-6299","authenticated-orcid":false,"given":"Mohamed El Bachir","family":"Menai","sequence":"additional","affiliation":[{"name":"Department of Computer Science, College of Computer and Information Sciences, King Saud University, Riyadh, P. O. Box 89638, Saudi Arabia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"219","published-online":{"date-parts":[[2019,8,22]]},"reference":[{"key":"S0218001420530043BIB002","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2017.01.002"},{"key":"S0218001420530043BIB003","doi-asserted-by":"publisher","DOI":"10.1016\/0020-0190(86)90091-8"},{"key":"S0218001420530043BIB005","doi-asserted-by":"publisher","DOI":"10.1109\/TSMCC.2011.2134847"},{"key":"S0218001420530043BIB006","doi-asserted-by":"publisher","DOI":"10.1613\/jair.2985"},{"key":"S0218001420530043BIB007","first-page":"167","volume-title":"Proc. 24th Int. Conf. Computational Linguistics (COLING 2012)","author":"B\u00e4r D.","year":"2012"},{"key":"S0218001420530043BIB008","doi-asserted-by":"publisher","DOI":"10.1162\/COLI_a_00153"},{"key":"S0218001420530043BIB009","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/S15-2004"},{"key":"S0218001420530043BIB010","doi-asserted-by":"publisher","DOI":"10.1162\/COLI_a_00166"},{"key":"S0218001420530043BIB011","first-page":"546","volume-title":"Proc. 2012 Joint Conf. Empirical Methods in Natural Language Processing and Computational Natural Language Learning","author":"Blacoe W.","year":"2012"},{"key":"S0218001420530043BIB012","first-page":"1","volume-title":"Proc. 3rd Int. Workshop on Paraphrasing","author":"Brockett C.","year":"2005"},{"key":"S0218001420530043BIB013","first-page":"449","volume-title":"Proc. 50th Annual Meeting of the Association for Computational Linguistics: Long Papers-Volume 1","author":"Bu F.","year":"2012"},{"key":"S0218001420530043BIB014","doi-asserted-by":"publisher","DOI":"10.1145\/2483669.2483676"},{"key":"S0218001420530043BIB015","doi-asserted-by":"publisher","DOI":"10.3115\/1613715.1613743"},{"key":"S0218001420530043BIB016","first-page":"164","volume-title":"Proc. Machine Translation Summit VIII","volume":"318","author":"Callison-Burch C.","year":"2001"},{"key":"S0218001420530043BIB017","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D15-1177"},{"key":"S0218001420530043BIB019","first-page":"152","volume-title":"Proc. 40th Annual Meeting on Association for Computational Linguistics","author":"Clough P.","year":"2002"},{"key":"S0218001420530043BIB020","doi-asserted-by":"publisher","DOI":"10.1162\/coli.08-003-R1-07-044"},{"key":"S0218001420530043BIB021","doi-asserted-by":"publisher","DOI":"10.3115\/1631862.1631865"},{"key":"S0218001420530043BIB022","doi-asserted-by":"publisher","DOI":"10.3115\/1687878.1687944"},{"key":"S0218001420530043BIB023","first-page":"449","volume-title":"Proc. LREC","volume":"6","author":"De Marneffe M.-C.","year":"2006"},{"key":"S0218001420530043BIB024","doi-asserted-by":"publisher","DOI":"10.3115\/1220355.1220406"},{"key":"S0218001420530043BIB026","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/S15-2011"},{"key":"S0218001420530043BIB027","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-71746-3_21"},{"key":"S0218001420530043BIB028","first-page":"45","volume-title":"Proc. 11th Annual Research Colloquium of the UK Special Interest Group for Computational Linguistics","author":"Fernando S.","year":"2008"},{"key":"S0218001420530043BIB029","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P15-1097"},{"key":"S0218001420530043BIB030","first-page":"17","volume-title":"Proc. Third Int. Workshop on Paraphrasing (IWP2005)","author":"Finch A.","year":"2005"},{"key":"S0218001420530043BIB032","first-page":"864","volume-title":"Proc. 50th Annual Meeting of the Association for Computational Linguistics: Long Papers \u2014 Volume 1, ACL \u201912","author":"Guo W.","year":"2012"},{"issue":"5","key":"S0218001420530043BIB033","doi-asserted-by":"crossref","DOI":"10.25103\/jestr.135.04","volume":"9","author":"Gupta D.","year":"2016","journal-title":"J. Eng. Sci. Technol. Rev."},{"key":"S0218001420530043BIB034","first-page":"44","volume-title":"* SEM@ NAACL-HLT","author":"Han L.","year":"2013"},{"key":"S0218001420530043BIB035","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D15-1181"},{"key":"S0218001420530043BIB036","first-page":"2042","volume-title":"Advances in Neural Information Processing Systems","author":"Hu B.","year":"2014"},{"key":"S0218001420530043BIB037","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46672-9_21"},{"key":"S0218001420530043BIB038","doi-asserted-by":"publisher","DOI":"10.1075\/cilt.309.18isl"},{"key":"S0218001420530043BIB039","first-page":"891","volume-title":"EMNLP","author":"Ji Y.","year":"2013"},{"key":"S0218001420530043BIB040","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/S15-2013"},{"key":"S0218001420530043BIB041","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/S15-2012"},{"key":"S0218001420530043BIB042","doi-asserted-by":"publisher","DOI":"10.1007\/11816508_52"},{"key":"S0218001420530043BIB043","doi-asserted-by":"publisher","DOI":"10.1214\/aoms\/1177729694"},{"key":"S0218001420530043BIB044","doi-asserted-by":"publisher","DOI":"10.1109\/ICAICTA.2016.7803127"},{"key":"S0218001420530043BIB045","doi-asserted-by":"publisher","DOI":"10.3115\/981574.981590"},{"issue":"1","key":"S0218001420530043BIB046","first-page":"19","volume":"34","author":"Lintean M. C.","year":"2010","journal-title":"Informatica"},{"key":"S0218001420530043BIB047","first-page":"263","volume-title":"FLAIRS Conf.","author":"Lintean M. C.","year":"2011"},{"key":"S0218001420530043BIB048","doi-asserted-by":"publisher","DOI":"10.1162\/coli_a_00002"},{"key":"S0218001420530043BIB049","first-page":"182","volume-title":"Proc. 2012 Conf. North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL HLT \u201912","author":"Madnani N.","year":"2012"},{"key":"S0218001420530043BIB050","doi-asserted-by":"publisher","DOI":"10.3115\/1667884.1667889"},{"key":"S0218001420530043BIB051","doi-asserted-by":"publisher","DOI":"10.1075\/nlp.3"},{"key":"S0218001420530043BIB052","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-59692-1_14"},{"key":"S0218001420530043BIB053","doi-asserted-by":"publisher","DOI":"10.1016\/S8755-4615(04)00039-8"},{"key":"S0218001420530043BIB054","first-page":"109","volume-title":"Proc. European Workshop on Natural Language Generation","author":"Marsi E.","year":"2005"},{"issue":"8","key":"S0218001420530043BIB055","first-page":"1050","volume":"12","author":"Maurer H. A.","year":"2006","journal-title":"J. Univers. Comput. Sci."},{"key":"S0218001420530043BIB056","doi-asserted-by":"crossref","first-page":"23","DOI":"10.1093\/oso\/9780198294252.003.0002","volume-title":"Phraseology. Theory, Analysis, and Applications","author":"Mel\u010duk I.","year":"1998"},{"issue":"10","key":"S0218001420530043BIB057","first-page":"80","volume":"4","author":"Menai M. E. B.","year":"2012","journal-title":"Int. J. Inf. Technol. Comput. Sci."},{"key":"S0218001420530043BIB058","first-page":"775","volume-title":"Proc. 21st National Conf. Artificial Intelligence \u2014 Volume 1, AAAI\u201906","author":"Mihalcea R.","year":"2006"},{"first-page":"708","volume-title":"Proc. 2014 Conf. Empirical Methods in Natural Language Processing (EMNLP)","author":"Milajevs D.","key":"S0218001420530043BIB059"},{"key":"S0218001420530043BIB060","doi-asserted-by":"publisher","DOI":"10.3726\/978-3-0352-0096-6"},{"key":"S0218001420530043BIB061","doi-asserted-by":"publisher","DOI":"10.1145\/219717.219748"},{"key":"S0218001420530043BIB062","doi-asserted-by":"publisher","DOI":"10.1016\/j.asoc.2011.12.021"},{"key":"S0218001420530043BIB063","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-35326-0_54"},{"issue":"2","key":"S0218001420530043BIB064","first-page":"135","volume":"32","author":"Osman N. S. A. H.","year":"2011","journal-title":"J. Theor. Appl. Inf. Technol."},{"key":"S0218001420530043BIB065","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1162"},{"key":"S0218001420530043BIB066","doi-asserted-by":"publisher","DOI":"10.1002\/asi.23593"},{"volume-title":"Proc. SEPLN 10 Workshop on Uncovering Plagiarism, Authorship and Social Software Misuse","year":"2010","author":"Potthast M.","key":"S0218001420530043BIB067"},{"key":"S0218001420530043BIB069","first-page":"997","volume-title":"23rd Int. Conf. Computational Linguistics (COLING 10)","author":"Potthast M.","year":"2010"},{"key":"S0218001420530043BIB070","doi-asserted-by":"publisher","DOI":"10.3115\/1610075.1610079"},{"key":"S0218001420530043BIB071","doi-asserted-by":"publisher","DOI":"10.1007\/s40979-016-0013-y"},{"key":"S0218001420530043BIB072","first-page":"2422","volume-title":"LREC","author":"Rus V.","year":"2014"},{"key":"S0218001420530043BIB073","first-page":"23","volume-title":"The 8th Language Resources and Evaluation Conf. (LREC 2012) Semantic Relations II: Enhancing Resources and Applications","author":"Rus V.","year":"2012"},{"key":"S0218001420530043BIB074","first-page":"201","volume-title":"FLAIRS Conf.","author":"Rus V.","year":"2008"},{"key":"S0218001420530043BIB075","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/S15-2009"},{"key":"S0218001420530043BIB076","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-50496-4_4"},{"key":"S0218001420530043BIB077","first-page":"801","volume-title":"Advances in Neural Information Processing Systems","author":"Socher R.","year":"2011"},{"issue":"22","key":"S0218001420530043BIB078","first-page":"4894","volume":"4","author":"Ul-Qayyum Z.","year":"2012","journal-title":"Res. J. Appl. Sci. Eng. Technol."},{"key":"S0218001420530043BIB079","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/S15-2007"},{"issue":"0","key":"S0218001420530043BIB080","first-page":"83","volume":"46","author":"Vila M.","year":"2010","journal-title":"Procesamiento Lenguaje Nat."},{"key":"S0218001420530043BIB081","doi-asserted-by":"publisher","DOI":"10.4236\/ojml.2014.41016"},{"key":"S0218001420530043BIB082","first-page":"131","volume-title":"Proc. Australasian Language Technology Workshop","volume":"2006","author":"Wan S.","year":"2006"},{"key":"S0218001420530043BIB083","first-page":"1340","volume-title":"26th Int. Conf. Computational Linguistics: Technical Papers Proc. COLING 2016","author":"Wang Z.","year":"2016"},{"key":"S0218001420530043BIB084","first-page":"901","volume-title":"Proc. 2015 Int. Conf. North America Chapter of the Association for Computational Linguistics: Human Language Technologyies","author":"Wenpeng Y.","year":"2015"},{"key":"S0218001420530043BIB085","first-page":"347","volume-title":"Proc. 19th ASCILITE Conf.","author":"Williams J. B.","year":"2002"},{"key":"S0218001420530043BIB087","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/S15-2001"},{"key":"S0218001420530043BIB088","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00194"},{"key":"S0218001420530043BIB089","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/N15-1154"},{"key":"S0218001420530043BIB090","first-page":"160","volume-title":"Proc. Australasian Language Technology Workshop","author":"Zhang Y.","year":"2005"},{"key":"S0218001420530043BIB091","doi-asserted-by":"publisher","DOI":"10.3115\/1690219.1690263"},{"key":"S0218001420530043BIB092","doi-asserted-by":"publisher","DOI":"10.1007\/11735106_66"}],"container-title":["International Journal of Pattern Recognition and Artificial Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.worldscientific.com\/doi\/pdf\/10.1142\/S0218001420530043","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,7,18]],"date-time":"2024-07-18T02:49:51Z","timestamp":1721270991000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.worldscientific.com\/doi\/abs\/10.1142\/S0218001420530043"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,8,22]]},"references-count":85,"journal-issue":{"issue":"04","published-print":{"date-parts":[[2020,4]]}},"alternative-id":["10.1142\/S0218001420530043"],"URL":"https:\/\/doi.org\/10.1142\/s0218001420530043","relation":{},"ISSN":["0218-0014","1793-6381"],"issn-type":[{"type":"print","value":"0218-0014"},{"type":"electronic","value":"1793-6381"}],"subject":[],"published":{"date-parts":[[2019,8,22]]}}}