{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T15:44:12Z","timestamp":1782834252358,"version":"3.54.5"},"reference-count":109,"publisher":"Privacy Enhancing Technologies Symposium Advisory Board","issue":"4","license":[{"start":{"date-parts":[[2021,7,23]],"date-time":"2021-07-23T00:00:00Z","timestamp":1626998400000},"content-version":"unspecified","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by-nc-nd\/3.0"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2021,10,1]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Privacy policies have become a focal point of privacy research. With their goal to reflect the privacy practices of a website, service, or app, they are often the starting point for researchers who analyze the accuracy of claimed data practices, user understanding of practices, or control mechanisms for users. Due to vast differences in structure, presentation, and content, it is often challenging to extract privacy policies from online resources like websites for analysis. In the past, researchers have relied on scrapers tailored to the specific analysis or task, which complicates comparing results across different studies.<\/jats:p><jats:p>To unify future research in this field, we developed a toolchain to process website privacy policies and prepare them for research purposes. The core part of this chain is a detector module for English and German, using natural language processing and machine learning to automatically determine whether given texts are privacy or cookie policies. We leverage multiple existing data sets to refine our approach, evaluate it on a recently published longitudinal corpus, and show that it contains a number of misclassified documents. We believe that unifying data preparation for the analysis of privacy policies can help make different studies more comparable and is a step towards more thorough analyses. In addition, we provide insights into common pitfalls that may lead to invalid analyses.<\/jats:p>","DOI":"10.2478\/popets-2021-0081","type":"journal-article","created":{"date-parts":[[2021,7,24]],"date-time":"2021-07-24T23:16:57Z","timestamp":1627168617000},"page":"480-499","source":"Crossref","is-referenced-by-count":11,"title":["Unifying Privacy Policy Detection"],"prefix":"10.56553","volume":"2021","author":[{"given":"Henry","family":"Hosseini","sequence":"first","affiliation":[{"name":"University of M\u00fcnster & Ruhr University Bochum"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Martin","family":"Degeling","sequence":"additional","affiliation":[{"name":"Ruhr University Bochum"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Christine","family":"Utz","sequence":"additional","affiliation":[{"name":"Ruhr University Bochum"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Thomas","family":"Hupperich","sequence":"additional","affiliation":[{"name":"University of M\u00fcnster"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"35752","published-online":{"date-parts":[[2021,7,23]]},"reference":[{"key":"2022060521170780611_j_popets-2021-0081_ref_001","doi-asserted-by":"crossref","unstructured":"[1] Kenneth D. Pimple. Emerging Pervasive Information and Communication Technologies (PICT). Springer, 2014.10.1007\/978-94-007-6833-8","DOI":"10.1007\/978-94-007-6833-8"},{"key":"2022060521170780611_j_popets-2021-0081_ref_002","unstructured":"[2] Willis H. Ware. Records, Computers and the Rights of Citizens. Technical report, The Rand Corporation, Santa Monica, California, 1973."},{"key":"2022060521170780611_j_popets-2021-0081_ref_003","unstructured":"[3] Christine Utz, Martin Degeling, Sascha Fahl, Florian Schaub, and Thorsten Holz. (Un)informed Consent: Studying GDPR Consent Notices in the Field. In Proceedings of the 2019 ACM SIGSAC Conference on Computer and Communications Security, pages 973\u2013990, 2019."},{"key":"2022060521170780611_j_popets-2021-0081_ref_004","doi-asserted-by":"crossref","unstructured":"[4] Julie M. Robillard, Tanya L. Feng, Arlo B. Sporn, Jen-Ai Lai, Cody Lo, Monica Ta, and Roland Nadler. Availability, readability, and content of privacy policies and terms of agreements of mental health apps. Internet Interventions, 17:100243, 2019.10.1016\/j.invent.2019.100243643003830949436","DOI":"10.1016\/j.invent.2019.100243"},{"key":"2022060521170780611_j_popets-2021-0081_ref_005","doi-asserted-by":"crossref","unstructured":"[5] Noriko Tomuro, Steven Lytinen, and Kurt Hornsburg. Automatic Summarization of Privacy Policies using Ensemble Learning. In Proceedings of the Sixth ACM Conference on Data and Application Security and Privacy, pages 133\u2013135, 2016.10.1145\/2857705.2857741","DOI":"10.1145\/2857705.2857741"},{"key":"2022060521170780611_j_popets-2021-0081_ref_006","doi-asserted-by":"crossref","unstructured":"[6] Razieh Nokhbeh Zaeem, Rachel L. German, and K. Suzanne Barber. PrivacyCheck: Automatic Summarization of Privacy Policies Using Data Mining. ACM Transactions on Internet Technology (TOIT), 18(4):1\u201318, 2018.","DOI":"10.1145\/3127519"},{"key":"2022060521170780611_j_popets-2021-0081_ref_007","doi-asserted-by":"crossref","unstructured":"[7] Dhiren A. Audich, Rozita Dara, and Blair Nonnecke. Extracting keyword and keyphrase from online privacy policies. In 2016 Eleventh International Conference on Digital Information Management (ICDIM), pages 127\u2013132. IEEE, 2016.10.1109\/ICDIM.2016.7829792","DOI":"10.1109\/ICDIM.2016.7829792"},{"key":"2022060521170780611_j_popets-2021-0081_ref_008","doi-asserted-by":"crossref","unstructured":"[8] Benjamin Fabian, Tatiana Ermakova, and Tino Lentz. Large-Scale Readability Analysis of Privacy Policies. In Proceedings of the International Conference on Web Intelligence, pages 18\u201325, 2017.10.1145\/3106426.3106427","DOI":"10.1145\/3106426.3106427"},{"key":"2022060521170780611_j_popets-2021-0081_ref_009","doi-asserted-by":"crossref","unstructured":"[9] Shomir Wilson, Florian Schaub, Aswarth Abhilash Dara, Frederick Liu, Sushain Cherivirala, Pedro Giovanni Leon, Mads Schaarup Andersen, Sebastian Zimmeck, Kanthashree Mysore Sathyendra, N. Cameron Russell, et al. The Creation and Analysis of a Website Privacy Policy Corpus. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1330\u20131340, 2016.10.18653\/v1\/P16-1126","DOI":"10.18653\/v1\/P16-1126"},{"key":"2022060521170780611_j_popets-2021-0081_ref_010","doi-asserted-by":"crossref","unstructured":"[10] Dhiren A. Audich, Rozita Dara, and Blair Nonnecke. Privacy Policy Annotation for Semi-automated Analysis: A Cost-Effective Approach. In IFIP International Conference on Trust Management, pages 29\u201344. Springer, 2018.10.1007\/978-3-319-95276-5_3","DOI":"10.1007\/978-3-319-95276-5_3"},{"key":"2022060521170780611_j_popets-2021-0081_ref_011","unstructured":"[11] Hamza Harkous, Kassem Fawaz, R\u00e9mi Lebret, Florian Schaub, Kang G. Shin, and Karl Aberer. Polisis: Automated Analysis and Presentation of Privacy Policies Using Deep Learning. In Proceedings of the 27th USENIX Security Symposium, pages 531\u2013548, 2018."},{"key":"2022060521170780611_j_popets-2021-0081_ref_012","doi-asserted-by":"crossref","unstructured":"[12] Tobias Urban, Martin Degeling, Thorsten Holz, and Nor-bert Pohlmann. \u201cYour Hashed IP Address: Ubuntu.\u201d Perspectives on Transparency Tools for Online Advertising. In Proceedings of the 35th Annual Computer Security Applications Conference, pages 702\u2013717, 2019.","DOI":"10.1145\/3359789.3359798"},{"key":"2022060521170780611_j_popets-2021-0081_ref_013","doi-asserted-by":"crossref","unstructured":"[13] Luca Bufalieri, Massimo La Morgia, Alessandro Mei, and Julinda Stefa. GDPR: When the Right to Access Personal Data Becomes a Threat. arXiv preprint arXiv:2005.01868, 2020.","DOI":"10.1109\/ICWS49710.2020.00017"},{"key":"2022060521170780611_j_popets-2021-0081_ref_014","doi-asserted-by":"crossref","unstructured":"[14] Coline Boniface, Imane Fouad, Nataliia Bielova, C\u00e9dric Lauradoux, and Cristiana Santos. Security Analysis of Subject Access Request Procedures. In Annual Privacy Forum, pages 182\u2013209. Springer, 2019.10.1007\/978-3-030-21752-5_12","DOI":"10.1007\/978-3-030-21752-5_12"},{"key":"2022060521170780611_j_popets-2021-0081_ref_015","doi-asserted-by":"crossref","unstructured":"[15] Martin Degeling, Christine Utz, Christopher Lentzsch, Henry Hosseini, Florian Schaub, and Thorsten Holz. We Value Your Privacy ... Now Take Some Cookies: Measuring the GDPR\u2019s Impact on Web Privacy. In Proceedings of the 26th Annual Network and Distributed System Security Symposium (NDSS 2019). The Internet Society, February 2019.10.14722\/ndss.2019.23378","DOI":"10.14722\/ndss.2019.23378"},{"key":"2022060521170780611_j_popets-2021-0081_ref_016","doi-asserted-by":"crossref","unstructured":"[16] Mukund Srinath, Shomir Wilson, and C. Lee Giles. Privacy at Scale: Introducing the PrivaSeer Corpus of Web Privacy Policies. arXiv preprint arXiv:2004.11131, 2020.","DOI":"10.18653\/v1\/2021.acl-long.532"},{"key":"2022060521170780611_j_popets-2021-0081_ref_017","unstructured":"[17] Fei Liu, Rohan Ramanath, Norman Sadeh, and Noah A. Smith. A Step Towards Usable Privacy Policy: Automatic Alignment of Privacy Statements. In Proceedings of COLING 2014, the 25th International Conference on Computational Linguistics: Technical Papers, pages 884\u2013894, 2014."},{"key":"2022060521170780611_j_popets-2021-0081_ref_018","doi-asserted-by":"crossref","unstructured":"[18] Le Yu, Xiapu Luo, Xule Liu, and Tao Zhang. Can We Trust the Privacy Policies of Android Apps? In 2016 46th Annual IEEE\/IFIP International Conference on Dependable Systems and Networks (DSN), pages 538\u2013549. IEEE, 2016.10.1109\/DSN.2016.55","DOI":"10.1109\/DSN.2016.55"},{"key":"2022060521170780611_j_popets-2021-0081_ref_019","doi-asserted-by":"crossref","unstructured":"[19] Abhijith Athreya Mysore Gopinath, Shomir Wilson, and Norman Sadeh. Supervised and Unsupervised Methods for Robust Separation of Section Titles and Prose Text in Web Documents. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 850\u2013855. Association for Computational Linguistics, 2018.10.18653\/v1\/D18-1099","DOI":"10.18653\/v1\/D18-1099"},{"key":"2022060521170780611_j_popets-2021-0081_ref_020","doi-asserted-by":"crossref","unstructured":"[20] Timothy Libert. An Automated Approach to Auditing Disclosure of Third-Party Data Collection in Website Privacy Policies. In Proceedings of the 2018 World Wide Web Conference, pages 207\u2013216, 2018.10.1145\/3178876.3186087","DOI":"10.1145\/3178876.3186087"},{"key":"2022060521170780611_j_popets-2021-0081_ref_021","doi-asserted-by":"crossref","unstructured":"[21] Keishiro Fukushima, Toru Nakamura, Daisuke Ikeda, and Shinsaku Kiyomoto. Challenges in Classifying Privacy Policies by Machine Learning with Word-based Features. In Proceedings of the 2nd International Conference on Cryptography, Security and Privacy (ICCSP 2018), pages 62\u201366, Guiyang, China, 2018. ACM.10.1145\/3199478.3199486","DOI":"10.1145\/3199478.3199486"},{"key":"2022060521170780611_j_popets-2021-0081_ref_022","doi-asserted-by":"crossref","unstructured":"[22] Tarun Ramadorai, Antoine Uettwiller, and Ansgar Walther. The Market for Data Privacy. https:\/\/dx.doi.org\/10.2139\/ssrn.3352175, 2019.10.2139\/ssrn.3352175","DOI":"10.2139\/ssrn.3352175"},{"key":"2022060521170780611_j_popets-2021-0081_ref_023","doi-asserted-by":"crossref","unstructured":"[23] Martin Boldt and Kaavya Rekanar. Analysis and Text Classification of Privacy Policies From Rogue and Top-100 Fortune Global Companies. International Journal of Information Security and Privacy (IJISP), 13(2):47\u201366, 2019.10.4018\/IJISP.2019040104","DOI":"10.4018\/IJISP.2019040104"},{"key":"2022060521170780611_j_popets-2021-0081_ref_024","doi-asserted-by":"crossref","unstructured":"[24] Sebastian Zimmeck, Peter Story, Daniel Smullen, Abhilasha Ravichander, Ziqi Wang, Joel Reidenberg, N. Cameron Russell, and Norman Sadeh. MAPS: Scaling Privacy Compliance Analysis to a Million Apps. Proceedings on Privacy Enhancing Technologies, 2019(3):66\u201386, 2019.","DOI":"10.2478\/popets-2019-0037"},{"key":"2022060521170780611_j_popets-2021-0081_ref_025","doi-asserted-by":"crossref","unstructured":"[25] David Sarne, Jonathan Schler, Alon Singer, Ayelet Sela, and Ittai Bar Siman Tov. Unsupervised Topic Extraction from Privacy Policies. In Companion Proceedings of The 2019 World Wide Web Conference, pages 563\u2013568. IW3C2 (International World Wide Web Conference Committee), 2019.10.1145\/3308560.3317585","DOI":"10.1145\/3308560.3317585"},{"key":"2022060521170780611_j_popets-2021-0081_ref_026","doi-asserted-by":"crossref","unstructured":"[26] Mitra Bokaie Hosseini, KC Pragyan, Irwin Reyes, and Serge Egelman. Identifying and Classifying Third-party Entities in Natural Language Privacy Policies. In Proceedings of the Second Workshop on Privacy in NLP, pages 18\u201327, 2020.10.18653\/v1\/2020.privatenlp-1.3","DOI":"10.18653\/v1\/2020.privatenlp-1.3"},{"key":"2022060521170780611_j_popets-2021-0081_ref_027","doi-asserted-by":"crossref","unstructured":"[27] Vinayshekhar Bannihatti Kumar, Roger Iyengar, Namita Nisal, Yuanyuan Feng, Hana Habib, Peter Story, Sushain Cherivirala, Margaret Hagan, Lorrie Cranor, Shomir Wilson, et al. Finding a Choice in a Haystack: Automatic Extraction of Opt-Out Statements from Privacy Policy Text. In Proceedings of The Web Conference 2020, 2020.10.1145\/3366423.3380262","DOI":"10.1145\/3366423.3380262"},{"key":"2022060521170780611_j_popets-2021-0081_ref_028","doi-asserted-by":"crossref","unstructured":"[28] Thomas Linden, Rishabh Khandelwal, Hamza Harkous, and Kassem Fawaz. The Privacy Policy Landscape After the GDPR. Proceedings on Privacy Enhancing Technologies, 2020(1):47\u201364, 2020.10.2478\/popets-2020-0004","DOI":"10.2478\/popets-2020-0004"},{"key":"2022060521170780611_j_popets-2021-0081_ref_029","doi-asserted-by":"crossref","unstructured":"[29] Yoon Kim. Convolutional Neural Networks for Sentence Classification. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1746\u20131751. Association for Computational Linguistics, 2014.10.3115\/v1\/D14-1181","DOI":"10.3115\/v1\/D14-1181"},{"key":"2022060521170780611_j_popets-2021-0081_ref_030","doi-asserted-by":"crossref","unstructured":"[30] Ryan Amos, Gunes Acar, Elena Lucherini, Mihir Kshirsagar, Arvind Narayanan, and Jonathan Mayer. Privacy Policies over Time: Curation and Analysis of a Million-Document Dataset. arXiv preprint, arXiv:2008.09159, 2020.","DOI":"10.1145\/3442381.3450048"},{"key":"2022060521170780611_j_popets-2021-0081_ref_031","unstructured":"[31] Leonard Richardson. Beautiful Soup. https:\/\/www.crummy.com\/software\/BeautifulSoup\/bs4\/doc\/, 2007. [Online; accessed 24 April 2020]."},{"key":"2022060521170780611_j_popets-2021-0081_ref_032","unstructured":"[32] Postlight. Mercury Parser \u2013 Extracting content from chaos. https:\/\/github.com\/postlight\/mercury-parser."},{"key":"2022060521170780611_j_popets-2021-0081_ref_033","unstructured":"[33] Stefan Behnel, Martijn Faassen, and Ian Bicking. lxml: Processing XML and HTML with Python. https:\/\/lxml.de\/, 2005. [Online; accessed 14 June 2021]."},{"key":"2022060521170780611_j_popets-2021-0081_ref_034","unstructured":"[34] Kanthashree Mysore Sathyendra, Abhilasha Ravichander, Peter Garth Story, Alan W. Black, and Norman Sadeh. Helping Users Understand Privacy Notices with Automated Query Answering Functionality: An Exploratory Study. Technical report, 2017."},{"key":"2022060521170780611_j_popets-2021-0081_ref_035","unstructured":"[35] Marco Lui and Timothy Baldwin. langid.py: An Off-the-shelf Language Identification Tool. In Proceedings of the ACL 2012 System Demonstrations, pages 25\u201330. Association for Computational Linguistics, 2012."},{"key":"2022060521170780611_j_popets-2021-0081_ref_036","doi-asserted-by":"crossref","unstructured":"[36] Christian Kohlsch\u00fctter, Peter Fankhauser, and Wolfgang Nejdl. Boilerplate Detection using Shallow Text Features. In Proceedings of the third ACM international conference on Web search and data mining, pages 441\u2013450, 2010.10.1145\/1718487.1718542","DOI":"10.1145\/1718487.1718542"},{"key":"2022060521170780611_j_popets-2021-0081_ref_037","doi-asserted-by":"crossref","unstructured":"[37] Rohan Ramanath, Fei Liu, Norman Sadeh, and Noah A. Smith. Unsupervised Alignment of Privacy Policies using Hidden Markov Models. In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 605\u2013610, 2014.10.3115\/v1\/P14-2099","DOI":"10.3115\/v1\/P14-2099"},{"key":"2022060521170780611_j_popets-2021-0081_ref_038","unstructured":"[38] Jim Plush and Robbie Coleman. Goose - Article Extractor. https:\/\/github.com\/goose3\/goose3, 2011. [Online; accessed 24 April 2020]."},{"key":"2022060521170780611_j_popets-2021-0081_ref_039","unstructured":"[39] Nakatani Shuyo. Language Detection Library for Java. http:\/\/code.google.com\/p\/language-detection\/, 2010."},{"key":"2022060521170780611_j_popets-2021-0081_ref_040","unstructured":"[40] Welderufael B. Tesfay, Peter Hofmann, Toru Nakamura, Shinsaku Kiyomoto, and Jetzabel Serna. PrivacyGuide: Towards an Implementation of the EU GDPR on Internet Privacy Policy Evaluation. In Proceedings of the Fourth ACM International Workshop on Security and Privacy Analytics, pages 15\u201321, 2018."},{"key":"2022060521170780611_j_popets-2021-0081_ref_041","doi-asserted-by":"crossref","unstructured":"[41] Matthew E. Peters and Dan Lecocq. Content Extraction Using Diverse Feature Sets. In Companion Publication of the 22nd International World Wide Web Conference, pages 89\u201390, 2013.10.1145\/2487788.2487828","DOI":"10.1145\/2487788.2487828"},{"key":"2022060521170780611_j_popets-2021-0081_ref_042","doi-asserted-by":"crossref","unstructured":"[42] Stuart Rose, Dave Engel, Nick Cramer, and Wendy Cowley. Automatic Keyword Extraction from Individual Documents. Text Mining: Applications and Theory, 1:1\u201320, 2010.10.1002\/9780470689646.ch1","DOI":"10.1002\/9780470689646.ch1"},{"key":"2022060521170780611_j_popets-2021-0081_ref_043","doi-asserted-by":"crossref","unstructured":"[43] Rada Mihalcea and Paul Tarau. TextRank: Bringing Order into Text. In Proceedings of the 2004 Conference on Empirical Methods in Natural Language Processing, pages 404\u2013411, 2004.","DOI":"10.3115\/1220575.1220627"},{"key":"2022060521170780611_j_popets-2021-0081_ref_044","unstructured":"[44] Jonathan Hedley. jsoup: Java HTML Parser. https:\/\/jsoup.org, 2009."},{"key":"2022060521170780611_j_popets-2021-0081_ref_045","doi-asserted-by":"crossref","unstructured":"[45] Elisa Costante, Yuanhao Sun, Milan Petkovi\u00a2, and Jerry den Hartog. A Machine Learning Solution to Assess Privacy Policy Completeness. In Proceedings of the 2012 ACM Workshop on Privacy in the Electronic Society, pages 91\u201396. ACM, 2012.10.1145\/2381966.2381979","DOI":"10.1145\/2381966.2381979"},{"key":"2022060521170780611_j_popets-2021-0081_ref_046","doi-asserted-by":"crossref","unstructured":"[46] Niharika Guntamukkala, Rozita Dara, and Gary Grewal. A Machine-Learning Based Approach for Measuring the Completeness of Online Privacy Policies. In 2015 IEEE 14th International Conference on Machine Learning and Applications (ICMLA), pages 289\u2013294. IEEE, 2015.10.1109\/ICMLA.2015.143","DOI":"10.1109\/ICMLA.2015.143"},{"key":"2022060521170780611_j_popets-2021-0081_ref_047","unstructured":"[47] Shuang Liu, Renjie Guo, Baiyang Zhao, Tao Chen, and Meishan Zhang. APPCorp: A Corpus for Android Privacy Policy Document Structure Analysis. arXiv preprint arXiv:2005.06945, 2020."},{"key":"2022060521170780611_j_popets-2021-0081_ref_048","doi-asserted-by":"crossref","unstructured":"[48] Cheng Chang, Huaxin Li, Yichi Zhang, Suguo Du, Hui Cao, and Haojin Zhu. Automated and Personalized Privacy Policy Extraction Under GDPR Consideration. In International Conference on Wireless Algorithms, Systems, and Applications, pages 43\u201354. Springer, 2019.10.1007\/978-3-030-23597-0_4","DOI":"10.1007\/978-3-030-23597-0_4"},{"key":"2022060521170780611_j_popets-2021-0081_ref_049","unstructured":"[49] Parvaneh Shayegh, Vijayanta Jain, Amin Rabinia, and Sepideh Ghanavati. Automated Approach to Improve IoT Privacy Policies. arXiv preprint arXiv:1910.04133, 2019."},{"key":"2022060521170780611_j_popets-2021-0081_ref_050","unstructured":"[50] Statista. Percentage of mobile device website traffic worldwide from 1st quarter 2015 to 1st quarter 2021. https:\/\/www.statista.com\/statistics\/277125\/share-of-website-traffic-coming-from-mobile-devices\/. [Online; accessed 14 June 2021]."},{"key":"2022060521170780611_j_popets-2021-0081_ref_051","doi-asserted-by":"crossref","unstructured":"[51] Pradeep K. Murukannaiah, Chinmaya Dabral, Karthik Sheshadri, Esha Sharma, and Jessica Staddon. Learning a Privacy Incidents Database. In Proceedings of the Hot Topics in Science of Security: Symposium and Bootcamp, pages 35\u201344, 2017.10.1145\/3055305.3055309","DOI":"10.1145\/3055305.3055309"},{"key":"2022060521170780611_j_popets-2021-0081_ref_052","unstructured":"[52] Aaron Swartz and Alireza Savand. HTML2Text. https:\/\/alir3z4.github.io\/html2text\/, 2011. [Online; accessed 20 April 2020]."},{"key":"2022060521170780611_j_popets-2021-0081_ref_053","unstructured":"[53] Albert Weichselbraun and Fabian Odoni. inscriptis \u2013 HTML to text conversion library, command line client and Web service. https:\/\/inscriptis.readthedocs.io\/en\/latest\/, 2016. [Online; accessed 20 April 2020]."},{"key":"2022060521170780611_j_popets-2021-0081_ref_054","unstructured":"[54] Mozilla. Readability.js. https:\/\/github.com\/mozilla\/readability, 2015. [Online; accessed 24 April 2020]."},{"key":"2022060521170780611_j_popets-2021-0081_ref_055","unstructured":"[55] Jorj X. McKie and Ruikai Liu. PyMuPDF. https:\/\/github.com\/pymupdf\/PyMuPDF, 2016. [Online; accessed 7 January 2021]."},{"key":"2022060521170780611_j_popets-2021-0081_ref_056","unstructured":"[56] The Apache Software Foundation. Apache Tika \u2013 a content analysis toolkit. https:\/\/tika.apache.org\/, 2019. Online; accessed 15 June 2021."},{"key":"2022060521170780611_j_popets-2021-0081_ref_057","unstructured":"[57] Dick Sites. Compact Language Detector 2. https:\/\/github.com\/CLD2Owners\/cld2, 2013. Online; accessed 15 June 2021."},{"key":"2022060521170780611_j_popets-2021-0081_ref_058","unstructured":"[58] Alex Salcianu, Andy Golding, Anton Bakalov, Chris Alberti, Daniel Andor, David Weiss, Emily Pitler, Greg Coppola, Jason Riesa, Kuzman Ganchev, et al. Compact Language Detector v3. https:\/\/github.com\/google\/cld3, 2018."},{"key":"2022060521170780611_j_popets-2021-0081_ref_059","unstructured":"[59] Kent Johnson and Phi-Long Do. Goose \u2013 Article Extractor. https:\/\/bitbucket.org\/spirit\/guess_language\/, 2008. [Online; accessed 24 April 2020]."},{"key":"2022060521170780611_j_popets-2021-0081_ref_060","unstructured":"[60] Burton DeWilde. textacy: NLP, before and after spaCy. https:\/\/github.com\/chartbeat-labs\/textacy, 2016. [Online; accessed 24 April 2020]."},{"key":"2022060521170780611_j_popets-2021-0081_ref_061","unstructured":"[61] Armand Joulin, Edouard Grave, Piotr Bojanowski, Matthijs Douze, H\u00e9rve J\u00e9gou, and Tomas Mikolov. FastText.zip: Compressing text classification models. arXiv preprint arXiv:1612.03651, 2016."},{"key":"2022060521170780611_j_popets-2021-0081_ref_062","doi-asserted-by":"crossref","unstructured":"[62] Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. Bag of Tricks for Efficient Text Classification. arXiv preprint arXiv:1607.01759, 2016.","DOI":"10.18653\/v1\/E17-2068"},{"key":"2022060521170780611_j_popets-2021-0081_ref_063","unstructured":"[63] Trang Ho and Allan Simon. Tatoeba: Collection of sentences and translations. https:\/\/tatoeba.org, 2016. [Online; accessed 15 June 2020]."},{"key":"2022060521170780611_j_popets-2021-0081_ref_064","unstructured":"[64] J\u00f6rg Tiedemann. Parallel Data, Tools and Interfaces in OPUS. In Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC \u201912), Istanbul, Turkey, May 2012. European Language Resources Association (ELRA)."},{"key":"2022060521170780611_j_popets-2021-0081_ref_065","unstructured":"[65] Dirk Goldhahn, Thomas Eckart, and Uwe Quasthoff. Building Large Monolingual Dictionaries at the Leipzig Corpora Collection: From 100 to 200 Languages. In Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC \u201912)."},{"key":"2022060521170780611_j_popets-2021-0081_ref_066","unstructured":"[66] Liling Tan, Marcos Zampieri, Nikola Ljube\u0161ic, and J\u00f6rg Tiedemann. Merging Comparable Data Sources for the Discrimination of Similar Languages: The DSL Corpus Collection. In Proceedings of the 7th Workshop on Building and Using Comparable Corpora (BUCC), pages 11\u201315, Reykjavik, Iceland, 2014."},{"key":"2022060521170780611_j_popets-2021-0081_ref_067","unstructured":"[67] Mitja Trampus. Evaluating language identification performance. https:\/\/blog.twitter.com\/engineering\/en_us\/a\/2015\/evaluating-language-identification-performance, 2015. [Online; accessed 15 April 2021]."},{"key":"2022060521170780611_j_popets-2021-0081_ref_068","unstructured":"[68] Ralf Steinberger, Bruno Pouliquen, Anna Widiger, Camelia Ignat, Tomaz Erjavec, Dan Tufis, and D\u00e1niel Varga. The JRC-Acquis: A Multilingual Aligned Parallel Corpus with 20+ Languages. arXiv preprint cs\/0609058, 2006."},{"key":"2022060521170780611_j_popets-2021-0081_ref_069","unstructured":"[69] Jamie Callan, Mark Hoy, Changkuk Yoo, and Le Zhao. The ClueWeb09 Dataset. http:\/\/boston.lti.cs.cmu.edu\/Data\/clueweb09, 2009. [Online; accessed 14 June 2021]."},{"key":"2022060521170780611_j_popets-2021-0081_ref_070","unstructured":"[70] David D. Lewis, Yiming Yang, Tony G. Rose, and Fan Li. RCV1: A New Benchmark Collection for Text Categorization Research. The Journal of Machine Learning Research, 5:361\u2013397, 2004."},{"key":"2022060521170780611_j_popets-2021-0081_ref_071","unstructured":"[71] Tomohiro Kubota. Introduction to i18n. https:\/\/www.debian.org\/doc\/manuals\/intro-i18n\/, 2003. Online; accessed 24 April 2021."},{"key":"2022060521170780611_j_popets-2021-0081_ref_072","unstructured":"[72] Quoc Le and Tomas Mikolov. Distributed representations of sentences and documents. In Proceedings of the 31st International Conference on Machine Learning (ICML\u201914), pages II\u20131188\u2013II\u20131196, 2014."},{"key":"2022060521170780611_j_popets-2021-0081_ref_073","unstructured":"[73] Peng Qi, Yuhao Zhang, Yuhui Zhang, Jason Bolton, and Christopher D. Manning. Stanza: A Python Natural Language Processing Toolkit for Many Human Languages. arXiv preprint arXiv:2003.07082, 2020."},{"key":"2022060521170780611_j_popets-2021-0081_ref_074","unstructured":"[74] Katrin Ortmann, Adam Roussel, and Stefanie Dipper. Evaluating Off-the-Shelf NLP Tools for German. In Proceedings of the 15th Conference on Natural Language Processing (KONVENS 2019), pages 212\u2013222, 2019."},{"key":"2022060521170780611_j_popets-2021-0081_ref_075","unstructured":"[75] Matthew Honnibal and Ines Montani. spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing. https:\/\/sentometrics-research.com\/publication\/72\/. [To appear]."},{"key":"2022060521170780611_j_popets-2021-0081_ref_076","unstructured":"[76] Radim \u0158eh\u016f\u0159ek and Petr Sojka. Software Framework for Topic Modelling with Large Corpora. In Proceedings of the LREC 2010 Workshop on New Challenges for NLP Frameworks, pages 45\u201350, Valletta, Malta, May 2010. ELRA."},{"key":"2022060521170780611_j_popets-2021-0081_ref_077","doi-asserted-by":"crossref","unstructured":"[77] Helena G\u00f3mez-Adorno, Juan-Pablo Posadas-Dur\u00e1n, Grigori Sidorov, and David Pinto. Document embeddings learned on various types of n-grams for cross-topic authorship attribution. Computing, 100(7):741\u2013756, 2018.10.1007\/s00607-018-0587-8","DOI":"10.1007\/s00607-018-0587-8"},{"key":"2022060521170780611_j_popets-2021-0081_ref_078","doi-asserted-by":"crossref","unstructured":"[78] Erich Schubert, J\u00f6rg Sander, Martin Ester, Hans Peter Kriegel, and Xiaowei Xu. DBSCAN Revisited, Revisited: Why and How You Should (Still) Use DBSCAN. ACM Transactions on Database Systems (TODS), 42(3):1\u201321, 2017.","DOI":"10.1145\/3068335"},{"key":"2022060521170780611_j_popets-2021-0081_ref_079","unstructured":"[79] Ian H. Witten, Gordon W. Paynter, Eibe Frank, Carl Gutwin, and Craig G. Nevill-Manning. KEA: Practical Automatic Keyphrase Extraction. arXiv preprint arXiv:cs\/9902007, 1999."},{"key":"2022060521170780611_j_popets-2021-0081_ref_080","doi-asserted-by":"crossref","unstructured":"[80] Xiaojun Wan and Jianguo Xiao. CollabRank: Towards a Collaborative Approach to Single-Document Keyphrase Extraction. In Proceedings of the 22nd International Conference on Computational Linguistics (Coling 2008), pages 969\u2013976, 2008.","DOI":"10.3115\/1599081.1599203"},{"key":"2022060521170780611_j_popets-2021-0081_ref_081","doi-asserted-by":"crossref","unstructured":"[81] Olena Medelyan, Eibe Frank, and Ian H. Witten. Human-competitive tagging using automatic keyphrase extraction. In Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing (EMNLP \u201909), pages 1318\u20131327, 2009.10.3115\/1699648.1699678","DOI":"10.3115\/1699648.1699678"},{"key":"2022060521170780611_j_popets-2021-0081_ref_082","unstructured":"[82] Samhaa R. El-Beltagy and Ahmed Rafea. KP-Miner: Participation in SemEval-2. In Proceedings of the 5th International Workshop on Semantic Evaluation, pages 190\u2013193. Association for Computational Linguistics, 2010."},{"key":"2022060521170780611_j_popets-2021-0081_ref_083","unstructured":"[83] Thuy Dung Nguyen and Minh-Thang Luong. WINGNUS: Keyphrase Extraction Utilizing Document Logical Structure. In Proceedings of the 5th International Workshop on Semantic Evaluation, pages 166\u2013169. Association for Computational Linguistics, 2010."},{"key":"2022060521170780611_j_popets-2021-0081_ref_084","unstructured":"[84] Adrien Bougouin, Florian Boudin, and B\u00e9atrice Daille. TopicRank: Graph-based Topic Ranking for Keyphrase Extraction. In Proceedings of the Sixth International Joint Conference on Natural Language Processing, pages 543\u2013551. Asian Federation of Natural Language Processing, 2013."},{"key":"2022060521170780611_j_popets-2021-0081_ref_085","doi-asserted-by":"crossref","unstructured":"[85] Lucas Sterckx, Thomas Demeester, Johannes Deleu, and Chris Develder. Topical Word Importance for Fast Keyphrase Extraction. In Proceedings of the 24th International Conference on World Wide Web (WWW \u201915 Companion), pages 121\u2013122, 2015.10.1145\/2740908.2742730","DOI":"10.1145\/2740908.2742730"},{"key":"2022060521170780611_j_popets-2021-0081_ref_086","doi-asserted-by":"crossref","unstructured":"[86] Soheil Danesh, Tamara Sumner, and James H. Martin. SGRank: Combining Statistical and Graphical Methods to Improve the State of the Art in Unsupervised Keyphrase Extraction. In Proceedings of the Fourth Joint Conference on Lexical and Computational Semantics, pages 117\u2013126. Association for Computational Linguistics, 2015.10.18653\/v1\/S15-1013","DOI":"10.18653\/v1\/S15-1013"},{"key":"2022060521170780611_j_popets-2021-0081_ref_087","doi-asserted-by":"crossref","unstructured":"[87] Corina Florescu and Cornelia Caragea. PositionRank: An Unsupervised Approach to Keyphrase Extraction from Scholarly Documents. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1105\u20131115. Association for Computational Linguistics, 2017.","DOI":"10.18653\/v1\/P17-1102"},{"key":"2022060521170780611_j_popets-2021-0081_ref_088","doi-asserted-by":"crossref","unstructured":"[88] Rui Meng, Sanqiang Zhao, Shuguang Han, Daqing He, Peter Brusilovsky, and Yu Chi. Deep Keyphrase Generation. arXiv preprint arXiv:1704.06879, 2017.","DOI":"10.18653\/v1\/P17-1054"},{"key":"2022060521170780611_j_popets-2021-0081_ref_089","doi-asserted-by":"crossref","unstructured":"[89] Florian Boudin. Unsupervised Keyphrase Extraction with Multipartite Graphs. arXiv preprint arXiv:1803.08721, 2018.10.18653\/v1\/N18-2105","DOI":"10.18653\/v1\/N18-2105"},{"key":"2022060521170780611_j_popets-2021-0081_ref_090","doi-asserted-by":"crossref","unstructured":"[90] Ricardo Campos, V\u00edtor Mangaravite, Arian Pasquali, Al\u00edpio M\u00e1rio Jorge, C\u00e9lia Nunes, and Adam Jatowt. A Text Feature Based Automatic Keyword Extraction Method for Single Documents. In European Conference on Information Retrieval, pages 684\u2013691. Springer, 2018.10.1007\/978-3-319-76941-7_63","DOI":"10.1007\/978-3-319-76941-7_63"},{"key":"2022060521170780611_j_popets-2021-0081_ref_091","doi-asserted-by":"crossref","unstructured":"[91] Ricardo Campos, V\u00edtor Mangaravite, Arian Pasquali, Al\u00edpio M\u00e1rio Jorge, C\u00e9lia Nunes, and Adam Jatowt. YAKE! Collection-independent Automatic Keyword Extractor. In European Conference on Information Retrieval, pages 806\u2013810. Springer, 2018.10.1007\/978-3-319-76941-7_80","DOI":"10.1007\/978-3-319-76941-7_80"},{"key":"2022060521170780611_j_popets-2021-0081_ref_092","doi-asserted-by":"crossref","unstructured":"[92] Ricardo Campos, V\u00edtor Mangaravite, Arian Pasquali, Al\u00edpio Jorge, C\u00e9lia Nunes, and Adam Jatowt. YAKE! Keyword extraction from single documents using multiple local features. Information Sciences, 509:257\u2013289, 2020.10.1016\/j.ins.2019.09.013","DOI":"10.1016\/j.ins.2019.09.013"},{"key":"2022060521170780611_j_popets-2021-0081_ref_093","doi-asserted-by":"crossref","unstructured":"[93] Swagata Duari and Vasudha Bhatnagar. sCAKE: Semantic Connectivity Aware Keyword Extraction. Information Sciences, 477:100\u2013117, 2019.","DOI":"10.1016\/j.ins.2018.10.034"},{"key":"2022060521170780611_j_popets-2021-0081_ref_094","doi-asserted-by":"crossref","unstructured":"[94] Claude Sammut and Geoffrey I. Webb. Tf-idf. In Encyclopedia of Machine Learning and Data Mining, pages 1274\u20131274. Springer US, Boston, MA, 2017.10.1007\/978-1-4899-7687-1_832","DOI":"10.1007\/978-1-4899-7687-1_832"},{"key":"2022060521170780611_j_popets-2021-0081_ref_095","unstructured":"[95] Gael Varoquaux. Joblib: running Python functions as pipeline jobs. https:\/\/joblib.readthedocs.io\/, 2020. [Online; accessed 15 June 2021]."},{"key":"2022060521170780611_j_popets-2021-0081_ref_096","doi-asserted-by":"crossref","unstructured":"[96] Joel Nothman, Hanmin Qin, and Roman Yurchak. Stop Word Lists in Free Open-source Software Packages. In Proceedings of Workshop for NLP Open Source Software (NLP-OSS), pages 7\u201312, 2018.10.18653\/v1\/W18-2502","DOI":"10.18653\/v1\/W18-2502"},{"key":"2022060521170780611_j_popets-2021-0081_ref_097","unstructured":"[97] Florian Boudin. pke: an open source python-based keyphrase extraction toolkit. In Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: System Demonstrations, pages 69\u201373, Osaka, Japan, December 2016."},{"key":"2022060521170780611_j_popets-2021-0081_ref_098","doi-asserted-by":"crossref","unstructured":"[98] Victor Le Pochat, Tom Van Goethem, Samaneh Tajalizadehkhoob, Maciej Korczy\u00abski, and Wouter Joosen. Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation. In Proceedings of the 26th Annual Network and Distributed System Security Symposium (NDSS 2019). The Internet Society, February 2019.10.14722\/ndss.2019.23386","DOI":"10.14722\/ndss.2019.23386"},{"key":"2022060521170780611_j_popets-2021-0081_ref_099","doi-asserted-by":"crossref","unstructured":"[99] Steven Englehardt and Arvind Narayanan. Online Tracking: A 1-million-site Measurement and Analysis. In Proceedings of the 26th ACM Conference on Computer and Communications Security, pages 1388\u20131401, 2016.10.1145\/2976749.2978313","DOI":"10.1145\/2976749.2978313"},{"key":"2022060521170780611_j_popets-2021-0081_ref_100","unstructured":"[100] Adam Cohen. FuzzyWuzzy: Fuzzy String Matching in Python. https:\/\/github.com\/seatgeek\/fuzzywuzzy, 2011. [Online; accessed 15 December 2020]."},{"key":"2022060521170780611_j_popets-2021-0081_ref_101","doi-asserted-by":"crossref","unstructured":"[101] Harald Cram\u00e9r. Mathematical Methods of Statistics. Princeton University Press, 1946.10.1515\/9781400883868","DOI":"10.1515\/9781400883868"},{"key":"2022060521170780611_j_popets-2021-0081_ref_102","unstructured":"[102] Rami Al-Rfou, Bryan Perozzi, and Steven Skiena. Polyglot: Distributed Word Representations for Multilingual NLP. In Proceedings of the Seventeenth Conference on Computational Natural Language Learning, pages 183\u2013192. Association for Computational Linguistics, 2013."},{"key":"2022060521170780611_j_popets-2021-0081_ref_103","unstructured":"[103] Sebastian Raschka. Python Machine Learning. Packt Publishing Ltd, 2015."},{"key":"2022060521170780611_j_popets-2021-0081_ref_104","doi-asserted-by":"crossref","unstructured":"[104] Fabrice Colas and Pavel Brazdil. Comparison of SVM and Some Older Classification Algorithms in Text Classification Tasks. In IFIP AI: International Conference on Artificial Intelligence in Theory and Practice, pages 169\u2013178. Springer, 2006.10.1007\/978-0-387-34747-9_18","DOI":"10.1007\/978-0-387-34747-9_18"},{"key":"2022060521170780611_j_popets-2021-0081_ref_105","doi-asserted-by":"crossref","unstructured":"[105] Kanish Shah, Henil Patel, Devanshi Sanghvi, and Manan Shah. A Comparative Analysis of Logistic Regression, Random Forest and KNN Models for the Text Classification. Augmented Human Research, 5(1):1\u201316, 2020.10.1007\/s41133-020-00032-0","DOI":"10.1007\/s41133-020-00032-0"},{"key":"2022060521170780611_j_popets-2021-0081_ref_106","doi-asserted-by":"crossref","unstructured":"[106] Leo Breiman. Random Forests. Machine Learning, 45(1):5\u201332, 2001.10.1023\/A:1010933404324","DOI":"10.1023\/A:1010933404324"},{"key":"2022060521170780611_j_popets-2021-0081_ref_107","doi-asserted-by":"crossref","unstructured":"[107] Ciyou Zhu, Richard H. Byrd, Peihuang Lu, and Jorge Nocedal. Algorithm 778: L-BFGS-B: Fortran subroutines for large-scale bound-constrained optimization. ACM Transactions on Mathematical Software (TOMS), 23(4):550\u2013560, 1997.","DOI":"10.1145\/279232.279236"},{"key":"2022060521170780611_j_popets-2021-0081_ref_108","unstructured":"[108] Pedro G. Fonseca and Hugo D. Lopes. Calibration of Machine Learning Classifiers for Probability of Default Modelling. arXiv preprint arXiv:1710.08901, 2017."},{"key":"2022060521170780611_j_popets-2021-0081_ref_109","unstructured":"[109] Scott M. Lundberg and Su-In Lee. A Unified Approach to Interpreting Model Predictions. In Proceedings of the 31st International Conference on Neural Information Processing Systems, page 4768\u20134777. ACM, 2017."}],"container-title":["Proceedings on Privacy Enhancing Technologies"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.sciendo.com\/pdf\/10.2478\/popets-2021-0081","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,9,5]],"date-time":"2024-09-05T00:08:53Z","timestamp":1725494933000},"score":1,"resource":{"primary":{"URL":"https:\/\/petsymposium.org\/popets\/2021\/popets-2021-0081.php"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,7,23]]},"references-count":109,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2021,7,23]]},"published-print":{"date-parts":[[2021,10,1]]}},"alternative-id":["10.2478\/popets-2021-0081"],"URL":"https:\/\/doi.org\/10.2478\/popets-2021-0081","relation":{},"ISSN":["2299-0984"],"issn-type":[{"value":"2299-0984","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,7,23]]}}}