{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,28]],"date-time":"2026-03-28T12:20:44Z","timestamp":1774700444419,"version":"3.50.1"},"reference-count":48,"publisher":"Cambridge University Press (CUP)","issue":"3","license":[{"start":{"date-parts":[[2012,12,14]],"date-time":"2012-12-14T00:00:00Z","timestamp":1355443200000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/www.cambridge.org\/core\/terms"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Nat. Lang. Eng."],"published-print":{"date-parts":[[2014,7]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>While human annotation is crucial for many natural language processing tasks, it is often very expensive and time-consuming. Inspired by previous work on crowdsourcing, we investigate the viability of using non-expert labels instead of gold standard annotations from experts for a machine learning approach to automatic readability prediction. In order to do so, we evaluate two different methodologies to assess the readability of a wide variety of text material: A more traditional setup in which expert readers make readability judgments and a crowdsourcing setup for users who are not necessarily experts. To this purpose two assessment tools were implemented: a tool where expert readers can rank a batch of texts based on readability, and a lightweight crowdsourcing tool, which invites users to provide pairwise comparisons. To validate this approach, readability assessments for a corpus of written Dutch generic texts were gathered. By collecting multiple assessments per text, we explicitly wanted to level out readers' background knowledge and attitude. Our findings show that the assessments collected through both methodologies are highly consistent and that crowdsourcing is a viable alternative to expert labeling. This is a good news as crowdsourcing is more lightweight to use and can have access to a much wider audience of potential annotators. By performing a set of basic machine learning experiments using a feature set that mainly encodes basic lexical and morpho-syntactic information, we further illustrate how the collected data can be used to perform text comparisons or to assign an absolute readability score to an individual text. We do not focus on optimising the algorithms to achieve the best possible results for the learning tasks, but carry them out to illustrate the various possibilities of our data sets. The results on different data sets, however, show that our system outperforms the readability formulas and a baseline language modelling approach. We conclude that readability assessment by comparing texts is a polyvalent methodology, which can be adapted to specific domains and target audiences if required.<\/jats:p>","DOI":"10.1017\/s1351324912000344","type":"journal-article","created":{"date-parts":[[2012,12,14]],"date-time":"2012-12-14T11:59:48Z","timestamp":1355486388000},"page":"293-325","source":"Crossref","is-referenced-by-count":26,"title":["Using the crowd for readability prediction"],"prefix":"10.1017","volume":"20","author":[{"given":"ORPH\u00c9E","family":"DE CLERCQ","sequence":"first","affiliation":[]},{"given":"V\u00c9RONIQUE","family":"HOSTE","sequence":"additional","affiliation":[]},{"given":"BART","family":"DESMET","sequence":"additional","affiliation":[]},{"given":"PHILIP","family":"VAN OOSTEN","sequence":"additional","affiliation":[]},{"given":"MARTINE","family":"DE COCK","sequence":"additional","affiliation":[]},{"given":"LIEVE","family":"MACKEN","sequence":"additional","affiliation":[]}],"member":"56","published-online":{"date-parts":[[2012,12,14]]},"reference":[{"key":"S1351324912000344_ref46","volume-title":"Proceedings of the seventh International Conference on Language Resources and Evaluation (LREC'10)","author":"van Oosten","year":"2010"},{"key":"S1351324912000344_ref37","first-page":"523","volume-title":"Proceedings of the 43rd Annual Meeting of the ACL","author":"Schwarm","year":"2005"},{"key":"S1351324912000344_ref36","first-page":"2471","volume-title":"Proceedings of the 7th International Conference on Language Resources and Evaluation (LREC'10)","author":"Schuurman","year":"2010"},{"key":"S1351324912000344_ref2","doi-asserted-by":"publisher","DOI":"10.1016\/S0271-5309(01)00005-2"},{"key":"S1351324912000344_ref33","first-page":"131","article-title":"The cloze procedure: its validity and utility","volume":"8","author":"Rankin","year":"1959","journal-title":"Eighth Yearbook of the National Reading Conference"},{"key":"S1351324912000344_ref41","volume-title":"Cito Leesbaarheidsindex voor het Basisonderwijs: Verslag van een Leesbaarheidsonderzoek","author":"Staphorsius","year":"1985"},{"key":"S1351324912000344_ref19","doi-asserted-by":"publisher","DOI":"10.3758\/BF03195564"},{"key":"S1351324912000344_ref11","article-title":"De leesbaarheid van landbouwbladen: een onderzoek naar en een toepassing van leesbaarheidsformules","volume":"17","author":"Douma","year":"1960","journal-title":"Bulletin"},{"key":"S1351324912000344_ref31","doi-asserted-by":"publisher","DOI":"10.3115\/1613715.1613742"},{"key":"S1351324912000344_ref42","volume-title":"Proceedings of the International Conference on Spoken Language Processing","author":"Stolcke","year":"2002"},{"key":"S1351324912000344_ref26","doi-asserted-by":"publisher","DOI":"10.5117\/TVT2009.2.LEES356"},{"key":"S1351324912000344_ref28","doi-asserted-by":"publisher","DOI":"10.1016\/j.ijmedinf.2010.02.002"},{"key":"S1351324912000344_ref44","first-page":"191","volume-title":"Selected Papers of the 17th Computational Linguistics in the Netherlands Meeting","author":"van den Bosch","year":"2007"},{"key":"S1351324912000344_ref20","volume-title":"The Technique of Clear Writing","author":"Gunning","year":"1952"},{"key":"S1351324912000344_ref29","unstructured":"McNamara D. S. , Kintsch E. , Songer N. B. , and Kintsch W. 1993. Are good texts always better? Interactions of text coherence, background knowledge, and levels of understanding in learning from text. Technical Report, Institute of Cognitive Science, University of Colorado, Boulder, CO, USA."},{"key":"S1351324912000344_ref8","doi-asserted-by":"crossref","first-page":"407","DOI":"10.1177\/002224378001700401","article-title":"The optimal number of response alternatives for a scale: a review","volume":"17","author":"Cox","year":"1980","journal-title":"Journal of Marketing Research"},{"key":"S1351324912000344_ref1","unstructured":"Anderson R. C. , and Davison A. 1986. Conceptual and empirical bases of readability formulas. Technical Report 392, University of Illinois at Urbana-Champaign, Urbana, IL, USA."},{"key":"S1351324912000344_ref39","first-page":"254","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing","author":"Snow","year":"2008"},{"key":"S1351324912000344_ref9","first-page":"11","article-title":"A formula for predicting readability","volume":"27","author":"Dale","year":"1948","journal-title":"Educational Research Bulletin"},{"key":"S1351324912000344_ref7","doi-asserted-by":"publisher","DOI":"10.1002\/asi.20243"},{"key":"S1351324912000344_ref10","doi-asserted-by":"publisher","DOI":"10.2307\/747483"},{"key":"S1351324912000344_ref3","doi-asserted-by":"publisher","DOI":"10.1145\/345508.345576"},{"key":"S1351324912000344_ref45","first-page":"147","volume-title":"Essential Speech and Language Technology for Dutch. Series: Theory and Applications of Natural Language Processing","author":"van Noord","year":"2012"},{"key":"S1351324912000344_ref4","first-page":"454","article-title":"Onderzoek naar de leesmoeilijkheden van Nederlands proza","volume":"40","author":"Brouwer","year":"1963","journal-title":"Pedagogische Studi\u00ebn"},{"key":"S1351324912000344_ref35","volume-title":"Automatic Text Processing: The Transformation, Analysis and Retrieval of Information by Computer","author":"Salton","year":"1989"},{"key":"S1351324912000344_ref13","volume-title":"Unlocking Language: The Classic Readability Sstudies","author":"DuBay","year":"2007"},{"key":"S1351324912000344_ref30","doi-asserted-by":"publisher","DOI":"10.1016\/j.csl.2008.04.003"},{"key":"S1351324912000344_ref34","first-page":"1","volume-title":"Proceedings of the Workshop on Comparing Corpora, 38th Annual Meeting of the Association for Computational Linguistics","author":"Rayson","year":"2000"},{"key":"S1351324912000344_ref6","volume-title":"Proceedings of HLT\/NAACL 2004","author":"Collins-Thompson","year":"2004"},{"key":"S1351324912000344_ref12","volume-title":"The Principles of Readability","author":"DuBay","year":"2004"},{"key":"S1351324912000344_ref15","first-page":"276","volume-title":"Proceedings of COLING 2010, Poster Vol. 23\u201327","author":"Feng","year":"2010"},{"key":"S1351324912000344_ref43","doi-asserted-by":"publisher","DOI":"10.1162\/coli.09-036-R2-08-050"},{"key":"S1351324912000344_ref17","doi-asserted-by":"publisher","DOI":"10.1037\/h0057532"},{"key":"S1351324912000344_ref21","volume-title":"The Third Workshop on Innovative Use of NLP for Building Educational Applications","author":"Heilman","year":"2008"},{"key":"S1351324912000344_ref18","volume-title":"Proceedings of the EACL 2009 Student Research Workshop","author":"Fran\u00e7ois","year":"2009"},{"key":"S1351324912000344_ref38","first-page":"574","volume-title":"Proceedings of the 10th International Conference on Information Knowledge Management","author":"Si","year":"2001"},{"key":"S1351324912000344_ref22","doi-asserted-by":"publisher","DOI":"10.1075\/term.16.1.01hos"},{"key":"S1351324912000344_ref23","doi-asserted-by":"publisher","DOI":"10.1145\/1498759.1498827"},{"key":"S1351324912000344_ref14","first-page":"229","volume-title":"Proceedings of the 12th Conference of the European Chapter of the ACL","author":"Feng","year":"2009"},{"key":"S1351324912000344_ref27","volume-title":"Proceeding of the International Conference on Asia-Pacific Digital Libraries (ICADL 2011)","author":"Leroy","year":"2011"},{"key":"S1351324912000344_ref25","doi-asserted-by":"crossref","unstructured":"Kincaid J. P. , Jr., R. P. F., Rogers R. L. , and Chissom B. S. 1975. Derivation of new readability formulas (Automated Readability Index, Fog Count and Flesch Reading Ease Formula) for navy-enlisted personnel. Technical Report, Naval Technical Training Command Millington Tenn Research Branch, Department of Navy, Washington, DC.","DOI":"10.21236\/ADA006655"},{"key":"S1351324912000344_ref16","first-page":"80","volume-title":"Proceedings of the NAACL HLT 2010 Workshop on Creating Speech and Language Data with Amazon(tm)s Mechanical Turk","author":"Finin","year":"2010"},{"key":"S1351324912000344_ref5","first-page":"22","article-title":"Word association norms, mutual information, and lexicography","volume":"16","author":"Church","year":"1990","journal-title":"Computational Linguistics"},{"key":"S1351324912000344_ref47","first-page":"429","article-title":"A readability checker with supervised learning using deep indicators","volume":"4","author":"vor der Br\u00fcck","year":"2008","journal-title":"Informatica"},{"key":"S1351324912000344_ref48","doi-asserted-by":"publisher","DOI":"10.1197\/jamia.M2592"},{"key":"S1351324912000344_ref40","volume-title":"Leesbaarheid en Leesvaardigheid. De Ontwikkeling van een Domeingericht Meetinstrument","author":"Staphorsius","year":"1994"},{"key":"S1351324912000344_ref32","volume-title":"Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC'08)","author":"Poesio","year":"2008"},{"key":"S1351324912000344_ref24","volume-title":"Proceedings of the 23rd International Conference on Computational Linguistics","author":"Kate","year":"2010"}],"container-title":["Natural Language Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.cambridge.org\/core\/services\/aop-cambridge-core\/content\/view\/S1351324912000344","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2019,7,7]],"date-time":"2019-07-07T02:43:30Z","timestamp":1562467410000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.cambridge.org\/core\/product\/identifier\/S1351324912000344\/type\/journal_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2012,12,14]]},"references-count":48,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2014,7]]}},"alternative-id":["S1351324912000344"],"URL":"https:\/\/doi.org\/10.1017\/s1351324912000344","relation":{},"ISSN":["1351-3249","1469-8110"],"issn-type":[{"value":"1351-3249","type":"print"},{"value":"1469-8110","type":"electronic"}],"subject":[],"published":{"date-parts":[[2012,12,14]]}}}