{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,5,29]],"date-time":"2025-05-29T13:40:03Z","timestamp":1748526003033,"version":"3.41.0"},"reference-count":29,"publisher":"MIT Press","issue":"3","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Computational Linguistics"],"published-print":{"date-parts":[[2015,9]]},"abstract":"<jats:p>Linguistic corpus design is a critical concern for building rich annotated corpora useful in different domains of applications. For example, speech technologies such as ASR (Automatic Speech Recognition) or TTS (Text-to-Speech) need a huge amount of speech data to train data-driven models or to produce synthetic speech. Collecting data is always related to costs (recording speech, verifying annotations, etc.), and as a rule of thumb, the more data you gather, the more costly your application will be. Within this context, we present in this article solutions to reduce the amount of linguistic text content while maintaining a sufficient level of linguistic richness required by a model or an application. This problem can be formalized as a Set Covering Problem (SCP) and we evaluate two algorithmic heuristics applied to design large text corpora in English and French for covering phonological information or POS labels. The first considered algorithm is a standard greedy solution with an agglomerative\/spitting strategy and we propose a second algorithm based on Lagrangian relaxation. The latter approach provides a lower bound to the cost of each covering solution. This lower bound can be used as a metric to evaluate the quality of a reduced corpus whatever the algorithm applied. Experiments show that a suboptimal algorithm like a greedy algorithm achieves good results; the cost of its solutions is not so far from the lower bound (about 4.35% for 3-phoneme coverings). Usually, constraints in SCP are binary; we proposed here a generalization where the constraints on each covering feature can be multi-valued.<\/jats:p>","DOI":"10.1162\/coli_a_00225","type":"journal-article","created":{"date-parts":[[2015,7,20]],"date-time":"2015-07-20T18:19:23Z","timestamp":1437416363000},"page":"355-383","source":"Crossref","is-referenced-by-count":3,"title":["Large Linguistic Corpus Reduction with SCP Algorithms"],"prefix":"10.1162","volume":"41","author":[{"given":"Nelly","family":"Barbot","sequence":"first","affiliation":[{"name":"IRISA, University of Rennes 1"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Olivier","family":"Bo\u00ebffard","sequence":"additional","affiliation":[{"name":"IRISA, University of Rennes 1"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jonathan","family":"Chevelu","sequence":"additional","affiliation":[{"name":"IRISA, University of Rennes 1"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Arnaud","family":"Delhay","sequence":"additional","affiliation":[{"name":"IRISA, University of Rennes 1"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"281","reference":[{"key":"R1","doi-asserted-by":"publisher","DOI":"10.1145\/1150334.1150336"},{"key":"R2","unstructured":"Barbot, Nelly, Olivier Bo\u00ebffard, and Arnaud Delhay. 2012. Comparing performance of different set-covering strategies for linguistic content optimization in speech corpora. In Proceedings of the International Conference on Language Resources and Evaluation (LREC), pages 969\u2013974, Istanbul."},{"key":"R3","unstructured":"B\u00e9chet, Fr\u00e9d\u00e9ric. 2001. Liaphon: un systeme complet de phon\u00e9tisation de textes. Traitement automatique des langues, 42(1):47\u201367."},{"key":"R4","unstructured":"Bo\u00ebffard, Olivier, Laure Charonnat, S\u00e9bastien Le Maguer, Damien Lolive, and Ga\u00eblle Vidal. 2012. Towards fully automatic annotation of audiobooks for TTS. In Proceedings of the International Conference on Language Resources and Evaluation (LREC), pages 975\u2013980, Istanbul."},{"key":"R6","unstructured":"Bunnell, H. Timothy. 2010. Crafting small databases for unit selection TTS: Effects on intelligibility. In Proceedings of the ISCA Tutorial and Research Workshop on Speech Synthesis (SSW7), pages 40\u201344, Kyoto."},{"key":"R7","unstructured":"Cadic, Didier, C\u00e9dric Boidin, and Christophe d'Alessandro. 2010. Towards optimal TTS corpora. In Proceedings of the International Conference on Language Resources and Evaluation (LREC), pages 99\u2013104, Malta."},{"key":"R8","unstructured":"Candito, Marie, Enrique Henestroza Anguiano, and Djam\u00e9 Seddah. 2011. A word clustering approach to domain adaptation: Effective parsing of biomedical texts. In Proceedings of the 12th International Conference on Parsing Technologies, pages 37\u201342, Dublin."},{"key":"R9","doi-asserted-by":"publisher","DOI":"10.1287\/opre.47.5.730"},{"key":"R10","doi-asserted-by":"publisher","DOI":"10.1023\/A:1019225027893"},{"key":"R11","doi-asserted-by":"publisher","DOI":"10.1007\/BF01581106"},{"key":"R12","unstructured":"Chevelu, Jonathan, Nelly Barbot, Olivier Bo\u00ebffard, and Arnaud Delhay. 2007. Lagrangian relaxation for optimal corpus design. In Proceedings of the ISCA Tutorial and Research Workshop on Speech Synthesis (SSW6), pages 211\u2013216, Bonn."},{"key":"R13","unstructured":"Chevelu, Jonathan, Nelly Barbot, Olivier Bo\u00ebffard, and Arnaud Delhay. 2008. Comparing set-covering strategies for optimal corpus design. In Proceedings of the International Conference on Language Resources and Evaluation (LREC), pages 2951\u20132956, Marrakech."},{"key":"R14","unstructured":"Combescure, Pierre. 1981. 20 listes de 10 phrases phon\u00e9tiquement \u00e9quilibr\u00e9es. Revue d'Acoustique, 56:34\u201338."},{"key":"R15","doi-asserted-by":"publisher","DOI":"10.1287\/mnsc.27.1.1"},{"key":"R16","doi-asserted-by":"crossref","unstructured":"Fran\u00e7ois, H\u00e9l\u00e8ne and Olivier Bo\u00ebffard. 2001. Design of an optimal continuous speech database for text-to-speech synthesis considered as a set covering problem. In Proceedings of the European Conference on Speech Communication and Technology (Eurospeech), pages 829\u2013832, Aalborg.","DOI":"10.21437\/Eurospeech.2001-255"},{"key":"R18","doi-asserted-by":"crossref","unstructured":"Gauvain, Jean-Luc, Lori Lamel, and Maxine Esk\u00e9nazi. 1990. Design considerations and text selection for Bref, a large French readspeech corpus. In Proceedings of the International Conference of Spoken Language Processing (ICSLP), pages 1097\u20131100, Kobe.","DOI":"10.21437\/ICSLP.1990-287"},{"key":"R19","doi-asserted-by":"crossref","unstructured":"Gotab, Pierre, Fr\u00e9d\u00e9ric B\u00e9chet, and G\u00e9raldine Damnati. 2009. Active learning for rule-based and corpus-based spoken language understanding models. In Proceedings of the IEEE workshop on Automatic Speech Recognition and Understanding (ASRU), pages 444\u2013449, Merano.","DOI":"10.1109\/ASRU.2009.5373377"},{"key":"R23","doi-asserted-by":"crossref","unstructured":"Kawai, Hisashi, Seiichi Yamamoto, Norio Higuchi, and Tohru Shimizu. 2000. A design method of speech corpus for text-to-speech synthesis taking account of prosody. In Proceedings of the International Conference on Spoken Language Processing (ICSLP), pages 420\u2013425, Beijing.","DOI":"10.21437\/ICSLP.2000-563"},{"key":"R25","doi-asserted-by":"publisher","DOI":"10.1016\/j.specom.2006.07.002"},{"key":"R26","unstructured":"Krul, Aleksandra, G\u00e9raldine Damnati, Fran\u00e7ois Yvon, C\u00e9dric Boidin, and Thierry Moudenc. 2007. Adaptive database reduction for domain specific speech synthesis. In Proceedings of the ISCA Tutorial and Research Workshop on Speech Synthesis (SSW6), pages 217\u2013222, Bonn."},{"key":"R27","doi-asserted-by":"crossref","unstructured":"Krul, Aleksandra, G\u00e9raldine Damnati, Fran\u00e7ois Yvon, and Thierry Moudenc. 2006. Corpus design based on the Kullback-Leibler divergence for Text-To-Speech synthesis application. In Proceedings of the International Conference on Spoken Language Processing (ICSLP), pages 2030\u20132033, Pittsburgh, PA.","DOI":"10.21437\/Interspeech.2006-397"},{"key":"R28","unstructured":"Neubig, Graham and Shinsuke Mori. 2010. Word-based partial annotation for efficient corpus construction. In Proceedings of the International Conference on Language Resources and Evaluation (LREC), pages 2723\u20132727, Malta."},{"key":"R29","doi-asserted-by":"publisher","DOI":"10.1145\/258533.258641"},{"key":"R30","unstructured":"Rojc, Matej and Zdravko Ka\u010di\u010d. 2000. Design of optimal Slovenian speech corpus for use in the concatenative speech synthesis system. In Proceedings of the International Conference on Language Resources and Evaluation (LREC), pages 321\u2013326, Athens."},{"key":"R34","doi-asserted-by":"crossref","unstructured":"Tian, Jilei and Jani Nurminen. 2009. Optimization of text database using hierachical clustering. In Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), pages 4269\u20134272, Taipei.","DOI":"10.1109\/ICASSP.2009.4960572"},{"key":"R35","doi-asserted-by":"crossref","unstructured":"Tian, Jilei, Jani Nurminen, and Imre Kiss. 2005. Optimal subset selection from text databases. In Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), pages 305\u2013308, Philadelphia, PA.","DOI":"10.1109\/ICASSP.2005.1415111"},{"key":"R36","doi-asserted-by":"publisher","DOI":"10.3115\/1564131.1564140"},{"key":"R37","doi-asserted-by":"crossref","unstructured":"Van Santen, Jan P. H. and Adam L. Buchsbaum. 1997. Methods for optimal text selection. In Proceedings of the European Conference on Speech Communication and Technology (Eurospeech), pages 553\u2013556, Rhodes.","DOI":"10.21437\/Eurospeech.1997-207"},{"key":"R38","doi-asserted-by":"publisher","DOI":"10.1093\/ietisy\/e91-d.3.615"}],"container-title":["Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mitpressjournals.org\/doi\/pdf\/10.1162\/COLI_a_00225","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,5,29]],"date-time":"2025-05-29T13:15:01Z","timestamp":1748524501000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/coli\/article\/41\/3\/355-383\/1525"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2015,9]]},"references-count":29,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2015,9]]}},"alternative-id":["10.1162\/COLI_a_00225"],"URL":"https:\/\/doi.org\/10.1162\/coli_a_00225","relation":{},"ISSN":["0891-2017","1530-9312"],"issn-type":[{"type":"print","value":"0891-2017"},{"type":"electronic","value":"1530-9312"}],"subject":[],"published":{"date-parts":[[2015,9]]}}}