{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,8]],"date-time":"2025-09-08T06:35:27Z","timestamp":1757313327503,"version":"3.37.3"},"reference-count":58,"publisher":"Springer Science and Business Media LLC","issue":"4","license":[{"start":{"date-parts":[[2021,9,4]],"date-time":"2021-09-04T00:00:00Z","timestamp":1630713600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,9,4]],"date-time":"2021-09-04T00:00:00Z","timestamp":1630713600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Spanish Ministry of Economy, Industry and Competitiveness, through the Ram\u00f3n y Cajal","award":["RYC-2015-17239"],"award-info":[{"award-number":["RYC-2015-17239"]}]},{"DOI":"10.13039\/100018967","name":"Universitat Pompeu Fabra","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100018967","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Lang Resources &amp; Evaluation"],"published-print":{"date-parts":[[2021,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Research on speech technologies necessitates spoken data, which is usually obtained through read recorded speech, and specifically adapted to the research needs. When the aim is to deal with the prosody involved in speech, the available data must reflect natural and conversational speech, which is usually costly and difficult to get. This paper presents a machine learning-oriented toolkit for collecting, handling, and visualization of speech data, using prosodic heuristic. We present two corpora resulting from these methodologies: PANTED corpus, containing 250 h of English speech from TED Talks, and Heroes corpus containing 8 h of parallel English and Spanish movie speech. We demonstrate their use in two deep learning-based applications: punctuation restoration and machine translation. The presented corpora are freely available to the research community.<\/jats:p>","DOI":"10.1007\/s10579-021-09556-2","type":"journal-article","created":{"date-parts":[[2021,9,4]],"date-time":"2021-09-04T09:02:56Z","timestamp":1630746176000},"page":"925-946","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Corpora compilation for prosody-informed speech processing"],"prefix":"10.1007","volume":"55","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-0700-1159","authenticated-orcid":false,"given":"Alp","family":"\u00d6ktem","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7160-9513","authenticated-orcid":false,"given":"Mireia","family":"Farr\u00fas","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6240-9915","authenticated-orcid":false,"given":"Antonio","family":"Bonafonte","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2021,9,4]]},"reference":[{"key":"9556_CR1","doi-asserted-by":"crossref","unstructured":"Adami, A. G., Mihaescu, R., Reynolds, D. A., & Godfrey, J. J. (2003). Modeling prosodic dynamics for speaker recognition. In 2003 IEEE international conference on acoustics, speech, and signal processing, 2003. Proceedings (ICASSP\u201903) (Vol. 4, pp. IV-788). IEEE.","DOI":"10.1109\/ICASSP.2003.1202761"},{"key":"9556_CR2","doi-asserted-by":"crossref","unstructured":"Almeman, K., Lee, M., & Almiman, A.A. (2013). Multi dialect Arabic speech parallel corpora. In 1st International conference on communications, signal processing, and their applications (ICCSPA) (pp. 1\u20136). IEEE.","DOI":"10.1109\/ICCSPA.2013.6487288"},{"key":"9556_CR3","unstructured":"Avanzi, M., Lacheret-Dujour, A., & Victorri, B. (2008). ANALOR. A tool for semi-automatic annotation of French prosodic structure."},{"key":"9556_CR4","unstructured":"Bahdanau, D., Cho, K., & Bengio, Y. (2014). Neural machine translation by jointly learning to align and translate. CoRR."},{"issue":"2","key":"9556_CR5","doi-asserted-by":"publisher","first-page":"474","DOI":"10.1109\/TASL.2011.2159594","volume":"20","author":"F Batista","year":"2012","unstructured":"Batista, F., Moniz, H., Trancoso, I., & Mamede, N. (2012). Bilingual experiments on automatic recovery of capitalization and punctuation of automatic speech transcripts. IEEE Transactions on Audio, Speech, and Language Processing, 20(2), 474\u2013485.","journal-title":"IEEE Transactions on Audio, Speech, and Language Processing"},{"key":"9556_CR6","unstructured":"Bendazzoli, C., & Sandrelli, A. (2005). An approach to corpus-based interpreting studies: Developing EPIC (European Parliament Interpreting Corpus). In MuTra 2005\u2013Challenges of multidimensional translation (pp. 1\u201312)."},{"key":"9556_CR7","unstructured":"Boersma, P., & Weenink, D. (2019). Praat: Doing phonetics by computer (computer program), version 6.0.46. Retrieved January 3, 2019, from http:\/\/www.praat.org\/ ."},{"key":"9556_CR8","unstructured":"Canavan, A., Graff, D., & Zipperlen, G. (1997). CALLHOME American English Speech LDC97S42. Web download. Linguistic Data Consortium. https:\/\/catalog.ldc.upenn.edu\/LDC97S42 ."},{"key":"9556_CR9","unstructured":"Cettolo, M., Girardi, C., & Federico, M. (2012). Wit3: Web inventory of transcribed and translated talks. In Proceedings of the international conference of the European Association for Machine Translation (EAMT), Trento, Italy (pp. 261\u2013268)."},{"key":"9556_CR10","doi-asserted-by":"crossref","unstructured":"Cho, E., Niehues, J., & Waibel, A. H. (2017). NMT-based segmentation and punctuation insertion for real-time spoken language translation. In Proceedings of Interspeech 2017.","DOI":"10.21437\/Interspeech.2017-1320"},{"key":"9556_CR11","unstructured":"Crystal, D. (2003). A dictionary of linguistics and phonetics. Blackwell Publishing Ltd."},{"key":"9556_CR12","unstructured":"Dom\u00ednguez, M., Farr\u00fas, M., & Wanner, L. (2016a). An automatic prosody tagger for spontaneous speech. In Proceedings of the 26th international conference on computational linguistics (COLING), Osaka, Japan (pp. 377\u2013386)."},{"key":"9556_CR13","unstructured":"Dom\u00ednguez, M., Latorre, I., Farr\u00fas, M., Codina-Filb\u00e0, J., & Wanner, L. (2016b). Praat on the web: An upgrade of Praat for semi-automatic speech annotation. In Proceedings of the 26th international conference on computational linguistics (COLING), Osaka, Japan (pp. 218\u2013222)."},{"key":"9556_CR14","doi-asserted-by":"crossref","unstructured":"Farr\u00fas, M., Lai, C., & Moore, J. D. (2016). Paragraph-based cues for speech synthesis applications. In Proceedings of the speech prosody, Boston, MA.","DOI":"10.21437\/SpeechProsody.2016-235"},{"key":"9556_CR15","doi-asserted-by":"crossref","unstructured":"Favre, B., Grishman, R., Hillard, D., Ji, H., Hakkani-Tur, D., & Ostendorf, M. (2008). Punctuating speech for information extraction. In IEEE international conference on acoustics, speech and signal processing, 2008. ICASSP 2008 (pp. 5013\u20135016). IEEE.","DOI":"10.1109\/ICASSP.2008.4518784"},{"key":"9556_CR16","unstructured":"Federmann, C., & Lewis, W. D. (2016). Microsoft Speech Language Translation (MSLT) corpus: The IWSLT 2016 release for English, French and German. In International workshop on spoken language translation."},{"key":"9556_CR17","doi-asserted-by":"crossref","unstructured":"Fujisaki, H. (1997). Prosody, models, and spontaneous speech. In Y. Sagisaka, N. Campbell & N. Higuchi (Eds.), Computing prosody: Computational models for processing spontaneous speech (pp. 27\u201342). Springer.","DOI":"10.1007\/978-1-4612-2258-3_3"},{"key":"9556_CR18","doi-asserted-by":"crossref","unstructured":"Garrido, J., Codina, M., & Fodge, K. (2018). TransDic, a public domain tool for the generation of phonetic dictionaries in standard and dialectal Spanish and Catalan. In Proceedings of Iberspeech, Barcelona, Spain (pp. 291\u2013295).","DOI":"10.21437\/IberSPEECH.2018-61"},{"key":"9556_CR19","unstructured":"Georgescu, A. L., Cucu, H., Buzo, A., & Burileanu, C. (2020). RSC: A Romanian read speech corpus for automatic speech recognition. In Proceedings of the 12th language resources and evaluation conference (pp. 6606\u20136612)."},{"key":"9556_CR20","unstructured":"Godfrey, J., & Holliman, E. (1993). Switchboard-1 release 2 ldc97s62. DVD. Linguistic Data Consortium."},{"key":"9556_CR21","doi-asserted-by":"crossref","unstructured":"Hermann, K. M., & Blunsom, P. (2014). Multilingual models for compositional distributional semantics. In Proceedings of the 52nd annual meeting of the Association for Computational Linguistics (ACL), Baltimore, Maryland, USA (pp. 58\u201368).","DOI":"10.3115\/v1\/P14-1006"},{"key":"9556_CR22","doi-asserted-by":"crossref","unstructured":"Hillard, D., Huang, Z., Ji, H., Grishman, R., Hakkani-Tur, D., Harper, M., Ostendorf, M., & Wang, W. (2006). Impact of automatic comma prediction on POS\/name tagging of speech. In Proceedings of the IEEE spoken language technology workshop, Palm Beach, Aruba (pp. 58\u201361).","DOI":"10.1109\/SLT.2006.326816"},{"key":"9556_CR23","unstructured":"Huang, Z., Chen, L., & Harper, M. (2006). An open source prosodic feature extraction tool. In Proceedings of the fifth international conference on language resources and evaluation (LREC), Genoa, Italy."},{"key":"9556_CR24","doi-asserted-by":"publisher","unstructured":"Jones, B. E. M. (1994). Exploring the role of punctuation in parsing natural text. In Proceedings of the 15th conference on computational linguistics, COLING \u201994 (Vol. 1, pp. 421\u2013425). Association for Computational Linguistics. https:\/\/doi.org\/10.3115\/991886.991960.","DOI":"10.3115\/991886.991960"},{"key":"9556_CR25","doi-asserted-by":"publisher","unstructured":"K\u00fclebi, B., & \u00d6ktem, A. (2018). Building an open source automatic speech recognition system for Catalan. In Proceedings of IberSPEECH 2018 (pp. 25\u201329). https:\/\/doi.org\/10.21437\/IberSPEECH.2018-6.","DOI":"10.21437\/IberSPEECH.2018-6"},{"key":"9556_CR26","unstructured":"K\u00fclebi, B., \u00d6ktem, A., Peir\u00f3-Lilja, A., Pascual, S., & Farr\u00fas, M. (2020). CATOTRON\u2014A neural text-to-speech system in Catalan. In Proceedings of Interspeech 2020 (pp. 490\u2013491)."},{"key":"9556_CR27","unstructured":"Lacheret, A., Kahane, S., Beliao, J., Dister, A., Gerdes, K., Goldman, J., Obin, N., Pietrandrea, P., & Tchobanov, A. (2014). Rhapsodie: A prosodic-syntactic treebank for spoken French. In N. Calzolari, K. Choukri, T. Declerck, H. Loftsson, B. Maegaard, J. Mariani, A. Moreno, J. Odijk & S. Piperidis (Eds.), Proceedings of the ninth international conference on language resources and evaluation, LREC 2014, Reykjavik, Iceland, May 26\u201331, 2014 (pp. 295\u2013301). European Language Resources Association (ELRA)."},{"key":"9556_CR28","unstructured":"Lison, P., & Tiedemann, J. (2016). Opensubtitles2016: Extracting large parallel corpora from movie and TV subtitles. In LREC 2016, tenth international conference on language resources and evaluation. European Language Resources Association."},{"key":"9556_CR29","unstructured":"Lison, P., Tiedemann, J., & Kouylekov, M. (2018). Open subtitles 2018: Statistical rescoring of sentence alignments in large, noisy parallel corpora. In LREC 2018, eleventh international conference on language resources and evaluation. European Language Resources Association (ELRA)."},{"key":"9556_CR30","unstructured":"Lu, W., & Ng, H. T. (2010). Better punctuation prediction with dynamic conditional random fields. In Proceedings of the 2010 conference on empirical methods in natural language processing (pp. 177\u2013186). Association for Computational Linguistics."},{"key":"9556_CR31","doi-asserted-by":"crossref","unstructured":"Ma, J., Zhang, Y., & Zhu, J. (2014). Punctuation processing for projective dependency parsing. In Proceedings of the 52nd annual meeting of the Association for Computational Linguistics: Short papers (Vol. 2, pp. 791\u2013796). Association for Computational Linguistics. http:\/\/www.aclweb.org\/anthology\/P14-2128.","DOI":"10.3115\/v1\/P14-2128"},{"key":"9556_CR32","unstructured":"Matusov, E., Mauser, A., & Ney, H. (2006). Automatic sentence segmentation and punctuation prediction for spoken language translation. In International workshop on spoken language translation (IWSLT) 2006."},{"key":"9556_CR33","unstructured":"McAuliffe, M., Socolof, M., Mihuc, S., Wagner, M., & Sonderegger, M. (2013). Montreal Forced Aligner: Trainable text-speech alignment using Kaldi. In Proceedings of tools and resources for the analysis of speech prosody (TRASP), Aix-en-Provence, France (pp. 7\u201310)."},{"key":"9556_CR34","doi-asserted-by":"crossref","unstructured":"Mertens, P. (2004). The Prosogram: Semi-automatic transcription of prosody based on a tonal perception model. In Proceedings of the 2nd international conference on speech prosody, Nara, Japan (pp. 549\u2013552).","DOI":"10.21437\/SpeechProsody.2004-127"},{"key":"9556_CR35","doi-asserted-by":"publisher","unstructured":"Nayak, S., Baumann, T., Bhattacharya, S., Karakanta, A., Negri, M., & Turchi, M. (2020). See me speaking? Differentiating on whether words are spoken on screen or off to optimize machine dubbing. In K. P. Truong, D. Heylen, M. Czerwinski, N. Berthouze, M. Chetouani & M. Nakano (Eds.), Companion publication of the 2020 international conference on multimodal interaction, ICMI companion 2020, virtual event, The Netherlands, October, 2020 (pp. 130\u2013134). ACM. https:\/\/doi.org\/10.1145\/3395035.3425640.","DOI":"10.1145\/3395035.3425640"},{"key":"9556_CR36","unstructured":"\u00d6ktem, A. (2019). Incorporating prosody into neural speech processing pipelines: Applications on automatic speech transcription and spoken language machine translation. PhD Thesis, Universitat Pompeu Fabra."},{"key":"9556_CR37","doi-asserted-by":"crossref","unstructured":"\u00d6ktem, A., Farr\u00fas, M., & Bonafonte, A. (2018). Bilingual prosodic dataset compilation for spoken language translation. In Proceedings of the Iberspeech, Barcelona, Spain (pp. 20\u201324).","DOI":"10.21437\/IberSPEECH.2018-5"},{"key":"9556_CR38","doi-asserted-by":"publisher","unstructured":"\u00d6ktem, A., Farr\u00fas, M., & Bonafonte, A. (2019). Prosodic phrase alignment for machine dubbing. In G. Kubin & Z. Kacic (Eds.), Interspeech 2019, 20th annual conference of the International Speech Communication Association, Graz, Austria, September 15\u201319, 2019 (pp. 4215\u20134219). ISCA. https:\/\/doi.org\/10.21437\/Interspeech.2019-1621.","DOI":"10.21437\/Interspeech.2019-1621"},{"key":"9556_CR39","doi-asserted-by":"crossref","unstructured":"\u00d6ktem, A., Farr\u00fas, M., & Wanner, L. (2017a). Attentional parallel RNNs for generating punctuation in transcribed speech. In Statistical language and speech processing (pp. 131\u2013142). Springer.","DOI":"10.1007\/978-3-319-68456-7_11"},{"key":"9556_CR40","unstructured":"\u00d6ktem, A., Farr\u00fas, M., & Wanner, L. (2017b). Prosograph: A tool for prosody visualisation of large speech corpora. In Proceedings of Interspeech, Stockholm, Sweden (pp. 809\u2013810)."},{"key":"9556_CR41","unstructured":"Ostendorf, M., Price, P., & Shattuck-Hufnagel, S. (1996). Boston University radio speech corpus LDC96S36. DVD. Linguistic Data Consortium. https:\/\/catalog.ldc.upenn.edu\/LDC96S36."},{"key":"9556_CR42","doi-asserted-by":"crossref","unstructured":"Papineni, K., Roukos, S., Ward, T., & Zhu, W. J. (2002). BLEU: A method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting on Association for Computational Linguistics (ACL), Philadelphia, Pennsylvania (pp. 311\u2013318).","DOI":"10.3115\/1073083.1073135"},{"key":"9556_CR43","doi-asserted-by":"crossref","unstructured":"Pappas, N., & Popescu-Belis, A. (2013) Sentiment analysis of user comments for one-class collaborative filtering over TED talks. In Proceedings of the 36th international ACM SIGIR conference on research and development in information retrieval, Dublin, Ireland (pp. 773\u2013776).","DOI":"10.1145\/2484028.2484116"},{"key":"9556_CR44","first-page":"1","volume":"54","author":"E Parada-Cabaleiro","year":"2019","unstructured":"Parada-Cabaleiro, E., Costantini, G., Batliner, A., Schmitt, M., & Schuller, B. W. (2019). DEMoS: An Italian emotional speech corpus. Language Resources and Evaluation, 54, 1\u201343.","journal-title":"Language Resources and Evaluation"},{"key":"9556_CR45","unstructured":"Peitz, S., Freitag, M., Mauser, A., & Ney, H. (2011). Modeling punctuation prediction as machine translation. In International workshop on spoken language translation (IWSLT) 2011."},{"key":"9556_CR46","doi-asserted-by":"crossref","unstructured":"Rosenberg, A. (2010). AuToBi\u2014A tool for automatic ToBi annotation. In Eleventh annual conference of the International Speech Communication Association.","DOI":"10.21437\/Interspeech.2010-71"},{"key":"9556_CR47","unstructured":"Rousseau, A., Del\u00e9glise, P., & Est\u00e8ve, Y. (2012). TED-LIUM: An automatic speech recognition dedicated corpus. In Proceedings of the eighth international conference on language resources and evaluation (LREC), Istanbul, Turkey."},{"key":"9556_CR48","unstructured":"Sloetjes, H., & Wittenburg, P. (2008). Annotation by category\u2014ELAN and ISO DCR. In 6th International conference on language resources and evaluation (LREC 2008)."},{"key":"9556_CR49","unstructured":"Spitkovsky, V. I., Alshawi, H., & Jurafsky, D. (2011). Punctuation: Making a point in unsupervised dependency parsing. In Proceedings of the fifteenth conference on computational natural language learning, CoNLL \u201911 (pp. 19\u201328). Association for Computational Linguistics. http:\/\/dl.acm.org\/citation.cfm?id=2018936.2018939."},{"key":"9556_CR50","unstructured":"Sutskever, I., Vinyals, O., & Le, Q. V. (2014). Sequence to sequence learning with neural networks. In Proceedings of the 27th international conference on neural information processing systems\u2014NIPS\u201914 (Vol. 2, pp. 3104\u20133112). MIT Press."},{"key":"9556_CR51","unstructured":"Takamichi, S., & Saruwatari, H. (2018). CPJD corpus: Crowdsourced parallel speech corpus of Japanese dialects. In Proceedings of the eleventh international conference on language resources and evaluation (LREC 2018). European Language Resources Association (ELRA). https:\/\/www.aclweb.org\/anthology\/L18-1067."},{"key":"9556_CR52","unstructured":"Tiedemann, J. (2012). Parallel data, tools and interfaces in OPUS. In LREC (Vol. 2012, pp. 2214\u20132218)."},{"key":"9556_CR53","doi-asserted-by":"crossref","unstructured":"Tilk, O., & Alum\u00e4e, T. (2016). Bidirectional recurrent neural network with attention mechanism for punctuation restoration. In Proceedings of Interspeech, San Francisco, CA, USA (pp. 3047\u20133051).","DOI":"10.21437\/Interspeech.2016-1517"},{"key":"9556_CR54","doi-asserted-by":"crossref","unstructured":"Tsiartas, A., Ghosh, P., Georgiou, P. G., & Narayanan, S. (2011). Bilingual audio-subtitle extraction using automatic segmentation of movie audio. In Proceedings of the international conference on acoustics, speech and signal processing (ICASSP), Prague, Czech Republic (pp. 5624\u20135627).","DOI":"10.1109\/ICASSP.2011.5947635"},{"key":"9556_CR55","unstructured":"Wester, M. (2010). The EMIME Bilingual Database. Technical report. The University of Edinburgh."},{"key":"9556_CR57","unstructured":"Xu, Y. (2013). ProsodyPro\u2014A tool for large-scale systematic prosody analysis. In Proceedings of tools and resources for the analysis of speech prosody (TRASP), Aix-en-Provence, France (pp. 7\u201310)."},{"issue":"7","key":"9556_CR56","doi-asserted-by":"publisher","first-page":"1063","DOI":"10.1007\/s11265-017-1289-8","volume":"90","author":"C Xu","year":"2017","unstructured":"Xu, C., Xie, L., & Xiao, X. (2017). A bidirectional LSTM approach with word embeddings for sentence boundary detection. Journal of Signal Processing Systems, 90(7), 1063\u20131075. https:\/\/doi.org\/10.1007\/s11265-017-1289-8.","journal-title":"Journal of Signal Processing Systems"},{"key":"9556_CR58","unstructured":"Zanon Boito, M., Havard, W. N., Garnerin, M., Le Ferrand, \u00c9., & Besacier, L. (2020). MaSS: A large and clean multilingual corpus of sentence-aligned spoken utterances extracted from the Bible. In Proceedings of the 12th language resources and evaluation conference, Marseille, France (pp. 6486\u20136493). https:\/\/hal.archives-ouvertes.fr\/hal-02611059."}],"container-title":["Language Resources and Evaluation"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10579-021-09556-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10579-021-09556-2\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10579-021-09556-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,9,7]],"date-time":"2024-09-07T19:51:44Z","timestamp":1725738704000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10579-021-09556-2"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,9,4]]},"references-count":58,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2021,12]]}},"alternative-id":["9556"],"URL":"https:\/\/doi.org\/10.1007\/s10579-021-09556-2","relation":{},"ISSN":["1574-020X","1574-0218"],"issn-type":[{"type":"print","value":"1574-020X"},{"type":"electronic","value":"1574-0218"}],"subject":[],"published":{"date-parts":[[2021,9,4]]},"assertion":[{"value":"29 July 2021","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"4 September 2021","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}