{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,10]],"date-time":"2026-06-10T16:29:26Z","timestamp":1781108966987,"version":"3.54.1"},"reference-count":69,"publisher":"IGI Global Scientific Publishing","issue":"2","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2011,4,1]]},"abstract":"<p>This paper proposes a paradigm for using speech to interact with computers, one that complements and extends traditional spoken dialogue systems: speech for content creation. The literature in automatic speech recognition (ASR), natural language processing (NLP), sentiment detection, and opinion mining is surveyed to argue that the time has come to use mobile devices to create content on-the-fly. Recent work in user modelling and recommender systems is examined to support the claim that using speech in this way can result in a useful interface to uniquely personalizable data. A data collection effort recently undertaken to help build a prototype system for spoken restaurant reviews is discussed. This vision critically depends on mobile technology, for enabling the creation of the content and for providing ancillary data to make its processing more relevant to individual users. This type of system can be of use where only limited speech processing is possible.<\/p>","DOI":"10.4018\/jmhci.2011040103","type":"journal-article","created":{"date-parts":[[2011,10,19]],"date-time":"2011-10-19T12:41:21Z","timestamp":1319028081000},"page":"35-49","source":"Crossref","is-referenced-by-count":0,"title":["Speech for Content Creation"],"prefix":"10.4018","volume":"3","author":[{"given":"Joseph","family":"Polifroni","sequence":"first","affiliation":[{"name":"Nokia Research Center, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Imre","family":"Kiss","sequence":"additional","affiliation":[{"name":"Nokia Research Center, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Stephanie","family":"Seneff","sequence":"additional","affiliation":[{"name":"MIT CSAIL, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"2432","reference":[{"key":"jmhci.2011040103-0","doi-asserted-by":"crossref","unstructured":"Abdul-Rahman, A., & Hailes, S. (1997). A distributed trust model. In Proceedings of the Workshop on New Security Paradigms (pp. 48-60).","DOI":"10.1145\/283699.283739"},{"key":"jmhci.2011040103-1","doi-asserted-by":"publisher","DOI":"10.1145\/1055709.1055714"},{"key":"jmhci.2011040103-2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2005.99"},{"key":"jmhci.2011040103-3","unstructured":"Appelt, D., & Martin, D. (1999). Named entity extraction from speech: Approach and results using the textpro system. In Proceedings of the DARPA Broadcast News Workshop (pp. 51-54)."},{"key":"jmhci.2011040103-4","doi-asserted-by":"publisher","DOI":"10.1145\/245108.245124"},{"key":"jmhci.2011040103-5","unstructured":"Barker, E., Polifroni, J., Walker, M., & Gaizauskas, R. (2009). Angle-seeking as a scenario for task-based evaluation of information access technology. Paper presented at the International Workshop on Intelligent User Interfaces, Sanibel Island, FL."},{"key":"jmhci.2011040103-6","doi-asserted-by":"crossref","unstructured":"Barnard, E., Davel, M., & van Heerden, C. (2009). ASR corpus design for resource-scarce languages. In Proceedings of the 10th Annual Conference of the International Speech Communication Association (pp. 2847-2850).","DOI":"10.21437\/Interspeech.2009-727"},{"key":"jmhci.2011040103-7","unstructured":"Barnard, E., Davel, M., & van Huyssteen, G. (2010). Speech technology for information access: A South African case study. In Proceedings of the AAAI Symposium on Artificial Intelligence, Palo Alto, CA (pp. 8-13)."},{"key":"jmhci.2011040103-8","doi-asserted-by":"publisher","DOI":"10.1016\/j.specom.2003.07.003"},{"key":"jmhci.2011040103-9","doi-asserted-by":"publisher","DOI":"10.1016\/0957-4174(95)00011-W"},{"key":"jmhci.2011040103-10","doi-asserted-by":"publisher","DOI":"10.1017\/S1351324904003468"},{"issue":"1","key":"jmhci.2011040103-11","doi-asserted-by":"crossref","first-page":"569","DOI":"10.1613\/jair.2633","article-title":"Learning document-level semantic properties from free-text annotations.","volume":"34","author":"S. R. K.Branavan","year":"2009","journal-title":"Journal of Artificial Intelligence Research"},{"key":"jmhci.2011040103-12","doi-asserted-by":"crossref","unstructured":"Carenini, G., & Moore, J. D. (2001). A strategy for evaluating generative arguments. In Proceedings of the First International Conference on Natural Language Generation (pp. 1307\u20131314).","DOI":"10.3115\/1118253.1118261"},{"key":"jmhci.2011040103-13","doi-asserted-by":"publisher","DOI":"10.1016\/j.artint.2006.05.003"},{"key":"jmhci.2011040103-14","doi-asserted-by":"crossref","unstructured":"Carenini, G., Ng, R. T., & Zwart, E. (2005). Extracting knowledge from evaluative text. In Proceedings of the 3rd International Conference on Knowledge Capture (pp. 11-18).","DOI":"10.1145\/1088622.1088626"},{"key":"jmhci.2011040103-15","doi-asserted-by":"crossref","unstructured":"Carenini, G., & Rizoli, L. (2009). A multimedia interface for facilitating comparisons of opinions. In Proceedings of the International Conference on Intelligent User Interfaces (pp. 325-334).","DOI":"10.1145\/1502650.1502696"},{"key":"jmhci.2011040103-16","doi-asserted-by":"crossref","unstructured":"Dave, K., Lawrence, S., & Pennock, D. M. (2003). Mining the peanut gallery: Opinion extraction and semantic classification of product reviews. In Proceedings of the 12th International Conference on World Wide Web (pp. 519-528).","DOI":"10.1145\/775152.775226"},{"key":"jmhci.2011040103-17","unstructured":"Demberg, V., & Moore, J. (2006). Information presentation in spoken dialogue systems. In Proceedings of the 11th International Conference of the European Chapter of the Association for Computational Linguistics."},{"key":"jmhci.2011040103-18","unstructured":"Di Fabbrizio, G., Gupta, N., Besana, S., & Mani, P. (2010). Have2eat: A restaurant finder with review summarization for mobile phones. In Proceedings of the 23rd International Conference on Computational Linguistics: Demonstrations (pp. 17-20)."},{"key":"jmhci.2011040103-19","unstructured":"Dowman, M., Tablan, V., Cunningham, H., Ursu, C., & Popov, B. (2005). Semantically enhanced television news through web and video integration. Paper presented at the Second European Semantic Web Conference Workshop, Crete, Greece."},{"key":"jmhci.2011040103-20","doi-asserted-by":"crossref","unstructured":"Favre, B., B\u00e9chet, F., & Noc\u00e9ra, P. (2005). Robust named entity extraction from large spoken archives. In Proceedings of the Conference on Human Language Technology and Empirical Methods in Natural Language Processing (pp. 491-498).","DOI":"10.3115\/1220575.1220637"},{"key":"jmhci.2011040103-21","doi-asserted-by":"crossref","unstructured":"Feng, J., Bangalore, S., & Gilbert, M. (2009). Role of natural language understanding in voice local search. In Proceedings of the 10th Annual Conference of the International Speech Communication Association (pp. 1859-1862).","DOI":"10.21437\/Interspeech.2009-540"},{"key":"jmhci.2011040103-22","doi-asserted-by":"crossref","unstructured":"Glass, J., Hazen, T., Cyphers, S., Malioutov, I., Huynh, D., & Barzilay, R. (2007). Recent progress in the MIT spoken lecture processing project. In Proceeding of the 8th Annual Conference of the International Communication Association (pp. 2553-2556).","DOI":"10.21437\/Interspeech.2007-678"},{"key":"jmhci.2011040103-23","doi-asserted-by":"crossref","unstructured":"Goldberg, A. B., & Zhu, X. (2006). Seeing stars when there aren\u2019t many stars: Graph-based semi-supervised learning for sentiment categorization. In Proceedings of TextGraphs: The First Workshop on Graph Based Methods for Natural Language Processing (pp. 45-52)","DOI":"10.3115\/1654758.1654769"},{"key":"jmhci.2011040103-24","unstructured":"Good, N., Schafer, J. B., Konstan, J. A., Borchers, A., Sarwar, B., Herlocker, J., et al. (1999). Combining collaborative filtering with personal agents for better recommendations. In Proceedings of the Sixteenth National Conference on Artificial Intelligence (pp. 439-446)."},{"key":"jmhci.2011040103-25","doi-asserted-by":"publisher","DOI":"10.1016\/S0167-6393(97)00040-X"},{"key":"jmhci.2011040103-26","unstructured":"Gupta, N., Di Fabbrizio, G., & Haffner, P. (2010). Capturing the stars: Predicting rankings for service and product reviews. In Proceedings of the NAACL HLT Workshop on Semantic Search (pp. 36-43)."},{"key":"jmhci.2011040103-27","doi-asserted-by":"crossref","unstructured":"Hillard, D., Huang, Z., Ji, H., Grishman, R., Hakkani-T\u00fcr, D., Harper, M., et al. (2006). Impact of automatic comma prediction on POS\/name tagging of speech. In Proceedings of the IEEE\/ACL Workshop on Spoken Language Technology (pp. 58-61).","DOI":"10.1109\/SLT.2006.326816"},{"key":"jmhci.2011040103-28","doi-asserted-by":"crossref","unstructured":"Horlock, J., & King, S. (2003). Discriminative methods for improving named entity extraction on speech data. In Proceedings of the 8th European Conference on Speech Communication and Technology (pp. 2765-2768).","DOI":"10.21437\/Eurospeech.2003-737"},{"key":"jmhci.2011040103-29","unstructured":"Hu, M., & Liu, B. (2004). Mining opinion features in customer reviews. In Proceedings of the 19th National Conference on Artificial Intelligence (pp. 755-760)."},{"key":"jmhci.2011040103-30","doi-asserted-by":"crossref","unstructured":"Huang, J., Zweig, G., & Padmanabhan, M. (2001). Information extraction from voicemail. In Proceedings of the Conference of the Association for Computational Linguistics (pp. 290-297).","DOI":"10.3115\/1073012.1073051"},{"key":"jmhci.2011040103-31","doi-asserted-by":"crossref","unstructured":"Hurst, M., & Nigam, K. (2004). Retrieving topical sentiments from online document collections. Document Recognition and Retrieval 11, 5296, 27-34.","DOI":"10.1117\/12.529422"},{"key":"jmhci.2011040103-32","unstructured":"International Telecommunications Union. (2009). ICT statistics. Retrieved from http:\/\/www.itu.int\/ITU-D\/ict\/statistics\/"},{"key":"jmhci.2011040103-33","doi-asserted-by":"crossref","unstructured":"Jansche, M., & Abney, S. P. (2002). Information extraction from voicemail transcripts. In Proceedings of the ACL Conference on Empirical Methods in Natural Language Processing (Vol. 10).","DOI":"10.3115\/1118693.1118734"},{"key":"jmhci.2011040103-34","first-page":"1159","article-title":"Combining key-phrase detection and subword-based verification for flexible speech understanding. In","volume":"2","author":"T.Kawahara","year":"1997","journal-title":"Proceedings of the International Conference on Acoustics, Speech, and Signal Processing"},{"key":"jmhci.2011040103-35","unstructured":"Kol\u00e1\u0159, \u015c., & Psutka, J. (2004), Automatic punctuation annotation in Czech broadcast news speech. In Proceedings of the International Speech Communication Association (pp. 319-325)."},{"key":"jmhci.2011040103-36","doi-asserted-by":"crossref","unstructured":"Liu, J., & Seneff, S. (2009). Review sentiment scoring via a parse-and-paraphrase paradigm. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (pp. 161-169).","DOI":"10.3115\/1699510.1699532"},{"key":"jmhci.2011040103-37","doi-asserted-by":"publisher","DOI":"10.1109\/TASL.2006.878255"},{"key":"jmhci.2011040103-38","doi-asserted-by":"crossref","unstructured":"Malioutov, I., & Barzilay, R. (2006). Minimum cut model for spoken lecture segmentation. In Proceedings of the 21st International Conference on Computational Linguistics and the 44th Annual Meeting of the Association for Computational Linguistics (pp. 25-32).","DOI":"10.3115\/1220175.1220179"},{"key":"jmhci.2011040103-39","doi-asserted-by":"publisher","DOI":"10.1145\/637848.637862"},{"key":"jmhci.2011040103-40","doi-asserted-by":"crossref","unstructured":"Massa, P., & Avesani, P. (2004). Trust-aware collaborative filtering for recommender systems. In Proceedings of the Federated International Conference on the Move to Meaningful Internet: CoopIS, DOA, ODBASE (pp. 492-508).","DOI":"10.1007\/978-3-540-30468-5_31"},{"key":"jmhci.2011040103-41","doi-asserted-by":"crossref","unstructured":"Miller, D., Schwartz, R., Weischedel, R., & Stone, R. (1999). Named entity extraction from broadcast news. In Proceedings of the DARPA Broadcast News Workshop (pp. 37-40).","DOI":"10.3115\/974147.974191"},{"key":"jmhci.2011040103-42","doi-asserted-by":"crossref","unstructured":"M\u00fcller, C., Grossmann-Hutter, B., Jameson, A., Rummer, R., & Wittig, F. (2001) Recognizing time pressure and cognitive load on the basis of speech: An experimental study. In Proceedings of the 8th International Conference on User Modeling (pp. 24-33).","DOI":"10.1007\/3-540-44566-8_3"},{"key":"jmhci.2011040103-43","doi-asserted-by":"crossref","unstructured":"Munteanu, C., Penn, G., & Zhu, X. (2009). Improving automatic speech recognition for lectures through transformation-based rules learned from minimal data. In Proceedings of the Joint Conference of the 47th Annual Meeting of the ACL and the 4th International Joint Conference on Natural Language Processing of the AFNLP (pp. 764-772).","DOI":"10.3115\/1690219.1690253"},{"key":"jmhci.2011040103-44","unstructured":"Nigam, K., & Hurst, M. (2004). Towards a robust metric of opinion. Paper presented at the AAAI Spring Symposium on Exploring Attitude and Affect in Text, Stanford, CA."},{"key":"jmhci.2011040103-45","unstructured":"Novotney, S., & Callison-Burch, C. (2010). Cheap, fast and good enough: Automatic speech recognition with non-expert transcription. In Proceedings of the Annual NAACL Conference on Human Language Technologies (pp. 207-215)."},{"key":"jmhci.2011040103-46","doi-asserted-by":"crossref","unstructured":"O\u2019Donovan, J., & Smyth, B. (2005). Trust in recommender systems. In Proceedings of the 10th International Conference on Intelligent User Interfaces (pp. 167-174).","DOI":"10.1145\/1040830.1040870"},{"key":"jmhci.2011040103-47","doi-asserted-by":"crossref","unstructured":"Oviatt, S. (2006). Human-centered design meets cognitive load theory: Designing interfaces that help people think. In Proceedings of the 14th Annual ACM International Conference on Multimedia (pp. 871-880).","DOI":"10.1145\/1180639.1180831"},{"key":"jmhci.2011040103-48","unstructured":"Ozowa, V. N. (1997). Information needs of small scale farmers in Africa: The Nigerian example. Consultative Group on International Agricultural Research News, 4(3)."},{"key":"jmhci.2011040103-49","doi-asserted-by":"crossref","unstructured":"Paksima, T., Georgila, K., & Moore, J. D. (2009). Evaluating the effectiveness of information presentation in a full end-to-end dialogue system. In Proceedings of the SIGDIAL Conference on the 10th Annual Meeting of the Special Interest Group on Discourse and Dialogue (pp. 1-10).","DOI":"10.3115\/1708376.1708377"},{"key":"jmhci.2011040103-50","doi-asserted-by":"crossref","unstructured":"Pang, B., & Lee, L. (2005). Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales. In Proceedings of the 43rd Annual Meeting on Association for Computational Linguistics (pp. 115-124).","DOI":"10.3115\/1219840.1219855"},{"key":"jmhci.2011040103-51","doi-asserted-by":"crossref","unstructured":"Pang, B., Lee, L., & Vaithyanathan, S. (2002). Thumbs up? Sentiment classification using machine learning techniques. In Proceedings of the ACL Conference on Empirical Methods in Natural Language Processing (pp. 79-86).","DOI":"10.3115\/1118693.1118704"},{"key":"jmhci.2011040103-52","doi-asserted-by":"crossref","unstructured":"Plauche, M., Nallasamy, U., Pal, J., Wooters, C., & Ramachandran, D. (2006). Speech recognition for illiterate access to information and technology. In Proceedings of the International Conference on Information and Communications Technologies and Development (pp. 83-92).","DOI":"10.1109\/ICTD.2006.301842"},{"key":"jmhci.2011040103-53","doi-asserted-by":"crossref","unstructured":"Polifroni, J., & Seneff, S. (2010). Combining word based features, statistical language models, and parsing for named entity recognition. In Proceedings of the 11th Annual Conference of the International Speech Communication Association (pp. 1289-1292).","DOI":"10.21437\/Interspeech.2010-404"},{"key":"jmhci.2011040103-54","doi-asserted-by":"crossref","unstructured":"Polifroni, J., Seneff, S., Branavan, S. R. K., Wang, C., & Barzilay, R. (2010). Good grief, I can speak it! Preliminary experiments in audio restaurant reviews. Paper presented at the IEEE Workshop on Spoken Language Technology, Berkley, CA.","DOI":"10.1109\/SLT.2010.5700828"},{"key":"jmhci.2011040103-55","unstructured":"Polifroni, J., & Walker, M. (2008). Intentional summaries as cooperative responses in dialogue: Automation and evaluation. In Proceedings of the ACL Conference on Human Language Technologies (pp. 479-487)."},{"key":"jmhci.2011040103-56","doi-asserted-by":"publisher","DOI":"10.1023\/A:1007649029923"},{"issue":"1","key":"jmhci.2011040103-57","first-page":"61","article-title":"TINA: A natural language system for spoken language applications.","volume":"18","author":"S.Seneff","year":"1992","journal-title":"Computational Linguistics"},{"key":"jmhci.2011040103-58","unstructured":"Sherwani, J., Ali, N., Ros\u00e9, C. P., & Rosenfeld, R. (2009). Orality-grounded HCI: Understanding the oral user. Information Technologies & International Development, 5(4)."},{"key":"jmhci.2011040103-59","doi-asserted-by":"crossref","unstructured":"Sherwani, J., Palijo, S., Mirza, S., Ahmed, T., Ali, N., & Rosenfeld, R. (2009). Speech vs. touch-tone: Telephony interfaces for information access by low literate users. In Proceedings of Information and Communications Technologies and Development (pp. 447-457).","DOI":"10.1109\/ICTD.2009.5426682"},{"key":"jmhci.2011040103-60","unstructured":"Snyder, B., & Barzilay, R. (2007). Multiple aspect ranking using the good grief algorithm. In Proceedings of the Human Language Technology Conference of the North American Chapter of the Association of Computational Linguistics (pp. 300-307)."},{"key":"jmhci.2011040103-61","doi-asserted-by":"crossref","unstructured":"Suzuki, Y., Fukumoto, F., & Sekiguchi, Y. (1998). Keyword extraction using term-domain interdependence for dictation of radio news. In Proceedings of the 17th International Conference on Computational Linguistics (pp. 1272-1276).","DOI":"10.3115\/980432.980776"},{"key":"jmhci.2011040103-62","unstructured":"Thong, J.-M. D., Goddeau, D., Litvinova, A., Logan, B., Moreno, P., & Swain, M. (2000). Speechbot: A speech recognition based audio indexing system for the web. In Proceedings of the 6th International Conference on Computer-Assisted Information Retrieval (pp. 106-115)."},{"key":"jmhci.2011040103-63","doi-asserted-by":"crossref","unstructured":"Titov, I., & McDonald, R. (2008). Modeling online reviews with multi-grain topic models. In Proceeding of the 17th International Conference on World Wide Web (pp. 111-120).","DOI":"10.1145\/1367497.1367513"},{"key":"jmhci.2011040103-64","doi-asserted-by":"crossref","unstructured":"van Heerden, C., Barnard, E., & Davel, M. (2009). Basic speech recognition for spoken dialogues. In Proceedings of the 10th Annual Conference of the International Speech Communication Association (pp. 3003-3006).","DOI":"10.21437\/Interspeech.2009-760"},{"key":"jmhci.2011040103-65","doi-asserted-by":"publisher","DOI":"10.1207\/s15516709cog2805_8"},{"key":"jmhci.2011040103-66","doi-asserted-by":"crossref","unstructured":"Ward, W. (1989). Understanding spontaneous speech. In Proceedings of the Workshop on Speech and Natural Language (pp. 137-141).","DOI":"10.3115\/100964.100975"},{"key":"jmhci.2011040103-67","doi-asserted-by":"crossref","unstructured":"Whitaker, S., Hirschberg, J., Amento, B., Stark, L., Bacchiani, M., Isenhour, P., et al. (2002). SCANMail: A voicemail interface that makes speech browsable, readable, and searchable. In Proceedings of the Conference on Human Factors in Computing Systems (pp. 275-282).","DOI":"10.1145\/503376.503426"},{"key":"jmhci.2011040103-68","doi-asserted-by":"crossref","unstructured":"Zhai, L., Fung, P., Schwartz, R., Carpuat, M., & Wu, D. (2004). Using N-best lists for named entity recognition from Chinese speech. In Proceedings of HLT-NAACL: Short Papers (pp. 37-40).","DOI":"10.3115\/1613984.1613994"}],"container-title":["International Journal of Mobile Human Computer Interaction"],"original-title":[],"language":"ng","link":[{"URL":"https:\/\/www.igi-global.com\/viewtitle.aspx?TitleId=53215","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,6,10]],"date-time":"2023-06-10T11:59:10Z","timestamp":1686398350000},"score":1,"resource":{"primary":{"URL":"https:\/\/services.igi-global.com\/resolvedoi\/resolve.aspx?doi=10.4018\/jmhci.2011040103"}},"subtitle":[""],"short-title":[],"issued":{"date-parts":[[2011,4,1]]},"references-count":69,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2011,4]]}},"URL":"https:\/\/doi.org\/10.4018\/jmhci.2011040103","relation":{},"ISSN":["1942-390X","1942-3918"],"issn-type":[{"value":"1942-390X","type":"print"},{"value":"1942-3918","type":"electronic"}],"subject":[],"published":{"date-parts":[[2011,4,1]]}}}