{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,21]],"date-time":"2026-07-21T22:36:52Z","timestamp":1784673412835,"version":"3.55.0"},"reference-count":32,"publisher":"Association for Computing Machinery (ACM)","issue":"9","license":[{"start":{"date-parts":[[2017,8,23]],"date-time":"2017-08-23T00:00:00Z","timestamp":1503446400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["IIS-1149709, IIS-1218209"],"award-info":[{"award-number":["IIS-1149709, IIS-1218209"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100007270","name":"University of Michigan","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100007270","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/100006785","name":"Google","doi-asserted-by":"crossref","award":["Alfred P. Sloan Foundation Fellowship, Microsoft Research Ph.D. Fellowship"],"award-info":[{"award-number":["Alfred P. Sloan Foundation Fellowship, Microsoft Research Ph.D. Fellowship"]}],"id":[{"id":"10.13039\/100006785","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Commun. ACM"],"published-print":{"date-parts":[[2017,8,23]]},"abstract":"<jats:p>Quickly converting speech to text allows deaf and hard of hearing people to interactively follow along with live speech. Doing so reliably requires a combination of perception, understanding, and speed that neither humans nor machines possess alone. In this article, we discuss how our Scribe system combines human labor and machine intelligence in real time to reliably convert speech to text with less than 4s latency. To achieve this speed while maintaining high accuracy, Scribe integrates automated assistance in two ways. First, its user interface directs workers to different portions of the audio stream, slows down the portion they are asked to type, and adaptively determines segment length based on typing speed. Second, it automatically merges the partial input of multiple workers into a single transcript using a custom version of multiple-sequence alignment. Scribe illustrates the broad potential for deeply interleaving human labor and machine intelligence to provide intelligent interactive services that neither can currently achieve alone.<\/jats:p>","DOI":"10.1145\/3068663","type":"journal-article","created":{"date-parts":[[2017,8,24]],"date-time":"2017-08-24T11:49:04Z","timestamp":1503575344000},"page":"93-100","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":16,"title":["Scribe"],"prefix":"10.1145","volume":"60","author":[{"given":"Walter S.","family":"Lasecki","sequence":"first","affiliation":[{"name":"University of Michigan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Christopher D.","family":"Miller","sequence":"additional","affiliation":[{"name":"University of Rochester"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Iftekhar","family":"Naim","sequence":"additional","affiliation":[{"name":"University of Rochester"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Raja","family":"Kushalnagar","sequence":"additional","affiliation":[{"name":"Gallaudet University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Adam","family":"Sadilek","sequence":"additional","affiliation":[{"name":"University of Rochester"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Daniel","family":"Gildea","sequence":"additional","affiliation":[{"name":"University of Rochester"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jeffrey P.","family":"Bigham","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2017,8,23]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/2047196.2047201"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/1866029.1866080"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0167-6393(00)00034-0"},{"key":"e_1_2_1_4_1","volume-title":"Time-scale modification algorithms for music audio signals. Master's thesis","author":"Driedger J.","year":"2011","unstructured":"Driedger , J. Time-scale modification algorithms for music audio signals. Master's thesis , Saarland University , 2011 . Driedger, J. Time-scale modification algorithms for music audio signals. Master's thesis, Saarland University, 2011."},{"key":"e_1_2_1_5_1","volume-title":"Muscle: multiple sequence alignment with high accuracy and high throughput. Nucleic acids research 32, 5","author":"Edgar R.","year":"2004","unstructured":"Edgar , R. Muscle: multiple sequence alignment with high accuracy and high throughput. Nucleic acids research 32, 5 ( 2004 ), 1792--1797. Edgar, R. Muscle: multiple sequence alignment with high accuracy and high throughput. Nucleic acids research 32, 5 (2004), 1792--1797."},{"key":"e_1_2_1_6_1","volume-title":"American Educational Research Association Annual Meeting","author":"Elliot L.B.","year":"2008","unstructured":"Elliot , L.B. , Stinson , M.S. , Easton , D. , Bourgeois , J. College students learning with C-print's education software and automatic speech recognition . In American Educational Research Association Annual Meeting ( New York, NY , 2008 ), AERA. Elliot, L.B., Stinson, M.S., Easton, D., Bourgeois, J. College students learning with C-print's education software and automatic speech recognition. In American Educational Research Association Annual Meeting (New York, NY, 2008), AERA."},{"key":"e_1_2_1_7_1","volume-title":"Salience in the performance of one speech act:the case of definitions. Discource Processes 15, 2 (Apr--June","author":"Flowerdew J.L.","year":"1992","unstructured":"Flowerdew , J.L. Salience in the performance of one speech act:the case of definitions. Discource Processes 15, 2 (Apr--June 1992 ), 165--181. Flowerdew, J.L. Salience in the performance of one speech act:the case of definitions. Discource Processes 15, 2 (Apr--June 1992), 165--181."},{"key":"e_1_2_1_8_1","volume-title":"Proceedings of INTERSPEECH","author":"Metze F.","year":"2016","unstructured":"Metze , F. , Gaur , Y. , Bigham , J. P. Manipulating word lattices to incorporate human corrections . In Proceedings of INTERSPEECH , ( 2016 ). Metze, F., Gaur, Y., Bigham, J. P. Manipulating word lattices to incorporate human corrections. In Proceedings of INTERSPEECH, (2016)."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/2899475.2899478"},{"key":"e_1_2_1_10_1","volume-title":"Interspeech","author":"Glass J.R.","year":"2007","unstructured":"Glass , J.R. , Hazen , T.J. , Cyphers , D.S. , Malioutov , I. , Huynh , D. , Barzilay , R. Recent progress in the MIT spoken lecture processing project . In Interspeech ( 2007 ), 2553--2556. Glass, J.R., Hazen, T.J., Cyphers, D.S., Malioutov, I., Huynh, D., Barzilay, R. Recent progress in the MIT spoken lecture processing project. In Interspeech (2007), 2553--2556."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/2815585.2815729"},{"key":"e_1_2_1_12_1","series-title":"Springer Handbook of Auditory Research","volume-title":"Perspectives on Auditory Research","author":"Gordon-Salant S.","year":"2014","unstructured":"Gordon-Salant , S. Aging , hearing loss, and speech recognition: stop shouting, i can't understand you . In Perspectives on Auditory Research , volume 50 of Springer Handbook of Auditory Research . A.N. Popper and R.R. Fay, eds. Springer New York , 2014 , 211--228. Gordon-Salant, S. Aging, hearing loss, and speech recognition: stop shouting, i can't understand you. In Perspectives on Auditory Research, volume 50 of Springer Handbook of Auditory Research. A.N. Popper and R.R. Fay, eds. Springer New York, 2014, 211--228."},{"key":"e_1_2_1_13_1","first-page":"4","article-title":"Closed-captioned television presentation speed and vocabulary","volume":"140","author":"Jensema C.","year":"1996","unstructured":"Jensema , C. , McCann , R. , Ramsey , S . Closed-captioned television presentation speed and vocabulary . In Am Ann Deaf 140 , 4 ( October 1996 ), 284--292. Jensema, C., McCann, R., Ramsey, S. Closed-captioned television presentation speed and vocabulary. In Am Ann Deaf 140, 4 (October 1996), 284--292.","journal-title":"Am Ann Deaf"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/67450.67472"},{"key":"e_1_2_1_15_1","volume-title":"IEEE European Signal Processing Conference","author":"Kadri H.","year":"2008","unstructured":"Kadri , H. , Davy , M. , Rabaoui , A. , Lachiri , Z. , Ellouze , N. , Robust audio speaker segmentation using one class SVMs . In IEEE European Signal Processing Conference ( Lausanne, Switzerland , 2008 ) ISSN: 2219--5491. Kadri, H., Davy, M., Rabaoui, A., Lachiri, Z., Ellouze, N., et al. Robust audio speaker segmentation using one class SVMs. In IEEE European Signal Processing Conference (Lausanne, Switzerland, 2008) ISSN: 2219--5491."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/2461121.2461142"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/2543578"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/2384916.2384942"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/2380116.2380122"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/2642918.2647367"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.15346\/hc.v1i1.5"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/2596695.2596701"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/2470654.2466269"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/2047196.2047200"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/2441776.2441912"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1093\/deafed\/eni011"},{"key":"e_1_2_1_27_1","volume-title":"Proceedings North American Chapter of the Association for Computational Linguistics (NAACL)","author":"Naim I.","year":"2013","unstructured":"Naim , I. , Gildea , D. , Lasecki , W.S. , Bigham , J.P. Text alignment for real-time crowd captioning . In Proceedings North American Chapter of the Association for Computational Linguistics (NAACL) ( 2013 ), 201--210. Naim, I., Gildea, D., Lasecki, W.S., Bigham, J.P. Text alignment for real-time crowd captioning. In Proceedings North American Chapter of the Association for Computational Linguistics (NAACL) (2013), 201--210."},{"key":"e_1_2_1_28_1","volume-title":"Proceedings of the 2015 International Conference on Autonomous Agents and Multiagent Systems. International Foundation for Autonomous Agents and Multiagent Systems","author":"Salisbury E.","year":"2015","unstructured":"Salisbury , E. , Stein , S. , Ramchurn , S. Real-time opinion aggregation methods for crowd robotics . In Proceedings of the 2015 International Conference on Autonomous Agents and Multiagent Systems. International Foundation for Autonomous Agents and Multiagent Systems ( 2015 ), 841--849. Salisbury, E., Stein, S., Ramchurn, S. Real-time opinion aggregation methods for crowd robotics. In Proceedings of the 2015 International Conference on Autonomous Agents and Multiagent Systems. International Foundation for Autonomous Agents and Multiagent Systems (2015), 841--849."},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1086\/456502"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1108\/17415650680000058"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1089\/cmb.1994.1.337"},{"key":"e_1_2_1_32_1","unstructured":"World Health Organization. Deafness and hearing loss fact sheet N300. http:\/\/www.who.int\/mediacentre\/factsheets\/fs300\/en\/ February 2014.  World Health Organization. Deafness and hearing loss fact sheet N300. http:\/\/www.who.int\/mediacentre\/factsheets\/fs300\/en\/ February 2014."}],"container-title":["Communications of the ACM"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3068663","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3068663","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3068663","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T03:03:44Z","timestamp":1750215824000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3068663"}},"subtitle":["deep integration of human and machine intelligence to caption speech in real time"],"short-title":[],"issued":{"date-parts":[[2017,8,23]]},"references-count":32,"journal-issue":{"issue":"9","published-print":{"date-parts":[[2017,8,23]]}},"alternative-id":["10.1145\/3068663"],"URL":"https:\/\/doi.org\/10.1145\/3068663","relation":{},"ISSN":["0001-0782","1557-7317"],"issn-type":[{"value":"0001-0782","type":"print"},{"value":"1557-7317","type":"electronic"}],"subject":[],"published":{"date-parts":[[2017,8,23]]},"assertion":[{"value":"2017-08-23","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}