{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,24]],"date-time":"2026-04-24T17:13:03Z","timestamp":1777050783700,"version":"3.51.4"},"reference-count":72,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:00:00Z","timestamp":1750291200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Comput. Sci."],"abstract":"<jats:sec><jats:title>Introduction<\/jats:title><jats:p>Access to non-speech information (NSI) in video content is essential to creating accessible and engaging video content, particularly for D\/deaf and Hard-of-Hearing (DHH) audiences. In this paper we present an overview of the current state of NSI captioning research, professional practice, and user preferences.<\/jats:p><\/jats:sec><jats:sec><jats:title>Methods<\/jats:title><jats:p>We utilized a comprehensive review approach that combined a systematic literature review methodology with a mixed-methods survey and interview study. 1276 papers were screened with 36 eligible for the final inductive best fit analysis. 168 DHH participants completed an online survey and 15 participated in semi-structured interviews. Additionally, 5 professional captioners participated in semi-structured interviews.<\/jats:p><\/jats:sec><jats:sec><jats:title>Results and discussion<\/jats:title><jats:p>We offer systematic insights into the current challenges related to NSI captioning faced by DHH users and professional captioners, trends in recent NSI captioning research, as well as opportunities for future work that enhance user agency, utilize integrated research methodologies, and broaden community involvement.<\/jats:p><\/jats:sec>","DOI":"10.3389\/fcomp.2025.1575176","type":"journal-article","created":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T17:29:24Z","timestamp":1750354164000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["\u201cChoices? That's the dream\u201d: challenges and opportunities in non-speech information closed-captioning"],"prefix":"10.3389","volume":"7","author":[{"given":"Lloyd","family":"May","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Michael","family":"Clemens","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Khang","family":"Dang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Keita","family":"Ohshiro","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sripathi","family":"Sridhar","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Pauline","family":"Wee","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Magdalena","family":"Fuentes","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sooyeon","family":"Lee","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mark","family":"Cartwright","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1965","published-online":{"date-parts":[[2025,6,19]]},"reference":[{"key":"B1","unstructured":"\u201c2017 state of captioning report,\u201d\n          \n          \n          Technical Report\n          \n          2017"},{"key":"B2","article-title":"\u201c2023 state of captioning report,\u201d","year":"2023","journal-title":"Technical Report"},{"key":"B3","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3517428.3544808","article-title":"\u201cBeyond subtitles: captioning and visualizing non-speech sounds to improve accessibility of user-generated videos,\u201d","volume-title":"Proceedings of the 24th International ACM SIGACCESS Conference on Computers and Accessibility","author":"Alonzo","year":"2022"},{"key":"B4","first-page":"50","article-title":"\u201cCustomization of closed captions via large language models,\u201d","volume-title":"International Conference on Computers Helping People with Special","author":"Arroyo Chavez","year":"2024"},{"key":"B5","author":"Austin","year":"1984","journal-title":"Hearing-Impaired Viewers of Prime-Time Television"},{"key":"B6","first-page":"1","article-title":"\u201cProduction and distribution workflow for closed captioning,\u201d","volume-title":"2010 International Conference on Distributed Frameworks for Multimedia Applications","author":"Barbero","year":"2010"},{"key":"B7","doi-asserted-by":"crossref","first-page":"155","DOI":"10.1145\/3132525.3132541","article-title":"\u201cDeaf and hard-of-hearing perspectives on imperfect automatic speech recognition for captioning one-on-one meetings,\u201d","volume-title":"Proceedings of the 19th International ACM SIGACCESS Conference on Computers and Accessibility","author":"Berke","year":"2017"},{"key":"B8","first-page":"1","volume-title":"Web Content Accessibility Guidelines (WCAG) 2.0","author":"Caldwell","year":"2008"},{"key":"B9","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/1471-2288-13-37","article-title":"\u201cbest fit\u201d framework synthesis: refining the method","volume":"13","author":"Carroll","year":"2013","journal-title":"BMC Med. Res. Methodol"},{"key":"B10","doi-asserted-by":"crossref","first-page":"25","DOI":"10.1007\/978-1-84800-050-6_3","article-title":"\u201cHearing impairments,\u201d","volume-title":"Web Accessibility","author":"Cavender","year":"2008"},{"key":"B11","first-page":"1","article-title":"\u201cA way for deaf and hard of hearing people to enjoy music by exploring and customizing cross-modal music concepts,\u201d","volume-title":"Proceedings of the CHI Conference on Human Factors in Computing Systems","author":"Choi","year":"2024"},{"key":"B12","doi-asserted-by":"publisher","first-page":"2819","DOI":"10.1109\/ACCESS.2020.3047377","article-title":"Vr360 subtitling: Requirements, technology and user experience","volume":"9","author":"Climent","year":"2021","journal-title":"IEEE Access"},{"key":"B13","doi-asserted-by":"publisher","first-page":"6","DOI":"10.1109\/TAFFC.2022.3174721","article-title":"Hidden bawls, whispers, and yelps: Can text convey the sound of speech, beyond words?","volume":"14","author":"de Lacerda Pataca","year":"2023","journal-title":"IEEE Trans. Affect. Comput"},{"key":"B14","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3613904.3642258","article-title":"\u201cCaption royale: exploring the design space of affective captions from the perspective of deaf and hard-of-hearing individuals,\u201d","volume-title":"Proceedings of the CHI Conference on Human Factors in Computing Systems","author":"de Lacerda Pataca","year":"2024"},{"key":"B15","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3544548.3581511","article-title":"\u201cVisualization of speech prosody and emotion in captions: accessibility for deaf and hard-of-hearing users,\u201d","volume-title":"Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems","author":"de Lacerda Pataca","year":"2023"},{"key":"B16","doi-asserted-by":"crossref","DOI":"10.1353\/book.3337","volume-title":"Closed Captioning: Subtitling, Stenography, and the Digital Convergence of Text With Television","author":"Downey","year":"2008"},{"key":"B17","doi-asserted-by":"crossref","first-page":"1","DOI":"10.4324\/9781003183624-1","article-title":"\u201cAural diversity: general introduction,\u201d","volume-title":"Aural Diversity","author":"Drever","year":"2022"},{"key":"B18","unstructured":"Evans\n              M. K.\n            \n          \n          Mountainview, CA\n          YouTube\n          Here's How Automatic Captions Earned Their Nickname\n          \n          2019"},{"key":"B19","article-title":"\u201cEmotive captioning and access to television,\u201d","volume-title":"AMCIS 2005 Proceedings","author":"Fels","year":"2005"},{"key":"B20","doi-asserted-by":"crossref","first-page":"536","DOI":"10.1353\/aad.2012.1342","article-title":"Closed-captioned television viewing preferences","volume":"156","author":"Fitzgerald","year":"1981","journal-title":"Am Ann Deaf."},{"key":"B21","volume-title":"Acoustics \u2014 Normal Equal-Loudness-Level Contours","author":"for Standardization","year":"2023"},{"key":"B22","doi-asserted-by":"crossref","first-page":"571","DOI":"10.1177\/154193120805200613","article-title":"\u201c\u201cthanks for pointing that out.\u201d making sarcasm accessible for all,\u201d","author":"Fourney","year":"2008","journal-title":"Proceedings of the Human Factors and Ergonomics Society Annual Meeting"},{"key":"B23","first-page":"1","article-title":"\u201cAdaptive subtitles: preferences and trade-offs in real-time media adaption,\u201d","author":"Gorman","year":"2021","journal-title":"Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems"},{"key":"B24","unstructured":"Harrenstien\n              K.\n            \n          \n          Mountainview, CA\n          Google LLC\n          Automatic Captions in Youtube\n          \n          2009"},{"key":"B25","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/2513383.2513413","article-title":"\u201cCrowd caption correction (CCC),\u201d","volume-title":"Proceedings of the 15th International ACM SIGACCESS Conference on Computers and Accessibility","author":"Harrington","year":"2013"},{"key":"B26","unstructured":"Chicago, IL\n          Hearing Loss Association of America\n          Hearing Loss Facts and Statistics\n          \n          2023"},{"key":"B27","author":"Henry","year":"2022","journal-title":"Captions\/Subtitles"},{"key":"B28","doi-asserted-by":"publisher","first-page":"2","DOI":"10.7790\/tja.v63i2.406","article-title":"Deaf people's experiences, attitudes and requirements of contextual subtitles: a two-country survey","volume":"63","author":"Hersh","year":"2013","journal-title":"Telecommun. J. Austral"},{"key":"B29","volume-title":"Introduction to American Deaf Culture","author":"Holcomb","year":"2013"},{"key":"B30","doi-asserted-by":"crossref","first-page":"80","DOI":"10.1145\/3462244.3479946","article-title":"\u201cTowards sound accessibility in virtual reality,\u201d","volume-title":"Proceedings of the 2021 International Conference on Multimodal Interaction","author":"Jain","year":"2021"},{"key":"B31","doi-asserted-by":"publisher","DOI":"10.1353\/aad.2012.0073","author":"Jensema","year":"1998","journal-title":"Viewer Reaction to Different Television Captioning Speeds"},{"key":"B32","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/IMCOM60618.2024.10418268","article-title":"\u201cEmotional subtitles through speech in films: a case study,\u201d","volume-title":"2024 18th International Conference on Ubiquitous Information Management and Communication (IMCOM)","author":"Jeon","year":"2024"},{"key":"B33","first-page":"1","article-title":"\u201cComparing the impact of professional and automatic closed captions on video-watching experience,\u201d","volume-title":"Extended Abstracts of the 2023 CHI Conference on Human Factors in Computing Systems","author":"Kim","year":""},{"key":"B34","doi-asserted-by":"crossref","DOI":"10.1145\/3544548.3581130","article-title":"\u201cVisible nuances: a caption system to visualize paralinguistic speech cues for deaf and hard-of-hearing individuals,\u201d","volume-title":"Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, CHI '23","author":"Kim","year":""},{"key":"B35","first-page":"185","article-title":"\u201cEnhancing caption accessibility through simultaneous multimodal information: visual-tactile captions,\u201d","author":"Kushalnagar","year":"2014","journal-title":"Proceedings of the 16th International ACM SIGACCESS Conference on Computers"},{"key":"B36","first-page":"1","article-title":"\u201cLegion scribe: Real-time captioning by the non-experts,\u201d","volume-title":"Proceedings of the 10th International Cross-Disciplinary Conference on Web","author":"Lasecki","year":"2013"},{"key":"B37","doi-asserted-by":"publisher","first-page":"11","DOI":"10.1145\/1281329.1281344","article-title":"Emotive captioning","volume":"5","author":"Lee","year":"2007","journal-title":"Comp. Entertain"},{"key":"B38","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3512922","article-title":"An exploration of captioning practices and challenges of individual content creators on youtube for people with hearing impairments","volume":"6","author":"Li","year":"2022","journal-title":"Proc. ACM Human-Comp. Inter"},{"key":"B39","first-page":"1","article-title":"\u201cCrossA11y: identifying video accessibility issues via cross-modal grounding,\u201d","volume-title":"Proceedings of the 35th Annual ACM Symposium on User Interface Software and Technology, UIST '22","author":"Liu","year":"2022"},{"key":"B40","doi-asserted-by":"publisher","first-page":"209","DOI":"10.1162\/leon_a_02490","article-title":"Experimental modalities: Crip representation and access with electronic arts intermix","volume":"57","author":"Martin","year":"2024","journal-title":"Leonardo"},{"key":"B41","first-page":"1","article-title":"\u201cUnspoken sound: identifying trends in non-speech audio captioning on youtube,\u201d","author":"May","year":"2024","journal-title":"Proceedings of the CHI Conference on Human Factors in Computing Systems"},{"key":"B42","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3597638.3608398","article-title":"\u201cEnhancing non-speech information communicated in closed captioning through critical design,\u201d","volume-title":"Proceedings of the 25th International ACM SIGACCESS Conference on Computers and Accessibility, ASSETS '23","author":"May","year":"2023"},{"key":"B43","first-page":"1","author":"McDonnell","year":"2024"},{"key":"B44","doi-asserted-by":"crossref","first-page":"270","DOI":"10.1145\/3064857.3079159","article-title":"\u201cCymasense: a real-time 3D cymatics-based sound visualisation tool,\u201d","volume-title":"Proceedings of the 2017 ACM Conference Companion Publication on Designing Interactive Systems","author":"McGowan","year":"2017"},{"key":"B45","doi-asserted-by":"crossref","DOI":"10.1186\/s13636-022-00259-2","article-title":"\u201cAutomated audio captioning: an overview of recent progress and new challenges,\u201d","volume-title":"EURASIP Journal on Audio, Speech, and Music Processing","author":"Mei","year":"2022"},{"key":"B46","first-page":"125","article-title":"\u201cCaption UI\/UX - display emotive and paralinguistic information in captions,\u201d","author":"Mendis","year":"2022","journal-title":"The Journal on Technology and Persons with Disabilities"},{"key":"B47","volume-title":"An Introduction to the Psychology of Hearing","author":"Moore","year":"2012"},{"key":"B48","doi-asserted-by":"crossref","first-page":"951","DOI":"10.1109\/TIC-STH.2009.5444362","article-title":"\u201cSeeing the music can animated lyrics provide access to the emotional content in music for people who are deaf or hard of hearing?,\u201d","volume-title":"2009 IEEE Toronto International Conference Science and Technology for Humanity (TIC-STH)","author":"Mori","year":"2009"},{"key":"B49","first-page":"201","article-title":"\u201cText alignment for real-time crowd captioning,\u201d","volume-title":"Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Naim","year":"2013"},{"key":"B50","doi-asserted-by":"publisher","first-page":"n71","DOI":"10.31222\/osf.io\/v7gm2","article-title":"The prisma 2020 statement: an updated guideline for reporting systematic reviews","volume":"372","author":"Page","year":"2021","journal-title":"BMJ"},{"key":"B51","first-page":"28492","article-title":"\u201cRobust speech recognition via large-scale weak supervision,\u201d","volume-title":"International Conference on Machine Learning","author":"Radford","year":"2023"},{"key":"B52","first-page":"24","volume-title":"Expressing Emotions Using Animated Text Captions, volume 4061 of Lecture Notes in Computer Science","author":"Rashid","year":"2006"},{"key":"B53","doi-asserted-by":"publisher","first-page":"505","DOI":"10.1080\/10447310802142342","article-title":"Dancing with words: Using animated text for captioning","volume":"24","author":"Rashid","year":"2008","journal-title":"Intl. J. Human-Comp. Interact"},{"key":"B54","doi-asserted-by":"crossref","DOI":"10.1145\/3685266","article-title":"\u201cAn umbrella review of reporting quality in chi systematic reviews: guiding questions and best practices for HCI,\u201d","volume-title":"ACM Transactions on Computer-Human Interaction","author":"Rogers","year":"2024"},{"key":"B55","year":"2010","journal-title":"Final Fantasy XIV: A Realm Reborn"},{"key":"B56","author":"Tripp","year":"2023","journal-title":"Coda Identity-Why Our Stories Are Important: A Qualitative Look at the Personal Narratives of Adult Hearing Children of Deaf Adults"},{"key":"B57","first-page":"233","article-title":"Making sound accessible: The labelling of soundeffects in subtitling for the deaf and hard-ofhearing","volume":"17","author":"Tsaousi","year":"2015","journal-title":"Hermeneus"},{"key":"B58","doi-asserted-by":"publisher","first-page":"207","DOI":"10.1080\/09544820903310691","article-title":"The rogue poster-children of universal design: closed captioning and audio description","volume":"21","author":"Udo","year":"2010","journal-title":"J. Eng. Design"},{"key":"B59","unstructured":"Suitland-Silver Hill, MD\n          United States Census Bureau\n          American Community Survey\n          \n          2021"},{"key":"B60","unstructured":"Washington, DC\n          United States Congress\n          Television Decoder Circuitry Act of 1990\n          \n          1990"},{"key":"B61","year":"1996","journal-title":"Telecommunications Act of 1996"},{"key":"B62","unstructured":"47 cfr 79.4 - Closed Captioning of Video Programming Delivered Using Internet Protocol\n          \n          2012"},{"key":"B63","first-page":"916","volume-title":"Using Avatars for Improving Speaker Identification in Captioning, volume 5727 of Lecture Notes in Computer Science","author":"Vy","year":"2009"},{"key":"B64","first-page":"247","volume-title":"Using Placement and Name for Speaker Identification in Captioning, volume 6179 of Lecture Notes in Computer Science","author":"Vy","year":"2010"},{"key":"B65","doi-asserted-by":"publisher","first-page":"2","DOI":"10.7790\/tja.v61i2.209","article-title":"Enhanced captioning - speaker identification: text vs. images","volume":"61","author":"Vy","year":"2011","journal-title":"Telecommunic. J. Austral"},{"key":"B66","first-page":"609","volume-title":"EnACT: A Software Tool for Creating Animated Text Captions, volume 5105 of Lecture Notes in Computer Science","author":"Vy","year":"2008"},{"key":"B67","doi-asserted-by":"crossref","first-page":"683","DOI":"10.1007\/11788713_100","article-title":"\u201cCaptioning for deaf and hard of hearing people by editing automatic speech recognition in real time,\u201d","author":"Wald","year":"2006","journal-title":"Computers Helping People with Special Needs"},{"key":"B68","doi-asserted-by":"publisher","first-page":"418","DOI":"10.1109\/TMM.2016.2613641","article-title":"Visualizing video sounds with sound word animation to enrich user experience","volume":"19","author":"Wang","year":"2016","journal-title":"IEEE Trans. Multimedia"},{"key":"B69","doi-asserted-by":"publisher","first-page":"95","DOI":"10.1109\/TASLP.2023.3321968","article-title":"Beyond the Status Quo: a contemporary survey of advances and challenges in audio captioning","volume":"32","author":"Xu","year":"2024","journal-title":"IEEE Trans. Multimedia"},{"key":"B70","doi-asserted-by":"publisher","first-page":"33","DOI":"10.18061\/dsq.v31i3.1667","article-title":"Which sounds are significant? towards a rhetoric of closed captioning","volume":"31","author":"Zdenek","year":"2011","journal-title":"Disab. Stud. Quart"},{"key":"B71","doi-asserted-by":"crossref","DOI":"10.7208\/chicago\/9780226312811.001.0001","article-title":"\u201cReading sounds,\u201d","volume-title":"Reading Sounds","author":"Zdenek","year":"2015"},{"key":"B72","author":"Zdenek","year":"2018","journal-title":"Logocentrism: The Tendency to Privilege Speech Over Non-Speech in Closed Captioning"}],"container-title":["Frontiers in Computer Science"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fcomp.2025.1575176\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T17:29:40Z","timestamp":1750354180000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fcomp.2025.1575176\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,19]]},"references-count":72,"alternative-id":["10.3389\/fcomp.2025.1575176"],"URL":"https:\/\/doi.org\/10.3389\/fcomp.2025.1575176","relation":{},"ISSN":["2624-9898"],"issn-type":[{"value":"2624-9898","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,6,19]]},"article-number":"1575176"}}