{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,18]],"date-time":"2026-07-18T13:52:23Z","timestamp":1784382743057,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":55,"publisher":"ACM","license":[{"start":{"date-parts":[[2019,5,2]],"date-time":"2019-05-02T00:00:00Z","timestamp":1556755200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2019,5,2]]},"DOI":"10.1145\/3290605.3300376","type":"proceedings-article","created":{"date-parts":[[2019,4,29]],"date-time":"2019-04-29T17:04:32Z","timestamp":1556557472000},"page":"1-11","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":117,"title":["SottoVoce"],"prefix":"10.1145","author":[{"given":"Naoki","family":"Kimura","sequence":"first","affiliation":[{"name":"The University of Tokyo, Bunkyo-ku, Tokyo, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Michinari","family":"Kono","sequence":"additional","affiliation":[{"name":"The University of Tokyo, Bunkyo-ku, Tokyo, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jun","family":"Rekimoto","sequence":"additional","affiliation":[{"name":"The University of Tokyo &amp; Sony Computer Science Laboratories, Inc., Bunkyo-ku, Tokyo, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2019,5,2]]},"reference":[{"key":"e_1_3_2_2_1_1","unstructured":"Mart\u00edn Abadi Ashish Agarwal Paul Barham Eugene Brevdo Zhifeng Chen Craig Citro Greg S. Corrado Andy Davis Jeffrey Dean Matthieu Devin Sanjay Ghemawat Ian Goodfellow Andrew Harp Geoffrey Irving Michael Isard Yangqing Jia Rafal Jozefowicz Lukasz Kaiser Manjunath Kudlur Josh Levenberg Dandelion Man\u00e9 Rajat Monga Sherry Moore Derek Murray Chris Olah Mike Schuster Jonathon Shlens Benoit Steiner Ilya Sutskever Kunal Talwar Paul Tucker Vincent Vanhoucke Vijay Vasudevan Fernanda Vi\u00e9gas Oriol Vinyals Pete Warden Martin Wattenberg Martin Wicke Yuan Yu and Xiaoqiang Zheng. 2015. TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems. https:\/\/www.tensorflow.org\/ Software available from tensorflow.org.  Mart\u00edn Abadi Ashish Agarwal Paul Barham Eugene Brevdo Zhifeng Chen Craig Citro Greg S. Corrado Andy Davis Jeffrey Dean Matthieu Devin Sanjay Ghemawat Ian Goodfellow Andrew Harp Geoffrey Irving Michael Isard Yangqing Jia Rafal Jozefowicz Lukasz Kaiser Manjunath Kudlur Josh Levenberg Dandelion Man\u00e9 Rajat Monga Sherry Moore Derek Murray Chris Olah Mike Schuster Jonathon Shlens Benoit Steiner Ilya Sutskever Kunal Talwar Paul Tucker Vincent Vanhoucke Vijay Vasudevan Fernanda Vi\u00e9gas Oriol Vinyals Pete Warden Martin Wattenberg Martin Wicke Yuan Yu and Xiaoqiang Zheng. 2015. TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems. https:\/\/www.tensorflow.org\/ Software available from tensorflow.org."},{"key":"e_1_3_2_2_2_1","unstructured":"Dabi Ahn. 2017. Voice Conversion with Non-Parallel Data. https: \/\/github.com\/andabi\/deep-voice-conversion.  Dabi Ahn. 2017. Voice Conversion with Non-Parallel Data. https: \/\/github.com\/andabi\/deep-voice-conversion."},{"key":"e_1_3_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/3126594.3126649"},{"key":"e_1_3_2_2_4_1","doi-asserted-by":"crossref","unstructured":"F Bocquelet Thomas Hueber Laurent Girin Pierre Badin and B Yvert. 2014. Robust articulatory speech synthesis using deep neural networks for BCI applications. 2288--2292 pages.  F Bocquelet Thomas Hueber Laurent Girin Pierre Badin and B Yvert. 2014. Robust articulatory speech synthesis using deep neural networks for BCI applications. 2288--2292 pages.","DOI":"10.21437\/Interspeech.2014-449"},{"key":"e_1_3_2_2_5_1","volume-title":"Real-Time Control of an Articulatory-Based Speech Synthesizer for Brain Computer Interfaces. PLOS Computational Biology 12, 11 (11","author":"Bocquelet Florent","year":"2016","unstructured":"Florent Bocquelet , Thomas Hueber , Laurent Girin , Christophe Savariaux , and Blaise Yvert . 2016. Real-Time Control of an Articulatory-Based Speech Synthesizer for Brain Computer Interfaces. PLOS Computational Biology 12, 11 (11 2016 ), 1--28. Florent Bocquelet, Thomas Hueber, Laurent Girin, Christophe Savariaux, and Blaise Yvert. 2016. Real-Time Control of an Articulatory-Based Speech Synthesizer for Brain Computer Interfaces. PLOS Computational Biology 12, 11 (11 2016), 1--28."},{"key":"e_1_3_2_2_6_1","unstructured":"Fran\u00e7ois Chollet et al. 2015. Keras. https:\/\/keras.io.  Fran\u00e7ois Chollet et al. 2015. Keras. https:\/\/keras.io."},{"key":"e_1_3_2_2_7_1","unstructured":"Tam\u00e1s G\u00e1bor Csap\u00f3 Tam\u00e1s Gr\u00f3sz G\u00e1bor Gosztolya L\u00e1szl\u00f3 T\u00f3th and Alexandra Mark\u00f3. 2017. DNN-Based Ultrasound-to-Speech Conversion for a Silent Speech Interface. In INTERSPEECH.  Tam\u00e1s G\u00e1bor Csap\u00f3 Tam\u00e1s Gr\u00f3sz G\u00e1bor Gosztolya L\u00e1szl\u00f3 T\u00f3th and Alexandra Mark\u00f3. 2017. DNN-Based Ultrasound-to-Speech Conversion for a Silent Speech Interface. In INTERSPEECH."},{"key":"e_1_3_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.specom.2009.08.002"},{"key":"e_1_3_2_2_9_1","volume-title":"2004 IEEE International Conference on Acoustics, Speech, and Signal Processing","volume":"1","author":"Denby B.","unstructured":"B. Denby and M. Stone . 2004. Speech synthesis from real time ultrasound images of the tongue . In 2004 IEEE International Conference on Acoustics, Speech, and Signal Processing , Vol. 1 . I--685. B. Denby and M. Stone. 2004. Speech synthesis from real time ultrasound images of the tongue. In 2004 IEEE International Conference on Acoustics, Speech, and Signal Processing, Vol. 1. I--685."},{"key":"e_1_3_2_2_10_1","volume-title":"2015 International Joint Conference on Neural Networks (IJCNN). 1--7.","author":"Diener L.","unstructured":"L. Diener , M. Janke , and T. Schultz . 2015. Direct conversion from facial myoelectric signals to speech using Deep Neural Networks . In 2015 International Joint Conference on Neural Networks (IJCNN). 1--7. L. Diener, M. Janke, and T. Schultz. 2015. Direct conversion from facial myoelectric signals to speech using Deep Neural Networks. In 2015 International Joint Conference on Neural Networks (IJCNN). 1--7."},{"key":"e_1_3_2_2_11_1","unstructured":"Roman Gr. Maev (ed.). 2013. Advances in Acoustic Microscopy and High Resolution Imaging: From Principles to Applications. Wiley-VCH.  Roman Gr. Maev (ed.). 2013. Advances in Acoustic Microscopy and High Resolution Imaging: From Principles to Applications. Wiley-VCH."},{"key":"e_1_3_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW.2017.61"},{"key":"e_1_3_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.specom.2017.08.002"},{"key":"e_1_3_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.medengphy.2007.05.003"},{"key":"e_1_3_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/223904.223966"},{"key":"e_1_3_2_2_16_1","volume-title":"CHI 2019","author":"Fiebrink Rebecca","year":"2017","unstructured":"Rebecca Fiebrink . 2017 . Machine Learning as Meta-Instrument: Human-Machine Partnerships Shaping Expressive Instrumental Creation. Springer Singapore, Singapore, 137--151 . CHI 2019 , May 4 --9 , 2019, Glasgow, Scotland Uk Kimura, Kono and Rekimoto Rebecca Fiebrink. 2017. Machine Learning as Meta-Instrument: Human-Machine Partnerships Shaping Expressive Instrumental Creation. Springer Singapore, Singapore, 137--151. CHI 2019, May 4--9, 2019, Glasgow, Scotland Uk Kimura, Kono and Rekimoto"},{"key":"e_1_3_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3242587.3242603"},{"key":"e_1_3_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/2702123.2702591"},{"key":"e_1_3_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.csl.2016.02.002"},{"key":"e_1_3_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASSP.1984.1164317"},{"key":"e_1_3_2_2_21_1","doi-asserted-by":"crossref","unstructured":"Tam\u00e1s Gr\u00f3sz G\u00e1bor Gosztolya L\u00e1szl\u00f3 T\u00f3th Tam\u00e1s Csap\u00f3 and Alexandra Mark\u00f3. 2018. F0 Estimation for DNN-Based Ultrasound Silent Speech Interfaces.  Tam\u00e1s Gr\u00f3sz G\u00e1bor Gosztolya L\u00e1szl\u00f3 T\u00f3th Tam\u00e1s Csap\u00f3 and Alexandra Mark\u00f3. 2018. F0 Estimation for DNN-Based Ultrasound Silent Speech Interfaces.","DOI":"10.1109\/ICASSP.2018.8461732"},{"key":"e_1_3_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.specom.2009.12.001"},{"key":"e_1_3_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.specom.2012.02.001"},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3170427.3188487"},{"key":"e_1_3_2_2_25_1","volume-title":"2007 IEEE International Conference on Acoustics, Speech and Signal Processing - ICASSP '07","volume":"1","author":"Hueber T.","unstructured":"T. Hueber , G. Aversano , G. Cholle , B. Denby , G. Dreyfus , Y. Oussar , P. Roussel , and M. Stone . 2007. Eigentongue Feature Extraction for an Ultrasound-Based Silent Speech Interface . In 2007 IEEE International Conference on Acoustics, Speech and Signal Processing - ICASSP '07 , Vol. 1 . I--1245--I--1248. T. Hueber, G. Aversano, G. Cholle, B. Denby, G. Dreyfus, Y. Oussar, P. Roussel, and M. Stone. 2007. Eigentongue Feature Extraction for an Ultrasound-Based Silent Speech Interface. In 2007 IEEE International Conference on Acoustics, Speech and Signal Processing - ICASSP '07, Vol. 1. I--1245--I--1248."},{"key":"e_1_3_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.specom.2009.11.004"},{"key":"e_1_3_2_2_27_1","doi-asserted-by":"crossref","unstructured":"Thomas Hueber Elie-Laurent Benaroya Bruce Denby and G\u00e9rard Chollet. 2011. Statistical Mapping Between Articulatory and Acoustic Data for an Ultrasound-Based Silent Speech Interface. In INTERSPEECH.  Thomas Hueber Elie-Laurent Benaroya Bruce Denby and G\u00e9rard Chollet. 2011. Statistical Mapping Between Articulatory and Acoustic Data for an Ultrasound-Based Silent Speech Interface. In INTERSPEECH.","DOI":"10.21437\/Interspeech.2011-239"},{"key":"e_1_3_2_2_28_1","volume-title":"Acquisition of ultrasound, video and acoustic speech data for a silentspeech interface application. (01","author":"Hueber Thomas","year":"2008","unstructured":"Thomas Hueber , Gerard Chollet , Bruce Denby , and M Stone . 2008. Acquisition of ultrasound, video and acoustic speech data for a silentspeech interface application. (01 2008 ). Thomas Hueber, Gerard Chollet, Bruce Denby, and M Stone. 2008. Acquisition of ultrasound, video and acoustic speech data for a silentspeech interface application. (01 2008)."},{"key":"e_1_3_2_2_29_1","unstructured":"Google Inc. {n. d.}. Clund Speech-to-Text. https:\/\/cloud.google.com\/ speech-to-text\/.  Google Inc. {n. d.}. Clund Speech-to-Text. https:\/\/cloud.google.com\/ speech-to-text\/."},{"key":"e_1_3_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2016-385"},{"key":"e_1_3_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.specom.2018.02.002"},{"key":"e_1_3_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/HICSS.2005.683"},{"key":"e_1_3_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3172944.3172977"},{"key":"e_1_3_2_2_34_1","volume-title":"Kingma and Jimmy Ba","author":"Diederik","year":"2014","unstructured":"Diederik P. Kingma and Jimmy Ba . 2014 . Adam : A Method for Stochastic Optimization. CoRR abs\/1412.6980 (2014). arXiv:1412.6980 http:\/\/arxiv.org\/abs\/1412.6980 Diederik P. Kingma and Jimmy Ba. 2014. Adam: A Method for Stochastic Optimization. CoRR abs\/1412.6980 (2014). arXiv:1412.6980 http:\/\/arxiv.org\/abs\/1412.6980"},{"key":"e_1_3_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/3064663.3064672"},{"key":"e_1_3_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/3196709.3196784"},{"key":"e_1_3_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/ASRU.2005.1566521"},{"key":"e_1_3_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/765891.765996"},{"key":"e_1_3_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/3025453.3025692"},{"key":"e_1_3_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/3025453.3025807"},{"key":"e_1_3_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/3173574.3173580"},{"key":"e_1_3_2_2_42_1","volume-title":"2003 IEEE International Conference on Acoustics, Speech, and Signal Processing, 2003. Proceedings. (ICASSP '03)","volume":"5","author":"Nakajima Y.","unstructured":"Y. Nakajima , H. Kashioka , K. Shikano , and N. Campbell . 2003. Nonaudible murmur recognition input interface using stethoscopic microphone attached to the skin . In 2003 IEEE International Conference on Acoustics, Speech, and Signal Processing, 2003. Proceedings. (ICASSP '03) ., Vol. 5 . V--708. Y. Nakajima, H. Kashioka, K. Shikano, and N. Campbell. 2003. Nonaudible murmur recognition input interface using stethoscopic microphone attached to the skin. In 2003 IEEE International Conference on Acoustics, Speech, and Signal Processing, 2003. Proceedings. (ICASSP '03)., Vol. 5. V--708."},{"key":"e_1_3_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/3210240.3210322"},{"key":"e_1_3_2_2_44_1","unstructured":"Anne Porbadnigk Marek Wester Jan Calliess and Tanja Schultz. 2009. EEG-based Speech Recognition - Impact of Temporal Effects. In BIOSIGNALS.  Anne Porbadnigk Marek Wester Jan Calliess and Tanja Schultz. 2009. EEG-based Speech Recognition - Impact of Temporal Effects. In BIOSIGNALS."},{"key":"e_1_3_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/3027063.3053246"},{"key":"e_1_3_2_2_46_1","volume-title":"U-Net: Convolutional Networks for Biomedical Image Segmentation. CoRR abs\/1505.04597","author":"Ronneberger Olaf","year":"2015","unstructured":"Olaf Ronneberger , Philipp Fischer , and Thomas Brox . 2015. U-Net: Convolutional Networks for Biomedical Image Segmentation. CoRR abs\/1505.04597 ( 2015 ). arXiv:1505.04597 http:\/\/arxiv.org\/abs\/1505. 04597 Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015. U-Net: Convolutional Networks for Biomedical Image Segmentation. CoRR abs\/1505.04597 (2015). arXiv:1505.04597 http:\/\/arxiv.org\/abs\/1505. 04597"},{"key":"e_1_3_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.3115\/100964.100972"},{"key":"e_1_3_2_2_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/2634317.2634322"},{"key":"e_1_3_2_2_49_1","volume-title":"Computers Helping People with Special Needs","author":"Schultz Tanja","unstructured":"Tanja Schultz . 2010. ICCHP Keynote: Recognizing Silent and Weak Speech Based on Electromyography . In Computers Helping People with Special Needs , Klaus Miesenberger, Joachim Klaus, Wolfgang Zagler, and Arthur Karshmer (Eds.). Springer Berlin Heidelberg , Berlin, Heidelberg , 595--604. Tanja Schultz. 2010. ICCHP Keynote: Recognizing Silent and Weak Speech Based on Electromyography. In Computers Helping People with Special Needs, Klaus Miesenberger, Joachim Klaus, Wolfgang Zagler, and Arthur Karshmer (Eds.). Springer Berlin Heidelberg, Berlin, Heidelberg, 595--604."},{"key":"e_1_3_2_2_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/3196709.3196772"},{"key":"e_1_3_2_2_51_1","doi-asserted-by":"crossref","unstructured":"L\u00e1szl\u00f3 T\u00f3th G\u00e1bor Gosztolya Tam\u00e1s Gr\u00f3sz Alexandra Mark\u00f3 and Tam\u00e1s Csap\u00f3. 2018. Multi-Task Learning of Speech Recognition and Speech Synthesis Parameters for Ultrasound-based Silent Speech Interfaces.  L\u00e1szl\u00f3 T\u00f3th G\u00e1bor Gosztolya Tam\u00e1s Gr\u00f3sz Alexandra Mark\u00f3 and Tam\u00e1s Csap\u00f3. 2018. Multi-Task Learning of Speech Recognition and Speech Synthesis Parameters for Ultrasound-based Silent Speech Interfaces.","DOI":"10.21437\/Interspeech.2018-1078"},{"key":"e_1_3_2_2_52_1","volume-title":"Lipreading with Long Short-Term Memory. CoRR abs\/1601.08188","author":"Wand Michael","year":"2016","unstructured":"Michael Wand , Jan Koutn\u00edk , and J\u00fcrgen Schmidhuber . 2016. Lipreading with Long Short-Term Memory. CoRR abs\/1601.08188 ( 2016 ). arXiv:1601.08188 http:\/\/arxiv.org\/abs\/1601.08188 Michael Wand, Jan Koutn\u00edk, and J\u00fcrgen Schmidhuber. 2016. Lipreading with Long Short-Term Memory. CoRR abs\/1601.08188 (2016). arXiv:1601.08188 http:\/\/arxiv.org\/abs\/1601.08188"},{"key":"e_1_3_2_2_53_1","volume-title":"Interactive Silent Speech Interface Based on Electromagnetic Articulograph. (06","author":"Wang Jun","year":"2014","unstructured":"Jun Wang , Ashok Samal , and Jordan Green . 2014. Preliminary Test of a Real-Time , Interactive Silent Speech Interface Based on Electromagnetic Articulograph. (06 2014 ). Jun Wang, Ashok Samal, and Jordan Green. 2014. Preliminary Test of a Real-Time, Interactive Silent Speech Interface Based on Electromagnetic Articulograph. (06 2014)."},{"key":"e_1_3_2_2_54_1","volume-title":"Saurous","author":"Wang Yuxuan","year":"2017","unstructured":"Yuxuan Wang , R. J. Skerry-Ryan , Daisy Stanton , Yonghui Wu , Ron J. Weiss , Navdeep Jaitly , Zongheng Yang , Ying Xiao , Zhifeng Chen , Samy Bengio , Quoc V. Le , Yannis Agiomyrgiannakis , Rob Clark , and Rif A . Saurous . 2017 . Tacotron : A Fully End-to-End Text-To-Speech Synthesis Model. CoRR abs\/1703.10135 (2017). arXiv:1703.10135 http:\/\/arxiv. org\/abs\/1703.10135 Yuxuan Wang, R. J. Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J. Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, Quoc V. Le, Yannis Agiomyrgiannakis, Rob Clark, and Rif A. Saurous. 2017. Tacotron: A Fully End-to-End Text-To-Speech Synthesis Model. CoRR abs\/1703.10135 (2017). arXiv:1703.10135 http:\/\/arxiv. org\/abs\/1703.10135"},{"key":"e_1_3_2_2_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/2556288.2556981"}],"event":{"name":"CHI '19: CHI Conference on Human Factors in Computing Systems","location":"Glasgow Scotland Uk","acronym":"CHI '19","sponsor":["SIGCHI ACM Special Interest Group on Computer-Human Interaction"]},"container-title":["Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3290605.3300376","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3290605.3300376","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T23:53:21Z","timestamp":1750204401000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3290605.3300376"}},"subtitle":["An Ultrasound Imaging-Based Silent Speech Interaction Using Deep Neural Networks"],"short-title":[],"issued":{"date-parts":[[2019,5,2]]},"references-count":55,"alternative-id":["10.1145\/3290605.3300376","10.1145\/3290605"],"URL":"https:\/\/doi.org\/10.1145\/3290605.3300376","relation":{},"subject":[],"published":{"date-parts":[[2019,5,2]]},"assertion":[{"value":"2019-05-02","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}