{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:20:34Z","timestamp":1750220434837,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":25,"publisher":"ACM","license":[{"start":{"date-parts":[[2020,10,12]],"date-time":"2020-10-12T00:00:00Z","timestamp":1602460800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2020,10,12]]},"DOI":"10.1145\/3394171.3413542","type":"proceedings-article","created":{"date-parts":[[2020,10,12]],"date-time":"2020-10-12T12:27:38Z","timestamp":1602505658000},"page":"367-374","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":7,"title":["SonoSpace"],"prefix":"10.1145","author":[{"given":"Naoki","family":"Kimura","sequence":"first","affiliation":[{"name":"The University of Tokyo, Tokyo, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Keisuke","family":"Shiro","sequence":"additional","affiliation":[{"name":"The University of Tokyo, Tokyo, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yota","family":"Takakura","sequence":"additional","affiliation":[{"name":"Innoqua Inc., Tokyo, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hiromi","family":"Nakamura","sequence":"additional","affiliation":[{"name":"The University of Tokyo, Tokyo, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jun","family":"Rekimoto","sequence":"additional","affiliation":[{"name":"The Univertsity of Tokyo &amp; Sony Computer Science Laboratories, Tokyo, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2020,10,12]]},"reference":[{"key":"e_1_3_2_2_1_1","volume-title":"Proceedings of the International Symposium on Computer Music Multidisciplinary Research (CMMR)","author":"Abe\u00dfer Jakob","year":"2013","unstructured":"Jakob Abe\u00dfer , Johannes Hasselhorn , Christian Dittmar , Andreas Lehmann , and Sascha Grollmisch . 2013 . Automatic quality assessment of vocal and instrumental performances of ninth-grade and tenth-grade pupils . In Proceedings of the International Symposium on Computer Music Multidisciplinary Research (CMMR) , Marseille, France. 975--988. Jakob Abe\u00dfer, Johannes Hasselhorn, Christian Dittmar, Andreas Lehmann, and Sascha Grollmisch. 2013. Automatic quality assessment of vocal and instrumental performances of ninth-grade and tenth-grade pupils. In Proceedings of the International Symposium on Computer Music Multidisciplinary Research (CMMR), Marseille, France. 975--988."},{"key":"e_1_3_2_2_2_1","volume-title":"Proceedings of the International Symposium on Computer Music Multidisciplinary Research (CMMR)","author":"Bozkurt Barics","year":"2017","unstructured":"Barics Bozkurt , Ozan Baysal , and D Yuret . 2017 . A Dataset and Baseline System for Singing Voice Assessment . In Proceedings of the International Symposium on Computer Music Multidisciplinary Research (CMMR) , Matosinhos, Portugal. 25--28. Barics Bozkurt, Ozan Baysal, and D Yuret. 2017. A Dataset and Baseline System for Singing Voice Assessment. In Proceedings of the International Symposium on Computer Music Multidisciplinary Research (CMMR), Matosinhos, Portugal. 25--28."},{"key":"e_1_3_2_2_3_1","unstructured":"Fran\u00e7ois Chollet.. Variational AutoEncoder. https:\/\/keras.io\/examples\/generative\/vae\/. (Accessed on 08\/10\/2020).  Fran\u00e7ois Chollet.. Variational AutoEncoder. https:\/\/keras.io\/examples\/generative\/vae\/. (Accessed on 08\/10\/2020)."},{"key":"e_1_3_2_2_4_1","volume-title":"Syntax-Directed Variational Autoencoder for Structured Data. In 6th International Conference on Learning Representations, ICLR","author":"Dai Hanjun","year":"2018","unstructured":"Hanjun Dai , Yingtao Tian , Bo Dai , Steven Skiena , and Le Song . 2018 . Syntax-Directed Variational Autoencoder for Structured Data. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada. Hanjun Dai, Yingtao Tian, Bo Dai, Steven Skiena, and Le Song. 2018. Syntax-Directed Variational Autoencoder for Structured Data. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada."},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.3389\/fpsyg.2019.00334"},{"key":"e_1_3_2_2_6_1","volume-title":"N.Y.)","author":"Hinton Geoffrey E","year":"2006","unstructured":"Geoffrey E Hinton and Ruslan R Salakhutdinov . 2006. Reducing the Dimensionality of Data with Neural Networks. Science (New York , N.Y.) , Vol. 313 (08 2006 ), 504--7. https:\/\/doi.org\/10.1126\/science.1127647 10.1126\/science.1127647 Geoffrey E Hinton and Ruslan R Salakhutdinov. 2006. Reducing the Dimensionality of Data with Neural Networks. Science (New York, N.Y.), Vol. 313 (08 2006), 504--7. https:\/\/doi.org\/10.1126\/science.1127647"},{"key":"e_1_3_2_2_7_1","volume-title":"Auto-Encoding Variational Bayes. In 2nd International Conference on Learning Representations, ICLR","author":"Diederik","year":"2014","unstructured":"Diederik P. Kingma and Max Welling. 2014 . Auto-Encoding Variational Bayes. In 2nd International Conference on Learning Representations, ICLR 2014 , Banff, AB, Canada,, Yoshua Bengio and Yann LeCun (Eds.). http:\/\/arxiv.org\/abs\/1312.6114 Diederik P. Kingma and Max Welling. 2014. Auto-Encoding Variational Bayes. In 2nd International Conference on Learning Representations, ICLR 2014, Banff, AB, Canada,, Yoshua Bengio and Yann LeCun (Eds.). http:\/\/arxiv.org\/abs\/1312.6114"},{"key":"e_1_3_2_2_8_1","volume-title":"Proceedings of the 12th International Society for Music Information Retrieval Conference, ISMIR 2011","author":"Knight Trevor","year":"2011","unstructured":"Trevor Knight , Finn Upham , and Ichiro Fujinaga . 2011 . The potential for automatic assessment of trumpet tone quality . In Proceedings of the 12th International Society for Music Information Retrieval Conference, ISMIR 2011 , Miami, Florida, USA, October 24-28,, Anssi Klapuri and Colby Leider (Eds.). University of Miami, 573--578. Trevor Knight, Finn Upham, and Ichiro Fujinaga. 2011. The potential for automatic assessment of trumpet tone quality. In Proceedings of the 12th International Society for Music Information Retrieval Conference, ISMIR 2011, Miami, Florida, USA, October 24-28,, Anssi Klapuri and Colby Leider (Eds.). University of Miami, 573--578."},{"key":"e_1_3_2_2_9_1","volume-title":"ATT Labs [Online]","volume":"2","author":"LeCun Yann","year":"2010","unstructured":"Yann LeCun , Corinna Cortes , and CJ Burges . 2010 . MNIST handwritten digit database . ATT Labs [Online] , Vol. 2 (2010). http:\/\/yann.lecun.com\/exdb\/mnist Yann LeCun, Corinna Cortes, and CJ Burges. 2010. MNIST handwritten digit database. ATT Labs [Online], Vol. 2 (2010). http:\/\/yann.lecun.com\/exdb\/mnist"},{"volume-title":"Proceedings of the 16th International Society for Music Information Retrieval Conference, ISMIR 2015, M\u00e1 laga, Spain, October, Meinard M\u00fc ller and Frans Wiering (Eds.). 809--815","author":"Li Pei-Ching","key":"e_1_3_2_2_10_1","unstructured":"Pei-Ching Li , Li Su , Yi-Hsuan Yang , and Alvin W. Y. Su . 2015. Analysis of Expressive Musical Terms in Violin Using Score-Informed and Expression-Based Audio Features . In Proceedings of the 16th International Society for Music Information Retrieval Conference, ISMIR 2015, M\u00e1 laga, Spain, October, Meinard M\u00fc ller and Frans Wiering (Eds.). 809--815 . Pei-Ching Li, Li Su, Yi-Hsuan Yang, and Alvin W. Y. Su. 2015. Analysis of Expressive Musical Terms in Violin Using Score-Informed and Expression-Based Audio Features. In Proceedings of the 16th International Society for Music Information Retrieval Conference, ISMIR 2015, M\u00e1 laga, Spain, October, Meinard M\u00fc ller and Frans Wiering (Eds.). 809--815."},{"key":"e_1_3_2_2_11_1","unstructured":"librosa development team. 2020. LibROSA. https:\/\/librosa.org\/doc\/latest\/index.html. (Accessed on 01\/30\/2020).  librosa development team. 2020. LibROSA. https:\/\/librosa.org\/doc\/latest\/index.html. (Accessed on 01\/30\/2020)."},{"key":"e_1_3_2_2_12_1","volume-title":"Proceedings of the 16th International Society for Music Information Retrieval Conference, ISMIR 2015, M\u00e1 laga, Spain, October 26-30","author":"Luo Yin-Jyun","year":"2015","unstructured":"Yin-Jyun Luo , Li Su , Yi-Hsuan Yang , and Tai-Shih Chi . 2015 . Detection of Common Mistakes in Novice Violin Playing . In Proceedings of the 16th International Society for Music Information Retrieval Conference, ISMIR 2015, M\u00e1 laga, Spain, October 26-30 ,, Meinard M\u00fc ller and Frans Wiering (Eds.). 316--322. Yin-Jyun Luo, Li Su, Yi-Hsuan Yang, and Tai-Shih Chi. 2015. Detection of Common Mistakes in Novice Violin Playing. In Proceedings of the 16th International Society for Music Information Retrieval Conference, ISMIR 2015, M\u00e1 laga, Spain, October 26-30,, Meinard M\u00fc ller and Frans Wiering (Eds.). 316--322."},{"key":"e_1_3_2_2_13_1","unstructured":"makemusic.. SmartMusic | Music Learning Software for Educators & Students. https:\/\/www.smartmusic.com\/. (Accessed on 01\/31\/2020).  makemusic.. SmartMusic | Music Learning Software for Educators & Students. https:\/\/www.smartmusic.com\/. (Accessed on 01\/31\/2020)."},{"key":"e_1_3_2_2_14_1","volume-title":"INTERSPEECH 2006 - ICSLP, Ninth International Conference on Spoken Language Processing","author":"Nakano Tomoyasu","year":"2006","unstructured":"Tomoyasu Nakano , Masataka Goto , and Yuzuru Hiraga . 2006 . An automatic singing skill evaluation method for unknown melodies using pitch interval accuracy and vibrato features . In INTERSPEECH 2006 - ICSLP, Ninth International Conference on Spoken Language Processing , Pittsburgh, PA, USA, September 17--21. ISCA. Tomoyasu Nakano, Masataka Goto, and Yuzuru Hiraga. 2006. An automatic singing skill evaluation method for unknown melodies using pitch interval accuracy and vibrato features. In INTERSPEECH 2006 - ICSLP, Ninth International Conference on Spoken Language Processing, Pittsburgh, PA, USA, September 17--21. ISCA."},{"key":"e_1_3_2_2_15_1","unstructured":"Oy. 2020. Yousician | Learn Guitar Piano Ukulele With The Songs you Love. https:\/\/yousician.com\/. (Accessed on 01\/31\/2020).  Oy. 2020. Yousician | Learn Guitar Piano Ukulele With The Songs you Love. https:\/\/yousician.com\/. (Accessed on 01\/31\/2020)."},{"key":"e_1_3_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.3390\/app8040507"},{"key":"e_1_3_2_2_17_1","unstructured":"Hubert Pham. 2006. PyAudio Documentation. https:\/\/people.csail.mit.edu\/hubert\/pyaudio\/docs\/. (Accessed on 01\/30\/2020).  Hubert Pham. 2006. PyAudio Documentation. https:\/\/people.csail.mit.edu\/hubert\/pyaudio\/docs\/. (Accessed on 01\/30\/2020)."},{"key":"e_1_3_2_2_18_1","unstructured":"Matt Prockup.. Percussion Dataset. http:\/\/www.mattprockup.com\/percussion-dataset (Accessed on 05\/20\/2020).  Matt Prockup.. Percussion Dataset. http:\/\/www.mattprockup.com\/percussion-dataset (Accessed on 05\/20\/2020)."},{"volume-title":"Proceedings of the 14th International Society for Music Information Retrieval Conference, ISMIR 2013","author":"Prockup Matthew","key":"e_1_3_2_2_19_1","unstructured":"Matthew Prockup , Erik M. Schmidt , Jeffrey J. Scott , and Youngmoo E. Kim . 2013. Toward Understanding Expressive Percussion Through Content Based Analysis . In Proceedings of the 14th International Society for Music Information Retrieval Conference, ISMIR 2013 , Curitiba, Brazil, November 4-8,, Alceu de Souza Britto Jr., Fabien Gouyon, and Simon Dixon (Eds.). 143--148. Matthew Prockup, Erik M. Schmidt, Jeffrey J. Scott, and Youngmoo E. Kim. 2013. Toward Understanding Expressive Percussion Through Content Based Analysis. In Proceedings of the 14th International Society for Music Information Retrieval Conference, ISMIR 2013, Curitiba, Brazil, November 4-8,, Alceu de Souza Britto Jr., Fabien Gouyon, and Simon Dixon (Eds.). 143--148."},{"key":"e_1_3_2_2_20_1","volume-title":"Audio Engineering Society Convention 138","author":"Picas Oriol Romani","year":"2015","unstructured":"Oriol Romani Picas , Hector Parra Rodriguez , Dara Dabiri , Hiroshi Tokuda , Wataru Hariya , Koji Oishi , and Xavier Serra . 2015 . A real-time system for measuring sound goodness in instrumental sounds . In Audio Engineering Society Convention 138 . Audio Engineering Society. Oriol Romani Picas, Hector Parra Rodriguez, Dara Dabiri, Hiroshi Tokuda, Wataru Hariya, Koji Oishi, and Xavier Serra. 2015. A real-time system for measuring sound goodness in instrumental sounds. In Audio Engineering Society Convention 138. Audio Engineering Society."},{"key":"e_1_3_2_2_21_1","volume-title":"University of Iowa Studies in the Psychology of Music. Iowa city: University of Iowa","author":"Seashore Carl","year":"1936","unstructured":"Carl Seashore . 1936. University of Iowa Studies in the Psychology of Music. Iowa city: University of Iowa ( 1936 ). Carl Seashore. 1936. University of Iowa Studies in the Psychology of Music. Iowa city: University of Iowa (1936)."},{"key":"e_1_3_2_2_22_1","volume-title":"Vincent Pagel, and Thierry Dutoit.","author":"Tits No\u00e9","year":"2019","unstructured":"No\u00e9 Tits , Fengna Wang , Kevin El Haddad , Vincent Pagel, and Thierry Dutoit. 2019 . Visualization and Interpretation of Latent Spaces for Controlling Expressive Speech Synthesis Through Audio Analysis . 4475--4479. https:\/\/doi.org\/10.21437\/Interspeech.2019-1426 10.21437\/Interspeech.2019-1426 No\u00e9 Tits, Fengna Wang, Kevin El Haddad, Vincent Pagel, and Thierry Dutoit. 2019. Visualization and Interpretation of Latent Spaces for Controlling Expressive Speech Synthesis Through Audio Analysis. 4475--4479. https:\/\/doi.org\/10.21437\/Interspeech.2019-1426"},{"key":"e_1_3_2_2_23_1","volume-title":"Audio Engineering Society Conference: 2017 AES International Conference on Semantic Audio. Audio Engineering Society.","author":"Vidwans Amruta","year":"2017","unstructured":"Amruta Vidwans , Siddharth Gururani , Chih-Wei Wu , Vinod Subramanian , Rupak Vignesh Swaminathan , and Alexander Lerch . 2017 . Objective descriptors for the assessment of student music performances . In Audio Engineering Society Conference: 2017 AES International Conference on Semantic Audio. Audio Engineering Society. Amruta Vidwans, Siddharth Gururani, Chih-Wei Wu, Vinod Subramanian, Rupak Vignesh Swaminathan, and Alexander Lerch. 2017. Objective descriptors for the assessment of student music performances. In Audio Engineering Society Conference: 2017 AES International Conference on Semantic Audio. Audio Engineering Society."},{"key":"e_1_3_2_2_24_1","volume-title":"VASC: dimension reduction and visualization of single-cell RNA-seq data by deep variational autoencoder. Genomics, proteomics & bioinformatics","author":"Wang Dongfang","year":"2018","unstructured":"Dongfang Wang and Jin Gu. 2018. VASC: dimension reduction and visualization of single-cell RNA-seq data by deep variational autoencoder. Genomics, proteomics & bioinformatics , Vol. 16 , 5 ( 2018 ), 320--331. Dongfang Wang and Jin Gu. 2018. VASC: dimension reduction and visualization of single-cell RNA-seq data by deep variational autoencoder. Genomics, proteomics & bioinformatics, Vol. 16, 5 (2018), 320--331."},{"key":"e_1_3_2_2_25_1","volume-title":"Deep Convolutional Variational Autoencoder as a 2D-Visualization Tool for Partial Discharge Source Classification in Hydrogenerators","author":"Zemouri Ryad","year":"2019","unstructured":"Ryad Zemouri , Melanie Levesque , Normand Amyot , Claude Hudon , Olivier Kokoko , and Antoine Tahan . 2019. Deep Convolutional Variational Autoencoder as a 2D-Visualization Tool for Partial Discharge Source Classification in Hydrogenerators . IEEE Access, Vol . PP ( 12 2019 ), 1--1. https:\/\/doi.org\/10.1109\/ACCESS.2019.2962775 10.1109\/ACCESS.2019.2962775 Ryad Zemouri, Melanie Levesque, Normand Amyot, Claude Hudon, Olivier Kokoko, and Antoine Tahan. 2019. Deep Convolutional Variational Autoencoder as a 2D-Visualization Tool for Partial Discharge Source Classification in Hydrogenerators. IEEE Access, Vol. PP (12 2019), 1--1. https:\/\/doi.org\/10.1109\/ACCESS.2019.2962775"}],"event":{"name":"MM '20: The 28th ACM International Conference on Multimedia","sponsor":["SIGMM ACM Special Interest Group on Multimedia"],"location":"Seattle WA USA","acronym":"MM '20"},"container-title":["Proceedings of the 28th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3394171.3413542","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3394171.3413542","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:47:13Z","timestamp":1750193233000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3394171.3413542"}},"subtitle":["Visual Feedback of Timbre with Unsupervised Learning"],"short-title":[],"issued":{"date-parts":[[2020,10,12]]},"references-count":25,"alternative-id":["10.1145\/3394171.3413542","10.1145\/3394171"],"URL":"https:\/\/doi.org\/10.1145\/3394171.3413542","relation":{},"subject":[],"published":{"date-parts":[[2020,10,12]]},"assertion":[{"value":"2020-10-12","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}