{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,20]],"date-time":"2026-01-20T02:22:02Z","timestamp":1768875722023,"version":"3.49.0"},"reference-count":37,"publisher":"Institute of Electronics, Information and Communications Engineers (IEICE)","issue":"10","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IEICE Trans. Inf. &amp; Syst."],"published-print":{"date-parts":[[2021,10,1]]},"DOI":"10.1587\/transinf.2021edp7016","type":"journal-article","created":{"date-parts":[[2021,9,30]],"date-time":"2021-09-30T22:41:49Z","timestamp":1633041709000},"page":"1734-1748","source":"Crossref","is-referenced-by-count":1,"title":["Diversity-Robust Acoustic Feature Signatures Based on Multiscale Fractal Dimension for Similarity Search of Environmental Sounds"],"prefix":"10.1587","volume":"E104.D","author":[{"given":"Motohiro","family":"SUNOUCHI","sequence":"first","affiliation":[{"name":"Design Department, Sapporo City University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Masaharu","family":"YOSHIOKA","sequence":"additional","affiliation":[{"name":"Graduate School of Information Science and Technology, Hokkaido University"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"532","reference":[{"key":"1","unstructured":"[1] V. Akkermans, F. Font, J. Funollet, B. de Jong, G. Roma, S. Togias, and X. Serra, \u201cFREESOUND 2.0: An Improved Platform for Sharing Audio Clips,\u201d 12th Int. Soc. Music Inf. Retr. Conf., 2011."},{"key":"2","unstructured":"[2] Music Technology Group of Universitat Pompeu Fabra, \u201cThe Freesound Project.,\u201d https:\/\/www.freesound.org\/ (Retrieved 2020-12-20)."},{"key":"3","unstructured":"[3] SoundCloud Limited, \u201cSoundCloud-Hear the world&apos;s sounds,\u201d https:\/\/soundcloud.com\/ (Retrieved 2020-12-20)."},{"key":"4","unstructured":"[4] F. L\u00f3pez, \u201cEnvironmental sound matter,\u201d 1998. [Online]. Available: http:\/\/www.franciscolopez.net\/pdf\/env.pdf [Accessed: 23-Apr-2021]."},{"key":"5","doi-asserted-by":"publisher","unstructured":"[5] D. Michael, \u201cToward a Dark Nature Recording,\u201d Organised Sound, vol.16, no.3, pp.206-210, 2011. 10.1017\/s1355771811000203","DOI":"10.1017\/S1355771811000203"},{"key":"6","unstructured":"[6] T.H. Park, J. Lee, J. You, M.-J. Yoo, and J. Turner, \u201cTowards Soundscape Information Retrieval (SIR),\u201d Proc. ICMC|SMC |2014, Athens, Greece, 2014, pp.1218-1225, 2014."},{"key":"7","doi-asserted-by":"publisher","unstructured":"[7] S. Chachada and C.-C.J. Kuo, \u201cEnvironmental sound recognition: A survey,\u201d APSIPA Trans. Signal Inf. Process., vol.3, 2014. 10.1017\/atsip.2014.12","DOI":"10.1017\/ATSIP.2014.12"},{"key":"8","doi-asserted-by":"publisher","unstructured":"[8] D. Stowell, D. Giannoulis, E. Benetos, M. Lagrange, and M.D. Plumbley, \u201cDetection and Classification of Acoustic Scenes and Events,\u201d IEEE Trans. Multimed., vol.17, no.10, pp.1733-1746, 2015. 10.1109\/tmm.2015.2428998","DOI":"10.1109\/TMM.2015.2428998"},{"key":"9","unstructured":"[9] K. Koutini, F. Henkel, H. Eghbal-zadeh, and G. Widmer, \u201cCP-JKU Submissions to DCASE&apos;20: Low-Complexity Cross-Device Acoustic Scene Classification with RF-Regularized CNNs,\u201d DCASE2020 Challenge, Tech. Rep., 2020."},{"key":"10","doi-asserted-by":"publisher","unstructured":"[10] M. Cowling and R. Sitte, \u201cComparison of techniques for environmental sound recognition,\u201d Pattern Recognit. Lett., vol.24, no.15, pp.2895-2907, 2003. 10.1016\/s0167-8655(03)00147-8","DOI":"10.1016\/S0167-8655(03)00147-8"},{"key":"11","doi-asserted-by":"publisher","unstructured":"[11] S. Chu, S. Narayanan, and C.-C. Kuo, \u201cEnvironmental Sound Recognition With Time-Frequency Audio Features,\u201d IEEE Trans. Audio. Speech. Lang. Processing, vol.17, no.6, pp.1142-1158, 2009. 10.1109\/tasl.2009.2017438","DOI":"10.1109\/TASL.2009.2017438"},{"key":"12","doi-asserted-by":"publisher","unstructured":"[12] R. Mogi and H. Kasai, \u201cNoise-Robust environmental sound classification method based on combination of ICA and MP features,\u201d Artif. Intell. Res., vol.2, no.1, p.107, 2012. 10.5430\/air.v2n1p107","DOI":"10.5430\/air.v2n1p107"},{"key":"13","doi-asserted-by":"crossref","unstructured":"[13] C. Bauge, M. Lagrange, J. Anden, and S. Mallat, \u201cRepresenting environmental sounds using the separable scattering transform,\u201d 2013 IEEE International Conference on Acoustics, Speech and Signal Processing, pp.8667-8671, 2013. 10.1109\/icassp.2013.6639358","DOI":"10.1109\/ICASSP.2013.6639358"},{"key":"14","doi-asserted-by":"crossref","unstructured":"[14] J. Xue, G. Wichern, H. Thornbug, and A. Spanias, \u201cFast query by example of environmental sounds via robust and efficient cluster-based indexing,\u201d 2008 IEEE International Conference on Acoustics, Speech and Signal Processing, pp.5-8, 2008. 10.1109\/icassp.2008.4517532","DOI":"10.1109\/ICASSP.2008.4517532"},{"key":"15","doi-asserted-by":"crossref","unstructured":"[15] G. Roma, J. Janer, S. Kersten, M. Schirosa, P. Herrera, and X. Serra, \u201cEcological Acoustics Perspective for Content-Based Retrieval of Environmental Sounds,\u201d EURASIP J. Audio, Speech, Music Process., vol.2010, pp.1-11, 2010.","DOI":"10.1155\/2010\/960863"},{"key":"16","doi-asserted-by":"crossref","unstructured":"[16] G. Chechik, E. Ie, M. Rehn, S. Bengio, and D. Lyon, \u201cLarge-scale content-based audio retrieval from text queries,\u201d Vancouver, British Columbia, Canada, pp.105-112, 2008. 10.1145\/1460096.1460115","DOI":"10.1145\/1460096.1460115"},{"key":"17","unstructured":"[17] M. Sunouchi and Y. Tanaka, \u201cSimilarity Search of Freesound Environmental Sound Based on Their Enhanced Multiscale Fractal Dimension,\u201d Sound Music Comput. Conf. 2013, SMC 2013, pp.715-721, 2013."},{"key":"18","doi-asserted-by":"crossref","unstructured":"[18] S. Handel, \u201cTimbre perception and auditory object identification,\u201d Hearing, Academic Press, p.468, 1995.","DOI":"10.1016\/B978-012505626-7\/50014-5"},{"key":"19","unstructured":"[19] D. Mitrovi\u0107, M. Zeppelzauer, and H. Eidenberger, \u201cOn Feature Selection in Environmental Sound Recognition,\u201d 51st Int. Symp. ELMAR, no.September, pp.28-30, 2009."},{"key":"20","doi-asserted-by":"crossref","unstructured":"[20] D. Mitrovi\u0107, M. Zeppelzauer, and C. Breiteneder, \u201cFeatures for Content-Based Audio Retrieval,\u201d Adv. Comput., 2010, vol.78, ch.3, pp.71-150, 2010. 10.1016\/s0065-2458(10)78003-7","DOI":"10.1016\/S0065-2458(10)78003-7"},{"key":"21","doi-asserted-by":"crossref","unstructured":"[21] S.G. Mallat and Z. Zhang, \u201cMatching pursuits with time-frequency dictionaries,\u201d IEEE Trans. Signal Process., vol.41, no.12, pp.3397-3415, 1993. 10.1109\/78.258082","DOI":"10.1109\/78.258082"},{"key":"22","doi-asserted-by":"publisher","unstructured":"[22] S. Innami and H. Kasai, \u201cNMF-based environmental sound source separation using time-variant gain features,\u201d Comput. Math. with Appl., vol.64, no.5, pp.1333-1342, 2012. 10.1016\/j.camwa.2012.03.077","DOI":"10.1016\/j.camwa.2012.03.077"},{"key":"23","unstructured":"[23] B. Mandelbrot, \u201cThe Fractal Geometry of Nature,\u201d W.H. Freeman and Company, 1982."},{"key":"24","doi-asserted-by":"publisher","unstructured":"[24] R.F. Voss and J. Clarke, \u201c \u20181\/f noise\u2019 in music and speech,\u201d Nature, vol.258, no.5533, pp.317-318, 1975. 10.1038\/258317a0","DOI":"10.1038\/258317a0"},{"key":"25","doi-asserted-by":"crossref","unstructured":"[25] K.J. Hsu and A.J. Hsu, \u201cFractal Geometry of Music,\u201d Proc. Natl. Acad. Sci., vol.87, no.3, pp.938-941, 1990. 10.1073\/pnas.87.3.938","DOI":"10.1073\/pnas.87.3.938"},{"key":"26","doi-asserted-by":"crossref","unstructured":"[26] P. Maragos, \u201cFractal aspects of speech signals: dimension and interpolation,\u201d [Proceedings] ICASSP 91: 1991 International Conference on Acoustics, Speech, and Signal Processing, pp.417-420, 1991. 10.1109\/icassp.1991.150365","DOI":"10.1109\/ICASSP.1991.150365"},{"key":"27","doi-asserted-by":"publisher","unstructured":"[27] P. Maragos and A. Potamianos, \u201cFractal dimensions of speech sounds: computation and application to automatic speech recognition,\u201d J. Acoust. Soc. Am., vol.105, no.3, pp.1925-1932, 1999. 10.1121\/1.426738","DOI":"10.1121\/1.426738"},{"key":"28","unstructured":"[28] A. Zlatintsi and P. Maragos, \u201cMusical Instruments Signal Analysis and Recognition Using Fractal Features,\u201d 19th Eur. Signal Process. Conf. (EUSIPCO 2011), Barcelona, Spain, no.Eusipco, pp.684-688, 2011."},{"key":"29","doi-asserted-by":"publisher","unstructured":"[29] A. Zlatintsi and P. Maragos, \u201cMultiscale Fractal Analysis of Musical Instrument Signals With Application to Recognition,\u201d IEEE Trans. Audio. Speech. Lang. Processing, vol.21, no.4, pp.737-748, 2013. 10.1109\/tasl.2012.2231073","DOI":"10.1109\/TASL.2012.2231073"},{"key":"30","doi-asserted-by":"crossref","unstructured":"[30] D.W. Scott, \u201cMultivariate Density Estimation: Theory, Practice, and Visualization,\u201d 1st ed. Wiley, 1992. 10.1002\/9780470316849","DOI":"10.1002\/9780470316849"},{"key":"31","unstructured":"[31] D. Keller and B. Truax, \u201cEcologically-based granular synthesis,\u201d Proc. Int. Comput. Music Conf., Ann Arbor, USA, 1998, pp.117-120, 1998."},{"key":"32","doi-asserted-by":"publisher","unstructured":"[32] J. Thorson, T. Weber, and F. Huber, \u201cAuditory behavior of the cricket,\u201d J. Comp. Physiol. A. Neuroethol. Sens. Neural. Behav. Physiol., vol.146, no.3, pp.361-378, 1982. 10.1007\/bf00612706","DOI":"10.1007\/BF00612706"},{"key":"33","unstructured":"[33] SPTK working group, \u201cSpeech Signal Processing Toolkit (SPTK),\u201d http:\/\/sp-tk.sourceforge.net\/, (Retrieved 2020-12-20)."},{"key":"34","doi-asserted-by":"publisher","unstructured":"[34] M.F. Porter, \u201cAn algorithm for suffix stripping,\u201d Progr. Electron. Libr. Inf. Syst., vol.14, no.3, pp.130-137, 1980. 10.1108\/eb046814","DOI":"10.1108\/eb046814"},{"key":"35","doi-asserted-by":"crossref","unstructured":"[35] E. Fonseca, M. Plakal, F. Font, D.P.W. Ellis, X. Favory, J. Pons, and X. Serra, \u201cGeneral-purpose tagging of Freesound audio with AudioSet labels: task description, dataset, and baseline,\u201d in Proceedings of the Detection and Classification of Acoustic Scenes and Events 2018 Workshop (DCASE2018), pp.69-73, 2018.","DOI":"10.33682\/w13e-5v06"},{"key":"36","unstructured":"[36] E. Fonseca, X. Favory, J. Pons, F. Font, M. Plakal, D.P.W. Ellis, and X. Serra, \u201cFSDKaggle2018,\u201d 29-Jan-2019. [Online]. Available: https:\/\/zenodo.org\/record\/2552860. [Accessed: 23-Apr-2021]."},{"key":"37","doi-asserted-by":"crossref","unstructured":"[37] J.F. Gemmeke, D.P.W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R.C. Moore, M. Plakal, and M. Ritter, \u201cAudio Set: An ontology and human-labeled dataset for audio events,\u201d 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.776-780, 2017, doi: 10.1109\/ICASSP.2017.7952261. 10.1109\/icassp.2017.7952261","DOI":"10.1109\/ICASSP.2017.7952261"}],"container-title":["IEICE Transactions on Information and Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/transinf\/E104.D\/10\/E104.D_2021EDP7016\/_pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,9,9]],"date-time":"2024-09-09T03:27:46Z","timestamp":1725852466000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/transinf\/E104.D\/10\/E104.D_2021EDP7016\/_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,10,1]]},"references-count":37,"journal-issue":{"issue":"10","published-print":{"date-parts":[[2021]]}},"URL":"https:\/\/doi.org\/10.1587\/transinf.2021edp7016","relation":{},"ISSN":["0916-8532","1745-1361"],"issn-type":[{"value":"0916-8532","type":"print"},{"value":"1745-1361","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,10,1]]},"article-number":"2021EDP7016"}}