{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,10]],"date-time":"2026-05-10T10:11:50Z","timestamp":1778407910401,"version":"3.51.4"},"reference-count":31,"publisher":"Springer Science and Business Media LLC","issue":"8","license":[{"start":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:00:00Z","timestamp":1750291200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:00:00Z","timestamp":1750291200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100006012","name":"Christian Doppler Forschungsgesellschaft","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100006012","id-type":"DOI","asserted-by":"publisher"}]},{"name":"TU Wien"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Mach Learn"],"published-print":{"date-parts":[[2025,8]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>State Space Models have achieved good performance on long sequence modeling tasks such as raw audio classification. Their definition in continuous time allows for discretization and operation of the network at different sampling rates. However, this property has not yet been utilized to decrease the computational demand on a per-layer basis. We propose a family of hardware-friendly S-Edge models with a layer-wise downsampling approach to adjust the temporal resolution between individual layers. Applying existing methods from linear control theory allows us to analyze state\/memory dynamics and provides an understanding of how and where to downsample. Evaluated on the Google Speech Command dataset, our autoregressive\/causal S-Edge models range from 8\u2013141k parameters at 90\u201395% test accuracy in comparison to a causal S5 model with 208k parameters at 95.8% test accuracy. Using our C++17 header-only implementation on an ARM Cortex-M4F the largest model requires 103\u00a0sec. inference time with 95.19% test accuracy, and the smallest model with 88.01% test accuracy, requires 0.29\u00a0sec. Our solutions cover a design space that spans 17x in model size, 358x in inference latency, and 7.18 percentage points in accuracy.<\/jats:p>","DOI":"10.1007\/s10994-025-06807-z","type":"journal-article","created":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T16:04:29Z","timestamp":1750349069000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Efficient and interpretable raw audio classification with diagonal state space models"],"prefix":"10.1007","volume":"114","author":[{"given":"Matthias","family":"Bittner","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Daniel","family":"Schn\u00f6ll","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Matthias","family":"Wess","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Axel","family":"Jantsch","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,6,19]]},"reference":[{"issue":"2","key":"6807_CR1","doi-asserted-by":"publisher","first-page":"157","DOI":"10.1109\/72.279181","volume":"5","author":"Y Bengio","year":"1994","unstructured":"Bengio, Y., Simard, P., & Frasconi, P. (1994). Learning long-term dependencies with gradient descent is difficult. IEEE Transactions on Neural Networks, 5(2), 157\u2013166. https:\/\/doi.org\/10.1109\/72.279181","journal-title":"IEEE Transactions on Neural Networks"},{"issue":"15","key":"6807_CR2","doi-asserted-by":"publisher","first-page":"247","DOI":"10.1016\/j.ifacol.2024.08.536","volume":"58","author":"F Bonassi","year":"2024","unstructured":"Bonassi, F., Andersson, C., Mattsson, P., & Sch\u00f6n, T. B. (2024). Structured state-space models are deep wiener models. IFAC-PapersOnLine, 58(15), 247\u2013252. https:\/\/doi.org\/10.1016\/j.ifacol.2024.08.536","journal-title":"IFAC-PapersOnLine"},{"key":"6807_CR3","doi-asserted-by":"publisher","unstructured":"Cho, K., Merri\u00ebnboer, B. V., Bahdanau, D., & Bengio, Y. (2014). On the properties of neural machine translation: Encoder-decoder approaches. In Proceedings Eighth Workshop on Syntax, Semantics and Structure in Statistical Translation. https:\/\/doi.org\/10.3115\/v1\/w14-4012","DOI":"10.3115\/v1\/w14-4012"},{"key":"6807_CR4","doi-asserted-by":"publisher","unstructured":"Gerum, C., Frischknecht, A., Hald, T., Bernardo, P. P., Lubeck, K., & Bringmann, O. (2022). Hardware accelerator and neural network co-optimization for ultra-low-power audio processing devices . In 2022 25th Euromicro Conference on Digital System Design (DSD). IEEE Computer Society, Los Alamitos, pp 365\u2013369, https:\/\/doi.org\/10.1109\/DSD57027.2022.00056, https:\/\/doi.ieeecomputersociety.org\/10.1109\/DSD57027.2022.00056","DOI":"10.1109\/DSD57027.2022.00056"},{"key":"6807_CR5","doi-asserted-by":"publisher","first-page":"121902","DOI":"10.1016\/j.eswa.2023.121902","volume":"238","author":"B Ding","year":"2024","unstructured":"Ding, B., Zhang, T., Wang, C., Liu, G., Liang, J., Ruimin, H., Yulin, W., & Guo, D. (2024). Acoustic scene classification: A comprehensive survey. Expert Systems with Applications, 238, 121902. https:\/\/doi.org\/10.1016\/j.eswa.2023.121902","journal-title":"Expert Systems with Applications"},{"issue":"2","key":"6807_CR6","doi-asserted-by":"publisher","first-page":"179","DOI":"10.1016\/0364-0213(90)90002-E","volume":"14","author":"JL Elman","year":"1990","unstructured":"Elman, J. L. (1990). Finding structure in time. Cognitive Science, 14(2), 179\u2013211. https:\/\/doi.org\/10.1016\/0364-0213(90)90002-E","journal-title":"Cognitive Science"},{"key":"6807_CR7","unstructured":"Gu, A., & Dao, T. (2023). Mamba: Linear-time sequence modeling with selective state spaces. CoRR arXiv:abs\/2312.00752"},{"key":"6807_CR8","unstructured":"Gu, A., Goel, K., & R\u00e9, C. (2022). Efficiently modeling long sequences with structured state spaces. In The International Conference on Learning Representations."},{"key":"6807_CR9","unstructured":"Gu, A., Johnson, I., Timalsina, A., Rudra, A., & Re, C. (2023). How to train your HIPPO: State space models with generalized orthogonal basis projections. In International Conference on Learning Representations"},{"key":"6807_CR10","unstructured":"Gu, A., Gupta, A., Goel, K., & R\u00e9, C. (2024). On the parameterization and initialization of diagonal state space models. In Proceedings of the 36th International Conference on Neural Information Processing Systems. Curran Associates Inc., Red Hook, NIPS \u201922."},{"key":"6807_CR11","unstructured":"Gupta, A., Gu, A., & Berant, J. (2024). Diagonal state spaces are as effective as structured state spaces. In Proceedings of the 36th International Conference on Neural Information Processing Systems. Curran Associates Inc., Red Hook, NIPS \u201922."},{"key":"6807_CR12","unstructured":"Hasani, R. M., Lechner, M., Wang, T. -H., Chahine, M., Amini, A., & Rus D. (2023). Liquid structural state-space models. In The Eleventh International Conference on Learning Representations, ICLR 2023."},{"issue":"8","key":"6807_CR13","doi-asserted-by":"publisher","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","volume":"9","author":"S Hochreiter","year":"1997","unstructured":"Hochreiter, S., & Schmidhuber, J. (1997). Long short-term memory. Neural Computation, 9(8), 1735\u20131780. https:\/\/doi.org\/10.1162\/neco.1997.9.8.1735","journal-title":"Neural Computation"},{"issue":"8","key":"6807_CR14","doi-asserted-by":"publisher","first-page":"2554","DOI":"10.1073\/pnas.79.8.2554","volume":"79","author":"JJ Hopfield","year":"1982","unstructured":"Hopfield, J. J. (1982). Neural networks and physical systems with emergent collective computational abilities. Proceedings of the National Academy of Sciences, 79(8), 2554\u20132558. https:\/\/doi.org\/10.1073\/pnas.79.8.2554","journal-title":"Proceedings of the National Academy of Sciences"},{"key":"6807_CR15","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511995569","volume-title":"A First Course in the Numerical Analysis of Differential Equations","author":"A Iserles","year":"2008","unstructured":"Iserles, A. (2008). A First Course in the Numerical Analysis of Differential Equations (2nd ed.). USA: Cambridge University Press.","edition":"2"},{"key":"6807_CR16","unstructured":"Kalchbrenner, N., Espeholt, L., Simonyan, K., van den Oord, A., Graves, A., & Kavukcuoglu, K. (2016). Neural machine translation in linear time. CoRR arXiv:abs\/1610.10099"},{"key":"6807_CR17","doi-asserted-by":"publisher","first-page":"6696","DOI":"10.1109\/ACCESS.2022.3140807","volume":"10","author":"M Mohaimenuzzaman","year":"2022","unstructured":"Mohaimenuzzaman, M., Bergmeir, C., & Meyer, B. (2022). Pruning vs xnor-net: A comprehensive study of deep learning for audio classification on edge-devices. IEEE Access, 10, 6696\u20136707. https:\/\/doi.org\/10.1109\/ACCESS.2022.3140807","journal-title":"IEEE Access"},{"key":"6807_CR18","doi-asserted-by":"publisher","first-page":"109025","DOI":"10.1016\/j.patcog.2022.109025","volume":"133","author":"Md Mohaimenuzzaman","year":"2023","unstructured":"Mohaimenuzzaman, Md., Bergmeir, C., West, I., & Meyer, B. (2023). Environmental sound classification on the edge: A pipeline for deep acoustic networks on extremely resource-constrained devices. Pattern Recognition, 133, 109025. https:\/\/doi.org\/10.1016\/j.patcog.2022.109025","journal-title":"Pattern Recognition"},{"key":"6807_CR19","unstructured":"Orvieto, A., Smith, S.L., Gu, A., Fernando, A., Gulcehre, C., Pascanu, R., & De, S. (2023). Resurrecting recurrent neural networks for long sequences. In Proceedings of the 40th International Conference on Machine Learning, ICML\u201923."},{"key":"6807_CR20","doi-asserted-by":"publisher","first-page":"3423","DOI":"10.1109\/ICASSP43922.2022.9746535","volume":"2022","author":"D Peter","year":"2022","unstructured":"Peter, D., Roth, W., & Pernkopf, F. (2022). End-to-end keyword spotting using neural architecture search and quantization. IEEE ICASSP, 2022, 3423\u20133427. https:\/\/doi.org\/10.1109\/ICASSP43922.2022.9746535","journal-title":"IEEE ICASSP"},{"key":"6807_CR21","doi-asserted-by":"publisher","unstructured":"Scherer, M., Cioflan, C., Magno, M., & Benini, L. (2024). Work in progress: Linear transformers for tinyml. In 2024 Design, Automation & Test in Europe Conference & Exhibition, pp. 1\u20132, https:\/\/doi.org\/10.23919\/DATE58400.2024.10546828","DOI":"10.23919\/DATE58400.2024.10546828"},{"key":"6807_CR22","doi-asserted-by":"publisher","first-page":"21","DOI":"10.9790\/3021-04812125","volume":"4","author":"P Singh","year":"2014","unstructured":"Singh, P., & Rani, P. (2014). An approach to extract feature using MFCC. IOSR Journal of Engineering, 4, 21\u201325. https:\/\/doi.org\/10.9790\/3021-04812125","journal-title":"IOSR Journal of Engineering"},{"key":"6807_CR23","unstructured":"Smith, J.T., Warrington, A., & Linderman, S. (2023). Simplified state space layers for sequence modeling. In The Eleventh International Conference on Learning Representations."},{"key":"6807_CR24","doi-asserted-by":"publisher","first-page":"3210","DOI":"10.1109\/ICPR56361.2022.9956211","volume":"2022","author":"A Smyth","year":"2022","unstructured":"Smyth, A., Lyons, N., Wada, T., Zopf, R., Pandey, A., & Santra, A. (2022). Robust representations for keyword spotting systems. ICPR, 2022, 3210\u20133215. https:\/\/doi.org\/10.1109\/ICPR56361.2022.9956211","journal-title":"ICPR"},{"key":"6807_CR25","unstructured":"Tay, Y., Dehghani, M., Abnar, S., Shen, Y., Bahri, D., Pham, P., Rao, J., Yang, L., Ruder, S., & Metzler, D. (2021). Long range arena : A benchmark for efficient transformers. In International Conference on Learning Representations."},{"key":"6807_CR26","doi-asserted-by":"publisher","first-page":"1100","DOI":"10.1109\/TASLP.2023.3244507","volume":"31","author":"AM Tripathi","year":"2023","unstructured":"Tripathi, A. M., & Pandey, O. J. (2023). Divide and distill: New outlooks on knowledge distillation for environmental sound classification. IEEE\/ACM Transactions on Audio, Speech, and Language Processing, 31, 1100\u20131113. https:\/\/doi.org\/10.1109\/TASLP.2023.3244507","journal-title":"IEEE\/ACM Transactions on Audio, Speech, and Language Processing"},{"key":"6807_CR27","doi-asserted-by":"publisher","unstructured":"Troeng, O., Bernhardsson, B., & Rivetta, C. (2017). Complex-coefficient systems in control. In 2017 American Control Conference (ACC), pp 1721\u20131727, https:\/\/doi.org\/10.23919\/ACC.2017.7963201","DOI":"10.23919\/ACC.2017.7963201"},{"key":"6807_CR28","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, \u0141., & Polosukhin, I. (2017). Attention is all you need. In Proceedings of the 31st International Conference on Neural Information Processing Systems. Curran Associates Inc., Red Hook, NIPS\u201917, p 6000-6010."},{"key":"6807_CR29","unstructured":"Wang, S., & Xue, B. (2023). State-space models with layer-wise nonlinearity are universal approximators with exponential decaying memory. In Thirty-seventh Conference on Neural Information Processing Systems."},{"key":"6807_CR30","unstructured":"Wang, X., Wang, S., Ding, Y., Li, Y., Wu, W., Rong, Y., Kong, W., Huang, J., Li, S., Yang, H., Wang, Z., Jiang, B., Li, C., Wang, Y., Tian, Y., & Tang, J. (2024). State space model for new-generation network alternative to transformers: A survey. CoRR arXiv:abs\/2404.09516"},{"key":"6807_CR31","unstructured":"Warden, P. (2018). Speech commands: A dataset for limited-vocabulary speech recognition. arXiv:1804.03209"}],"container-title":["Machine Learning"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-025-06807-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10994-025-06807-z\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-025-06807-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,6]],"date-time":"2025-09-06T20:24:56Z","timestamp":1757190296000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10994-025-06807-z"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,19]]},"references-count":31,"journal-issue":{"issue":"8","published-print":{"date-parts":[[2025,8]]}},"alternative-id":["6807"],"URL":"https:\/\/doi.org\/10.1007\/s10994-025-06807-z","relation":{},"ISSN":["0885-6125","1573-0565"],"issn-type":[{"value":"0885-6125","type":"print"},{"value":"1573-0565","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,6,19]]},"assertion":[{"value":"9 April 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"9 April 2025","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"26 May 2025","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"19 June 2025","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Not applicable.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical Approval"}},{"value":"The authors declare no conflict of interest.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflicts of Interest"}},{"value":"Training and inference code will be available after acceptance within GitHub repositories. S-Edge:  Cpp-NN:","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Code availability"}},{"value":"The authors of this manuscript consent to its publication.","order":5,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}}],"article-number":"175"}}