{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,22]],"date-time":"2026-04-22T17:53:54Z","timestamp":1776880434539,"version":"3.51.2"},"reference-count":27,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2025,5,19]],"date-time":"2025-05-19T00:00:00Z","timestamp":1747612800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["DMS-1952644"],"award-info":[{"award-number":["DMS-1952644"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["DMS-2151235"],"award-info":[{"award-number":["DMS-2151235"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100005144","name":"Qualcomm","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100005144","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Appl. Math. Stat."],"abstract":"<jats:p>We study a fast local-global window-based attention method to accelerate Informer for long sequence time-series forecasting (LSTF) in a robust manner. While window attention being local is a considerable computational saving, it lacks the ability to capture global token information which is compensated by a subsequent Fourier transform block. Our method, named FWin, does not rely on query sparsity hypothesis and an empirical approximation underlying the ProbSparse attention of Informer. Experiments on univariate and multivariate datasets show that FWin transformers improve the overall prediction accuracies of Informer while accelerating its inference speeds by 1.6 to 2 times. <jats:italic>On strongly non-stationary data (power grid and dengue disease data), FWin outperforms Informer and recent SOTAs thereby demonstrating its superior robustness<\/jats:italic>. We give mathematical definition of FWin attention, and prove its equivalency to the canonical full attention under the block diagonal invertibility (BDI) condition of the attention matrix. The BDI is verified to hold with high probability on benchmark datasets experimentally.<\/jats:p>","DOI":"10.3389\/fams.2025.1600136","type":"journal-article","created":{"date-parts":[[2025,5,19]],"date-time":"2025-05-19T05:15:56Z","timestamp":1747631756000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Fourier-mixed window attention for efficient and robust long sequence time-series forecasting"],"prefix":"10.3389","volume":"11","author":[{"given":"Nhat Thanh","family":"Tran","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jack","family":"Xin","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1965","published-online":{"date-parts":[[2025,5,19]]},"reference":[{"key":"B1","doi-asserted-by":"publisher","first-page":"11106","DOI":"10.1609\/aaai.v35i12.17325","article-title":"Informer: beyond efficient transformer for long sequence time-series forecasting","author":"Zhou","year":"2021","journal-title":"Proceedings of the Association for the Advancement of Artificial Intelligence"},{"key":"B2","article-title":"FEDformer: frequency enhanced decomposed transformer for long-term series forecasting","author":"Zhou","year":"2022","journal-title":"International Conference on Machine Learning"},{"key":"B3","article-title":"Autoformer: decomposition transformers with auto-correlation for long-term series forecasting","author":"Wu","year":"2021","journal-title":"Advances in Neural Information Processing Systems"},{"key":"B4","article-title":"A Time series is worth 64 words: long-term forecasting with transformers","author":"Nie","year":"2023","journal-title":"The Eleventh International Conference on Learning Representations"},{"key":"B5","article-title":"iTransformer: inverted transformers are effective for time series forecasting","author":"Liu","year":"2024","journal-title":"The Twelfth International Conference on Learning Representations"},{"key":"B6","first-page":"30","article-title":"Attention is all you need","author":"Vaswani","year":"2017","journal-title":"Advances in Neural Information Processing Systems"},{"key":"B7","first-page":"4296","article-title":"FNet: Mixing tokens with Fourier transforms","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Lee-Thorp","year":"2022"},{"key":"B8","article-title":"MLP-mixer: An all-MLP architecture for vision","author":"Tolstikhin","year":"2021","journal-title":"Advances in Neural Information Processing Systems"},{"key":"B9","first-page":"980","article-title":"Global filter networks for image classification","author":"Rao","year":"2021","journal-title":"Advances in Neural Information Processing Systems"},{"key":"B10","doi-asserted-by":"publisher","first-page":"12009","DOI":"10.1109\/CVPR52688.2022.00320","article-title":"Swin transformer: hierarchical vision transformer using shifted windows","author":"Liu","year":"2022","journal-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition"},{"key":"B11","doi-asserted-by":"publisher","first-page":"103886","DOI":"10.1016\/j.artint.2023.103886","article-title":"Expanding the prediction capacity in long sequence time-series forecasting","volume":"318","author":"Zhou","year":"2023","journal-title":"Artif Intell"},{"key":"B12","unstructured":"Chow\n              J\n            \n            \n              Rogers\n              G\n            \n          \n          Power System Toolbox\n          \n          2008"},{"key":"B13","doi-asserted-by":"publisher","first-page":"3968","DOI":"10.1109\/ICASSP43922.2022.9747394","article-title":"Glassoformer: a query-sparse transformer for post-fault power grid voltage prediction","author":"Zheng","year":"2022","journal-title":"Proceedings of IEEE International Conference on Acoustics, Speech, and Signal Processing"},{"key":"B14","doi-asserted-by":"crossref","first-page":"160","DOI":"10.1007\/978-3-031-82484-5_12","article-title":"Fwin transformer for dengue prediction under climate and ocean influence","volume-title":"Machine Learning, Optimization, and Data Science","author":"Tran","year":"2025"},{"key":"B15","article-title":"Efficient token mixing for transformers via adaptive Fourier neural operators","author":"Guibas","year":"2022","journal-title":"International Conference on Learning Representations"},{"key":"B16","article-title":"MOAT: alternating mobile convolution and attention brings strong vision models","author":"Yang","year":"2022","journal-title":"arXiv preprint arXiv:221001820"},{"key":"B17","article-title":"Mobilevit: Light-weight, general-purpose, and mobile-friendly vision transformer","author":"Mehta","year":"2022","journal-title":"ICLR"},{"key":"B18","article-title":"ETSformer: exponential smoothing transformers for time-series forecasting","author":"Woo","year":"2022","journal-title":"arXiv:2202.01381"},{"key":"B19","volume-title":"The Hartley Transform","author":"Bracewell","year":"1986"},{"key":"B20","article-title":"Modeling long- and short-term temporal patterns with deep neural networks","author":"Lai","year":"2018","journal-title":"arXiv:1703.07015"},{"key":"B21","article-title":"Generalizable memory-driven transformer for multivariate long sequence time-series forecasting","author":"Li","year":"2022","journal-title":"arXiv preprint arXiv:220707827"},{"key":"B22","doi-asserted-by":"publisher","DOI":"10.1109\/SmartGridComm47815.2020.9302975","article-title":"Machine-learning-based online transient analysis via iterative computation of generator dynamics","author":"Li","year":"2020","journal-title":"Proceedings of the IEEE SmartGridComm"},{"key":"B23","article-title":"Fourierformer: transformer meets generalized Fourier integral theorem","author":"Nguyen","year":"2022","journal-title":"Advances in Neural Information Processing Systems"},{"key":"B24","doi-asserted-by":"publisher","first-page":"141","DOI":"10.1137\/1109020","article-title":"On estimating regression","volume":"9","author":"Nadaraya","year":"1964","journal-title":"Theory Probab Applic"},{"key":"B25","doi-asserted-by":"publisher","first-page":"1065","DOI":"10.1214\/aoms\/1177704472","article-title":"On estimation of a probability density function and mode","volume":"33","author":"Parzen","year":"1962","journal-title":"Ann Mathem Stat"},{"key":"B26","doi-asserted-by":"publisher","first-page":"832","DOI":"10.1214\/aoms\/1177728190","article-title":"Remarks on some nonparametric estimates of a density function","volume":"27","author":"Rosenblatt","year":"1956","journal-title":"Ann Mathem Stat"},{"key":"B27","article-title":"Fourier-mixed window attention: accelerating informer for long sequence time-series forecasting","author":"Tran","year":"2024","journal-title":"arXiv:2307.00493"}],"container-title":["Frontiers in Applied Mathematics and Statistics"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fams.2025.1600136\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,5,19]],"date-time":"2025-05-19T05:16:02Z","timestamp":1747631762000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fams.2025.1600136\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,5,19]]},"references-count":27,"alternative-id":["10.3389\/fams.2025.1600136"],"URL":"https:\/\/doi.org\/10.3389\/fams.2025.1600136","relation":{},"ISSN":["2297-4687"],"issn-type":[{"value":"2297-4687","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,5,19]]},"article-number":"1600136"}}