{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T14:49:47Z","timestamp":1782398987309,"version":"3.54.5"},"reference-count":38,"publisher":"Springer Science and Business Media LLC","issue":"5","license":[{"start":{"date-parts":[[2026,3,24]],"date-time":"2026-03-24T00:00:00Z","timestamp":1774310400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T00:00:00Z","timestamp":1782345600000},"content-version":"vor","delay-in-days":93,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J. King Saud Univ. Comput. Inf. Sci."],"published-print":{"date-parts":[[2026,7]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>This paper introduces spectral attention, which filters the attention score matrix directly in the frequency domain via FFT\/IFFT with learnable, per-head masks. This complements the time-domain view by enabling explicit control over low-, mid-, and high-frequency components of attention patterns. We study nine variants, including an adaptive mechanism that modulates masks from input content. On WikiText-2, Penn Treebank, and WikiText-103, the adaptive spectral variant consistently improves over standard attention, reducing perplexity by 10.7% on WikiText-2 and 15.3% on WikiText-103 in our setup. Analysis shows low-frequency components carry the most useful signal and that learned frequency preferences outperform fixed low\/high\/band-pass filters. These results indicate that frequency-domain processing is an effective complement for autoregressive transformer language modeling in our evaluated settings.<\/jats:p>","DOI":"10.1007\/s44443-026-00599-5","type":"journal-article","created":{"date-parts":[[2026,3,24]],"date-time":"2026-03-24T13:52:13Z","timestamp":1774360333000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Spectral attention for transformers: frequency-domain filtering of attention maps"],"prefix":"10.1007","volume":"38","author":[{"given":"Zhigao","family":"Huang","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Pinghui","family":"Wu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Musheng","family":"Chen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Quanfa","family":"Li","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Miao","family":"Pan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,3,24]]},"reference":[{"key":"599_CR1","unstructured":"Ba JL, Kiros JR, Hinton GE (2016) Layer normalization. arXiv preprint arXiv:1607.06450"},{"key":"599_CR2","unstructured":"Beltagy I, Peters ME, Cohan A (2020) Longformer: The long-document transformer. arXiv preprint arXiv:2004.05150"},{"key":"599_CR3","first-page":"1877","volume":"33","author":"T Brown","year":"2020","unstructured":"Brown T, Mann B, Ryder N, Subbiah M, Kaplan JD, Dhariwal P, Neelakantan A, Shyam P, Girish R, Amodei D et al (2020) Language models are few-shot learners. Adv Neural Inf Process Syst 33:1877\u20131901","journal-title":"Adv Neural Inf Process Syst"},{"key":"599_CR4","doi-asserted-by":"crossref","unstructured":"Chen Y, Dai X, Liu M, Chen D, Yuan L, Liu Z (2020) Dynamic convolution: Attention over convolution kernels. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (2020):11030\u201311039","DOI":"10.1109\/CVPR42600.2020.01104"},{"key":"599_CR5","unstructured":"Child R, Gray S, Radford A, Sutskever I (2019) Generating long sequences with sparse transformers. arXiv preprint arXiv:1904.10509"},{"issue":"240","key":"599_CR6","first-page":"1","volume":"24","author":"A Chowdhery","year":"2022","unstructured":"Chowdhery A, Narang S, Devlin J, Bosma M, Mishra G, Roberts A, Barham P, Chung HW, Sutton C, Gehrmann S et al (2022) Palm: Scaling language modeling with pathways. J Mach Learn Res 24(240):1\u2013113","journal-title":"J Mach Learn Res"},{"key":"599_CR7","doi-asserted-by":"crossref","unstructured":"Clark K, Khandelwal U, Levy O, Manning CD (2019) What does bert look at? an analysis of bert\u2019s attention. arXiv preprint arXiv:1906.04341","DOI":"10.18653\/v1\/W19-4828"},{"key":"599_CR8","unstructured":"Dao T (2023) Flashattention-2: Faster attention with better parallelism and work partitioning (2023). http:\/\/arxiv.org\/abs\/2307.08691arXiv:2307.08691"},{"key":"599_CR9","doi-asserted-by":"publisher","first-page":"16344","DOI":"10.52202\/068431-1189","volume":"35","author":"T Dao","year":"2022","unstructured":"Dao T, Fu D, Ermon S, Rudra A, R\u00e9 C (2022) Flashattention: Fast and memory-efficient exact attention with io-awareness. Adv Neural Inf Process Syst 35:16344\u201316359","journal-title":"Adv Neural Inf Process Syst"},{"key":"599_CR10","doi-asserted-by":"crossref","unstructured":"Dettmers T, Pagnoni A, Holtzman A, Zettlemoyer L (2023) Qlora: Efficient finetuning of quantized llms (2023). arXiv:2305.14314","DOI":"10.52202\/075280-0441"},{"key":"599_CR11","unstructured":"Gu A, Goel K, R\u00e9 C (2021) Efficiently modeling long sequences with structured state spaces. arXiv preprint arXiv:2111.00396"},{"key":"599_CR12","unstructured":"Gu A, Dao T (2023) Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752"},{"key":"599_CR13","doi-asserted-by":"crossref","unstructured":"Hoffmann J, Borgeaud S, Mensch A, Buchatskaya E, Cai T, Rutherford E, Casas DdL, Hendricks LA, Welbl J, Clark A et al (2022) Training compute-optimal large language models. arXiv preprint arXiv:2203.15556","DOI":"10.52202\/068431-2176"},{"key":"599_CR14","doi-asserted-by":"publisher","DOI":"10.1109\/access.2025.3590604","author":"Z Huang","year":"2025","unstructured":"Huang Z, Chen M (2025) Optimizing the learnable rope theta parameter in transformers. IEEE Access. https:\/\/doi.org\/10.1109\/access.2025.3590604","journal-title":"IEEE Access"},{"key":"599_CR15","doi-asserted-by":"publisher","DOI":"10.1016\/j.asoc.2025.113637","author":"Z Huang","year":"2025","unstructured":"Huang Z, Chen M, Zheng S (2025) Transformer spectral optimization: From gradient frequency analysis to adaptive spectral integration. Appl Soft Comput. https:\/\/doi.org\/10.1016\/j.asoc.2025.113637","journal-title":"Appl Soft Comput"},{"key":"599_CR16","doi-asserted-by":"publisher","DOI":"10.3390\/sym17101648","author":"Z Huang","year":"2025","unstructured":"Huang Z, Gong N, Li Q, Wu T, Zheng S, Pan M (2025) Restoring spectral symmetry in gradients: A normalization approach for efficient neural network training. Symmetry. https:\/\/doi.org\/10.3390\/sym17101648","journal-title":"Symmetry"},{"key":"599_CR17","unstructured":"Jiang AQ et al (2023) Mistral 7b (2023). arXiv:2310.06825"},{"key":"599_CR18","unstructured":"Kaplan J, McCandlish S, Henighan T, Brown TB, Chess B, Child R, Gray S, Radford A, Wu J, Amodei D (2020) Scaling laws for neural language models. arXiv preprint arXiv:2001.08361"},{"key":"599_CR19","unstructured":"Katharopoulos A, Vyas A, Pappas N, Fleuret F (2020) Transformers are rnns: Fast autoregressive transformers with linear attention. Int Conf Mach Learn (2020):5156\u20135165"},{"key":"599_CR20","doi-asserted-by":"crossref","unstructured":"Lee-Thorp J, Ainslie J, Eckstein I, Ontanon S (2021) Fnet: Mixing tokens with fourier transforms. arXiv preprint arXiv:2105.03824","DOI":"10.18653\/v1\/2022.naacl-main.319"},{"key":"599_CR21","unstructured":"Liu Y, Zhang X, Zhang X, Wang Q, Huang C, Yang F, Wang Z, Zhao T (2023) Retentive network: A successor to transformer for large language models. arXiv preprint arXiv:2307.08621"},{"key":"599_CR22","unstructured":"Mathieu M, Henaff M, LeCun Y (2013) Fast training of convolutional networks through ffts. arXiv preprint arXiv:1312.5851"},{"key":"599_CR23","unstructured":"Miyato T, Kataoka T, Koyama M, Yoshida Y (2018) Spectral normalization for generative adversarial networks. arXiv preprint arXiv:1802.05957"},{"key":"599_CR24","unstructured":"OpenAI (2024) Gpt-4 technical report. arXiv preprint. arXiv:2303.08774"},{"key":"599_CR25","doi-asserted-by":"crossref","unstructured":"Peng B, Alcaide E, Anthony Q, Albalak A, Arcadinho S, Cao H, Cheng X, Chung M, Grella M, Kiran K GV et al (2023) Rwkv: Reinventing rnns for the transformer era, arXiv preprint arXiv:2305.13048","DOI":"10.18653\/v1\/2023.findings-emnlp.936"},{"key":"599_CR26","unstructured":"Poli M, Massaroli S, Nguyen E, Fu DY, Dao T, Baccus S, Bengio Y, Ermon S, R\u00e9 C (2023) Hyena hierarchy: Towards larger convolutional language models. http:\/\/arxiv.org\/abs\/2302.10866arXiv:2302.10866"},{"key":"599_CR27","unstructured":"Press O, Smith NA, Lewis M (2021) Train short, test long: Attention with linear biases enables input length extrapolation. arXiv preprint arXiv:2108.12409"},{"key":"599_CR28","unstructured":"Qin Z, Sun W, Deng H, Li D, Wei Y, Lv B, Yan J, Kong L, Zhong Y (2022) cosformer: Rethinking softmax in attention. arXiv preprint arXiv:2202.08791"},{"key":"599_CR29","unstructured":"Rahaman N, Baratin A, Arpit D, Draxler F, Lin M, Hamprecht F, Bengio Y, Courville A (2019) On the spectral bias of neural networks. Int Conf Mach Learn (2019):5301\u20135310"},{"key":"599_CR30","first-page":"842","volume":"8","author":"A Rogers","year":"2020","unstructured":"Rogers A, Kovaleva O, Rumshisky A (2020) A primer in bertology: What we know about how bert works, Transactions of the Association for. Comput Linguist 8:842\u2013866","journal-title":"Comput Linguist"},{"key":"599_CR31","unstructured":"Su J, Lu Y, Pan S, Wen B, Liu Y (2021) Roformer: Enhanced transformer with rotary position embedding. arXiv preprint arXiv:2104.09864"},{"key":"599_CR32","doi-asserted-by":"crossref","unstructured":"Sukhbaatar S, Grave E, Bojanowski P, Joulin A (2019) Adaptive attention span in transformers. In: Proceedings of the 57th annual meeting of the association for computational linguistics (2019):331\u2013335","DOI":"10.18653\/v1\/P19-1032"},{"issue":"6","key":"599_CR33","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3530811","volume":"55","author":"Y Tay","year":"2022","unstructured":"Tay Y, Dehghani M, Rao J, Fedus W, Abnar S, Chung HW, Narang S, Yogatama D, Vaswani A, Metzler D (2022) Efficient transformers: A survey. ACM Comput Surv 55(6):1\u201328","journal-title":"ACM Comput Surv"},{"key":"599_CR34","unstructured":"Touvron H et al (2023) Llama 2: Open foundation and fine-tuned chat models (2023). arXiv:2307.09288"},{"key":"599_CR35","unstructured":"Touvron H, Lavril T, Izacard G, Martinet X, Lachaux M-A, Lacroix T, Rozi\u00e8re B, Goyal N, Hambro E, Azhar F et al (2023) Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971"},{"key":"599_CR36","unstructured":"Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser \u0141, Polosukhin I (2017) Attention is all you need. Adv Neural Inf Process Syst 30"},{"key":"599_CR37","unstructured":"Wei J, Tay Y, Bommasani R, Raffel C, Zoph B, Borgeaud S, Yogatama D, Bosma M, Zhou D, Metzler D et al (2022) Emergent abilities of large language models. arXiv preprint arXiv:2206.07682"},{"key":"599_CR38","unstructured":"Zhou T, Ma Z, Wen Q, Wang X, Sun L, Jin R (2022) Fedformer: Frequency enhanced decomposed transformer for long-term series forecasting. Int Conf Mach Learn (2022):27268\u201327286"}],"container-title":["Journal of King Saud University Computer and Information Sciences"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s44443-026-00599-5","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44443-026-00599-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44443-026-00599-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T14:10:55Z","timestamp":1782396655000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s44443-026-00599-5"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,24]]},"references-count":38,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2026,7]]}},"alternative-id":["599"],"URL":"https:\/\/doi.org\/10.1007\/s44443-026-00599-5","relation":{},"ISSN":["1319-1578","2213-1248"],"issn-type":[{"value":"1319-1578","type":"print"},{"value":"2213-1248","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,24]]},"assertion":[{"value":"1 December 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"18 February 2026","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"24 March 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"261"}}