{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,3]],"date-time":"2026-02-03T17:46:59Z","timestamp":1770140819280,"version":"3.49.0"},"reference-count":111,"publisher":"Oxford University Press (OUP)","issue":"1","license":[{"start":{"date-parts":[[2026,2,3]],"date-time":"2026-02-03T00:00:00Z","timestamp":1770076800000},"content-version":"vor","delay-in-days":7,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2026,1,27]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Transformers models were originally designed for the processing of textual data. In the last years they have been extended to handle different modalities of data including image and video, audio, tabular, and even multimodal data. Adapting the vanilla Transformer architecture is necessary to optimize the performance for each data type. The vast amount of new architectures that have emerged makes it difficult to detect and understand the differences with the original Transformer. This paper provides an overview of Transformer applications for various input modalities, and recommendations to guide the development and use of models.<\/jats:p>","DOI":"10.1093\/jigpal\/jzaf040","type":"journal-article","created":{"date-parts":[[2025,5,21]],"date-time":"2025-05-21T08:58:27Z","timestamp":1747817907000},"source":"Crossref","is-referenced-by-count":0,"title":["A didactic overview of Transformer applications: model variations and user guidelines"],"prefix":"10.1093","volume":"34","author":[{"given":"MI","family":"Cabrera-Bermejo","sequence":"first","affiliation":[{"name":"Intelligent Systems and Data Mining Group , Department of Computer Science, University of Ja\u00e9n, 23071, Ja\u00e9n, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"M","family":"Germ\u00e1n","sequence":"additional","affiliation":[{"name":"Intelligent Systems and Data Mining Group , Department of Computer Science, University of Ja\u00e9n, 23071, Ja\u00e9n, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"D","family":"de la Rosa","sequence":"additional","affiliation":[{"name":"Intelligent Systems and Data Mining Group , Department of Computer Science, University of Ja\u00e9n, 23071, Ja\u00e9n, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"MJ Del","family":"Jesus","sequence":"additional","affiliation":[{"name":"Intelligent Systems and Data Mining Group , Department of Computer Science, University of Ja\u00e9n, 23071, Ja\u00e9n, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"AJ","family":"Rivera","sequence":"additional","affiliation":[{"name":"Intelligent Systems and Data Mining Group , Department of Computer Science, University of Ja\u00e9n, 23071, Ja\u00e9n, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"D","family":"Elizondo","sequence":"additional","affiliation":[{"name":"Department of Computer Technology , De Montfort University, LE1 9BH, Leicester, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"F","family":"Charte","sequence":"additional","affiliation":[{"name":"Intelligent Systems and Data Mining Group , Department of Computer Science, University of Ja\u00e9n, 23071, Ja\u00e9n, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"MD P\u00e9rez","family":"Godoy","sequence":"additional","affiliation":[{"name":"Intelligent Systems and Data Mining Group , Department of Computer Science, University of Ja\u00e9n, 23071, Ja\u00e9n, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2026,2,3]]},"reference":[{"key":"2026020301213684200_ref1","doi-asserted-by":"publisher","first-page":"12462","DOI":"10.1609\/aaai.v35i14.17478","article-title":"Gate: graph attention transformer encoder for cross-lingual relation and event extraction","volume":"35","author":"Ahmad","year":"2021","journal-title":"Proc AAAI"},{"key":"2026020301213684200_ref2","first-page":"9577","article-title":"Transformer-encoder detector module: using context to improve robustness to adversarial attacks on object detection","author":"Alamri","year":"2021","journal-title":"Proc ICPR"},{"key":"2026020301213684200_ref3","article-title":"Practical and optimal lsh for angular distance","volume-title":"Advances in Neural Information Processing Systems","author":"Andoni","year":"2015"},{"key":"2026020301213684200_ref4","article-title":"A survey on transformers in nlp with focus on efficiency","author":"Ansar","year":"2024"},{"key":"2026020301213684200_ref5","first-page":"6836","article-title":"Vivit: a video vision transformer","author":"Arnab","year":"2021","journal-title":"Proc ICCV"},{"key":"2026020301213684200_ref6","first-page":"1538","article-title":"Simple, scalable adaptation for neural machine translation","author":"Bapna","year":"2019","journal-title":"Proc EMNLP IJCNLP"},{"key":"2026020301213684200_ref7","article-title":"Language models are few-shot learners","volume-title":"Advances in Neural Information Processing Systems","author":"Brown","year":"2020"},{"key":"2026020301213684200_ref8","first-page":"205","article-title":"Swin-unet: Unet-like pure transformer for medical image segmentation","author":"Hu","year":"2023","journal-title":"Proc ECCV"},{"key":"2026020301213684200_ref9","first-page":"213","article-title":"End-to-end object detection with transformers","author":"Carion","year":"2020","journal-title":"Proc ECCV"},{"key":"2026020301213684200_ref10","doi-asserted-by":"publisher","first-page":"445","DOI":"10.1016\/j.inffus.2022.10.030","article-title":"Shape-former: bridging cnn and transformer via shapeconv for multimodal image matching","volume":"91","author":"Chen","year":"2023","journal-title":"Inform Fusion"},{"key":"2026020301213684200_ref11","first-page":"6897","article-title":"Key-sparse transformer for multimodal speech emotion recognition","author":"Chen","year":"2022","journal-title":"Proc ICASSP"},{"key":"2026020301213684200_ref12","first-page":"5904","article-title":"Developing real-time streaming transformer transducer for speech recognition on large-scale dataset","author":"Chen","year":"2021","journal-title":"Proc IEEE ICASSP"},{"key":"2026020301213684200_ref13","doi-asserted-by":"publisher","first-page":"85","DOI":"10.26599\/BDMA.2022.9020017","article-title":"Wtasr: wavelet transformer for automatic speech recognition of indian languages","volume":"6","author":"Choudhary","year":"2023","journal-title":"Big Data Min Anal"},{"key":"2026020301213684200_ref14","first-page":"10575","article-title":"Meshed-memory transformer for image captioning","author":"Cornia","year":"2020","journal-title":"Proc CVPR"},{"key":"2026020301213684200_ref15","first-page":"2026","article-title":"Edited media understanding frames: reasoning about the intent and implications of visual misinformation","author":"Da","year":"2020","journal-title":"Proc ACL IJCNLP"},{"key":"2026020301213684200_ref16","first-page":"6857","article-title":"Dpt-fsnet: dual-path transformer based full-band and sub-band fusion network for speech enhancement","author":"Dang","year":"2022","journal-title":"Proc ICASSP"},{"key":"2026020301213684200_ref17","doi-asserted-by":"publisher","first-page":"549","DOI":"10.1016\/j.neunet.2023.09.039","article-title":"Sttre: a spatio-temporal transformer with relative embeddings for multivariate time series forecasting","volume":"168","author":"Deihim","year":"2023","journal-title":"Neural Netw"},{"key":"2026020301213684200_ref18","first-page":"4171","article-title":"BERT: pre-training of deep bidirectional transformers for language understanding","author":"Devlin","year":"2019","journal-title":"Proc NAACL"},{"key":"2026020301213684200_ref19","first-page":"5884","article-title":"Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition","author":"Dong","year":"2018","journal-title":"Proc IEEE ICASSP"},{"key":"2026020301213684200_ref20","article-title":"An image is worth 16x16 words: transformers for image recognition at scale","author":"Dosovitskiy","year":"2020"},{"key":"2026020301213684200_ref21","doi-asserted-by":"publisher","first-page":"2705","DOI":"10.1109\/JSTARS.2023.3344215","article-title":"Pan-sharpening via multiscale embedding and dual attention transformers","volume":"17","author":"Fan","year":"2024","journal-title":"IEEE J Sel Top Appl Earth Obs Remote Sens"},{"key":"2026020301213684200_ref22","doi-asserted-by":"crossref","first-page":"22","DOI":"10.1007\/s10618-023-00948-2","article-title":"Improving position encoding of transformers for multivariate time series classification","volume":"38","author":"Foumani","year":"2024","journal-title":"Data Min Knowl Discov"},{"key":"2026020301213684200_ref23","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3586074","article-title":"A practical survey on faster and lighter transformers","volume":"55","author":"Fournier","year":"2023","journal-title":"ACM Comput Surv"},{"key":"2026020301213684200_ref24","first-page":"2251","article-title":"Fashionbert: text and image matching with adaptive loss for cross-modal retrieval","author":"Gao","year":"2020","journal-title":"Proc ACM SIGIR"},{"key":"2026020301213684200_ref25","doi-asserted-by":"publisher","first-page":"127651","DOI":"10.1016\/j.neucom.2024.127651","article-title":"Show, tell and rectify: boost image caption generation via an output rectifier","volume":"585","author":"Ge","year":"2024","journal-title":"Neurocomputing"},{"key":"2026020301213684200_ref26","article-title":"Gemini: a family of multimodal foundation models","author":"Google Gemini Team","year":"2023"},{"key":"2026020301213684200_ref27","article-title":"The reversible residual network: backpropagation without storing activations","volume-title":"Advances in Neural Information Processing Systems","author":"Gomez","year":"2017"},{"key":"2026020301213684200_ref28","article-title":"Non-autoregressive neural machine translation","author":"Gu","year":"2018","journal-title":"Proc ICLR"},{"key":"2026020301213684200_ref29","first-page":"956","article-title":"KAT: a knowledge augmented transformer for vision-and-language","author":"Gui","year":"2022","journal-title":"Proc NAACL"},{"key":"2026020301213684200_ref30","first-page":"5036","article-title":"Conformer: convolution-augmented transformer for speech recognition","author":"Gulati","year":"2020","journal-title":"Proc Interspeech"},{"key":"2026020301213684200_ref31","first-page":"2214","article-title":"Learning shared semantic space for speech-to-text translation","author":"Han","year":"2021","journal-title":"ProcACL IJCNLP"},{"key":"2026020301213684200_ref32","doi-asserted-by":"publisher","first-page":"87","DOI":"10.1109\/TPAMI.2022.3152247","article-title":"A survey on vision transformer","volume":"45","author":"Han","year":"2023","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"2026020301213684200_ref33","doi-asserted-by":"crossref","first-page":"12972","DOI":"10.1609\/aaai.v35i14.17534","article-title":"Humor knowledge enriched transformer for understanding multimodal humor","volume":"14B","author":"Hasan","year":"2021","journal-title":"Proc AAAI"},{"key":"2026020301213684200_ref34","first-page":"1748","article-title":"Unetr: transformers for 3d medical image segmentation","author":"Hatamizadeh","year":"2022","journal-title":"Proc IEEE\/CVF WACV"},{"key":"2026020301213684200_ref35","doi-asserted-by":"crossref","first-page":"5931","DOI":"10.1609\/aaai.v35i7.16741","article-title":"Actionbert: leveraging user actions for semantic understanding of user interfaces","volume":"7","author":"He","year":"2021","journal-title":"Proc AAAI"},{"key":"2026020301213684200_ref36","article-title":"Tabpfn: a transformer that solves small tabular classification problems in a second","author":"Hollmann","year":"2022"},{"key":"2026020301213684200_ref37","first-page":"9989","article-title":"Iterative answer prediction with pointer-augmented multimodal transformers for textvqa","author":"Hu","year":"2020","journal-title":"Proc CVPR"},{"key":"2026020301213684200_ref38","article-title":"Music transformer","author":"Huang","year":"2018"},{"key":"2026020301213684200_ref39","first-page":"470","article-title":"Multimodal pretraining for dense video captioning","author":"Huang","year":"2020","journal-title":"Proc AACL"},{"key":"2026020301213684200_ref40","first-page":"2012","article-title":"Reformer-TTS: neural speech synthesis with reformer network","author":"Ihm","year":"2020","journal-title":"Proc Interspeech"},{"key":"2026020301213684200_ref41","article-title":"Samformer: unlocking the potential of transformers in time series forecasting with sharpness-aware minimization and channel-wise attention","author":"Ilbert","year":"2024"},{"key":"2026020301213684200_ref42","first-page":"110393","article-title":"Bts-st: Swin transformer network for segmentation and classification of multimodality breast cancer images","volume":"267","author":"Iqbal","year":"2023","journal-title":"KBS"},{"key":"2026020301213684200_ref43","doi-asserted-by":"publisher","first-page":"122666","DOI":"10.1016\/j.eswa.2023.122666","article-title":"A comprehensive survey on applications of transformers for deep learning tasks","volume":"241","author":"Islam","year":"2024","journal-title":"Exp Syst Appl"},{"key":"2026020301213684200_ref44","doi-asserted-by":"publisher","first-page":"102804","DOI":"10.1016\/j.artmed.2024.102804","article-title":"Prognostic prediction of sepsis patient using transformer with skip connected token for tabular data","volume":"149","author":"Jee-Woo","year":"2024","journal-title":"Artif Intell Med"},{"key":"2026020301213684200_ref45","doi-asserted-by":"publisher","first-page":"1655","DOI":"10.1609\/aaai.v35i2.16258","article-title":"Improving image captioning by leveraging intra- and inter-layer global representation in transformer network","volume":"35","author":"Ji","year":"2021","journal-title":"Proc AAAI"},{"key":"2026020301213684200_ref46","doi-asserted-by":"publisher","DOI":"10.3390\/electronics12061330","article-title":"Low complexity speech enhancement network based on frame-level swin transformer","volume":"12","author":"Jiang","year":"2023","journal-title":"Electronics"},{"key":"2026020301213684200_ref47","first-page":"14745","article-title":"Transgan: two pure transformers can make one strong Gan, and that can scale up","volume":"34","author":"Jiang","year":"2021","journal-title":"Proc NIPS"},{"key":"2026020301213684200_ref48","article-title":"Ammus: a survey of transformer-based pretrained models in natural language processing","author":"Kalyan","year":"2021"},{"key":"2026020301213684200_ref49","doi-asserted-by":"publisher","first-page":"111507","DOI":"10.1016\/j.knosys.2024.111507","article-title":"Transformer-based multivariate time series anomaly detection using inter-variable attention mechanism","volume":"290","author":"Kang","year":"2024","journal-title":"Knowl-Based Syst"},{"key":"2026020301213684200_ref50","doi-asserted-by":"publisher","first-page":"123640","DOI":"10.1016\/j.eswa.2024.123640","article-title":"Video2music: suitable music generation from videos using an affective multimodal transformer model","volume":"249","author":"Kang","year":"2024","journal-title":"Exp Syst Appl"},{"key":"2026020301213684200_ref51","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3505244","article-title":"Transformers in vision: a survey","volume":"54","author":"Khan","year":"2022","journal-title":"ACM Comput Surv"},{"key":"2026020301213684200_ref52","first-page":"6649","article-title":"T-gsa: transformer with gaussian-weighted self-attention for speech enhancement","author":"Kim","year":"2020","journal-title":"Proc IEEE ICASSP"},{"key":"2026020301213684200_ref53","first-page":"9361","article-title":"Squeezeformer: an efficient transformer for automatic speech recognition","volume-title":"Advances in Neural Information Processing Systems","author":"Kim","year":"2022"},{"key":"2026020301213684200_ref54","first-page":"344","article-title":"Albert: a lite bert for self-supervised learning of language representations","author":"Lan","year":"2020","journal-title":"Proc ICLR"},{"key":"2026020301213684200_ref55","article-title":"Bart: denoising sequence-to-sequence pre-training for natural language processing","author":"Lewis","year":"2020"},{"key":"2026020301213684200_ref56","doi-asserted-by":"publisher","first-page":"286","DOI":"10.1609\/aaai.v35i1.16103","article-title":"Two-stream convolution augmented transformer for human activity recognition","volume":"35","author":"Li","year":"2021","journal-title":"Proc AAAI"},{"key":"2026020301213684200_ref57","doi-asserted-by":"publisher","first-page":"471","DOI":"10.1109\/LSP.2024.3358754","article-title":"Esaformer: enhanced self-attention for automatic speech recognition","volume":"31","author":"Li","year":"2024","journal-title":"IEEE Signal Process Lett"},{"key":"2026020301213684200_ref58","article-title":"Visualbert: a simple and performant baseline for vision and language","author":"Li","year":"2019"},{"key":"2026020301213684200_ref59","doi-asserted-by":"publisher","first-page":"6706","DOI":"10.1609\/aaai.v33i01.33016706","article-title":"Neural speech synthesis with transformer network","volume":"33","author":"Li","year":"2019","journal-title":"Proc AAAI"},{"key":"2026020301213684200_ref60","first-page":"2592","article-title":"UNIMO: towards unified-modal understanding and generation via cross-modal contrastive learning","author":"Li","year":"2021","journal-title":"Proc ACL IJCNLP"},{"key":"2026020301213684200_ref61","doi-asserted-by":"publisher","first-page":"371","DOI":"10.1109\/LSP.2024.3353039","article-title":"Transformer-based end-to-end speech translation with rotary position embedding","volume":"31","author":"Li","year":"2024","journal-title":"IEEE Signal Process Lett"},{"key":"2026020301213684200_ref62","first-page":"1293","article-title":"Forecaster: a graph transformer for forecasting spatial and time-dependent data","volume":"325","author":"Li","year":"2020","journal-title":"Front Artif Intell Appl"},{"key":"2026020301213684200_ref63","doi-asserted-by":"publisher","first-page":"13315","DOI":"10.1609\/aaai.v35i15.17572","article-title":"An efficient transformer decoder with compressed sub-layers","volume":"35","author":"Li","year":"2021","journal-title":"Proc AAAI"},{"key":"2026020301213684200_ref64","doi-asserted-by":"publisher","first-page":"111","DOI":"10.1016\/j.aiopen.2022.10.001","article-title":"A survey of transformers","volume":"3","author":"Lin","year":"2022","journal-title":"AI Open"},{"key":"2026020301213684200_ref65","doi-asserted-by":"publisher","first-page":"124616","DOI":"10.1016\/j.eswa.2024.124616","article-title":"Lightweight multimodal cycle-attention transformer towards cancer diagnosis","volume":"255","author":"Liu","year":"2024","journal-title":"Exp Syst Appl"},{"key":"2026020301213684200_ref66","doi-asserted-by":"publisher","first-page":"7478","DOI":"10.1109\/TNNLS.2022.3227717","article-title":"A survey of visual transformers","volume":"35","author":"Liu","year":"2024","journal-title":"IEEE Trans Neural Netw Learn Syst"},{"key":"2026020301213684200_ref67","first-page":"10012","article-title":"Swin transformer: hierarchical vision transformer using shifted windows","author":"Liu","year":"2021","journal-title":"Proc ICCV"},{"key":"2026020301213684200_ref68","doi-asserted-by":"publisher","first-page":"2286","DOI":"10.1609\/aaai.v35i3.16328","article-title":"Dual-level collaborative transformer for image captioning","volume":"35","author":"Luo","year":"2021","journal-title":"Proc AAAI"},{"key":"2026020301213684200_ref69","article-title":"Molecule attention transformer","author":"Maziarka","year":"2020"},{"key":"2026020301213684200_ref70","first-page":"1744","article-title":"UmlsBERT: clinical domain knowledge augmentation of contextual embeddings using the unified medical language system Metathesaurus","author":"Michalopoulos","year":"2021","journal-title":"Proc NAACL"},{"key":"2026020301213684200_ref71","article-title":"Transformers with convolutional context for asr","author":"Mohamed","year":"2019"},{"key":"2026020301213684200_ref72","doi-asserted-by":"publisher","first-page":"10690","DOI":"10.1109\/ACCESS.2024.3354972","article-title":"Lf-transformer: latent factorizer transformer for tabular learning","volume":"12","author":"Na","year":"2024","journal-title":"IEEE Access"},{"key":"2026020301213684200_ref73","article-title":"Gpt-4 technical report","author":"OpenAI","year":"2023"},{"key":"2026020301213684200_ref74","first-page":"4055","article-title":"Image transformer","volume":"80","author":"Parmar","year":"2018","journal-title":"Proc ICML"},{"key":"2026020301213684200_ref75","doi-asserted-by":"publisher","first-page":"453","DOI":"10.1609\/aaai.v35i1.16122","article-title":"Rarebert: transformer architecture for rare disease patient identification using administrative claims","volume":"35","author":"Prakash","year":"2021","journal-title":"Proc AAAI"},{"key":"2026020301213684200_ref76","article-title":"Imagebert: cross-modal pre-training with large-scale weak-supervised image-text data","author":"Qi","year":"2020"},{"key":"2026020301213684200_ref77","article-title":"cosformer: rethinking softmax in attention","author":"Qin","year":"2022"},{"key":"2026020301213684200_ref78","first-page":"9","article-title":"Language models are unsupervised multitask learners","volume":"1","author":"Radford","year":"2019","journal-title":"OpenAI"},{"key":"2026020301213684200_ref79","doi-asserted-by":"publisher","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster R-CNN: towards real-time object detection with region proposal networks","volume":"39","author":"Ren","year":"2017","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"2026020301213684200_ref80","doi-asserted-by":"publisher","first-page":"53","DOI":"10.1162\/tacl_a_00353","article-title":"Efficient content-based sparse attention with routing transformers","volume":"9","author":"Roy","year":"2021","journal-title":"TACL"},{"key":"2026020301213684200_ref81","doi-asserted-by":"publisher","first-page":"12922","DOI":"10.1109\/TPAMI.2023.3243465","article-title":"Video transformers: a survey","volume":"45","author":"Selva","year":"2023","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"2026020301213684200_ref82","article-title":"Saint: improved neural networks for tabular data via row attention and contrastive pre-training","author":"Somepalli","year":"2021"},{"key":"2026020301213684200_ref83","first-page":"4091","article-title":"Attend and diagnose: clinical time series analysis using attention models","author":"Song","year":"2018","journal-title":"Proc AAAI"},{"key":"2026020301213684200_ref84","doi-asserted-by":"publisher","first-page":"103721","DOI":"10.1016\/j.cviu.2023.103721","article-title":"Transformer-based image generation from scene graphs","volume":"233","author":"Sortino","year":"2023","journal-title":"Computer Vision and Image Understanding"},{"key":"2026020301213684200_ref85","article-title":"VL-BERT: pre-training of generic visual-linguistic representations","author":"Su","year":"2019"},{"key":"2026020301213684200_ref86","first-page":"21","article-title":"Attention is all you need in speech separation","author":"Subakan","year":"2021","journal-title":"Proc IEEE ICASSP"},{"key":"2026020301213684200_ref87","first-page":"7463","article-title":"Videobert: a joint model for video and language representation learning","author":"Sun","year":"2019","journal-title":"Proc ICCV"},{"key":"2026020301213684200_ref88","doi-asserted-by":"crossref","first-page":"13860","DOI":"10.1609\/aaai.v35i15.17633","article-title":"Rpbert: a text-image relation propagation-based bert model for multimodal ner","volume":"15","author":"Sun","year":"2021","journal-title":"Proc AAAI"},{"key":"2026020301213684200_ref89","first-page":"908","article-title":"LCD\u2014line clustering and description for place recognition","author":"Taubner","year":"2020","journal-title":"Proc 3DV"},{"key":"2026020301213684200_ref90","article-title":"Llama: open and efficient foundation language models","author":"Touvron","year":"2023"},{"key":"2026020301213684200_ref91","first-page":"5999","article-title":"Attention is all you need","volume":"30","author":"Vaswani","year":"2017","journal-title":"Proc NIPS"},{"key":"2026020301213684200_ref92","first-page":"479","article-title":"Semi-autoregressive neural machine translation","author":"Wang","year":"2018","journal-title":"Proc EMNLP"},{"key":"2026020301213684200_ref93","doi-asserted-by":"crossref","first-page":"5377","DOI":"10.1609\/aaai.v33i01.33015377","article-title":"Non-autoregressive machine translation with auxiliary regularization","author":"Wang","year":"2019","journal-title":"Proc AAAI"},{"key":"2026020301213684200_ref94","article-title":"Transformers in time series: a survey","author":"Wen","year":"2022"},{"key":"2026020301213684200_ref95","article-title":"Transfertransfo: a transfer learning approach for neural network based conversational agents","author":"Wolf","year":"2019"},{"key":"2026020301213684200_ref96","first-page":"38","article-title":"Transformers: state-of-the-art natural language processing","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations","author":"Wolf","year":"2020"},{"key":"2026020301213684200_ref97","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1038\/s41524-023-01016-5","article-title":"Transpolymer: a transformer-based language model for polymer property predictions","volume":"9","author":"Xu","year":"2023","journal-title":"Npj Comput Mater"},{"key":"2026020301213684200_ref98","article-title":"Tener: adapting transformer encoder for named entity recognition","author":"Yan","year":"2019"},{"key":"2026020301213684200_ref99","doi-asserted-by":"publisher","first-page":"14257","DOI":"10.1609\/aaai.v35i16.17677","article-title":"Contrastive triple extraction with generative transformer","volume":"35","author":"Ye","year":"2021","journal-title":"Proc AAAI"},{"key":"2026020301213684200_ref100","article-title":"Fast and accurate reading comprehension by combining self-attention and convolution","author":"Yu","year":"2018","journal-title":"Proc ICLR"},{"key":"2026020301213684200_ref101","doi-asserted-by":"crossref","first-page":"1152","DOI":"10.1007\/s12559-020-09817-2","article-title":"Setransformer: speech enhancement transformer","volume":"14","author":"Yu","year":"2022","journal-title":"Cogn Comput"},{"key":"2026020301213684200_ref102","doi-asserted-by":"publisher","DOI":"10.3389\/fpls.2023.1273029","article-title":"Dic-transformer: interpretation of plant disease classification results using image caption generation technology","volume":"14","author":"Zeng","year":"2024","journal-title":"Front Plant Sci"},{"key":"2026020301213684200_ref103","first-page":"917","article-title":"Token shift transformer for video classification","author":"Zhang","year":"2021","journal-title":"Proc ACM MM"},{"key":"2026020301213684200_ref104","doi-asserted-by":"publisher","first-page":"1141","DOI":"10.1007\/s11263-022-01739-w","article-title":"Vitaev2: vision transformer advanced by exploring inductive bias for image recognition and beyond","volume":"131","author":"Zhang","year":"2023","journal-title":"IJCV"},{"key":"2026020301213684200_ref105","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1109\/TGRS.2023.3297858","article-title":"A lightweight transformer network for hyperspectral image classification","volume":"61","author":"Zhang","year":"2023","journal-title":"IEEE Trans Geosci Remote Sens"},{"key":"2026020301213684200_ref106","doi-asserted-by":"publisher","first-page":"123241","DOI":"10.1016\/j.eswa.2024.123241","article-title":"Regularizing cross-attention learning for end-to-end speech translation with asr and mt attention matrices","volume":"247","author":"Zhao","year":"2024","journal-title":"Exp Syst Appl"},{"key":"2026020301213684200_ref107","first-page":"6734","article-title":"Improving end-to-end speech synthesis with local recurrent neural network enhanced transformer","author":"Zheng","year":"2020","journal-title":"Proc IEEE ICASSP"},{"key":"2026020301213684200_ref108","doi-asserted-by":"publisher","first-page":"11106","DOI":"10.1609\/aaai.v35i12.17325","article-title":"Informer: beyond efficient transformer for long sequence time-series forecasting","volume":"35","author":"Zhou","year":"2021","journal-title":"Proc AAAI"},{"key":"2026020301213684200_ref109","first-page":"3797","article-title":"Deep features fusion with mutual attention transformer for skin lesion diagnosis","author":"Zhou","year":"2021","journal-title":"Proc ICIP"},{"key":"2026020301213684200_ref110","first-page":"1","article-title":"Deformable DETR: deformable transformers for end-to-end object detection","author":"Zhu","year":"2021","journal-title":"Proc ICLR"},{"key":"2026020301213684200_ref111","doi-asserted-by":"crossref","DOI":"10.24963\/ijcai.2023\/764","article-title":"A survey on efficient training of transformers","author":"Zhuang","year":"2023"}],"container-title":["Logic Journal of the IGPL"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/jigpal\/article-pdf\/34\/1\/jzaf040\/66714218\/jzaf040.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/jigpal\/article-pdf\/34\/1\/jzaf040\/66714218\/jzaf040.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,3]],"date-time":"2026-02-03T06:22:00Z","timestamp":1770099720000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/jigpal\/article\/doi\/10.1093\/jigpal\/jzaf040\/8455720"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,1,27]]},"references-count":111,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,1,27]]}},"URL":"https:\/\/doi.org\/10.1093\/jigpal\/jzaf040","relation":{},"ISSN":["1367-0751","1368-9894"],"issn-type":[{"value":"1367-0751","type":"print"},{"value":"1368-9894","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2026,2]]},"published":{"date-parts":[[2026,1,27]]},"article-number":"jzaf040"}}