{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T01:21:57Z","timestamp":1760059317909,"version":"build-2065373602"},"reference-count":45,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2025,6,5]],"date-time":"2025-06-05T00:00:00Z","timestamp":1749081600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Walailak University","award":["WU67219"],"award-info":[{"award-number":["WU67219"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Algorithms"],"abstract":"<jats:p>Transformer-based time series models are being increasingly employed for time series data analysis. However, their training remains memory intensive, especially with high-dimensional data and extended look-back windows, while model-level memory optimizations are well studied, the batch formation process remains an underexplored factor to performance inefficiency. This paper introduces a memory-efficient batching framework based on view-based sliding windows operating directly on GPU-resident tensors. This approach eliminates redundant data materialization caused by tensor stacking and reduces data transfer volumes without modifying model architectures. We present two variants of our solution: (1) per-batch optimization for datasets exceeding GPU memory, and (2) dataset-wise optimization for in-memory workloads. We evaluate our proposed batching framework systematically using peak GPU memory consumption and epoch runtime as efficiency metrics across varying batch sizes, sequence lengths, feature dimensions, and model architectures. Results show consistent memory savings, averaging 90% and runtime improvements of up to 33% across multiple transformer-based models (Informer, Autoformer, Transformer, and PatchTST) and a linear baseline (DLinear) without compromising model accuracy. We extensively validate our method using synthetic and standard real-world benchmarks, demonstrating accuracy preservation and practical scalability in distributed GPU environments. The proposed method highlights batch formation process as a critical component for improving training efficiency.<\/jats:p>","DOI":"10.3390\/a18060350","type":"journal-article","created":{"date-parts":[[2025,6,5]],"date-time":"2025-06-05T11:04:06Z","timestamp":1749121446000},"page":"350","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Memory-Efficient Batching for Time Series Transformer Training: A Systematic Evaluation"],"prefix":"10.3390","volume":"18","author":[{"ORCID":"https:\/\/orcid.org\/0009-0006-4423-3860","authenticated-orcid":false,"given":"Phanwadee","family":"Sinthong","sequence":"first","affiliation":[{"name":"Informatics Innovation Center of Excellence, School of Informatics, Walailak University, Nakhon Si Thammarat 80160, Thailand"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nam","family":"Nguyen","sequence":"additional","affiliation":[{"name":"Capital One, New York, NY 10171, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Vijay","family":"Ekambaram","sequence":"additional","affiliation":[{"name":"IBM Research India, Banglore 560045, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Arindam","family":"Jati","sequence":"additional","affiliation":[{"name":"IBM Research India, Banglore 560045, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jayant","family":"Kalagnanam","sequence":"additional","affiliation":[{"name":"IBM TJ Watson Research Center, Yorktown Heights, NY 10598, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7667-2573","authenticated-orcid":false,"given":"Peeravit","family":"Koad","sequence":"additional","affiliation":[{"name":"Informatics Innovation Center of Excellence, School of Informatics, Walailak University, Nakhon Si Thammarat 80160, Thailand"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2025,6,5]]},"reference":[{"key":"ref_1","unstructured":"Thoppilan, R., De Freitas, D., Hall, J., Shazeer, N., Kulshreshtha, A., Cheng, H.T., Jin, A., Bos, T., Baker, L., and Du, Y. (2022). Lamda: Language models for dialog applications. arXiv."},{"key":"ref_2","unstructured":"Adiwardana, D., Luong, M.T., So, D.R., Hall, J., Fiedel, N., Thoppilan, R., Yang, Z., Kulshreshtha, A., Nemade, G., and Lu, Y. (2020). Towards a human-like open-domain chatbot. arXiv."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Sun, S., Galley, M., Chen, Y.C., Brockett, C., Gao, X., Gao, J., Liu, J., and Dolan, B. (2019). Dialogpt: Large-scale generative pre-training for conversational response generation. arXiv.","DOI":"10.18653\/v1\/2020.acl-demos.30"},{"key":"ref_4","first-page":"4839","article-title":"Beyond english-centric multilingual machine translation","volume":"22","author":"Fan","year":"2021","journal-title":"J. Mach. Learn. Res."},{"key":"ref_5","unstructured":"Costa-juss\u00e0, M.R., Cross, J., \u00c7elebi, O., Elbayad, M., Heafield, K., Heffernan, K., Kalbassi, E., Lam, J., Licht, D., and Maillard, J. (2022). No language left behind: Scaling human-centered machine translation. arXiv."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Xue, L., Constant, N., Roberts, A., Kale, M., Al-Rfou, R., Siddhant, A., Barua, A., and Raffel, C. (2021, January 6\u201311). mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, Online.","DOI":"10.18653\/v1\/2021.naacl-main.41"},{"key":"ref_7","first-page":"1877","article-title":"Language models are few-shot learners","volume":"33","author":"Brown","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst. (NeurIPS)"},{"key":"ref_8","first-page":"1","article-title":"Palm: Scaling language modeling with pathways","volume":"24","author":"Chowdhery","year":"2023","journal-title":"J. Mach. Learn. Res."},{"key":"ref_9","unstructured":"Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.A., Lacroix, T., Rozi\u00e8re, B., Goyal, N., Hambro, E., and Azhar, F. (2023). Llama: Open and efficient foundation language models (2023). arXiv."},{"key":"ref_10","unstructured":"Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. (2019). Roberta: A robustly optimized bert pretraining approach. arXiv."},{"key":"ref_11","unstructured":"Kenton, J.D.M.W.C., and Toutanova, L.K. (2019, January 2\u20137). Bert: Pre-training of deep bidirectional transformers for language understanding. Proceedings of the naacL-HLT, Minneapolis, MN, USA."},{"key":"ref_12","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2021, January 3\u20137). An image is worth 16x16 words: Transformers for image recognition at scale. Proceedings of the International Conference on Learning Representations (ICLR), Vienna, Austria."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021, January 11\u201317). Swin transformer: Hierarchical vision transformer using shifted windows. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Online.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_14","unstructured":"Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and J\u00e9gou, H. (2021, January 18\u201324). Training data-efficient image transformers & distillation through attention. Proceedings of the International Conference on Machine Learning (ICML), Online."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Wang, W., Xie, E., Li, X., Fan, D.P., Song, K., Liang, D., Lu, T., Luo, P., and Shao, L. (2021, January 11\u201317). Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Online.","DOI":"10.1109\/ICCV48922.2021.00061"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Feng, Z., Guo, D., Tang, D., Duan, N., Feng, X., Gong, M., Shou, L., Qin, B., Liu, T., and Jiang, D. (2020, January 16\u201320). Codebert: A pre-trained model for programming and natural languages. Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), Online.","DOI":"10.18653\/v1\/2020.findings-emnlp.139"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Ahmad, W.U., Chakraborty, S., Ray, B., and Chang, K.W. (2021, January 6\u201311). Unified pre-training for program understanding and generation. Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Online.","DOI":"10.18653\/v1\/2021.naacl-main.211"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Wang, Y., Wang, W., Joty, S., and Hoi, S.C. (2021, January 7\u201311). Codet5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation. Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), Online.","DOI":"10.18653\/v1\/2021.emnlp-main.685"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Zerveas, G., Jayaraman, S., Patel, D., Bhamidipaty, A., and Eickhoff, C. (2021, January 14\u201318). A transformer-based framework for multivariate time series representation learning. Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Singapore.","DOI":"10.1145\/3447548.3467401"},{"key":"ref_20","unstructured":"Zhou, T., Ma, Z., Wen, Q., Wang, X., Sun, L., and Jin, R. (2022, January 17\u201323). Fedformer: Frequency enhanced decomposed transformer for long-term series forecasting. Proceedings of the International Conference on Machine Learning (ICML), Baltimore, MA, USA."},{"key":"ref_21","unstructured":"Xu, J., Wu, H., Wang, J., and Long, M. (2021, January 3\u20137). Anomaly transformer: Time series anomaly detection with association discrepancy. Proceedings of the International Conference on Learning Representations (ICLR), Vienna, Austria."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Zhou, H., Zhang, S., Peng, J., Zhang, S., Li, J., Xiong, H., and Zhang, W. (2021, January 2\u20139). Informer: Beyond efficient transformer for long sequence time-series forecasting. Proceedings of the AAAI Conference on Artificial Intelligence, Online.","DOI":"10.1609\/aaai.v35i12.17325"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"1201","DOI":"10.14778\/3514061.3514067","article-title":"TranAD: Deep transformer networks for anomaly detection in multivariate time series data","volume":"15","author":"Tuli","year":"2022","journal-title":"Proc. VLDB Endow."},{"key":"ref_24","unstructured":"Yang, C.H.H., Tsai, Y.Y., and Chen, P.Y. (2021, January 18\u201324). Voice2series: Reprogramming acoustic models for time series classification. Proceedings of the International Conference on Machine Learning (ICML), Online."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"20200209","DOI":"10.1098\/rsta.2020.0209","article-title":"Time-series forecasting with deep learning: A survey","volume":"379","author":"Lim","year":"2021","journal-title":"Philos. Trans. R. Soc. A"},{"key":"ref_26","first-page":"22419","article-title":"Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting","volume":"34","author":"Wu","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_27","unstructured":"Chowdhury, S.P., Solomou, A., Dubey, A., and Sachan, M. (2021). On learning the transformer kernel. arXiv."},{"key":"ref_28","first-page":"12267","article-title":"Tempo: Accelerating transformer-based model training through memory footprint reduction","volume":"35","author":"Andoorveedu","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"2150","DOI":"10.14778\/3415478.3415530","article-title":"Pytorch distributed: Experiences on accelerating data parallel training","volume":"13","author":"Li","year":"2020","journal-title":"Proc. VLDB Endow."},{"key":"ref_30","unstructured":"Nie, Y., Nguyen, N.H., Sinthong, P., and Kalagnanam, J. (2023, January 1\u20135). A Time Series is Worth 64 Words: Long-term Forecasting with Transformers. Proceedings of the International Conference on Learning Representations (ICLR), Kigali, Rwanda."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"1748","DOI":"10.1016\/j.ijforecast.2021.03.012","article-title":"Temporal fusion transformers for interpretable multi-horizon time series forecasting","volume":"37","author":"Lim","year":"2021","journal-title":"Int. J. Forecast."},{"key":"ref_32","unstructured":"Zeng, A., Chen, M., Zhang, L., and Xu, Q. (2023, January 7\u201314). Are transformers effective for time series forecasting?. Proceedings of the AAAI Conference on Artificial Intelligence, Washington, DC, USA."},{"key":"ref_33","first-page":"28092","article-title":"Post-training quantization for vision transformer","volume":"34","author":"Liu","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_34","first-page":"27168","article-title":"Zeroquant: Efficient and affordable post-training quantization for large-scale transformers","volume":"35","author":"Yao","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_35","unstructured":"Zhou, S., Wu, Y., Ni, Z., Zhou, X., Wen, H., and Zou, Y. (2016). Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients. arXiv."},{"key":"ref_36","unstructured":"Esser, S.K., McKinstry, J.L., Bablani, D., Appuswamy, R., and Modha, D.S. (2019). Learned step size quantization. arXiv."},{"key":"ref_37","first-page":"16344","article-title":"Flashattention: Fast and memory-efficient exact attention with io-awareness","volume":"35","author":"Dao","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_38","unstructured":"Dao, T. (2023). Flashattention-2: Faster attention with better parallelism and work partitioning. arXiv."},{"key":"ref_39","unstructured":"Hu, E.J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. (2021). Lora: Low-rank adaptation of large language models. arXiv."},{"key":"ref_40","first-page":"10088","article-title":"QLORA: Efficient finetuning of quantized LLMs","volume":"36","author":"Dettmers","year":"2024","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Huang, C.C., Jin, G., and Li, J. (2020, January 16\u201320). Swapadvisor: Pushing deep learning beyond the gpu memory limit via smart swapping. Proceedings of the Twenty-Fifth International Conference on Architectural Support for Programming Languages and Operating Systems, Lausanne, Switzerland.","DOI":"10.1145\/3373376.3378530"},{"key":"ref_42","unstructured":"Chen, T., Xu, B., Zhang, C., and Guestrin, C. (2016). Training deep nets with sublinear memory cost. arXiv."},{"key":"ref_43","unstructured":"Kirisame, M., Lyubomirsky, S., Haan, A., Brennan, J., He, M., Roesch, J., Chen, T., and Tatlock, Z. (2020). Dynamic tensor rematerialization. arXiv."},{"key":"ref_44","first-page":"497","article-title":"Checkmate: Breaking the memory wall with optimal tensor rematerialization","volume":"2","author":"Jain","year":"2020","journal-title":"Proc. Mach. Learn. Syst."},{"key":"ref_45","unstructured":"Nie, Y. (2025, March 03). PatchTST. Available online: https:\/\/github.com\/yuqinie98\/PatchTST."}],"container-title":["Algorithms"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-4893\/18\/6\/350\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T17:47:16Z","timestamp":1760032036000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-4893\/18\/6\/350"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,5]]},"references-count":45,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2025,6]]}},"alternative-id":["a18060350"],"URL":"https:\/\/doi.org\/10.3390\/a18060350","relation":{},"ISSN":["1999-4893"],"issn-type":[{"type":"electronic","value":"1999-4893"}],"subject":[],"published":{"date-parts":[[2025,6,5]]}}}