{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,22]],"date-time":"2025-10-22T00:16:40Z","timestamp":1761092200021,"version":"build-2065373602"},"reference-count":169,"publisher":"MDPI AG","issue":"10","license":[{"start":{"date-parts":[[2025,10,19]],"date-time":"2025-10-19T00:00:00Z","timestamp":1760832000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100007706","name":"Ministero dello Sviluppo Economico","doi-asserted-by":"crossref","award":["2022-CUP B69J23000500005"],"award-info":[{"award-number":["2022-CUP B69J23000500005"]}],"id":[{"id":"10.13039\/501100007706","id-type":"DOI","asserted-by":"crossref"}]},{"name":"National Recovery and Resilience Plan (PNRR) of the Italian Ministry of University and Research","award":["PE00000013_1-CUP E63C22002150007"],"award-info":[{"award-number":["PE00000013_1-CUP E63C22002150007"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>The raster is the most common type of spatio-temporal data, and it can be either regularly or irregularly spaced. Spatio-temporal prediction on regular raster data is crucial for modelling and understanding dynamics in disparate realms, such as environment, traffic, astronomy, remote sensing, gaming and video processing, to name a few. Historically, statistical and classical machine learning methods have been used to model spatio-temporal data, and, in recent years, deep learning has shown outstanding results in regular raster spatio-temporal prediction. This work provides a self-contained review about effective deep learning methods for the prediction of regular raster spatio-temporal data. Each deep learning technique is described in detail, underlining its advantages and drawbacks. Finally, a discussion of relevant aspects and further developments in deep learning for regular raster spatio-temporal prediction is presented.<\/jats:p>","DOI":"10.3390\/info16100917","type":"journal-article","created":{"date-parts":[[2025,10,20]],"date-time":"2025-10-20T13:54:41Z","timestamp":1760968481000},"page":"917","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Deep Learning for Regular Raster Spatio-Temporal Prediction: An Overview"],"prefix":"10.3390","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1606-217X","authenticated-orcid":false,"given":"Vincenzo","family":"Capone","sequence":"first","affiliation":[{"name":"Department of Science and Technology, Parthenope University of Naples, Centro Direzionale Isola C4, 80143 Naples, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7577-6765","authenticated-orcid":false,"given":"Angelo","family":"Casolaro","sequence":"additional","affiliation":[{"name":"Department of Science and Technology, Parthenope University of Naples, Centro Direzionale Isola C4, 80143 Naples, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4439-7583","authenticated-orcid":false,"given":"Francesco","family":"Camastra","sequence":"additional","affiliation":[{"name":"Department of Science and Technology, Parthenope University of Naples, Centro Direzionale Isola C4, 80143 Naples, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2025,10,19]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"104502","DOI":"10.1016\/j.envsoft.2019.104502","article-title":"A spatiotemporal deep learning model for sea surface temperature field prediction using time-series satellite data","volume":"120","author":"Xiao","year":"2019","journal-title":"Environ. Model. Softw."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"23070","DOI":"10.1109\/ACCESS.2021.3055554","article-title":"Attentive Spatial Temporal Graph CNN for Land Cover Mapping From Multi Temporal Remote Sensing Data","volume":"9","author":"Censi","year":"2021","journal-title":"IEEE Access"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"117137","DOI":"10.1016\/j.neuroimage.2020.117137","article-title":"Connectome spectral analysis to track EEG task dynamics on a subsecond scale","volume":"221","author":"Glomb","year":"2020","journal-title":"NeuroImage"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Pillai, K.G., Angryk, R.A., Banda, J.M., Schuh, M.A., and Wylie, T. (2012, January 10\u201313). Spatio-temporal Co-occurrence Pattern Mining in Data Sets with Evolving Regions. Proceedings of the 2012 IEEE 12th International Conference on Data Mining Workshops, Brussels, Belgium.","DOI":"10.1109\/ICDMW.2012.130"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"A8","DOI":"10.1051\/0004-6361\/202243461","article-title":"Spatio-temporal analysis of chromospheric heating in a plage region","volume":"664","author":"Morosin","year":"2022","journal-title":"Astron. Astrophys."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Cressie, N.A.C. (1993). Statistics for Spatial Data, Wiley.","DOI":"10.1002\/9781119115151"},{"key":"ref_7","unstructured":"Cressie, N.A.C., and Wikle, C.K. (2011). Statistics for Spatio-Temporal Data, Wiley."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3161602","article-title":"Spatio-Temporal Data Mining: A Survey of Problems and Methods","volume":"51","author":"Atluri","year":"2019","journal-title":"Acm Comput. Surv."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"3681","DOI":"10.1109\/TKDE.2020.3025580","article-title":"Deep Learning for Spatio-Temporal Data Mining: A Survey","volume":"34","author":"Wang","year":"2020","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"ref_10","unstructured":"Goodfellow, I., Bengio, Y., and Courville, A. (2016). Deep Learning, MIT Press."},{"key":"ref_11","first-page":"6000","article-title":"Attention is All you Need","volume":"Volume 30","author":"Guyon","year":"2017","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"ref_12","first-page":"802","article-title":"Convolutional LSTM Network: A Machine Learning Approach for Precipitation Nowcasting","volume":"Volume 28","author":"Cortes","year":"2015","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"ref_13","first-page":"25390","article-title":"Earthformer: Exploring Space-Time Transformers for Earth System Forecasting","volume":"Volume 35","author":"Koyejo","year":"2022","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"2208","DOI":"10.1109\/TPAMI.2022.3165153","article-title":"PredRNN: A Recurrent Neural Network for Spatiotemporal Predictive Learning","volume":"45","author":"Wang","year":"2023","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Benson, V., Robin, C., Requena-Mesa, C., Alonso, L., Carvalhais, N., Cort\u00e9s, J., Gao, Z., Linscheid, N., Weynants, M., and Reichstein, M. (2024, January 17\u201321). Multi-modal Learning for Geospatial Vegetation Forecasting. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.02625"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"924","DOI":"10.1109\/TGRS.2018.2863224","article-title":"Learning Spectral-Spatial-Temporal Features via a Recurrent Convolutional Neural Network for Change Detection in Multispectral Imagery","volume":"57","author":"Mou","year":"2019","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_17","first-page":"102731","article-title":"Space-time super-resolution for satellite video: A joint framework based on multi-scale spatial-temporal transformer","volume":"108","author":"Xiao","year":"2022","journal-title":"Int. J. Appl. Earth Obs. Geoinf."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Ebel, P., Fare Garnot, V.S., Schmitt, M., Wegner, J.D., and Zhu, X.X. (2023, January 17\u201324). UnCRtainTS: Uncertainty Quantification for Cloud Removal in Optical Satellite Time Series. Proceedings of the 2023 IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Vancouver, BC, Canada.","DOI":"10.1109\/CVPRW59228.2023.00202"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"51617","DOI":"10.1109\/ACCESS.2025.3551782","article-title":"SatDiff: A Stable Diffusion Framework for Inpainting Very High-Resolution Satellite Imagery","volume":"13","author":"Panboonyuen","year":"2025","journal-title":"IEEE Access"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"3913","DOI":"10.1109\/TITS.2019.2906365","article-title":"Deep Spatial\u2013Temporal 3D Convolutional Neural Networks for Traffic Data Forecasting","volume":"20","author":"Guo","year":"2019","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"102674","DOI":"10.1016\/j.trc.2020.102674","article-title":"Stacked bidirectional and unidirectional LSTM recurrent neural network for forecasting network-wide traffic state with missing values","volume":"118","author":"Cui","year":"2020","journal-title":"Transp. Res. Part Emerg. Technol."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"22386","DOI":"10.1109\/TITS.2021.3102983","article-title":"Learning Dynamic and Hierarchical Traffic Spatiotemporal Features with Transformer","volume":"23","author":"Yan","year":"2022","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"124187","DOI":"10.1016\/j.eswa.2024.124187","article-title":"Modeling dynamic spatio-temporal correlations and transitions with time window partitioning for traffic flow prediction","volume":"252","author":"Yu","year":"2024","journal-title":"Expert Syst. Appl."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"449","DOI":"10.1111\/rssc.12540","article-title":"Forecasting High-Frequency Spatio-Temporal Wind Power with Dimensionally Reduced Echo State Networks","volume":"71","author":"Huang","year":"2022","journal-title":"J. R. Stat. Soc. Ser. Appl. Stat."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"119868","DOI":"10.1016\/j.renene.2023.119868","article-title":"High-resolution spatiotemporal assessment of solar potential from remote sensing data using deep learning","volume":"222","author":"Mongus","year":"2024","journal-title":"Renew. Energy"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Girdhar, R., Joao Carreira, J., Doersch, C., and Zisserman, A. (2019, January 15\u201320). Video Action Transformer Network. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00033"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Arnab, A., Dehghani, M., Heigold, G., Sun, C., Lu\u010di\u0107, M., and Schmid, C. (2021, January 10\u201317). ViViT: A Video Vision Transformer. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00676"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Liu, Z., Ning, J., Cao, Y., Wei, Y., Zhang, Z., Lin, S., and Hu, H. (2022, January 18\u201324). Video Swin Transformer. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00320"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Zhai, S., Ye, Z., Liu, J., Xie, W., Hu, J., Peng, Z., Xue, H., Chen, D., Wang, X., and Yang, L. (2025, January 17\u201324). StarGen: A Spatiotemporal Autoregression Framework with Video Diffusion Model for Scalable and Controllable Scene Generation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52734.2025.02498"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Wikle, C.K., Zammit Mangion, A., and Cressie, N.A.C. (2019). Spatio-Temporal Statistics with R, Taylor & Francis Group. Chapman & Hall\/CRC: The R Series.","DOI":"10.1201\/9781351769723"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"35","DOI":"10.2307\/1268381","article-title":"A Three-Stage Iterative Procedure for Space-Time Modeling","volume":"22","author":"Pfeifer","year":"1980","journal-title":"Technometrics"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"117","DOI":"10.1111\/j.1538-4632.1981.tb00720.x","article-title":"Seasonal Space-Time ARIMA Modeling","volume":"13","author":"Pfeifer","year":"1981","journal-title":"Geogr. Anal."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"762","DOI":"10.1080\/01621459.1986.10478333","article-title":"Estimation and Identification of Space-Time ARMAX Models in the Presence of Missing Data","volume":"81","author":"Stoffer","year":"1986","journal-title":"J. Am. Stat. Assoc."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Vapnik, V.N. (2000). The Nature of Statistical Learning Theory, Springer.","DOI":"10.1007\/978-1-4757-3264-1"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"5","DOI":"10.1023\/A:1010933404324","article-title":"Random Forests","volume":"45","author":"Breiman","year":"2001","journal-title":"Mach. Learn."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Bishop, C.M., and Bishop, H. (2024). Deep Learning: Foundations and Concepts, Springer International Publishing. [1st ed.].","DOI":"10.1007\/978-3-031-45468-4"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep Residual Learning for Image Recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Lea, C., Flynn, M.D., Vidal, R., Reiter, A., and Hager, G.D. (2017, January 21\u201326). Temporal Convolutional Networks for Action Segmentation and Detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.113"},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"147","DOI":"10.1016\/j.artint.2018.03.002","article-title":"Predicting citywide crowd flows using deep spatio-temporal residual networks","volume":"259","author":"Zhang","year":"2018","journal-title":"Artif. Intell."},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"568","DOI":"10.1038\/s41586-019-1559-7","article-title":"Deep learning for multi-year ENSO forecasts","volume":"573","author":"Ham","year":"2019","journal-title":"Nature"},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"186","DOI":"10.1016\/j.procs.2019.02.036","article-title":"All convolutional neural networks for radar-based precipitation nowcasting","volume":"150","author":"Ayzel","year":"2019","journal-title":"Procedia Comput. Sci."},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"100408","DOI":"10.1016\/j.spasta.2020.100408","article-title":"Deep integro-difference equation models for spatio-temporal forecasting","volume":"37","author":"Wikle","year":"2020","journal-title":"Spat. Stat."},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"5124","DOI":"10.1038\/s41467-021-25257-4","article-title":"Seasonal Arctic sea ice forecasting with probabilistic deep learning","volume":"12","author":"Andersson","year":"2021","journal-title":"Nat. Commun."},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"221","DOI":"10.1109\/TPAMI.2012.59","article-title":"3D Convolutional Neural Networks for Human Action Recognition","volume":"35","author":"Ji","year":"2013","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Karpathy, A., Toderici, G., Shetty, S., Leung, T., Sukthankar, R., and Fei-Fei, L. (2014, January 23\u201328). Large-scale Video Classification with Convolutional Neural Networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.223"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Tran, D., Bourdev, L., Fergus, R., Torresani, L., and Paluri, M. (2015, January 7\u201313). Learning Spatiotemporal Features with 3D Convolutional Networks. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.510"},{"key":"ref_47","unstructured":"Dumoulin, V., and Visin, F. (2016). A guide to convolution arithmetic for deep learning. arXiv."},{"key":"ref_48","first-page":"234","article-title":"U-Net: Convolutional Networks for Biomedical Image Segmentation","volume":"Volume 9351","author":"Navab","year":"2015","journal-title":"Medical Image Computing and Computer-Assisted Intervention\u2014MICCAI 2015"},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Wu, P., Yin, Z., Yang, H., Wu, Y., and Ma, X. (2019). Reconstructing Geostationary Satellite Land Surface Temperature Imagery Based on a Multiscale Feature Connected Convolutional Neural Network. Remote Sens., 11.","DOI":"10.3390\/rs11030300"},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"140","DOI":"10.1007\/978-3-642-15567-3_11","article-title":"Convolutional Learning of Spatio-temporal Features","volume":"Volume 6316","author":"Daniilidis","year":"2010","journal-title":"Computer Vision\u2014ECCV 2010"},{"key":"ref_51","unstructured":"Feichtenhofer, C., Fan, H., Malik, J., and He, K. (November, January 27). SlowFast Networks for Video Recognition. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea."},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Gao, Z., Tan, C., Wu, L., and Li, S.Z. (2022, January 18\u201324). SimVP: Simpler yet Better Video Prediction. Proceedings of the 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00317"},{"key":"ref_53","doi-asserted-by":"crossref","first-page":"18","DOI":"10.1109\/MSP.2017.2693418","article-title":"Geometric Deep Learning: Going beyond Euclidean data","volume":"34","author":"Bronstein","year":"2017","journal-title":"IEEE Signal Process. Mag."},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Dai, J., Qi, H., Xiong, Y., Li, Y., Zhang, G., Hu, H., and Wei, Y. (2017, January 22\u201329). Deformable Convolutional Networks. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.89"},{"key":"ref_55","doi-asserted-by":"crossref","first-page":"425","DOI":"10.1175\/JCLI-D-11-00175.1","article-title":"Systematic Comparison of ENSO Teleconnection Patterns between Models and Observations","volume":"25","author":"Yang","year":"2012","journal-title":"J. Clim."},{"key":"ref_56","doi-asserted-by":"crossref","first-page":"179","DOI":"10.1207\/s15516709cog1402_1","article-title":"Finding Structure in Time","volume":"14","author":"Elman","year":"1990","journal-title":"Cogn. Sci."},{"key":"ref_57","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long Short-Term Memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Comput."},{"key":"ref_58","doi-asserted-by":"crossref","unstructured":"Cho, K., van Merrienboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y. (2014). Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation. arXiv.","DOI":"10.3115\/v1\/D14-1179"},{"key":"ref_59","first-page":"13","article-title":"The \u201cecho state\u201d approach to analysing and training recurrent neural networks-with an erratum note","volume":"148","author":"Jaeger","year":"2001","journal-title":"Bonn Ger. Ger. Natl. Res. Cent. Inf. Technol. Gmd Tech. Rep."},{"key":"ref_60","doi-asserted-by":"crossref","first-page":"839","DOI":"10.1109\/TCYB.2017.2788081","article-title":"Spatial\u2013Temporal Recurrent Neural Network for Emotion Recognition","volume":"49","author":"Zhang","year":"2019","journal-title":"IEEE Trans. Cybern."},{"key":"ref_61","doi-asserted-by":"crossref","unstructured":"Jain, A., Zamir, A.R., Savarese, S., and Saxena, A. (2016, January 27\u201330). Structural-RNN: Deep Learning on Spatio-Temporal Graphs. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.573"},{"key":"ref_62","doi-asserted-by":"crossref","first-page":"130400","DOI":"10.1016\/j.neucom.2025.130400","article-title":"Spatio-temporal prediction using graph neural networks: A survey","volume":"643","author":"Capone","year":"2025","journal-title":"Neurocomputing"},{"key":"ref_63","doi-asserted-by":"crossref","first-page":"315","DOI":"10.1002\/sta4.160","article-title":"An ensemble quadratic echo state network for non-linear spatio-temporal forecasting","volume":"6","author":"McDermott","year":"2017","journal-title":"Stat"},{"key":"ref_64","doi-asserted-by":"crossref","first-page":"e2553","DOI":"10.1002\/env.2553","article-title":"Deep echo state networks with uncertainty quantification for spatio-temporal forecasting","volume":"30","author":"McDermott","year":"2019","journal-title":"Environmetrics"},{"key":"ref_65","doi-asserted-by":"crossref","unstructured":"Fragkiadaki, K., Levine, S., Felsen, P., and Malik, J. (2015, January 7\u201313). Recurrent Network Models for Human Dynamics. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.494"},{"key":"ref_66","unstructured":"Srivastava, N., Mansimov, E., and Salakhudinov, R. (2015, January 6\u201311). Unsupervised Learning of Video Representations using LSTMs. Proceedings of the 32nd International Conference on Machine Learning, Lille, France."},{"key":"ref_67","doi-asserted-by":"crossref","unstructured":"Jia, X., Khandelwal, A., Nayak, G., Gerber, J., Carlson, K., West, P., and Kumar, V. (2017, January 27\u201329). Predict Land Covers with Transition Modeling and Incremental Learning. Proceedings of the 2017 SIAM International Conference on Data Mining (SDM), Houston, TX, USA.","DOI":"10.1137\/1.9781611974973.20"},{"key":"ref_68","doi-asserted-by":"crossref","unstructured":"Jia, X., Khandelwal, A., Nayak, G., Gerber, J., Carlson, K., West, P., and Kumar, V. (2017, January 13\u201317). Incremental Dual-memory LSTM in Land Cover Prediction. Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Halifax, NS, Canada.","DOI":"10.1145\/3097983.3098112"},{"key":"ref_69","doi-asserted-by":"crossref","first-page":"409","DOI":"10.1007\/s40808-018-0431-3","article-title":"Prediction of vegetation dynamics using NDVI time series data and LSTM","volume":"4","author":"Reddy","year":"2018","journal-title":"Model. Earth Syst. Environ."},{"key":"ref_70","doi-asserted-by":"crossref","unstructured":"Ndikumana, E., Ho Tong Minh, D., Baghdadi, N., Courault, D., and Hossard, L. (2018). Deep Recurrent Neural Network for Agricultural Classification using multitemporal SAR Sentinel-1 for Camargue, France. Remote Sens., 10.","DOI":"10.3390\/rs10081217"},{"key":"ref_71","doi-asserted-by":"crossref","unstructured":"McDermott, P.L., and Wikle, C.K. (2019). Bayesian Recurrent Neural Network Models for Forecasting and Quantifying Uncertainty in Spatial-Temporal Data. Entropy, 21.","DOI":"10.3390\/e21020184"},{"key":"ref_72","doi-asserted-by":"crossref","first-page":"191","DOI":"10.1016\/j.neunet.2020.02.016","article-title":"Backpropagation algorithms and Reservoir Computing in Recurrent Neural Networks for the forecasting of complex spatiotemporal dynamics","volume":"126","author":"Vlachas","year":"2020","journal-title":"Neural Netw."},{"key":"ref_73","doi-asserted-by":"crossref","unstructured":"Lees, T., Tseng, G., Atzberger, C., Reece, S., and Dadson, S. (2022). Deep Learning for Vegetation Health Forecasting: A Case Study in Kenya. Remote Sens., 14.","DOI":"10.3390\/rs14030698"},{"key":"ref_74","first-page":"e220002","article-title":"Machine Learning Crop Yield Models Based on Meteorological Features and Comparison with a Process-Based Model","volume":"1","author":"Liu","year":"2022","journal-title":"Artif. Intell. Earth Syst."},{"key":"ref_75","doi-asserted-by":"crossref","first-page":"91","DOI":"10.1016\/j.isprsjprs.2019.01.011","article-title":"DuPLO: A DUal view Point deep Learning architecture for time series classificatiOn","volume":"149","author":"Interdonato","year":"2019","journal-title":"Isprs J. Photogramm. Remote Sens."},{"key":"ref_76","doi-asserted-by":"crossref","first-page":"151","DOI":"10.1016\/j.isprsjprs.2019.05.004","article-title":"Local climate zone-based urban land cover classification from multi-seasonal Sentinel-2 images with a recurrent residual network","volume":"154","author":"Qiu","year":"2019","journal-title":"Isprs J. Photogramm. Remote Sens."},{"key":"ref_77","doi-asserted-by":"crossref","unstructured":"Kaur, A., Goyal, P., Sharma, K., Sharma, L., and Goyal, N. (2022, January 17\u201320). A Generalized Multimodal Deep Learning Model for Early Crop Yield Prediction. Proceedings of the 2022 IEEE International Conference on Big Data (Big Data), Osaka, Japan.","DOI":"10.1109\/BigData55660.2022.10020917"},{"key":"ref_78","doi-asserted-by":"crossref","unstructured":"Graves, A. (2013). Generating Sequences with Recurrent Neural Networks. arXiv.","DOI":"10.1007\/978-3-642-24797-2_3"},{"key":"ref_79","first-page":"5622","article-title":"Deep Learning for Precipitation Nowcasting: A Benchmark and A New Model","volume":"Volume 30","author":"Guyon","year":"2017","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"ref_80","unstructured":"Guyon, I., Luxburg, U.V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., and Garnett, R. (2017, January 4\u20139). PredRNN: Recurrent Neural Networks for Predictive Learning using Spatiotemporal LSTMs. Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA."},{"key":"ref_81","doi-asserted-by":"crossref","unstructured":"Wang, Y., Zhang, J., Zhu, H., Long, M., Wang, J., and Yu, P.S. (2019, January 15\u201320). Memory in Memory: A Predictive Neural Network for Learning Higher-Order Non-Stationarity From Spatiotemporal Dynamics. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00937"},{"key":"ref_82","doi-asserted-by":"crossref","first-page":"207","DOI":"10.1109\/LGRS.2017.2780843","article-title":"A CFCC-LSTM Model for Sea Surface Temperature Prediction","volume":"15","author":"Yang","year":"2018","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_83","doi-asserted-by":"crossref","first-page":"11","DOI":"10.1016\/j.isprsjprs.2019.09.016","article-title":"Combining Sentinel-1 and Sentinel-2 Satellite Image Time Series for land cover mapping via a multi-source deep learning architecture","volume":"158","author":"Ienco","year":"2019","journal-title":"Isprs J. Photogramm. Remote Sens."},{"key":"ref_84","doi-asserted-by":"crossref","first-page":"101325","DOI":"10.1016\/j.ecoinf.2021.101325","article-title":"A novel CNN-LSTM-based approach to predict urban expansion","volume":"64","author":"Boulila","year":"2021","journal-title":"Ecol. Inform."},{"key":"ref_85","doi-asserted-by":"crossref","unstructured":"Wu, H., Yao, Z., Wang, J., and Long, M. (2021, January 20\u201325). MotionRNN: A Flexible Model for Video Prediction with Spacetime-Varying Motions. Proceedings of the 2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01518"},{"key":"ref_86","unstructured":"Robin, C., Requena-Mesa, C., Benson, V., Alonso, L., Poehls, J., Carvalhais, N., and Reichstein, M. (2022). Learning to Forecast Vegetation Greenness at Fine Resolution over Africa with ConvLSTMs. arXiv."},{"key":"ref_87","unstructured":"Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (2018). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv."},{"key":"ref_88","unstructured":"Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I. (Preprint, 2018). Improving language understanding by generative pre-training, Preprint, in press."},{"key":"ref_89","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020). An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. arXiv."},{"key":"ref_90","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021, January 10\u201317). Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_91","first-page":"813","article-title":"Is Space-Time Attention All You Need for Video Understanding?","volume":"Volume 139","author":"Meila","year":"2021","journal-title":"Proceedings of the 38th International Conference on Machine Learning"},{"key":"ref_92","doi-asserted-by":"crossref","unstructured":"Yan, S., Xiong, X., Arnab, A., Lu, Z., Zhang, M., Sun, C., and Schmid, C. (2022, January 18\u201324). Multiview Transformers for Video Recognition. Proceedings of the 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00333"},{"key":"ref_93","unstructured":"Li, K., Wang, Y., Peng, G., Song, G., Liu, Y., Li, H., and Qiao, Y. (2022, January 25\u201329). UniFormer: Unified Transformer for Efficient Spatial-Temporal Representation Learning. Proceedings of the International Conference on Learning Representations, Virtual."},{"key":"ref_94","doi-asserted-by":"crossref","unstructured":"Neimark, D., Bar, O., Zohar, M., and Asselmann, D. (2021, January 11\u201317). Video Transformer Network. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV) Workshops, Montreal, BC, Canada.","DOI":"10.1109\/ICCVW54120.2021.00355"},{"key":"ref_95","unstructured":"Beltagy, I., Peters, M.E., and Cohan, A. (2020). Longformer: The Long-Document Transformer. arXiv."},{"key":"ref_96","doi-asserted-by":"crossref","unstructured":"Meinhardt, T., Kirillov, A., Leal-Taix\u00e9, L., and Feichtenhofer, C. (2022, January 18\u201324). TrackFormer: Multi-Object Tracking with Transformers. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00864"},{"key":"ref_97","first-page":"35946","article-title":"Masked Autoencoders As Spatiotemporal Learners","volume":"Volume 35","author":"Koyejo","year":"2022","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"ref_98","doi-asserted-by":"crossref","unstructured":"Caballero, J., Ledig, C., Aitken, A., Acosta, A., Totz, J., Wang, Z., and Shi, W. (2017, January 21\u201326). Real-Time Video Super-Resolution with Spatio-Temporal Networks and Motion Compensation. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.304"},{"key":"ref_99","doi-asserted-by":"crossref","first-page":"528","DOI":"10.1007\/978-3-030-58517-4_31","article-title":"Learning Joint Spatial-Temporal Transformations for Video Inpainting","volume":"Volume 12361","author":"Vedaldi","year":"2020","journal-title":"Computer Vision\u2014ECCV 2020"},{"key":"ref_100","doi-asserted-by":"crossref","unstructured":"Yan, B., Peng, H., Fu, J., Wang, D., and Lu, H. (2021, January 10\u201317). Learning Spatio-Temporal Transformer for Visual Tracking. Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.01028"},{"key":"ref_101","doi-asserted-by":"crossref","first-page":"2496","DOI":"10.1109\/TNNLS.2022.3190367","article-title":"An Effective Video Transformer with Synchronized Spatiotemporal and Spatial Self-Attention for Action Recognition","volume":"35","author":"Alfasly","year":"2022","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_102","doi-asserted-by":"crossref","unstructured":"Lin, K., Li, L., Lin, C.C., Ahmed, F., Gan, Z., Liu, Z., Lu, Y., and Wang, L. (2022, January 18\u201324). SwinBERT: End-to-End Transformers with Sparse Attention for Video Captioning. Proceedings of the 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01742"},{"key":"ref_103","doi-asserted-by":"crossref","first-page":"4462","DOI":"10.1109\/TCSVT.2023.3281448","article-title":"MSVT: Multiple Spatiotemporal Views Transformer for DeepFake Video Detection","volume":"33","author":"Yu","year":"2023","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_104","doi-asserted-by":"crossref","first-page":"42","DOI":"10.1109\/TMRB.2023.3237867","article-title":"SVT-SDE: Spatiotemporal Vision Transformers-Based Self-Supervised Depth Estimation in Stereoscopic Surgical Videos","volume":"5","author":"Tao","year":"2023","journal-title":"IEEE Trans. Med. Robot. Bionics"},{"key":"ref_105","doi-asserted-by":"crossref","first-page":"126582","DOI":"10.1016\/j.neucom.2023.126582","article-title":"TSDTVOS: Target-guided spatiotemporal dual-stream transformers for video object segmentation","volume":"555","author":"Zhou","year":"2023","journal-title":"Neurocomputing"},{"key":"ref_106","doi-asserted-by":"crossref","first-page":"3013","DOI":"10.1109\/TIP.2023.3275069","article-title":"Video Summarization with Spatiotemporal Vision Transformer","volume":"32","author":"Hsu","year":"2023","journal-title":"IEEE Transactions on Image Processing"},{"key":"ref_107","unstructured":"Gupta, A., Tian, S., Zhang, Y., Wu, J., Mart\u00edn-Mart\u00edn, R., and Fei-Fei, L. (2023, January 1\u20135). MaskViT: Masked Visual Pre-Training for Video Prediction. Proceedings of the The Eleventh International Conference on Learning Representations, Kigali, Rwanda."},{"key":"ref_108","doi-asserted-by":"crossref","first-page":"2171","DOI":"10.1109\/TIP.2024.3372454","article-title":"VRT: A Video Restoration Transformer","volume":"33","author":"Liang","year":"2024","journal-title":"IEEE Trans. Image Process."},{"key":"ref_109","doi-asserted-by":"crossref","first-page":"6055","DOI":"10.1109\/TPAMI.2024.3377192","article-title":"A Semantic and Motion-Aware Spatiotemporal Transformer Network for Action Detection","volume":"46","author":"Korban","year":"2024","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_110","doi-asserted-by":"crossref","first-page":"214","DOI":"10.1109\/THMS.2024.3370582","article-title":"RTSformer: A Robust Toroidal Transformer with Spatiotemporal Features for Visual Tracking","volume":"54","author":"Gu","year":"2024","journal-title":"IEEE Trans.-Hum.-Mach. Syst."},{"key":"ref_111","doi-asserted-by":"crossref","first-page":"28","DOI":"10.1007\/978-3-031-53305-1_3","article-title":"Spatiotemporal Representation Enhanced ViT for Video Recognition","volume":"Volume 14554","author":"Rudinac","year":"2024","journal-title":"MultiMedia Modeling"},{"key":"ref_112","doi-asserted-by":"crossref","unstructured":"Lin, F., Crawford, S., Guillot, K., Zhang, Y., Chen, Y., Yuan, X., Chen, L., Williams, S., Minvielle, R., and Xiao, X. (2023, January 4\u20136). MMST-ViT: Climate Change-aware Crop Yield Prediction via Multi-Modal Spatial-Temporal Vision Transformer. Proceedings of the 2023 IEEE\/CVF International Conference on Computer Vision (ICCV), Paris, France.","DOI":"10.1109\/ICCV51070.2023.00531"},{"key":"ref_113","doi-asserted-by":"crossref","unstructured":"Wang, W., Xie, E., Li, X., Fan, D.P., Song, K., Liang, D., Lu, T., Luo, P., and Shao, L. (2021, January 10\u201317). Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions. Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00061"},{"key":"ref_114","doi-asserted-by":"crossref","first-page":"415","DOI":"10.1007\/s41095-022-0274-8","article-title":"PVT v2: Improved baselines with pyramid vision transformer","volume":"8","author":"Wang","year":"2022","journal-title":"Comput. Vis. Media"},{"key":"ref_115","unstructured":"Tseng, G., Cartuyvels, R., Zvonkov, I., Purohit, M., Rolnick, D., and Kerner, H. (2023). Lightweight, Pre-trained Transformers for Remote Sensing Timeseries. arXiv."},{"key":"ref_116","doi-asserted-by":"crossref","unstructured":"Tang, S., Li, C., Zhang, P., and Tang, R. (2023, January 4\u20136). SwinLSTM: Improving Spatiotemporal Prediction Accuracy using Swin Transformer and LSTM. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Paris, France.","DOI":"10.1109\/ICCV51070.2023.01239"},{"key":"ref_117","doi-asserted-by":"crossref","first-page":"847","DOI":"10.1109\/JSTARS.2020.2971763","article-title":"A CNN-Transformer Hybrid Approach for Crop Classification Using Multitemporal Multisensor Images","volume":"13","author":"Li","year":"2020","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_118","doi-asserted-by":"crossref","unstructured":"Aksan, E., Kaufmann, M., Cao, P., and Hilliges, O. (2021, January 1\u20133). A Spatio-temporal Transformer for 3D Human Motion Prediction. Proceedings of the 2021 International Conference on 3D Vision (3DV), London, UK.","DOI":"10.1109\/3DV53792.2021.00066"},{"key":"ref_119","first-page":"5607514","article-title":"Remote Sensing Image Change Detection with Transformers","volume":"60","author":"Chen","year":"2022","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_120","doi-asserted-by":"crossref","unstructured":"Huang, L., Mao, F., Zhang, K., and Li, Z. (2022). Spatial-Temporal Convolutional Transformer Network for Multivariate Time Series Forecasting. Sensors, 22.","DOI":"10.3390\/s22030841"},{"key":"ref_121","first-page":"1","article-title":"SwinSUNet: Pure Transformer Network for Remote Sensing Image Change Detection","volume":"60","author":"Zhang","year":"2022","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_122","first-page":"5536814","article-title":"Spectral\u2013Spatial\u2013Temporal Transformers for Hyperspectral Image Change Detection","volume":"60","author":"Wang","year":"2022","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_123","doi-asserted-by":"crossref","first-page":"160446","DOI":"10.1016\/j.scitotenv.2022.160446","article-title":"Predicting hourly PM2.5 concentrations in wildfire-prone areas using a SpatioTemporal Transformer model","volume":"860","author":"Yu","year":"2023","journal-title":"Sci. Total Environ."},{"key":"ref_124","doi-asserted-by":"crossref","unstructured":"Yi, Z., Zhang, H., Tan, P., and Gong, M. (2017, January 22\u201329). DualGAN: Unsupervised Dual Learning for Image-to-Image Translation. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy. Version Number: 4.","DOI":"10.1109\/ICCV.2017.310"},{"key":"ref_125","doi-asserted-by":"crossref","unstructured":"Karras, T., Laine, S., and Aila, T. (2019, January 15\u201320). A Style-Based Generator Architecture for Generative Adversarial Networks. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA. Version Number: 3.","DOI":"10.1109\/CVPR.2019.00453"},{"key":"ref_126","unstructured":"Brock, A., Donahue, J., and Simonyan, K. (2018). Large Scale GAN Training for High Fidelity Natural Image Synthesis. arXiv."},{"key":"ref_127","unstructured":"Wallach, H., Larochelle, H., Beygelzimer, A., Alch\u00e9-Buc, F.d., Fox, E., and Garnett, R. (2019, January 8\u201314). Time-series Generative Adversarial Networks. Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada."},{"key":"ref_128","unstructured":"Liao, S., Ni, H., Szpruch, L., Wiese, M., Sabate-Vidales, M., and Xiao, B. (2020). Conditional Sig-Wasserstein GANs for Time Series Generation. arXiv."},{"key":"ref_129","first-page":"133","article-title":"TTS-GAN: A Transformer-Based Time-Series Generative Adversarial Network","volume":"Volume 13263","author":"Michalowski","year":"2022","journal-title":"Artificial Intelligence in Medicine"},{"key":"ref_130","unstructured":"Ghahramani, Z., Welling, M., Cortes, C., Lawrence, N., and Weinberger, K.Q. (2014, January 8\u201313). Generative Adversarial Nets. Proceedings of the Advances in Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_131","unstructured":"Luo, C. (2022). Understanding Diffusion Models: A Unified Perspective. arXiv."},{"key":"ref_132","first-page":"6840","article-title":"Denoising Diffusion Probabilistic Models","volume":"Volume 33","author":"Larochelle","year":"2020","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"ref_133","unstructured":"Yuan, H., Zhou, S., and Yu, S. (2023). EHRDiff: Exploring Realistic EHR Synthesis with Diffusion Models. arXiv."},{"key":"ref_134","first-page":"45259","article-title":"DYffusion: A Dynamics-informed Diffusion Model for Spatiotemporal Forecasting","volume":"Volume 36","author":"Oh","year":"2023","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"ref_135","first-page":"18100","article-title":"CARD: Classification and Regression Diffusion Models","volume":"Volume 35","author":"Koyejo","year":"2022","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"ref_136","doi-asserted-by":"crossref","unstructured":"Awasthi, A., Ly, S.T., Nizam, J., Mehta, V., Ahmad, S., Nemani, R., Prasad, S., and Nguyen, H.V. (2024, January 2\u20134). Anomaly Detection in Satellite Videos Using Diffusion Models. Proceedings of the 2024 IEEE 26th International Workshop on Multimedia Signal Processing (MMSP), West Lafayette, IN, USA.","DOI":"10.1109\/MMSP61759.2024.10743372"},{"key":"ref_137","unstructured":"Kingma, D.P., and Welling, M. (2013). Auto-Encoding Variational Bayes. arXiv."},{"key":"ref_138","unstructured":"Song, J., Meng, C., and Ermon, S. (2021, January 3\u20137). Denoising Diffusion Implicit Models. Proceedings of the International Conference on Learning Representations, Virtual."},{"key":"ref_139","first-page":"6666","article-title":"STDiff: Spatio-Temporal Diffusion for Continuous Stochastic Video Prediction","volume":"38","author":"Ye","year":"2024","journal-title":"Proc. Aaai Conf. Artif. Intell."},{"key":"ref_140","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/TGRS.2024.3510693","article-title":"Advancing Realistic Precipitation Nowcasting with a Spatiotemporal Transformer-Based Denoising Diffusion Model","volume":"62","author":"Zhao","year":"2024","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_141","first-page":"1","article-title":"DiffCR: A Fast Conditional Diffusion Framework for Cloud Removal From Optical Satellite Images","volume":"62","author":"Zou","year":"2024","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_142","doi-asserted-by":"crossref","unstructured":"Liu, H., Liu, J., Hu, T., and Ma, H. (2025). Spatio-Temporal Probabilistic Forecasting of Wind Speed Using Transformer-Based Diffusion Models. IEEE Trans. Sustain. Energy, 1\u201313.","DOI":"10.1109\/TSTE.2025.3591920"},{"key":"ref_143","doi-asserted-by":"crossref","unstructured":"Yao, S., Zhang, X., Liu, X., Liu, M., and Cui, Z. (2025, January 10\u201317). STDD: Spatio-Temporal Dual Diffusion for Video Generation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR52734.2025.01173"},{"key":"ref_144","doi-asserted-by":"crossref","first-page":"313","DOI":"10.1016\/0304-4149(82)90051-5","article-title":"Reverse-time diffusion equation models","volume":"12","author":"Anderson","year":"1982","journal-title":"Stoch. Processes Their Appl."},{"key":"ref_145","unstructured":"Lipman, Y., Chen, R.T.Q., Ben-Hamu, H., Nickel, M., and Le, M. (2022). Flow Matching for Generative Modeling. arXiv."},{"key":"ref_146","unstructured":"Lipman, Y., Havasi, M., Holderrieth, P., Shaul, N., Le, M., Karrer, B., Chen, R.T.Q., Lopez-Paz, D., Ben-Hamu, H., and Gat, I. (2024). Flow Matching Guide and Code. arXiv."},{"key":"ref_147","doi-asserted-by":"crossref","first-page":"730","DOI":"10.1007\/s11633-025-1562-4","article-title":"DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models","volume":"22","author":"Lu","year":"2025","journal-title":"Mach. Intell. Res."},{"key":"ref_148","unstructured":"Holderrieth, P., Havasi, M., Yim, J., Shaul, N., Gat, I., Jaakkola, T., Karrer, B., Chen, R.T.Q., and Lipman, Y. (2025, January 24\u201328). Generator Matching: Generative modeling with arbitrary Markov processes. Proceedings of the International Conference on Representation Learning, Singapore."},{"key":"ref_149","doi-asserted-by":"crossref","unstructured":"Peebles, W., and Xie, S. (2023, January 2\u20136). Scalable Diffusion Models with Transformers. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Paris, France.","DOI":"10.1109\/ICCV51070.2023.00387"},{"key":"ref_150","unstructured":"Gori, M., Monfardini, G., and Scarselli, F. (August, January 31). A new model for learning in graph domains. Proceedings of the 2005 IEEE International Joint Conference on Neural Networks, Montreal, QC, Canada."},{"key":"ref_151","doi-asserted-by":"crossref","first-page":"61","DOI":"10.1109\/TNN.2008.2005605","article-title":"The Graph Neural Network Model","volume":"20","author":"Scarselli","year":"2009","journal-title":"IEEE Trans. Neural Netw."},{"key":"ref_152","doi-asserted-by":"crossref","first-page":"2923","DOI":"10.5194\/hess-26-2923-2022","article-title":"Quantifying the uncertainty of precipitation forecasting using probabilistic deep learning","volume":"26","author":"Xu","year":"2022","journal-title":"Hydrol. Earth Syst. Sci."},{"key":"ref_153","doi-asserted-by":"crossref","first-page":"103097","DOI":"10.1016\/j.ecoinf.2025.103097","article-title":"Predicting ground-level nitrogen dioxide concentrations using the BaYesian attention-based deep neural network","volume":"87","author":"Casolaro","year":"2025","journal-title":"Ecol. Inform."},{"key":"ref_154","doi-asserted-by":"crossref","first-page":"105654","DOI":"10.1016\/j.envsoft.2023.105654","article-title":"Cyclone trajectory and intensity prediction with uncertainty quantification using variational recurrent neural networks","volume":"162","author":"Kapoor","year":"2023","journal-title":"Environ. Model. Softw."},{"key":"ref_155","doi-asserted-by":"crossref","first-page":"448","DOI":"10.1162\/neco.1992.4.3.448","article-title":"A Practical Bayesian Framework for Backpropagation Networks","volume":"4","author":"MacKay","year":"1992","journal-title":"Neural Comput."},{"key":"ref_156","doi-asserted-by":"crossref","first-page":"686","DOI":"10.1016\/j.jcp.2018.10.045","article-title":"Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations","volume":"378","author":"Raissi","year":"2019","journal-title":"J. Comput. Phys."},{"key":"ref_157","doi-asserted-by":"crossref","first-page":"828","DOI":"10.1038\/s42256-022-00540-1","article-title":"Physically constrained generative adversarial networks for improving precipitation fields from Earth system models","volume":"4","author":"Hess","year":"2022","journal-title":"Nat. Mach. Intell."},{"key":"ref_158","first-page":"1","article-title":"Hard-Constrained Deep Learning for Climate Downscaling","volume":"24","author":"Harder","year":"2023","journal-title":"J. Mach. Learn. Res."},{"key":"ref_159","doi-asserted-by":"crossref","first-page":"122151","DOI":"10.1016\/j.apenergy.2023.122151","article-title":"Explainable Spatio-Temporal Graph Neural Networks for multi-site photovoltaic energy production","volume":"353","author":"Verdone","year":"2024","journal-title":"Appl. Energy"},{"key":"ref_160","unstructured":"Peters, J., Janzing, D., and Sch\u00f6lkopf, B. (2017). Elements of Causal Inference: Foundations and Learning Algorithms, The MIT Press."},{"key":"ref_161","doi-asserted-by":"crossref","first-page":"2245","DOI":"10.1109\/TPAMI.2024.3506283","article-title":"Foundation Models Defining a New Era in Vision: A Survey and Outlook","volume":"47","author":"Awais","year":"2025","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_162","doi-asserted-by":"crossref","first-page":"1180","DOI":"10.1038\/s41586-025-09005-y","article-title":"A foundation model for the Earth system","volume":"641","author":"Bodnar","year":"2025","journal-title":"Nature"},{"key":"ref_163","unstructured":"Szwarcman, D., Roy, S., Fraccaro, P., Gislason, T.E., Blumenstiel, B., Ghosal, R., de Oliveira, P.H., Almeida, J.L.d.S., Sedona, R., and Kang, Y. (2024). Prithvi-EO-2.0: A Versatile Multi-Temporal Foundation Model for Earth Observation Applications. arXiv."},{"key":"ref_164","unstructured":"(2025, October 18). Clay. Clay Foundation Model. Available online: https:\/\/madewithclay.org\/."},{"key":"ref_165","unstructured":"Xiong, Z., Wang, Y., Zhang, F., Stewart, A.J., Hanna, J., Borth, D., Papoutsis, I., Saux, B.L., Camps-Valls, G., and Zhu, X.X. (2024). Neural Plasticity-Inspired Multimodal Foundation Model for Earth Observation. arXiv."},{"key":"ref_166","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/TGRS.2022.3146246","article-title":"SEN12MS-CR-TS: A Remote-Sensing Data Set for Multimodal Multitemporal Cloud Removal","volume":"60","author":"Ebel","year":"2022","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_167","doi-asserted-by":"crossref","first-page":"13838","DOI":"10.1109\/JIOT.2024.3524030","article-title":"Spatiotemporal Pretrained Large Language Model for Forecasting with Missing Values","volume":"12","author":"Fang","year":"2025","journal-title":"IEEE Internet Things J."},{"key":"ref_168","unstructured":"Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J. (2017, January 24\u201326). Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. Proceedings of the International Conference on Learning Representations, Toulon, France."},{"key":"ref_169","first-page":"1","article-title":"Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity","volume":"23","author":"Fedus","year":"2022","journal-title":"J. Mach. Learn. Res."}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/16\/10\/917\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,21]],"date-time":"2025-10-21T04:12:30Z","timestamp":1761019950000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/16\/10\/917"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,19]]},"references-count":169,"journal-issue":{"issue":"10","published-online":{"date-parts":[[2025,10]]}},"alternative-id":["info16100917"],"URL":"https:\/\/doi.org\/10.3390\/info16100917","relation":{},"ISSN":["2078-2489"],"issn-type":[{"type":"electronic","value":"2078-2489"}],"subject":[],"published":{"date-parts":[[2025,10,19]]}}}