{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,4]],"date-time":"2026-08-04T22:12:46Z","timestamp":1785881566538,"version":"3.56.0"},"reference-count":34,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2022,7,30]],"date-time":"2022-07-30T00:00:00Z","timestamp":1659139200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/"}],"funder":[{"DOI":"10.13039\/100000001","name":"US National Science Foundation","doi-asserted-by":"crossref","award":["IIS-1838730"],"award-info":[{"award-number":["IIS-1838730"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Knowl. Discov. Data"],"published-print":{"date-parts":[[2022,12,31]]},"abstract":"<jats:p>\n            Multivariate time-series data are frequently observed in critical care settings and are typically characterized by sparsity (missing information) and irregular time intervals. Existing approaches for learning representations in this domain handle these challenges by either aggregation or imputation of values, which in-turn suppresses the fine-grained information and adds undesirable noise\/overhead into the machine learning model. To tackle this problem, we propose a\n            <jats:bold>S<\/jats:bold>\n            elf-supervised\n            <jats:bold>Tra<\/jats:bold>\n            nsformer for\n            <jats:bold>T<\/jats:bold>\n            ime-\n            <jats:bold>S<\/jats:bold>\n            eries (STraTS) model, which overcomes these pitfalls by treating time-series as a set of observation triplets instead of using the standard dense matrix representation. It employs a novel Continuous Value Embedding technique to encode continuous time and variable values without the need for discretization. It is composed of a Transformer component with multi-head attention layers, which enable it to learn contextual triplet embeddings while avoiding the problems of recurrence and vanishing gradients that occur in recurrent architectures. In addition, to tackle the problem of limited availability of labeled data (which is typically observed in many healthcare applications), STraTS utilizes self-supervision by leveraging unlabeled data to learn better representations by using time-series forecasting as an auxiliary proxy task. Experiments on real-world multivariate clinical time-series benchmark datasets demonstrate that STraTS has better prediction performance than state-of-the-art methods for mortality prediction, especially when labeled data is limited. Finally, we also present an interpretable version of STraTS, which can identify important measurements in the time-series data. Our data preprocessing and model implementation codes are available at\n            <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/github.com\/sindhura97\/STraTS\">https:\/\/github.com\/sindhura97\/STraTS<\/jats:ext-link>\n            .\n          <\/jats:p>","DOI":"10.1145\/3516367","type":"journal-article","created":{"date-parts":[[2022,6,24]],"date-time":"2022-06-24T10:05:22Z","timestamp":1656065122000},"page":"1-17","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":117,"title":["Self-Supervised Transformer for Sparse and Irregularly Sampled Multivariate Clinical Time-Series"],"prefix":"10.1145","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-5502-1616","authenticated-orcid":false,"given":"Sindhu","family":"Tipirneni","sequence":"first","affiliation":[{"name":"Virginia Tech, Arlington, Virginia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2839-3662","authenticated-orcid":false,"given":"Chandan K.","family":"Reddy","sequence":"additional","affiliation":[{"name":"Virginia Tech, Arlington, Virginia"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,7,30]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"Shaojie Bai J. Zico Kolter and Vladlen Koltun. 2018. An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling. CoRR abs\/1803.01271 (2018). arXiv:1803.01271 http:\/\/arxiv.org\/abs\/1803.01271"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/3097983.3097997"},{"key":"e_1_3_2_4_2","first-page":"153","volume-title":"Proceedings of the 21st Annual Conference on Advances in Neural Information Processing Systems","author":"Bonilla Edwin V.","year":"2007","unstructured":"Edwin V. Bonilla, Kian Ming Adam Chai, and Christopher K. I. Williams. 2007. Multi-task gaussian process prediction. In Proceedings of the 21st Annual Conference on Advances in Neural Information Processing Systems. Curran Associates, Inc., 153\u2013160."},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41598-018-24271-9"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2018.00111"},{"key":"e_1_3_2_7_2","first-page":"3504","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 29: Annual Conference on Neural Information Processing Systems","author":"Choi Edward","year":"2016","unstructured":"Edward Choi, Mohammad Taha Bahadori, Jimeng Sun, Joshua Kulas, Andy Schuetz, and Walter F. Stewart. 2016. RETAIN: An interpretable predictive model for healthcare using reverse time attention mechanism. In Proceedings of the Advances in Neural Information Processing Systems 29: Annual Conference on Neural Information Processing Systems. 3504\u20133512."},{"key":"e_1_3_2_8_2","unstructured":"Junyoung Chung \u00c7aglar G\u00fcl\u00e7ehre KyungHyun Cho and Yoshua Bengio. 2014. Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling. CoRR abs\/1412.3555 (2014). arXiv:1412.3555 http:\/\/arxiv.org\/abs\/1412.3555"},{"issue":"1","key":"e_1_3_2_9_2","first-page":"1","article-title":"Prognostic implications of blood lactate concentrations after cardiac arrest: a retrospective study","volume":"7","author":"Dell\u2019Anna Antonio Maria","year":"2017","unstructured":"Antonio Maria Dell\u2019Anna, Claudio Sandroni, Irene Lamanna, Ilaria Belloni, Katia Donadello, Jacques Creteur, Jean-Louis Vincent, and Fabio Silvio Taccone. 2017. Prognostic implications of blood lactate concentrations after cardiac arrest: a retrospective study. Annals of Intensive Care 7, 1 (2017), 1\u20139.","journal-title":"Annals of Intensive Care"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/n19-1423"},{"key":"e_1_3_2_11_2","article-title":"Biochemistry, Lactate Dehydrogenase.[Updated 2020 May 17]","author":"Farhana A.","year":"2021","unstructured":"A. Farhana and S. L. Lappin. 2021. Biochemistry, Lactate Dehydrogenase.[Updated 2020 May 17]. StatPearls [Internet]. Treasure Island (FL): StatPearls Publishing (2021).","journal-title":"StatPearls [Internet]. Treasure Island (FL): StatPearls Publishing"},{"key":"e_1_3_2_12_2","first-page":"1174","volume-title":"Proceedings of the 34th International Conference on Machine Learning, (ICML\u201917)","author":"Futoma Joseph","year":"2017","unstructured":"Joseph Futoma, Sanjay Hariharan, and Katherine A. Heller. 2017. Learning to detect sepsis with a multitask gaussian process RNN classifier. In Proceedings of the 34th International Conference on Machine Learning, (ICML\u201917). PMLR, 1174\u20131182."},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1161\/01.CIR.101.23.e215"},{"key":"e_1_3_2_14_2","first-page":"4353","volume-title":"Proceedings of the 37th International Conference on Machine Learning, (ICML\u201920)","volume":"119","author":"Horn Max","year":"2020","unstructured":"Max Horn, Michael Moor, Christian Bock, Bastian Rieck, and Karsten M. Borgwardt. 2020. Set functions for time series. In Proceedings of the 37th International Conference on Machine Learning, (ICML\u201920), Vol. 119. PMLR, 4353\u20134363."},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-47426-3_39"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2020.2992393"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1038\/sdata.2016.35"},{"key":"e_1_3_2_18_2","volume-title":"Proceedings of the 3rd International Conference on Learning Representations, (ICLR\u201915)","author":"Kingma Diederik P.","year":"2015","unstructured":"Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In Proceedings of the 3rd International Conference on Learning Representations, (ICLR\u201915)."},{"key":"e_1_3_2_19_2","first-page":"484","volume-title":"Proceedings of the 31st Conference on Uncertainty in Artificial Intelligence","author":"Li Steven Cheng-Xian","year":"2015","unstructured":"Steven Cheng-Xian Li and Benjamin M. Marlin. 2015. Classification of sparse and irregularly sampled time series with mixtures of expected gaussian kernels and random features. In Proceedings of the 31st Conference on Uncertainty in Artificial Intelligence. AUAI Press, 484\u2013493."},{"key":"e_1_3_2_20_2","first-page":"1804","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 29: Annual Conference on Neural Information Processing Systems","author":"Li Steven Cheng-Xian","year":"2016","unstructured":"Steven Cheng-Xian Li and Benjamin M. Marlin. 2016. A scalable end-to-end Gaussian process adapter for irregularly sampled time series classification. In Proceedings of the Advances in Neural Information Processing Systems 29: Annual Conference on Neural Information Processing Systems (2016). 1804\u20131812."},{"key":"e_1_3_2_21_2","unstructured":"Zachary Chase Lipton David C. Kale and Randall C. Wetzel. 2015. Phenotyping of Clinical Time Series with LSTM Recurrent Neural Networks. CoRR abs\/1510.07641 (2015). arXiv:1510.07641 http:\/\/arxiv.org\/abs\/1510.07641"},{"key":"e_1_3_2_22_2","series-title":"Proceedings of the 1st Machine Learning in Health Care, MLHC 2016","first-page":"253","volume":"56","author":"Lipton Zachary C.","year":"2016","unstructured":"Zachary C. Lipton, David C. Kale, and Randall C. Wetzel. 2016. Directly modeling missing data in sequences with RNNs: Improved classification of clinical time series. In Proceedings of the 1st Machine Learning in Health Care, MLHC 2016(JMLR Workshop and Conference Proceedings, Vol. 56). JMLR.org, 253\u2013270."},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2021.3090866"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/1390156.1390235"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-28650-9_4"},{"key":"e_1_3_2_26_2","first-page":"5321","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems, (2019) NeurIPS\u201919","author":"Rubanova Yulia","year":"2019","unstructured":"Yulia Rubanova, Tian Qi Chen, and David Duvenaud. 2019. Latent ordinary differential equations for irregularly-sampled time series. In Proceedings of the Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems, (2019) NeurIPS\u201919. 5321\u20135331."},{"key":"e_1_3_2_27_2","volume-title":"Proceedings of the 7th International Conference on Learning Representations, (ICLR\u201919)","author":"Shukla Satya Narayan","year":"2019","unstructured":"Satya Narayan Shukla and Benjamin M. Marlin. 2019. Interpolation-prediction networks for irregularly sampled time series. In Proceedings of the 7th International Conference on Learning Representations, (ICLR\u201919). Retrieved from OpenReview.net. https:\/\/openreview.net\/forum?id=r1efr3C9Ym."},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11635"},{"key":"e_1_3_2_29_2","first-page":"3104","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 27: Annual Conference on Neural Information Processing Systems","author":"Sutskever Ilya","year":"2014","unstructured":"Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014. Sequence to sequence learning with neural networks. In Proceedings of the Advances in Neural Information Processing Systems 27: Annual Conference on Neural Information Processing Systems (2014). 3104\u20133112."},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i01.5440"},{"key":"e_1_3_2_31_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 30. Curran Associates, Inc."},{"key":"e_1_3_2_32_2","first-page":"5754","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems (NeurIPS\u201919)","author":"Yang Zhilin","year":"2019","unstructured":"Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019. XLNet: Generalized autoregressive pretraining for language understanding. In Proceedings of the Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems (NeurIPS\u201919). 5754\u20135764."},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394486.3403129"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/3447548.3467401"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394486.3403087"}],"container-title":["ACM Transactions on Knowledge Discovery from Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3516367","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3516367","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:30:21Z","timestamp":1750188621000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3516367"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,7,30]]},"references-count":34,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2022,12,31]]}},"alternative-id":["10.1145\/3516367"],"URL":"https:\/\/doi.org\/10.1145\/3516367","relation":{},"ISSN":["1556-4681","1556-472X"],"issn-type":[{"value":"1556-4681","type":"print"},{"value":"1556-472X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,7,30]]},"assertion":[{"value":"2021-08-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-01-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-07-30","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}