{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,18]],"date-time":"2026-07-18T22:12:56Z","timestamp":1784412776320,"version":"3.55.0"},"reference-count":87,"publisher":"Springer Science and Business Media LLC","issue":"6","license":[{"start":{"date-parts":[[2023,2,15]],"date-time":"2023-02-15T00:00:00Z","timestamp":1676419200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,2,15]],"date-time":"2023-02-15T00:00:00Z","timestamp":1676419200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"University of Oulu including Oulu University Hospital"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Comput Vis"],"published-print":{"date-parts":[[2023,6]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Remote photoplethysmography (rPPG), which aims at measuring heart activities and physiological signals from facial video without any contact, has great potential in many applications (e.g., remote healthcare and affective computing). Recent deep learning approaches focus on mining subtle rPPG clues using convolutional neural networks with limited spatio-temporal receptive fields, which neglect the long-range spatio-temporal perception and interaction for rPPG modeling. In this paper, we propose two end-to-end video transformer based architectures, namely PhysFormer and PhysFormer++, to adaptively aggregate both local and global spatio-temporal features for rPPG representation enhancement. As key modules in PhysFormer, the temporal difference transformers first enhance the quasi-periodic rPPG features with temporal difference guided global attention, and then refine the local spatio-temporal representation against interference. To better exploit the temporal contextual and periodic rPPG clues, we also extend the PhysFormer to the two-pathway SlowFast based PhysFormer++ with temporal difference periodic and cross-attention transformers. Furthermore, we propose the label distribution learning and a curriculum learning inspired dynamic constraint in frequency domain, which provide elaborate supervisions for PhysFormer and PhysFormer++ and alleviate overfitting. Comprehensive experiments are performed on four benchmark datasets to show our superior performance on both intra- and cross-dataset testings. Unlike most transformer networks needed pretraining from large-scale datasets, the proposed PhysFormer family can be easily trained from scratch on rPPG datasets, which makes it promising as a novel transformer baseline for the rPPG community.<\/jats:p>","DOI":"10.1007\/s11263-023-01758-1","type":"journal-article","created":{"date-parts":[[2023,2,15]],"date-time":"2023-02-15T17:13:29Z","timestamp":1676481209000},"page":"1307-1330","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":129,"title":["PhysFormer++: Facial Video-Based Physiological Measurement with SlowFast Temporal Difference Transformer"],"prefix":"10.1007","volume":"131","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-6505-3304","authenticated-orcid":false,"given":"Zitong","family":"Yu","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yuming","family":"Shen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jingang","family":"Shi","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hengshuang","family":"Zhao","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yawen","family":"Cui","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jiehua","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Philip","family":"Torr","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Guoying","family":"Zhao","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2023,2,15]]},"reference":[{"key":"1758_CR1","doi-asserted-by":"crossref","unstructured":"Arnab, A., Dehghani, M., Heigold, G., Sun, C., Lucic, M., & Schmid, C. (2021). Vivit: A video vision transformer. In 2021 IEEE\/CVF international conference on computer vision (ICCV) (pp. 6816\u20136826).","DOI":"10.1109\/ICCV48922.2021.00676"},{"key":"1758_CR2","doi-asserted-by":"crossref","unstructured":"Bengio, Y., Louradour, J., Collobert, R., & Weston, J. (2009). Curriculum learning. In Proceedings of the 26th annual international conference on machine learning (pp. 41\u201348). Association for Computing Machinery.","DOI":"10.1145\/1553374.1553380"},{"key":"1758_CR3","unstructured":"Bertasius, G., Wang, H., & Torresani, L. (2021). Is space\u2013time attention all you need for video understanding? In ICML (Vol.\u00a02, p.\u00a04)."},{"key":"1758_CR4","unstructured":"Bulat, A., P\u00e9rez-R\u00faa, J.-M., Sudhakaran, S., Mart\u00edez, B., & Tzimiropoulos, G. (2021). Advances in neural information processing systems. In M. Ranzato, A. Beygelzimer, Y. Dauphin, P. S. Liang, & J. Wortman Vaughan (Eds.), Space-time mixing attention for video transformer (vol. 34, pp. 19594\u201319607). Curran Associates, Inc."},{"key":"1758_CR5","unstructured":"Cao, J., Li, Y., Zhang, K., & Gool, L.\u00a0V. (2021). Video super-resolution transformer. ArXiv:2106.06847"},{"key":"1758_CR6","doi-asserted-by":"crossref","unstructured":"Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., & Zagoruyko, S. (2020). End-to-end object detection with transformers. ArXiv:2005.12872","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"1758_CR7","doi-asserted-by":"crossref","unstructured":"Carreira, J. & Zisserman, A. (2017). Quo vadis, action recognition? A new model and the kinetics dataset. In 2017 IEEE conference on computer vision and pattern recognition (CVPR) (pp. 4724\u20134733).","DOI":"10.1109\/CVPR.2017.502"},{"key":"1758_CR8","doi-asserted-by":"crossref","unstructured":"Chen, C.-F.\u00a0R., Fan, Q., & Panda, R. (2021a). Crossvit: Cross-attention multi-scale vision transformer for image classification. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 357\u2013366).","DOI":"10.1109\/ICCV48922.2021.00041"},{"key":"1758_CR9","unstructured":"Chen, H., Tang, H., Sebe, N., & Zhao, G. (2021b). Aniformer: Data-driven 3d animation with transformer. ArXiv:2110.10533"},{"key":"1758_CR10","doi-asserted-by":"crossref","unstructured":"Chen, H., Tang, H., Yu, Z., Sebe, N., & Zhao, G. (2021c). Geometry-contrastive transformer for generalized 3d pose transfer. In AAAI conference on artificial intelligence.","DOI":"10.1609\/aaai.v36i1.19901"},{"key":"1758_CR11","doi-asserted-by":"crossref","unstructured":"Chen, W. & McDuff, D. (2018). Deepphys: Video-based physiological measurement using convolutional attention networks. In Proceedings of the European conference on computer vision (ECCV) (pp. 349\u2013365).","DOI":"10.1007\/978-3-030-01216-8_22"},{"issue":"10","key":"1758_CR12","doi-asserted-by":"publisher","first-page":"3600","DOI":"10.1109\/TIM.2018.2879706","volume":"68","author":"X Chen","year":"2018","unstructured":"Chen, X., Cheng, J., Song, R., Liu, Y., Ward, R., & Wang, Z. J. (2018). Video-based heart rate measurement: Recent advances and future prospects. IEEE Transactions on Instrumentation and Measurement, 68(10), 3600\u20133615.","journal-title":"IEEE Transactions on Instrumentation and Measurement"},{"key":"1758_CR13","doi-asserted-by":"crossref","unstructured":"Cho, K., Van\u00a0Merri\u00ebnboer, B., Bahdanau, D., & Bengio, Y. (2014). On the properties of neural machine translation: Encoder\u2013decoder approaches. ArXiv:1409.1259","DOI":"10.3115\/v1\/W14-4012"},{"issue":"10","key":"1758_CR14","doi-asserted-by":"publisher","first-page":"2878","DOI":"10.1109\/TBME.2013.2266196","volume":"60","author":"G De Haan","year":"2013","unstructured":"De Haan, G., & Jeanne, V. (2013). Robust pulse rate from chrominance-based rPPG. IEEE Transactions on Biomedical Engineering, 60(10), 2878\u20132886.","journal-title":"IEEE Transactions on Biomedical Engineering"},{"key":"1758_CR15","doi-asserted-by":"crossref","unstructured":"Ding, M., Lian, X., Yang, L., Wang, P., Jin, X., Lu, Z., & Luo, P. (2021). Hr-nas: Searching efficient high-resolution neural architectures with lightweight transformers. In 2021 IEEE\/CVF conference on computer vision and pattern recognition (CVPR) (pp. 2981\u20132991).","DOI":"10.1109\/CVPR46437.2021.00300"},{"key":"1758_CR16","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., & Gelly, S., Uszkoreit, J. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. ArXiv:2010.11929"},{"key":"1758_CR17","doi-asserted-by":"crossref","unstructured":"Fan, H., Xiong, B., Mangalam, K., Li, Y., Yan, Z., Malik, J., & Feichtenhofer, C. (2021). Multiscale vision transformers. In 2021 IEEE\/CVF international conference on computer vision (ICCV) (pp. 6804\u20136815).","DOI":"10.1109\/ICCV48922.2021.00675"},{"key":"1758_CR18","doi-asserted-by":"crossref","unstructured":"Feichtenhofer, C., Fan, H., Malik, J., & He, K. (2018). Slowfast networks for video recognition. In 2019 IEEE\/CVF international conference on computer vision (ICCV) (pp. 6201\u20136210).","DOI":"10.1109\/ICCV.2019.00630"},{"key":"1758_CR19","doi-asserted-by":"publisher","first-page":"2825","DOI":"10.1109\/TIP.2017.2689998","volume":"26","author":"B-B Gao","year":"2016","unstructured":"Gao, B.-B., Xing, C., Xie, C.-W., Wu, J., & Geng, X. (2016). Deep label distribution learning with label ambiguity. IEEE Transactions on Image Processing, 26, 2825\u20132838.","journal-title":"IEEE Transactions on Image Processing"},{"key":"1758_CR20","doi-asserted-by":"crossref","unstructured":"Gao, B.-B., Zhou, H.-Y., Wu, J., & Geng, X. (2018). Age estimation using expectation of label distribution learning. In International joint conference on artificial intelligence.","DOI":"10.24963\/ijcai.2018\/99"},{"key":"1758_CR21","doi-asserted-by":"publisher","first-page":"2401","DOI":"10.1109\/TPAMI.2013.51","volume":"35","author":"X Geng","year":"2010","unstructured":"Geng, X., Smith-Miles, K., & Zhou, Z.-H. (2010). Facial age estimation by learning from label distributions. IEEE Transactions on Pattern Analysis and Machine Intelligence, 35, 2401\u20132412.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"1758_CR22","doi-asserted-by":"crossref","unstructured":"Gideon, J. & Stent, S. (2021). The way to my heart is through contrastive learning: Remote photoplethysmography from unlabelled video. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 3995\u20134004).","DOI":"10.1109\/ICCV48922.2021.00396"},{"key":"1758_CR23","doi-asserted-by":"crossref","unstructured":"Girdhar, R., Carreira, J., Doersch, C., & Zisserman, A. (2018). Video action transformer network. In 2019 IEEE\/CVF conference on computer vision and pattern recognition (CVPR) (pp. 244\u2013253).","DOI":"10.1109\/CVPR.2019.00033"},{"key":"1758_CR24","doi-asserted-by":"publisher","first-page":"87","DOI":"10.1109\/TPAMI.2022.3152247","volume":"45","author":"K Han","year":"2022","unstructured":"Han, K., Wang, Y., Chen, H., Chen, X., Guo, J., Liu, Z., & Yang, Z. (2022). A survey on vision transformer. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45, 87\u2013110.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"1758_CR25","first-page":"15908","volume":"34","author":"K Han","year":"2021","unstructured":"Han, K., Xiao, A., Wu, E., Guo, J., Xu, C., & Wang, Y. (2021). Transformer in transformer. Advances in Neural Information Processing Systems, 34, 15908\u201315919.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"1758_CR26","unstructured":"Hassani, A., Walton, S., Shah, N., Abuduweili, A., Li, J., & Shi, H. (2021). Escaping the big data paradigm with compact transformers. ArXiv:2104.05704"},{"key":"1758_CR27","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In CVPR.","DOI":"10.1109\/CVPR.2016.90"},{"key":"1758_CR28","doi-asserted-by":"crossref","unstructured":"He, S., Luo, H., Wang, P., Wang, F., Li, H., & Jiang, W. (2021). Transreid: Transformer-based object re-identification. In 2021 IEEE\/CVF international conference on computer vision (ICCV) (pp. 14993\u201315002).","DOI":"10.1109\/ICCV48922.2021.01474"},{"key":"1758_CR29","doi-asserted-by":"crossref","unstructured":"Hsu, G.-S., Ambikapathi, A., & Chen, M.-S. (2017). Deep learning with time-frequency representation for pulse estimation from facial videos. In 2017 IEEE international joint conference on biometrics (IJCB) (pp. 383\u2013389). IEEE.","DOI":"10.1109\/BTAS.2017.8272721"},{"key":"1758_CR30","unstructured":"Huang, C.-Z.\u00a0A., Vaswani, A., Uszkoreit, J., Simon, I., Hawthorne, C., Shazeer, N.\u00a0M., Dai, A.\u00a0M., Hoffman, M.\u00a0D., Dinculescu, M., & Eck, D. (2019). Music transformer: Generating music with long-term structure. In International conference on learning representations."},{"key":"1758_CR31","doi-asserted-by":"crossref","unstructured":"Kazakos, E., Nagrani, A., Zisserman, A., & Damen, D. (2021). Slow-fast auditory streams for audio recognition. In ICASSP 2021\u20142021 IEEE international conference on acoustics, speech and signal processing (ICASSP) (pp. 855\u2013859).","DOI":"10.1109\/ICASSP39728.2021.9413376"},{"key":"1758_CR32","unstructured":"Khan, S., Naseer, M., Hayat, M., Zamir, S.\u00a0W., Khan, F.\u00a0S., & Shah, M. (2021). Transformers in vision: A survey. ArXiv:2101.01169"},{"key":"1758_CR33","doi-asserted-by":"crossref","unstructured":"Lam, A. & Kuno, Y. (2015). Robust heart rate measurement from video using select random patches. In Proceedings of the IEEE international conference on computer vision (pp. 3640\u20133648).","DOI":"10.1109\/ICCV.2015.415"},{"key":"1758_CR34","doi-asserted-by":"crossref","unstructured":"Lee, E., Chen, E., & Lee, C.-Y. (2020). Meta-rppg: Remote heart rate estimation using a transductive meta-learner. In European conference on computer vision.","DOI":"10.1007\/978-3-030-58583-9_24"},{"key":"1758_CR35","doi-asserted-by":"crossref","unstructured":"Li, X., Alikhani, I., Shi, J., Sepp\u00e4nen, T., Junttila, J.\u00a0M., Majamaa-Voltti, K., Tulppo, M.\u00a0P., & Zhao, G. (2018). The obf database: A large face video database for remote physiological signal measurement and atrial fibrillation detection. In 2018 13th IEEE international conference on automatic face & gesture recognition (FG 2018) (pp. 242\u2013249).","DOI":"10.1109\/FG.2018.00043"},{"key":"1758_CR36","doi-asserted-by":"crossref","unstructured":"Li, X., Chen, J., Zhao, G., & Pietikainen, M. (2014). Remote heart rate measurement from face videos under realistic situations. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 4264\u20134271).","DOI":"10.1109\/CVPR.2014.543"},{"key":"1758_CR37","doi-asserted-by":"crossref","unstructured":"Lin, J., Gan, C., & Han, S. (2018). Tsm: Temporal shift module for efficient video understanding. In 2019 IEEE\/CVF international conference on computer vision (ICCV) (pp. 7082\u20137092).","DOI":"10.1109\/ICCV.2019.00718"},{"key":"1758_CR38","doi-asserted-by":"crossref","unstructured":"Lin, T., Wang, Y., Liu, X., & Qiu, X. (2022). A survey of transformers. AI Open.","DOI":"10.1016\/j.aiopen.2022.10.001"},{"key":"1758_CR39","unstructured":"Lin, Y., Zhang, T., Sun, P., Li, Z., & Zhou, S. (2021). Fq-vit: Fully quantized vision transformer without retraining. ArXiv:2111.13824"},{"key":"1758_CR40","doi-asserted-by":"crossref","unstructured":"Liu, R., Deng, H., Huang, Y., Shi, X., Lu, L., Sun, W., Wang, X., Dai, J., & Li, H. (2021a). Fuseformer: Fusing fine-grained information in transformers for video inpainting. In 2021 IEEE\/CVF international conference on computer vision (ICCV) (pp. 14020\u201314029).","DOI":"10.1109\/ICCV48922.2021.01378"},{"key":"1758_CR41","first-page":"19400","volume":"33","author":"X Liu","year":"2020","unstructured":"Liu, X., Fromm, J., Patel, S., & McDuff, D. (2020). Multi-task temporal shift attention networks for on-device contactless vitals measurement. Advances in Neural Information Processing Systems, 33, 19400\u201319411.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"1758_CR42","doi-asserted-by":"crossref","unstructured":"Liu, X., Hill, B., Jiang, Z., Patel, S., & McDuff, D. (2023). Efficientphys: Enabling simple, fast and accurate camera-based cardiac measurement. In Proceedings of the IEEE\/CVF winter conference on applications of computer vision (pp. 5008\u20135017).","DOI":"10.1109\/WACV56688.2023.00498"},{"key":"1758_CR43","unstructured":"Liu, X., Patel, S., & McDuff, D. (2021b). Camera-based physiological sensing: Challenges and future directions. ArXiv:2110.13362"},{"key":"1758_CR44","doi-asserted-by":"crossref","unstructured":"Liu, X., Wang, Q., Hu, Y., Tang, X., Zhang, S., Bai, S., & Bai, X. (2021c). End-to-end temporal action detection with transformer. IEEE Transactions on Image Processing, 31, 5427\u20135441.","DOI":"10.1109\/TIP.2022.3195321"},{"key":"1758_CR45","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., & Guo, B. (2021d). Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 10012\u201310022).","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"1758_CR46","doi-asserted-by":"crossref","unstructured":"Liu, Z., Ning, J., Cao, Y., Wei, Y., Zhang, Z., Lin, S., & Hu, H. (2021e). Video swin transformer. In 2022 IEEE\/CVF conference on computer vision and pattern recognition (CVPR) (pp. 3192\u20133201).","DOI":"10.1109\/CVPR52688.2022.00320"},{"issue":"1","key":"1758_CR47","doi-asserted-by":"publisher","first-page":"33","DOI":"10.1016\/j.vrih.2020.10.002","volume":"3","author":"H Lu","year":"2021","unstructured":"Lu, H., & Han, H. (2021). Nas-hr: Neural architecture search for heart rate estimation from face videos. Virtual Reality & Intelligent Hardware, 3(1), 33\u201342.","journal-title":"Virtual Reality & Intelligent Hardware"},{"key":"1758_CR48","doi-asserted-by":"crossref","unstructured":"Lu, H., Han, H., & Zhou, S.\u00a0K. (2021). Dual-gan: Joint bvp and noise modeling for remote physiological measurement. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 12404\u201312413).","DOI":"10.1109\/CVPR46437.2021.01222"},{"key":"1758_CR49","doi-asserted-by":"crossref","unstructured":"Magdalena\u00a0Nowara, E., Marks, T.\u00a0K., Mansour, H., & Veeraraghavan, A. (2018). Sparseppg: Towards driver monitoring using camera-based vital signs estimation in near-infrared. In Proceedings of the IEEE conference on computer vision and pattern recognition workshops (pp. 1272\u20131281).","DOI":"10.1109\/CVPRW.2018.00174"},{"key":"1758_CR50","doi-asserted-by":"crossref","unstructured":"Neimark, D., Bar, O., Zohar, M., & Asselmann, D. (2021). Video transformer network. In 2021 IEEE\/CVF international conference on computer vision workshops (ICCVW) (pp. 3156\u20133165).","DOI":"10.1109\/ICCVW54120.2021.00355"},{"key":"1758_CR51","doi-asserted-by":"crossref","unstructured":"Niu, X., Han, H., Shan, S., & Chen, X. (2017). Continuous heart rate measurement from face: A robust rppg approach with distribution learning. In 2017 IEEE international joint conference on biometrics (IJCB) (pp. 642\u2013650).","DOI":"10.1109\/BTAS.2017.8272752"},{"key":"1758_CR52","doi-asserted-by":"crossref","unstructured":"Niu, X., Han, H., Shan, S., & Chen, X. (2018). Synrhythm: Learning a deep heart rate estimator from general to specific. In 2018 24th international conference on pattern recognition (ICPR) (pp. 3580\u20133585). IEEE.","DOI":"10.1109\/ICPR.2018.8546321"},{"key":"1758_CR53","doi-asserted-by":"crossref","unstructured":"Niu, X., Shan, S., Han, H., & Chen, X. (2019a). Rhythmnet: End-to-end heart rate estimation from face via spatial-temporal representation. IEEE Transactions on Image Processing, 29, 2409\u20132423.","DOI":"10.1109\/TIP.2019.2947204"},{"key":"1758_CR54","doi-asserted-by":"crossref","unstructured":"Niu, X., Yu, Z., Han, H., Li, X., Shan, S., & Zhao, G. (2020). Video-based remote physiological measurement via cross-verified feature disentangling. In ECCV (pp. 295\u2013310). Springer.","DOI":"10.1007\/978-3-030-58536-5_18"},{"key":"1758_CR55","doi-asserted-by":"crossref","unstructured":"Niu, X., Zhao, X., Han, H., Das, A., Dantcheva, A., Shan, S., & Chen, X. (2019b). Robust remote heart rate estimation from face utilizing spatial-temporal attention. In 2019 14th IEEE international conference on automatic face & gesture recognition (FG 2019) (pp. 1\u20138). IEEE.","DOI":"10.1109\/FG.2019.8756554"},{"key":"1758_CR56","doi-asserted-by":"crossref","unstructured":"Nowara, E.\u00a0M., McDuff, D., & Veeraraghavan, A. (2021). The benefit of distraction: Denoising camera-based physiological measurements using inverse attention. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 4955\u20134964).","DOI":"10.1109\/ICCV48922.2021.00491"},{"key":"1758_CR57","doi-asserted-by":"crossref","unstructured":"Poh, M.-Z., McDuff, D. J., & Picard, R. W. (2010a). Advancements in noncontact, multiparameter physiological measurements using a webcam. IEEE Transactions on Biomedical Engineering, 58(1), 7\u201311.","DOI":"10.1109\/TBME.2010.2086456"},{"key":"1758_CR58","doi-asserted-by":"crossref","unstructured":"Poh, M.-Z., McDuff, D. J., & Picard, R. W. (2010b). Non-contact, automated cardiac pulse measurements using video imaging and blind source separation. Optics Express, 18(10), 10762\u201310774.","DOI":"10.1364\/OE.18.010762"},{"key":"1758_CR59","unstructured":"Qin, H., Ding, Y., Zhang, M., Yan, Q., Liu, A., Dang, Q., Liu, Z., & Liu, X. (2022). Bibert: Accurate fully binarized bert. In ICLR."},{"issue":"7","key":"1758_CR60","doi-asserted-by":"publisher","first-page":"1778","DOI":"10.1109\/TMM.2018.2883866","volume":"21","author":"Y Qiu","year":"2018","unstructured":"Qiu, Y., Liu, Y., Arteaga-Falconi, J., Dong, H., & El Saddik, A. (2018). EVM-CNN: Real-time contactless heart rate estimation from facial video. IEEE Transactions on Multimedia, 21(7), 1778\u20131787.","journal-title":"IEEE Transactions on Multimedia"},{"key":"1758_CR61","doi-asserted-by":"crossref","unstructured":"Revanur, A., Dasari, A., Tucker, C.\u00a0S., & Jeni, L.\u00a0A. (2022). Instantaneous physiological estimation using video transformers. ArXiv:2202.12368","DOI":"10.1007\/978-3-031-14771-5_22"},{"key":"1758_CR62","doi-asserted-by":"crossref","unstructured":"Shaw, P., Uszkoreit, J., & Vaswani, A. (2018). Self-attention with relative position representations. In North American chapter of the Association for Computational Linguistics.","DOI":"10.18653\/v1\/N18-2074"},{"key":"1758_CR63","doi-asserted-by":"publisher","first-page":"42","DOI":"10.1109\/T-AFFC.2011.25","volume":"3","author":"M Soleymani","year":"2012","unstructured":"Soleymani, M., Lichtenauer, J., Pun, T., & Pantic, M. (2012). A multimodal database for affect recognition and implicit tagging. IEEE Transactions on Affective Computing, 3, 42\u201355.","journal-title":"IEEE Transactions on Affective Computing"},{"key":"1758_CR64","unstructured":"\u0160petl\u00edk, R., Franc, V., & Matas, J. (2018). Visual heart rate estimation with convolutional neural network. In Proceedings of the British machine vision conference, Newcastle, UK (pp. 3\u20136)."},{"key":"1758_CR65","unstructured":"Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., & J\u00e9gou, H. (2021). Training data-efficient image transformers & distillation through attention. In International conference on machine learning (pp. 10347\u201310357). PMLR."},{"key":"1758_CR66","doi-asserted-by":"crossref","unstructured":"Tulyakov, S., Alameda-Pineda, X., Ricci, E., Yin, L., Cohn, J.\u00a0F., & Sebe, N. (2016). Self-adaptive matrix completion for heart rate estimation from face videos under realistic conditions. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 2396\u20132404).","DOI":"10.1109\/CVPR.2016.263"},{"key":"1758_CR67","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.\u00a0N., Kaiser, \u0141., & Polosukhin, I. (2017). Attention is all you need. In Advances in neural information processing systems, 30."},{"issue":"26","key":"1758_CR68","doi-asserted-by":"publisher","first-page":"21434","DOI":"10.1364\/OE.16.021434","volume":"16","author":"W Verkruysse","year":"2008","unstructured":"Verkruysse, W., Svaasand, L. O., & Nelson, J. S. (2008). Remote plethysmographic imaging using ambient light. Optics Express, 16(26), 21434\u201321445.","journal-title":"Optics Express"},{"key":"1758_CR69","unstructured":"Wang, L., Yang, H., Wu, W., Yao, H., & Huang, H. (2021a). Temporal action proposal generation with transformers. ArXiv:2105.12043"},{"issue":"7","key":"1758_CR70","doi-asserted-by":"publisher","first-page":"1479","DOI":"10.1109\/TBME.2016.2609282","volume":"64","author":"W Wang","year":"2016","unstructured":"Wang, W., Den Brinker, A. C., Stuijk, S., & De Haan, G. (2016). Algorithmic principles of remote PPG. IEEE Transactions on Biomedical Engineering, 64(7), 1479\u20131491.","journal-title":"IEEE Transactions on Biomedical Engineering"},{"key":"1758_CR71","doi-asserted-by":"crossref","unstructured":"Wang, W., Xie, E., Li, X., Fan, D.-P., Song, K., Liang, D., Lu, T., Luo, P., & Shao, L. (2021b). Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 568\u2013578).","DOI":"10.1109\/ICCV48922.2021.00061"},{"key":"1758_CR72","doi-asserted-by":"crossref","unstructured":"Wu, K., Peng, H., Chen, M., Fu, J., & Chao, H. (2021). Rethinking and improving relative position encoding for vision transformer. In 2021 IEEE\/CVF international conference on computer vision (ICCV) (pp. 10013\u201310021).","DOI":"10.1109\/ICCV48922.2021.00988"},{"key":"1758_CR73","unstructured":"Xiao, T., Singh, M., Mintun, E., Darrell, T., Doll\u00e1r, P., & Girshick, R.\u00a0B. (2021). Early convolutions help transformers see better. In Neural information processing systems."},{"key":"1758_CR74","unstructured":"Xu, M., Xiong, Y., Chen, H., Li, X., Xia, W., Tu, Z., & Soatto, S. (2021). Long short-term transformer for online action detection. ArXiv:2107.03377"},{"key":"1758_CR75","doi-asserted-by":"crossref","unstructured":"Yu, Z., Li, X., Niu, X., Shi, J., & Zhao, G. (2020). Autohr: A strong end-to-end baseline for remote heart rate measurement with neural searching. IEEE Signal Processing Letters, 27, 1245\u20131249.","DOI":"10.1109\/LSP.2020.3007086"},{"key":"1758_CR76","doi-asserted-by":"publisher","first-page":"1290","DOI":"10.1109\/LSP.2021.3089908","volume":"28","author":"Z Yu","year":"2021","unstructured":"Yu, Z., Li, X., Wang, P., & Zhao, G. (2021). Transrppg: Remote photoplethysmography transformer for 3d mask face presentation attack detection. IEEE Signal Processing Letters, 28, 1290\u20131294.","journal-title":"IEEE Signal Processing Letters"},{"key":"1758_CR77","unstructured":"Yu, Z., Li, X., & Zhao, G. (2019a). Remote photoplethysmograph signal measurement from facial videos using spatio-temporal networks. In British machine vision conference (pp. 277\u2013289)."},{"issue":"6","key":"1758_CR78","doi-asserted-by":"publisher","first-page":"50","DOI":"10.1109\/MSP.2021.3106285","volume":"38","author":"Z Yu","year":"2021","unstructured":"Yu, Z., Li, X., & Zhao, G. (2021). Facial-video-based physiological signal measurement: Recent advances and affective applications. IEEE Signal Processing Magazine, 38(6), 50\u201358.","journal-title":"IEEE Signal Processing Magazine"},{"key":"1758_CR79","doi-asserted-by":"crossref","unstructured":"Yu, Z., Peng, W., Li, X., Hong, X., & Zhao, G. (2019b). Remote heart rate measurement from highly compressed facial videos: an end-to-end deep learning solution with video enhancement. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 151\u2013160).","DOI":"10.1109\/ICCV.2019.00024"},{"key":"1758_CR80","doi-asserted-by":"crossref","unstructured":"Yu, Z., Qin, Y., Li, X., Zhao, C., Lei, Z., & Zhao, G. (2021). Deep learning for face anti-spoofing: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence.","DOI":"10.1109\/TPAMI.2022.3215850"},{"key":"1758_CR81","doi-asserted-by":"crossref","unstructured":"Yu, Z., Shen, Y., Shi, J., Zhao, H., Torr, P.\u00a0H., & Zhao, G. (2022). Physformer: facial video-based physiological measurement with temporal difference transformer. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 4186\u20134196).","DOI":"10.1109\/CVPR52688.2022.00415"},{"key":"1758_CR82","doi-asserted-by":"publisher","unstructured":"Yu, Z., Zhou, B., Wan, J., Wang, P., Chen, H., Liu, X., Li, S. Z., & Zhao, G. (2022). Deep learning for face anti-spoofing: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence. https:\/\/doi.org\/10.1109\/TPAMI.2022.3215850.","DOI":"10.1109\/TPAMI.2022.3215850"},{"key":"1758_CR83","doi-asserted-by":"crossref","unstructured":"Yuan, L., Chen, Y., Wang, T., Yu, W., Shi, Y., Jiang, Z.-H., Tay, F.\u00a0E., Feng, J., & Yan, S. (2021). Tokens-to-token vit: Training vision transformers from scratch on imagenet. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 558\u2013567).","DOI":"10.1109\/ICCV48922.2021.00060"},{"key":"1758_CR84","doi-asserted-by":"crossref","unstructured":"Zeng, Y., Fu, J., & Chao, H. (2020). Learning joint spatial-temporal transformations for video inpainting. ArXiv:2007.10247","DOI":"10.1007\/978-3-030-58517-4_31"},{"key":"1758_CR85","first-page":"1499","volume":"23","author":"K Zhang","year":"2016","unstructured":"Zhang, K., Zhang, Z., Li, Z., & Qiao, Y. (2016). Joint face detection and alignment using multitask cascaded convolutional networks. IEEE SPL, 23, 1499\u20131503.","journal-title":"IEEE SPL"},{"key":"1758_CR86","doi-asserted-by":"crossref","unstructured":"Zhao, J., Li, X., Liu, C., Shuai, B., Chen, H., Snoek, C. G.\u00a0M., & Tighe, J. (2021). Tuber: Tube-transformer for action detection. ArXiv:2104.00969","DOI":"10.1109\/CVPR52688.2022.01323"},{"key":"1758_CR87","doi-asserted-by":"crossref","unstructured":"Zheng, S., Lu, J., Zhao, H., Zhu, X., Luo, Z., Wang, Y., Fu, Y., Feng, J., Xiang, T., Torr, P. H.\u00a0S., & Zhang, L. (2020). Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In 2021 IEEE\/CVF conference on computer vision and pattern recognition (CVPR) (pp. 6877\u20136886).","DOI":"10.1109\/CVPR46437.2021.00681"}],"container-title":["International Journal of Computer Vision"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-023-01758-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11263-023-01758-1\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-023-01758-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,4,18]],"date-time":"2023-04-18T05:05:35Z","timestamp":1681794335000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11263-023-01758-1"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,2,15]]},"references-count":87,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2023,6]]}},"alternative-id":["1758"],"URL":"https:\/\/doi.org\/10.1007\/s11263-023-01758-1","relation":{},"ISSN":["0920-5691","1573-1405"],"issn-type":[{"value":"0920-5691","type":"print"},{"value":"1573-1405","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,2,15]]},"assertion":[{"value":"27 April 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"6 January 2023","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 February 2023","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}