{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,23]],"date-time":"2026-04-23T01:25:46Z","timestamp":1776907546175,"version":"3.51.2"},"reference-count":48,"publisher":"Association for Computing Machinery (ACM)","issue":"1s","license":[{"start":{"date-parts":[[2023,1,23]],"date-time":"2023-01-23T00:00:00Z","timestamp":1674432000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"CAPES - Brazil","award":["88881.188744\/2018-01"],"award-info":[{"award-number":["88881.188744\/2018-01"]}]},{"name":"Petrobras\/Fundunesp","award":["2662\/2017"],"award-info":[{"award-number":["2662\/2017"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2023,2,28]]},"abstract":"<jats:p>In this article, we propose a hybrid framework for cross-resolution 3D face recognition which utilizes a Streamed Attention Network (SAN) that combines handcrafted features with Convolutional Neural Networks (CNNs). It consists of two main stages: first, we process the depth images to extract low-level surface descriptors and derive the corresponding Descriptor Images (DIs), represented as four-channel images. To build the DIs, we propose a variation of the 3D Local Binary Pattern (3DLBP) operator that encodes depth differences using a sigmoid function. Then, we design a CNN that learns from these DIs. The peculiarity of our solution consists in processing each channel of the input image separately, and fusing the contribution of each channel by means of both self- and cross-attention mechanisms. This strategy showed two main advantages over the direct application of Deep-CNN to depth images of the face; on the one hand, the DIs can reduce the diversity between high- and low-resolution data by encoding surface properties that are robust to resolution differences. On the other, it allows a better exploitation of the richer information provided by low-level features, resulting in improved recognition. We evaluated the proposed architecture in a challenging cross-dataset, cross-resolution scenario. To this aim, we first train the network on scanner-resolution 3D data. Next, we utilize the pre-trained network as feature extractor on low-resolution data, where the output of the last fully connected layer is used as face descriptor. Other than standard benchmarks, we also perform experiments on a newly collected dataset of paired high- and low-resolution 3D faces. We use the high-resolution data as gallery, while low-resolution faces are used as probe, allowing us to assess the real gap existing between these two types of data. Extensive experiments on low-resolution 3D face benchmarks show promising results with respect to state-of-the-art methods.<\/jats:p>","DOI":"10.1145\/3527158","type":"journal-article","created":{"date-parts":[[2022,3,25]],"date-time":"2022-03-25T13:08:35Z","timestamp":1648213715000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":10,"title":["Learning Streamed Attention Network from Descriptor Images for Cross-Resolution 3D Face Recognition"],"prefix":"10.1145","volume":"19","author":[{"given":"Jo\u00e3o Baptista Cardia","family":"Neto","sequence":"first","affiliation":[{"name":"S\u00e3o Paulo State Technological College (FATEC), S\u00e3o Paulo, Brazil"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Claudio","family":"Ferrari","sequence":"additional","affiliation":[{"name":"Department of Architecture and Engineering, University of Parma and University of Florence, Firenze FI, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Aparecido Nilceu","family":"Marana","sequence":"additional","affiliation":[{"name":"RECOGNA Laboratory, S\u00e3o Paulo State University (UNESP), S\u00e3o Paulo, Brazil"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Stefano","family":"Berretti","sequence":"additional","affiliation":[{"name":"MICC, University of Florence, Firenze FI, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Alberto","family":"Del Bimbo","sequence":"additional","affiliation":[{"name":"MICC, University of Florence, Firenze FI, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,1,23]]},"reference":[{"issue":"4","key":"e_1_3_2_2_1","doi-asserted-by":"crossref","first-page":"831","DOI":"10.1007\/s00371-020-01833-5","article-title":"Efficient object tracking using hierarchical convolutional features model and correlation filters","volume":"37","author":"Abbass Mohammed Y.","year":"2021","unstructured":"Mohammed Y. Abbass, Ki-Chul Kwon, Nam Kim, Safey A. Abdelwahab, Fathi E. Abd El-Samie, and Ashraf A. M. Khalaf. 2021. Efficient object tracking using hierarchical convolutional features model and correlation filters. The Visual Computer 37, 4 (2021), 831\u2013842.","journal-title":"The Visual Computer"},{"key":"e_1_3_2_3_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00521-018-3649-0"},{"key":"e_1_3_2_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2010.43"},{"key":"e_1_3_2_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/34.121791"},{"key":"e_1_3_2_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIFS.2016.2601059"},{"key":"e_1_3_2_7_1","volume-title":"3D Face Recognition Using Kinect","author":"Neto J. B. Cardia","year":"2014","unstructured":"J. B. Cardia Neto and A. N. Marana. 2014. 3D Face Recognition Using Kinect. Master\u2019s thesis. S\u00e3o Paulo State University (UNESP), Bauru SP 17033-360, Brazil."},{"key":"e_1_3_2_8_1","doi-asserted-by":"crossref","first-page":"135","DOI":"10.1007\/978-3-319-75193-1_17","volume-title":"Progress in Pattern Recognition, Image Analysis, Computer Vision, and Applications (CIARP)","author":"Neto J. B. Cardia","year":"2018","unstructured":"J. B. Cardia Neto and A. N. Marana. 2018. Utilizing deep learning and 3DLBP for 3D face recognition. In Progress in Pattern Recognition, Image Analysis, Computer Vision, and Applications (CIARP). 135\u2013142."},{"key":"e_1_3_2_9_1","volume-title":"Eurographics Workshop on 3D Object Retrieval (3DOR\u201919)","author":"Neto J. B. Cardia","year":"2019","unstructured":"J. B. Cardia Neto, A. N. Marana, C. Ferrari, S. Berretti, and A. Del Bimbo. 2019. Depth based face recognition by learning from 3D-LBP images. In Eurographics Workshop on 3D Object Retrieval (3DOR\u201919)."},{"key":"e_1_3_2_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/tpami.2013.48"},{"key":"e_1_3_2_11_1","first-page":"1","volume-title":"International Conference of the BIOSIG Special Interest Group","author":"Drosou A.","year":"2013","unstructured":"A. Drosou, P. Moschonas, and D. Tzovaras. 2013. Robust 3D face recognition from low resolution images. In International Conference of the BIOSIG Special Interest Group. 1\u20138."},{"key":"e_1_3_2_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIFS.2007.916287"},{"key":"e_1_3_2_13_1","article-title":"A sparse and locally coherent morphable face model for dense semantic correspondence across heterogeneous 3D faces","author":"Ferrari Claudio","year":"2021","unstructured":"Claudio Ferrari, Stefano Berretti, Pietro Pala, and Alberto Del Bimbo. 2021. A sparse and locally coherent morphable face model for dense semantic correspondence across heterogeneous 3D faces. IEEE Transactions on Pattern Analysis and Machine Intelligence (to appear).","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_2_14_1","article-title":"The MICC-3D face dataset","author":"Ferrari Claudio","year":"2022","unstructured":"Claudio Ferrari, Stefano Berretti, Pietro Pala, and Alberto Del Bimbo. 2022. The MICC-3D face dataset. Sensors (to appear).","journal-title":"Sensors"},{"key":"e_1_3_2_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2018.2861359"},{"key":"e_1_3_2_16_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2019.05.002"},{"key":"e_1_3_2_17_1","first-page":"1896","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201918)","author":"Gilani Syed Zulqarnain","year":"2018","unstructured":"Syed Zulqarnain Gilani and Ajmal Mian. 2018. Learning from millions of 3D scans for large-scale 3D face recognition. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201918). 1896\u20131905."},{"key":"e_1_3_2_18_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2017.04.013"},{"key":"e_1_3_2_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/BTAS.2013.6712717"},{"key":"e_1_3_2_20_1","unstructured":"Kaiming He Xiangyu Zhang Shaoqing Ren and Jian Sun. 2015. Deep residual learning for image recognition. arxiv:1512.03385 [cs.CV]. https:\/\/arxiv.org\/abs\/1512.03385."},{"key":"e_1_3_2_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2961900"},{"key":"e_1_3_2_22_1","first-page":"1995","volume-title":"European Signal Processing Conference (EUSIPCO\u201912)","author":"Hernandez M.","year":"2012","unstructured":"M. Hernandez, J. Choi, and G. Medioni. 2012. Laser scan quality 3-D face modeling using a low-cost depth camera. In European Signal Processing Conference (EUSIPCO\u201912). 1995\u20131999."},{"key":"e_1_3_2_23_1","doi-asserted-by":"crossref","unstructured":"Andrew Howard Mark Sandler Grace Chu Liang-Chieh Chen Bo Chen Mingxing Tan Weijun Wang Yukun Zhu Ruoming Pang Vijay Vasudevan Quoc V. Le and Hartwig Adam. 2019. Searching for MobileNetV3. arXiv:1905.02244 [cs.CV]. https:\/\/arxiv.org\/abs\/1905.02244.","DOI":"10.1109\/ICCV.2019.00140"},{"key":"e_1_3_2_24_1","first-page":"0","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201919) Workshops","author":"Hu Zhenguo","year":"2019","unstructured":"Zhenguo Hu, Qijun Zhao, and Feng Liu. 2019. Revisiting depth-based face recognition from a quality perspective. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201919) Workshops. 0\u20130."},{"key":"e_1_3_2_25_1","doi-asserted-by":"publisher","DOI":"10.5244\/C.20.90"},{"key":"e_1_3_2_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.291928"},{"key":"e_1_3_2_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/BTAS.2017.8272691"},{"key":"e_1_3_2_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/WACV.2013.6475017"},{"key":"e_1_3_2_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICIP40778.2020.9190677"},{"key":"e_1_3_2_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/LSP.2016.2553784"},{"key":"e_1_3_2_31_1","first-page":"1739","volume-title":"International Conference on Pattern Recognition (ICPR\u201912)","author":"Min R.","year":"2012","unstructured":"R. Min, J. Choi, G. Medioni, and J. Dugelay. 2012. Real-time 3D face identification from a depth camera. In International Conference on Pattern Recognition (ICPR\u201912). 1739\u20131742."},{"key":"e_1_3_2_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSMC.2014.2331215"},{"key":"e_1_3_2_33_1","first-page":"5773","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Mu Guodong","year":"2019","unstructured":"Guodong Mu, Di Huang, Guosheng Hu, Jia Sun, and Yunhong Wang. 2019. Led3D: A lightweight and efficient deep approach to recognizing low-quality 3D faces. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 5773\u20135782."},{"key":"e_1_3_2_34_1","doi-asserted-by":"publisher","DOI":"10.1016\/0031-3203(95)00067-4"},{"key":"e_1_3_2_35_1","first-page":"1","volume-title":"2017 14th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS\u201917)","author":"Parchami Mostafa","year":"2017","unstructured":"Mostafa Parchami, Saman Bashbaghi, and Eric Granger. 2017. CNNs with cross-correlation matching for face recognition in video surveillance using a single training sample per person. In 2017 14th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS\u201917). IEEE, 1\u20136."},{"key":"e_1_3_2_36_1","first-page":"6","volume-title":"British Machine Vision Conference (BMVC\u201915)","author":"Parkhi Omkar M.","year":"2015","unstructured":"Omkar M. Parkhi, Andrea Vedaldi, and Andrew Zisserman. 2015. Deep face recognition. In British Machine Vision Conference (BMVC\u201915). 6."},{"key":"e_1_3_2_37_1","first-page":"947","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201905)","volume":"1","author":"Phillips P. Jonathon","year":"2005","unstructured":"P. Jonathon Phillips, Patrick J. Flynn, Todd Scruggs, Kevin W. Bowyer, Jin Chang, Kevin Hoffman, Joe Marques, Jaesik Min, and William Worek. 2005. Overview of the face recognition grand challenge. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201905), Vol. 1. 947\u2013954."},{"key":"e_1_3_2_38_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-89991-4_6"},{"key":"e_1_3_2_39_1","first-page":"815","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201915)","author":"Schroff F.","year":"2015","unstructured":"F. Schroff, D. Kalenichenko, and J. Philbin. 2015. FaceNet: A unified embedding for face recognition and clustering. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201915). 815\u2013823."},{"key":"e_1_3_2_40_1","first-page":"592","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201918) Workshops","author":"Singh Maneet","year":"2018","unstructured":"Maneet Singh, Shruti Nagpal, Mayank Vatsa, Richa Singh, and Angshul Majumdar. 2018. Identity aware synthesis for cross resolution face recognition. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201918) Workshops. 592\u201359209."},{"key":"e_1_3_2_41_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-011-0426-2"},{"key":"e_1_3_2_42_1","first-page":"5998","volume-title":"Advances in Neural Information Processing Systems","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems. 5998\u20136008."},{"key":"e_1_3_2_43_1","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201920)","author":"Wang Heng","year":"2020","unstructured":"Heng Wang, Du Tran, Lorenzo Torresani, and Matt Feiszli. 2020a. Video modeling with correlation networks. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201920)."},{"key":"e_1_3_2_44_1","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201920)","author":"Wang Qiangchang","year":"2020","unstructured":"Qiangchang Wang, Tianyi Wu, He Zheng, and Guodong Guo. 2020b. Hierarchical pyramid diverse attention networks for face recognition. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201920)."},{"key":"e_1_3_2_45_1","first-page":"3876","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201921)","author":"Wang Qiang","year":"2021","unstructured":"Qiang Wang, Yun Zheng, Pan Pan, and Yinghui Xu. 2021. Multiple object tracking with correlation learning. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201921). 3876\u20133886."},{"key":"e_1_3_2_46_1","first-page":"141","volume-title":"International Symposium on Benchmarking, Measuring and Optimization","author":"Xiong Xingwang","year":"2019","unstructured":"Xingwang Xiong, Xu Wen, and Cheng Huang. 2019. Improving RGB-D face recognition via transfer learning from a pretrained 2D network. In International Symposium on Benchmarking, Measuring and Optimization. Springer, 141\u2013148."},{"key":"e_1_3_2_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2017.2765830"},{"key":"e_1_3_2_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00676"},{"key":"e_1_3_2_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.319"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3527158","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3527158","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:51:01Z","timestamp":1750182661000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3527158"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,1,23]]},"references-count":48,"journal-issue":{"issue":"1s","published-print":{"date-parts":[[2023,2,28]]}},"alternative-id":["10.1145\/3527158"],"URL":"https:\/\/doi.org\/10.1145\/3527158","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,1,23]]},"assertion":[{"value":"2021-09-22","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-03-14","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-01-23","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}