{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,26]],"date-time":"2025-10-26T14:56:40Z","timestamp":1761490600933,"version":"build-2065373602"},"reference-count":30,"publisher":"MDPI AG","issue":"8","license":[{"start":{"date-parts":[[2018,8,14]],"date-time":"2018-08-14T00:00:00Z","timestamp":1534204800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>We present a multi-task learning-based convolutional neural network (MTL-CNN) able to estimate multiple tags describing face images simultaneously. In total, the model is able to estimate up to 74 different face attributes belonging to three distinct recognition tasks: age group, gender and visual attributes (such as hair color, face shape and the presence of makeup). The proposed model shares all the CNN\u2019s parameters among tasks and deals with task-specific estimation through the introduction of two components: (i) a gating mechanism to control activations\u2019 sharing and to adaptively route them across different face attributes; (ii) a module to post-process the predictions in order to take into account the correlation among face attributes. The model is trained by fusing multiple databases for increasing the number of face attributes that can be estimated and using a center loss for disentangling representations among face attributes in the embedding space. Extensive experiments validate the effectiveness of the proposed approach.<\/jats:p>","DOI":"10.3390\/s18082666","type":"journal-article","created":{"date-parts":[[2018,8,14]],"date-time":"2018-08-14T10:31:16Z","timestamp":1534242676000},"page":"2666","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":9,"title":["Fine-Grained Face Annotation Using Deep Multi-Task CNN"],"prefix":"10.3390","volume":"18","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5925-2646","authenticated-orcid":false,"given":"Luigi","family":"Celona","sequence":"first","affiliation":[{"name":"Department of Informatics, Systems and Communication, University of Milano-Bicocca, viale Sarca, 336 Milano, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7070-1545","authenticated-orcid":false,"given":"Simone","family":"Bianco","sequence":"additional","affiliation":[{"name":"Department of Informatics, Systems and Communication, University of Milano-Bicocca, viale Sarca, 336 Milano, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7461-1451","authenticated-orcid":false,"given":"Raimondo","family":"Schettini","sequence":"additional","affiliation":[{"name":"Department of Informatics, Systems and Communication, University of Milano-Bicocca, viale Sarca, 336 Milano, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2018,8,14]]},"reference":[{"key":"ref_1","first-page":"6","article-title":"Deep Face Recognition","volume":"1","author":"Parkhi","year":"2015","journal-title":"BMVC"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"36","DOI":"10.1016\/j.patrec.2017.03.006","article-title":"Large Age-Gap Face Verification by Feature Injection in Deep Networks","volume":"90","author":"Bianco","year":"2017","journal-title":"Pattern Recognit. Lett."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"063002","DOI":"10.1117\/1.JEI.25.6.063002","article-title":"Robust smile detection using convolutional neural networks","volume":"25","author":"Bianco","year":"2016","journal-title":"J. Electron. Imaging"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Mavani, V., Raman, S., and Miyapuram, K.P. (2017, January 22\u201329). Facial Expression Recognition Using Visual Saliency and Deep Learning. Proceedings of the 2017 International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCVW.2017.327"},{"key":"ref_5","unstructured":"Zhu, X., and Ramanan, D. (2012, January 16\u201321). Face detection, pose estimation, and landmark localization in the wild. Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Ehrlich, M., Shields, T.J., Almaev, T., and Amer, M.R. (July, January 26). Facial attributes classification using multi-task representation learning. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Las Vegas, NV, USA.","DOI":"10.1109\/CVPRW.2016.99"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Kalayeh, M.M., Gong, B., and Shah, M. (2017, January 21\u201326). Improving Facial Attribute Prediction Using Semantic Segmentation. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.450"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Caruana, R. (1998). Multitask learning. Learning to Learn, Springer.","DOI":"10.1007\/978-1-4615-5529-2_5"},{"key":"ref_9","unstructured":"Ranjan, R., Sankaranarayanan, S., Castillo, C.D., and Chellappa, R. (June, January 30). An all-in-one convolutional neural network for face analysis. Proceedings of the 12th International Conference on Automatic Face & Gesture Recognition (FG), Washington, DC, USA."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Han, H., Jain, A.K., Shan, S., and Chen, X. (2017). Heterogeneous Face Attribute Estimation: A Deep Multi-Task Learning Approach. IEEE Trans. Pattern Anal. Mach. Intell.","DOI":"10.1109\/TPAMI.2017.2738004"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"2170","DOI":"10.1109\/TIFS.2014.2359646","article-title":"Age and gender estimation of unfiltered faces","volume":"9","author":"Eidinger","year":"2014","journal-title":"Trans. Inf. Forensics Secur."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Hassner, T., Harel, S., Paz, E., and Enbar, R. (2015, January 7\u201312). Effective face frontalization in unconstrained images. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7299058"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Levi, G., and Hassner, T. (2015, January 7\u201312). Age and gender classification using convolutional neural networks. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Boston, MA, USA.","DOI":"10.1109\/CVPRW.2015.7301352"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Van de Wolfshaar, J., Karaaba, M.F., and Wiering, M.A. (2015, January 7\u201310). Deep convolutional neural networks and support vector machines for gender recognition. Proceedings of the 2015 IEEE Symposium Series on Computational Intelligence, Cape Town, South Africa.","DOI":"10.1109\/SSCI.2015.37"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"144","DOI":"10.1007\/s11263-016-0940-3","article-title":"Deep expectation of real and apparent age from a single image without facial landmarks","volume":"126","author":"Rothe","year":"2018","journal-title":"Int. J. Comput. Vis."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Kumar, N., Belhumeur, P., and Nayar, S. (2008, January 12\u201318). Facetracer: A search engine for large collections of images with faces. Proceedings of the 10th European Conference on Computer Vision, Marseille, France.","DOI":"10.1007\/978-3-540-88693-8_25"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Zhang, N., Paluri, M., Ranzato, M., Darrell, T., and Bourdev, L. (2014, January 23\u201328). Panda: Pose aligned networks for deep attribute modeling. Proceedings of the 2014 IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.212"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Liu, Z., Luo, P., Wang, X., and Tang, X. (2015, January 7\u201313). Deep learning face attributes in the wild. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.425"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Rozsa, A., G\u00fcnther, M., Rudd, E.M., and Boult, T.E. (2016, January 4\u20138). Are facial attributes adversarially robust?. Proceedings of the 2016 23rd International Conference on Pattern Recognition (ICPR), Cancun, Mexico.","DOI":"10.1109\/ICPR.2016.7900114"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Rudd, E.M., G\u00fcnther, M., and Boult, T.E. (2016, January 8\u201316). Moon: A mixed objective optimization network for the recognition of facial attributes. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46454-1_2"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Uric\u00e1r, M., Timofte, R., Rothe, R., Matas, J., and Van Gool, L. (July, January 26). Structured output svm prediction of apparent age, gender and smile from deep features. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Las Vegas, NV, USA.","DOI":"10.1109\/CVPRW.2016.96"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Hand, E.M., and Chellappa, R. (2017, January 4\u20139). Attributes for Improved Attributes: A Multi-Task Network Utilizing Implicit and Explicit Relationships for Facial Attribute Classification. Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence (AAAI-17), San Francisco, CA, USA.","DOI":"10.1609\/aaai.v31i1.11229"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Li, Y., Wang, Q., Nie, L., and Cheng, H. (2017, January 25\u201327). Face Attributes Recognition via Deep Multi-Task Cascade. Proceedings of the International Conference on Data Mining, Communications and Information Technology, Phuket, Thailand.","DOI":"10.1145\/3089871.3089878"},{"key":"ref_24","unstructured":"Wang, F., Han, H., Shan, S., and Chen, X. (June, January 30). Deep Multi-Task Learning for Joint Prediction of Heterogeneous Face Attributes. Proceedings of the 12th International Conference on Automatic Face & Gesture Recognition (FG), Washington, DC, USA."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas Valley, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Comput."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Torralba, A., and Efros, A.A. (2011, January 20\u201325). Unbiased look at dataset bias. Proceedings of the 2011 Computer Vision and Pattern Recognition (CVPR), Colorado Springs, CO, USA.","DOI":"10.1109\/CVPR.2011.5995347"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Tommasi, T., Patricia, N., Caputo, B., and Tuytelaars, T. (2017). A deeper look at dataset bias. Domain Adaptation in Computer Vision Applications, Springer.","DOI":"10.1007\/978-3-319-58347-1_2"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Wen, Y., Zhang, K., Li, Z., and Qiao, Y. (2016, January 8\u201316). A discriminative feature learning approach for deep face recognition. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46478-7_31"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"758","DOI":"10.1109\/TIFS.2017.2766583","article-title":"An Ensemble CNN2ELM for Age Estimation","volume":"13","author":"Duan","year":"2018","journal-title":"Trans. Inf. Forensics Secur."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/18\/8\/2666\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T15:18:35Z","timestamp":1760195915000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/18\/8\/2666"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,8,14]]},"references-count":30,"journal-issue":{"issue":"8","published-online":{"date-parts":[[2018,8]]}},"alternative-id":["s18082666"],"URL":"https:\/\/doi.org\/10.3390\/s18082666","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2018,8,14]]}}}