{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,12]],"date-time":"2026-05-12T19:26:41Z","timestamp":1778614001344,"version":"3.51.4"},"reference-count":29,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2017,5,17]],"date-time":"2017-05-17T00:00:00Z","timestamp":1494979200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["31301086"],"award-info":[{"award-number":["31301086"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>Face verification for unrestricted faces in the wild is a challenging task. This paper proposes a method based on two deep convolutional neural networks (CNN) for face verification. In this work, we explore using identification signals to supervise one CNN and the combination of semi-verification and identification to train the other one. In order to estimate semi-verification loss at a low computation cost, a circle, which is composed of all faces, is used for selecting face pairs from pairwise samples. In the process of face normalization, we propose using different landmarks of faces to solve the problems caused by poses. In addition, the final face representation is formed by the concatenating feature of each deep CNN after principal component analysis (PCA) reduction. Furthermore, each feature is a combination of multi-scale representations through making use of auxiliary classifiers. For the final verification, we only adopt the face representation of one region and one resolution of a face jointing Joint Bayesian classifier. Experiments show that our method can extract effective face representation with a small training dataset and our algorithm achieves 99.71% verification accuracy on Labeled Faces in the Wild (LFW) dataset.<\/jats:p>","DOI":"10.3390\/e19050228","type":"journal-article","created":{"date-parts":[[2017,5,17]],"date-time":"2017-05-17T11:13:17Z","timestamp":1495019597000},"page":"228","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["Face Verification with Multi-Task and Multi-Scale Feature Fusion"],"prefix":"10.3390","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2051-4290","authenticated-orcid":false,"given":"Xiaojun","family":"Lu","sequence":"first","affiliation":[{"name":"College of Sciences, Northeastern University, Shenyang 110819, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yue","family":"Yang","sequence":"additional","affiliation":[{"name":"College of Sciences, Northeastern University, Shenyang 110819, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Weilin","family":"Zhang","sequence":"additional","affiliation":[{"name":"Department of Mathematics, New York University Shanghai, 1555 Century Ave, Pudong, Shanghai 200122, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qi","family":"Wang","sequence":"additional","affiliation":[{"name":"College of Sciences, Northeastern University, Shenyang 110819, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yang","family":"Wang","sequence":"additional","affiliation":[{"name":"College of Sciences, Northeastern University, Shenyang 110819, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2017,5,17]]},"reference":[{"key":"ref_1","unstructured":"Ren, S., He, K., Girshick, R., and Sun, J. (arXiv, 2015). Faster R-CNN: Towards real-time object detection with region proposal networks, arXiv."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Long, J., Shelhamer, E., and Darrell, T. (2015, January 7\u201312). Fully convolutional networks for semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298965"},{"key":"ref_3","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, G.E. (2012, January 3\u20138). Imagenet classification with deep convolutional neural networks. Proceedings of the Twenty-Sixth Annual Conference on Neural Information Processing Systems (NIPS), Stateline, NV, USA."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Schroff, F., Kalenichenko, D., and Philbin, J. (2015, January 7\u201312). Facenet: A unified embedding for face recognition and clustering. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298682"},{"key":"ref_5","unstructured":"Liu, J., Deng, Y., Bai, T., Wei, Z.P., and Huang, H. (arXiv, 2015). Targeting ultimate accuracy: Face recognition via deep embedding, arXiv."},{"key":"ref_6","unstructured":"Sun, Y., Wang, X., and Tang, X. (arXiv, 2014). Deep learning face representation by joint identification-verification, arXiv."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Sun, Y., Wang, X., and Tang, X. (2014, January 23\u201328). Deep learning face representation from predicting 10,000 classes. Proceedings of the 2014 IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.244"},{"key":"ref_8","unstructured":"Simonyan, K., and Zisserman, A. (arXiv, 2014). Very deep convolutional networks for large-scale image recognition, arXiv."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A. (2015, January 7\u201312). Going deeper with convolutions. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_11","unstructured":"Chopra, S., Hadsell, R., and Lee, C.Y. (2005, January 20\u201325). Learning a similarity metric discriminatively, with application to face verification. Proceedings of the 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR\u201905), San Diego, CA, USA."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Wen, Y., Zhang, K., Li, Z., and Qiao, Y. (2016, January 11\u201314). A Discriminative Feature Learning Approach for Deep Face Recognition. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46478-7_31"},{"key":"ref_13","unstructured":"Yi, D., Lei, Z., Liao, S., and Li, S.Z. (arXiv, 2014). Learning face representation from scratch, arXiv."},{"key":"ref_14","unstructured":"Lee, C.Y., Xie, S., Gallagher, P., Zhang, Z., and Tu, Z. (arXiv, 2014). Deeply-Supervised Nets, arXiv."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Song, H.O., Xiang, Y., Jegelka, S., and Savarese, S. (arXiv, 2015). Deep metric learning via lifted structured feature embedding, arXiv.","DOI":"10.1109\/CVPR.2016.434"},{"key":"ref_16","unstructured":"Ioffe, S., and Szegedy, C. (arXiv, 2015). Batch normalization:Accelerating deep network training by reducing internal covariate shift, arXiv."},{"key":"ref_17","unstructured":"Nair, V., and Hinton, G.E. (2010, January 21\u201324). Rectified linear units improve restricted boltzmann machines. Proceedings of the 27th International Conference on Machine Learning, Haifa, Israel."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Ioffe, S., Vanhoucke, V., and Alemi, A. (arXiv, 2016). Inception-v4, inception-resnet and the impact of residual connections on learning, arXiv.","DOI":"10.1609\/aaai.v31i1.11231"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Yang, S., Luo, P., Loy, C.C., and Tang, X. (2015, January 7\u201313). From facial parts responses to face detection: A deep learning approach. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.419"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Ye, M., Wang, X., Yang, R., Ren, L., and Pollefeys, M. (2011, January 6\u201313). Accurate 3d pose estimation from a single depth image. Proceedings of the 2011 International Conference on Computer Vision, Barcelona, Spain.","DOI":"10.1109\/ICCV.2011.6126310"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Chen, D., Cao, X., Wang, L., Wen, F., and Sun, J. (2012, January 7\u201313). Bayesian face revisited: A joint formulation. Proceedings of the 12th European Conference on Computer Vision, Florence, Italy.","DOI":"10.1007\/978-3-642-33712-3_41"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Jia, Y., Shelhamer, E., Donahue, J., Karayev, S., Long, J., Girshick, R., Guadarrama, S., and Darrell, T. (arXiv, 2014). Caffe: Convolutional architecture for fast feature embedding, arXiv.","DOI":"10.1145\/2647868.2654889"},{"key":"ref_23","unstructured":"Huang, G.B., Ramesh, M., Berg, T., and Learned-Miller, E. (2007). Labeled Faces in the Wild: A Database for Studying Face Recognition in Unconstrained Environments, University of Massachusetts. Technical Report."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Wolf, L., Hassner, T., and Maoz, I. (2011, January 20\u201325). Face recognition in unconstrained videos with matched background similarity. Proceedings of the 2011 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Colorado Springs, CO, USA.","DOI":"10.1109\/CVPR.2011.5995566"},{"key":"ref_25","unstructured":"Sun, Y., Liang, D., Wang, X., and Tang, X. (arXiv, 2015). Deepid3: Face recognition with very deep neural networks, arXiv."},{"key":"ref_26","unstructured":"Zhou, E., Cao, Z., and Yin, Q. (arXiv, 2015). Naive-deep face recognition: Touching the limit of LFW benchmark or not?, arXiv."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Sun, Y., Wang, X., and Tang, X. (2015, January 7\u201312). Deeply learned face representations are sparse, selective, and robust. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298907"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Parkhi, O.M., Vedaldi, A., and Zisserman, A. (2015, January 7\u201310). Deep Face Recognition. Proceedings of the British Machine Vision Conference (BMVC), Swansea, UK.","DOI":"10.5244\/C.29.41"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Taigman, Y., Yang, M., Ranzato, M.A., and Wolf, L. (2014, January 23\u201328). Deepface: Closing the gap to human-level performance in face verification. Proceedings of the 2014 IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.220"}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/19\/5\/228\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T18:36:09Z","timestamp":1760207769000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/19\/5\/228"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017,5,17]]},"references-count":29,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2017,5]]}},"alternative-id":["e19050228"],"URL":"https:\/\/doi.org\/10.3390\/e19050228","relation":{},"ISSN":["1099-4300"],"issn-type":[{"value":"1099-4300","type":"electronic"}],"subject":[],"published":{"date-parts":[[2017,5,17]]}}}