{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,11]],"date-time":"2026-06-11T15:30:57Z","timestamp":1781191857901,"version":"3.54.1"},"reference-count":45,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2024,5,20]],"date-time":"2024-05-20T00:00:00Z","timestamp":1716163200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>Lightweight convolutional neural networks are widely used for face detection due to their ability to learn local representations through spatial induction bias and translational invariance. However, convolutional face detectors have limitations in detecting faces under challenging conditions like occlusion, blurring, or changes in facial poses, primarily attributed to fixed-size receptive fields and a lack of global modeling. Transformer-based models have advantages on learning global representations but are insensitive to capture local patterns. To address these limitations, we propose an efficient face detector that combines convolutional neural network and transformer architectures. We introduce a bi-stream structure that integrates convolutional neural network and transformer blocks within the backbone network, enabling the preservation of local pattern features and the extraction of global context. To further preserve the local details captured by convolutional neural networks, we propose a feature enhancement convolution block in a hierarchical backbone structure. Additionally, we devise a multiscale feature aggregation module to enhance obscured and blurred facial features. Experimental results demonstrate that our method has achieved improved lightweight face detection accuracy with an average precision of 95.30%, 94.20%, and 87.56% across the easy, medium, and hard subdatasets of WIDER FACE, respectively. Therefore, we believe our method will be a useful supplement to the collection of current artificial intelligence models and benefit the engineering applications of face detection.<\/jats:p>","DOI":"10.3390\/info15050290","type":"journal-article","created":{"date-parts":[[2024,5,20]],"date-time":"2024-05-20T11:06:41Z","timestamp":1716203201000},"page":"290","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["A Lightweight Face Detector via Bi-Stream Convolutional Neural Network and Vision Transformer"],"prefix":"10.3390","volume":"15","author":[{"given":"Zekun","family":"Zhang","sequence":"first","affiliation":[{"name":"College of Electronic Information, Qingdao University, Qingdao 260000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Qingqing","family":"Chao","sequence":"additional","affiliation":[{"name":"College of Electronic Information, Qingdao University, Qingdao 260000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shijie","family":"Wang","sequence":"additional","affiliation":[{"name":"College of Electronic Information, Qingdao University, Qingdao 260000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Teng","family":"Yu","sequence":"additional","affiliation":[{"name":"College of Electronic Information, Qingdao University, Qingdao 260000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2024,5,20]]},"reference":[{"key":"ref_1","unstructured":"Zhang, S., Zhu, R., Wang, X., Shi, H., Fu, T., Wang, S., Mei, T., and Li, S. (2019). Improved selective refinement network for face detection. arXiv."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Kuzdeuov, A., Koishigarina, D., and Varol, H.A. (2023, January 13\u201316). Anyface: A data-centric approach for input-agnostic face detection. Proceedings of the 2023 IEEE International Conference on Big Data and Smart Computing(BigComp), Jeju, Republic of Korea.","DOI":"10.1109\/BigComp57234.2023.00042"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Howard, A.G., Sandler, M., Chu, G., Chen, L.-C., Chen, B., Tan, M., Wang, W., Zhu, Y., Pang, R., and Vasudevan, V. (November, January 27). Searching for mobilenetv3. Proceedings of the 2019 IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea.","DOI":"10.1109\/ICCV.2019.00140"},{"key":"ref_4","unstructured":"Wang, H., Li, Z., Ji, X., and Wang, Y. (2017). Face r-cnn. arXiv."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"He, K., Gkioxari, G., Doll\u00e1r, P., and Girshick, R.B. (2017). Mask r-cnn. arXiv.","DOI":"10.1109\/ICCV.2017.322"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"1499","DOI":"10.1109\/LSP.2016.2603342","article-title":"Joint face detection and alignment using multitask cascaded convolutional networks","volume":"23","author":"Zhang","year":"2016","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Liu, Y., Wang, F., Sun, B., and Li, H. (2022, January 18\u201324). Mogface: Towards a deeper appreciation on face detection. Proceedings of the 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00406"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Deng, J., Guo, J., Ververas, E., Kotsia, I., and Zafeiriou, S. (2020, January 13\u201319). Retinaface: Single-shot multi-level face localisation in the wild. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00525"},{"key":"ref_9","unstructured":"Zhang, F., Fan, X., Ai, G., Song, J., Qin, Y., and Wu, J. (2019). Accurate face detection for high performance. arXiv."},{"key":"ref_10","unstructured":"Vaswani, A., Shazeer, N.M., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., and Polosukhin, I. (2017, January 4\u20139). Attention is all you need. Proceedings of the Neural Information Processing Systems 2017, Long Beach, CA, USA."},{"key":"ref_11","unstructured":"Wang, S., Li, B.Z., Khabsa, M., Fang, H., and Ma, H. (2020). Linformer: Self-attention with linear complexity. arXiv."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Wang, L., and Koniusz, P. (2023, January 17\u201324). 3mformer: Multi-order multi-mode transformer for skeletal action recognition. Proceedings of the 2023 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.00544"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Li, Y., Yu, Z., Choy, C.B., Xiao, C., \u00c1lvarez, J.M., Fidler, S., Feng, C., and Anandkumar, A. (2023, January 17\u201324). Voxformer: Sparse voxel transformer for camera-based 3d semantic scene completion. Proceedings of the 2023 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.00877"},{"key":"ref_14","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv."},{"key":"ref_15","unstructured":"Chu, X., Tian, Z., Zhang, B., Wang, X., and Shen, C. (2021, January 3\u20137). Conditional positional encodings for vision transformers. Proceedings of the International Conference on Learning Representations, Virtual Event, Austria."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021, January 10\u201317). Swin transformer: Hierarchical vision transformer using shifted windows. Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_17","unstructured":"Mehta, S., and Rastegari, M. (2021). Mobilevit: Light-weight, general-purpose, and mobile-friendly vision transformer. arXiv."},{"key":"ref_18","unstructured":"Dai, Z., Liu, H., Le, Q.V., and Tan, M. (2021). Coatnet: Marrying convolution and attention for all data sizes. arXiv."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Lin, T.-Y., Doll\u00e1r, P., Girshick, R.B., He, K., Hariharan, B., and Belongie, S.J. (2017, January 21\u201326). Feature pyramid networks for object detection. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Pan, X., Ge, C., Lu, R., Song, S., Chen, G., Huang, Z., and Huang, G. (2022, January 18\u201324). On the integration of self-attention and convolution. Proceedings of the 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00089"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Zhang, H., Hu, W., and Wang, X. (2022, January 23\u201327). Parc-net: Position aware circular convolution with merits from convnets and transformer. Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel.","DOI":"10.1007\/978-3-031-19809-0_35"},{"key":"ref_22","unstructured":"Chu, X., Tian, Z., Wang, Y., Zhang, B., Ren, H., Wei, X., Xia, H., and Shen, C. (2021, January 7\u201310). Twins: Revisiting the design of spatial attention in vision transformers. Proceedings of the Advances in Neural Information Processing Systems 34 (NeurIPS 2021), Online."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"106184","DOI":"10.1016\/j.engappai.2023.106184","article-title":"Multi-level learning counting via pyramid vision transformer and cnn","volume":"123","author":"Liu","year":"2023","journal-title":"Eng. Appl. Artif. Intell."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Woo, S., Park, J., Lee, J.-Y., and Kweon, I.-S. (2018). Cbam: Convolutional block attention module. arXiv.","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"ref_25","unstructured":"Yang, S., Xiong, Y., Loy, C.C., and Tang, X. (2017). Face detection through scale-friendly deep convolutional networks. arXiv."},{"key":"ref_26","unstructured":"Zhang, C., Xu, X., and Tu, D. (2018). Face detection using improved faster rcnn. arXiv."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S.E., Fu, C.-Y., and Berg, A.C. (2015, January 7\u201313). Ssd: Single shot multibox detector. Proceedings of the European Conference on Computer Vision, Santiago, Chile.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster r-cnn: Towards real-time object detection with region proposal networks","volume":"39","author":"Ren","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Li, J., Wang, Y., Wang, C., Tai, Y., Qian, J., Yang, J., Wang, C., Li, J., and Huang, F. (2019, January 15\u201320). Dsfd: Dual shot face detector. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00520"},{"key":"ref_30","unstructured":"He, Y., Xu, D., Wu, L., Jian, M., Xiang, S., and Pan, C. (2019). Lffd: A light and fast face detector for edge devices. arXiv."},{"key":"ref_31","unstructured":"Qi, D., Tan, W., Yao, Q., and Liu, J. (2021). Computer Vision\u2013ECCV 2022 Workshops. ECCV 2022, Springer."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Wang, G.Q., Li, J.Y., Wu, Z., Xu, J., Shen, J., and Yang, W. (2023). Efficientface: An efficient deep network with feature enhancement for accurate face detection. arXiv.","DOI":"10.1007\/s00530-023-01134-6"},{"key":"ref_33","unstructured":"Yoo, Y.J., Han, D., and Yun, S. (2019). Extd: Extremely tiny face detector via iterative filter reuse. arXiv."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S. (2020). End-to-end object detection with transformers. arXiv.","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"ref_35","unstructured":"Mehta, S., and Rastegari, M. (2022). Separable self-attention for mobile vision transformers. arXiv."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Hu, H., Gu, J., Zhang, Z., Dai, J., and Wei, Y. (2018, January 18\u201323). Relation networks for object detection. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00378"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Zhang, H., Wu, C., Zhang, Z., Zhu, Y., Zhang, Z.-L., Lin, H., Sun, Y., He, T., Mueller, J.W., and Manmatha, R. (2022, January 19\u201320). Resnest: Split-attention networks. Proceedings of the 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), New Orleans, LA, USA.","DOI":"10.1109\/CVPRW56347.2022.00309"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Yang, S., Luo, P., Loy, C.C., and Tang, X. (2016, January 27\u201330). Wider face: A face detection benchmark. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.596"},{"key":"ref_39","unstructured":"Jain, V., and Learned-Miller, E.G. (2010). Fddb: A Benchmark for Face Detection in Unconstrained Settings, UMass Amherst."},{"key":"ref_40","unstructured":"Loshchilov, I., and Hutter, F. (2017). Decoupled weight decay regularization. arXiv."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Tan, M., Pang, R., and Le, Q.V. (2020, January 13\u201319). Efficientdet: Scalable and efficient object detection. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01079"},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"34153","DOI":"10.1007\/s11042-020-09143-7","article-title":"Os-lffd: A light and fast face detector with ommateum structure","volume":"80","author":"Xu","year":"2020","journal-title":"Multimed. Tools Appl."},{"key":"ref_43","unstructured":"Guo, J., Deng, J., Lattas, A., and Zafeiriou, S. (2021). Sample and computation redistribution for efficient face detection. arXiv."},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Jiang, C., Ma, H., and Li, L. (2022, January 26\u201328). Irnet: An improved retinanet model for face detection. Proceedings of the 2022 7th International Conference on Image, Vision and Computing (ICIVC), Xi\u2019an, China.","DOI":"10.1109\/ICIVC55077.2022.9886975"},{"key":"ref_45","unstructured":"Zhu, Y., Cai, H., Zhang, S., Wang, C., and Xiong, Y. (2020). Tinaface: Strong but simple baseline for face detection. arXiv."}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/15\/5\/290\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T14:45:18Z","timestamp":1760107518000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/15\/5\/290"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,5,20]]},"references-count":45,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2024,5]]}},"alternative-id":["info15050290"],"URL":"https:\/\/doi.org\/10.3390\/info15050290","relation":{},"ISSN":["2078-2489"],"issn-type":[{"value":"2078-2489","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,5,20]]}}}