{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,3]],"date-time":"2026-06-03T04:02:38Z","timestamp":1780459358436,"version":"3.54.1"},"reference-count":39,"publisher":"MDPI AG","issue":"13","license":[{"start":{"date-parts":[[2024,7,4]],"date-time":"2024-07-04T00:00:00Z","timestamp":1720051200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>In the field of autofocus for optical systems, although passive focusing methods are widely used due to their cost-effectiveness, fixed focusing windows and evaluation functions in certain scenarios can still lead to focusing failures. Additionally, the lack of datasets limits the extensive research of deep learning methods. In this work, we propose a neural network autofocus method with the capability of dynamically selecting the region of interest (ROI). Our main work is as follows: first, we construct a dataset for automatic focusing of grayscale images; second, we transform the autofocus issue into an ordinal regression problem and propose two focusing strategies: full-stack search and single-frame prediction; and third, we construct a MobileViT network with a linear self-attention mechanism to achieve automatic focusing on dynamic regions of interest. The effectiveness of the proposed focusing method is verified through experiments, and the results show that the focusing MAE of the full-stack search can be as low as 0.094, with a focusing time of 27.8 ms, and the focusing MAE of the single-frame prediction can be as low as 0.142, with a focusing time of 27.5 ms.<\/jats:p>","DOI":"10.3390\/s24134336","type":"journal-article","created":{"date-parts":[[2024,7,4]],"date-time":"2024-07-04T03:52:58Z","timestamp":1720065178000},"page":"4336","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":14,"title":["Deep Learning-Based Dynamic Region of Interest Autofocus Method for Grayscale Image"],"prefix":"10.3390","volume":"24","author":[{"given":"Yao","family":"Wang","sequence":"first","affiliation":[{"name":"Changchun Institute of Optics, Fine Mechanics and Physics, Chinese Academy of Sciences, Changchun 130033, China"},{"name":"University of Chinese Academy of Sciences, Beijing 100049, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chuan","family":"Wu","sequence":"additional","affiliation":[{"name":"Changchun Institute of Optics, Fine Mechanics and Physics, Chinese Academy of Sciences, Changchun 130033, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8857-0259","authenticated-orcid":false,"given":"Yunlong","family":"Gao","sequence":"additional","affiliation":[{"name":"Changchun Institute of Optics, Fine Mechanics and Physics, Chinese Academy of Sciences, Changchun 130033, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-7871-2088","authenticated-orcid":false,"given":"Huiying","family":"Liu","sequence":"additional","affiliation":[{"name":"Changchun Institute of Optics, Fine Mechanics and Physics, Chinese Academy of Sciences, Changchun 130033, China"},{"name":"University of Chinese Academy of Sciences, Beijing 100049, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2024,7,4]]},"reference":[{"key":"ref_1","first-page":"63","article-title":"Autofocus area design of digital imaging system","volume":"31","author":"Li","year":"2002","journal-title":"Acta Photonica Sin."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Herrmann, C., Bowen, R.S., Wadhwa, N., Garg, R., He, Q.R., Barron, J.T., and Zabih, R. (2020, January 14\u201319). Learning to Autofocus. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Electr Network, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00230"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1237","DOI":"10.1109\/TCSVT.2008.924105","article-title":"Enhanced Autofocus Algorithm Using Robust Focus Measure and Fuzzy Reasoning","volume":"18","author":"Lee","year":"2008","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Pech-Pacheco, J.L., Crist\u00f3bal, G., Chamorro-Mart\u00ednez, J., and Fern\u00e1ndez-Valdivia, J. (2000, January 3\u20137). Diatom autofocusing in brightfield microscopy: A comparative study. Proceedings of the 15th International Conference on Pattern Recognition (ICPR-2000), Barcelona, Spain.","DOI":"10.1109\/ICPR.2000.903548"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"64837","DOI":"10.1109\/ACCESS.2019.2914186","article-title":"A Novel Auto-Focus Method for Image Processing Using Laser Triangulation","volume":"7","author":"Zhang","year":"2019","journal-title":"IEEE Access"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Yousefi, S., Rahman, M., Kehtarnavaz, N., and Gamadia, M. (2011, January 9\u201312). A New Auto-Focus Sharpness Function for Digital and Smart-Phone Cameras. Proceedings of the IEEE International Conference on Consumer Electronics (ICCE 2011), Las Vegas, NV, USA.","DOI":"10.1109\/ICCE.2011.5722691"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Evers, A., and Jackson, J.A. (2020, January 28\u201330). A Comparison of Autofocus Algorithms for Backprojection Synthetic Aperture Radar. Proceedings of the IEEE International Radar Conference (RADAR), Electr Network, Washington, DC, USA.","DOI":"10.1109\/RADAR42522.2020.9114579"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"773","DOI":"10.1002\/jemt.24332","article-title":"Fast autofocus method for piezoelectric microscopy system for high interaction scenes","volume":"86","author":"Hao","year":"2023","journal-title":"Microsc. Res. Tech."},{"key":"ref_9","first-page":"24","article-title":"Contrast optimization autofocus algorithm","volume":"25","author":"Liu","year":"2003","journal-title":"J. Electron. Inf. Technol."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Rigling, B.D. (2006, January 17\u201320). Multistage entropy minimization for SAR image autofocus\u2014Art. no. 62370J. Proceedings of the Conference on Algorithms for Synthetic Aperture Radar Imagery XIII, Kissimmee, FL, USA.","DOI":"10.1117\/12.669957"},{"key":"ref_11","unstructured":"Yang, G., and Nelson, B.J. (2003, January 27\u201331). Wavelet-based autofocusing and unsupervised segmentation of microscopic images. Proceedings of the IEEE\/RSJ International Conference on Intelligent Robots and Systems, Las Vegas, NV, USA."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Liu, K.H., and Munson, D.C. (2008, January 26\u201329). Fourier-Domain Multichannel Autofocus for Synthetic Aperture Radar. Proceedings of the 42nd Asilomar Conference on Signals, Systems and Computers, Pacific Grove, CA, USA.","DOI":"10.1109\/ACSSC.2008.5074529"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"133","DOI":"10.1109\/LSP.2008.2008938","article-title":"Reduced Energy-Ratio Measure for Robust Autofocusing in Digital Camera","volume":"16","author":"Lee","year":"2009","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"8281","DOI":"10.3390\/s110908281","article-title":"Robust Automatic Focus Algorithm for Low Contrast Images Using a New Contrast Measure","volume":"11","author":"Xu","year":"2011","journal-title":"Sensors"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"257","DOI":"10.1109\/TCE.2003.1209511","article-title":"Modified fast climbing search auto-focus algorithm with adaptive step size searching technique for digital camera","volume":"49","author":"He","year":"2003","journal-title":"IEEE Trans. Consum. Electron."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"1575","DOI":"10.1109\/TIP.2017.2698924","article-title":"Analysis of Disparity Error for Stereo Autofocus","volume":"27","author":"Yang","year":"2018","journal-title":"IEEE Trans. Image Process."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"599","DOI":"10.1007\/s10732-015-9291-4","article-title":"An autofocus heuristic for digital cameras based on supervised machine learning","volume":"21","author":"Mir","year":"2015","journal-title":"J. Heuristics"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"480","DOI":"10.1364\/BOE.379780","article-title":"Whole slide imaging system using deep learning-based automated focusing","volume":"11","author":"Dastidar","year":"2020","journal-title":"Biomed. Opt. Express"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"428","DOI":"10.1109\/LSP.2023.3265569","article-title":"Deep Ordinal Regression Framework for No-Reference Image Quality Assessment","volume":"30","author":"Wang","year":"2023","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"794","DOI":"10.1364\/OPTICA.6.000794","article-title":"Deep learning for single-shot autofocus microscopy","volume":"6","author":"Pinkard","year":"2019","journal-title":"Optica"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"314","DOI":"10.1364\/BOE.446928","article-title":"Deep learning-based single-shot autofocus method for digital microscopy","volume":"13","author":"Liao","year":"2022","journal-title":"Biomed. Opt. Express"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"1601","DOI":"10.1364\/BOE.9.001601","article-title":"Transform- and multi-domain deep learning for single-frame rapid autofocusing in whole slide imaging","volume":"9","author":"Jiang","year":"2018","journal-title":"Biomed. Opt. Express"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"258","DOI":"10.1109\/TCI.2021.3059497","article-title":"Deep Learning for Camera Autofocus","volume":"7","author":"Wang","year":"2021","journal-title":"IEEE Trans. Comput. Imaging"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"13","DOI":"10.1109\/TGRS.2022.3217063","article-title":"AFnet and PAFnet: Fast and Accurate SAR Autofocus Based on Deep Learning","volume":"60","author":"Liu","year":"2022","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Sakurikar, P., Mehta, I., Balasubramanian, V.N., and Narayanan, P.J. (2018, January 8\u201314). RefocusGAN: Scene Refocusing Using a Single Image. Proceedings of the 15th European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01225-0_31"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"1280","DOI":"10.1049\/el.2018.5101","article-title":"Image ordinal classification with deep multi-view learning","volume":"54","author":"Zhang","year":"2018","journal-title":"Electron. Lett."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Xun, L.N., Zhang, H.C., Yan, Q., Wu, Q., and Zhang, J. (2022). VISOR-NET: Visibility Estimation Based on Deep Ordinal Relative Learning under Discrete-Level Labels. Sensors, 22.","DOI":"10.3390\/s22166227"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Niu, Z.X., Zhou, M., Wang, L., Gao, X.B., and Hua, G. (2016, January 27\u201330). Ordinal Regression with Multiple Output CNN for Age Estimation. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR.2016.532"},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"325","DOI":"10.1016\/j.patrec.2020.11.008","article-title":"Rank consistent ordinal regression for neural networks with application to age estimation","volume":"140","author":"Cao","year":"2020","journal-title":"Pattern Recognit. Lett."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"318","DOI":"10.1109\/TPAMI.2018.2858826","article-title":"Focal Loss for Dense Object Detection","volume":"42","author":"Lin","year":"2020","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"D\u00edaz, R., Marathe, A., and Soc, I.C. (2019, January 16\u201320). Soft Labels for Ordinal Regression. Proceedings of the 32nd IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00487"},{"key":"ref_32","unstructured":"Wu, H.X., Wu, J.L., Xu, J.H., Wang, J.M., and Long, M.S. (2022, January 17\u201323). Flowformer: Linearizing Transformers with Conservation Flows. Proceedings of the 39th International Conference on Machine Learning (ICML), Baltimore, MD, USA."},{"key":"ref_33","unstructured":"Mehta, S., and Rastegari, M. (2021). MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer. arXiv."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"259","DOI":"10.1016\/0167-2789(92)90242-F","article-title":"Nonlinear total variation based noise removal algorithms","volume":"60","author":"Rudin","year":"1992","journal-title":"Physica D"},{"key":"ref_35","unstructured":"Tenenbaum, J.M. (1971). Accommodation in Computer Vision, Stanford University."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"151","DOI":"10.1109\/TIP.2008.2007049","article-title":"Improvements in Shape-From-Focus for Holographic Reconstructions With Regard to Focus Operators, Neighborhood-Size, and Height Value Interpolation","volume":"18","author":"Thelen","year":"2009","journal-title":"IEEE Trans. Image Process."},{"key":"ref_37","unstructured":"Howard, A., Sandler, M., Chu, G., Chen, L.C., Chen, B., Tan, M.X., Wang, W.J., Zhu, Y.K., Pang, R.M., and Vasudevan, V. (November, January 27). Searching for MobileNetV3. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Ma, N.N., Zhang, X.Y., Zheng, H.T., and Sun, J. (2018, January 8\u201314). ShuffleNet V2: Practical Guidelines for Efficient CNN Architecture Design. Proceedings of the 15th European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01264-9_8"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Sandler, M., Howard, A., Zhu, M.L., Zhmoginov, A., Chen, L.C., and IEEE (2018, January 18\u201323). MobileNetV2: Inverted Residuals and Linear Bottlenecks. In Proceedings of the 31st IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00474"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/13\/4336\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T15:09:57Z","timestamp":1760108997000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/13\/4336"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,7,4]]},"references-count":39,"journal-issue":{"issue":"13","published-online":{"date-parts":[[2024,7]]}},"alternative-id":["s24134336"],"URL":"https:\/\/doi.org\/10.3390\/s24134336","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,7,4]]}}}