{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T02:43:44Z","timestamp":1760150624397,"version":"build-2065373602"},"reference-count":43,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2023,12,19]],"date-time":"2023-12-19T00:00:00Z","timestamp":1702944000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Sichuan Science and Technology Planning Project","award":["2019ZDZX0007","2021YFQ0056","2021YJ0372"],"award-info":[{"award-number":["2019ZDZX0007","2021YFQ0056","2021YJ0372"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>For the synthesis of ultra-large scene and ultra-high resolution videos, in order to obtain high-quality large-scene videos, high-quality video stitching and fusion are achieved through multi-scale unstructured array cameras. This paper proposes a network model image feature point extraction algorithm based on symmetric auto-encoding and scale feature fusion. By using the principle of symmetric auto-encoding, the hierarchical restoration of image feature location information is incorporated into the corresponding scale feature, along with deep separable convolution image feature extraction, which not only improves the performance of feature point detection but also significantly reduces the computational complexity of the network model. Based on the calculated high-precision feature point pairing information, a new image localization method is proposed based on area ratio and homography matrix scaling, which improves the speed and accuracy of the array camera image scale alignment and positioning, realizes high-definition perception of local details in large scenes, and obtains clearer synthesis effects of large scenes and high-quality stitched images. The experimental results show that the feature point extraction algorithm proposed in this paper has been experimentally compared with four typical algorithms using the HPatches dataset. The performance of feature point detection has been improved by an average of 4.9%, the performance of homography estimation has been improved by an average of 2.5%, the amount of computation has been reduced by 18%, the number of network model parameters has been reduced by 47%, and the synthesis of billion-pixel videos has been achieved, demonstrating practicality and robustness.<\/jats:p>","DOI":"10.3390\/s24010005","type":"journal-article","created":{"date-parts":[[2023,12,19]],"date-time":"2023-12-19T06:03:11Z","timestamp":1702965791000},"page":"5","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["SASFF: A Video Synthesis Algorithm for Unstructured Array Cameras Based on Symmetric Auto-Encoding and Scale Feature Fusion"],"prefix":"10.3390","volume":"24","author":[{"ORCID":"https:\/\/orcid.org\/0009-0008-6732-4102","authenticated-orcid":false,"given":"Linliang","family":"Zhang","sequence":"first","affiliation":[{"name":"School of Information Science and Technology, Southwest Jiaotong University, Chengdu 611756, China"},{"name":"Shanxi Intelligent Transportation Institute Co., Ltd., Taiyuan 030036, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lianshan","family":"Yan","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology, Southwest Jiaotong University, Chengdu 611756, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shuo","family":"Li","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology, Southwest Jiaotong University, Chengdu 611756, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Saifei","family":"Li","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology, Southwest Jiaotong University, Chengdu 611756, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,12,19]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"6441","DOI":"10.1109\/TIP.2023.3333547","article-title":"MCSfM: Multi-Camera Based Incremental Structure-from-Motion","volume":"32","author":"Cui","year":"2023","journal-title":"IEEE Trans. Image Process."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"6004","DOI":"10.1109\/TIP.2023.3327912","article-title":"BVI-VFI: A Video Quality Database for Video Frame Interpolation","volume":"32","author":"Danier","year":"2023","journal-title":"IEEE Trans. Image Process."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"6115","DOI":"10.1109\/TIP.2023.3328478","article-title":"From Global to Local: Multi-scale Out-of-distribution Detection","volume":"32","author":"Zhang","year":"2023","journal-title":"IEEE Trans. Image Process."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"6090","DOI":"10.1109\/TIP.2023.3328471","article-title":"Multi-level Content-aware Boundary Detection for Temporal Action Proposal Generation","volume":"32","author":"Su","year":"2023","journal-title":"IEEE Trans. Image Process."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"6075","DOI":"10.1109\/TIP.2023.3328486","article-title":"Optimization-Inspired Learning with Architecture Augmentations and Control Mechanisms for Low-Level Vision","volume":"32","author":"Liu","year":"2023","journal-title":"IEEE Trans. Image Process."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"6061","DOI":"10.1109\/TIP.2023.3328230","article-title":"Self-Supervised 3D Behavior Representation Learning Based on Homotopic Hyperbolic Embedding","volume":"32","author":"Chen","year":"2023","journal-title":"IEEE Trans. Image Process."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"433","DOI":"10.1109\/TIP.2021.3130538","article-title":"Toward Unaligned Guided Thermal Super-Resolution","volume":"31","author":"Gupta","year":"2021","journal-title":"IEEE Trans. Image Process."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"6032","DOI":"10.1109\/TIP.2023.3327924","article-title":"Clip-driven fine-grained text-image person re-identification","volume":"32","author":"Yan","year":"2023","journal-title":"IEEE Trans. Image Process."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Ashraf, M.W., Sultani, W., and Shah, M. (2021, January 20\u201325). Dogfight: Detecting drones from drones videos. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00699"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Dai, Z., Cai, B., Lin, Y., and Chen, J. (2021, January 20\u201325). Up-detr: Unsupervised pre-training for object detection with transformers. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00165"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"127503","DOI":"10.1016\/j.optcom.2021.127503","article-title":"Calibration of a camera-array-based microscopic system with spatiotemporal structured light encoding","volume":"504","author":"Hu","year":"2022","journal-title":"Opt. Commun."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"3195","DOI":"10.1109\/TNNLS.2021.3053249","article-title":"New Generation Deep Learning for Video Object Detection: A Survey","volume":"33","author":"Jiao","year":"2021","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"166520","DOI":"10.1016\/j.ijleo.2021.166520","article-title":"A high-quality stitching algorithm based on fisheye images","volume":"238","author":"Xue","year":"2021","journal-title":"Optik"},{"key":"ref_14","first-page":"1811013","article-title":"Multi-Camera System: Imaging Enhancement and Application","volume":"58","author":"Guo","year":"2021","journal-title":"Laser Optoelectron. Prog."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Kim, Y., Koh, Y.J., Lee, C., Kim, S., and Kim, C.-S. (2015, January 27\u201330). Dark image enhancement based onpairwise target contrast and multi-scale detail boosting. Proceedings of the 2015 IEEE International Conference on Image Processing (ICIP), Quebec City, QC, Canada.","DOI":"10.1109\/ICIP.2015.7351031"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Lee, S., Seong, H., Lee, S., and Kim, E. (2022, January 18\u201324). Correlation verification for image retrieval. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00530"},{"key":"ref_17","unstructured":"Zhang, R., and Wang, L. (2011, January 21\u201323). An image matching evolutionary algorithm based on Hu invariant moments. Proceedings of the 2011 International Conference on Image Analysis and Signal Processing, Wuhan, China."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"901","DOI":"10.1049\/iet-ipr.2019.1157","article-title":"Robust image hashing with visual attention model and invariant moments","volume":"14","author":"Tang","year":"2020","journal-title":"IET Image Process."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Yi, K.M., Trulls, E., Lepetit, V., and Fua, P. (2016, January 11\u201314). Lift: Learned invariant feature transform. Proceedings of the Computer Vision\u2013ECCV 2016: 14th European Conference, Amsterdam, The Netherlands. Proceedings, Part VI 14.","DOI":"10.1007\/978-3-319-46466-4_28"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"142","DOI":"10.1016\/j.mri.2019.05.037","article-title":"Learning image-based spatial transformations via convolutional neural networks: A review","volume":"64","author":"Tustison","year":"2019","journal-title":"Magn. Reson. Imaging"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"DeTone, D., Malisiewicz, T., and Rabinovich, A. (2018, January 18\u201322). Superpoint: Self-supervised interest point detection and description. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPRW.2018.00060"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Bian, J.W., Lin, W.Y., Matsushita, Y., Yeung, S.-K., Nguyen, T.-D., and Cheng, M.-M. (2017, January 21\u201326). Gms: Grid-based motion statistics for fast, ultra-robust feature correspondence. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.302"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Lindenberger, P., Sarlin, P.E., and Pollefeys, M. (2023). LightGlue: Local Feature Matching at Light Speed. arXiv.","DOI":"10.1109\/ICCV51070.2023.01616"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Cossairt, O.S., Miau, D., and Nayar, S.K. (2011, January 8\u201310). Gigapixel computational imaging. Proceedings of the 2011 IEEE International Conference on Computational Photography (ICCP), Pittsburgh, PA, USA.","DOI":"10.1109\/ICCPHOT.2011.5753115"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"386","DOI":"10.1038\/nature11150","article-title":"Multiscale gigapixel photography","volume":"486","author":"Brady","year":"2012","journal-title":"Nature"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"93","DOI":"10.1145\/1276377.1276494","article-title":"Capturing and viewing gigapixel images","volume":"26","author":"Kopf","year":"2007","journal-title":"ACM Trans. Graph."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"1474","DOI":"10.1086\/508573","article-title":"Astrometry in Wide-Field Surveys","volume":"118","author":"Bakos","year":"2006","journal-title":"Publ. Astron. Soc. Pac."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"354","DOI":"10.1017\/S174392130802663X","article-title":"HAT-South: A Global network of southern Hemisphere automated telescopes to detect transiting exoplanets","volume":"4","author":"Bakos","year":"2008","journal-title":"Proc. Int. Astron. Union"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Takahashi, I., Tsunashima, K., Tatsuhito, T., Saori, O., Kazutaka, Y., and Yoshida, A. (2010, January 30). Optical wide field monitor AROMA-W using multiple digital single-lens reflex cameras. Proceedings of the the First Year of MAXI: Monitoring Variable X-ray Sources, Tokyo, Japan.","DOI":"10.1155\/2010\/214604"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"765","DOI":"10.1145\/1073204.1073259","article-title":"High performance imaging using large camera arrays","volume":"24","author":"Wilburn","year":"2005","journal-title":"ACM Trans. Graph."},{"key":"ref_31","unstructured":"Nomura, Y., Zhang, L., and Nayar, S.K. (2007, January 25\u201327). Scene collages and flexible camera arrays. Proceedings of the 18th Eurographics Conference on Rendering Techniques, Grenoble, France."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"37","DOI":"10.1038\/s41377-021-00485-x","article-title":"A modular hierarchical array camera","volume":"10","author":"Yuan","year":"2021","journal-title":"Light Sci. Appl."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., and Chen, L. (2018, January 18\u201323). Mobilenetv2: Inverted residuals and linear bottlenecks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00474"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Balntas, V., Lenc, K., Vedaldi, A., and Mikolajczyk, K. (2017, January 21\u201326). HPatches: A benchmark and evaluation of handcrafted and learned local descriptors. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.410"},{"key":"ref_35","unstructured":"Rosten, E., and Drummond, T. (2006). European Conference on Computer Vision, Springer."},{"key":"ref_36","first-page":"10","article-title":"A combined corner and edge detector","volume":"15","author":"Harris","year":"1988","journal-title":"Alvey Vis. Conf."},{"key":"ref_37","unstructured":"Shi, J. (1994, January 21\u201323). Good features to track. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Rublee, E., Rabaud, V., Konolige, K., and Bradski, G. (2011, January 6\u201313). ORB: An efficient alternative to SIFT or SURF. Proceedings of the IEEE International Conference on Computer Vision, ICCV 2011, Barcelona, Spain.","DOI":"10.1109\/ICCV.2011.6126544"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Lowe, D.G. (1999, January 20\u201327). Object recognition from local scale-invariant features. Proceedings of the Seventh IEEE International Conference on Computer Vision, Kerkyra, Greece.","DOI":"10.1109\/ICCV.1999.790410"},{"key":"ref_40","unstructured":"Bay, H., Tuytelaars, T., and Gool, L.V. (2006). European Conference on Computer Vision, Springer."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Sun, J., Shen, Z., Wang, Y., Bao, H., and Zhou, X. (2021, January 20\u201325). LoFTR: Detector-free local feature matching with transformers. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00881"},{"key":"ref_42","first-page":"14254","article-title":"DISK: Learning local features with policy gradient","volume":"33","author":"Tyszkiewicz","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"809","DOI":"10.1038\/s41566-019-0474-7","article-title":"Video-rate imaging of biological dynamics at centimetre scale and micrometre resolution","volume":"13","author":"Fan","year":"2019","journal-title":"Nat. Photonics"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/1\/5\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T21:41:21Z","timestamp":1760132481000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/1\/5"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,12,19]]},"references-count":43,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2024,1]]}},"alternative-id":["s24010005"],"URL":"https:\/\/doi.org\/10.3390\/s24010005","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2023,12,19]]}}}