{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,4]],"date-time":"2026-06-04T13:46:16Z","timestamp":1780580776627,"version":"3.54.1"},"reference-count":47,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2025,4,7]],"date-time":"2025-04-07T00:00:00Z","timestamp":1743984000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62172032 and 62120106009"],"award-info":[{"award-number":["62172032 and 62120106009"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Fundamental Research Program of Shanxi Province","award":["202403021222025"],"award-info":[{"award-number":["202403021222025"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,4,30]]},"abstract":"<jats:p>Numerous significant progress on fisheye image rectification has been achieved through CNN. Nevertheless, constrained by a fixed receptive field, the global distribution and the local symmetry of the distortion have not been fully exploited. To leverage these two characteristics, we introduce FishFormer that processes the fisheye image as a sequence to enhance global and local perception. We tune the Transformer according to the structural properties of fisheye images. First, the uneven distortion distribution in patches generated by the existing square slicing method hinders the understanding of the global structure. Therefore, we propose an annulus slicing method to maintain the consistency of the distortion in each patch; thus, the applicability of the Transformer is expanded to perceive the distortion distribution efficiently. Second, the distortion of adjacent patches is progressive. Such explicit correlations in local regions need to be rapidly constructed and maintained, but Transformer has a weakness in local area perception. Hence, a novel layer attention mechanism is introduced to enhance the local perception and feature interaction. Our network simultaneously implements global perception and focused local perception. Extensive experiments demonstrate that our method provides superior performance compared with state-of-the-art methods.<\/jats:p>","DOI":"10.1145\/3719348","type":"journal-article","created":{"date-parts":[[2025,2,26]],"date-time":"2025-02-26T15:11:06Z","timestamp":1740582666000},"page":"1-16","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["FishFormer: Annulus Slicing-based Transformer for Fisheye Rectification"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0009-0008-3419-9501","authenticated-orcid":false,"given":"Shangrong","family":"Yang","sequence":"first","affiliation":[{"name":"Institute of Information Science, Beijing Jiaotong University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2847-0349","authenticated-orcid":false,"given":"Chunyu","family":"Lin","sequence":"additional","affiliation":[{"name":"Institute of Information Science, Beijing Jiaotong University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9429-1096","authenticated-orcid":false,"given":"Kang","family":"Liao","sequence":"additional","affiliation":[{"name":"Institute of Information Science, Beijing Jiaotong University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8581-9554","authenticated-orcid":false,"given":"Yao","family":"Zhao","sequence":"additional","affiliation":[{"name":"Institute of Information Science; Beijing Key Laboratory of Advanced Information Science and Network Technology, Beijing Jiaotong University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,4,7]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1177\/0278364920903774"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/3526024"},{"issue":"1","key":"e_1_3_1_4_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3462219","article-title":"Multi-feature fusion VoteNet for 3D object detection","volume":"18","author":"Wang Z.","year":"2002","unstructured":"Z. Wang, Q. Xie, M. Wei, K. Long, and J. Wang. 2002. Multi-feature fusion VoteNet for 3D object detection. ACM Trans. Multimedia Comput. Commun. Appl. 18, 1 (2002), 1\u201317.","journal-title":"ACM Trans. Multimedia Comput. Commun. Appl"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/3584362"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/3545609"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2012.10.006"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3430257"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/3557896"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/7.55557"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/3159170"},{"key":"e_1_3_1_12_2","first-page":"2301","volume-title":"International Conference on Intelligent Robots and Systems","author":"Zhang Q.","year":"2004","unstructured":"Q. Zhang and R. Pless. 2004. Extrinsic calibration of a camera and laser range finder (improves camera calibration). In International Conference on Intelligent Robots and Systems, 2301\u20132306."},{"key":"e_1_3_1_13_2","first-page":"666","volume-title":"IEEE\/CVF International Conference on Computer Vision (ICCV)","author":"Zhang Z.","year":"1999","unstructured":"Z. Zhang. 1999. Flexible camera calibration by viewing a plane from unknown orientations. In IEEE\/CVF International Conference on Computer Vision (ICCV), 666\u2013673."},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2013.137"},{"issue":"1","key":"e_1_3_1_15_2","doi-asserted-by":"crossref","first-page":"215","DOI":"10.1017\/S0033822200013904","article-title":"Extended 14c data base and revised calib 3.0 14c age calibration program","volume":"35","author":"Stuiver M.","year":"1993","unstructured":"M. Stuiver and P. Reimer. 1993. Extended 14c data base and revised calib 3.0 14c age calibration program. Radiocarbon 35, 1 (1993), 215\u2013230.","journal-title":"Radiocarbon"},{"key":"e_1_3_1_16_2","first-page":"537","volume-title":"IEEE\/CVF International Conference on Computer Vision (ICCV)","author":"Melo R.","year":"2013","unstructured":"R. Melo, M. Antunes, J. P. Barreto, G. Falc\u00e3o Paiva Fernandes, and N. Gon\u00e7alves. 2013. Unsupervised intrinsic calibration from a single frame using a \u201cplumb-line\u201d approach. In IEEE\/CVF International Conference on Computer Vision (ICCV), 537\u2013544."},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/3065386"},{"key":"e_1_3_1_18_2","first-page":"35","volume-title":"Asian Conference on Computer Vision (ACCV)","volume":"10113","author":"Rong J.","year":"2016","unstructured":"J. Rong, S. Huang, Z. Shang, and X. Ying. 2016. Radial lens distortion correction using convolutional neural networks trained with synthesized images. In Asian Conference on Computer Vision (ACCV), Vol. 10113, 35\u201349."},{"key":"e_1_3_1_19_2","first-page":"475","volume-title":"European Conference Computer Vision (ECCV)","volume":"11214","author":"Yin X.","year":"2018","unstructured":"X. Yin, X. Wang, J. Yu, M. Zhang, P. Fua, and D. Tao. 2018. Fisheyerecnet: A multi-context collaborative deep network for fisheye image rectification. In European Conference Computer Vision (ECCV), Vol. 11214, 475\u2013490."},{"key":"e_1_3_1_20_2","first-page":"1643","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Xue Z.","year":"2019","unstructured":"Z. Xue, N. Xue, G. S. Xia, and W. Shen. 2019. Learning to calibrate straight lines for fisheye image rectification. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 1643\u20131651."},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2019.2897984"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2020.2964523"},{"key":"e_1_3_1_23_2","first-page":"4855","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Li X.","year":"2019","unstructured":"X. Li, B. Zhang, P. V. Sander, and J. Liao. 2019. Blind geometric distortion correction on images through deep learning. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 4855\u20134864."},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/ROBOT.2007.364084"},{"key":"e_1_3_1_25_2","first-page":"1195","volume-title":"IEEE\/CVF International Conference on Computer Vision (ICCV)","author":"Gasparini S.","year":"2009","unstructured":"S. Gasparini, P. F. Sturm, and J. P. Barreto. 2009. Plane-based calibration of central catadioptric cameras. In IEEE\/CVF International Conference on Computer Vision (ICCV), 1195\u20131202."},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-010-0411-1"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/34.888718"},{"key":"e_1_3_1_28_2","first-page":"4137","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Zhang M.","year":"2015","unstructured":"M. Zhang, J. Yao, M. Xia, K. Li, Y. Zhang, and Y. Liu. 2015. Line-based multi-label energy optimization for fisheye image rectification and calibration. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 4137\u20134145."},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2005.163"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.rse.2009.01.007"},{"issue":"1","key":"e_1_3_1_31_2","doi-asserted-by":"crossref","first-page":"14","DOI":"10.1007\/PL00013269","article-title":"Straight lines have to be straight: Automatic calibration and removal of distortion from scenes of structured environments","volume":"13","author":"Devernay F.","year":"2001","unstructured":"F. Devernay and O. Faugeras. 2001. Straight lines have to be straight: Automatic calibration and removal of distortion from scenes of structured environments. Machine Vision Appl. 13, 1 (2001), 14\u201324.","journal-title":"Machine Vision Appl"},{"key":"e_1_3_1_32_2","first-page":"12384","volume-title":"IEEE\/CVF International Conference on Computer Vision (ICCV)","author":"Feng H.","year":"2023","unstructured":"H. Feng, W. Wang, J. Deng, W. Zhou, L. Li, and L. Li. 2023. Simfir: A simple framework for fisheye image rectification with self-supervised representation learning. In IEEE\/CVF International Conference on Computer Vision (ICCV), 12384\u201312393."},{"key":"e_1_3_1_33_2","first-page":"12653","volume-title":"IEEE\/CVF International Conference on Computer Vision (ICCV)","author":"Yang S.","year":"2023","unstructured":"S. Yang, C. Lin, K. Liao, and Y. Zhao. 2023. Innovating real fisheye image correction with dual diffusion architecture. In IEEE\/CVF International Conference on Computer Vision (ICCV), 12653\u201312662."},{"key":"e_1_3_1_34_2","first-page":"6344","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Yang S.","year":"2021","unstructured":"S. Yang, C. Lin, K. Liao, C. Zhang, and Y. Zhao. 2021. Progressively complementary network for fisheye image rectification using appearance flow. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 6344\u20136353."},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1016\/0167-8655(94)00115-J"},{"key":"e_1_3_1_36_2","first-page":"17662","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Wang Z.","year":"2021","unstructured":"Z. Wang, X. Cun, J. Bao, and J. Liu. 2021. Uformer: A general u-shaped transformer for image restoration. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 17662\u201317672."},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2021.3119293"},{"key":"e_1_3_1_38_2","first-page":"6:1","volume-title":"European Conference on Visual Media Production (CVMP)","author":"Bogdan O.","year":"2018","unstructured":"O. Bogdan, V. Eckstein, F. Rameau, and J. Bazin. 2018. Deepcalib: A deep learning approach for automatic intrinsic calibration of wide field-of-view cameras. In European Conference on Visual Media Production (CVMP), 6:1\u20136:10."},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2017.2723009"},{"key":"e_1_3_1_40_2","first-page":"3730","volume-title":"IEEE International Conference on Computer Vision (ICCV)","author":"Liu Z.","year":"2014","unstructured":"Z. Liu, P. Luo, X. Wang, and X. Tang. 2014. Deep learning face attributes in the wild. In IEEE International Conference on Computer Vision (ICCV), 3730\u20133738."},{"key":"e_1_3_1_41_2","first-page":"1541","volume-title":"IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Eichenseer A.","year":"2016","unstructured":"A. Eichenseer and Andr\u00e9 Kaup. 2016. A data set providing synthetic and real-world fisheye video sequences. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 1541\u20131545."},{"key":"e_1_3_1_42_2","first-page":"9308","volume-title":"IEEE\/CVF International Conference on Computer Vision (ICCV)","author":"Yogamani S.","year":"2019","unstructured":"S. Yogamani, C. Witt, H. Rashed, S. Nayak, S. Mansoor, P. Varley, X. Perrotton, D. Odea, P. Perez, C. Hughes, et al. 2019. Woodscape: A multi-task, multi-camera fisheye dataset for autonomous driving. In IEEE\/CVF International Conference on Computer Vision (ICCV), 9308\u20139318."},{"key":"e_1_3_1_43_2","first-page":"1398","volume-title":"Conference Record of the 37th Asilomar Conference on Signals, Systems and Computers","author":"Wang Z.","year":"2003","unstructured":"Z. Wang, E. P. Simoncelli, and A. C. Bovik. 2003. Multi-scale structural similarity for image quality assessment. In Conference Record of the 37th Asilomar Conference on Signals, Systems and Computers, 1398\u20131402."},{"key":"e_1_3_1_44_2","first-page":"6626","article-title":"GANs trained by a two time-scale update rule converge to a local Nash equilibrium","author":"Martin H.","year":"2017","unstructured":"H. Martin, R. Hubert, U. Thomas, N. Bernhard, and S. Hochreiter. 2017. GANs trained by a two time-scale update rule converge to a local Nash equilibrium. In Neural Information Processing Systems (NIPS), 6626\u20136637.","journal-title":"Neural Information Processing Systems (NIPS)"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2009.2025923"},{"key":"e_1_3_1_46_2","volume-title":"Pattern Classification and Scene Analysis","author":"Duda R. O.","year":"1974","unstructured":"R. O. Duda and P. E. Hart. 1974. Pattern Classification and Scene Analysis. Wiley-Interscience, New York."},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1023\/B:VISI.0000029664.99615.94"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/358669.358692"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3719348","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3719348","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T18:43:21Z","timestamp":1750272201000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3719348"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,4,7]]},"references-count":47,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2025,4,30]]}},"alternative-id":["10.1145\/3719348"],"URL":"https:\/\/doi.org\/10.1145\/3719348","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,4,7]]},"assertion":[{"value":"2023-11-14","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-02-17","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-04-07","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}