{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,23]],"date-time":"2025-08-23T05:25:17Z","timestamp":1755926717299,"version":"3.41.0"},"reference-count":45,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2022,11,30]],"date-time":"2022-11-30T00:00:00Z","timestamp":1669766400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Centre for Perceptual and Interactive Intelligence","award":["14207319"],"award-info":[{"award-number":["14207319"]}]},{"name":"Centre for Perceptual and Interactive Intelligence","award":["14204021"],"award-info":[{"award-number":["14204021"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2022,12]]},"abstract":"<jats:p>We tackle the problem of estimating correspondences from a general marker, such as a movie poster, to an image that captures such a marker. Conventionally, this problem is addressed by fitting a homography model based on sparse feature matching. However, they are only able to handle plane-like markers and the sparse features do not sufficiently utilize appearance information. In this paper, we propose a novel framework NeuralMarker, training a neural network estimating dense marker correspondences under various challenging conditions, such as marker deformation, harsh lighting, etc. Deep learning has presented an excellent performance in correspondence learning once provided with sufficient training data. However, annotating pixel-wise dense correspondence for training marker correspondence is too expensive. We observe that the challenges of marker correspondence estimation come from two individual aspects: geometry variation and appearance variation. We, therefore, design two components addressing these two challenges in NeuralMarker. First, we create a synthetic dataset FlyingMarkers containing marker-image pairs with ground truth dense correspondences. By training with FlyingMarkers, the neural network is encouraged to capture various marker motions. Second, we propose the novel Symmetric Epipolar Distance (SED) loss, which enables learning dense correspondence from posed images. Learning with the SED loss and the cross-lighting posed images collected by Structure-from-Motion (SfM), NeuralMarker is remarkably robust in harsh lighting environments and avoids synthetic image bias. Besides, we also propose a novel marker correspondence evaluation method circumstancing annotations on real marker-image pairs and create a new benchmark. We show that NeuralMarker significantly outperforms previous methods and enables new interesting applications, including Augmented Reality (AR) and video editing.<\/jats:p>","DOI":"10.1145\/3550454.3555468","type":"journal-article","created":{"date-parts":[[2022,11,30]],"date-time":"2022-11-30T21:19:07Z","timestamp":1669843147000},"page":"1-10","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":7,"title":["NeuralMarker"],"prefix":"10.1145","volume":"41","author":[{"given":"Zhaoyang","family":"Huang","sequence":"first","affiliation":[{"name":"The Chinese University of Hong Kong and NVIDIA, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaokun","family":"Pan","sequence":"additional","affiliation":[{"name":"Zhejiang University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Weihong","family":"Pan","sequence":"additional","affiliation":[{"name":"Zhejiang University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Weikang","family":"Bian","sequence":"additional","affiliation":[{"name":"The Chinese University of Hong Kong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yan","family":"Xu","sequence":"additional","affiliation":[{"name":"The Chinese University of Hong Kong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ka Chun","family":"Cheung","sequence":"additional","affiliation":[{"name":"NVIDIA, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Guofeng","family":"Zhang","sequence":"additional","affiliation":[{"name":"Zhejiang University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hongsheng","family":"Li","sequence":"additional","affiliation":[{"name":"The Chinese University of Hong Kong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,11,30]]},"reference":[{"key":"e_1_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/2001269.2001293"},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2005.475"},{"key":"e_1_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-33783-3_44"},{"key":"e_1_2_2_5_1","volume-title":"Dzmitry Bahdanau, and Yoshua Bengio.","author":"Cho Kyunghyun","year":"2014","unstructured":"Kyunghyun Cho, Bart Van Merri\u00ebnboer, Dzmitry Bahdanau, and Yoshua Bengio. 2014. On the properties of neural machine translation: Encoder-decoder approaches. arXiv preprint arXiv:1409.1259 (2014)."},{"key":"e_1_2_2_6_1","volume-title":"Twins: Revisiting the design of spatial attention in vision transformers. Advances in Neural Information Processing Systems 34","author":"Chu Xiangxiang","year":"2021","unstructured":"Xiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang, Haibing Ren, Xiaolin Wei, Huaxia Xia, and Chunhua Shen. 2021. Twins: Revisiting the design of spatial attention in vision transformers. Advances in Neural Information Processing Systems 34 (2021)."},{"key":"e_1_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.164"},{"key":"e_1_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2018.00060"},{"key":"e_1_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.316"},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00828"},{"key":"e_1_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/358669.358692"},{"key":"e_1_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1177\/0278364913491297"},{"key":"e_1_2_2_13_1","volume-title":"Learnable visual markers. Advances In Neural Information Processing Systems 29","author":"Grinchuk Oleg","year":"2016","unstructured":"Oleg Grinchuk, Vadim Lebedev, and Victor Lempitsky. 2016. Learnable visual markers. Advances In Neural Information Processing Systems 29 (2016)."},{"key":"e_1_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00863"},{"key":"e_1_2_2_15_1","volume-title":"Hongwei Qin, Jifeng Dai, and Hongsheng Li.","author":"Huang Zhaoyang","year":"2022","unstructured":"Zhaoyang Huang, Xiaoyu Shi, Chao Zhang, Qiang Wang, Ka Chun Cheung, Hongwei Qin, Jifeng Dai, and Hongsheng Li. 2022. FlowFormer: A Transformer Architecture for Optical Flow. arXiv preprint arXiv:2203.16194 (2022)."},{"key":"e_1_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00218"},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1023\/B:VISI.0000029664.99615.94"},{"key":"e_1_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2020.3023565"},{"key":"e_1_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.438"},{"key":"e_1_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/WACV.2019.00115"},{"key":"e_1_2_2_21_1","volume-title":"Dynamic projection mapping onto deforming non-rigid surface using deformable dot cluster marker","author":"Narita Gaku","year":"2016","unstructured":"Gaku Narita, Yoshihiro Watanabe, and Masatoshi Ishikawa. 2016. Dynamic projection mapping onto deforming non-rigid surface using deformable dot cluster marker. IEEE transactions on visualization and computer graphics 23, 3 (2016), 1235--1248."},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2011.5979561"},{"key":"e_1_2_2_23_1","unstructured":"OpenCV. 2015. Open Source Computer Vision Library."},{"key":"e_1_2_2_24_1","volume-title":"E2ETag: An End-to-End Trainable Method for Generating and Detecting Fiducial Markers. arXiv preprint arXiv:2105.14184","author":"Peace J Brennan","year":"2021","unstructured":"J Brennan Peace, Eric Psota, Yanfeng Liu, and Lance C P\u00e9rez. 2021. E2ETag: An End-to-End Trainable Method for Generating and Detecting Fiducial Markers. arXiv preprint arXiv:2105.14184 (2021)."},{"key":"e_1_2_2_25_1","volume-title":"GAN-Supervised Dense Visual Alignment. arXiv preprint arXiv:2112.05143","author":"Peebles William","year":"2021","unstructured":"William Peebles, Jun-Yan Zhu, Richard Zhang, Antonio Torralba, Alexei Efros, and Eli Shechtman. 2021. GAN-Supervised Dense Visual Alignment. arXiv preprint arXiv:2112.05143 (2021)."},{"key":"e_1_2_2_26_1","volume-title":"Garnett (Eds.)","volume":"32","author":"Revaud Jerome","year":"2019","unstructured":"Jerome Revaud, Cesar De Souza, Martin Humenberger, and Philippe Weinzaepfel. 2019. R2D2: Reliable and Repeatable Detector and Descriptor. In Advances in Neural Information Processing Systems, H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alch\u00e9-Buc, E. Fox, and R. Garnett (Eds.), Vol. 32. Curran Associates, Inc. https:\/\/proceedings.neurips.cc\/paper\/2019\/file\/3198dfd0aef271d22f7bcddd6f12f5cb-Paper.pdf"},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.12"},{"key":"e_1_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00499"},{"key":"e_1_2_2_29_1","series-title":"Journal of Physics: Conference Series","volume-title":"Developing augmented reality based application for character education using unity with Vuforia SDK","author":"Sarosa M","year":"2035","unstructured":"M Sarosa, A Chalim, S Suhari, Z Sari, and HB Hakim. 2019. Developing augmented reality based application for character education using unity with Vuforia SDK. In Journal of Physics: Conference Series, Vol. 1375. IOP Publishing, 012035."},{"key":"e_1_2_2_30_1","volume-title":"Structure-from-Motion Revisited. In Conference on Computer Vision and Pattern Recognition (CVPR).","author":"Sch\u00f6nberger Johannes Lutz","year":"2016","unstructured":"Johannes Lutz Sch\u00f6nberger and Jan-Michael Frahm. 2016. Structure-from-Motion Revisited. In Conference on Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_2_2_31_1","volume-title":"RANSAC-Flow: generic two-stage image alignment. arXiv preprint arXiv:2004.01526","author":"Shen Xi","year":"2020","unstructured":"Xi Shen, Fran\u00e7ois Darmon, Alexei A Efros, and Mathieu Aubry. 2020. RANSAC-Flow: generic two-stage image alignment. arXiv preprint arXiv:2004.01526 (2020)."},{"key":"e_1_2_2_32_1","unstructured":"Alexandro Simonetti Ibanez and Josep Paredes Figueras. 2013. Vuforia v1. 5 SDK: Analysis and evaluation of capabilities. Master's thesis. Universitat Polit\u00e8cnica de Catalunya."},{"key":"e_1_2_2_33_1","doi-asserted-by":"crossref","unstructured":"Richard Szeliski et al. 2007. Image alignment and stitching: A tutorial. Foundations and Trends\u00ae in Computer Graphics and Vision 2 1 (2007) 1--104.","DOI":"10.1561\/0600000009"},{"key":"e_1_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58536-5_24"},{"key":"e_1_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00629"},{"key":"e_1_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00566"},{"key":"e_1_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISMAR.2011.6092394"},{"key":"e_1_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-69321-5_32"},{"key":"e_1_2_2_39_1","volume-title":"Attention is all you need. Advances in neural information processing systems 30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017)."},{"key":"e_1_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/IROS.2016.7759617"},{"key":"e_1_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58452-8_44"},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2009.5459375"},{"key":"e_1_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/CRV.2011.13"},{"key":"e_1_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3450626.3459762"},{"key":"e_1_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2013.243"},{"key":"e_1_2_2_46_1","volume-title":"CVPR Workshops. 1--7.","author":"Yang Guandao","year":"2019","unstructured":"Guandao Yang, Tomasz Malisiewicz, Serge J Belongie, Erez Farhan, Sungsoo Ha, Yuewei Lin, Xiaojing Huang, Hanfei Yan, and Wei Xu. 2019. Learning Data-Adaptive Interest Points through Epipolar Adaptation.. In CVPR Workshops. 1--7."}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3550454.3555468","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3550454.3555468","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:49:11Z","timestamp":1750182551000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3550454.3555468"}},"subtitle":["A Framework for Learning General Marker Correspondence"],"short-title":[],"issued":{"date-parts":[[2022,11,30]]},"references-count":45,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2022,12]]}},"alternative-id":["10.1145\/3550454.3555468"],"URL":"https:\/\/doi.org\/10.1145\/3550454.3555468","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"type":"print","value":"0730-0301"},{"type":"electronic","value":"1557-7368"}],"subject":[],"published":{"date-parts":[[2022,11,30]]},"assertion":[{"value":"2022-11-30","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}