{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,2]],"date-time":"2026-06-02T09:29:53Z","timestamp":1780392593105,"version":"3.54.1"},"reference-count":49,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2014,3,1]],"date-time":"2014-03-01T00:00:00Z","timestamp":1393632000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100000811","name":"European Institute of Innovation and Technology","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100000811","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100006785","name":"Google","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100006785","id-type":"DOI","asserted-by":"publisher"}]},{"name":"MSR-INRIA Laboratory"},{"DOI":"10.13039\/501100001665","name":"Agence Nationale de la Recherche","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100001665","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2014,3]]},"abstract":"<jats:p>\n            This article describes a technique that can reliably align arbitrary 2D depictions of an architectural site, including drawings, paintings, and historical photographs, with a 3D model of the site. This is a tremendously difficult task, as the appearance and scene structure in the 2D depictions can be very different from the appearance and geometry of the 3D model, for example, due to the specific rendering style, drawing error, age, lighting, or change of seasons. In addition, we face a hard search problem: the number of possible alignments of the painting to a large 3D model, such as a partial reconstruction of a city, is huge. To address these issues, we develop a new compact representation of complex 3D scenes. The 3D model of the scene is represented by a small set of\n            <jats:italic>discriminative visual elements<\/jats:italic>\n            that are automatically learned from rendered views. Similar to object detection, the set of visual elements, as well as the weights of individual features for each element, are learned in a discriminative fashion. We show that the learned visual elements are reliably matched in 2D depictions of the scene despite large variations in rendering style (e.g., watercolor, sketch, historical photograph) and structural changes (e.g., missing scene parts, large occluders) of the scene. We demonstrate an application of the proposed approach to automatic rephotography to find an approximate viewpoint of historical paintings and photographs with respect to a 3D model of the site. The proposed alignment procedure is validated via a human user study on a new database of paintings and sketches spanning several sites. The results demonstrate that our algorithm produces significantly better alignments than several baseline methods.\n          <\/jats:p>","DOI":"10.1145\/2591009","type":"journal-article","created":{"date-parts":[[2014,4,15]],"date-time":"2014-04-15T17:50:29Z","timestamp":1397584229000},"page":"1-14","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":68,"title":["Painting-to-3D model alignment via discriminative visual elements"],"prefix":"10.1145","volume":"33","author":[{"given":"Mathieu","family":"Aubry","sequence":"first","affiliation":[{"name":"INRIA and TU M\u00fcnchen, France"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Bryan C.","family":"Russell","sequence":"additional","affiliation":[{"name":"Intel Labs"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Josef","family":"Sivic","sequence":"additional","affiliation":[{"name":"INRIA, France"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2014,4,8]]},"reference":[{"key":"e_1_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2007.1024"},{"key":"e_1_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.5555\/2964398.2964437"},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2011.5995727"},{"key":"e_1_2_2_4_1","volume-title":"Diffrac: A discriminative and flexible framework for clustering. In Advances in Neural Information Processing Systems.","author":"Bach F.","year":"2008"},{"key":"e_1_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/1805964.1805968"},{"key":"e_1_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/1778765.1778824"},{"key":"e_1_2_2_7_1","volume-title":"Pattern Recognition and Machine Learning","author":"Bishop C. M."},{"key":"e_1_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.aei.2009.08.006"},{"key":"e_1_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2006.125"},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2005.177"},{"key":"e_1_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2013.237"},{"key":"e_1_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/237170.237191"},{"key":"e_1_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/2185520.2185597"},{"key":"e_1_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.5555\/1390681.1442794"},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2009.167"},{"key":"e_1_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/358669.358692"},{"key":"e_1_2_2_17_1","volume-title":"Proceedings of the International Conference on Computer Vision.","author":"Frome A."},{"key":"e_1_2_2_18_1","volume-title":"Proceedings of the Conference on Computer Vision and Pattern Recognition.","author":"Furukawa Y."},{"key":"e_1_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2009.161"},{"key":"e_1_2_2_20_1","unstructured":"M. Gharbi T. Malisiewicz S. Paris and F. Durand. 2012. A Gaussian approximation of feature space for fast image similarity. Tech. rep. MIT-CSAIL-TR-2012-032. http:\/\/people.csail.mit.edu\/tomasz\/papers\/gharbi_techreport_2012.pdf.  M. Gharbi T. Malisiewicz S. Paris and F. Durand. 2012. A Gaussian approximation of feature space for fast image similarity. Tech. rep. MIT-CSAIL-TR-2012-032. http:\/\/people.csail.mit.edu\/tomasz\/papers\/gharbi_techreport_2012.pdf."},{"key":"e_1_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-33765-9_33"},{"key":"e_1_2_2_22_1","doi-asserted-by":"crossref","unstructured":"R. I. Hartley and A. Zisserman. 2004. Multiple View Geometry in Computer Vision 2nd Ed. Cambridge University Press.   R. I. Hartley and A. Zisserman. 2004. Multiple View Geometry in Computer Vision 2 nd Ed. Cambridge University Press.","DOI":"10.1017\/CBO9780511811685"},{"key":"e_1_2_2_23_1","volume-title":"Proceedings of the Conference on Computer Vision and Pattern Recognition.","author":"Hauagge D."},{"key":"e_1_2_2_24_1","volume-title":"Proceedings of the International Conference on Computer Vision.","author":"Huttenlocher D. P."},{"key":"e_1_2_2_25_1","volume-title":"Proceedings of the Conference on Computer Vision and Pattern Recognition.","author":"Irschara A."},{"key":"e_1_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2013.332"},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2013.124"},{"key":"e_1_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCOM.1967.1089532"},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/1409060.1409069"},{"key":"e_1_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/253607.253631"},{"key":"e_1_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-33718-5_2"},{"key":"e_1_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1007\/BF00128526"},{"key":"e_1_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1023\/B:VISI.0000029664.99615.94"},{"key":"e_1_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2011.6126229"},{"key":"e_1_2_2_35_1","volume-title":"et al","author":"Musialski P.","year":"2012"},{"key":"e_1_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1011139631724"},{"key":"e_1_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1080\/13602360802573868"},{"key":"e_1_2_2_38_1","volume-title":"Proceedings of the IEEE Workshop on 3D Representation for Recognition (3dRR'11)","author":"Russell B. C."},{"key":"e_1_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2011.6126302"},{"key":"e_1_2_2_40_1","volume-title":"Proceedings of the Conference on Computer Vision and Pattern Recognition.","author":"Schindler G."},{"key":"e_1_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10107-010-0420-4"},{"key":"e_1_2_2_42_1","volume-title":"Proceedings of the Conference on Computer Vision and Pattern Recognition.","author":"Shechtman E."},{"key":"e_1_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/2070781.2024188"},{"key":"e_1_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.5555\/2964398.2964405"},{"key":"e_1_2_2_45_1","volume-title":"Proceedings of the International Conference on Computer Vision.","author":"Sivic J."},{"key":"e_1_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/1141911.1141964"},{"key":"e_1_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1561\/0600000009"},{"key":"e_1_2_2_48_1","volume-title":"European Workshop on 3D Structure from Multiple Images of Large-Scale Environments (SMILE'98)","author":"Szeliski R."},{"key":"e_1_2_2_49_1","volume-title":"Proceedings of the Conference on Computer Vision and Pattern Recognition.","author":"Wu C."}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2591009","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2591009","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T20:01:13Z","timestamp":1750276873000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2591009"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2014,3]]},"references-count":49,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2014,3]]}},"alternative-id":["10.1145\/2591009"],"URL":"https:\/\/doi.org\/10.1145\/2591009","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2014,3]]},"assertion":[{"value":"2013-09-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2013-11-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2014-04-08","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}