{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,13]],"date-time":"2026-08-13T12:28:52Z","timestamp":1786624132260,"version":"3.56.0"},"reference-count":66,"publisher":"Wiley","license":[{"start":{"date-parts":[[2026,8,13]],"date-time":"2026-08-13T00:00:00Z","timestamp":1786579200000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"},{"start":{"date-parts":[[2026,8,13]],"date-time":"2026-08-13T00:00:00Z","timestamp":1786579200000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/doi.wiley.com\/10.1002\/tdm_license_1.1"}],"funder":[{"DOI":"10.13039\/501100002954","name":"Universit\u00e0 degli Studi di Milano-Bicocca","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100002954","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Computer Graphics Forum"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Transformers are sequence\u2010to\u2010sequence architectures originally designed to handle structurally rigid and order\u2010sensitive data, such as text and images. At their core, they exploit the attention mechanism, which is permutation\u2010equivariant and relies on computing token\u2010to\u2010token relationships. These models have been applied to 3D geometry in several instances, achieving discrete success across tasks such as shape generation, segmentation, classification, shape matching, and registration. While existing 3D geometry methods use transformers as traditional learners, we present an approach that reinterprets the transformer as an optimization pipeline for shape correspondence. By fitting the model directly to a shape pair, our method eliminates the need for large training datasets, providing a category\u2010agnostic solution. In particular, we focus on the use of attention weights, tailored to encode token\u2010to\u2010token information, to inject and extract point\u2010to\u2010point information during the processing of one or more meshes. We demonstrate, that self\u2010 and cross\u2010attention mechanisms can, by design, serve as feature extractors and matching solvers, respectively. Furthermore, instead of deriving correspondence from the final output of the network, we exploit the cross\u2010attention weights directly as the permutation matrix. This framework not only achieves robust shape matching and registration but also provides a theoretically grounded, interpretable approach to attention for unstructured 3D data. Notably, our work represents an approach that leverages the transformer architecture as an end\u2010to\u2010end pipeline for shape correspondence, operating effectively without requiring additional training data.<\/jats:p>","DOI":"10.1111\/cgf.70525","type":"journal-article","created":{"date-parts":[[2026,8,13]],"date-time":"2026-08-13T11:36:22Z","timestamp":1786620982000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Attention Based Optimization for 3D Shape Registration"],"prefix":"10.1111","author":[{"ORCID":"https:\/\/orcid.org\/0009-0007-0177-0604","authenticated-orcid":false,"given":"A.","family":"Riva","sequence":"first","affiliation":[{"name":"University of Milano\u2010Bicocca  Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-7290-3549","authenticated-orcid":false,"given":"L.","family":"Olearo","sequence":"additional","affiliation":[{"name":"University of Milano\u2010Bicocca  Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2790-9591","authenticated-orcid":false,"given":"S.","family":"Melzi","sequence":"additional","affiliation":[{"name":"University of Milano\u2010Bicocca  Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"311","published-online":{"date-parts":[[2026,8,13]]},"reference":[{"key":"e_1_2_9_2_2","doi-asserted-by":"crossref","unstructured":"ArnabA. DehghaniM. HeigoldG. SunC. Lu\u010di\u0107M. SchmidC.: Vivit: A video vision transformer. In2021IEEE\/CVF International Conference on Computer Vision (ICCV)(2021) pp.6816\u20136826. 3","DOI":"10.1109\/ICCV48922.2021.00676"},{"key":"e_1_2_9_3_2","doi-asserted-by":"crossref","first-page":"1626","DOI":"10.1109\/ICCVW.2011.6130444","volume-title":"Computer Vision Workshops (ICCV Workshops), 2011 IEEE International Conference on","author":"Aubry M.","year":"2011"},{"key":"e_1_2_9_4_2","first-page":"561","volume-title":"European conference on computer vision","author":"Bogo F.","year":"2016"},{"key":"e_1_2_9_5_2","first-page":"1877","volume-title":"Advances in Neural Information Processing Systems","author":"Brown T.","year":"2020"},{"issue":"3","key":"e_1_2_9_6_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/1531326.1531379","article-title":"A benchmark for 3d mesh segmentation","volume":"28","author":"Chen X.","year":"2009","journal-title":"Acm transactions on graphics (tog)"},{"issue":"4","key":"e_1_2_9_7_2","article-title":"Unsupervised learning of robust spectral shape matching","volume":"42","author":"Cao D.","year":"2023","journal-title":"ACM Trans. Graph."},{"key":"e_1_2_9_8_2","first-page":"35549","volume":"2024","author":"Dao T.","year":"2024","journal-title":"International Conference on Learning Representations"},{"key":"e_1_2_9_9_2","doi-asserted-by":"crossref","unstructured":"DonatiN. CormanE. OvsjanikovM.: Deep orientation\u2010aware functional maps: Tackling symmetry issues in shape matching.CVPR(2022). 3 10","DOI":"10.1109\/CVPR52688.2022.00082"},{"key":"e_1_2_9_10_2","doi-asserted-by":"crossref","first-page":"28","DOI":"10.1016\/j.cag.2020.08.008","article-title":"Shrec'20: Shape correspondence with non\u2010isometric deformations","volume":"92","author":"Dyke R. M.","year":"2020","journal-title":"Computers & Graphics"},{"key":"e_1_2_9_11_2","doi-asserted-by":"crossref","unstructured":"DuttN. S. MuralikrishnanS. MitraN. J.: Diffusion 3d features (diff3f): Decorating untextured shapes with distilled semantic features. InProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(2024) pp.4494\u20134504. 5 9","DOI":"10.1109\/CVPR52733.2024.00430"},{"key":"e_1_2_9_12_2","unstructured":"DosovitskiyA.: An image is worth 16x16 words: Transformers for image recognition at scale.arXiv preprint arXiv:2010.11929(2020). 3"},{"key":"e_1_2_9_13_2","unstructured":"HuangQ. et al.: Transformers in medical imaging: A survey.Medical Image Analysis(2023). 3"},{"key":"e_1_2_9_14_2","doi-asserted-by":"crossref","first-page":"2734","DOI":"10.1109\/TIP.2023.3272821","article-title":"Hierarchical shape\u2010consistent transformer for unsupervised point cloud shape correspondence","volume":"32","author":"He J.","year":"2023","journal-title":"IEEE Transactions on Image Processing"},{"key":"e_1_2_9_15_2","doi-asserted-by":"crossref","first-page":"252","DOI":"10.1109\/3DV50981.2020.00035","volume-title":"2020 International Conference on 3D Vision (3DV)","author":"Holzschuh B.","year":"2020"},{"issue":"5","key":"e_1_2_9_16_2","doi-asserted-by":"crossref","first-page":"265","DOI":"10.1111\/cgf.14084","article-title":"Consistent zoomout: Efficient spectral map synchronization","volume":"39","author":"Huang R.","year":"2020","journal-title":"Computer Graphics Forum"},{"key":"e_1_2_9_17_2","doi-asserted-by":"crossref","first-page":"536","DOI":"10.1109\/ICARA60736.2024.10552931","volume-title":"2024 10th International Conference on Automation, Robotics and Applications (ICARA)","author":"Huang H.","year":"2024"},{"key":"e_1_2_9_18_2","doi-asserted-by":"crossref","unstructured":"JiY. et al.: Dnabert: pre\u2010trained bidirectional encoder representations from transformers for dna\u2010language in genome.Bioinformatics(2021). 3","DOI":"10.1101\/2020.09.17.301879"},{"key":"e_1_2_9_19_2","unstructured":"JumperJ. et al.: Highly accurate protein structure prediction with alphafold.Nature(2021). 3"},{"key":"e_1_2_9_20_2","first-page":"4651","volume-title":"International conference on machine learning","author":"Jaegle A.","year":"2021"},{"key":"e_1_2_9_21_2","first-page":"79","volume":"30","author":"Kim V. G.","year":"2011","journal-title":"ACM Transactions on Graphics (TOG)"},{"key":"e_1_2_9_22_2","unstructured":"LiY. et al.: Competition\u2010level code generation with alphacode. InScience(2022). 3"},{"key":"e_1_2_9_23_2","doi-asserted-by":"crossref","first-page":"27757","DOI":"10.52202\/068431-2013","article-title":"Non\u2010rigid point cloud registration with neural deformation pyramid","volume":"35","author":"Li Y.","year":"2022","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_9_24_2","unstructured":"LiZ. LiuX. DrenkowN. DingA. CreightonF. X. TaylorR. H. UnberathM.: Revisiting stereo depth estimation from a sequence\u2010to\u2010sequence perspective with transformers. InProceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV)(October2021) pp.6197\u20136206. 2"},{"key":"e_1_2_9_25_2","unstructured":"LinJ. LongH. GuoH. ZhangJ. YangJ. GuoT. YangY. LiJ. ZhangW. NiessnerM. YangW.: Meshripple: Structured autoregressive generation of artist\u2010meshes. InIEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)(2026). 3"},{"key":"e_1_2_9_26_2","doi-asserted-by":"crossref","unstructured":"LinY. LinC. PanP. YanH. FengY. MuY. FragkiadakiK.: Partcrafter: Structured 3d mesh generation via compositional latent diffusion transformers. InAdvances in Neural Information Processing Systems (NeurIPS)(2025). 3","DOI":"10.52202\/085713-1189"},{"issue":"6","key":"e_1_2_9_27_2","first-page":"248:1","article-title":"SMPL: A skinned multi\u2010person linear model","volume":"34","author":"Loper M.","year":"2015","journal-title":"ACM Trans. Graphics (Proc. SIGGRAPH Asia)"},{"key":"e_1_2_9_28_2","first-page":"55","volume-title":"Eurographics Workshop on 3D Object Retrieval, EG 3DOR","author":"L\u00e4hner Z.","year":"2016"},{"key":"e_1_2_9_29_2","doi-asserted-by":"crossref","unstructured":"LiangD. ZhouX. XuW. ZhuX. ZouZ. YeX. TanX. BaiX.: Pointmamba: A simple state space model for point cloud analysis. InAdvances in Neural Information Processing Systems(2024). 3","DOI":"10.52202\/079017-1026"},{"key":"e_1_2_9_30_2","first-page":"183","volume-title":"European Conference on Computer Vision","author":"Maggioli F.","year":"2024"},{"issue":"1","key":"e_1_2_9_31_2","doi-asserted-by":"crossref","first-page":"4:1","DOI":"10.1145\/3144454","article-title":"Discrete time evolution process descriptor for shape analysis and matching","volume":"37","author":"Melzi S.","year":"2018","journal-title":"ACM Transactions on Graphics (TOG)"},{"key":"e_1_2_9_32_2","volume-title":"Correspondence learning via linearly\u2010invariant embedding","author":"Marin R."},{"issue":"6","key":"e_1_2_9_33_2","doi-asserted-by":"crossref","first-page":"155:1","DOI":"10.1145\/3355089.3356524","article-title":"Zoomout: Spectral upsampling for efficient shape correspondence","volume":"38","author":"Melzi S.","year":"2019","journal-title":"ACM Transactions on Graphics (TOG)"},{"key":"e_1_2_9_34_2","doi-asserted-by":"crossref","unstructured":"M\u00fcllerN. SiddiquiY. PorziL. BuloS. R. KontschiederP. NiessnerM.: Diffrf: Rendering\u2010guided 3d radiance field diffusion. InProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(2023) pp.4328\u20134338. 2 3","DOI":"10.1109\/CVPR52729.2023.00421"},{"issue":"2","key":"e_1_2_9_35_2","doi-asserted-by":"crossref","first-page":"179","DOI":"10.1111\/cgf.13352","article-title":"Improved functional mappings via product preservation","volume":"37","author":"Nogneng D.","year":"2018","journal-title":"Computer Graphics Forum"},{"issue":"2","key":"e_1_2_9_36_2","doi-asserted-by":"crossref","first-page":"259","DOI":"10.1111\/cgf.13124","article-title":"Informative descriptor preservation via commutativity for shape matching","volume":"36","author":"Nogneng D.","year":"2017","journal-title":"Computer Graphics Forum"},{"issue":"4","key":"e_1_2_9_37_2","doi-asserted-by":"crossref","first-page":"30:1","DOI":"10.1145\/2185520.2185526","article-title":"Functional maps: a flexible representation of maps between shapes","volume":"31","author":"Ovsjanikov M.","year":"2012","journal-title":"ACM Transactions on Graphics (TOG)"},{"key":"e_1_2_9_38_2","first-page":"394","volume-title":"Computer graphics forum","author":"Panine M.","year":"2022"},{"key":"e_1_2_9_39_2","doi-asserted-by":"crossref","first-page":"384","DOI":"10.1109\/CVPR46437.2021.00045","volume-title":"2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Pai G.","year":"2021"},{"key":"e_1_2_9_40_2","unstructured":"PardalosP. M. RendlF. WolkowiczH.:The quadratic assignment problem: A survey and recent developments. 3"},{"key":"e_1_2_9_41_2","doi-asserted-by":"crossref","first-page":"604","DOI":"10.1007\/978-3-031-20086-1_35","volume-title":"Computer Vision\u2013ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23\u201327, 2022, Proceedings, Part II","author":"Pang Y.","year":"2022"},{"key":"e_1_2_9_42_2","unstructured":"RadfordA. et al.: Language models are unsupervised multitask learners.OpenAI Blog(2019). 3"},{"key":"e_1_2_9_43_2","unstructured":"RongY. et al.: Self\u2010supervised graph transformer on large\u2010scale molecular data. InNeurIPS(2020). 3"},{"issue":"5","key":"e_1_2_9_44_2","doi-asserted-by":"crossref","first-page":"81","DOI":"10.1111\/cgf.14359","article-title":"Discrete optimization for shape matching","volume":"40","author":"Ren J.","year":"2021","journal-title":"Computer Graphics Forum"},{"key":"e_1_2_9_45_2","first-page":"e14912","volume-title":"Computer Graphics Forum","author":"Raganato A.","year":"2023"},{"issue":"6","key":"e_1_2_9_46_2","article-title":"Continuous and orientation\u2010preserving correspondences via functional maps","volume":"37","author":"Ren J.","year":"2018","journal-title":"ACM Transactions on Graphics (TOG)"},{"key":"e_1_2_9_47_2","article-title":"Random features for large\u2010scale kernel machines","volume":"20","author":"Rahimi A.","year":"2007","journal-title":"Advances in neural information processing systems"},{"key":"e_1_2_9_48_2","doi-asserted-by":"publisher","DOI":"10.2312\/stag.20241345"},{"key":"e_1_2_9_49_2","doi-asserted-by":"crossref","unstructured":"RoetzerP. SwobodaP. CremersD. BernardF.: A scalable combinatorial solver for elastic geometrically consistent 3d shape matching. InProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(2022) pp.428\u2013438. 3","DOI":"10.1109\/CVPR52688.2022.00052"},{"key":"e_1_2_9_50_2","doi-asserted-by":"crossref","unstructured":"SchwallerP. et al.: Molecular transformer: A model for uncertainty\u2010calibrated chemical reaction prediction.Chemical Science(2019). 2","DOI":"10.26434\/chemrxiv.7297379.v2"},{"issue":"3","key":"e_1_2_9_51_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3507905","article-title":"Diffusionnet: Discretization agnostic learning on surfaces","volume":"41","author":"Sharp N.","year":"2022","journal-title":"ACM Transactions on Graphics (ToG)"},{"key":"e_1_2_9_52_2","doi-asserted-by":"crossref","first-page":"127063","DOI":"10.1016\/j.neucom.2023.127063","article-title":"Roformer: Enhanced transformer with rotary position embedding","volume":"568","author":"Su J.","year":"2024","journal-title":"Neurocomputing"},{"key":"e_1_2_9_53_2","doi-asserted-by":"crossref","unstructured":"ShinJ. KimY. HongS. LeeJ.: Learning dual hierarchical representation for 3d surface reconstruction. InProceedings of the Asian Conference on Computer Vision (ACCV)(December2024) pp.4422\u20134438. 3","DOI":"10.1007\/978-981-96-0969-7_18"},{"issue":"5","key":"e_1_2_9_54_2","doi-asserted-by":"crossref","first-page":"1383","DOI":"10.1111\/j.1467-8659.2009.01515.x","article-title":"A concise and provably informative multi\u2010scale signature based on heat diffusion","volume":"28","author":"Sun J.","year":"2009","journal-title":"Computer graphics forum"},{"key":"e_1_2_9_55_2","first-page":"5731","article-title":"Shape registration in the time of transformers","volume":"34","author":"Trappolini G.","year":"2021","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_9_56_2","first-page":"356","volume-title":"Proc. ECCV","author":"Tombari F.","year":"2010"},{"key":"e_1_2_9_57_2","doi-asserted-by":"crossref","first-page":"517","DOI":"10.1109\/3DV.2017.00065","volume-title":"2017 international conference on 3D vision (3DV)","author":"Vestner M.","year":"2017"},{"key":"e_1_2_9_58_2","unstructured":"Vigan\u00f2G. LongariG. PereiraL. F. MiolaneN. MelziS.:Geomfum: A python package for machine learning with functional maps. 6"},{"key":"e_1_2_9_59_2","doi-asserted-by":"crossref","unstructured":"VestnerM. LitmanR. Rodol\u00e0E. BronsteinA. CremersD.: Product manifold filter: Non\u2010rigid shape correspondence via kernel density estimation in the product space. InProc. CVPR(2017) pp.6681\u20136690. 3","DOI":"10.1109\/CVPR.2017.707"},{"issue":"4","key":"e_1_2_9_60_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3730943","article-title":"Nam: Neural adjoint maps for refining shape correspondences","volume":"44","author":"Vigan\u00f2 G.","year":"2025","journal-title":"ACM Transactions on Graphics (TOG)"},{"key":"e_1_2_9_61_2","article-title":"Attention is all you need","volume":"30","author":"Vaswani A.","year":"2017","journal-title":"Advances in neural information processing systems"},{"key":"e_1_2_9_62_2","doi-asserted-by":"crossref","unstructured":"WuX. JiangL. WangP.\u2010S. LiuZ. LiuX. QiaoY. OuyangW. HeT. ZhaoH.: Point transformer v3: Simpler faster stronger. InProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)(2024). 3","DOI":"10.1109\/CVPR52733.2024.00463"},{"key":"e_1_2_9_63_2","doi-asserted-by":"crossref","unstructured":"WuX. LaoY. JiangL. LiuX. ZhaoH.: Point transformer v2: Grouped vector attention and partition\u2010based pooling. InAdvances in Neural Information Processing Systems(2022). 3","DOI":"10.52202\/068431-2415"},{"key":"e_1_2_9_64_2","unstructured":"YueY. RobertD. WangJ. HongS. WegnerJ. D. RupprechtC. SchindlerK.: LitePT: Lighter Yet Stronger Point Transformer. InIEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)(2026). 3"},{"key":"e_1_2_9_65_2","doi-asserted-by":"crossref","unstructured":"ZhangZ. GirdharR. JoulinA. MisraI.: Self\u2010supervised pretraining of 3d features on any point\u2010cloud. InICCV(2021). 3","DOI":"10.1109\/ICCV48922.2021.01009"},{"key":"e_1_2_9_66_2","unstructured":"ZhaoH. JiangL. JiaJ. TorrP. H. KoltunV.: Point transformer. InProceedings of the IEEE\/CVF International Conference on Computer Vision(2021) pp.16259\u201316268. 3"},{"key":"e_1_2_9_67_2","doi-asserted-by":"crossref","unstructured":"ZuffiS. KanazawaA. JacobsD. W. BlackM. J.: 3d menagerie: Modeling the 3d shape and pose of animals. InProceedings of the IEEE conference on computer vision and pattern recognition(2017) pp.6365\u20136373. 6","DOI":"10.1109\/CVPR.2017.586"}],"container-title":["Computer Graphics Forum"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/cgf.70525","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/full-xml\/10.1111\/cgf.70525","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/cgf.70525","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,8,13]],"date-time":"2026-08-13T11:36:39Z","timestamp":1786620999000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1111\/cgf.70525"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,8,13]]},"references-count":66,"alternative-id":["10.1111\/cgf.70525"],"URL":"https:\/\/doi.org\/10.1111\/cgf.70525","archive":["Portico"],"relation":{},"ISSN":["0167-7055","1467-8659"],"issn-type":[{"value":"0167-7055","type":"print"},{"value":"1467-8659","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,8,13]]},"assertion":[{"value":"2026-08-13","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e70525"}}