{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T08:18:54Z","timestamp":1783066734843,"version":"3.54.6"},"reference-count":42,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T00:00:00Z","timestamp":1783036800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"DOI":"10.13039\/100000001","name":"NSF","doi-asserted-by":"publisher","award":["IIS2335492"],"award-info":[{"award-number":["IIS2335492"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000001","name":"NSF","doi-asserted-by":"publisher","award":["OAC2403239"],"award-info":[{"award-number":["OAC2403239"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"name":"CSAIL Future of Data"},{"name":"MIT\u2013IBM Watson AI Laboratory"},{"name":"FinTechAI programs"},{"name":"MIT Generative AI Impact Consortium"},{"name":"Toyota\u2013CSAIL Joint Research Center"},{"name":"Schmidt Sciences"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2026,7,3]]},"abstract":"<jats:p>We present a GPU-based system for automatic differentiation (AD) of functions defined on triangle meshes, designed to exploit the locality and sparsity in mesh-based computation. Our system evaluates derivatives using perelement forward-mode AD, confining all computation to registers and shared memory and assembling global gradients, sparse Jacobians, and sparse Hessians directly on the GPU. By avoiding global computation graphs, intermediate buffers, and device-host synchronization, our approach minimizes memory traffic and enables efficient differentiation under both static and dynamically changing sparsity. Our programming model lets users express energy terms over mesh neighborhoods, while our system automatically manages parallel execution, derivative propagation, sparse assembly, and matrix-free operations such as Hessian-vector products. Our system supports both scalar- and vector-valued objectives, dynamic interaction-driven sparsity updates, and seamless integration with external GPU sparse linear solvers. We evaluate our system on applications including elastic and cloth simulation, surface parameterization, mesh smoothing, frame field design, ARAP deformation, and spherical manifold optimization. Across these tasks, our system consistently outperforms state-of-the-art differentiation frameworks, including PyTorch, JAX, Warp, Dr.JIT, EnzymeAD, and Thallo. We demonstrate speedups across a range of solver types, from Newton and Gauss-Newton for nonlinear least squares to L-BFGS and gradient descent, and across different derivative usage modes, including Hessian-vector products as well as full sparse Hessian and Jacobian construction. Our system is available as open source at https:\/\/github.com\/owensgroup\/RXMesh.<\/jats:p>","DOI":"10.1145\/3811338","type":"journal-article","created":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T07:05:51Z","timestamp":1783062351000},"page":"1-18","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Locality-Aware Automatic Differentiation on the GPU for Mesh-Based Computations"],"prefix":"10.1145","volume":"45","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1857-913X","authenticated-orcid":false,"given":"Ahmed","family":"Mahmoud","sequence":"first","affiliation":[{"name":"Massachusetts Institute of Technology (MIT), Cambridge, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9564-4022","authenticated-orcid":false,"given":"Rahul","family":"Goel","sequence":"additional","affiliation":[{"name":"Massachusetts Institute of Technology (MIT), Cambridge, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6243-9543","authenticated-orcid":false,"given":"Jonathan","family":"Ragan-Kelley","sequence":"additional","affiliation":[{"name":"Massachusetts Institute of Technology (MIT), Cambridge, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7701-7586","authenticated-orcid":false,"given":"Justin","family":"Solomon","sequence":"additional","affiliation":[{"name":"Massachusetts Institute of Technology (MIT), Cambridge, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,7,3]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/3620665.3640366"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.5555\/3600270.3600648"},{"key":"e_1_2_1_3_1","volume-title":"Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang.","author":"Bradbury James","year":"2018","unstructured":"James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang. 2018. JAX: composable transformations of Python+NumPy programs. http:\/\/github.com\/jax-ml\/jax"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/73833.73858"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/3132188"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/2766906"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/3764928"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/229473.229474"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.5555\/1455489"},{"key":"e_1_2_1_10_1","volume-title":"Discrete Shells. In Proceedings of the 2003 ACM SIGGRAPH\/Eurographics Symposium on Computer Animation","author":"Grinspun Eitan","year":"2003","unstructured":"Eitan Grinspun, Anil N. Hirani, Mathieu Desbrun, and Peter Schr\u00f6der. 2003. Discrete Shells. In Proceedings of the 2003 ACM SIGGRAPH\/Eurographics Symposium on Computer Animation (San Diego, California) (SCA '03). Eurographics Association, Goslar, DEU, 62\u201367."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/3687986"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/3520484"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/3657648"},{"key":"e_1_2_1_14_1","unstructured":"Wenzel Jakob. 2019. Enoki: structured vectorization and differentiation on modern processor architectures. https:\/\/github.com\/mitsuba-renderer\/enoki."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/3528223.3530099"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/3388769.3407490"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3532720.3535628"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3386569.3392425"},{"key":"e_1_2_1_19_1","unstructured":"Minchen Li Chenfanfu Jiang and Zhaofeng Luo. 2024. Physics-Based Simulation. https:\/\/phys-sim-book.github.io\/"},{"key":"e_1_2_1_20_1","volume-title":"NVIDIA GPU Technology Conference (GTC).","author":"Macklin Miles","year":"2022","unstructured":"Miles Macklin. 2022. Warp: A High-performance Python Framework for GPU Simulation and Graphics. https:\/\/github.com\/nvidia\/warp. NVIDIA GPU Technology Conference (GTC)."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/3450626.3459748"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3731162"},{"key":"e_1_2_1_23_1","volume-title":"Introduction to Solid Modeling","author":"M\u00e4ntyl\u00e4 M.","unstructured":"M. M\u00e4ntyl\u00e4. 1988. Introduction to Solid Modeling. W. H. Freeman & Co., New York, NY, USA."},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3453986"},{"key":"e_1_2_1_25_1","volume-title":"Martins and Andrew Ning","author":"Joaquim R. R.","year":"2021","unstructured":"Joaquim R. R. A. Martins and Andrew Ning. 2021. Engineering Design Optimization. Cambridge University Press."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3458817.3476165"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.5555\/3571885.3571964"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1137\/1.9781611972078"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-0-387-40065-5"},{"key":"e_1_2_1_30_1","volume-title":"CUB: Cooperative primitives for CUDA C++. https:\/\/nvidia.github.io\/cccl\/cub\/.","author":"NVIDIA Corporation","year":"2025","unstructured":"NVIDIA Corporation. 2025a. CUB: Cooperative primitives for CUDA C++. https:\/\/nvidia.github.io\/cccl\/cub\/."},{"issue":"7","key":"e_1_2_1_31_1","first-page":"1","article-title":"cuDSS","volume":"0","author":"NVIDIA Corporation","year":"2025","unstructured":"NVIDIA Corporation. 2025b. cuDSS: Release 0.7.1. https:\/\/developer.nvidia.com\/cudss\/.","journal-title":"Release"},{"key":"e_1_2_1_32_1","unstructured":"NVIDIA Corporation. 2025c. Thrust. https:\/\/nvidia.github.io\/cccl\/thrust\/index.html."},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.20380\/GI1995.17"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1111\/cgf.14607"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/1015706.1015812"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/2343483.2343501"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/2766947"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.5555\/1281991.1282006"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.gmod.2013.07.001"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/1073368.1073394"},{"key":"e_1_2_1_41_1","unstructured":"Ingo Wald et al. 2026. cuBQL: A CUDA BVH Build-and-Query Library. https:\/\/github.com\/NVIDIA\/cuBQL."},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/3550454.3555430"}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/abs\/10.1145\/3811338","content-type":"text\/html","content-version":"vor","intended-application":"syndication"}],"deposited":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T07:51:02Z","timestamp":1783065062000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3811338"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,7,3]]},"references-count":42,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,7,3]]}},"alternative-id":["10.1145\/3811338"],"URL":"https:\/\/doi.org\/10.1145\/3811338","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,7,3]]},"assertion":[{"value":"2026-01-21","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-27","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-07-03","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}