Saved in:
Bibliographic Details
Main Authors: Mahmoud, Ahmed H., Goel, Rahul, Ragan-Kelley, Jonathan, Solomon, Justin
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2509.00406
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917237572501504
author Mahmoud, Ahmed H.
Goel, Rahul
Ragan-Kelley, Jonathan
Solomon, Justin
author_facet Mahmoud, Ahmed H.
Goel, Rahul
Ragan-Kelley, Jonathan
Solomon, Justin
contents We present a GPU-based system for automatic differentiation (AD) of functions defined on triangle meshes, designed to exploit the locality and sparsity in mesh-based computation. Our system evaluates derivatives using per-element forward-mode AD, confining all computation to registers and shared memory and assembling global gradients, sparse Jacobians, and sparse Hessians directly on the GPU. By avoiding global computation graphs, intermediate buffers, and device-host synchronization, our approach minimizes memory traffic and enables efficient differentiation under both static and dynamically changing sparsity. Our programming model lets users express energy terms over mesh neighborhoods, while our system automatically manages parallel execution, derivative propagation, sparse assembly, and matrix-free operations such as Hessian-vector products. Our system supports both scalar and vector-valued objectives, dynamic interaction-driven sparsity updates, and seamless integration with external GPU sparse linear solvers. We evaluate our system on applications including elastic and cloth simulation, surface parameterization, mesh smoothing, frame field design, ARAP deformation, and spherical manifold optimization. Across these tasks, our system consistently outperforms state-of-the-art differentiation frameworks, including PyTorch, JAX, Warp, Dr.JIT, and Thallo. We demonstrate speedups across a range of solver types, from Newton and Gauss-Newton for nonlinear least squares to L-BFGS and gradient descent, and across different derivative usage modes, including Hessian-vector products as well as full sparse Hessian and Jacobian construction.
format Preprint
id arxiv_https___arxiv_org_abs_2509_00406
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Locality-Aware Automatic Differentiation on the GPU for Mesh-Based Computations
Mahmoud, Ahmed H.
Goel, Rahul
Ragan-Kelley, Jonathan
Solomon, Justin
Graphics
We present a GPU-based system for automatic differentiation (AD) of functions defined on triangle meshes, designed to exploit the locality and sparsity in mesh-based computation. Our system evaluates derivatives using per-element forward-mode AD, confining all computation to registers and shared memory and assembling global gradients, sparse Jacobians, and sparse Hessians directly on the GPU. By avoiding global computation graphs, intermediate buffers, and device-host synchronization, our approach minimizes memory traffic and enables efficient differentiation under both static and dynamically changing sparsity. Our programming model lets users express energy terms over mesh neighborhoods, while our system automatically manages parallel execution, derivative propagation, sparse assembly, and matrix-free operations such as Hessian-vector products. Our system supports both scalar and vector-valued objectives, dynamic interaction-driven sparsity updates, and seamless integration with external GPU sparse linear solvers. We evaluate our system on applications including elastic and cloth simulation, surface parameterization, mesh smoothing, frame field design, ARAP deformation, and spherical manifold optimization. Across these tasks, our system consistently outperforms state-of-the-art differentiation frameworks, including PyTorch, JAX, Warp, Dr.JIT, and Thallo. We demonstrate speedups across a range of solver types, from Newton and Gauss-Newton for nonlinear least squares to L-BFGS and gradient descent, and across different derivative usage modes, including Hessian-vector products as well as full sparse Hessian and Jacobian construction.
title Locality-Aware Automatic Differentiation on the GPU for Mesh-Based Computations
topic Graphics
url https://arxiv.org/abs/2509.00406