TopoMesh: High-Fidelity Mesh Autoencoding via Topological Unification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luo, Guan, Li, Xiu, Chen, Rui, Yi, Xuanyu, Lin, Jing, Chen, Chia-Hao, Liu, Jiahang, Zhang, Song-Hai, Zhang, Jianfeng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914422454222848
author Luo, Guan
Li, Xiu
Chen, Rui
Yi, Xuanyu
Lin, Jing
Chen, Chia-Hao
Liu, Jiahang
Zhang, Song-Hai
Zhang, Jianfeng
author_facet Luo, Guan
Li, Xiu
Chen, Rui
Yi, Xuanyu
Lin, Jing
Chen, Chia-Hao
Liu, Jiahang
Zhang, Song-Hai
Zhang, Jianfeng
contents The dominant paradigm for high-fidelity 3D generation relies on a VAE-Diffusion pipeline, where the VAE's reconstruction capability sets a firm upper bound on generation quality. A fundamental challenge limiting existing VAEs is the representation mismatch between ground-truth meshes and network predictions: GT meshes have arbitrary, variable topology, while VAEs typically predict fixed-structure implicit fields (\eg, SDF on regular grids). This inherent misalignment prevents establishing explicit mesh-level correspondences, forcing prior work to rely on indirect supervision signals such as SDF or rendering losses. Consequently, fine geometric details, particularly sharp features, are poorly preserved during reconstruction. To address this, we introduce TopoMesh, a sparse voxel-based VAE that unifies both GT and predicted meshes under a shared Dual Marching Cubes (DMC) topological framework. Specifically, we convert arbitrary input meshes into DMC-compliant representations via a remeshing algorithm that preserves sharp edges using an L$\infty$ distance metric. Our decoder outputs meshes in the same DMC format, ensuring that both predicted and target meshes share identical topological structures. This establishes explicit correspondences at the vertex and face level, allowing us to derive explicit mesh-level supervision signals for topology, vertex positions, and face orientations with clear gradients. Our sparse VAE architecture employs this unified framework and is trained with Teacher Forcing and progressive resolution training for stable and efficient convergence. Extensive experiments demonstrate that TopoMesh significantly outperforms existing VAEs in reconstruction fidelity, achieving superior preservation of sharp features and geometric details.
format Preprint
id arxiv_https___arxiv_org_abs_2603_24278
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TopoMesh: High-Fidelity Mesh Autoencoding via Topological Unification
Luo, Guan
Li, Xiu
Chen, Rui
Yi, Xuanyu
Lin, Jing
Chen, Chia-Hao
Liu, Jiahang
Zhang, Song-Hai
Zhang, Jianfeng
Computer Vision and Pattern Recognition
The dominant paradigm for high-fidelity 3D generation relies on a VAE-Diffusion pipeline, where the VAE's reconstruction capability sets a firm upper bound on generation quality. A fundamental challenge limiting existing VAEs is the representation mismatch between ground-truth meshes and network predictions: GT meshes have arbitrary, variable topology, while VAEs typically predict fixed-structure implicit fields (\eg, SDF on regular grids). This inherent misalignment prevents establishing explicit mesh-level correspondences, forcing prior work to rely on indirect supervision signals such as SDF or rendering losses. Consequently, fine geometric details, particularly sharp features, are poorly preserved during reconstruction. To address this, we introduce TopoMesh, a sparse voxel-based VAE that unifies both GT and predicted meshes under a shared Dual Marching Cubes (DMC) topological framework. Specifically, we convert arbitrary input meshes into DMC-compliant representations via a remeshing algorithm that preserves sharp edges using an L$\infty$ distance metric. Our decoder outputs meshes in the same DMC format, ensuring that both predicted and target meshes share identical topological structures. This establishes explicit correspondences at the vertex and face level, allowing us to derive explicit mesh-level supervision signals for topology, vertex positions, and face orientations with clear gradients. Our sparse VAE architecture employs this unified framework and is trained with Teacher Forcing and progressive resolution training for stable and efficient convergence. Extensive experiments demonstrate that TopoMesh significantly outperforms existing VAEs in reconstruction fidelity, achieving superior preservation of sharp features and geometric details.
title TopoMesh: High-Fidelity Mesh Autoencoding via Topological Unification
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.24278