FACE: A Face-based Autoregressive Representation for High-Fidelity and Efficient Mesh Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Hanxiao, Guo, Yuan-Chen, Liu, Ying-Tian, Zou, Zi-Xin, Zhang, Biao, Quan, Weize, Liang, Ding, Cao, Yan-Pei, Yan, Dong-Ming
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914364334800896
author Wang, Hanxiao
Guo, Yuan-Chen
Liu, Ying-Tian
Zou, Zi-Xin
Zhang, Biao
Quan, Weize
Liang, Ding
Cao, Yan-Pei
Yan, Dong-Ming
author_facet Wang, Hanxiao
Guo, Yuan-Chen
Liu, Ying-Tian
Zou, Zi-Xin
Zhang, Biao
Quan, Weize
Liang, Ding
Cao, Yan-Pei
Yan, Dong-Ming
contents Autoregressive models for 3D mesh generation suffer from a fundamental limitation: they flatten meshes into long vertex-coordinate sequences. This results in prohibitive computational costs, hindering the efficient synthesis of high-fidelity geometry. We argue this bottleneck stems from operating at the wrong semantic level. We introduce FACE, a novel Autoregressive Autoencoder (ARAE) framework that reconceptualizes the task by generating meshes at the face level. Our one-face-one-token strategy treats each triangle face, the fundamental building block of a mesh, as a single, unified token. This simple yet powerful design reduces the sequence length by a factor of nine, leading to an unprecedented compression ratio of 0.11, halving the previous state-of-the-art. This dramatic efficiency gain does not compromise quality; by pairing our face-level decoder with a powerful VecSet encoder, FACE achieves state-of-the-art reconstruction quality on standard benchmarks. The versatility of the learned latent space is further demonstrated by training a latent diffusion model that achieves high-fidelity, single-image-to-mesh generation. FACE provides a simple, scalable, and powerful paradigm that lowers the barrier to high-quality structured 3D content creation.
format Preprint
id arxiv_https___arxiv_org_abs_2603_01515
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FACE: A Face-based Autoregressive Representation for High-Fidelity and Efficient Mesh Generation
Wang, Hanxiao
Guo, Yuan-Chen
Liu, Ying-Tian
Zou, Zi-Xin
Zhang, Biao
Quan, Weize
Liang, Ding
Cao, Yan-Pei
Yan, Dong-Ming
Computer Vision and Pattern Recognition
Autoregressive models for 3D mesh generation suffer from a fundamental limitation: they flatten meshes into long vertex-coordinate sequences. This results in prohibitive computational costs, hindering the efficient synthesis of high-fidelity geometry. We argue this bottleneck stems from operating at the wrong semantic level. We introduce FACE, a novel Autoregressive Autoencoder (ARAE) framework that reconceptualizes the task by generating meshes at the face level. Our one-face-one-token strategy treats each triangle face, the fundamental building block of a mesh, as a single, unified token. This simple yet powerful design reduces the sequence length by a factor of nine, leading to an unprecedented compression ratio of 0.11, halving the previous state-of-the-art. This dramatic efficiency gain does not compromise quality; by pairing our face-level decoder with a powerful VecSet encoder, FACE achieves state-of-the-art reconstruction quality on standard benchmarks. The versatility of the learned latent space is further demonstrated by training a latent diffusion model that achieves high-fidelity, single-image-to-mesh generation. FACE provides a simple, scalable, and powerful paradigm that lowers the barrier to high-quality structured 3D content creation.
title FACE: A Face-based Autoregressive Representation for High-Fidelity and Efficient Mesh Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.01515