Helix4D: Complex 4D Mesh Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yenphraphai, Jiraphon, Chen, Jianqi, Wang, Jian, Qian, Gordon, Tulyakov, Sergey, Abdal, Rameen, Yeh, Raymond A., Wonka, Peter, Wang, Chaoyang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914600263352320
author Yenphraphai, Jiraphon
Chen, Jianqi
Wang, Jian
Qian, Gordon
Tulyakov, Sergey
Abdal, Rameen
Yeh, Raymond A.
Wonka, Peter
Wang, Chaoyang
author_facet Yenphraphai, Jiraphon
Chen, Jianqi
Wang, Jian
Qian, Gordon
Tulyakov, Sergey
Abdal, Rameen
Yeh, Raymond A.
Wonka, Peter
Wang, Chaoyang
contents Current video-to-4D methods struggle with complex topology changes, transparent materials, thin structures, and inner surfaces. We present Helix4D, a dynamic mesh generation framework by inheriting the expressive representation of Trellis2, adapting it from image-to-3D to video-conditioned 4D generation. Our design arises from two key questions: (a) how to enable Trellis2's frame-local attention to share information across frames while preserving its pretrained quality on rare cases such as transparent objects and inner surfaces, and (b) how to inject temporal information into a purely 3D positional encoding without breaking pretrained capabilities. We address (a) with a sliding-window cross-frame attention and anchor on the first frame. The first frame is generated by the base Trellis2 model and injected into our model, letting it inherit Trellis2's quality in rare cases through cross-frame attention. We address (b) with a 4D temporal encoding that repurposes redundant low-frequency spatial RoPE bands for time, extending the encoding from 3D with no additional parameters. Extensive experiments show the effectiveness of Helix4D for high-quality dynamic mesh generation on ActionBench and our own challenging complex dynamics set.
format Preprint
id arxiv_https___arxiv_org_abs_2605_26109
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Helix4D: Complex 4D Mesh Generation
Yenphraphai, Jiraphon
Chen, Jianqi
Wang, Jian
Qian, Gordon
Tulyakov, Sergey
Abdal, Rameen
Yeh, Raymond A.
Wonka, Peter
Wang, Chaoyang
Computer Vision and Pattern Recognition
Current video-to-4D methods struggle with complex topology changes, transparent materials, thin structures, and inner surfaces. We present Helix4D, a dynamic mesh generation framework by inheriting the expressive representation of Trellis2, adapting it from image-to-3D to video-conditioned 4D generation. Our design arises from two key questions: (a) how to enable Trellis2's frame-local attention to share information across frames while preserving its pretrained quality on rare cases such as transparent objects and inner surfaces, and (b) how to inject temporal information into a purely 3D positional encoding without breaking pretrained capabilities. We address (a) with a sliding-window cross-frame attention and anchor on the first frame. The first frame is generated by the base Trellis2 model and injected into our model, letting it inherit Trellis2's quality in rare cases through cross-frame attention. We address (b) with a 4D temporal encoding that repurposes redundant low-frequency spatial RoPE bands for time, extending the encoding from 3D with no additional parameters. Extensive experiments show the effectiveness of Helix4D for high-quality dynamic mesh generation on ActionBench and our own challenging complex dynamics set.
title Helix4D: Complex 4D Mesh Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.26109