Saved in:
Bibliographic Details
Main Author: Xu, Philip
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2512.22294
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908797453205504
author Xu, Philip
author_facet Xu, Philip
contents We introduce Uni4D, a unified framework for large scale open vocabulary 3D retrieval and controlled 4D generation based on structured three level alignment across text, 3D models, and image modalities. Built upon the Align3D 130 dataset, Uni4D employs a 3D text multi head attention and search model to optimize text to 3D retrieval through improved semantic alignment. The framework further strengthens cross modal alignment through three components: precise text to 3D retrieval, multi view 3D to image alignment, and image to text alignment for generating temporally consistent 4D assets. Experimental results demonstrate that Uni4D achieves high quality 3D retrieval and controllable 4D generation, advancing dynamic multimodal understanding and practical applications.
format Preprint
id arxiv_https___arxiv_org_abs_2512_22294
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Three-Level Alignment Framework for Large-Scale 3D Retrieval and Controlled 4D Generation
Xu, Philip
Computer Vision and Pattern Recognition
We introduce Uni4D, a unified framework for large scale open vocabulary 3D retrieval and controlled 4D generation based on structured three level alignment across text, 3D models, and image modalities. Built upon the Align3D 130 dataset, Uni4D employs a 3D text multi head attention and search model to optimize text to 3D retrieval through improved semantic alignment. The framework further strengthens cross modal alignment through three components: precise text to 3D retrieval, multi view 3D to image alignment, and image to text alignment for generating temporally consistent 4D assets. Experimental results demonstrate that Uni4D achieves high quality 3D retrieval and controllable 4D generation, advancing dynamic multimodal understanding and practical applications.
title A Three-Level Alignment Framework for Large-Scale 3D Retrieval and Controlled 4D Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.22294