IGGT: Instance-Grounded Geometry Transformer for Semantic 3D Reconstruction
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Hao, Zou, Zhengyu, Liu, Fangfu, Zhang, Xuanyang, Hong, Fangzhou, Cao, Yukang, Lan, Yushi, Zhang, Manyuan, Yu, Gang, Zhang, Dingwen, Liu, Ziwei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
IRIS-SLAM: Unified Geo-Instance Representations for Robust Semantic Localization and Mapping
by: Xiao, Tingyang, et al.
Published: (2026)
by: Xiao, Tingyang, et al.
Published: (2026)
OmniVGGT: Omni-Modality Driven Visual Geometry Grounded Transformer
by: Peng, Haosong, et al.
Published: (2025)
by: Peng, Haosong, et al.
Published: (2025)
HSImul3R: Physics-in-the-Loop Reconstruction of Simulation-Ready Human-Scene Interactions
by: Cao, Yukang, et al.
Published: (2026)
by: Cao, Yukang, et al.
Published: (2026)
DiffTF++: 3D-aware Diffusion Transformer for Large-Vocabulary 3D Generation
by: Cao, Ziang, et al.
Published: (2024)
by: Cao, Ziang, et al.
Published: (2024)
Reconstructing 4D Spatial Intelligence: A Survey
by: Cao, Yukang, et al.
Published: (2025)
by: Cao, Yukang, et al.
Published: (2025)
Holi-Spatial: Evolving Video Streams into Holistic 3D Spatial Intelligence
by: Gao, Yuanyuan, et al.
Published: (2026)
by: Gao, Yuanyuan, et al.
Published: (2026)
StructLDM: Structured Latent Diffusion for 3D Human Generation
by: Hu, Tao, et al.
Published: (2024)
by: Hu, Tao, et al.
Published: (2024)
SurMo: Surface-based 4D Motion Modeling for Dynamic Human Rendering
by: Hu, Tao, et al.
Published: (2024)
by: Hu, Tao, et al.
Published: (2024)
STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer
by: Lan, Yushi, et al.
Published: (2025)
by: Lan, Yushi, et al.
Published: (2025)
GS-VTON: Controllable 3D Virtual Try-on with Gaussian Splatting
by: Cao, Yukang, et al.
Published: (2024)
by: Cao, Yukang, et al.
Published: (2024)
CityGS-X: A Scalable Architecture for Efficient and Geometrically Accurate Large-Scale Scene Reconstruction
by: Gao, Yuanyuan, et al.
Published: (2025)
by: Gao, Yuanyuan, et al.
Published: (2025)
MonoArt: Progressive Structural Reasoning for Monocular Articulated 3D Reconstruction
by: Li, Haitian, et al.
Published: (2026)
by: Li, Haitian, et al.
Published: (2026)
PhysX-Anything: Simulation-Ready Physical 3D Assets from Single Image
by: Cao, Ziang, et al.
Published: (2025)
by: Cao, Ziang, et al.
Published: (2025)
Compositional Generative Model of Unbounded 4D Cities
by: Xie, Haozhe, et al.
Published: (2025)
by: Xie, Haozhe, et al.
Published: (2025)
Generative Gaussian Splatting for Unbounded 3D City Generation
by: Xie, Haozhe, et al.
Published: (2024)
by: Xie, Haozhe, et al.
Published: (2024)
CityDreamer: Compositional Generative Model of Unbounded 3D Cities
by: Xie, Haozhe, et al.
Published: (2023)
by: Xie, Haozhe, et al.
Published: (2023)
GeneMAN: Generalizable Single-Image 3D Human Reconstruction from Multi-Source Human Data
by: Wang, Wentao, et al.
Published: (2024)
by: Wang, Wentao, et al.
Published: (2024)
FashionEngine: Interactive 3D Human Generation and Editing via Multimodal Controls
by: Hu, Tao, et al.
Published: (2024)
by: Hu, Tao, et al.
Published: (2024)
ScenePainter: Semantically Consistent Perpetual 3D Scene Generation with Concept Relation Alignment
by: Xia, Chong, et al.
Published: (2025)
by: Xia, Chong, et al.
Published: (2025)
InstructLayout: Instruction-Driven 2D and 3D Layout Synthesis with Semantic Graph Prior
by: Lin, Chenguo, et al.
Published: (2024)
by: Lin, Chenguo, et al.
Published: (2024)
ResVG: Enhancing Relation and Semantic Understanding in Multiple Instances for Visual Grounding
by: Zheng, Minghang, et al.
Published: (2024)
by: Zheng, Minghang, et al.
Published: (2024)
Calibrating Undisciplined Over-Smoothing in Transformer for Weakly Supervised Semantic Segmentation
by: Cheng, Lechao, et al.
Published: (2023)
by: Cheng, Lechao, et al.
Published: (2023)
Semantic One-Dimensional Tokenizer for Image Reconstruction and Generation
by: Qu, Yunpeng, et al.
Published: (2026)
by: Qu, Yunpeng, et al.
Published: (2026)
FreeMorph: Tuning-Free Generalized Image Morphing with Diffusion Model
by: Cao, Yukang, et al.
Published: (2025)
by: Cao, Yukang, et al.
Published: (2025)
Unsupervised Pre-training with Language-Vision Prompts for Low-Data Instance Segmentation
by: Zhang, Dingwen, et al.
Published: (2024)
by: Zhang, Dingwen, et al.
Published: (2024)
Adaptive Semantic Communication for UAV/UGV Cooperative Path Planning
by: Zhao, Fangzhou, et al.
Published: (2025)
by: Zhao, Fangzhou, et al.
Published: (2025)
3D Scene Generation: A Survey
by: Wen, Beichen, et al.
Published: (2025)
by: Wen, Beichen, et al.
Published: (2025)
CrowdMoGen: Zero-Shot Text-Driven Collective Motion Generation
by: Cao, Yukang, et al.
Published: (2024)
by: Cao, Yukang, et al.
Published: (2024)
A Comprehensive Framework for Semantic Similarity Analysis of Human and AI-Generated Text Using Transformer Architectures and Ensemble Techniques
by: Gao, Lifu, et al.
Published: (2025)
by: Gao, Lifu, et al.
Published: (2025)
3DTopia-XL: Scaling High-quality 3D Asset Generation via Primitive Diffusion
by: Chen, Zhaoxi, et al.
Published: (2024)
by: Chen, Zhaoxi, et al.
Published: (2024)
PhysX-3D: Physical-Grounded 3D Asset Generation
by: Cao, Ziang, et al.
Published: (2025)
by: Cao, Ziang, et al.
Published: (2025)
GSAlign: Geometric and Semantic Alignment Network for Aerial-Ground Person Re-Identification
by: Li, Qiao, et al.
Published: (2025)
by: Li, Qiao, et al.
Published: (2025)
AvatarGO: Zero-shot 4D Human-Object Interaction Generation and Animation
by: Cao, Yukang, et al.
Published: (2024)
by: Cao, Yukang, et al.
Published: (2024)
GGPT: Geometry Grounded Point Transformer
by: Chen, Yutong, et al.
Published: (2026)
by: Chen, Yutong, et al.
Published: (2026)
InfiniteVGGT: Visual Geometry Grounded Transformer for Endless Streams
by: Yuan, Shuai, et al.
Published: (2026)
by: Yuan, Shuai, et al.
Published: (2026)
VDNeRF: Vision-only Dynamic Neural Radiance Field for Urban Scenes
by: Zou, Zhengyu, et al.
Published: (2025)
by: Zou, Zhengyu, et al.
Published: (2025)
OnlineX: Unified Online 3D Reconstruction and Understanding with Active-to-Stable State Evolution
by: Xia, Chong, et al.
Published: (2026)
by: Xia, Chong, et al.
Published: (2026)
SimRecon: SimReady Compositional Scene Reconstruction from Real Videos
by: Xia, Chong, et al.
Published: (2026)
by: Xia, Chong, et al.
Published: (2026)
Wukong's 72 Transformations: High-fidelity Textured 3D Morphing via Flow Models
by: Yin, Minghao, et al.
Published: (2025)
by: Yin, Minghao, et al.
Published: (2025)
PhysX-Omni: Unified Simulation-Ready Physical 3D Generation for Rigid, Deformable, and Articulated Objects
by: Cao, Ziang, et al.
Published: (2026)
by: Cao, Ziang, et al.
Published: (2026)
Similar Items
-
IRIS-SLAM: Unified Geo-Instance Representations for Robust Semantic Localization and Mapping
by: Xiao, Tingyang, et al.
Published: (2026) -
OmniVGGT: Omni-Modality Driven Visual Geometry Grounded Transformer
by: Peng, Haosong, et al.
Published: (2025) -
HSImul3R: Physics-in-the-Loop Reconstruction of Simulation-Ready Human-Scene Interactions
by: Cao, Yukang, et al.
Published: (2026) -
DiffTF++: 3D-aware Diffusion Transformer for Large-Vocabulary 3D Generation
by: Cao, Ziang, et al.
Published: (2024) -
Reconstructing 4D Spatial Intelligence: A Survey
by: Cao, Yukang, et al.
Published: (2025)