PAGE-4D: VGGT-4D Perception via Disentangled Pose and Geometry Estimation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Kaichen, Wang, Yuhan, Chen, Grace, Chang, Xinhai, Beaudouin, Gaspard, Zhan, Fangneng, Liang, Paul Pu, Wang, Mengyu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Stream3D: Sequential Multi-View 3D Generation via Evidential Memory
by: Zhou, Kaichen, et al.
Published: (2026)
by: Zhou, Kaichen, et al.
Published: (2026)
GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation
by: Zhou, Kaichen, et al.
Published: (2026)
by: Zhou, Kaichen, et al.
Published: (2026)
Abstract 3D Perception for Spatial Intelligence in Vision-Language Models
by: Liu, Yifan, et al.
Published: (2025)
by: Liu, Yifan, et al.
Published: (2025)
Neural Surface Reconstruction from Sparse Views Using Epipolar Geometry
by: Chang, Xinhai, et al.
Published: (2024)
by: Chang, Xinhai, et al.
Published: (2024)
4D-VGGT: A General Foundation Model with SpatioTemporal Awareness for Dynamic Scene Geometry Estimation
by: Wang, Haonan, et al.
Published: (2025)
by: Wang, Haonan, et al.
Published: (2025)
Memorization in 3D Shape Generation: An Empirical Study
by: Pu, Shu, et al.
Published: (2025)
by: Pu, Shu, et al.
Published: (2025)
4DLangVGGT: 4D Language-Visual Geometry Grounded Transformer
by: Wu, Xianfeng, et al.
Published: (2025)
by: Wu, Xianfeng, et al.
Published: (2025)
Delta Rectified Flow Sampling for Text-to-Image Editing
by: Beaudouin, Gaspard, et al.
Published: (2025)
by: Beaudouin, Gaspard, et al.
Published: (2025)
VGGT4D: Mining Motion Cues in Visual Geometry Transformers for 4D Scene Reconstruction
by: Hu, Yu, et al.
Published: (2025)
by: Hu, Yu, et al.
Published: (2025)
GeoWorld-VLM: Geometry from World Models for Vision-Language Models
by: Gu, Renjie, et al.
Published: (2026)
by: Gu, Renjie, et al.
Published: (2026)
LiteVGGT: Boosting Vanilla VGGT via Geometry-aware Cached Token Merging
by: Shu, Zhijian, et al.
Published: (2025)
by: Shu, Zhijian, et al.
Published: (2025)
SplitFlow: Flow Decomposition for Inversion-Free Text-to-Image Editing
by: Yoon, Sung-Hoon, et al.
Published: (2025)
by: Yoon, Sung-Hoon, et al.
Published: (2025)
TransPose: 6D Object Pose Estimation with Geometry-Aware Transformer
by: Lin, Xiao, et al.
Published: (2023)
by: Lin, Xiao, et al.
Published: (2023)
VGGT-HPE: Reframing Head Pose Estimation as Relative Pose Prediction
by: Vasileiou, Vasiliki, et al.
Published: (2026)
by: Vasileiou, Vasiliki, et al.
Published: (2026)
VGGT-World: Transforming VGGT into an Autoregressive Geometry World Model
by: Sun, Xiangyu, et al.
Published: (2026)
by: Sun, Xiangyu, et al.
Published: (2026)
VGGT-Occ: Geometry-Grounded and Density-Aware Gated Fusion for 3D Occupancy Prediction
by: Chen, Xun, et al.
Published: (2026)
by: Chen, Xun, et al.
Published: (2026)
AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance
by: Xu, Tianling, et al.
Published: (2025)
by: Xu, Tianling, et al.
Published: (2025)
VGGT: Visual Geometry Grounded Transformer
by: Wang, Jianyuan, et al.
Published: (2025)
by: Wang, Jianyuan, et al.
Published: (2025)
RAD: A Dataset and Benchmark for Real-Life Anomaly Detection with Robotic Observations
by: Zhou, Kaichen, et al.
Published: (2024)
by: Zhou, Kaichen, et al.
Published: (2024)
DiffAge3D: Diffusion-based 3D-aware Face Aging
by: Wahid, Junaid, et al.
Published: (2024)
by: Wahid, Junaid, et al.
Published: (2024)
Lifelong Domain Adaptive 3D Human Pose Estimation
by: Peng, Qucheng, et al.
Published: (2025)
by: Peng, Qucheng, et al.
Published: (2025)
DynamicVGGT: Learning Dynamic Point Maps for 4D Scene Reconstruction in Autonomous Driving
by: He, Zhuolin, et al.
Published: (2026)
by: He, Zhuolin, et al.
Published: (2026)
VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detection
by: Cao, Yang, et al.
Published: (2026)
by: Cao, Yang, et al.
Published: (2026)
DecomPose: Disentangling Cross-Category Optimization Contention for Category-Level 6D Object Pose Estimation
by: Gao, Yifan, et al.
Published: (2026)
by: Gao, Yifan, et al.
Published: (2026)
Geometry-Aware Sparse Depth Sampling for High-Fidelity RGB-D Depth Completion in Robotic Systems
by: Salloom, Tony, et al.
Published: (2025)
by: Salloom, Tony, et al.
Published: (2025)
SOGS: Second-Order Anchor for Advanced 3D Gaussian Splatting
by: Zhang, Jiahui, et al.
Published: (2025)
by: Zhang, Jiahui, et al.
Published: (2025)
PCIE_Pose Solution for EgoExo4D Pose and Proficiency Estimation Challenge
by: Chen, Feng, et al.
Published: (2025)
by: Chen, Feng, et al.
Published: (2025)
FrameVGGT: Geometry-Aligned Frame-Level Memory for Bounded Streaming VGGT
by: Xu, Zhisong, et al.
Published: (2026)
by: Xu, Zhisong, et al.
Published: (2026)
Avatar4D: Synthesizing Domain-Specific 4D Humans for Real-World Pose Estimation
by: Bright, Jerrin, et al.
Published: (2025)
by: Bright, Jerrin, et al.
Published: (2025)
BioPose: Biomechanically-accurate 3D Pose Estimation from Monocular Videos
by: Koleini, Farnoosh, et al.
Published: (2025)
by: Koleini, Farnoosh, et al.
Published: (2025)
Improved Cryo-EM Pose Estimation and 3D Classification through Latent-Space Disentanglement
by: Chen, Weijie, et al.
Published: (2023)
by: Chen, Weijie, et al.
Published: (2023)
VGGT-360: Geometry-Consistent Zero-Shot Panoramic Depth Estimation
by: Yuan, Jiayi, et al.
Published: (2026)
by: Yuan, Jiayi, et al.
Published: (2026)
Advances in Feed-Forward 3D Reconstruction and View Synthesis: A Survey
by: Zhang, Jiahui, et al.
Published: (2025)
by: Zhang, Jiahui, et al.
Published: (2025)
Advances in Feed‐Forward 3D Reconstruction and View Synthesis: A Survey
by: Jiahui Zhang, et al.
Published: (2026)
by: Jiahui Zhang, et al.
Published: (2026)
MoViD: View-Invariant 3D Human Pose Estimation via Motion-View Disentanglement
by: Liu, Yejia, et al.
Published: (2026)
by: Liu, Yejia, et al.
Published: (2026)
MuSASplat: Efficient Sparse-View 3D Gaussian Splats via Lightweight Multi-Scale Adaptation
by: Xu, Muyu, et al.
Published: (2025)
by: Xu, Muyu, et al.
Published: (2025)
S ee 4D: Pose‐Free 4D Generation via Auto‐Regressive Video Inpainting
by: Dongyue Lu, et al.
Published: (2026)
by: Dongyue Lu, et al.
Published: (2026)
See4D: Pose-Free 4D Generation via Auto-Regressive Video Inpainting
by: Lu, Dongyue, et al.
Published: (2025)
by: Lu, Dongyue, et al.
Published: (2025)
Reloc-VGGT: Visual Re-localization with Geometry Grounded Transformer
by: Deng, Tianchen, et al.
Published: (2025)
by: Deng, Tianchen, et al.
Published: (2025)
Base de données de présidents d'universités
by: Beaudouin, Pierre-Yves
Published: (2025)
by: Beaudouin, Pierre-Yves
Published: (2025)
Similar Items
-
Stream3D: Sequential Multi-View 3D Generation via Evidential Memory
by: Zhou, Kaichen, et al.
Published: (2026) -
GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation
by: Zhou, Kaichen, et al.
Published: (2026) -
Abstract 3D Perception for Spatial Intelligence in Vision-Language Models
by: Liu, Yifan, et al.
Published: (2025) -
Neural Surface Reconstruction from Sparse Views Using Epipolar Geometry
by: Chang, Xinhai, et al.
Published: (2024) -
4D-VGGT: A General Foundation Model with SpatioTemporal Awareness for Dynamic Scene Geometry Estimation
by: Wang, Haonan, et al.
Published: (2025)