Uni4D-LLM: A Unified SpatioTemporal-Aware VLM for 4D Understanding and Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Hanyu, Lee, Gim Hee |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding
by: Zhou, Hanyu, et al.
Published: (2025)
by: Zhou, Hanyu, et al.
Published: (2025)
VLA-4D: Embedding 4D Awareness into Vision-Language-Action Models for SpatioTemporally Coherent Robotic Manipulation
by: Zhou, Hanyu, et al.
Published: (2025)
by: Zhou, Hanyu, et al.
Published: (2025)
4D-VGGT: A General Foundation Model with SpatioTemporal Awareness for Dynamic Scene Geometry Estimation
by: Wang, Haonan, et al.
Published: (2025)
by: Wang, Haonan, et al.
Published: (2025)
STD-GS: Exploring Frame-Event Interaction for SpatioTemporal-Disentangled Gaussian Splatting to Reconstruct High-Dynamic Scene
by: Zhou, Hanyu, et al.
Published: (2025)
by: Zhou, Hanyu, et al.
Published: (2025)
Motion4D: Learning 3D-Consistent Motion and Semantics for 4D Scene Understanding
by: Zhou, Haoran, et al.
Published: (2025)
by: Zhou, Haoran, et al.
Published: (2025)
LLaFEA: Frame-Event Complementary Fusion for Fine-Grained Spatiotemporal Understanding in LMMs
by: Zhou, Hanyu, et al.
Published: (2025)
by: Zhou, Hanyu, et al.
Published: (2025)
ST-$π$: Structured SpatioTemporal VLA for Robotic Manipulation
by: Ma, Chuanhao, et al.
Published: (2026)
by: Ma, Chuanhao, et al.
Published: (2026)
MotionScale: Reconstructing Appearance, Geometry, and Motion of Dynamic Scenes with Scalable 4D Gaussian Splatting
by: Zhou, Haoran, et al.
Published: (2026)
by: Zhou, Haoran, et al.
Published: (2026)
Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
Flow4DGS-SLAM: Optical Flow-Guided 4D Gaussian Splatting SLAM
by: Wang, Yunsong, et al.
Published: (2026)
by: Wang, Yunsong, et al.
Published: (2026)
B4DL: A Benchmark for 4D LiDAR LLM in Spatio-Temporal Understanding
by: Choi, Changho, et al.
Published: (2025)
by: Choi, Changho, et al.
Published: (2025)
3DGS-Calib: 3D Gaussian Splatting for Multimodal SpatioTemporal Calibration
by: Herau, Quentin, et al.
Published: (2024)
by: Herau, Quentin, et al.
Published: (2024)
SpatioTemporal Difference Network for Video Depth Super-Resolution
by: Wang, Zhengxue, et al.
Published: (2025)
by: Wang, Zhengxue, et al.
Published: (2025)
VIRST: Video-Instructed Reasoning Assistant for SpatioTemporal Segmentation
by: Hong, Jihwan, et al.
Published: (2026)
by: Hong, Jihwan, et al.
Published: (2026)
TriRelVLA: Triadic Relational Structure for Generalizable Embodied Manipulation
by: Zhou, Hanyu, et al.
Published: (2026)
by: Zhou, Hanyu, et al.
Published: (2026)
Unified Geometry and Color Compression Framework for Point Clouds via Generative Diffusion Priors
by: Huang, Tianxin, et al.
Published: (2025)
by: Huang, Tianxin, et al.
Published: (2025)
Enhancing Generalizability of Representation Learning for Data-Efficient 3D Scene Understanding
by: Wang, Yunsong, et al.
Published: (2024)
by: Wang, Yunsong, et al.
Published: (2024)
SpatioTemporal Learning for Human Pose Estimation in Sparsely-Labeled Videos
by: Jiao, Yingying, et al.
Published: (2025)
by: Jiao, Yingying, et al.
Published: (2025)
Remote Sensing SpatioTemporal Vision-Language Models: A Comprehensive Survey
by: Liu, Chenyang, et al.
Published: (2024)
by: Liu, Chenyang, et al.
Published: (2024)
UniTS: Unified Spatio-Temporal Generative Model for Remote Sensing
by: Zhang, Yuxiang, et al.
Published: (2025)
by: Zhang, Yuxiang, et al.
Published: (2025)
IAAO: Interactive Affordance Learning for Articulated Objects in 3D Environments
by: Zhang, Can, et al.
Published: (2025)
by: Zhang, Can, et al.
Published: (2025)
4D3R: Motion-Aware Neural Reconstruction and Rendering of Dynamic Scenes from Monocular Videos
by: Guo, Mengqi, et al.
Published: (2025)
by: Guo, Mengqi, et al.
Published: (2025)
Does SpatioTemporal information benefit Two video summarization benchmarks?
by: Ganesh, Aashutosh, et al.
Published: (2024)
by: Ganesh, Aashutosh, et al.
Published: (2024)
econSG: Efficient and Multi-view Consistent Open-Vocabulary 3D Semantic Gaussians
by: Zhang, Can, et al.
Published: (2025)
by: Zhang, Can, et al.
Published: (2025)
DOGS: Distributed-Oriented Gaussian Splatting for Large-Scale 3D Reconstruction Via Gaussian Consensus
by: Chen, Yu, et al.
Published: (2024)
by: Chen, Yu, et al.
Published: (2024)
EmbodiedSplat: Online Feed-Forward Semantic 3DGS for Open-Vocabulary 3D Scene Understanding
by: Lee, Seungjun, et al.
Published: (2026)
by: Lee, Seungjun, et al.
Published: (2026)
STR-Match: Matching SpatioTemporal Relevance Score for Training-Free Video Editing
by: Lee, Junsung, et al.
Published: (2025)
by: Lee, Junsung, et al.
Published: (2025)
SCAPO: Self-Supervised Category-Level Articulated Pose Estimation from a Single 3D Observation
by: Zhang, Can, et al.
Published: (2026)
by: Zhang, Can, et al.
Published: (2026)
X-Ray: A Sequential 3D Representation For Generation
by: Hu, Tao, et al.
Published: (2024)
by: Hu, Tao, et al.
Published: (2024)
Towards Gradient-based Time-Series Explanations through a SpatioTemporal Attention Network
by: Lee, Min Hun
Published: (2024)
by: Lee, Min Hun
Published: (2024)
HandMCM: Multi-modal Point Cloud-based Correspondence State Space Model for 3D Hand Pose Estimation
by: Cheng, Wencan, et al.
Published: (2026)
by: Cheng, Wencan, et al.
Published: (2026)
UniMesh: Unifying 3D Mesh Understanding and Generation
by: Huang, Peng, et al.
Published: (2026)
by: Huang, Peng, et al.
Published: (2026)
ST-Gen4D: Embedding 4D Spatiotemporal Cognition into World Model for 4D Generation
by: Wang, Haonan, et al.
Published: (2026)
by: Wang, Haonan, et al.
Published: (2026)
4D-GSW: Kinematic-Aware Spatio-Temporal Consistent Watermarking for 4D Gaussian Splatting
by: Zhou, Sifan, et al.
Published: (2026)
by: Zhou, Sifan, et al.
Published: (2026)
ST-Booster: An Iterative SpatioTemporal Perception Booster for Vision-and-Language Navigation in Continuous Environments
by: Yue, Lu, et al.
Published: (2025)
by: Yue, Lu, et al.
Published: (2025)
ChatSplat: 3D Conversational Gaussian Splatting
by: Chen, Hanlin, et al.
Published: (2024)
by: Chen, Hanlin, et al.
Published: (2024)
g3D-LF: Generalizable 3D-Language Feature Fields for Embodied Tasks
by: Wang, Zihan, et al.
Published: (2024)
by: Wang, Zihan, et al.
Published: (2024)
Animate124: Animating One Image to 4D Dynamic Scene
by: Zhao, Yuyang, et al.
Published: (2023)
by: Zhao, Yuyang, et al.
Published: (2023)
Fast 4D Mesh Generation by Spatio-Temporal Attention Chains
by: Samuel, Dvir, et al.
Published: (2026)
by: Samuel, Dvir, et al.
Published: (2026)
Syn-to-Real Unsupervised Domain Adaptation for Indoor 3D Object Detection
by: Wang, Yunsong, et al.
Published: (2024)
by: Wang, Yunsong, et al.
Published: (2024)
Similar Items
-
LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding
by: Zhou, Hanyu, et al.
Published: (2025) -
VLA-4D: Embedding 4D Awareness into Vision-Language-Action Models for SpatioTemporally Coherent Robotic Manipulation
by: Zhou, Hanyu, et al.
Published: (2025) -
4D-VGGT: A General Foundation Model with SpatioTemporal Awareness for Dynamic Scene Geometry Estimation
by: Wang, Haonan, et al.
Published: (2025) -
STD-GS: Exploring Frame-Event Interaction for SpatioTemporal-Disentangled Gaussian Splatting to Reconstruct High-Dynamic Scene
by: Zhou, Hanyu, et al.
Published: (2025) -
Motion4D: Learning 3D-Consistent Motion and Semantics for 4D Scene Understanding
by: Zhou, Haoran, et al.
Published: (2025)