LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Hanyu, Lee, Gim Hee |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Uni4D-LLM: A Unified SpatioTemporal-Aware VLM for 4D Understanding and Generation
by: Zhou, Hanyu, et al.
Published: (2025)
by: Zhou, Hanyu, et al.
Published: (2025)
VLA-4D: Embedding 4D Awareness into Vision-Language-Action Models for SpatioTemporally Coherent Robotic Manipulation
by: Zhou, Hanyu, et al.
Published: (2025)
by: Zhou, Hanyu, et al.
Published: (2025)
LLaFEA: Frame-Event Complementary Fusion for Fine-Grained Spatiotemporal Understanding in LMMs
by: Zhou, Hanyu, et al.
Published: (2025)
by: Zhou, Hanyu, et al.
Published: (2025)
STD-GS: Exploring Frame-Event Interaction for SpatioTemporal-Disentangled Gaussian Splatting to Reconstruct High-Dynamic Scene
by: Zhou, Hanyu, et al.
Published: (2025)
by: Zhou, Hanyu, et al.
Published: (2025)
4D-VGGT: A General Foundation Model with SpatioTemporal Awareness for Dynamic Scene Geometry Estimation
by: Wang, Haonan, et al.
Published: (2025)
by: Wang, Haonan, et al.
Published: (2025)
Motion4D: Learning 3D-Consistent Motion and Semantics for 4D Scene Understanding
by: Zhou, Haoran, et al.
Published: (2025)
by: Zhou, Haoran, et al.
Published: (2025)
3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer
by: Deng, Jiajun, et al.
Published: (2025)
by: Deng, Jiajun, et al.
Published: (2025)
LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness
by: Zhu, Chenming, et al.
Published: (2024)
by: Zhu, Chenming, et al.
Published: (2024)
LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
by: Shu, Fangxun, et al.
Published: (2024)
by: Shu, Fangxun, et al.
Published: (2024)
TG-LLaVA: Text Guided LLaVA via Learnable Latent Embeddings
by: Yan, Dawei, et al.
Published: (2024)
by: Yan, Dawei, et al.
Published: (2024)
MotionScale: Reconstructing Appearance, Geometry, and Motion of Dynamic Scenes with Scalable 4D Gaussian Splatting
by: Zhou, Haoran, et al.
Published: (2026)
by: Zhou, Haoran, et al.
Published: (2026)
LLaVA$^3$: Representing 3D Scenes like a Cubist Painter to Boost 3D Scene Understanding of VLMs
by: Petit, Doriand, et al.
Published: (2025)
by: Petit, Doriand, et al.
Published: (2025)
ST-$π$: Structured SpatioTemporal VLA for Robotic Manipulation
by: Ma, Chuanhao, et al.
Published: (2026)
by: Ma, Chuanhao, et al.
Published: (2026)
Enhance Image-to-Image Generation with LLaVA-generated Prompts
by: Ding, Zhicheng, et al.
Published: (2024)
by: Ding, Zhicheng, et al.
Published: (2024)
Enhancing Generalizability of Representation Learning for Data-Efficient 3D Scene Understanding
by: Wang, Yunsong, et al.
Published: (2024)
by: Wang, Yunsong, et al.
Published: (2024)
Flow4DGS-SLAM: Optical Flow-Guided 4D Gaussian Splatting SLAM
by: Wang, Yunsong, et al.
Published: (2026)
by: Wang, Yunsong, et al.
Published: (2026)
TC-LLaVA: Rethinking the Transfer from Image to Video Understanding with Temporal Considerations
by: Gao, Mingze, et al.
Published: (2024)
by: Gao, Mingze, et al.
Published: (2024)
LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding
by: Li, Hongyu, et al.
Published: (2025)
by: Li, Hongyu, et al.
Published: (2025)
NOVO: Bridging LLaVA and SAM with Visual-only Prompts for Reasoning Segmentation
by: Yoon, Kyung-Yoon, et al.
Published: (2025)
by: Yoon, Kyung-Yoon, et al.
Published: (2025)
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler
by: Zhang, Xingjian, et al.
Published: (2025)
by: Zhang, Xingjian, et al.
Published: (2025)
Cosmos-LLaVA: Chatting with the Visual Cosmos-LLaVA: Görselle Sohbet Etmek
by: Zeer, Ahmed, et al.
Published: (2024)
by: Zeer, Ahmed, et al.
Published: (2024)
EmbodiedSplat: Online Feed-Forward Semantic 3DGS for Open-Vocabulary 3D Scene Understanding
by: Lee, Seungjun, et al.
Published: (2026)
by: Lee, Seungjun, et al.
Published: (2026)
LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs?
by: Fang, Kechen, et al.
Published: (2026)
by: Fang, Kechen, et al.
Published: (2026)
ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts
by: Cai, Mu, et al.
Published: (2023)
by: Cai, Mu, et al.
Published: (2023)
Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
by: Yuan, Haobo, et al.
Published: (2025)
by: Yuan, Haobo, et al.
Published: (2025)
Animate124: Animating One Image to 4D Dynamic Scene
by: Zhao, Yuyang, et al.
Published: (2023)
by: Zhao, Yuyang, et al.
Published: (2023)
LLaVA-Critic: Learning to Evaluate Multimodal Models
by: Xiong, Tianyi, et al.
Published: (2024)
by: Xiong, Tianyi, et al.
Published: (2024)
LLaVA-c: Continual Improved Visual Instruction Tuning
by: Liu, Wenzhuo, et al.
Published: (2025)
by: Liu, Wenzhuo, et al.
Published: (2025)
Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA
by: Kanjula, Karthik Reddy, et al.
Published: (2025)
by: Kanjula, Karthik Reddy, et al.
Published: (2025)
OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
by: Zhang, Tao, et al.
Published: (2024)
by: Zhang, Tao, et al.
Published: (2024)
3DGS-Calib: 3D Gaussian Splatting for Multimodal SpatioTemporal Calibration
by: Herau, Quentin, et al.
Published: (2024)
by: Herau, Quentin, et al.
Published: (2024)
SpatioTemporal Difference Network for Video Depth Super-Resolution
by: Wang, Zhengxue, et al.
Published: (2025)
by: Wang, Zhengxue, et al.
Published: (2025)
VIRST: Video-Instructed Reasoning Assistant for SpatioTemporal Segmentation
by: Hong, Jihwan, et al.
Published: (2026)
by: Hong, Jihwan, et al.
Published: (2026)
LLaVA-Video: Video Instruction Tuning With Synthetic Data
by: Zhang, Yuanhan, et al.
Published: (2024)
by: Zhang, Yuanhan, et al.
Published: (2024)
TriRelVLA: Triadic Relational Structure for Generalizable Embodied Manipulation
by: Zhou, Hanyu, et al.
Published: (2026)
by: Zhou, Hanyu, et al.
Published: (2026)
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
by: Xu, Lin, et al.
Published: (2024)
by: Xu, Lin, et al.
Published: (2024)
LLaVA-SLT: Visual Language Tuning for Sign Language Translation
by: Liang, Han, et al.
Published: (2024)
by: Liang, Han, et al.
Published: (2024)
LLaVA-LE: Large Language-and-Vision Assistant for Lunar Exploration
by: Inal, Gokce, et al.
Published: (2026)
by: Inal, Gokce, et al.
Published: (2026)
MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning
by: Zhao, Xiangyu, et al.
Published: (2024)
by: Zhao, Xiangyu, et al.
Published: (2024)
LLaVA-SG: Leveraging Scene Graphs as Visual Semantic Expression in Vision-Language Models
by: Wang, Jingyi, et al.
Published: (2024)
by: Wang, Jingyi, et al.
Published: (2024)
Similar Items
-
Uni4D-LLM: A Unified SpatioTemporal-Aware VLM for 4D Understanding and Generation
by: Zhou, Hanyu, et al.
Published: (2025) -
VLA-4D: Embedding 4D Awareness into Vision-Language-Action Models for SpatioTemporally Coherent Robotic Manipulation
by: Zhou, Hanyu, et al.
Published: (2025) -
LLaFEA: Frame-Event Complementary Fusion for Fine-Grained Spatiotemporal Understanding in LMMs
by: Zhou, Hanyu, et al.
Published: (2025) -
STD-GS: Exploring Frame-Event Interaction for SpatioTemporal-Disentangled Gaussian Splatting to Reconstruct High-Dynamic Scene
by: Zhou, Hanyu, et al.
Published: (2025) -
4D-VGGT: A General Foundation Model with SpatioTemporal Awareness for Dynamic Scene Geometry Estimation
by: Wang, Haonan, et al.
Published: (2025)