From Priors to Perception: Grounding Video-LLMs in Physical Reality
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Zicheng, Gan, Chaofan, Li, Shijie, Lin, Weiyao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DAC: 2D-3D Retrieval with Noisy Labels via Divide-and-Conquer Alignment and Correction
by: Gan, Chaofan, et al.
Published: (2024)
by: Gan, Chaofan, et al.
Published: (2024)
CogStream: Context-guided Streaming Video Question Answering
by: Zhao, Zicheng, et al.
Published: (2025)
by: Zhao, Zicheng, et al.
Published: (2025)
Massive Activations are the Key to Local Detail Synthesis in Diffusion Transformers
by: Gan, Chaofan, et al.
Published: (2025)
by: Gan, Chaofan, et al.
Published: (2025)
VidLaDA: Bidirectional Diffusion Large Language Models for Efficient Video Understanding
by: He, Zhihao, et al.
Published: (2026)
by: He, Zhihao, et al.
Published: (2026)
Enhancing Video Large Language Models with Structured Multi-Video Collaborative Reasoning
by: He, Zhihao, et al.
Published: (2025)
by: He, Zhihao, et al.
Published: (2025)
Looking Beyond Visible Cues: Implicit Video Question Answering via Dual-Clue Reasoning
by: Chen, Tieyuan, et al.
Published: (2025)
by: Chen, Tieyuan, et al.
Published: (2025)
MCA: 2D-3D Retrieval with Noisy Labels via Multi-level Adaptive Correction and Alignment
by: Zou, Gui, et al.
Published: (2025)
by: Zou, Gui, et al.
Published: (2025)
MECD+: Unlocking Event-Level Causal Graph Discovery for Video Reasoning
by: Chen, Tieyuan, et al.
Published: (2025)
by: Chen, Tieyuan, et al.
Published: (2025)
Unleashing Diffusion Transformers for Visual Correspondence by Modulating Massive Activations
by: Gan, Chaofan, et al.
Published: (2025)
by: Gan, Chaofan, et al.
Published: (2025)
MECD: Unlocking Multi-Event Causal Discovery in Video Reasoning
by: Chen, Tieyuan, et al.
Published: (2024)
by: Chen, Tieyuan, et al.
Published: (2024)
Region-Adaptive Video Sharpening via Rate-Perception Optimization
by: Pang, Yingxue, et al.
Published: (2025)
by: Pang, Yingxue, et al.
Published: (2025)
Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding
by: Zheng, Zelin, et al.
Published: (2026)
by: Zheng, Zelin, et al.
Published: (2026)
IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs
by: Ma, David, et al.
Published: (2025)
by: Ma, David, et al.
Published: (2025)
VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding
by: Guo, Yongxin, et al.
Published: (2024)
by: Guo, Yongxin, et al.
Published: (2024)
VideoAesBench: Benchmarking the Video Aesthetics Perception Capabilities of Large Multimodal Models
by: Li, Yunhao, et al.
Published: (2026)
by: Li, Yunhao, et al.
Published: (2026)
WaterFlow: Explicit Physics-Prior Rectified Flow for Underwater Saliency Mask Generation
by: Li, Runting, et al.
Published: (2025)
by: Li, Runting, et al.
Published: (2025)
Adaptive High-Frequency Preprocessing for Video Coding
by: Pang, Yingxue, et al.
Published: (2025)
by: Pang, Yingxue, et al.
Published: (2025)
Grounding Creativity in Physics: A Brief Survey of Physical Priors in AIGC
by: Meng, Siwei, et al.
Published: (2025)
by: Meng, Siwei, et al.
Published: (2025)
An Empirical Study on How Video-LLMs Answer Video Questions
by: Gou, Chenhui, et al.
Published: (2025)
by: Gou, Chenhui, et al.
Published: (2025)
Autoregressive Image Generation Needs Only a Few Lines of Cached Tokens
by: Qin, Ziran, et al.
Published: (2025)
by: Qin, Ziran, et al.
Published: (2025)
RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
by: Huang, Haifeng, et al.
Published: (2025)
by: Huang, Haifeng, et al.
Published: (2025)
From Statics to Dynamics: Physics-Aware Image Editing with Latent Transition Priors
by: Zhao, Liangbing, et al.
Published: (2026)
by: Zhao, Liangbing, et al.
Published: (2026)
Grounding Video Reasoning in Physical Signals
by: Osmanli, Alibay, et al.
Published: (2026)
by: Osmanli, Alibay, et al.
Published: (2026)
Contrast-Unity for Partially-Supervised Temporal Sentence Grounding
by: Wang, Haicheng, et al.
Published: (2025)
by: Wang, Haicheng, et al.
Published: (2025)
DreamShot: Personalized Storyboard Synthesis with Video Diffusion Prior
by: Huang, Junjia, et al.
Published: (2026)
by: Huang, Junjia, et al.
Published: (2026)
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding
by: Li, Rong, et al.
Published: (2024)
by: Li, Rong, et al.
Published: (2024)
Taming Video Diffusion Prior with Scene-Grounding Guidance for 3D Gaussian Splatting from Sparse Inputs
by: Zhong, Yingji, et al.
Published: (2025)
by: Zhong, Yingji, et al.
Published: (2025)
Spatio-Temporal Distortion Aware Omnidirectional Video Super-Resolution
by: An, Hongyu, et al.
Published: (2024)
by: An, Hongyu, et al.
Published: (2024)
CodePercept: Code-Grounded Visual STEM Perception for MLLMs
by: Guan, Tongkun, et al.
Published: (2026)
by: Guan, Tongkun, et al.
Published: (2026)
DreamPhysics: Learning Physics-Based 3D Dynamics with Video Diffusion Priors
by: Huang, Tianyu, et al.
Published: (2024)
by: Huang, Tianyu, et al.
Published: (2024)
HawkEye: Training Video-Text LLMs for Grounding Text in Videos
by: Wang, Yueqian, et al.
Published: (2024)
by: Wang, Yueqian, et al.
Published: (2024)
ParticleGS: Learning Neural Gaussian Particle Dynamics from Videos for Prior-free Physical Motion Extrapolation
by: Quan, Jinsheng, et al.
Published: (2025)
by: Quan, Jinsheng, et al.
Published: (2025)
Q-Ground: Image Quality Grounding with Large Multi-modality Models
by: Chen, Chaofeng, et al.
Published: (2024)
by: Chen, Chaofeng, et al.
Published: (2024)
Pause and Think: A Dataset and Benchmark for Video-Grounded Assistive Action Suggestion
by: Singh, Shivam, et al.
Published: (2026)
by: Singh, Shivam, et al.
Published: (2026)
KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question Answering
by: Li, Zhiyang, et al.
Published: (2026)
by: Li, Zhiyang, et al.
Published: (2026)
Enrich and Detect: Video Temporal Grounding with Multimodal LLMs
by: Pramanick, Shraman, et al.
Published: (2025)
by: Pramanick, Shraman, et al.
Published: (2025)
SpatialReasoner: Active Perception for Large-Scale 3D Scene Understanding
by: Zheng, Hongpei, et al.
Published: (2025)
by: Zheng, Hongpei, et al.
Published: (2025)
ID-Crafter: VLM-Grounded Online RL for Compositional Multi-Subject Video Generation
by: Pan, Panwang, et al.
Published: (2025)
by: Pan, Panwang, et al.
Published: (2025)
Spatial Degradation-Aware and Temporal Consistent Diffusion Model for Compressed Video Super-Resolution
by: An, Hongyu, et al.
Published: (2025)
by: An, Hongyu, et al.
Published: (2025)
Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction
by: Qian, Rui, et al.
Published: (2025)
by: Qian, Rui, et al.
Published: (2025)
Similar Items
-
DAC: 2D-3D Retrieval with Noisy Labels via Divide-and-Conquer Alignment and Correction
by: Gan, Chaofan, et al.
Published: (2024) -
CogStream: Context-guided Streaming Video Question Answering
by: Zhao, Zicheng, et al.
Published: (2025) -
Massive Activations are the Key to Local Detail Synthesis in Diffusion Transformers
by: Gan, Chaofan, et al.
Published: (2025) -
VidLaDA: Bidirectional Diffusion Large Language Models for Efficient Video Understanding
by: He, Zhihao, et al.
Published: (2026) -
Enhancing Video Large Language Models with Structured Multi-Video Collaborative Reasoning
by: He, Zhihao, et al.
Published: (2025)