Inferring Dynamic Physical Properties from Video Foundation Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhan, Guanqi, Ma, Xianzheng, Xie, Weidi, Zisserman, Andrew |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A General Protocol to Probe Large Vision Models for 3D Physical Understanding
by: Zhan, Guanqi, et al.
Published: (2023)
by: Zhan, Guanqi, et al.
Published: (2023)
ELIP: Enhanced Visual-Language Foundation Models for Image Retrieval
by: Zhan, Guanqi, et al.
Published: (2025)
by: Zhan, Guanqi, et al.
Published: (2025)
Amodal Ground Truth and Completion in the Wild
by: Zhan, Guanqi, et al.
Published: (2023)
by: Zhan, Guanqi, et al.
Published: (2023)
Made to Order: Discovering monotonic temporal changes via self-supervised video ordering
by: Yang, Charig, et al.
Published: (2024)
by: Yang, Charig, et al.
Published: (2024)
Synchformer: Efficient Synchronization from Sparse Cues
by: Iashin, Vladimir, et al.
Published: (2024)
by: Iashin, Vladimir, et al.
Published: (2024)
Appearance-Based Refinement for Object-Centric Motion Segmentation
by: Xie, Junyu, et al.
Published: (2023)
by: Xie, Junyu, et al.
Published: (2023)
Moving Object Segmentation: All You Need Is SAM (and Flow)
by: Xie, Junyu, et al.
Published: (2024)
by: Xie, Junyu, et al.
Published: (2024)
GMOS: Grounding Moving Object Segmentation in 3D Space and Time
by: Xie, Junyu, et al.
Published: (2026)
by: Xie, Junyu, et al.
Published: (2026)
OVR: A Dataset for Open Vocabulary Temporal Repetition Counting in Videos
by: Dwibedi, Debidatta, et al.
Published: (2024)
by: Dwibedi, Debidatta, et al.
Published: (2024)
Perception Test 2024: Challenge Summary and a Novel Hour-Long VideoQA Benchmark
by: Heyward, Joseph, et al.
Published: (2024)
by: Heyward, Joseph, et al.
Published: (2024)
Character-Centric Understanding of Animated Movies
by: Gui, Zhongrui, et al.
Published: (2025)
by: Gui, Zhongrui, et al.
Published: (2025)
A Short Note on Evaluating RepNet for Temporal Repetition Counting in Videos
by: Dwibedi, Debidatta, et al.
Published: (2024)
by: Dwibedi, Debidatta, et al.
Published: (2024)
Understanding Particles From Video: Property Estimation of Granular Materials via Visuo-Haptic Learning
by: Zhang, Zeqing, et al.
Published: (2024)
by: Zhang, Zeqing, et al.
Published: (2024)
Character-aware audio-visual subtitling in context
by: Huh, Jaesung, et al.
Published: (2024)
by: Huh, Jaesung, et al.
Published: (2024)
3D-Aware Instance Segmentation and Tracking in Egocentric Videos
by: Bhalgat, Yash, et al.
Published: (2024)
by: Bhalgat, Yash, et al.
Published: (2024)
The Sound of Water: Inferring Physical Properties from Pouring Liquids
by: Bagad, Piyush, et al.
Published: (2024)
by: Bagad, Piyush, et al.
Published: (2024)
Progressive Autoregressive Video Diffusion Models
by: Xie, Desai, et al.
Published: (2024)
by: Xie, Desai, et al.
Published: (2024)
VMDT: Decoding the Trustworthiness of Video Foundation Models
by: Potter, Yujin, et al.
Published: (2025)
by: Potter, Yujin, et al.
Published: (2025)
Self-supervised Learning on Camera Trap Footage Yields a Strong Universal Face Embedder
by: Iashin, Vladimir, et al.
Published: (2025)
by: Iashin, Vladimir, et al.
Published: (2025)
VideoSAM: A Large Vision Foundation Model for High-Speed Video Segmentation
by: Maduabuchi, Chika, et al.
Published: (2024)
by: Maduabuchi, Chika, et al.
Published: (2024)
World Simulation with Video Foundation Models for Physical AI
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
AutoAD III: The Prequel -- Back to the Pixels
by: Han, Tengda, et al.
Published: (2024)
by: Han, Tengda, et al.
Published: (2024)
Chirality in Action: Time-Aware Video Representation Learning by Latent Straightening
by: Bagad, Piyush, et al.
Published: (2025)
by: Bagad, Piyush, et al.
Published: (2025)
Graph2Video: Leveraging Video Models to Model Dynamic Graph Evolution
by: Liu, Hua, et al.
Published: (2026)
by: Liu, Hua, et al.
Published: (2026)
AutoAD-Zero: A Training-Free Framework for Zero-Shot Audio Description
by: Xie, Junyu, et al.
Published: (2024)
by: Xie, Junyu, et al.
Published: (2024)
Foundation Model-oriented Robustness: Robust Image Model Evaluation with Pretrained Models
by: Zhang, Peiyan, et al.
Published: (2023)
by: Zhang, Peiyan, et al.
Published: (2023)
SciVid: Cross-Domain Evaluation of Video Models in Scientific Applications
by: Hasson, Yana, et al.
Published: (2025)
by: Hasson, Yana, et al.
Published: (2025)
Video Prediction of Dynamic Physical Simulations With Pixel-Space Spatiotemporal Transformers
by: Slack, Dean L, et al.
Published: (2025)
by: Slack, Dean L, et al.
Published: (2025)
Can Visual Foundation Models Achieve Long-term Point Tracking?
by: Aydemir, Görkay, et al.
Published: (2024)
by: Aydemir, Görkay, et al.
Published: (2024)
Training Video Foundation Models with NVIDIA NeMo
by: Patel, Zeeshan, et al.
Published: (2025)
by: Patel, Zeeshan, et al.
Published: (2025)
Grounded Question-Answering in Long Egocentric Videos
by: Di, Shangzhe, et al.
Published: (2023)
by: Di, Shangzhe, et al.
Published: (2023)
Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer
by: Cui, Jiahao, et al.
Published: (2024)
by: Cui, Jiahao, et al.
Published: (2024)
AVID: Adapting Video Diffusion Models to World Models
by: Rigter, Marc, et al.
Published: (2024)
by: Rigter, Marc, et al.
Published: (2024)
Video Diffusion Models: A Survey
by: Melnik, Andrew, et al.
Published: (2024)
by: Melnik, Andrew, et al.
Published: (2024)
FlexCap: Describe Anything in Images in Controllable Detail
by: Dwibedi, Debidatta, et al.
Published: (2024)
by: Dwibedi, Debidatta, et al.
Published: (2024)
VideoOrion: Tokenizing Object Dynamics in Videos
by: Feng, Yicheng, et al.
Published: (2024)
by: Feng, Yicheng, et al.
Published: (2024)
Sora Detector: A Unified Hallucination Detection for Large Text-to-Video Models
by: Chu, Zhixuan, et al.
Published: (2024)
by: Chu, Zhixuan, et al.
Published: (2024)
What's on Your Plate? Inferring Chinese Cuisine Intake from Wearable IMUs
by: Yin, Jiaxi, et al.
Published: (2025)
by: Yin, Jiaxi, et al.
Published: (2025)
LightCache: Memory-Efficient, Training-Free Acceleration for Video Generation
by: Xiao, Yang, et al.
Published: (2025)
by: Xiao, Yang, et al.
Published: (2025)
Adapting MLLMs for Nuanced Video Retrieval
by: Bagad, Piyush, et al.
Published: (2025)
by: Bagad, Piyush, et al.
Published: (2025)
Similar Items
-
A General Protocol to Probe Large Vision Models for 3D Physical Understanding
by: Zhan, Guanqi, et al.
Published: (2023) -
ELIP: Enhanced Visual-Language Foundation Models for Image Retrieval
by: Zhan, Guanqi, et al.
Published: (2025) -
Amodal Ground Truth and Completion in the Wild
by: Zhan, Guanqi, et al.
Published: (2023) -
Made to Order: Discovering monotonic temporal changes via self-supervised video ordering
by: Yang, Charig, et al.
Published: (2024) -
Synchformer: Efficient Synchronization from Sparse Cues
by: Iashin, Vladimir, et al.
Published: (2024)