A General Protocol to Probe Large Vision Models for 3D Physical Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Zhan, Guanqi, Zheng, Chuanxia, Xie, Weidi, Zisserman, Andrew |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Amodal Ground Truth and Completion in the Wild
by: Zhan, Guanqi, et al.
Published: (2023)
by: Zhan, Guanqi, et al.
Published: (2023)
Inferring Dynamic Physical Properties from Video Foundation Models
by: Zhan, Guanqi, et al.
Published: (2025)
by: Zhan, Guanqi, et al.
Published: (2025)
ELIP: Enhanced Visual-Language Foundation Models for Image Retrieval
by: Zhan, Guanqi, et al.
Published: (2025)
by: Zhan, Guanqi, et al.
Published: (2025)
GMOS: Grounding Moving Object Segmentation in 3D Space and Time
by: Xie, Junyu, et al.
Published: (2026)
by: Xie, Junyu, et al.
Published: (2026)
Appearance-Based Refinement for Object-Centric Motion Segmentation
by: Xie, Junyu, et al.
Published: (2023)
by: Xie, Junyu, et al.
Published: (2023)
Character-Centric Understanding of Animated Movies
by: Gui, Zhongrui, et al.
Published: (2025)
by: Gui, Zhongrui, et al.
Published: (2025)
Moving Object Segmentation: All You Need Is SAM (and Flow)
by: Xie, Junyu, et al.
Published: (2024)
by: Xie, Junyu, et al.
Published: (2024)
Made to Order: Discovering monotonic temporal changes via self-supervised video ordering
by: Yang, Charig, et al.
Published: (2024)
by: Yang, Charig, et al.
Published: (2024)
Free3D: Consistent Novel View Synthesis without 3D Representation
by: Zheng, Chuanxia, et al.
Published: (2023)
by: Zheng, Chuanxia, et al.
Published: (2023)
AutoAD-Zero: A Training-Free Framework for Zero-Shot Audio Description
by: Xie, Junyu, et al.
Published: (2024)
by: Xie, Junyu, et al.
Published: (2024)
AutoAD III: The Prequel -- Back to the Pixels
by: Han, Tengda, et al.
Published: (2024)
by: Han, Tengda, et al.
Published: (2024)
Shot-by-Shot: Film-Grammar-Aware Training-Free Audio Description Generation
by: Xie, Junyu, et al.
Published: (2025)
by: Xie, Junyu, et al.
Published: (2025)
SoccerMaster: A Vision Foundation Model for Soccer Understanding
by: Yang, Haolin, et al.
Published: (2025)
by: Yang, Haolin, et al.
Published: (2025)
DSO: Aligning 3D Generators with Simulation Feedback for Physical Soundness
by: Li, Ruining, et al.
Published: (2025)
by: Li, Ruining, et al.
Published: (2025)
The Manga Whisperer: Automatically Generating Transcriptions for Comics
by: Sachdeva, Ragav, et al.
Published: (2024)
by: Sachdeva, Ragav, et al.
Published: (2024)
From Panels to Prose: Generating Literary Narratives from Comics
by: Sachdeva, Ragav, et al.
Published: (2025)
by: Sachdeva, Ragav, et al.
Published: (2025)
Synchformer: Efficient Synchronization from Sparse Cues
by: Iashin, Vladimir, et al.
Published: (2024)
by: Iashin, Vladimir, et al.
Published: (2024)
CountGD++: Generalized Prompting for Open-World Counting
by: Amini-Naieni, Niki, et al.
Published: (2025)
by: Amini-Naieni, Niki, et al.
Published: (2025)
SPATIALALIGN: Aligning Dynamic Spatial Relationships in Video Generation
by: Liu, Fengming, et al.
Published: (2026)
by: Liu, Fengming, et al.
Published: (2026)
AnyHand: A Large-Scale Synthetic Dataset for RGB(-D) Hand Pose Estimation
by: Si, Chen, et al.
Published: (2026)
by: Si, Chen, et al.
Published: (2026)
ClusteringSDF: Self-Organized Neural Implicit Surfaces for 3D Decomposition
by: Wu, Tianhao, et al.
Published: (2024)
by: Wu, Tianhao, et al.
Published: (2024)
Amodal3R: Amodal 3D Reconstruction from Occluded 2D Images
by: Wu, Tianhao, et al.
Published: (2025)
by: Wu, Tianhao, et al.
Published: (2025)
NOVA3R: Non-pixel-aligned Visual Transformer for Amodal 3D Reconstruction
by: Chen, Weirong, et al.
Published: (2026)
by: Chen, Weirong, et al.
Published: (2026)
EchoSight: Advancing Visual-Language Models with Wiki Knowledge
by: Yan, Yibin, et al.
Published: (2024)
by: Yan, Yibin, et al.
Published: (2024)
Personalizing Retrieval using Joint Embeddings or "the Return of Fluffy"
by: Korbar, Bruno, et al.
Published: (2025)
by: Korbar, Bruno, et al.
Published: (2025)
Chirality in Action: Time-Aware Video Representation Learning by Latent Straightening
by: Bagad, Piyush, et al.
Published: (2025)
by: Bagad, Piyush, et al.
Published: (2025)
Grounded 3D-Aware Spatial Vision-Language Modeling
by: Cheng, An-Chieh, et al.
Published: (2026)
by: Cheng, An-Chieh, et al.
Published: (2026)
Understanding Co-speech Gestures in-the-wild
by: Hegde, Sindhu B, et al.
Published: (2025)
by: Hegde, Sindhu B, et al.
Published: (2025)
3D Spine Shape Estimation from Single 2D DXA
by: Bourigault, Emmanuelle, et al.
Published: (2024)
by: Bourigault, Emmanuelle, et al.
Published: (2024)
Text-Conditioned Resampler For Long Form Video Understanding
by: Korbar, Bruno, et al.
Published: (2023)
by: Korbar, Bruno, et al.
Published: (2023)
Mesh4D: 4D Mesh Reconstruction and Tracking from Monocular Video
by: Jiang, Zeren, et al.
Published: (2026)
by: Jiang, Zeren, et al.
Published: (2026)
Universal Video Temporal Grounding with Generative Multi-modal Large Language Models
by: Li, Zeqian, et al.
Published: (2025)
by: Li, Zeqian, et al.
Published: (2025)
SceneGen: Single-Image 3D Scene Generation in One Feedforward Pass
by: Meng, Yanxu, et al.
Published: (2025)
by: Meng, Yanxu, et al.
Published: (2025)
PanoDiffusion: 360-degree Panorama Outpainting via Diffusion
by: Wu, Tianhao, et al.
Published: (2023)
by: Wu, Tianhao, et al.
Published: (2023)
One-shot Human Motion Transfer via Occlusion-Robust Flow Prediction and Neural Texturing
by: Ji, Yuzhu, et al.
Published: (2024)
by: Ji, Yuzhu, et al.
Published: (2024)
Flash3D: Feed-Forward Generalisable 3D Scene Reconstruction from a Single Image
by: Szymanowicz, Stanislaw, et al.
Published: (2024)
by: Szymanowicz, Stanislaw, et al.
Published: (2024)
Grounded Question-Answering in Long Egocentric Videos
by: Di, Shangzhe, et al.
Published: (2023)
by: Di, Shangzhe, et al.
Published: (2023)
Geo4D: Leveraging Video Generators for Geometric 4D Scene Reconstruction
by: Jiang, Zeren, et al.
Published: (2025)
by: Jiang, Zeren, et al.
Published: (2025)
Can Visual Foundation Models Achieve Long-term Point Tracking?
by: Aydemir, Görkay, et al.
Published: (2024)
by: Aydemir, Görkay, et al.
Published: (2024)
Recognising BSL Fingerspelling in Continuous Signing Sequences
by: Chan, Alyssa, et al.
Published: (2026)
by: Chan, Alyssa, et al.
Published: (2026)
Similar Items
-
Amodal Ground Truth and Completion in the Wild
by: Zhan, Guanqi, et al.
Published: (2023) -
Inferring Dynamic Physical Properties from Video Foundation Models
by: Zhan, Guanqi, et al.
Published: (2025) -
ELIP: Enhanced Visual-Language Foundation Models for Image Retrieval
by: Zhan, Guanqi, et al.
Published: (2025) -
GMOS: Grounding Moving Object Segmentation in 3D Space and Time
by: Xie, Junyu, et al.
Published: (2026) -
Appearance-Based Refinement for Object-Centric Motion Segmentation
by: Xie, Junyu, et al.
Published: (2023)