Video models are zero-shot learners and reasoners
Fuente:
arXiv
Saved in:
| Main Authors: | Wiedemer, Thaddäus, Li, Yuxuan, Vicol, Paul, Gu, Shixiang Shane, Matarese, Nick, Swersky, Kevin, Kim, Been, Jaini, Priyank, Geirhos, Robert |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Do generative video models understand physical principles?
by: Motamed, Saman, et al.
Published: (2025)
by: Motamed, Saman, et al.
Published: (2025)
Intriguing properties of generative classifiers
by: Jaini, Priyank, et al.
Published: (2023)
by: Jaini, Priyank, et al.
Published: (2023)
Directly Fine-Tuning Diffusion Models on Differentiable Rewards
by: Clark, Kevin, et al.
Published: (2023)
by: Clark, Kevin, et al.
Published: (2023)
We Can't Understand AI Using our Existing Vocabulary
by: Hewitt, John, et al.
Published: (2025)
by: Hewitt, John, et al.
Published: (2025)
Neologism Learning for Controllability and Self-Verbalization
by: Hewitt, John, et al.
Published: (2025)
by: Hewitt, John, et al.
Published: (2025)
Towards flexible perception with visual memory
by: Geirhos, Robert, et al.
Published: (2024)
by: Geirhos, Robert, et al.
Published: (2024)
Collective Intelligence for 2D Push Manipulations with Mobile Robots
by: Kuroki, So, et al.
Published: (2022)
by: Kuroki, So, et al.
Published: (2022)
MATH-Beyond: A Benchmark for RL to Expand Beyond the Base Model
by: Mayilvahanan, Prasanna, et al.
Published: (2025)
by: Mayilvahanan, Prasanna, et al.
Published: (2025)
Let people fail! Exploring the influence of explainable virtual and robotic agents in learning-by-doing tasks
by: Matarese, Marco, et al.
Published: (2024)
by: Matarese, Marco, et al.
Published: (2024)
VL-Explore: Zero-shot Vision-Language Exploration and Target Discovery by Mobile Robots
by: Zhang, Yuxuan, et al.
Published: (2025)
by: Zhang, Yuxuan, et al.
Published: (2025)
Don't trust your eyes: on the (un)reliability of feature visualizations
by: Geirhos, Robert, et al.
Published: (2023)
by: Geirhos, Robert, et al.
Published: (2023)
Does CLIP's Generalization Performance Mainly Stem from High Train-Test Similarity?
by: Mayilvahanan, Prasanna, et al.
Published: (2023)
by: Mayilvahanan, Prasanna, et al.
Published: (2023)
LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws
by: Mayilvahanan, Prasanna, et al.
Published: (2025)
by: Mayilvahanan, Prasanna, et al.
Published: (2025)
Pretraining Frequency Predicts Compositional Generalization of CLIP on Real-World Tasks
by: Wiedemer, Thaddäus, et al.
Published: (2025)
by: Wiedemer, Thaddäus, et al.
Published: (2025)
TC-IDM: Grounding Video Generation for Executable Zero-shot Robot Motion
by: Mi, Weishi, et al.
Published: (2026)
by: Mi, Weishi, et al.
Published: (2026)
Zero-shot Reconstruction of In-Scene Object Manipulation from Video
by: Lin, Dixuan, et al.
Published: (2025)
by: Lin, Dixuan, et al.
Published: (2025)
Meta-reasoning Using Attention Maps and Its Applications in Cloud Robotics
by: Lendinez, Adrian, et al.
Published: (2025)
by: Lendinez, Adrian, et al.
Published: (2025)
Federated reinforcement learning for robot motion planning with zero-shot generalization
by: Yuan, Zhenyuan, et al.
Published: (2024)
by: Yuan, Zhenyuan, et al.
Published: (2024)
How can reasoning capability empower the AI copilot robot in endoscopic surgery
by: Wang, Guankun, et al.
Published: (2026)
by: Wang, Guankun, et al.
Published: (2026)
ZEST: Zero-shot Embodied Skill Transfer for Athletic Robot Control
by: Sleiman, Jean Pierre, et al.
Published: (2026)
by: Sleiman, Jean Pierre, et al.
Published: (2026)
One-shot Video Imitation via Parameterized Symbolic Abstraction Graphs
by: Wang, Jianren, et al.
Published: (2024)
by: Wang, Jianren, et al.
Published: (2024)
Provable Compositional Generalization for Object-Centric Learning
by: Wiedemer, Thaddäus, et al.
Published: (2023)
by: Wiedemer, Thaddäus, et al.
Published: (2023)
GeoNav: Empowering MLLMs with dual-scale geospatial reasoning for language-goal aerial navigation
by: Xu, Haotian, et al.
Published: (2025)
by: Xu, Haotian, et al.
Published: (2025)
Obstruction reasoning for robotic grasping
by: Jiao, Runyu, et al.
Published: (2025)
by: Jiao, Runyu, et al.
Published: (2025)
Zero-shot Interactive Perception
by: Sripada, Venkatesh, et al.
Published: (2026)
by: Sripada, Venkatesh, et al.
Published: (2026)
Recent Advances in Variable‐Stiffness Robotic Systems Enabled by Phase‐Change Materials
by: Sukrit Gaira, et al.
Published: (2025)
by: Sukrit Gaira, et al.
Published: (2025)
Make LLMs better zero-shot reasoners: Structure-orientated autonomous reasoning
by: He, Pengfei, et al.
Published: (2024)
by: He, Pengfei, et al.
Published: (2024)
AED: Adaptable Error Detection for Few-shot Imitation Policy
by: Yeh, Jia-Fong, et al.
Published: (2024)
by: Yeh, Jia-Fong, et al.
Published: (2024)
Navigation under uncertainty: Trajectory prediction and occlusion reasoning with switching dynamical systems
by: Wei, Ran, et al.
Published: (2024)
by: Wei, Ran, et al.
Published: (2024)
SurgiPose: Estimating Surgical Tool Kinematics from Monocular Video for Surgical Robot Learning
by: Chen, Juo-Tung, et al.
Published: (2025)
by: Chen, Juo-Tung, et al.
Published: (2025)
A Clinical Tuning Framework for Continuous Kinematic and Impedance Control of a Powered Knee-Ankle Prosthesis
by: Reznick, Emma, et al.
Published: (2024)
by: Reznick, Emma, et al.
Published: (2024)
Say, Dream, and Act: Learning Video World Models for Instruction-Driven Robot Manipulation
by: Gu, Songen, et al.
Published: (2026)
by: Gu, Songen, et al.
Published: (2026)
Scalable Policy Evaluation with Video World Models
by: Tseng, Wei-Cheng, et al.
Published: (2025)
by: Tseng, Wei-Cheng, et al.
Published: (2025)
Show and Grasp: Few-shot Semantic Segmentation for Robot Grasping through Zero-shot Foundation Models
by: Barcellona, Leonardo, et al.
Published: (2024)
by: Barcellona, Leonardo, et al.
Published: (2024)
VGGSounder: Audio-Visual Evaluations for Foundation Models
by: Zverev, Daniil, et al.
Published: (2025)
by: Zverev, Daniil, et al.
Published: (2025)
Imagine2Real: Towards Zero-shot Humanoid-Object Interaction via Video Generative Priors
by: Chen, Jiahe, et al.
Published: (2026)
by: Chen, Jiahe, et al.
Published: (2026)
Embodied AI with Two Arms: Zero-shot Learning, Safety and Modularity
by: Varley, Jake, et al.
Published: (2024)
by: Varley, Jake, et al.
Published: (2024)
MR.ScaleMaster: Scale-Consistent Collaborative Mapping from Crowd-Sourced Monocular Videos
by: Ju, Hyoseok, et al.
Published: (2026)
by: Ju, Hyoseok, et al.
Published: (2026)
ROS-LLM: A ROS framework for embodied AI with task feedback and structured reasoning
by: Mower, Christopher E., et al.
Published: (2024)
by: Mower, Christopher E., et al.
Published: (2024)
Disentangling perception and reasoning for improving data efficiency in learning cloth manipulation without demonstrations
by: Delehelle, Donatien, et al.
Published: (2026)
by: Delehelle, Donatien, et al.
Published: (2026)
Similar Items
-
Do generative video models understand physical principles?
by: Motamed, Saman, et al.
Published: (2025) -
Intriguing properties of generative classifiers
by: Jaini, Priyank, et al.
Published: (2023) -
Directly Fine-Tuning Diffusion Models on Differentiable Rewards
by: Clark, Kevin, et al.
Published: (2023) -
We Can't Understand AI Using our Existing Vocabulary
by: Hewitt, John, et al.
Published: (2025) -
Neologism Learning for Controllability and Self-Verbalization
by: Hewitt, John, et al.
Published: (2025)