Learning to Grasp Anything by Playing with Random Toys
Fuente:
arXiv
Saved in:
| Main Authors: | Niu, Dantong, Sharma, Yuvan, Shi, Baifeng, Ding, Rachel, Gioia, Matteo, Xue, Haoru, Tsai, Henry, Kallidromitis, Konstantinos, Pai, Anirudh, Regan, Caitlin, Sastry, Shankar, Darrell, Trevor, Malik, Jitendra, Herzig, Roei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning
by: Niu, Dantong, et al.
Published: (2024)
by: Niu, Dantong, et al.
Published: (2024)
In-Context Learning Enables Robot Action Prediction in LLMs
by: Yin, Yida, et al.
Published: (2024)
by: Yin, Yida, et al.
Published: (2024)
Pre-training Auto-regressive Robotic Models with 4D Representations
by: Niu, Dantong, et al.
Published: (2025)
by: Niu, Dantong, et al.
Published: (2025)
Recursive Visual Programming
by: Ge, Jiaxin, et al.
Published: (2023)
by: Ge, Jiaxin, et al.
Published: (2023)
Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
by: Mitra, Chancharik, et al.
Published: (2025)
by: Mitra, Chancharik, et al.
Published: (2025)
Do What? Teaching Vision-Language-Action Models to Reject the Impossible
by: Hsieh, Wen-Han, et al.
Published: (2025)
by: Hsieh, Wen-Han, et al.
Published: (2025)
Compositional Chain-of-Thought Prompting for Large Multimodal Models
by: Mitra, Chancharik, et al.
Published: (2023)
by: Mitra, Chancharik, et al.
Published: (2023)
From Generated Human Videos to Physically Plausible Robot Trajectories
by: Ni, James, et al.
Published: (2025)
by: Ni, James, et al.
Published: (2025)
Navigating the Labyrinth: Evaluating LLMs' Ability to Reason About Search Problems
by: Borazjanizadeh, Nasim, et al.
Published: (2024)
by: Borazjanizadeh, Nasim, et al.
Published: (2024)
TraveLER: A Modular Multi-LMM Agent Framework for Video Question-Answering
by: Shang, Chuyi, et al.
Published: (2024)
by: Shang, Chuyi, et al.
Published: (2024)
LeVERB: Humanoid Whole-Body Control with Latent Vision-Language Instruction
by: Xue, Haoru, et al.
Published: (2025)
by: Xue, Haoru, et al.
Published: (2025)
DAVE: A VLM Vision Encoder for Document Understanding and Web Agents
by: Huang, Brandon, et al.
Published: (2025)
by: Huang, Brandon, et al.
Published: (2025)
Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning
by: Huang, Brandon, et al.
Published: (2024)
by: Huang, Brandon, et al.
Published: (2024)
Visualizing Thought: Conceptual Diagrams Enable Robust Planning in LMMs
by: Borazjanizadeh, Nasim, et al.
Published: (2025)
by: Borazjanizadeh, Nasim, et al.
Published: (2025)
Latent Implicit Visual Reasoning
by: Li, Kelvin, et al.
Published: (2025)
by: Li, Kelvin, et al.
Published: (2025)
Humanoid Locomotion as Next Token Prediction
by: Radosavovic, Ilija, et al.
Published: (2024)
by: Radosavovic, Ilija, et al.
Published: (2024)
Learning Humanoid Locomotion over Challenging Terrain
by: Radosavovic, Ilija, et al.
Published: (2024)
by: Radosavovic, Ilija, et al.
Published: (2024)
Segment Anything without Supervision
by: Wang, XuDong, et al.
Published: (2024)
by: Wang, XuDong, et al.
Published: (2024)
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
by: Qin, Yiming, et al.
Published: (2025)
by: Qin, Yiming, et al.
Published: (2025)
Hyperbolic Active Learning for Semantic Segmentation under Domain Shift
by: Franco, Luca, et al.
Published: (2023)
by: Franco, Luca, et al.
Published: (2023)
LLM-grounded Video Diffusion Models
by: Lian, Long, et al.
Published: (2023)
by: Lian, Long, et al.
Published: (2023)
When Do We Not Need Larger Vision Models?
by: Shi, Baifeng, et al.
Published: (2024)
by: Shi, Baifeng, et al.
Published: (2024)
Toy Libraries: Learning through Play with Toys.
by: Kapellaka, Urania
Published: (1992)
by: Kapellaka, Urania
Published: (1992)
FoundationMotion: Auto-Labeling and Reasoning about Spatial Movement in Videos
by: Gan, Yulu, et al.
Published: (2025)
by: Gan, Yulu, et al.
Published: (2025)
Independent and Decentralized Learning in Markov Potential Games
by: Maheshwari, Chinmay, et al.
Published: (2022)
by: Maheshwari, Chinmay, et al.
Published: (2022)
SegLLM: Multi-round Reasoning Segmentation
by: Wang, XuDong, et al.
Published: (2024)
by: Wang, XuDong, et al.
Published: (2024)
xT: Nested Tokenization for Larger Context in Large Images
by: Gupta, Ritwik, et al.
Published: (2024)
by: Gupta, Ritwik, et al.
Published: (2024)
UnSAMv2: Self-Supervised Learning Enables Segment Anything at Any Granularity
by: Yu, Junwei, et al.
Published: (2025)
by: Yu, Junwei, et al.
Published: (2025)
Toy Libraries: Promoting Play, Toys, and Family Support Internationally.
by: Mayfield, Margie I.
Published: (1993)
by: Mayfield, Margie I.
Published: (1993)
Bibliography of Children's Play and Toys.
by: Quilitch, H. Robert
Published: (1974)
by: Quilitch, H. Robert
Published: (1974)
Whole-Body Conditioned Egocentric Video Prediction
by: Bai, Yutong, et al.
Published: (2025)
by: Bai, Yutong, et al.
Published: (2025)
Cellular Learning: Scattered Data Regression in High Dimensions via Voronoi Cells
by: Sastry, Shankar Prasad
Published: (2025)
by: Sastry, Shankar Prasad
Published: (2025)
Play and Learn with Toys; A Bibliography of Toys that "Teach Institutionalized Children".
Published: (1976)
Published: (1976)
TULIP: Towards Unified Language-Image Pretraining
by: Tang, Zineng, et al.
Published: (2025)
by: Tang, Zineng, et al.
Published: (2025)
Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features
by: Mitra, Chancharik, et al.
Published: (2024)
by: Mitra, Chancharik, et al.
Published: (2024)
HITTER: A HumanoId Table TEnnis Robot via Hierarchical Planning and Learning
by: Su, Zhi, et al.
Published: (2025)
by: Su, Zhi, et al.
Published: (2025)
Evaluación del comportamiento productivo en cerdos en crecimiento alimentados con una dieta no convencional
by: Yuván Contino-Esquijerosa
Published: (2017)
by: Yuván Contino-Esquijerosa
Published: (2017)
Adopción de nuevas prácticas agroecológicas en tres unidades básicas de producción cooperativa
by: Yuván Contino Esquijerosa
Published: (2018)
by: Yuván Contino Esquijerosa
Published: (2018)
Play for All Children: The Toy Library Solution.
by: Jackson, Sara C, et al.
Published: (1991)
by: Jackson, Sara C, et al.
Published: (1991)
Activation Reward Models for Few-Shot Model Alignment
by: Chai, Tianning, et al.
Published: (2025)
by: Chai, Tianning, et al.
Published: (2025)
Similar Items
-
LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning
by: Niu, Dantong, et al.
Published: (2024) -
In-Context Learning Enables Robot Action Prediction in LLMs
by: Yin, Yida, et al.
Published: (2024) -
Pre-training Auto-regressive Robotic Models with 4D Representations
by: Niu, Dantong, et al.
Published: (2025) -
Recursive Visual Programming
by: Ge, Jiaxin, et al.
Published: (2023) -
Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
by: Mitra, Chancharik, et al.
Published: (2025)