Generative Image as Action Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shridhar, Mohit, Lo, Yat Long, James, Stephen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PerAct2: Benchmarking and Learning for Robotic Bimanual Manipulation Tasks
von: Grotz, Markus, et al.
Veröffentlicht: (2024)
von: Grotz, Markus, et al.
Veröffentlicht: (2024)
SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts
von: Zhou, Gengze, et al.
Veröffentlicht: (2024)
von: Zhou, Gengze, et al.
Veröffentlicht: (2024)
CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
von: Li, Qixiu, et al.
Veröffentlicht: (2024)
von: Li, Qixiu, et al.
Veröffentlicht: (2024)
Render and Diffuse: Aligning Image and Action Spaces for Diffusion-based Behaviour Cloning
von: Vosylius, Vitalis, et al.
Veröffentlicht: (2024)
von: Vosylius, Vitalis, et al.
Veröffentlicht: (2024)
GenSim: Generating Robotic Simulation Tasks via Large Language Models
von: Wang, Lirui, et al.
Veröffentlicht: (2023)
von: Wang, Lirui, et al.
Veröffentlicht: (2023)
SlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video Generation
von: Hong, Yining, et al.
Veröffentlicht: (2024)
von: Hong, Yining, et al.
Veröffentlicht: (2024)
LangGap: Diagnosing and Closing the Language Gap in Vision-Language-Action Models
von: Hou, Yuchen, et al.
Veröffentlicht: (2026)
von: Hou, Yuchen, et al.
Veröffentlicht: (2026)
MAPS: Preserving Vision-Language Representations via Module-Wise Proximity Scheduling for Better Vision-Language-Action Generalization
von: Huang, Chengyue, et al.
Veröffentlicht: (2025)
von: Huang, Chengyue, et al.
Veröffentlicht: (2025)
ViPRA: Video Prediction for Robot Actions
von: Routray, Sandeep, et al.
Veröffentlicht: (2025)
von: Routray, Sandeep, et al.
Veröffentlicht: (2025)
Redundancy-aware Action Spaces for Robot Learning
von: Mazzaglia, Pietro, et al.
Veröffentlicht: (2024)
von: Mazzaglia, Pietro, et al.
Veröffentlicht: (2024)
ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop
von: Hong, Yining, et al.
Veröffentlicht: (2026)
von: Hong, Yining, et al.
Veröffentlicht: (2026)
Pre-trained Text-to-Image Diffusion Models Are Versatile Representation Learners for Control
von: Gupta, Gunshi, et al.
Veröffentlicht: (2024)
von: Gupta, Gunshi, et al.
Veröffentlicht: (2024)
EMMA: End-to-End Multimodal Model for Autonomous Driving
von: Hwang, Jyh-Jing, et al.
Veröffentlicht: (2024)
von: Hwang, Jyh-Jing, et al.
Veröffentlicht: (2024)
Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations
von: Grover, Shresth, et al.
Veröffentlicht: (2025)
von: Grover, Shresth, et al.
Veröffentlicht: (2025)
Critiques of World Models
von: Xing, Eric, et al.
Veröffentlicht: (2025)
von: Xing, Eric, et al.
Veröffentlicht: (2025)
LanguageMPC: Large Language Models as Decision Makers for Autonomous Driving
von: Sha, Hao, et al.
Veröffentlicht: (2023)
von: Sha, Hao, et al.
Veröffentlicht: (2023)
Energy-based Models are Zero-Shot Planners for Compositional Scene Rearrangement
von: Gkanatsios, Nikolaos, et al.
Veröffentlicht: (2023)
von: Gkanatsios, Nikolaos, et al.
Veröffentlicht: (2023)
AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents
von: Ahn, Michael, et al.
Veröffentlicht: (2024)
von: Ahn, Michael, et al.
Veröffentlicht: (2024)
Vision-Language Model Fine-Tuning via Simple Parameter-Efficient Modification
von: Li, Ming, et al.
Veröffentlicht: (2024)
von: Li, Ming, et al.
Veröffentlicht: (2024)
Embodied AI with Foundation Models for Mobile Service Robots: A Systematic Review
von: Lisondra, Matthew, et al.
Veröffentlicht: (2025)
von: Lisondra, Matthew, et al.
Veröffentlicht: (2025)
PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
von: Chow, Wei, et al.
Veröffentlicht: (2025)
von: Chow, Wei, et al.
Veröffentlicht: (2025)
Language and Planning in Robotic Navigation: A Multilingual Evaluation of State-of-the-Art Models
von: Mansour, Malak, et al.
Veröffentlicht: (2025)
von: Mansour, Malak, et al.
Veröffentlicht: (2025)
Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene Understanding
von: Man, Yunze, et al.
Veröffentlicht: (2024)
von: Man, Yunze, et al.
Veröffentlicht: (2024)
A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges
von: Li, Zongxia, et al.
Veröffentlicht: (2025)
von: Li, Zongxia, et al.
Veröffentlicht: (2025)
MultiPLY: A Multisensory Object-Centric Embodied Large Language Model in 3D World
von: Hong, Yining, et al.
Veröffentlicht: (2024)
von: Hong, Yining, et al.
Veröffentlicht: (2024)
From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors
von: Zhang, Zhengshen, et al.
Veröffentlicht: (2025)
von: Zhang, Zhengshen, et al.
Veröffentlicht: (2025)
Diagnostic Benchmark and Iterative Inpainting for Layout-Guided Image Generation
von: Cho, Jaemin, et al.
Veröffentlicht: (2023)
von: Cho, Jaemin, et al.
Veröffentlicht: (2023)
Hybrid Training for Vision-Language-Action Models
von: Mazzaglia, Pietro, et al.
Veröffentlicht: (2025)
von: Mazzaglia, Pietro, et al.
Veröffentlicht: (2025)
SELMA: Learning and Merging Skill-Specific Text-to-Image Experts with Auto-Generated Data
von: Li, Jialu, et al.
Veröffentlicht: (2024)
von: Li, Jialu, et al.
Veröffentlicht: (2024)
A Survey on Efficient Vision-Language-Action Models
von: Yu, Zhaoshu, et al.
Veröffentlicht: (2025)
von: Yu, Zhaoshu, et al.
Veröffentlicht: (2025)
Interactive Post-Training for Vision-Language-Action Models
von: Tan, Shuhan, et al.
Veröffentlicht: (2025)
von: Tan, Shuhan, et al.
Veröffentlicht: (2025)
CorrectNav: Self-Correction Flywheel Empowers Vision-Language-Action Navigation Model
von: Yu, Zhuoyuan, et al.
Veröffentlicht: (2025)
von: Yu, Zhuoyuan, et al.
Veröffentlicht: (2025)
Grounding Video Models to Actions through Goal Conditioned Exploration
von: Luo, Yunhao, et al.
Veröffentlicht: (2024)
von: Luo, Yunhao, et al.
Veröffentlicht: (2024)
AdaWorld: Learning Adaptable World Models with Latent Actions
von: Gao, Shenyuan, et al.
Veröffentlicht: (2025)
von: Gao, Shenyuan, et al.
Veröffentlicht: (2025)
UAV-VLA: Vision-Language-Action System for Large Scale Aerial Mission Generation
von: Sautenkov, Oleg, et al.
Veröffentlicht: (2025)
von: Sautenkov, Oleg, et al.
Veröffentlicht: (2025)
MVSA-Net: Multi-View State-Action Recognition for Robust and Deployable Trajectory Generation
von: Asali, Ehsan, et al.
Veröffentlicht: (2023)
von: Asali, Ehsan, et al.
Veröffentlicht: (2023)
Video-Language Critic: Transferable Reward Functions for Language-Conditioned Robotics
von: Alakuijala, Minttu, et al.
Veröffentlicht: (2024)
von: Alakuijala, Minttu, et al.
Veröffentlicht: (2024)
Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos
von: Chen, Yi, et al.
Veröffentlicht: (2024)
von: Chen, Yi, et al.
Veröffentlicht: (2024)
Teaching Embodied Reinforcement Learning Agents: Informativeness and Diversity of Language Use
von: Xi, Jiajun, et al.
Veröffentlicht: (2024)
von: Xi, Jiajun, et al.
Veröffentlicht: (2024)
DecisionNCE: Embodied Multimodal Representations via Implicit Preference Learning
von: Li, Jianxiong, et al.
Veröffentlicht: (2024)
von: Li, Jianxiong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
PerAct2: Benchmarking and Learning for Robotic Bimanual Manipulation Tasks
von: Grotz, Markus, et al.
Veröffentlicht: (2024) -
SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts
von: Zhou, Gengze, et al.
Veröffentlicht: (2024) -
CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
von: Li, Qixiu, et al.
Veröffentlicht: (2024) -
Render and Diffuse: Aligning Image and Action Spaces for Diffusion-based Behaviour Cloning
von: Vosylius, Vitalis, et al.
Veröffentlicht: (2024) -
GenSim: Generating Robotic Simulation Tasks via Large Language Models
von: Wang, Lirui, et al.
Veröffentlicht: (2023)