Salvato in:
| Autori principali: | Shridhar, Mohit, Lo, Yat Long, James, Stephen |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2407.07875 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
PerAct2: Benchmarking and Learning for Robotic Bimanual Manipulation Tasks
di: Grotz, Markus, et al.
Pubblicazione: (2024)
di: Grotz, Markus, et al.
Pubblicazione: (2024)
SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts
di: Zhou, Gengze, et al.
Pubblicazione: (2024)
di: Zhou, Gengze, et al.
Pubblicazione: (2024)
CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
di: Li, Qixiu, et al.
Pubblicazione: (2024)
di: Li, Qixiu, et al.
Pubblicazione: (2024)
GenSim: Generating Robotic Simulation Tasks via Large Language Models
di: Wang, Lirui, et al.
Pubblicazione: (2023)
di: Wang, Lirui, et al.
Pubblicazione: (2023)
SlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video Generation
di: Hong, Yining, et al.
Pubblicazione: (2024)
di: Hong, Yining, et al.
Pubblicazione: (2024)
LangGap: Diagnosing and Closing the Language Gap in Vision-Language-Action Models
di: Hou, Yuchen, et al.
Pubblicazione: (2026)
di: Hou, Yuchen, et al.
Pubblicazione: (2026)
Render and Diffuse: Aligning Image and Action Spaces for Diffusion-based Behaviour Cloning
di: Vosylius, Vitalis, et al.
Pubblicazione: (2024)
di: Vosylius, Vitalis, et al.
Pubblicazione: (2024)
MAPS: Preserving Vision-Language Representations via Module-Wise Proximity Scheduling for Better Vision-Language-Action Generalization
di: Huang, Chengyue, et al.
Pubblicazione: (2025)
di: Huang, Chengyue, et al.
Pubblicazione: (2025)
ViPRA: Video Prediction for Robot Actions
di: Routray, Sandeep, et al.
Pubblicazione: (2025)
di: Routray, Sandeep, et al.
Pubblicazione: (2025)
ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop
di: Hong, Yining, et al.
Pubblicazione: (2026)
di: Hong, Yining, et al.
Pubblicazione: (2026)
Redundancy-aware Action Spaces for Robot Learning
di: Mazzaglia, Pietro, et al.
Pubblicazione: (2024)
di: Mazzaglia, Pietro, et al.
Pubblicazione: (2024)
Pre-trained Text-to-Image Diffusion Models Are Versatile Representation Learners for Control
di: Gupta, Gunshi, et al.
Pubblicazione: (2024)
di: Gupta, Gunshi, et al.
Pubblicazione: (2024)
EMMA: End-to-End Multimodal Model for Autonomous Driving
di: Hwang, Jyh-Jing, et al.
Pubblicazione: (2024)
di: Hwang, Jyh-Jing, et al.
Pubblicazione: (2024)
Critiques of World Models
di: Xing, Eric, et al.
Pubblicazione: (2025)
di: Xing, Eric, et al.
Pubblicazione: (2025)
LanguageMPC: Large Language Models as Decision Makers for Autonomous Driving
di: Sha, Hao, et al.
Pubblicazione: (2023)
di: Sha, Hao, et al.
Pubblicazione: (2023)
Energy-based Models are Zero-Shot Planners for Compositional Scene Rearrangement
di: Gkanatsios, Nikolaos, et al.
Pubblicazione: (2023)
di: Gkanatsios, Nikolaos, et al.
Pubblicazione: (2023)
AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents
di: Ahn, Michael, et al.
Pubblicazione: (2024)
di: Ahn, Michael, et al.
Pubblicazione: (2024)
Vision-Language Model Fine-Tuning via Simple Parameter-Efficient Modification
di: Li, Ming, et al.
Pubblicazione: (2024)
di: Li, Ming, et al.
Pubblicazione: (2024)
Embodied AI with Foundation Models for Mobile Service Robots: A Systematic Review
di: Lisondra, Matthew, et al.
Pubblicazione: (2025)
di: Lisondra, Matthew, et al.
Pubblicazione: (2025)
PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
di: Chow, Wei, et al.
Pubblicazione: (2025)
di: Chow, Wei, et al.
Pubblicazione: (2025)
Language and Planning in Robotic Navigation: A Multilingual Evaluation of State-of-the-Art Models
di: Mansour, Malak, et al.
Pubblicazione: (2025)
di: Mansour, Malak, et al.
Pubblicazione: (2025)
Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations
di: Grover, Shresth, et al.
Pubblicazione: (2025)
di: Grover, Shresth, et al.
Pubblicazione: (2025)
Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene Understanding
di: Man, Yunze, et al.
Pubblicazione: (2024)
di: Man, Yunze, et al.
Pubblicazione: (2024)
A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges
di: Li, Zongxia, et al.
Pubblicazione: (2025)
di: Li, Zongxia, et al.
Pubblicazione: (2025)
MultiPLY: A Multisensory Object-Centric Embodied Large Language Model in 3D World
di: Hong, Yining, et al.
Pubblicazione: (2024)
di: Hong, Yining, et al.
Pubblicazione: (2024)
From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors
di: Zhang, Zhengshen, et al.
Pubblicazione: (2025)
di: Zhang, Zhengshen, et al.
Pubblicazione: (2025)
Diagnostic Benchmark and Iterative Inpainting for Layout-Guided Image Generation
di: Cho, Jaemin, et al.
Pubblicazione: (2023)
di: Cho, Jaemin, et al.
Pubblicazione: (2023)
Hybrid Training for Vision-Language-Action Models
di: Mazzaglia, Pietro, et al.
Pubblicazione: (2025)
di: Mazzaglia, Pietro, et al.
Pubblicazione: (2025)
SELMA: Learning and Merging Skill-Specific Text-to-Image Experts with Auto-Generated Data
di: Li, Jialu, et al.
Pubblicazione: (2024)
di: Li, Jialu, et al.
Pubblicazione: (2024)
Video-Language Critic: Transferable Reward Functions for Language-Conditioned Robotics
di: Alakuijala, Minttu, et al.
Pubblicazione: (2024)
di: Alakuijala, Minttu, et al.
Pubblicazione: (2024)
Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos
di: Chen, Yi, et al.
Pubblicazione: (2024)
di: Chen, Yi, et al.
Pubblicazione: (2024)
Teaching Embodied Reinforcement Learning Agents: Informativeness and Diversity of Language Use
di: Xi, Jiajun, et al.
Pubblicazione: (2024)
di: Xi, Jiajun, et al.
Pubblicazione: (2024)
DecisionNCE: Embodied Multimodal Representations via Implicit Preference Learning
di: Li, Jianxiong, et al.
Pubblicazione: (2024)
di: Li, Jianxiong, et al.
Pubblicazione: (2024)
SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding
di: Jia, Baoxiong, et al.
Pubblicazione: (2024)
di: Jia, Baoxiong, et al.
Pubblicazione: (2024)
LLaRA: Supercharging Robot Learning Data for Vision-Language Policy
di: Li, Xiang, et al.
Pubblicazione: (2024)
di: Li, Xiang, et al.
Pubblicazione: (2024)
ET tu, CLIP? Addressing Common Object Errors for Unseen Environments
di: Byun, Ye Won, et al.
Pubblicazione: (2024)
di: Byun, Ye Won, et al.
Pubblicazione: (2024)
Hierarchical Open-Vocabulary 3D Scene Graphs for Language-Grounded Robot Navigation
di: Werby, Abdelrhman, et al.
Pubblicazione: (2024)
di: Werby, Abdelrhman, et al.
Pubblicazione: (2024)
3D-GRAND: A Million-Scale Dataset for 3D-LLMs with Better Grounding and Less Hallucination
di: Yang, Jianing, et al.
Pubblicazione: (2024)
di: Yang, Jianing, et al.
Pubblicazione: (2024)
RT-Affordance: Affordances are Versatile Intermediate Representations for Robot Manipulation
di: Nasiriany, Soroush, et al.
Pubblicazione: (2024)
di: Nasiriany, Soroush, et al.
Pubblicazione: (2024)
LEGENT: Open Platform for Embodied Agents
di: Cheng, Zhili, et al.
Pubblicazione: (2024)
di: Cheng, Zhili, et al.
Pubblicazione: (2024)
Documenti analoghi
-
PerAct2: Benchmarking and Learning for Robotic Bimanual Manipulation Tasks
di: Grotz, Markus, et al.
Pubblicazione: (2024) -
SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts
di: Zhou, Gengze, et al.
Pubblicazione: (2024) -
CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
di: Li, Qixiu, et al.
Pubblicazione: (2024) -
GenSim: Generating Robotic Simulation Tasks via Large Language Models
di: Wang, Lirui, et al.
Pubblicazione: (2023) -
SlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video Generation
di: Hong, Yining, et al.
Pubblicazione: (2024)