VLM Agents Generate Their Own Memories: Distilling Experience into Embodied Programs of Thought
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sarch, Gabriel, Jang, Lawrence, Tarr, Michael J., Cohen, William W., Marino, Kenneth, Fragkiadaki, Katerina |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HELPER-X: A Unified Instructable Embodied Agent to Tackle Four Interactive Vision-Language Domains with Memory-Augmented Language Models
von: Sarch, Gabriel, et al.
Veröffentlicht: (2024)
von: Sarch, Gabriel, et al.
Veröffentlicht: (2024)
Grounded Reinforcement Learning for Visual Reasoning
von: Sarch, Gabriel, et al.
Veröffentlicht: (2025)
von: Sarch, Gabriel, et al.
Veröffentlicht: (2025)
ODIN: A Single Model for 2D and 3D Segmentation
von: Jain, Ayush, et al.
Veröffentlicht: (2024)
von: Jain, Ayush, et al.
Veröffentlicht: (2024)
DreamScene4D: Dynamic Multi-Object Scene Generation from Monocular Videos
von: Chu, Wen-Hsuan, et al.
Veröffentlicht: (2024)
von: Chu, Wen-Hsuan, et al.
Veröffentlicht: (2024)
Reanimating Images using Neural Representations of Dynamic Stimuli
von: Yeung, Jacob, et al.
Veröffentlicht: (2024)
von: Yeung, Jacob, et al.
Veröffentlicht: (2024)
TAPIP3D: Tracking Any Point in Persistent 3D Geometry
von: Zhang, Bowei, et al.
Veröffentlicht: (2025)
von: Zhang, Bowei, et al.
Veröffentlicht: (2025)
ADVEDM:Fine-grained Adversarial Attack against VLM-based Embodied Agents
von: Wang, Yichen, et al.
Veröffentlicht: (2025)
von: Wang, Yichen, et al.
Veröffentlicht: (2025)
Learning to Assist: Physics-Grounded Human-Human Control via Multi-Agent Reinforcement Learning
von: Shibata, Yuto, et al.
Veröffentlicht: (2026)
von: Shibata, Yuto, et al.
Veröffentlicht: (2026)
3D Diffuser Actor: Policy Diffusion with 3D Scene Representations
von: Ke, Tsung-Wei, et al.
Veröffentlicht: (2024)
von: Ke, Tsung-Wei, et al.
Veröffentlicht: (2024)
Dex4D: Task-Agnostic Point Track Policy for Sim-to-Real Dexterous Manipulation
von: Kuang, Yuxuan, et al.
Veröffentlicht: (2026)
von: Kuang, Yuxuan, et al.
Veröffentlicht: (2026)
Generative 4D Scene Gaussian Splatting with Object View-Synthesis Priors
von: Chu, Wen-Hsuan, et al.
Veröffentlicht: (2025)
von: Chu, Wen-Hsuan, et al.
Veröffentlicht: (2025)
GTR: Guided Thought Reinforcement Prevents Thought Collapse in RL-based VLM Agent Training
von: Wei, Tong, et al.
Veröffentlicht: (2025)
von: Wei, Tong, et al.
Veröffentlicht: (2025)
PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion Transformers
von: Lin, Yuchen, et al.
Veröffentlicht: (2025)
von: Lin, Yuchen, et al.
Veröffentlicht: (2025)
Aligning Text-to-Image Diffusion Models with Reward Backpropagation
von: Prabhudesai, Mihir, et al.
Veröffentlicht: (2023)
von: Prabhudesai, Mihir, et al.
Veröffentlicht: (2023)
Ella: Embodied Social Agents with Lifelong Memory
von: Zhang, Hongxin, et al.
Veröffentlicht: (2025)
von: Zhang, Hongxin, et al.
Veröffentlicht: (2025)
Zero-Shot Open-Vocabulary Tracking with Large Pre-Trained Models
von: Chu, Wen-Hsuan, et al.
Veröffentlicht: (2023)
von: Chu, Wen-Hsuan, et al.
Veröffentlicht: (2023)
Enhancing Video-LLM Reasoning via Agent-of-Thoughts Distillation
von: Shi, Yudi, et al.
Veröffentlicht: (2024)
von: Shi, Yudi, et al.
Veröffentlicht: (2024)
VLM-KD: Knowledge Distillation from VLM for Long-Tail Visual Recognition
von: Zhang, Zaiwei, et al.
Veröffentlicht: (2024)
von: Zhang, Zaiwei, et al.
Veröffentlicht: (2024)
EMemBench: Interactive Benchmarking of Episodic Memory for VLM Agents
von: Li, Xinze, et al.
Veröffentlicht: (2026)
von: Li, Xinze, et al.
Veröffentlicht: (2026)
FindingDory: A Benchmark to Evaluate Memory in Embodied Agents
von: Yadav, Karmesh, et al.
Veröffentlicht: (2025)
von: Yadav, Karmesh, et al.
Veröffentlicht: (2025)
RoboTAG: End-to-end Robot Configuration Estimation via Topological Alignment Graph
von: Liu, Yifan, et al.
Veröffentlicht: (2025)
von: Liu, Yifan, et al.
Veröffentlicht: (2025)
Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding
von: Fan, Yue, et al.
Veröffentlicht: (2024)
von: Fan, Yue, et al.
Veröffentlicht: (2024)
AtlasVA: Self-Evolving Visual Skill Memory for Teacher-Free VLM Agents
von: Wang, Pan, et al.
Veröffentlicht: (2026)
von: Wang, Pan, et al.
Veröffentlicht: (2026)
Vero: An Open RL Recipe for General Visual Reasoning
von: Sarch, Gabriel, et al.
Veröffentlicht: (2026)
von: Sarch, Gabriel, et al.
Veröffentlicht: (2026)
Grounding Task Assistance with Multimodal Cues from a Single Demonstration
von: Sarch, Gabriel, et al.
Veröffentlicht: (2025)
von: Sarch, Gabriel, et al.
Veröffentlicht: (2025)
Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames
von: Ravi, Sahithya, et al.
Veröffentlicht: (2025)
von: Ravi, Sahithya, et al.
Veröffentlicht: (2025)
Video Diffusion Alignment via Reward Gradients
von: Prabhudesai, Mihir, et al.
Veröffentlicht: (2024)
von: Prabhudesai, Mihir, et al.
Veröffentlicht: (2024)
Diffusion Beats Autoregressive in Data-Constrained Settings
von: Prabhudesai, Mihir, et al.
Veröffentlicht: (2025)
von: Prabhudesai, Mihir, et al.
Veröffentlicht: (2025)
Unified Multimodal Discrete Diffusion
von: Swerdlow, Alexander, et al.
Veröffentlicht: (2025)
von: Swerdlow, Alexander, et al.
Veröffentlicht: (2025)
BEAT: Visual Backdoor Attacks on VLM-based Embodied Agents via Contrastive Trigger Learning
von: Zhan, Qiusi, et al.
Veröffentlicht: (2025)
von: Zhan, Qiusi, et al.
Veröffentlicht: (2025)
Embodied Tree of Thoughts: Deliberate Manipulation Planning with Embodied World Model
von: Xu, Wenjiang, et al.
Veröffentlicht: (2025)
von: Xu, Wenjiang, et al.
Veröffentlicht: (2025)
RoboStressBench: Benchmarking VLM Robustness to Physical Visual Stress in Embodied Scenes
von: Wu, Leyi, et al.
Veröffentlicht: (2026)
von: Wu, Leyi, et al.
Veröffentlicht: (2026)
HomeGuard: VLM-based Embodied Safeguard for Identifying Contextual Risk in Household Task
von: Lu, Xiaoya, et al.
Veröffentlicht: (2026)
von: Lu, Xiaoya, et al.
Veröffentlicht: (2026)
Memo: Training Memory-Efficient Embodied Agents with Reinforcement Learning
von: Gupta, Gunshi, et al.
Veröffentlicht: (2025)
von: Gupta, Gunshi, et al.
Veröffentlicht: (2025)
IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household Tasks
von: Lu, Xiaoya, et al.
Veröffentlicht: (2025)
von: Lu, Xiaoya, et al.
Veröffentlicht: (2025)
GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation
von: Chen, Sixiang, et al.
Veröffentlicht: (2026)
von: Chen, Sixiang, et al.
Veröffentlicht: (2026)
Video Depth without Video Models
von: Ke, Bingxin, et al.
Veröffentlicht: (2024)
von: Ke, Bingxin, et al.
Veröffentlicht: (2024)
MoVieS: Motion-Aware 4D Dynamic View Synthesis in One Second
von: Lin, Chenguo, et al.
Veröffentlicht: (2025)
von: Lin, Chenguo, et al.
Veröffentlicht: (2025)
Do We Really Need a Complex Agent System? Distill Embodied Agent into a Single Model
von: Zhao, Zhonghan, et al.
Veröffentlicht: (2024)
von: Zhao, Zhonghan, et al.
Veröffentlicht: (2024)
MyVLM: Personalizing VLMs for User-Specific Queries
von: Alaluf, Yuval, et al.
Veröffentlicht: (2024)
von: Alaluf, Yuval, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
HELPER-X: A Unified Instructable Embodied Agent to Tackle Four Interactive Vision-Language Domains with Memory-Augmented Language Models
von: Sarch, Gabriel, et al.
Veröffentlicht: (2024) -
Grounded Reinforcement Learning for Visual Reasoning
von: Sarch, Gabriel, et al.
Veröffentlicht: (2025) -
ODIN: A Single Model for 2D and 3D Segmentation
von: Jain, Ayush, et al.
Veröffentlicht: (2024) -
DreamScene4D: Dynamic Multi-Object Scene Generation from Monocular Videos
von: Chu, Wen-Hsuan, et al.
Veröffentlicht: (2024) -
Reanimating Images using Neural Representations of Dynamic Stimuli
von: Yeung, Jacob, et al.
Veröffentlicht: (2024)