Guardado en:
| Autores principales: | Sarch, Gabriel, Jang, Lawrence, Tarr, Michael J., Cohen, William W., Marino, Kenneth, Fragkiadaki, Katerina |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2406.14596 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
HELPER-X: A Unified Instructable Embodied Agent to Tackle Four Interactive Vision-Language Domains with Memory-Augmented Language Models
por: Sarch, Gabriel, et al.
Publicado: (2024)
por: Sarch, Gabriel, et al.
Publicado: (2024)
Grounded Reinforcement Learning for Visual Reasoning
por: Sarch, Gabriel, et al.
Publicado: (2025)
por: Sarch, Gabriel, et al.
Publicado: (2025)
ODIN: A Single Model for 2D and 3D Segmentation
por: Jain, Ayush, et al.
Publicado: (2024)
por: Jain, Ayush, et al.
Publicado: (2024)
DreamScene4D: Dynamic Multi-Object Scene Generation from Monocular Videos
por: Chu, Wen-Hsuan, et al.
Publicado: (2024)
por: Chu, Wen-Hsuan, et al.
Publicado: (2024)
Reanimating Images using Neural Representations of Dynamic Stimuli
por: Yeung, Jacob, et al.
Publicado: (2024)
por: Yeung, Jacob, et al.
Publicado: (2024)
TAPIP3D: Tracking Any Point in Persistent 3D Geometry
por: Zhang, Bowei, et al.
Publicado: (2025)
por: Zhang, Bowei, et al.
Publicado: (2025)
3D Diffuser Actor: Policy Diffusion with 3D Scene Representations
por: Ke, Tsung-Wei, et al.
Publicado: (2024)
por: Ke, Tsung-Wei, et al.
Publicado: (2024)
Learning to Assist: Physics-Grounded Human-Human Control via Multi-Agent Reinforcement Learning
por: Shibata, Yuto, et al.
Publicado: (2026)
por: Shibata, Yuto, et al.
Publicado: (2026)
Dex4D: Task-Agnostic Point Track Policy for Sim-to-Real Dexterous Manipulation
por: Kuang, Yuxuan, et al.
Publicado: (2026)
por: Kuang, Yuxuan, et al.
Publicado: (2026)
Generative 4D Scene Gaussian Splatting with Object View-Synthesis Priors
por: Chu, Wen-Hsuan, et al.
Publicado: (2025)
por: Chu, Wen-Hsuan, et al.
Publicado: (2025)
Aligning Text-to-Image Diffusion Models with Reward Backpropagation
por: Prabhudesai, Mihir, et al.
Publicado: (2023)
por: Prabhudesai, Mihir, et al.
Publicado: (2023)
ADVEDM:Fine-grained Adversarial Attack against VLM-based Embodied Agents
por: Wang, Yichen, et al.
Publicado: (2025)
por: Wang, Yichen, et al.
Publicado: (2025)
PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion Transformers
por: Lin, Yuchen, et al.
Publicado: (2025)
por: Lin, Yuchen, et al.
Publicado: (2025)
GTR: Guided Thought Reinforcement Prevents Thought Collapse in RL-based VLM Agent Training
por: Wei, Tong, et al.
Publicado: (2025)
por: Wei, Tong, et al.
Publicado: (2025)
Zero-Shot Open-Vocabulary Tracking with Large Pre-Trained Models
por: Chu, Wen-Hsuan, et al.
Publicado: (2023)
por: Chu, Wen-Hsuan, et al.
Publicado: (2023)
Vero: An Open RL Recipe for General Visual Reasoning
por: Sarch, Gabriel, et al.
Publicado: (2026)
por: Sarch, Gabriel, et al.
Publicado: (2026)
RoboTAG: End-to-end Robot Configuration Estimation via Topological Alignment Graph
por: Liu, Yifan, et al.
Publicado: (2025)
por: Liu, Yifan, et al.
Publicado: (2025)
Video Diffusion Alignment via Reward Gradients
por: Prabhudesai, Mihir, et al.
Publicado: (2024)
por: Prabhudesai, Mihir, et al.
Publicado: (2024)
Diffusion Beats Autoregressive in Data-Constrained Settings
por: Prabhudesai, Mihir, et al.
Publicado: (2025)
por: Prabhudesai, Mihir, et al.
Publicado: (2025)
Unified Multimodal Discrete Diffusion
por: Swerdlow, Alexander, et al.
Publicado: (2025)
por: Swerdlow, Alexander, et al.
Publicado: (2025)
Grounding Task Assistance with Multimodal Cues from a Single Demonstration
por: Sarch, Gabriel, et al.
Publicado: (2025)
por: Sarch, Gabriel, et al.
Publicado: (2025)
Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames
por: Ravi, Sahithya, et al.
Publicado: (2025)
por: Ravi, Sahithya, et al.
Publicado: (2025)
Ella: Embodied Social Agents with Lifelong Memory
por: Zhang, Hongxin, et al.
Publicado: (2025)
por: Zhang, Hongxin, et al.
Publicado: (2025)
Enhancing Video-LLM Reasoning via Agent-of-Thoughts Distillation
por: Shi, Yudi, et al.
Publicado: (2024)
por: Shi, Yudi, et al.
Publicado: (2024)
EMemBench: Interactive Benchmarking of Episodic Memory for VLM Agents
por: Li, Xinze, et al.
Publicado: (2026)
por: Li, Xinze, et al.
Publicado: (2026)
VLM-KD: Knowledge Distillation from VLM for Long-Tail Visual Recognition
por: Zhang, Zaiwei, et al.
Publicado: (2024)
por: Zhang, Zaiwei, et al.
Publicado: (2024)
FindingDory: A Benchmark to Evaluate Memory in Embodied Agents
por: Yadav, Karmesh, et al.
Publicado: (2025)
por: Yadav, Karmesh, et al.
Publicado: (2025)
Iterative Refinement Improves Compositional Image Generation
por: Jaiswal, Shantanu, et al.
Publicado: (2026)
por: Jaiswal, Shantanu, et al.
Publicado: (2026)
Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding
por: Fan, Yue, et al.
Publicado: (2024)
por: Fan, Yue, et al.
Publicado: (2024)
AtlasVA: Self-Evolving Visual Skill Memory for Teacher-Free VLM Agents
por: Wang, Pan, et al.
Publicado: (2026)
por: Wang, Pan, et al.
Publicado: (2026)
BEAT: Visual Backdoor Attacks on VLM-based Embodied Agents via Contrastive Trigger Learning
por: Zhan, Qiusi, et al.
Publicado: (2025)
por: Zhan, Qiusi, et al.
Publicado: (2025)
Video Depth without Video Models
por: Ke, Bingxin, et al.
Publicado: (2024)
por: Ke, Bingxin, et al.
Publicado: (2024)
MoVieS: Motion-Aware 4D Dynamic View Synthesis in One Second
por: Lin, Chenguo, et al.
Publicado: (2025)
por: Lin, Chenguo, et al.
Publicado: (2025)
Embodied Tree of Thoughts: Deliberate Manipulation Planning with Embodied World Model
por: Xu, Wenjiang, et al.
Publicado: (2025)
por: Xu, Wenjiang, et al.
Publicado: (2025)
IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household Tasks
por: Lu, Xiaoya, et al.
Publicado: (2025)
por: Lu, Xiaoya, et al.
Publicado: (2025)
Memo: Training Memory-Efficient Embodied Agents with Reinforcement Learning
por: Gupta, Gunshi, et al.
Publicado: (2025)
por: Gupta, Gunshi, et al.
Publicado: (2025)
Energy-based Models are Zero-Shot Planners for Compositional Scene Rearrangement
por: Gkanatsios, Nikolaos, et al.
Publicado: (2023)
por: Gkanatsios, Nikolaos, et al.
Publicado: (2023)
TimeWarp: Evaluating Web Agents by Revisiting the Past
por: Ishmam, Md Farhan, et al.
Publicado: (2026)
por: Ishmam, Md Farhan, et al.
Publicado: (2026)
RoboStressBench: Benchmarking VLM Robustness to Physical Visual Stress in Embodied Scenes
por: Wu, Leyi, et al.
Publicado: (2026)
por: Wu, Leyi, et al.
Publicado: (2026)
HomeGuard: VLM-based Embodied Safeguard for Identifying Contextual Risk in Household Task
por: Lu, Xiaoya, et al.
Publicado: (2026)
por: Lu, Xiaoya, et al.
Publicado: (2026)
Ejemplares similares
-
HELPER-X: A Unified Instructable Embodied Agent to Tackle Four Interactive Vision-Language Domains with Memory-Augmented Language Models
por: Sarch, Gabriel, et al.
Publicado: (2024) -
Grounded Reinforcement Learning for Visual Reasoning
por: Sarch, Gabriel, et al.
Publicado: (2025) -
ODIN: A Single Model for 2D and 3D Segmentation
por: Jain, Ayush, et al.
Publicado: (2024) -
DreamScene4D: Dynamic Multi-Object Scene Generation from Monocular Videos
por: Chu, Wen-Hsuan, et al.
Publicado: (2024) -
Reanimating Images using Neural Representations of Dynamic Stimuli
por: Yeung, Jacob, et al.
Publicado: (2024)