Selective Visual Representations Improve Convergence and Generalization for Embodied AI
Fuente:
arXiv
Saved in:
| Main Authors: | Eftekhar, Ainaz, Zeng, Kuo-Hao, Duan, Jiafei, Farhadi, Ali, Kembhavi, Ani, Krishna, Ranjay |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The One RING: a Robotic Indoor Navigation Generalist
by: Eftekhar, Ainaz, et al.
Published: (2024)
by: Eftekhar, Ainaz, et al.
Published: (2024)
Generate Any Scene: Scene Graph Driven Data Synthesis for Visual Generation Training
by: Gao, Ziqi, et al.
Published: (2024)
by: Gao, Ziqi, et al.
Published: (2024)
Convergent Functions, Divergent Forms
by: Jeon, Hyeonseong, et al.
Published: (2025)
by: Jeon, Hyeonseong, et al.
Published: (2025)
Iterated Learning Improves Compositionality in Large Vision-Language Models
by: Zheng, Chenhao, et al.
Published: (2024)
by: Zheng, Chenhao, et al.
Published: (2024)
SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models
by: Ray, Arijit, et al.
Published: (2024)
by: Ray, Arijit, et al.
Published: (2024)
Task Me Anything
by: Zhang, Jieyu, et al.
Published: (2024)
by: Zhang, Jieyu, et al.
Published: (2024)
Holodeck: Language Guided Generation of 3D Embodied AI Environments
by: Yang, Yue, et al.
Published: (2023)
by: Yang, Yue, et al.
Published: (2023)
One Diffusion to Generate Them All
by: Le, Duong H., et al.
Published: (2024)
by: Le, Duong H., et al.
Published: (2024)
Eval3D: Interpretable and Fine-grained Evaluation for 3D Generation
by: Duggal, Shivam, et al.
Published: (2025)
by: Duggal, Shivam, et al.
Published: (2025)
RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
by: Yuan, Wentao, et al.
Published: (2024)
by: Yuan, Wentao, et al.
Published: (2024)
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing
by: Cheng, Long, et al.
Published: (2025)
by: Cheng, Long, et al.
Published: (2025)
VLS: Steering Pretrained Robot Policies via Vision-Language Models
by: Liu, Shuo, et al.
Published: (2026)
by: Liu, Shuo, et al.
Published: (2026)
MolmoSpaces: A Large-Scale Open Ecosystem for Robot Navigation and Manipulation
by: Kim, Yejin, et al.
Published: (2026)
by: Kim, Yejin, et al.
Published: (2026)
Seeing the Unseen: Visual Common Sense for Semantic Placement
by: Ramrakhya, Ram, et al.
Published: (2024)
by: Ramrakhya, Ram, et al.
Published: (2024)
SPARO: Selective Attention for Robust and Compositional Transformer Encodings for Vision
by: Vani, Ankit, et al.
Published: (2024)
by: Vani, Ankit, et al.
Published: (2024)
Contrastive Flow Matching
by: Stoica, George, et al.
Published: (2025)
by: Stoica, George, et al.
Published: (2025)
SPOC: Imitating Shortest Paths in Simulation Enables Effective Navigation and Manipulation in the Real World
by: Ehsani, Kiana, et al.
Published: (2023)
by: Ehsani, Kiana, et al.
Published: (2023)
FLaRe: Achieving Masterful and Adaptive Robot Policies with Large-Scale Reinforcement Learning Fine-Tuning
by: Hu, Jiaheng, et al.
Published: (2024)
by: Hu, Jiaheng, et al.
Published: (2024)
MosaicThinker: On-Device Visual Spatial Reasoning for Embodied AI via Iterative Construction of Space Representation
by: Wang, Haoming, et al.
Published: (2026)
by: Wang, Haoming, et al.
Published: (2026)
Videoshop: Localized Semantic Video Editing with Noise-Extrapolated Diffusion Inversion
by: Fan, Xiang, et al.
Published: (2024)
by: Fan, Xiang, et al.
Published: (2024)
Manipulate-Anything: Automating Real-World Robots using Vision-Language Models
by: Duan, Jiafei, et al.
Published: (2024)
by: Duan, Jiafei, et al.
Published: (2024)
MIMIC: Masked Image Modeling with Image Correspondences
by: Marathe, Kalyani, et al.
Published: (2023)
by: Marathe, Kalyani, et al.
Published: (2023)
Visual Representations inside the Language Model
by: Liu, Benlin, et al.
Published: (2025)
by: Liu, Benlin, et al.
Published: (2025)
VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models
by: Wu, Kui, et al.
Published: (2025)
by: Wu, Kui, et al.
Published: (2025)
MiraGe: Multimodal Discriminative Representation Learning for Generalizable AI-Generated Image Detection
by: Shi, Kuo, et al.
Published: (2025)
by: Shi, Kuo, et al.
Published: (2025)
BiGR: Harnessing Binary Latent Codes for Image Generation and Improved Visual Representation Capabilities
by: Hao, Shaozhe, et al.
Published: (2024)
by: Hao, Shaozhe, et al.
Published: (2024)
You Only Judge Once: Multi-response Reward Modeling in a Single Forward Pass
by: Yang, Yinuo, et al.
Published: (2026)
by: Yang, Yinuo, et al.
Published: (2026)
Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos
by: Seyfioglu, Mehmet Saygin, et al.
Published: (2023)
by: Seyfioglu, Mehmet Saygin, et al.
Published: (2023)
Empowering Embodied Visual Tracking with Visual Foundation Models and Offline RL
by: Zhong, Fangwei, et al.
Published: (2024)
by: Zhong, Fangwei, et al.
Published: (2024)
Synthetic Visual Genome
by: Park, Jae Sung, et al.
Published: (2025)
by: Park, Jae Sung, et al.
Published: (2025)
ActionAtlas: A VideoQA Benchmark for Domain-specialized Action Recognition
by: Salehi, Mohammadreza, et al.
Published: (2024)
by: Salehi, Mohammadreza, et al.
Published: (2024)
From Sight to Insight: Improving Visual Reasoning Capabilities of Multimodal Models via Reinforcement Learning
by: Sharif, Omar, et al.
Published: (2026)
by: Sharif, Omar, et al.
Published: (2026)
Counting Circuits: Mechanistic Interpretability of Visual Reasoning in Large Vision-Language Models
by: Che, Liwei, et al.
Published: (2026)
by: Che, Liwei, et al.
Published: (2026)
CoGR-MoE: Concept-Guided Expert Routing with Consistent Selection and Flexible Reasoning for Visual Question Answering
by: Zeng, Xiyin, et al.
Published: (2026)
by: Zeng, Xiyin, et al.
Published: (2026)
RoboEval: Where Robotic Manipulation Meets Structured and Scalable Evaluation
by: Wang, Yi Ru, et al.
Published: (2025)
by: Wang, Yi Ru, et al.
Published: (2025)
Improving Visual Representation Alignment Generation with GRPO
by: Mo, Shentong, et al.
Published: (2026)
by: Mo, Shentong, et al.
Published: (2026)
MetaUrban: An Embodied AI Simulation Platform for Urban Micromobility
by: Wu, Wayne, et al.
Published: (2024)
by: Wu, Wayne, et al.
Published: (2024)
Posterior Augmented Flow Matching
by: Stoica, George, et al.
Published: (2026)
by: Stoica, George, et al.
Published: (2026)
RefDecoder: Enhancing Visual Generation with Conditional Video Decoding
by: Fan, Xiang, et al.
Published: (2026)
by: Fan, Xiang, et al.
Published: (2026)
Perception Tokens Enhance Visual Reasoning in Multimodal Language Models
by: Bigverdi, Mahtab, et al.
Published: (2024)
by: Bigverdi, Mahtab, et al.
Published: (2024)
Similar Items
-
The One RING: a Robotic Indoor Navigation Generalist
by: Eftekhar, Ainaz, et al.
Published: (2024) -
Generate Any Scene: Scene Graph Driven Data Synthesis for Visual Generation Training
by: Gao, Ziqi, et al.
Published: (2024) -
Convergent Functions, Divergent Forms
by: Jeon, Hyeonseong, et al.
Published: (2025) -
Iterated Learning Improves Compositionality in Large Vision-Language Models
by: Zheng, Chenhao, et al.
Published: (2024) -
SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models
by: Ray, Arijit, et al.
Published: (2024)