VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Hao, Wei, Xiaobao, He, Jingyang, Bai, Chengyu, Fan, Chun-Kai, Cao, Jiajun, Chen, Jintao, Li, Ying, Rong, Shanyu, Lu, Ming, Ju, Xiaozhu, Tang, Jian, Zhang, Shanghang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
UniEdit-I: Training-free Image Editing for Unified VLM via Iterative Understanding, Editing and Verifying
di: Bai, Chengyu, et al.
Pubblicazione: (2025)
di: Bai, Chengyu, et al.
Pubblicazione: (2025)
SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics
di: Liu, Mengzhen, et al.
Pubblicazione: (2026)
di: Liu, Mengzhen, et al.
Pubblicazione: (2026)
MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders
di: Cao, Jiajun, et al.
Pubblicazione: (2025)
di: Cao, Jiajun, et al.
Pubblicazione: (2025)
TC-IDM: Grounding Video Generation for Executable Zero-shot Robot Motion
di: Mi, Weishi, et al.
Pubblicazione: (2026)
di: Mi, Weishi, et al.
Pubblicazione: (2026)
ManipDreamer: Boosting Robotic Manipulation World Model with Action Tree and Visual Guidance
di: Li, Ying, et al.
Pubblicazione: (2025)
di: Li, Ying, et al.
Pubblicazione: (2025)
EmbodiedOcc++: Boosting Embodied 3D Occupancy Prediction with Plane Regularization and Uncertainty Sampler
di: Wang, Hao, et al.
Pubblicazione: (2025)
di: Wang, Hao, et al.
Pubblicazione: (2025)
Action-Sketcher: From Reasoning to Action via Visual Sketches for Long-Horizon Robotic Manipulation
di: Tan, Huajie, et al.
Pubblicazione: (2026)
di: Tan, Huajie, et al.
Pubblicazione: (2026)
RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics
di: Zhou, Enshen, et al.
Pubblicazione: (2025)
di: Zhou, Enshen, et al.
Pubblicazione: (2025)
I-MedSAM: Implicit Medical Image Segmentation with Segment Anything
di: Wei, Xiaobao, et al.
Pubblicazione: (2023)
di: Wei, Xiaobao, et al.
Pubblicazione: (2023)
EgoActor: Grounding Task Planning into Spatial-aware Egocentric Actions for Humanoid Robots via Visual-Language Models
di: Bai, Yu, et al.
Pubblicazione: (2026)
di: Bai, Yu, et al.
Pubblicazione: (2026)
FastInit: Fast Noise Initialization for Temporally Consistent Video Generation
di: Bai, Chengyu, et al.
Pubblicazione: (2025)
di: Bai, Chengyu, et al.
Pubblicazione: (2025)
RoboTracer: Mastering Spatial Trace with Reasoning in Vision-Language Models for Robotics
di: Zhou, Enshen, et al.
Pubblicazione: (2025)
di: Zhou, Enshen, et al.
Pubblicazione: (2025)
Grounded Forcing: Bridging Time-Independent Semantics and Proximal Dynamics in Autoregressive Video Synthesis
di: Chen, Jintao, et al.
Pubblicazione: (2026)
di: Chen, Jintao, et al.
Pubblicazione: (2026)
TIGeR: Tool-Integrated Geometric Reasoning in Vision-Language Models for Robotics
di: Han, Yi, et al.
Pubblicazione: (2025)
di: Han, Yi, et al.
Pubblicazione: (2025)
Grounding Emotion Recognition with Visual Prototypes: VEGA -- Revisiting CLIP in MERC
di: Hu, Guanyu, et al.
Pubblicazione: (2025)
di: Hu, Guanyu, et al.
Pubblicazione: (2025)
From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors
di: Zhang, Zhengshen, et al.
Pubblicazione: (2025)
di: Zhang, Zhengshen, et al.
Pubblicazione: (2025)
RoboArmGS: High-Quality Robotic Arm Splatting via Bézier Curve Refinement
di: Wang, Hao, et al.
Pubblicazione: (2025)
di: Wang, Hao, et al.
Pubblicazione: (2025)
Layer-wise Alignment: Examining Safety Alignment Across Image Encoder Layers in Vision Language Models
di: Bachu, Saketh, et al.
Pubblicazione: (2024)
di: Bachu, Saketh, et al.
Pubblicazione: (2024)
ROCKET: Residual-Oriented Multi-Layer Alignment for Spatially-Aware Vision-Language-Action Models
di: Sun, Guoheng, et al.
Pubblicazione: (2026)
di: Sun, Guoheng, et al.
Pubblicazione: (2026)
StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and Compression
di: Chen, Yilong, et al.
Pubblicazione: (2025)
di: Chen, Yilong, et al.
Pubblicazione: (2025)
EPIC-Bench: A Perception-Centric Benchmark for Fine-Grained Embodied Visual Grounding in Vision-Language Models
di: Shan, Haozhe, et al.
Pubblicazione: (2026)
di: Shan, Haozhe, et al.
Pubblicazione: (2026)
Grounding Hierarchical Vision-Language-Action Models Through Explicit Language-Action Alignment
di: Wulff, Theodor, et al.
Pubblicazione: (2026)
di: Wulff, Theodor, et al.
Pubblicazione: (2026)
VLA-IAP: Training-Free Visual Token Pruning via Interaction Alignment for Vision-Language-Action Models
di: Cheng, Jintao, et al.
Pubblicazione: (2026)
di: Cheng, Jintao, et al.
Pubblicazione: (2026)
EmbodiedVSR: Dynamic Scene Graph-Guided Chain-of-Thought Reasoning for Visual Spatial Tasks
di: Zhang, Yi, et al.
Pubblicazione: (2025)
di: Zhang, Yi, et al.
Pubblicazione: (2025)
Naturally Supervised 3D Visual Grounding with Language-Regularized Concept Learners
di: Feng, Chun, et al.
Pubblicazione: (2024)
di: Feng, Chun, et al.
Pubblicazione: (2024)
Prototype-Aware Multimodal Alignment for Open-Vocabulary Visual Grounding
di: Xie, Jiangnan, et al.
Pubblicazione: (2025)
di: Xie, Jiangnan, et al.
Pubblicazione: (2025)
Beyond the Vision Encoder: Identifying and Mitigating Spatial Bias in Large Vision-Language Models
di: Zhu, Yingjie, et al.
Pubblicazione: (2025)
di: Zhu, Yingjie, et al.
Pubblicazione: (2025)
Dual Attribute-Spatial Relation Alignment for 3D Visual Grounding
di: Xu, Yue, et al.
Pubblicazione: (2024)
di: Xu, Yue, et al.
Pubblicazione: (2024)
ConceptWeaver: Weaving Disentangled Concepts with Flow
di: Chen, Jintao, et al.
Pubblicazione: (2026)
di: Chen, Jintao, et al.
Pubblicazione: (2026)
Grounded 3D-Aware Spatial Vision-Language Modeling
di: Cheng, An-Chieh, et al.
Pubblicazione: (2026)
di: Cheng, An-Chieh, et al.
Pubblicazione: (2026)
Reshaping Action Error Distributions for Reliable Vision-Language-Action Models
di: Bai, Shuanghao, et al.
Pubblicazione: (2026)
di: Bai, Shuanghao, et al.
Pubblicazione: (2026)
SA-VLA: Spatially-Aware Flow-Matching for Vision-Language-Action Reinforcement Learning
di: Pan, Xu, et al.
Pubblicazione: (2026)
di: Pan, Xu, et al.
Pubblicazione: (2026)
WristWorld: Generating Wrist-Views via 4D World Models for Robotic Manipulation
di: Qian, Zezhong, et al.
Pubblicazione: (2025)
di: Qian, Zezhong, et al.
Pubblicazione: (2025)
Equilibrium in Style: A Modeling Framework on the Cash Flow and the Life Cycle of a Consumer Store
di: Han, Shanyu, et al.
Pubblicazione: (2024)
di: Han, Shanyu, et al.
Pubblicazione: (2024)
Bootstrapping Action-Grounded Visual Dynamics in Unified Vision-Language Models
di: Qiu, Yifu, et al.
Pubblicazione: (2025)
di: Qiu, Yifu, et al.
Pubblicazione: (2025)
MC-LLaVA: Multi-Concept Personalized Vision-Language Model
di: An, Ruichuan, et al.
Pubblicazione: (2025)
di: An, Ruichuan, et al.
Pubblicazione: (2025)
SGANet: Semantic and Geometric Alignment for Multimodal Multi-view Anomaly Detection
di: Bai, Letian, et al.
Pubblicazione: (2026)
di: Bai, Letian, et al.
Pubblicazione: (2026)
Grounding-Aware Token Pruning: Recovering from Drastic Performance Drops in Visual Grounding Caused by Pruning
di: Chien, Tzu-Chun, et al.
Pubblicazione: (2025)
di: Chien, Tzu-Chun, et al.
Pubblicazione: (2025)
Improving Concept Alignment in Vision-Language Concept Bottleneck Models
di: Selvaraj, Nithish Muthuchamy, et al.
Pubblicazione: (2024)
di: Selvaraj, Nithish Muthuchamy, et al.
Pubblicazione: (2024)
TAG: Thinking with Action Unit Grounding for Facial Expression Recognition
di: Lin, Haobo, et al.
Pubblicazione: (2026)
di: Lin, Haobo, et al.
Pubblicazione: (2026)
Documenti analoghi
-
UniEdit-I: Training-free Image Editing for Unified VLM via Iterative Understanding, Editing and Verifying
di: Bai, Chengyu, et al.
Pubblicazione: (2025) -
SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics
di: Liu, Mengzhen, et al.
Pubblicazione: (2026) -
MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders
di: Cao, Jiajun, et al.
Pubblicazione: (2025) -
TC-IDM: Grounding Video Generation for Executable Zero-shot Robot Motion
di: Mi, Weishi, et al.
Pubblicazione: (2026) -
ManipDreamer: Boosting Robotic Manipulation World Model with Action Tree and Visual Guidance
di: Li, Ying, et al.
Pubblicazione: (2025)