3D CAVLA: Leveraging Depth and 3D Context to Generalize Vision Language Action Models for Unseen Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Bhat, Vineet, Lan, Yu-Hsiang, Krishnamurthy, Prashanth, Karri, Ramesh, Khorrami, Farshad |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HiFi-CS: Towards Open Vocabulary Visual Grounding For Robotic Grasping Using Vision-Language Models
by: Bhat, Vineet, et al.
Published: (2024)
by: Bhat, Vineet, et al.
Published: (2024)
MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping
by: Bhat, Vineet, et al.
Published: (2025)
by: Bhat, Vineet, et al.
Published: (2025)
Grounding Large Language Models for Robot Task Planning Using Closed‐Loop State Feedback
by: Vineet Bhat, et al.
Published: (2025)
by: Vineet Bhat, et al.
Published: (2025)
Grounding LLMs For Robot Task Planning Using Closed-loop State Feedback
by: Bhat, Vineet, et al.
Published: (2024)
by: Bhat, Vineet, et al.
Published: (2024)
BOP-ASK: Object-Interaction Reasoning for Vision-Language Models
by: Bhat, Vineet, et al.
Published: (2025)
by: Bhat, Vineet, et al.
Published: (2025)
RAZER: Robust Accelerated Zero-Shot 3D Open-Vocabulary Panoptic Reconstruction with Spatio-Temporal Aggregation
by: Patel, Naman, et al.
Published: (2025)
by: Patel, Naman, et al.
Published: (2025)
Efficient and Distributed Large-Scale 3D Map Registration using Tomographic Features
by: Unlu, Halil Utku, et al.
Published: (2024)
by: Unlu, Halil Utku, et al.
Published: (2024)
OSVI-WM: One-Shot Visual Imitation for Unseen Tasks using World-Model-Guided Trajectory Generation
by: Goswami, Raktim Gautam, et al.
Published: (2025)
by: Goswami, Raktim Gautam, et al.
Published: (2025)
RoboPEPP: Vision-Based Robot Pose and Joint Angle Estimation through Embedding Predictive Pre-Training
by: Goswami, Raktim Gautam, et al.
Published: (2024)
by: Goswami, Raktim Gautam, et al.
Published: (2024)
Enabling Deep Visibility into VxWorks-Based Embedded Controllers in Cyber-Physical Systems for Anomaly Detection
by: Krishnamurthy, Prashanth, et al.
Published: (2025)
by: Krishnamurthy, Prashanth, et al.
Published: (2025)
SCAMPER -- Synchrophasor Covert chAnnel for Malicious and Protective ERrands
by: Krishnamurthy, Prashanth, et al.
Published: (2025)
by: Krishnamurthy, Prashanth, et al.
Published: (2025)
Real-Time Multi-Modal Subcomponent-Level Measurements for Trustworthy System Monitoring and Malware Detection
by: Khorrami, Farshad, et al.
Published: (2025)
by: Khorrami, Farshad, et al.
Published: (2025)
Safe Multi-Robotic Arm Interaction via 3D Convex Shapes
by: Kaypak, Ali Umut, et al.
Published: (2025)
by: Kaypak, Ali Umut, et al.
Published: (2025)
SALSA: Swift Adaptive Lightweight Self-Attention for Enhanced LiDAR Place Recognition
by: Goswami, Raktim Gautam, et al.
Published: (2024)
by: Goswami, Raktim Gautam, et al.
Published: (2024)
Sailing Through Point Clouds: Safe Navigation Using Point Cloud Based Control Barrier Functions
by: Dai, Bolun, et al.
Published: (2024)
by: Dai, Bolun, et al.
Published: (2024)
Differentiable Optimization Based Time-Varying Control Barrier Functions for Dynamic Obstacle Avoidance
by: Dai, Bolun, et al.
Published: (2023)
by: Dai, Bolun, et al.
Published: (2023)
RESCORE: LLM-Driven Simulation Recovery in Control Systems Research Papers
by: Bhat, Vineet, et al.
Published: (2026)
by: Bhat, Vineet, et al.
Published: (2026)
Tracking Real-time Anomalies in Cyber-Physical Systems Through Dynamic Behavioral Analysis
by: Krishnamurthy, Prashanth, et al.
Published: (2024)
by: Krishnamurthy, Prashanth, et al.
Published: (2024)
Proactive Hierarchical Control Barrier Function-Based Safety Prioritization in Close Human-Robot Interaction Scenarios
by: Maithani, Patanjali, et al.
Published: (2025)
by: Maithani, Patanjali, et al.
Published: (2025)
SENTAUR: Security EnhaNced Trojan Assessment Using LLMs Against Undesirable Revisions
by: Bhandari, Jitendra, et al.
Published: (2024)
by: Bhandari, Jitendra, et al.
Published: (2024)
REMaQE: Reverse Engineering Math Equations from Executables
by: Udeshi, Meet, et al.
Published: (2023)
by: Udeshi, Meet, et al.
Published: (2023)
Distributed Inverse Dynamics Control for Quadruped Robots using Geometric Optimization
by: Khandelwal, Nimesh, et al.
Published: (2024)
by: Khandelwal, Nimesh, et al.
Published: (2024)
A Differentiable Distance Metric for Robotics Through Generalized Alternating Projection
by: Gonçalves, Vinicius M., et al.
Published: (2025)
by: Gonçalves, Vinicius M., et al.
Published: (2025)
Compliant Control of Quadruped Robots for Assistive Load Carrying
by: Khandelwal, Nimesh, et al.
Published: (2025)
by: Khandelwal, Nimesh, et al.
Published: (2025)
CLIPScope: Enhancing Zero-Shot OOD Detection with Bayesian Scoring
by: Fu, Hao, et al.
Published: (2024)
by: Fu, Hao, et al.
Published: (2024)
SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding
by: Ghazanfari, Sara, et al.
Published: (2026)
by: Ghazanfari, Sara, et al.
Published: (2026)
MultiTalk: Introspective and Extrospective Dialogue for Human-Environment-LLM Alignment
by: Devarakonda, Venkata Naren, et al.
Published: (2024)
by: Devarakonda, Venkata Naren, et al.
Published: (2024)
3D-VLA: A 3D Vision-Language-Action Generative World Model
by: Zhen, Haoyu, et al.
Published: (2024)
by: Zhen, Haoyu, et al.
Published: (2024)
GST-VLA: Structured Gaussian Spatial Tokens for 3D Depth-Aware Vision-Language-Action Models
by: Sarowar, Md Selim, et al.
Published: (2026)
by: Sarowar, Md Selim, et al.
Published: (2026)
FlashMix: Fast Map-Free LiDAR Localization via Feature Mixing and Contrastive-Constrained Accelerated Training
by: Goswami, Raktim Gautam, et al.
Published: (2024)
by: Goswami, Raktim Gautam, et al.
Published: (2024)
Ransomware 3.0: Self-Composing and LLM-Orchestrated
by: Raz, Md, et al.
Published: (2025)
by: Raz, Md, et al.
Published: (2025)
World Models for Learning Dexterous Hand-Object Interactions from Human Videos
by: Goswami, Raktim Gautam, et al.
Published: (2025)
by: Goswami, Raktim Gautam, et al.
Published: (2025)
SaMOSA: Sandbox for Malware Orchestration and Side-Channel Analysis
by: Udeshi, Meet, et al.
Published: (2025)
by: Udeshi, Meet, et al.
Published: (2025)
SHIELD: A Host-Independent Framework for Ransomware Detection using Deep Filesystem Features
by: Raz, Md, et al.
Published: (2025)
by: Raz, Md, et al.
Published: (2025)
3DVLA: Enhancing Vision-Language-Action Models via 3D Spatial and Instance Understanding
by: Xia, Zhongyu, et al.
Published: (2026)
by: Xia, Zhongyu, et al.
Published: (2026)
Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model
by: Lin, Tao, et al.
Published: (2026)
by: Lin, Tao, et al.
Published: (2026)
OG-VLA: Orthographic Image Generation for 3D-Aware Vision-Language Action Model
by: Singh, Ishika, et al.
Published: (2025)
by: Singh, Ishika, et al.
Published: (2025)
NextBestPath: Efficient 3D Mapping of Unseen Environments
by: Li, Shiyao, et al.
Published: (2025)
by: Li, Shiyao, et al.
Published: (2025)
PointVLA: Injecting the 3D World into Vision-Language-Action Models
by: Li, Chengmeng, et al.
Published: (2025)
by: Li, Chengmeng, et al.
Published: (2025)
Open-Architecture End-to-End System for Real-World Autonomous Robot Navigation
by: Devarakonda, Venkata Naren, et al.
Published: (2024)
by: Devarakonda, Venkata Naren, et al.
Published: (2024)
Similar Items
-
HiFi-CS: Towards Open Vocabulary Visual Grounding For Robotic Grasping Using Vision-Language Models
by: Bhat, Vineet, et al.
Published: (2024) -
MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping
by: Bhat, Vineet, et al.
Published: (2025) -
Grounding Large Language Models for Robot Task Planning Using Closed‐Loop State Feedback
by: Vineet Bhat, et al.
Published: (2025) -
Grounding LLMs For Robot Task Planning Using Closed-loop State Feedback
by: Bhat, Vineet, et al.
Published: (2024) -
BOP-ASK: Object-Interaction Reasoning for Vision-Language Models
by: Bhat, Vineet, et al.
Published: (2025)