Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding
Fuente:
arXiv
Salvato in:
| Autori principali: | Vo, Khoa, Hanyu, Taisei, Ikebe, Yuki, Pham, Trong Thang, Chung, Nhat, Vu, Minh Nhat, Minh, Duy Nguyen Ho, Nguyen, Anh, Gunderman, Anthony, Rainwater, Chase, Le, Ngan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CodeGraphVLP: Code-as-Planner Meets Semantic-Graph State for Non-Markovian Vision-Language-Action Models
di: Vo, Khoa, et al.
Pubblicazione: (2026)
di: Vo, Khoa, et al.
Pubblicazione: (2026)
SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation
di: Hanyu, Taisei, et al.
Pubblicazione: (2025)
di: Hanyu, Taisei, et al.
Pubblicazione: (2025)
Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective
di: Chung, Nhat, et al.
Pubblicazione: (2025)
di: Chung, Nhat, et al.
Pubblicazione: (2025)
Robotic-CLIP: Fine-tuning CLIP on Action Data for Robotic Applications
di: Nguyen, Nghia, et al.
Pubblicazione: (2024)
di: Nguyen, Nghia, et al.
Pubblicazione: (2024)
Lightweight Language-driven Grasp Detection using Conditional Consistency Model
di: Nguyen, Nghia, et al.
Pubblicazione: (2024)
di: Nguyen, Nghia, et al.
Pubblicazione: (2024)
ReFineVLA: Reasoning-Aware Teacher-Guided Transfer Fine-Tuning
di: Van Vo, Tuan, et al.
Pubblicazione: (2025)
di: Van Vo, Tuan, et al.
Pubblicazione: (2025)
EgoMusic-driven Human Dance Motion Estimation with Skeleton Mamba
di: Nguyen, Quang, et al.
Pubblicazione: (2025)
di: Nguyen, Quang, et al.
Pubblicazione: (2025)
Online Trajectory Replanner for Dynamically Grasping Irregular Objects
di: Vu, Minh Nhat, et al.
Pubblicazione: (2025)
di: Vu, Minh Nhat, et al.
Pubblicazione: (2025)
Publicly Verifiable Secret Sharing: Generic Constructions and Lattice-Based Instantiations in the Standard Model
di: Minh, Pham Nhat, et al.
Pubblicazione: (2025)
di: Minh, Pham Nhat, et al.
Pubblicazione: (2025)
XFlowMP: Task-Conditioned Motion Fields for Generative Robot Planning with Schrodinger Bridges
di: Nguyen, Khang, et al.
Pubblicazione: (2025)
di: Nguyen, Khang, et al.
Pubblicazione: (2025)
Learning Human Motion with Temporally Conditional Mamba
di: Nguyen, Quang, et al.
Pubblicazione: (2025)
di: Nguyen, Quang, et al.
Pubblicazione: (2025)
Language-driven Grasp Detection with Mask-guided Attention
di: Van Vo, Tuan, et al.
Pubblicazione: (2024)
di: Van Vo, Tuan, et al.
Pubblicazione: (2024)
Post-Quantum Secure Decentralized Random Number Generation Protocol with Two Rounds of Communication in the Standard Model
di: Minh, Pham Nhat, et al.
Pubblicazione: (2025)
di: Minh, Pham Nhat, et al.
Pubblicazione: (2025)
Multi-tracers, multi-surveys: a joint Fisher analysis of DESI+PFS
di: Nguyen, Nhat-Minh
Pubblicazione: (2026)
di: Nguyen, Nhat-Minh
Pubblicazione: (2026)
Physics Is All You Need? A Case Study in Physicist-Supervised AI Development of Scientific Software
di: Nguyen, Nhat-Minh
Pubblicazione: (2026)
di: Nguyen, Nhat-Minh
Pubblicazione: (2026)
Language-Driven 6-DoF Grasp Detection Using Negative Prompt Guidance
di: Nguyen, Toan, et al.
Pubblicazione: (2024)
di: Nguyen, Toan, et al.
Pubblicazione: (2024)
DuFal: Dual-Frequency-Aware Learning for High-Fidelity Extremely Sparse-view CBCT Reconstruction
di: Van, Cuong Tran, et al.
Pubblicazione: (2026)
di: Van, Cuong Tran, et al.
Pubblicazione: (2026)
ReFineVLA: Multimodal Reasoning-Aware Generalist Robotic Policies via Teacher-Guided Fine-Tuning
di: Van Vo, Tuan, et al.
Pubblicazione: (2026)
di: Van Vo, Tuan, et al.
Pubblicazione: (2026)
Hierarchical Motion Planning and Offline Robust Model Predictive Control for Autonomous Vehicles
di: Nguyen, Hung Duy, et al.
Pubblicazione: (2024)
di: Nguyen, Hung Duy, et al.
Pubblicazione: (2024)
Determination of Shear Wave Velocity Using Multichannel Analysis of Surface Wave in M'Drak District, Dak Lak Province, Vietnam
di: Ngan Nhat Kim Nguyen, et al.
Pubblicazione: (2025)
di: Ngan Nhat Kim Nguyen, et al.
Pubblicazione: (2025)
Amodal Instance Segmentation with Diffusion Shape Prior Estimation
di: Tran, Minh, et al.
Pubblicazione: (2024)
di: Tran, Minh, et al.
Pubblicazione: (2024)
ShapeFormer: Shape Prior Visible-to-Amodal Transformer-based Amodal Instance Segmentation
di: Tran, Minh, et al.
Pubblicazione: (2024)
di: Tran, Minh, et al.
Pubblicazione: (2024)
MedSteer: Counterfactual Endoscopic Synthesis via Training-Free Activation Steering
di: Pham, Trong-Thang, et al.
Pubblicazione: (2026)
di: Pham, Trong-Thang, et al.
Pubblicazione: (2026)
GraspMamba: A Mamba-based Language-driven Grasp Detection Framework with Hierarchical Feature Learning
di: Nguyen, Huy Hoang, et al.
Pubblicazione: (2024)
di: Nguyen, Huy Hoang, et al.
Pubblicazione: (2024)
DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving
di: Vo, Hao, et al.
Pubblicazione: (2026)
di: Vo, Hao, et al.
Pubblicazione: (2026)
Language-driven Grasp Detection
di: Vuong, An Dinh, et al.
Pubblicazione: (2024)
di: Vuong, An Dinh, et al.
Pubblicazione: (2024)
GPU-Accelerated Motion Planning of an Underactuated Forestry Crane in Cluttered Environments
di: Vu, Minh Nhat, et al.
Pubblicazione: (2025)
di: Vu, Minh Nhat, et al.
Pubblicazione: (2025)
Do stock markets care about climate change: A public media perspective
di: Minh Nhat Nguyen, et al.
Pubblicazione: (2024)
di: Minh Nhat Nguyen, et al.
Pubblicazione: (2024)
GazeQwen: Lightweight Gaze-Conditioned LLM Modulation for Streaming Video Understanding
di: Pham, Trong Thang, et al.
Pubblicazione: (2026)
di: Pham, Trong Thang, et al.
Pubblicazione: (2026)
Improving Robotic Manipulation with Efficient Geometry-Aware Vision Encoder
di: Vuong, An Dinh, et al.
Pubblicazione: (2025)
di: Vuong, An Dinh, et al.
Pubblicazione: (2025)
Self-intersections of arcs on a pair of pants
di: Doan, Nhat Minh, et al.
Pubblicazione: (2024)
di: Doan, Nhat Minh, et al.
Pubblicazione: (2024)
Self‐intersections of arcs on a pair of pants
di: Nhat Minh Doan, et al.
Pubblicazione: (2025)
di: Nhat Minh Doan, et al.
Pubblicazione: (2025)
DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion
di: Nguyen, Khang, et al.
Pubblicazione: (2025)
di: Nguyen, Khang, et al.
Pubblicazione: (2025)
CT-ScanGaze: A Dataset and Baselines for 3D Volumetric Scanpath Modeling
di: Pham, Trong-Thang, et al.
Pubblicazione: (2025)
di: Pham, Trong-Thang, et al.
Pubblicazione: (2025)
Revisit Visual Prompt Tuning: The Expressiveness of Prompt Experts
di: Le, Minh, et al.
Pubblicazione: (2025)
di: Le, Minh, et al.
Pubblicazione: (2025)
Whisper based Cross-Lingual Phoneme Recognition between Vietnamese and English
di: Minh, Nguyen Huu Nhat, et al.
Pubblicazione: (2025)
di: Minh, Nguyen Huu Nhat, et al.
Pubblicazione: (2025)
Action Tokenizer Matters in In-Context Imitation Learning
di: Vuong, An Dinh, et al.
Pubblicazione: (2025)
di: Vuong, An Dinh, et al.
Pubblicazione: (2025)
Equivariant Polynomial Functional Networks
di: Vo, Thieu N., et al.
Pubblicazione: (2024)
di: Vo, Thieu N., et al.
Pubblicazione: (2024)
Equivariant Neural Functional Networks for Transformers
di: Tran, Viet-Hoang, et al.
Pubblicazione: (2024)
di: Tran, Viet-Hoang, et al.
Pubblicazione: (2024)
Synthesis of 1 H ‐pyrazole frameworks from chalcones using p ‐toluenesulfonic acid as an efficient catalyst
di: Nhat Minh Nguyen, et al.
Pubblicazione: (2025)
di: Nhat Minh Nguyen, et al.
Pubblicazione: (2025)
Documenti analoghi
-
CodeGraphVLP: Code-as-Planner Meets Semantic-Graph State for Non-Markovian Vision-Language-Action Models
di: Vo, Khoa, et al.
Pubblicazione: (2026) -
SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation
di: Hanyu, Taisei, et al.
Pubblicazione: (2025) -
Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective
di: Chung, Nhat, et al.
Pubblicazione: (2025) -
Robotic-CLIP: Fine-tuning CLIP on Action Data for Robotic Applications
di: Nguyen, Nghia, et al.
Pubblicazione: (2024) -
Lightweight Language-driven Grasp Detection using Conditional Consistency Model
di: Nguyen, Nghia, et al.
Pubblicazione: (2024)