Lightweight Language-driven Grasp Detection using Conditional Consistency Model
Fuente:
arXiv
Saved in:
| Main Authors: | Nguyen, Nghia, Vu, Minh Nhat, Huang, Baoru, Vuong, An, Le, Ngan, Vo, Thieu, Nguyen, Anh |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Language-driven Grasp Detection with Mask-guided Attention
by: Van Vo, Tuan, et al.
Published: (2024)
by: Van Vo, Tuan, et al.
Published: (2024)
Language-Driven 6-DoF Grasp Detection Using Negative Prompt Guidance
by: Nguyen, Toan, et al.
Published: (2024)
by: Nguyen, Toan, et al.
Published: (2024)
Language-driven Grasp Detection
by: Vuong, An Dinh, et al.
Published: (2024)
by: Vuong, An Dinh, et al.
Published: (2024)
Robotic-CLIP: Fine-tuning CLIP on Action Data for Robotic Applications
by: Nguyen, Nghia, et al.
Published: (2024)
by: Nguyen, Nghia, et al.
Published: (2024)
GraspMamba: A Mamba-based Language-driven Grasp Detection Framework with Hierarchical Feature Learning
by: Nguyen, Huy Hoang, et al.
Published: (2024)
by: Nguyen, Huy Hoang, et al.
Published: (2024)
Learning Human Motion with Temporally Conditional Mamba
by: Nguyen, Quang, et al.
Published: (2025)
by: Nguyen, Quang, et al.
Published: (2025)
GraspMAS: Zero-Shot Language-driven Grasp Detection with Multi-Agent System
by: Nguyen, Quang, et al.
Published: (2025)
by: Nguyen, Quang, et al.
Published: (2025)
EgoMusic-driven Human Dance Motion Estimation with Skeleton Mamba
by: Nguyen, Quang, et al.
Published: (2025)
by: Nguyen, Quang, et al.
Published: (2025)
Autonomous Catheterization with Open-source Simulator and Expert Trajectory
by: Jianu, Tudor, et al.
Published: (2024)
by: Jianu, Tudor, et al.
Published: (2024)
HabiCrowd: A High Performance Simulator for Crowd-Aware Visual Navigation
by: Vuong, An Dinh, et al.
Published: (2023)
by: Vuong, An Dinh, et al.
Published: (2023)
ReFineVLA: Multimodal Reasoning-Aware Generalist Robotic Policies via Teacher-Guided Fine-Tuning
by: Van Vo, Tuan, et al.
Published: (2026)
by: Van Vo, Tuan, et al.
Published: (2026)
SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation
by: Hanyu, Taisei, et al.
Published: (2025)
by: Hanyu, Taisei, et al.
Published: (2025)
Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective
by: Chung, Nhat, et al.
Published: (2025)
by: Chung, Nhat, et al.
Published: (2025)
Online Trajectory Replanner for Dynamically Grasping Irregular Objects
by: Vu, Minh Nhat, et al.
Published: (2025)
by: Vu, Minh Nhat, et al.
Published: (2025)
Language-Driven Closed-Loop Grasping with Model-Predictive Trajectory Replanning
by: Nguyen, Huy Hoang, et al.
Published: (2024)
by: Nguyen, Huy Hoang, et al.
Published: (2024)
Lang2Lift: A Language-Guided Autonomous Forklift System for Outdoor Industrial Pallet Handling
by: Nguyen, Huy Hoang, et al.
Published: (2025)
by: Nguyen, Huy Hoang, et al.
Published: (2025)
Planning Robot Placement for Object Grasping
by: Saini, Manish, et al.
Published: (2024)
by: Saini, Manish, et al.
Published: (2024)
Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding
by: Vo, Khoa, et al.
Published: (2025)
by: Vo, Khoa, et al.
Published: (2025)
XFlowMP: Task-Conditioned Motion Fields for Generative Robot Planning with Schrodinger Bridges
by: Nguyen, Khang, et al.
Published: (2025)
by: Nguyen, Khang, et al.
Published: (2025)
RoboDesign1M: A Large-scale Dataset for Robot Design Understanding
by: Le, Tri, et al.
Published: (2025)
by: Le, Tri, et al.
Published: (2025)
ShapeFormer: Shape Prior Visible-to-Amodal Transformer-based Amodal Instance Segmentation
by: Tran, Minh, et al.
Published: (2024)
by: Tran, Minh, et al.
Published: (2024)
AffordMatcher: Affordance Learning in 3D Scenes from Visual Signifiers
by: Vu, Nghia, et al.
Published: (2026)
by: Vu, Nghia, et al.
Published: (2026)
Amodal Instance Segmentation with Diffusion Shape Prior Estimation
by: Tran, Minh, et al.
Published: (2024)
by: Tran, Minh, et al.
Published: (2024)
SIGMA: A Physics-Based Benchmark for Gas Chimney Understanding in Seismic Images
by: Truong, Bao, et al.
Published: (2026)
by: Truong, Bao, et al.
Published: (2026)
More Reliable Pseudo-labels, Better Performance: A Generalized Approach to Single Positive Multi-label Learning
by: Tran, Luong, et al.
Published: (2025)
by: Tran, Luong, et al.
Published: (2025)
FedEFM: Federated Endovascular Foundation Model with Unseen Data
by: Do, Tuong, et al.
Published: (2025)
by: Do, Tuong, et al.
Published: (2025)
GazeQwen: Lightweight Gaze-Conditioned LLM Modulation for Streaming Video Understanding
by: Pham, Trong Thang, et al.
Published: (2026)
by: Pham, Trong Thang, et al.
Published: (2026)
DextER: Language-driven Dexterous Grasp Generation with Embodied Reasoning
by: Lee, Junha, et al.
Published: (2026)
by: Lee, Junha, et al.
Published: (2026)
ReFineVLA: Reasoning-Aware Teacher-Guided Transfer Fine-Tuning
by: Van Vo, Tuan, et al.
Published: (2025)
by: Van Vo, Tuan, et al.
Published: (2025)
Revisit Visual Prompt Tuning: The Expressiveness of Prompt Experts
by: Le, Minh, et al.
Published: (2025)
by: Le, Minh, et al.
Published: (2025)
A Distributed Multi-Modal Sensing Approach for Human Activity Recognition in Real-Time Human-Robot Collaboration
by: Belcamino, Valerio, et al.
Published: (2026)
by: Belcamino, Valerio, et al.
Published: (2026)
SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object Detection
by: Vo, Hao, et al.
Published: (2026)
by: Vo, Hao, et al.
Published: (2026)
Graspness Discovery in Clutters for Fast and Accurate Grasp Detection
by: Wang, Chenxi, et al.
Published: (2024)
by: Wang, Chenxi, et al.
Published: (2024)
BiMa: Towards Biases Mitigation for Text-Video Retrieval via Scene Element Guidance
by: Le, Huy, et al.
Published: (2025)
by: Le, Huy, et al.
Published: (2025)
WAVER: Writing-style Agnostic Text-Video Retrieval via Distilling Vision-Language Models Through Open-Vocabulary Knowledge
by: Le, Huy, et al.
Published: (2023)
by: Le, Huy, et al.
Published: (2023)
DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion
by: Nguyen, Khang, et al.
Published: (2025)
by: Nguyen, Khang, et al.
Published: (2025)
Your Vision-Language-Action Model Already Has Attention Heads For Path Deviation Detection
by: Jeong, Jaehwan, et al.
Published: (2026)
by: Jeong, Jaehwan, et al.
Published: (2026)
Improving Robotic Manipulation with Efficient Geometry-Aware Vision Encoder
by: Vuong, An Dinh, et al.
Published: (2025)
by: Vuong, An Dinh, et al.
Published: (2025)
CodeGraphVLP: Code-as-Planner Meets Semantic-Graph State for Non-Markovian Vision-Language-Action Models
by: Vo, Khoa, et al.
Published: (2026)
by: Vo, Khoa, et al.
Published: (2026)
Language-Guided Grasp Detection with Coarse-to-Fine Learning for Robotic Manipulation
by: Jiang, Zebin, et al.
Published: (2025)
by: Jiang, Zebin, et al.
Published: (2025)
Similar Items
-
Language-driven Grasp Detection with Mask-guided Attention
by: Van Vo, Tuan, et al.
Published: (2024) -
Language-Driven 6-DoF Grasp Detection Using Negative Prompt Guidance
by: Nguyen, Toan, et al.
Published: (2024) -
Language-driven Grasp Detection
by: Vuong, An Dinh, et al.
Published: (2024) -
Robotic-CLIP: Fine-tuning CLIP on Action Data for Robotic Applications
by: Nguyen, Nghia, et al.
Published: (2024) -
GraspMamba: A Mamba-based Language-driven Grasp Detection Framework with Hierarchical Feature Learning
by: Nguyen, Huy Hoang, et al.
Published: (2024)