Robotic-CLIP: Fine-tuning CLIP on Action Data for Robotic Applications
Fuente:
arXiv
Saved in:
| Main Authors: | Nguyen, Nghia, Vu, Minh Nhat, Ta, Tung D., Huang, Baoru, Vo, Thieu, Le, Ngan, Nguyen, Anh |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Lightweight Language-driven Grasp Detection using Conditional Consistency Model
by: Nguyen, Nghia, et al.
Published: (2024)
by: Nguyen, Nghia, et al.
Published: (2024)
Language-driven Grasp Detection with Mask-guided Attention
by: Van Vo, Tuan, et al.
Published: (2024)
by: Van Vo, Tuan, et al.
Published: (2024)
Language-Driven 6-DoF Grasp Detection Using Negative Prompt Guidance
by: Nguyen, Toan, et al.
Published: (2024)
by: Nguyen, Toan, et al.
Published: (2024)
Learning Human Motion with Temporally Conditional Mamba
by: Nguyen, Quang, et al.
Published: (2025)
by: Nguyen, Quang, et al.
Published: (2025)
Language-driven Grasp Detection
by: Vuong, An Dinh, et al.
Published: (2024)
by: Vuong, An Dinh, et al.
Published: (2024)
ReFineVLA: Multimodal Reasoning-Aware Generalist Robotic Policies via Teacher-Guided Fine-Tuning
by: Van Vo, Tuan, et al.
Published: (2026)
by: Van Vo, Tuan, et al.
Published: (2026)
EgoMusic-driven Human Dance Motion Estimation with Skeleton Mamba
by: Nguyen, Quang, et al.
Published: (2025)
by: Nguyen, Quang, et al.
Published: (2025)
Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective
by: Chung, Nhat, et al.
Published: (2025)
by: Chung, Nhat, et al.
Published: (2025)
GraspMAS: Zero-Shot Language-driven Grasp Detection with Multi-Agent System
by: Nguyen, Quang, et al.
Published: (2025)
by: Nguyen, Quang, et al.
Published: (2025)
SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation
by: Hanyu, Taisei, et al.
Published: (2025)
by: Hanyu, Taisei, et al.
Published: (2025)
RoboDesign1M: A Large-scale Dataset for Robot Design Understanding
by: Le, Tri, et al.
Published: (2025)
by: Le, Tri, et al.
Published: (2025)
Autonomous Catheterization with Open-source Simulator and Expert Trajectory
by: Jianu, Tudor, et al.
Published: (2024)
by: Jianu, Tudor, et al.
Published: (2024)
GraspMamba: A Mamba-based Language-driven Grasp Detection Framework with Hierarchical Feature Learning
by: Nguyen, Huy Hoang, et al.
Published: (2024)
by: Nguyen, Huy Hoang, et al.
Published: (2024)
CADKnitter: Compositional CAD Generation from Text and Geometry Guidance
by: Le, Tri, et al.
Published: (2025)
by: Le, Tri, et al.
Published: (2025)
HabiCrowd: A High Performance Simulator for Crowd-Aware Visual Navigation
by: Vuong, An Dinh, et al.
Published: (2023)
by: Vuong, An Dinh, et al.
Published: (2023)
CLIP-Fields: Weakly Supervised Semantic Fields for Robotic Memory
by: Shafiullah, Nur Muhammad Mahi, et al.
Published: (2022)
by: Shafiullah, Nur Muhammad Mahi, et al.
Published: (2022)
XFlowMP: Task-Conditioned Motion Fields for Generative Robot Planning with Schrodinger Bridges
by: Nguyen, Khang, et al.
Published: (2025)
by: Nguyen, Khang, et al.
Published: (2025)
A Distributed Multi-Modal Sensing Approach for Human Activity Recognition in Real-Time Human-Robot Collaboration
by: Belcamino, Valerio, et al.
Published: (2026)
by: Belcamino, Valerio, et al.
Published: (2026)
SplineFormer: An Explainable Transformer-Based Approach for Autonomous Endovascular Navigation
by: Jianu, Tudor, et al.
Published: (2025)
by: Jianu, Tudor, et al.
Published: (2025)
Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding
by: Vo, Khoa, et al.
Published: (2025)
by: Vo, Khoa, et al.
Published: (2025)
ReFineVLA: Reasoning-Aware Teacher-Guided Transfer Fine-Tuning
by: Van Vo, Tuan, et al.
Published: (2025)
by: Van Vo, Tuan, et al.
Published: (2025)
Planning Robot Placement for Object Grasping
by: Saini, Manish, et al.
Published: (2024)
by: Saini, Manish, et al.
Published: (2024)
MulCLIP: A Multi-level Alignment Framework for Enhancing Fine-grained Long-context CLIP
by: Truong, Chau, et al.
Published: (2025)
by: Truong, Chau, et al.
Published: (2025)
Attention-Guided Integration of CLIP and SAM for Precise Object Masking in Robotic Manipulation
by: Muttaqien, Muhammad A., et al.
Published: (2025)
by: Muttaqien, Muhammad A., et al.
Published: (2025)
BiMa: Towards Biases Mitigation for Text-Video Retrieval via Scene Element Guidance
by: Le, Huy, et al.
Published: (2025)
by: Le, Huy, et al.
Published: (2025)
Lang2Lift: A Language-Guided Autonomous Forklift System for Outdoor Industrial Pallet Handling
by: Nguyen, Huy Hoang, et al.
Published: (2025)
by: Nguyen, Huy Hoang, et al.
Published: (2025)
ShapeFormer: Shape Prior Visible-to-Amodal Transformer-based Amodal Instance Segmentation
by: Tran, Minh, et al.
Published: (2024)
by: Tran, Minh, et al.
Published: (2024)
FedEFM: Federated Endovascular Foundation Model with Unseen Data
by: Do, Tuong, et al.
Published: (2025)
by: Do, Tuong, et al.
Published: (2025)
AffordMatcher: Affordance Learning in 3D Scenes from Visual Signifiers
by: Vu, Nghia, et al.
Published: (2026)
by: Vu, Nghia, et al.
Published: (2026)
Online Trajectory Replanner for Dynamically Grasping Irregular Objects
by: Vu, Minh Nhat, et al.
Published: (2025)
by: Vu, Minh Nhat, et al.
Published: (2025)
Semantic-aware Adversarial Fine-tuning for CLIP
by: Zhang, Jiacheng, et al.
Published: (2026)
by: Zhang, Jiacheng, et al.
Published: (2026)
Fine-Grained Action Segmentation for Renorrhaphy in Robot-Assisted Partial Nephrectomy
by: Dai, Jiaheng, et al.
Published: (2026)
by: Dai, Jiaheng, et al.
Published: (2026)
Amodal Instance Segmentation with Diffusion Shape Prior Estimation
by: Tran, Minh, et al.
Published: (2024)
by: Tran, Minh, et al.
Published: (2024)
WAVER: Writing-style Agnostic Text-Video Retrieval via Distilling Vision-Language Models Through Open-Vocabulary Knowledge
by: Le, Huy, et al.
Published: (2023)
by: Le, Huy, et al.
Published: (2023)
SIGMA: A Physics-Based Benchmark for Gas Chimney Understanding in Seismic Images
by: Truong, Bao, et al.
Published: (2026)
by: Truong, Bao, et al.
Published: (2026)
More Reliable Pseudo-labels, Better Performance: A Generalized Approach to Single Positive Multi-label Learning
by: Tran, Luong, et al.
Published: (2025)
by: Tran, Luong, et al.
Published: (2025)
AeroScene: Progressive Scene Synthesis for Aerial Robotics
by: Vu, Nghia, et al.
Published: (2026)
by: Vu, Nghia, et al.
Published: (2026)
FlowMP: Learning Motion Fields for Robot Planning with Conditional Flow Matching
by: Nguyen, Khang, et al.
Published: (2025)
by: Nguyen, Khang, et al.
Published: (2025)
Weakly-Supervised Learning via Multi-Lateral Decoder Branching for Tool Segmentation in Robot-Assisted Cardiovascular Catheterization
by: Omisore, Olatunji Mumini, et al.
Published: (2024)
by: Omisore, Olatunji Mumini, et al.
Published: (2024)
Volumetric Mapping with Panoptic Refinement via Kernel Density Estimation for Mobile Robots
by: Nguyen, Khang, et al.
Published: (2024)
by: Nguyen, Khang, et al.
Published: (2024)
Similar Items
-
Lightweight Language-driven Grasp Detection using Conditional Consistency Model
by: Nguyen, Nghia, et al.
Published: (2024) -
Language-driven Grasp Detection with Mask-guided Attention
by: Van Vo, Tuan, et al.
Published: (2024) -
Language-Driven 6-DoF Grasp Detection Using Negative Prompt Guidance
by: Nguyen, Toan, et al.
Published: (2024) -
Learning Human Motion with Temporally Conditional Mamba
by: Nguyen, Quang, et al.
Published: (2025) -
Language-driven Grasp Detection
by: Vuong, An Dinh, et al.
Published: (2024)