Language-driven Grasp Detection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Vuong, An Dinh, Vu, Minh Nhat, Huang, Baoru, Nguyen, Nghia, Le, Hieu, Vo, Thieu, Nguyen, Anh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Lightweight Language-driven Grasp Detection using Conditional Consistency Model
von: Nguyen, Nghia, et al.
Veröffentlicht: (2024)
von: Nguyen, Nghia, et al.
Veröffentlicht: (2024)
Language-driven Grasp Detection with Mask-guided Attention
von: Van Vo, Tuan, et al.
Veröffentlicht: (2024)
von: Van Vo, Tuan, et al.
Veröffentlicht: (2024)
Language-Driven 6-DoF Grasp Detection Using Negative Prompt Guidance
von: Nguyen, Toan, et al.
Veröffentlicht: (2024)
von: Nguyen, Toan, et al.
Veröffentlicht: (2024)
EgoMusic-driven Human Dance Motion Estimation with Skeleton Mamba
von: Nguyen, Quang, et al.
Veröffentlicht: (2025)
von: Nguyen, Quang, et al.
Veröffentlicht: (2025)
Robotic-CLIP: Fine-tuning CLIP on Action Data for Robotic Applications
von: Nguyen, Nghia, et al.
Veröffentlicht: (2024)
von: Nguyen, Nghia, et al.
Veröffentlicht: (2024)
GraspMamba: A Mamba-based Language-driven Grasp Detection Framework with Hierarchical Feature Learning
von: Nguyen, Huy Hoang, et al.
Veröffentlicht: (2024)
von: Nguyen, Huy Hoang, et al.
Veröffentlicht: (2024)
Learning Human Motion with Temporally Conditional Mamba
von: Nguyen, Quang, et al.
Veröffentlicht: (2025)
von: Nguyen, Quang, et al.
Veröffentlicht: (2025)
HabiCrowd: A High Performance Simulator for Crowd-Aware Visual Navigation
von: Vuong, An Dinh, et al.
Veröffentlicht: (2023)
von: Vuong, An Dinh, et al.
Veröffentlicht: (2023)
More Reliable Pseudo-labels, Better Performance: A Generalized Approach to Single Positive Multi-label Learning
von: Tran, Luong, et al.
Veröffentlicht: (2025)
von: Tran, Luong, et al.
Veröffentlicht: (2025)
Fourier-Attentive Representation Learning: A Fourier-Guided Framework for Few-Shot Generalization in Vision-Language Models
von: Pham, Hieu Dinh Trung, et al.
Veröffentlicht: (2025)
von: Pham, Hieu Dinh Trung, et al.
Veröffentlicht: (2025)
Bridging the Training-Deployment Gap: Gated Encoding and Multi-Scale Refinement for Efficient Quantization-Aware Image Enhancement
von: To-Thanh, Dat, et al.
Veröffentlicht: (2026)
von: To-Thanh, Dat, et al.
Veröffentlicht: (2026)
SAMURAI: Shape-Aware Multimodal Retrieval for 3D Object Identification
von: Vo, Dinh-Khoi, et al.
Veröffentlicht: (2025)
von: Vo, Dinh-Khoi, et al.
Veröffentlicht: (2025)
EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic Captions
von: Vo, Dinh-Khoi, et al.
Veröffentlicht: (2025)
von: Vo, Dinh-Khoi, et al.
Veröffentlicht: (2025)
AffordMatcher: Affordance Learning in 3D Scenes from Visual Signifiers
von: Vu, Nghia, et al.
Veröffentlicht: (2026)
von: Vu, Nghia, et al.
Veröffentlicht: (2026)
Autonomous Catheterization with Open-source Simulator and Expert Trajectory
von: Jianu, Tudor, et al.
Veröffentlicht: (2024)
von: Jianu, Tudor, et al.
Veröffentlicht: (2024)
Shape2Animal: Creative Animal Generation from Natural Silhouettes
von: Tran, Quoc-Duy, et al.
Veröffentlicht: (2025)
von: Tran, Quoc-Duy, et al.
Veröffentlicht: (2025)
PANDORA: Pixel-wise Attention Dissolution and Latent Guidance for Zero-Shot Object Removal
von: Vo, Dinh-Khoi, et al.
Veröffentlicht: (2026)
von: Vo, Dinh-Khoi, et al.
Veröffentlicht: (2026)
GraspMAS: Zero-Shot Language-driven Grasp Detection with Multi-Agent System
von: Nguyen, Quang, et al.
Veröffentlicht: (2025)
von: Nguyen, Quang, et al.
Veröffentlicht: (2025)
ReFineVLA: Multimodal Reasoning-Aware Generalist Robotic Policies via Teacher-Guided Fine-Tuning
von: Van Vo, Tuan, et al.
Veröffentlicht: (2026)
von: Van Vo, Tuan, et al.
Veröffentlicht: (2026)
FedEFM: Federated Endovascular Foundation Model with Unseen Data
von: Do, Tuong, et al.
Veröffentlicht: (2025)
von: Do, Tuong, et al.
Veröffentlicht: (2025)
Hierarchical Neural Collapse Detection Transformer for Class Incremental Object Detection
von: Pham, Duc Thanh, et al.
Veröffentlicht: (2025)
von: Pham, Duc Thanh, et al.
Veröffentlicht: (2025)
Revisit Visual Prompt Tuning: The Expressiveness of Prompt Experts
von: Le, Minh, et al.
Veröffentlicht: (2025)
von: Le, Minh, et al.
Veröffentlicht: (2025)
VisionGuard: Synergistic Framework for Helmet Violation Detection
von: Nguyen, Lam-Huy, et al.
Veröffentlicht: (2025)
von: Nguyen, Lam-Huy, et al.
Veröffentlicht: (2025)
CPAM: Context-Preserving Adaptive Manipulation for Zero-Shot Real Image Editing
von: Vo, Dinh-Khoi, et al.
Veröffentlicht: (2025)
von: Vo, Dinh-Khoi, et al.
Veröffentlicht: (2025)
Brain Tumor Segmentation in MRI Images with 3D U-Net and Contextual Transformer
von: Nguyen, Thien-Qua T., et al.
Veröffentlicht: (2024)
von: Nguyen, Thien-Qua T., et al.
Veröffentlicht: (2024)
OpenEvents V1: Large-Scale Benchmark Dataset for Multimodal Event Grounding
von: Nguyen, Hieu, et al.
Veröffentlicht: (2025)
von: Nguyen, Hieu, et al.
Veröffentlicht: (2025)
CADKnitter: Compositional CAD Generation from Text and Geometry Guidance
von: Le, Tri, et al.
Veröffentlicht: (2025)
von: Le, Tri, et al.
Veröffentlicht: (2025)
iCONTRA: Toward Thematic Collection Design Via Interactive Concept Transfer
von: Vo, Dinh-Khoi, et al.
Veröffentlicht: (2024)
von: Vo, Dinh-Khoi, et al.
Veröffentlicht: (2024)
ShapeFormer: Shape Prior Visible-to-Amodal Transformer-based Amodal Instance Segmentation
von: Tran, Minh, et al.
Veröffentlicht: (2024)
von: Tran, Minh, et al.
Veröffentlicht: (2024)
SIGMA: A Physics-Based Benchmark for Gas Chimney Understanding in Seismic Images
von: Truong, Bao, et al.
Veröffentlicht: (2026)
von: Truong, Bao, et al.
Veröffentlicht: (2026)
The Art of Camouflage: Few-Shot Learning for Animal Detection and Segmentation
von: Nguyen, Thanh-Danh, et al.
Veröffentlicht: (2023)
von: Nguyen, Thanh-Danh, et al.
Veröffentlicht: (2023)
Ensemble Learning for Vietnamese Scene Text Spotting in Urban Environments
von: Nguyen, Hieu, et al.
Veröffentlicht: (2024)
von: Nguyen, Hieu, et al.
Veröffentlicht: (2024)
Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention
von: Nguyen, Nhi Ngoc-Yen, et al.
Veröffentlicht: (2026)
von: Nguyen, Nhi Ngoc-Yen, et al.
Veröffentlicht: (2026)
ReCap: Event-Aware Image Captioning with Article Retrieval and Semantic Gaussian Normalization
von: Nguyen, Thinh-Phuc, et al.
Veröffentlicht: (2025)
von: Nguyen, Thinh-Phuc, et al.
Veröffentlicht: (2025)
Beyond Traditional Approaches: Multi-Task Network for Breast Ultrasound Diagnosis
von: Chung, Dat T., et al.
Veröffentlicht: (2024)
von: Chung, Dat T., et al.
Veröffentlicht: (2024)
SHREC 2025: Retrieval of Optimal Objects for Multi-modal Enhanced Language and Spatial Assistance (ROOMELSA)
von: Nguyen, Trong-Thuan, et al.
Veröffentlicht: (2025)
von: Nguyen, Trong-Thuan, et al.
Veröffentlicht: (2025)
TaleForge: Interactive Multimodal System for Personalized Story Creation
von: Nguyen, Minh-Loi, et al.
Veröffentlicht: (2025)
von: Nguyen, Minh-Loi, et al.
Veröffentlicht: (2025)
NeIn: Telling What You Don't Want
von: Bui, Nhat-Tan, et al.
Veröffentlicht: (2024)
von: Bui, Nhat-Tan, et al.
Veröffentlicht: (2024)
Multi-scale Coarse-to-fine Modeling for Test-time Human Motion Control
von: Le, Nhat, et al.
Veröffentlicht: (2026)
von: Le, Nhat, et al.
Veröffentlicht: (2026)
ViConsFormer: Constituting Meaningful Phrases of Scene Texts using Transformer-based Method in Vietnamese Text-based Visual Question Answering
von: Nguyen, Nghia Hieu, et al.
Veröffentlicht: (2024)
von: Nguyen, Nghia Hieu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Lightweight Language-driven Grasp Detection using Conditional Consistency Model
von: Nguyen, Nghia, et al.
Veröffentlicht: (2024) -
Language-driven Grasp Detection with Mask-guided Attention
von: Van Vo, Tuan, et al.
Veröffentlicht: (2024) -
Language-Driven 6-DoF Grasp Detection Using Negative Prompt Guidance
von: Nguyen, Toan, et al.
Veröffentlicht: (2024) -
EgoMusic-driven Human Dance Motion Estimation with Skeleton Mamba
von: Nguyen, Quang, et al.
Veröffentlicht: (2025) -
Robotic-CLIP: Fine-tuning CLIP on Action Data for Robotic Applications
von: Nguyen, Nghia, et al.
Veröffentlicht: (2024)