Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Chung, Nhat, Hanyu, Taisei, Nguyen, Toan, Le, Huy, Bumgarner, Frederick, Nguyen, Duy Minh Ho, Vo, Khoa, Yamazaki, Kashu, Rainwater, Chase, Kieu, Tung, Nguyen, Anh, Le, Ngan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation
by: Hanyu, Taisei, et al.
Published: (2025)
by: Hanyu, Taisei, et al.
Published: (2025)
Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding
by: Vo, Khoa, et al.
Published: (2025)
by: Vo, Khoa, et al.
Published: (2025)
CodeGraphVLP: Code-as-Planner Meets Semantic-Graph State for Non-Markovian Vision-Language-Action Models
by: Vo, Khoa, et al.
Published: (2026)
by: Vo, Khoa, et al.
Published: (2026)
BiMa: Towards Biases Mitigation for Text-Video Retrieval via Scene Element Guidance
by: Le, Huy, et al.
Published: (2025)
by: Le, Huy, et al.
Published: (2025)
WAVER: Writing-style Agnostic Text-Video Retrieval via Distilling Vision-Language Models Through Open-Vocabulary Knowledge
by: Le, Huy, et al.
Published: (2023)
by: Le, Huy, et al.
Published: (2023)
HENASY: Learning to Assemble Scene-Entities for Egocentric Video-Language Model
by: Vo, Khoa, et al.
Published: (2024)
by: Vo, Khoa, et al.
Published: (2024)
UNO: Unifying One-stage Video Scene Graph Generation via Object-Centric Visual Representation Learning
by: Le, Huy, et al.
Published: (2025)
by: Le, Huy, et al.
Published: (2025)
AISFormer: Amodal Instance Segmentation with Transformer
by: Tran, Minh, et al.
Published: (2022)
by: Tran, Minh, et al.
Published: (2022)
ShapeFormer: Shape Prior Visible-to-Amodal Transformer-based Amodal Instance Segmentation
by: Tran, Minh, et al.
Published: (2024)
by: Tran, Minh, et al.
Published: (2024)
Amodal Instance Segmentation with Diffusion Shape Prior Estimation
by: Tran, Minh, et al.
Published: (2024)
by: Tran, Minh, et al.
Published: (2024)
Language-Driven 6-DoF Grasp Detection Using Negative Prompt Guidance
by: Nguyen, Toan, et al.
Published: (2024)
by: Nguyen, Toan, et al.
Published: (2024)
Robotic-CLIP: Fine-tuning CLIP on Action Data for Robotic Applications
by: Nguyen, Nghia, et al.
Published: (2024)
by: Nguyen, Nghia, et al.
Published: (2024)
Learning Human Motion with Temporally Conditional Mamba
by: Nguyen, Quang, et al.
Published: (2025)
by: Nguyen, Quang, et al.
Published: (2025)
DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving
by: Vo, Hao, et al.
Published: (2026)
by: Vo, Hao, et al.
Published: (2026)
EgoMusic-driven Human Dance Motion Estimation with Skeleton Mamba
by: Nguyen, Quang, et al.
Published: (2025)
by: Nguyen, Quang, et al.
Published: (2025)
Revisit Visual Prompt Tuning: The Expressiveness of Prompt Experts
by: Le, Minh, et al.
Published: (2025)
by: Le, Minh, et al.
Published: (2025)
Lightweight Language-driven Grasp Detection using Conditional Consistency Model
by: Nguyen, Nghia, et al.
Published: (2024)
by: Nguyen, Nghia, et al.
Published: (2024)
Co-NAML-LSTUR: A Combined Model with Attentive Multi-View Learning and Long- and Short-term User Representations for News Recommendation
by: Nguyen, Minh Hoang, et al.
Published: (2025)
by: Nguyen, Minh Hoang, et al.
Published: (2025)
Language-driven Grasp Detection with Mask-guided Attention
by: Van Vo, Tuan, et al.
Published: (2024)
by: Van Vo, Tuan, et al.
Published: (2024)
RoboDesign1M: A Large-scale Dataset for Robot Design Understanding
by: Le, Tri, et al.
Published: (2025)
by: Le, Tri, et al.
Published: (2025)
Neural Network‐Driven Adaptive Control for Swing‐Free Trajectory Tracking in Double‐Link Overhead Cranes With Uncertain Dynamics
by: Manh Cuong Nguyen, et al.
Published: (2026)
by: Manh Cuong Nguyen, et al.
Published: (2026)
Reinforcing Trustworthiness in Multimodal Emotional Support Systems
by: Le, Huy M., et al.
Published: (2025)
by: Le, Huy M., et al.
Published: (2025)
Rethinking Top Probability from Multi-view for Distracted Driver Behaviour Localization
by: Nguyen, Quang Vinh, et al.
Published: (2024)
by: Nguyen, Quang Vinh, et al.
Published: (2024)
FurniMAS: Language-Guided Furniture Decoration using Multi-Agent System
by: Nguyen, Toan, et al.
Published: (2025)
by: Nguyen, Toan, et al.
Published: (2025)
SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object Detection
by: Vo, Hao, et al.
Published: (2026)
by: Vo, Hao, et al.
Published: (2026)
Folding model approach to the elastic $p+^{12,13}$C scattering at low energies and radiative capture $^{12,13}$C$(p,γ)$ reactions
by: Anh, Nguyen Le, et al.
Published: (2020)
by: Anh, Nguyen Le, et al.
Published: (2020)
One-Prompt Strikes Back: Sparse Mixture of Experts for Prompt-based Continual Learning
by: Le, Minh, et al.
Published: (2025)
by: Le, Minh, et al.
Published: (2025)
Asymmetric time‐varying integral barrier Lyapunov control‐based trajectory tracking of autonomous vehicles with input magnitude and rate constraints
by: Nhu Toan Nguyen, et al.
Published: (2025)
by: Nhu Toan Nguyen, et al.
Published: (2025)
GraspMAS: Zero-Shot Language-driven Grasp Detection with Multi-Agent System
by: Nguyen, Quang, et al.
Published: (2025)
by: Nguyen, Quang, et al.
Published: (2025)
Revisiting Prefix-tuning: Statistical Benefits of Reparameterization among Prompts
by: Le, Minh, et al.
Published: (2024)
by: Le, Minh, et al.
Published: (2024)
RepLoRA: Reparameterizing Low-Rank Adaptation via the Perspective of Mixture of Experts
by: Truong, Tuan, et al.
Published: (2025)
by: Truong, Tuan, et al.
Published: (2025)
ReFineVLA: Multimodal Reasoning-Aware Generalist Robotic Policies via Teacher-Guided Fine-Tuning
by: Van Vo, Tuan, et al.
Published: (2026)
by: Van Vo, Tuan, et al.
Published: (2026)
The Art of Camouflage: Few-Shot Learning for Animal Detection and Segmentation
by: Nguyen, Thanh-Danh, et al.
Published: (2023)
by: Nguyen, Thanh-Danh, et al.
Published: (2023)
Enhanced Multimodal Video Retrieval System: Integrating Query Expansion and Cross-modal Temporal Event Retrieval
by: Vo, Van-Thinh, et al.
Published: (2025)
by: Vo, Van-Thinh, et al.
Published: (2025)
Secrecy Offloading Analysis of UAV-assisted NOMA-MEC Incorporating WPT in IoT Networks
by: Nguyen, Gia-Huy, et al.
Published: (2025)
by: Nguyen, Gia-Huy, et al.
Published: (2025)
Universal Multi-Domain Translation via Diffusion Routers
by: Kieu, Duc, et al.
Published: (2025)
by: Kieu, Duc, et al.
Published: (2025)
Unified Interactive Multimodal Moment Retrieval via Cascaded Embedding-Reranking and Temporal-Aware Score Fusion
by: Thanh, Toan Le Ngo, et al.
Published: (2025)
by: Thanh, Toan Le Ngo, et al.
Published: (2025)
ReFineVLA: Reasoning-Aware Teacher-Guided Transfer Fine-Tuning
by: Van Vo, Tuan, et al.
Published: (2025)
by: Van Vo, Tuan, et al.
Published: (2025)
Variational Flow Models: Flowing in Your Style
by: Do, Kien, et al.
Published: (2024)
by: Do, Kien, et al.
Published: (2024)
Statistical Inference for Clustering-based Anomaly Detection
by: Phu, Nguyen Thi Minh, et al.
Published: (2025)
by: Phu, Nguyen Thi Minh, et al.
Published: (2025)
Similar Items
-
SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation
by: Hanyu, Taisei, et al.
Published: (2025) -
Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding
by: Vo, Khoa, et al.
Published: (2025) -
CodeGraphVLP: Code-as-Planner Meets Semantic-Graph State for Non-Markovian Vision-Language-Action Models
by: Vo, Khoa, et al.
Published: (2026) -
BiMa: Towards Biases Mitigation for Text-Video Retrieval via Scene Element Guidance
by: Le, Huy, et al.
Published: (2025) -
WAVER: Writing-style Agnostic Text-Video Retrieval via Distilling Vision-Language Models Through Open-Vocabulary Knowledge
by: Le, Huy, et al.
Published: (2023)