ReMem-VLA: Empowering Vision-Language-Action Model with Memory via Dual-Level Recurrent Queries
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Hang, Shen, Fengyi, Chen, Dong, Yang, Liudi, Wang, Xudong, Shi, Jinkui, Bing, Zhenshan, Liu, Ziyuan, Knoll, Alois |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ReMem: Mutual Information-Aware Fine-tuning of Pretrained Vision Transformers for Effective Knowledge Distillation
von: Dong, Chengyu, et al.
Veröffentlicht: (2025)
von: Dong, Chengyu, et al.
Veröffentlicht: (2025)
DualGazeNet: A Biologically Inspired Dual-Gaze Query Network for Salient Object Detection
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
Intelligent Transportation Systems Using External Infrastructure: A Literature Survey
von: Creß, Christian, et al.
Veröffentlicht: (2021)
von: Creß, Christian, et al.
Veröffentlicht: (2021)
CoVAR: Co-generation of Video and Action for Robotic Manipulation via Multi-Modal Diffusion
von: Yang, Liudi, et al.
Veröffentlicht: (2025)
von: Yang, Liudi, et al.
Veröffentlicht: (2025)
CE-NPBG: Connectivity Enhanced Neural Point-Based Graphics for Novel View Synthesis in Autonomous Driving Scenes
von: Altillawi, Mohammad, et al.
Veröffentlicht: (2025)
von: Altillawi, Mohammad, et al.
Veröffentlicht: (2025)
ControlUDA: Controllable Diffusion-assisted Unsupervised Domain Adaptation for Cross-Weather Semantic Segmentation
von: Shen, Fengyi, et al.
Veröffentlicht: (2024)
von: Shen, Fengyi, et al.
Veröffentlicht: (2024)
VideoWeaver: Multimodal Multi-View Video-to-Video Transfer for Embodied Agents
von: Eskandar, George, et al.
Veröffentlicht: (2026)
von: Eskandar, George, et al.
Veröffentlicht: (2026)
OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2025)
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2025)
Deep Fuzzy Optimization for Batch-Size and Nearest Neighbors in Optimal Robot Motion Planning
von: Zhang, Liding, et al.
Veröffentlicht: (2025)
von: Zhang, Liding, et al.
Veröffentlicht: (2025)
Genetic Informed Trees (GIT*): Path Planning via Reinforced Genetic Programming Heuristics
von: Zhang, Liding, et al.
Veröffentlicht: (2025)
von: Zhang, Liding, et al.
Veröffentlicht: (2025)
Dreaming the Unseen: World Model-regularized Diffusion Policy for Out-of-Distribution Robustness
von: Hu, Ziou, et al.
Veröffentlicht: (2026)
von: Hu, Ziou, et al.
Veröffentlicht: (2026)
Dynamic event‐triggered sliding mode control of networked switched systems with imperfect transmissions
von: Chunlian Wang, et al.
Veröffentlicht: (2024)
von: Chunlian Wang, et al.
Veröffentlicht: (2024)
Contact Energy Based Hindsight Experience Prioritization
von: Sayar, Erdi, et al.
Veröffentlicht: (2023)
von: Sayar, Erdi, et al.
Veröffentlicht: (2023)
DRAW2ACT: Turning Depth-Encoded Trajectories into Robotic Demonstration Videos
von: Bai, Yang, et al.
Veröffentlicht: (2025)
von: Bai, Yang, et al.
Veröffentlicht: (2025)
RoboSwap: A GAN-driven Video Diffusion Framework For Unsupervised Robot Arm Swapping
von: Bai, Yang, et al.
Veröffentlicht: (2025)
von: Bai, Yang, et al.
Veröffentlicht: (2025)
Language-Enhanced Mobile Manipulation for Efficient Object Search in Indoor Environments
von: Zhang, Liding, et al.
Veröffentlicht: (2025)
von: Zhang, Liding, et al.
Veröffentlicht: (2025)
Optimizing Dynamic Balance in a Rat Robot via the Lateral Flexion of a Soft Actuated Spine
von: Huang, Yuhong, et al.
Veröffentlicht: (2024)
von: Huang, Yuhong, et al.
Veröffentlicht: (2024)
Tree-Based Grafting Approach for Bidirectional Motion Planning with Local Subsets Optimization
von: Zhang, Liding, et al.
Veröffentlicht: (2025)
von: Zhang, Liding, et al.
Veröffentlicht: (2025)
Growing with Your Embodied Agent: A Human-in-the-Loop Lifelong Code Generation Framework for Long-Horizon Manipulation Skills
von: Meng, Yuan, et al.
Veröffentlicht: (2025)
von: Meng, Yuan, et al.
Veröffentlicht: (2025)
GeoVLA: Empowering 3D Representations in Vision-Language-Action Models
von: Sun, Lin, et al.
Veröffentlicht: (2025)
von: Sun, Lin, et al.
Veröffentlicht: (2025)
RobotDancing: Residual-Action Reinforcement Learning Enables Robust Long-Horizon Humanoid Motion Tracking
von: Sun, Zhenguo, et al.
Veröffentlicht: (2025)
von: Sun, Zhenguo, et al.
Veröffentlicht: (2025)
LoopVLA: Learning Sufficiency in Recurrent Refinement for Vision-Language-Action Models
von: Shen, Boyang, et al.
Veröffentlicht: (2026)
von: Shen, Boyang, et al.
Veröffentlicht: (2026)
MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation
von: Liu, Yang, et al.
Veröffentlicht: (2026)
von: Liu, Yang, et al.
Veröffentlicht: (2026)
ConfCtrl: Enabling Precise Camera Control in Video Diffusion via Confidence-Aware Interpolation
von: Yang, Liudi, et al.
Veröffentlicht: (2026)
von: Yang, Liudi, et al.
Veröffentlicht: (2026)
VLA-Touch: Enhancing Vision-Language-Action Models with Dual-Level Tactile Feedback
von: Bi, Jianxin, et al.
Veröffentlicht: (2025)
von: Bi, Jianxin, et al.
Veröffentlicht: (2025)
MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
von: Shi, Hao, et al.
Veröffentlicht: (2025)
von: Shi, Hao, et al.
Veröffentlicht: (2025)
LEMMo-Plan: LLM-Enhanced Learning from Multi-Modal Demonstration for Planning Sequential Contact-Rich Manipulation Tasks
von: Chen, Kejia, et al.
Veröffentlicht: (2024)
von: Chen, Kejia, et al.
Veröffentlicht: (2024)
ContactDexNet: Multi-fingered Robotic Hand Grasping in Cluttered Environments through Hand-object Contact Semantic Mapping
von: Zhang, Lei, et al.
Veröffentlicht: (2024)
von: Zhang, Lei, et al.
Veröffentlicht: (2024)
Pretrained Bayesian Non-parametric Knowledge Prior in Robotic Long-Horizon Reinforcement Learning
von: Meng, Yuan, et al.
Veröffentlicht: (2025)
von: Meng, Yuan, et al.
Veröffentlicht: (2025)
Language-Conditioned Imitation Learning with Base Skill Priors under Unstructured Data
von: Zhou, Hongkuan, et al.
Veröffentlicht: (2023)
von: Zhou, Hongkuan, et al.
Veröffentlicht: (2023)
Locomotion Generation for a Rat Robot based on Environmental Changes via Reinforcement Learning
von: Shan, Xinhui, et al.
Veröffentlicht: (2024)
von: Shan, Xinhui, et al.
Veröffentlicht: (2024)
Real-Time Adaptive Safety-Critical Control with Gaussian Processes in High-Order Uncertain Models
von: Zhang, Yu, et al.
Veröffentlicht: (2024)
von: Zhang, Yu, et al.
Veröffentlicht: (2024)
Real-time Contact State Estimation in Shape Control of Deformable Linear Objects under Small Environmental Constraints
von: Chen, Kejia, et al.
Veröffentlicht: (2024)
von: Chen, Kejia, et al.
Veröffentlicht: (2024)
OmniDexVLG: Learning Dexterous Grasp Generation from Vision Language Model-Guided Grasp Semantics, Taxonomy and Functional Affordance
von: Zhang, Lei, et al.
Veröffentlicht: (2025)
von: Zhang, Lei, et al.
Veröffentlicht: (2025)
GesVLA: Gesture-Aware Vision-Language-Action Model Embedded Representations
von: Guo, Wenxuan, et al.
Veröffentlicht: (2026)
von: Guo, Wenxuan, et al.
Veröffentlicht: (2026)
RoboEnvision: A Long-Horizon Video Generation Model for Multi-Task Robot Manipulation
von: Yang, Liudi, et al.
Veröffentlicht: (2025)
von: Yang, Liudi, et al.
Veröffentlicht: (2025)
DropVLA: An Action-Level Backdoor Attack on Vision-Language-Action Models
von: Xu, Zonghuan, et al.
Veröffentlicht: (2025)
von: Xu, Zonghuan, et al.
Veröffentlicht: (2025)
SwiftMem: Fast Agentic Memory via Query-aware Indexing
von: Tian, Anxin, et al.
Veröffentlicht: (2026)
von: Tian, Anxin, et al.
Veröffentlicht: (2026)
LLM-Empowered Functional Safety and Security by Design in Automotive Systems
von: Petrovic, Nenad, et al.
Veröffentlicht: (2026)
von: Petrovic, Nenad, et al.
Veröffentlicht: (2026)
GPT-4V as Traffic Assistant: An In-depth Look at Vision Language Model on Complex Traffic Events
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2024)
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ReMem: Mutual Information-Aware Fine-tuning of Pretrained Vision Transformers for Effective Knowledge Distillation
von: Dong, Chengyu, et al.
Veröffentlicht: (2025) -
DualGazeNet: A Biologically Inspired Dual-Gaze Query Network for Salient Object Detection
von: Zhang, Yu, et al.
Veröffentlicht: (2025) -
Intelligent Transportation Systems Using External Infrastructure: A Literature Survey
von: Creß, Christian, et al.
Veröffentlicht: (2021) -
CoVAR: Co-generation of Video and Action for Robotic Manipulation via Multi-Modal Diffusion
von: Yang, Liudi, et al.
Veröffentlicht: (2025) -
CE-NPBG: Connectivity Enhanced Neural Point-Based Graphics for Novel View Synthesis in Autonomous Driving Scenes
von: Altillawi, Mohammad, et al.
Veröffentlicht: (2025)