AnoleVLA: Lightweight Vision-Language-Action Model with Deep State Space Models for Mobile Manipulation
Fuente:
arXiv
Guardado en:
| Autores principales: | Takagi, Yusuke, Kambara, Motonari, Yashima, Daichi, Seno, Koki, Tokura, Kento, Sugiura, Komei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Pre-Manipulation Alignment Prediction with Parallel Deep State-Space and Transformer Models
por: Kambara, Motonari, et al.
Publicado: (2025)
por: Kambara, Motonari, et al.
Publicado: (2025)
Mobile Manipulation Instruction Generation from Multiple Images with Automatic Metric Enhancement
por: Katsumata, Kei, et al.
Publicado: (2025)
por: Katsumata, Kei, et al.
Publicado: (2025)
Future Success Prediction in Open-Vocabulary Object Manipulation Tasks Based on End-Effector Trajectories
por: Kambara, Motonari, et al.
Publicado: (2024)
por: Kambara, Motonari, et al.
Publicado: (2024)
LILAC: Language-Conditioned Object-Centric Optical Flow for Open-Loop Trajectory Generation
por: Kambara, Motonari, et al.
Publicado: (2026)
por: Kambara, Motonari, et al.
Publicado: (2026)
HiFlow: Tokenization-Free Scale-Wise Autoregressive Policy Learning via Flow Matching
por: Yashima, Daichi, et al.
Publicado: (2026)
por: Yashima, Daichi, et al.
Publicado: (2026)
Task Success Prediction for Open-Vocabulary Manipulation Based on Multi-Level Aligned Representations
por: Goko, Miyu, et al.
Publicado: (2024)
por: Goko, Miyu, et al.
Publicado: (2024)
Object Segmentation from Open-Vocabulary Manipulation Instructions Based on Optimal Transport Polygon Matching with Multimodal Foundation Models
por: Nishimura, Takayuki, et al.
Publicado: (2024)
por: Nishimura, Takayuki, et al.
Publicado: (2024)
Open-Vocabulary Mobile Manipulation Based on Double Relaxed Contrastive Learning with Dense Labeling
por: Yashima, Daichi, et al.
Publicado: (2024)
por: Yashima, Daichi, et al.
Publicado: (2024)
Nearest Neighbor Future Captioning: Generating Descriptions for Possible Collisions in Object Placement Tasks
por: Komatsu, Takumi, et al.
Publicado: (2024)
por: Komatsu, Takumi, et al.
Publicado: (2024)
FLARE-SSM: Deep State Space Models with Influence-Balanced Loss for 72-Hour Solar Flare Prediction
por: Takagi, Yusuke, et al.
Publicado: (2025)
por: Takagi, Yusuke, et al.
Publicado: (2025)
ReMoRa: Multimodal Large Language Model based on Refined Motion Representation for Long-Video Understanding
por: Yashima, Daichi, et al.
Publicado: (2026)
por: Yashima, Daichi, et al.
Publicado: (2026)
SG-VLA: Learning Spatially-Grounded Vision-Language-Action Models for Mobile Manipulation
por: Tu, Ruisen, et al.
Publicado: (2026)
por: Tu, Ruisen, et al.
Publicado: (2026)
MLLM-as-a-Judge Exhibits Model Preference Bias
por: Koyama, Shuitsu, et al.
Publicado: (2026)
por: Koyama, Shuitsu, et al.
Publicado: (2026)
Affordance RAG: Hierarchical Multimodal Retrieval with Affordance-Aware Embodied Memory for Mobile Manipulation
por: Korekata, Ryosuke, et al.
Publicado: (2025)
por: Korekata, Ryosuke, et al.
Publicado: (2025)
DAM-VLA: A Dynamic Action Model-Based Vision-Language-Action Framework for Robot Manipulation
por: Peng, Xiongfeng, et al.
Publicado: (2026)
por: Peng, Xiongfeng, et al.
Publicado: (2026)
ABMAMBA: Multimodal Large Language Model with Aligned Hierarchical Bidirectional Scan for Efficient Video Captioning
por: Yashima, Daichi, et al.
Publicado: (2026)
por: Yashima, Daichi, et al.
Publicado: (2026)
AIR-VLA: Vision-Language-Action Systems for Aerial Manipulation
por: Sun, Jianli, et al.
Publicado: (2026)
por: Sun, Jianli, et al.
Publicado: (2026)
Co-Scale Cross-Attentional Transformer for Rearrangement Target Detection
por: Matsuo, Haruka, et al.
Publicado: (2024)
por: Matsuo, Haruka, et al.
Publicado: (2024)
Bi-VLA: Vision-Language-Action Model-Based System for Bimanual Robotic Dexterous Manipulations
por: Gbagbe, Koffivi Fidèle, et al.
Publicado: (2024)
por: Gbagbe, Koffivi Fidèle, et al.
Publicado: (2024)
Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation
por: Fan, Yiguo, et al.
Publicado: (2025)
por: Fan, Yiguo, et al.
Publicado: (2025)
VLA^2: Empowering Vision-Language-Action Models with an Agentic Framework for Unseen Concept Manipulation
por: Zhao, Han, et al.
Publicado: (2025)
por: Zhao, Han, et al.
Publicado: (2025)
TA-VLA: Elucidating the Design Space of Torque-aware Vision-Language-Action Models
por: Zhang, Zongzheng, et al.
Publicado: (2025)
por: Zhang, Zongzheng, et al.
Publicado: (2025)
BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
por: Wang, Hongyu, et al.
Publicado: (2025)
por: Wang, Hongyu, et al.
Publicado: (2025)
DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation
por: Xie, Haozhe, et al.
Publicado: (2026)
por: Xie, Haozhe, et al.
Publicado: (2026)
Audio-VLA: Adding Contact Audio Perception to Vision-Language-Action Model for Robotic Manipulation
por: Wei, Xiangyi, et al.
Publicado: (2025)
por: Wei, Xiangyi, et al.
Publicado: (2025)
Shake-VLA: Vision-Language-Action Model-Based System for Bimanual Robotic Manipulations and Liquid Mixing
por: Khan, Muhamamd Haris, et al.
Publicado: (2025)
por: Khan, Muhamamd Haris, et al.
Publicado: (2025)
TaF-VLA: Tactile-Force Alignment in Vision-Language-Action Models for Force-aware Manipulation
por: Huang, Yuzhe, et al.
Publicado: (2026)
por: Huang, Yuzhe, et al.
Publicado: (2026)
ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models
por: Zhong, Linqing, et al.
Publicado: (2026)
por: Zhong, Linqing, et al.
Publicado: (2026)
AC^2-VLA: Action-Context-Aware Adaptive Computation in Vision-Language-Action Models for Efficient Robotic Manipulation
por: Yu, Wenda, et al.
Publicado: (2026)
por: Yu, Wenda, et al.
Publicado: (2026)
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
por: Deng, Shengliang, et al.
Publicado: (2025)
por: Deng, Shengliang, et al.
Publicado: (2025)
DeMaVLA: A Vision-Language-Action Foundation Model for Generalizable Deformable Manipulation
por: Su, Taiyi, et al.
Publicado: (2026)
por: Su, Taiyi, et al.
Publicado: (2026)
MoManipVLA: Transferring Vision-language-action Models for General Mobile Manipulation
por: Wu, Zhenyu, et al.
Publicado: (2025)
por: Wu, Zhenyu, et al.
Publicado: (2025)
MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
por: Shi, Hao, et al.
Publicado: (2025)
por: Shi, Hao, et al.
Publicado: (2025)
MAP-VLA: Memory-Augmented Prompting for Vision-Language-Action Model in Robotic Manipulation
por: Li, Runhao, et al.
Publicado: (2025)
por: Li, Runhao, et al.
Publicado: (2025)
EveryDayVLA: A Vision-Language-Action Model for Affordable Robotic Manipulation
por: Chopra, Samarth, et al.
Publicado: (2025)
por: Chopra, Samarth, et al.
Publicado: (2025)
FD-VLA: Force-Distilled Vision-Language-Action Model for Contact-Rich Manipulation
por: Zhao, Ruiteng, et al.
Publicado: (2026)
por: Zhao, Ruiteng, et al.
Publicado: (2026)
AnchorVLA4D: an Anchor-Based Spatial-Temporal Vision-Language-Action Model for Robotic Manipulation
por: Zhu, Juan, et al.
Publicado: (2026)
por: Zhu, Juan, et al.
Publicado: (2026)
ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic Manipulation
por: Liu, Zhenyang, et al.
Publicado: (2026)
por: Liu, Zhenyang, et al.
Publicado: (2026)
EchoVLA: Synergistic Declarative Memory for VLA-Driven Mobile Manipulation
por: Lin, Min, et al.
Publicado: (2025)
por: Lin, Min, et al.
Publicado: (2025)
TwinVLA: Data-Efficient Bimanual Manipulation with Twin Single-Arm Vision-Language-Action Models
por: Im, Hokyun, et al.
Publicado: (2025)
por: Im, Hokyun, et al.
Publicado: (2025)
Ejemplares similares
-
Pre-Manipulation Alignment Prediction with Parallel Deep State-Space and Transformer Models
por: Kambara, Motonari, et al.
Publicado: (2025) -
Mobile Manipulation Instruction Generation from Multiple Images with Automatic Metric Enhancement
por: Katsumata, Kei, et al.
Publicado: (2025) -
Future Success Prediction in Open-Vocabulary Object Manipulation Tasks Based on End-Effector Trajectories
por: Kambara, Motonari, et al.
Publicado: (2024) -
LILAC: Language-Conditioned Object-Centric Optical Flow for Open-Loop Trajectory Generation
por: Kambara, Motonari, et al.
Publicado: (2026) -
HiFlow: Tokenization-Free Scale-Wise Autoregressive Policy Learning via Flow Matching
por: Yashima, Daichi, et al.
Publicado: (2026)