Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment
Fuente:
arXiv
Salvato in:
| Autori principali: | Lin, Tao, Zhong, Yilei, Du, Yuxin, Zhang, Jingjing, Liu, Jiting, Chen, Yinxinyu, Gu, Encheng, Liu, Ziyan, Cai, Hongyi, Zou, Yanwen, Zou, Lixing, Zhou, Zhaoye, Li, Gen, Zhao, Bo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Evo-0: Vision-Language-Action Model with Implicit Spatial Understanding
di: Lin, Tao, et al.
Pubblicazione: (2025)
di: Lin, Tao, et al.
Pubblicazione: (2025)
Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model
di: Lin, Tao, et al.
Pubblicazione: (2026)
di: Lin, Tao, et al.
Pubblicazione: (2026)
Bridging the Semantic-Action Gap in Visual Token Pruning for Efficient VLA Inference
di: Liu, Ziyan, et al.
Pubblicazione: (2025)
di: Liu, Ziyan, et al.
Pubblicazione: (2025)
EvoVLA: Self-Evolving Vision-Language-Action Model
di: Liu, Zeting, et al.
Pubblicazione: (2025)
di: Liu, Zeting, et al.
Pubblicazione: (2025)
U-ARM : Ultra low-cost general teleoperation interface for robot manipulation
di: Zou, Yanwen, et al.
Pubblicazione: (2025)
di: Zou, Yanwen, et al.
Pubblicazione: (2025)
Semantic Alignment and Reinforcement for Data-Free Quantization of Vision Transformers
di: Zhong, Yunshan, et al.
Pubblicazione: (2024)
di: Zhong, Yunshan, et al.
Pubblicazione: (2024)
Coarse-to-Fine Monocular Re-Localization in OpenStreetMap via Semantic Alignment
di: Zou, Yuchen, et al.
Pubblicazione: (2026)
di: Zou, Yuchen, et al.
Pubblicazione: (2026)
When Vision Meets Texts in Listwise Reranking
di: Cai, Hongyi
Pubblicazione: (2026)
di: Cai, Hongyi
Pubblicazione: (2026)
LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models
di: Fei, Senyu, et al.
Pubblicazione: (2025)
di: Fei, Senyu, et al.
Pubblicazione: (2025)
NoTVLA: Semantics-Preserving Robot Adaptation via Narrative Action Interfaces
di: Huang, Zheng, et al.
Pubblicazione: (2025)
di: Huang, Zheng, et al.
Pubblicazione: (2025)
ROSA: Harnessing Robot States for Vision-Language and Action Alignment
di: Wen, Yuqing, et al.
Pubblicazione: (2025)
di: Wen, Yuqing, et al.
Pubblicazione: (2025)
G-EvoNAS: Evolutionary Neural Architecture Search Based on Network Growth
di: Zou, Juan, et al.
Pubblicazione: (2024)
di: Zou, Juan, et al.
Pubblicazione: (2024)
Semantic Direct Modeling
di: Zou, Qiang, et al.
Pubblicazione: (2025)
di: Zou, Qiang, et al.
Pubblicazione: (2025)
EvoFormer: Learning Dynamic Graph-Level Representations with Structural and Temporal Bias Correction
di: Zhong, Haodi, et al.
Pubblicazione: (2025)
di: Zhong, Haodi, et al.
Pubblicazione: (2025)
World-aware Planning Narratives Enhance Large Vision-Language Model Planner
di: Shi, Junhao, et al.
Pubblicazione: (2025)
di: Shi, Junhao, et al.
Pubblicazione: (2025)
Lightweight Frequency Masker for Cross-Domain Few-Shot Semantic Segmentation
di: Tong, Jintao, et al.
Pubblicazione: (2024)
di: Tong, Jintao, et al.
Pubblicazione: (2024)
Enhancing Low‐Light Pedestrian Detection With Lightweight Vision Transformers
di: Xiang Gu, et al.
Pubblicazione: (2025)
di: Xiang Gu, et al.
Pubblicazione: (2025)
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks
di: Hu, Yuanze, et al.
Pubblicazione: (2025)
di: Hu, Yuanze, et al.
Pubblicazione: (2025)
Locality Alignment Improves Vision-Language Models
di: Covert, Ian, et al.
Pubblicazione: (2024)
di: Covert, Ian, et al.
Pubblicazione: (2024)
EvoEvolver/Forest: zenodo
di: Zijian Zhang, et al.
Pubblicazione: (2025)
di: Zijian Zhang, et al.
Pubblicazione: (2025)
EvoSpec: Evolving Speculative Decoding via Real-Time Vocabulary and Parameter Adaptation
di: Zhang, Shuyu, et al.
Pubblicazione: (2026)
di: Zhang, Shuyu, et al.
Pubblicazione: (2026)
EvoLM: In Search of Lost Language Model Training Dynamics
di: Qi, Zhenting, et al.
Pubblicazione: (2025)
di: Qi, Zhenting, et al.
Pubblicazione: (2025)
Transcriptome Insight Into the Identification of Novel Luteolin 7‐ O ‐Glucosyltransferase From Lonicerae Japonicae Flos Under Salt Stress
di: Zhichen Cai, et al.
Pubblicazione: (2026)
di: Zhichen Cai, et al.
Pubblicazione: (2026)
Focusable Monocular Depth Estimation
di: Du, Yuxin, et al.
Pubblicazione: (2026)
di: Du, Yuxin, et al.
Pubblicazione: (2026)
SemanticFace: Semantic Facial Action Estimation via Semantic Distillation in Interpretable Space
di: Kang, Zejian, et al.
Pubblicazione: (2026)
di: Kang, Zejian, et al.
Pubblicazione: (2026)
EVLF: Early Vision-Language Fusion for Generative Dataset Distillation
di: Cai, Wenqi, et al.
Pubblicazione: (2026)
di: Cai, Wenqi, et al.
Pubblicazione: (2026)
Asymmetric Visual Semantic Embedding Framework for Efficient Vision-Language Alignment
di: Liu, Yang, et al.
Pubblicazione: (2025)
di: Liu, Yang, et al.
Pubblicazione: (2025)
CliniQ: A Multi-faceted Benchmark for Electronic Health Record Retrieval with Semantic Match Assessment
di: Zhao, Zhengyun, et al.
Pubblicazione: (2025)
di: Zhao, Zhengyun, et al.
Pubblicazione: (2025)
LARY: A Latent Action Representation Yielding Benchmark for Generalizable Vision-to-Action Alignment
di: Nie, Dujun, et al.
Pubblicazione: (2026)
di: Nie, Dujun, et al.
Pubblicazione: (2026)
Dynamic Pattern Alignment Learning for Pretraining Lightweight Human-Centric Vision Models
di: Wang, Xuanhan, et al.
Pubblicazione: (2025)
di: Wang, Xuanhan, et al.
Pubblicazione: (2025)
Continuous Vision-Language-Action Co-Learning with Semantic-Physical Alignment for Behavioral Cloning
di: Qi, Xiuxiu, et al.
Pubblicazione: (2025)
di: Qi, Xiuxiu, et al.
Pubblicazione: (2025)
EvoEdit: Evolving Null-space Alignment for Robust and Efficient Knowledge Editing
di: Lyu, Sicheng, et al.
Pubblicazione: (2025)
di: Lyu, Sicheng, et al.
Pubblicazione: (2025)
Balancing Preservation and Modification: A Region and Semantic Aware Metric for Instruction-Based Image Editing
di: Li, Zhuoying, et al.
Pubblicazione: (2025)
di: Li, Zhuoying, et al.
Pubblicazione: (2025)
Personalized Federated Learning for Spatio-Temporal Forecasting: A Dual Semantic Alignment-Based Contrastive Approach
di: Liu, Qingxiang, et al.
Pubblicazione: (2024)
di: Liu, Qingxiang, et al.
Pubblicazione: (2024)
MS‐based multi‐dimensional metabolomics reveals protective effect of Polygalae Radix against metabolic disturbances in Alzheimer's disease mice
di: Yanwen Chen, et al.
Pubblicazione: (2025)
di: Yanwen Chen, et al.
Pubblicazione: (2025)
Storyboard guided Alignment for Fine-grained Video Action Recognition
di: Liu, Enqi, et al.
Pubblicazione: (2024)
di: Liu, Enqi, et al.
Pubblicazione: (2024)
Lightweight Vision Model-based Multi-user Semantic Communication Systems
di: Jiang, Feibo, et al.
Pubblicazione: (2025)
di: Jiang, Feibo, et al.
Pubblicazione: (2025)
ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models
di: Zhong, Linqing, et al.
Pubblicazione: (2026)
di: Zhong, Linqing, et al.
Pubblicazione: (2026)
SplatCo: Structure-View Collaborative Gaussian Splatting for Detail-Preserving Rendering of Large-Scale Unbounded Scenes
di: Xiao, Haihong, et al.
Pubblicazione: (2025)
di: Xiao, Haihong, et al.
Pubblicazione: (2025)
ZeroDDI: A Zero-Shot Drug-Drug Interaction Event Prediction Method with Semantic Enhanced Learning and Dual-Modal Uniform Alignment
di: Wang, Ziyan, et al.
Pubblicazione: (2024)
di: Wang, Ziyan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Evo-0: Vision-Language-Action Model with Implicit Spatial Understanding
di: Lin, Tao, et al.
Pubblicazione: (2025) -
Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model
di: Lin, Tao, et al.
Pubblicazione: (2026) -
Bridging the Semantic-Action Gap in Visual Token Pruning for Efficient VLA Inference
di: Liu, Ziyan, et al.
Pubblicazione: (2025) -
EvoVLA: Self-Evolving Vision-Language-Action Model
di: Liu, Zeting, et al.
Pubblicazione: (2025) -
U-ARM : Ultra low-cost general teleoperation interface for robot manipulation
di: Zou, Yanwen, et al.
Pubblicazione: (2025)