MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Yi, Xu, Xiao, Xu, Zeyu, Zhang, Meng, Li, Yibo, Chen, Haoyu, Zhang, Junkang, Wang, Qiang, Sun, Jifa, Lin, Siling, Cheng, Shengxun, Zhang, Lingshu, Wang, Kang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
E-VRAG: Enhancing Long Video Understanding with Resource-Efficient Retrieval Augmented Generation
von: Xu, Zeyu, et al.
Veröffentlicht: (2025)
von: Xu, Zeyu, et al.
Veröffentlicht: (2025)
Vision-EKIPL: External Knowledge-Infused Policy Learning for Visual Reasoning
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025)
Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders
von: Zhang, Boqiang, et al.
Veröffentlicht: (2026)
von: Zhang, Boqiang, et al.
Veröffentlicht: (2026)
VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning
von: Yuan, Ruifeng, et al.
Veröffentlicht: (2025)
von: Yuan, Ruifeng, et al.
Veröffentlicht: (2025)
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning
von: Liu, Yanqing, et al.
Veröffentlicht: (2025)
von: Liu, Yanqing, et al.
Veröffentlicht: (2025)
JW-VL: A Vision-Language Model for Solar Physics
von: Shao, Mingfu, et al.
Veröffentlicht: (2026)
von: Shao, Mingfu, et al.
Veröffentlicht: (2026)
OpenVision 3: A Family of Unified Visual Encoder for Both Understanding and Generation
von: Zhang, Letian, et al.
Veröffentlicht: (2026)
von: Zhang, Letian, et al.
Veröffentlicht: (2026)
Recall: Empowering Multimodal Embedding for Edge Devices
von: Cai, Dongqi, et al.
Veröffentlicht: (2024)
von: Cai, Dongqi, et al.
Veröffentlicht: (2024)
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
von: Wang, Junyang, et al.
Veröffentlicht: (2024)
von: Wang, Junyang, et al.
Veröffentlicht: (2024)
DeepSeek-VL: Towards Real-World Vision-Language Understanding
von: Lu, Haoyu, et al.
Veröffentlicht: (2024)
von: Lu, Haoyu, et al.
Veröffentlicht: (2024)
VL-Explore: Zero-shot Vision-Language Exploration and Target Discovery by Mobile Robots
von: Zhang, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zhang, Yuxuan, et al.
Veröffentlicht: (2025)
A Cross-Hierarchical Difference Feature Fusion Network Based on Multiscale Encoder-Decoder for Hyperspectral Change Detection
von: Sheng, Mingshuai, et al.
Veröffentlicht: (2025)
von: Sheng, Mingshuai, et al.
Veröffentlicht: (2025)
CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning
von: Deria, Ankan, et al.
Veröffentlicht: (2026)
von: Deria, Ankan, et al.
Veröffentlicht: (2026)
EdgeMoE: Empowering Sparse Large Language Models on Mobile Devices
von: Yi, Rongjie, et al.
Veröffentlicht: (2023)
von: Yi, Rongjie, et al.
Veröffentlicht: (2023)
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion
von: Chen, Jiuhai, et al.
Veröffentlicht: (2024)
von: Chen, Jiuhai, et al.
Veröffentlicht: (2024)
AgriGPT-VL: Agricultural Vision-Language Understanding Suite
von: Yang, Bo, et al.
Veröffentlicht: (2025)
von: Yang, Bo, et al.
Veröffentlicht: (2025)
HyperVL: An Efficient and Dynamic Multimodal Large Language Model for Edge Devices
von: HyperAI Team, et al.
Veröffentlicht: (2025)
von: HyperAI Team, et al.
Veröffentlicht: (2025)
Towards Lightweight Graph Neural Network Search with Curriculum Graph Sparsification
von: Xie, Beini, et al.
Veröffentlicht: (2024)
von: Xie, Beini, et al.
Veröffentlicht: (2024)
VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training
von: Zhang, Jipeng, et al.
Veröffentlicht: (2025)
von: Zhang, Jipeng, et al.
Veröffentlicht: (2025)
ReDas: A Lightweight Architecture for Supporting Fine-Grained Reshaping and Multiple Dataflows on Systolic Array
von: Han, Meng, et al.
Veröffentlicht: (2023)
von: Han, Meng, et al.
Veröffentlicht: (2023)
MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices
von: Chu, Xiangxiang, et al.
Veröffentlicht: (2023)
von: Chu, Xiangxiang, et al.
Veröffentlicht: (2023)
Image Recognition with Online Lightweight Vision Transformer: A Survey
von: Zhang, Zherui, et al.
Veröffentlicht: (2025)
von: Zhang, Zherui, et al.
Veröffentlicht: (2025)
MobileMamba: Lightweight Multi-Receptive Visual Mamba Network
von: He, Haoyang, et al.
Veröffentlicht: (2024)
von: He, Haoyang, et al.
Veröffentlicht: (2024)
MobileIE: An Extremely Lightweight and Effective ConvNet for Real-Time Image Enhancement on Mobile Devices
von: Yan, Hailong, et al.
Veröffentlicht: (2025)
von: Yan, Hailong, et al.
Veröffentlicht: (2025)
TinyChemVL: Advancing Chemical Vision-Language Models via Efficient Visual Token Reduction and Complex Reaction Tasks
von: Zhao, Xuanle, et al.
Veröffentlicht: (2025)
von: Zhao, Xuanle, et al.
Veröffentlicht: (2025)
Transformers in Pseudo-Random Number Generation: A Dual Perspective on Theory and Practice
von: Li, Ran, et al.
Veröffentlicht: (2025)
von: Li, Ran, et al.
Veröffentlicht: (2025)
A Novel ViDAR Device With Visual Inertial Encoder Odometry and Reinforcement Learning-Based Active SLAM Method
von: Xin, Zhanhua, et al.
Veröffentlicht: (2025)
von: Xin, Zhanhua, et al.
Veröffentlicht: (2025)
VL-DPO: Vision-Language-Guided Finetuning for Preference-Aligned Autonomous Driving
von: Xu, Zhefan, et al.
Veröffentlicht: (2026)
von: Xu, Zhefan, et al.
Veröffentlicht: (2026)
Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision
von: Wei, Zhixiang, et al.
Veröffentlicht: (2026)
von: Wei, Zhixiang, et al.
Veröffentlicht: (2026)
ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models
von: Yi, Jingwei, et al.
Veröffentlicht: (2025)
von: Yi, Jingwei, et al.
Veröffentlicht: (2025)
Explicit Semantic-Base-Empowered Communications for 6G Mobile Networks
von: Wang, Fengyu, et al.
Veröffentlicht: (2024)
von: Wang, Fengyu, et al.
Veröffentlicht: (2024)
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking
von: Meng, Desen, et al.
Veröffentlicht: (2025)
von: Meng, Desen, et al.
Veröffentlicht: (2025)
Multiple Information Prompt Learning for Cloth-Changing Person Re-Identification
von: Wei, Shengxun, et al.
Veröffentlicht: (2024)
von: Wei, Shengxun, et al.
Veröffentlicht: (2024)
Qianfan-VL: Domain-Enhanced Universal Vision-Language Models
von: Dong, Daxiang, et al.
Veröffentlicht: (2025)
von: Dong, Daxiang, et al.
Veröffentlicht: (2025)
InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression
von: Lu, Dongchen, et al.
Veröffentlicht: (2025)
von: Lu, Dongchen, et al.
Veröffentlicht: (2025)
Code2Worlds: Empowering Coding LLMs for 4D World Generation
von: Zhang, Yi, et al.
Veröffentlicht: (2026)
von: Zhang, Yi, et al.
Veröffentlicht: (2026)
Re-Parameterization of Lightweight Transformer for On-Device Speech Emotion Recognition
von: Zhang, Zixing, et al.
Veröffentlicht: (2024)
von: Zhang, Zixing, et al.
Veröffentlicht: (2024)
A-VL: Adaptive Attention for Large Vision-Language Models
von: Zhang, Junyang, et al.
Veröffentlicht: (2024)
von: Zhang, Junyang, et al.
Veröffentlicht: (2024)
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
von: Chen, Zhe, et al.
Veröffentlicht: (2023)
von: Chen, Zhe, et al.
Veröffentlicht: (2023)
Uniform large deviations and metastability of random dynamical systems
von: Jiang, Jifa, et al.
Veröffentlicht: (2024)
von: Jiang, Jifa, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
E-VRAG: Enhancing Long Video Understanding with Resource-Efficient Retrieval Augmented Generation
von: Xu, Zeyu, et al.
Veröffentlicht: (2025) -
Vision-EKIPL: External Knowledge-Infused Policy Learning for Visual Reasoning
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025) -
Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders
von: Zhang, Boqiang, et al.
Veröffentlicht: (2026) -
VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning
von: Yuan, Ruifeng, et al.
Veröffentlicht: (2025) -
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning
von: Liu, Yanqing, et al.
Veröffentlicht: (2025)