Gespeichert in:
| Hauptverfasser: | Zhu, Minjie, Zhu, Yichen, Liu, Xin, Liu, Ning, Xu, Zhiyuan, Shen, Chaomin, Peng, Yaxin, Ou, Zhicai, Feng, Feifei, Tang, Jian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2403.06199 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model
von: Zhu, Yichen, et al.
Veröffentlicht: (2024)
von: Zhu, Yichen, et al.
Veröffentlicht: (2024)
ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025)
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025)
Language-Conditioned Robotic Manipulation with Fast and Slow Thinking
von: Zhu, Minjie, et al.
Veröffentlicht: (2024)
von: Zhu, Minjie, et al.
Veröffentlicht: (2024)
MMRo: Are Multimodal LLMs Eligible as the Brain for In-Home Robotics?
von: Li, Jinming, et al.
Veröffentlicht: (2024)
von: Li, Jinming, et al.
Veröffentlicht: (2024)
TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
Object-Centric Instruction Augmentation for Robotic Manipulation
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
ObjectVLA: End-to-End Open-World Object Manipulation Without Demonstration
von: Zhu, Minjie, et al.
Veröffentlicht: (2025)
von: Zhu, Minjie, et al.
Veröffentlicht: (2025)
Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
Fresh-CL: Feature Realignment through Experts on Hypersphere in Continual Learning
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025)
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025)
Scaling Diffusion Policy in Transformer to 1 Billion Parameters for Robotic Manipulation
von: Zhu, Minjie, et al.
Veröffentlicht: (2024)
von: Zhu, Minjie, et al.
Veröffentlicht: (2024)
DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control
von: Wen, Junjie, et al.
Veröffentlicht: (2025)
von: Wen, Junjie, et al.
Veröffentlicht: (2025)
EDT: An Efficient Diffusion Transformer Framework Inspired by Human-like Sketching
von: Chen, Xinwang, et al.
Veröffentlicht: (2024)
von: Chen, Xinwang, et al.
Veröffentlicht: (2024)
Efficient Feature Fusion for UAV Object Detection
von: Wang, Xudong, et al.
Veröffentlicht: (2025)
von: Wang, Xudong, et al.
Veröffentlicht: (2025)
PointVLA: Injecting the 3D World into Vision-Language-Action Models
von: Li, Chengmeng, et al.
Veröffentlicht: (2025)
von: Li, Chengmeng, et al.
Veröffentlicht: (2025)
Visual Robotic Manipulation with Depth-Aware Pretraining
von: Wang, Wanying, et al.
Veröffentlicht: (2024)
von: Wang, Wanying, et al.
Veröffentlicht: (2024)
dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought
von: Wen, Junjie, et al.
Veröffentlicht: (2025)
von: Wen, Junjie, et al.
Veröffentlicht: (2025)
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025)
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025)
Retrieval-Augmented Embodied Agents
von: Zhu, Yichen, et al.
Veröffentlicht: (2024)
von: Zhu, Yichen, et al.
Veröffentlicht: (2024)
WorldEval: World Model as Real-World Robot Policies Evaluator
von: Li, Yaxuan, et al.
Veröffentlicht: (2025)
von: Li, Yaxuan, et al.
Veröffentlicht: (2025)
Safety of Multimodal Large Language Models on Images and Texts
von: Liu, Xin, et al.
Veröffentlicht: (2024)
von: Liu, Xin, et al.
Veröffentlicht: (2024)
CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance
von: Li, Jinming, et al.
Veröffentlicht: (2024)
von: Li, Jinming, et al.
Veröffentlicht: (2024)
Res-Bench: Benchmarking the Robustness of Multimodal Large Language Models to Dynamic Resolution Input
von: Li, Chenxu, et al.
Veröffentlicht: (2025)
von: Li, Chenxu, et al.
Veröffentlicht: (2025)
Look Carefully: Adaptive Visual Reinforcements in Multimodal Large Language Models for Hallucination Mitigation
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models
von: Liu, Xin, et al.
Veröffentlicht: (2023)
von: Liu, Xin, et al.
Veröffentlicht: (2023)
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents
von: Luo, Yaxin, et al.
Veröffentlicht: (2025)
von: Luo, Yaxin, et al.
Veröffentlicht: (2025)
Dynamic Multimodal Prototype Learning in Vision-Language Models
von: Zhu, Xingyu, et al.
Veröffentlicht: (2025)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2025)
Student-Oriented Teacher Knowledge Refinement for Knowledge Distillation
von: Shen, Chaomin, et al.
Veröffentlicht: (2024)
von: Shen, Chaomin, et al.
Veröffentlicht: (2024)
ERGeoBench:A Comprehensive Benchmark for Embodied Reasoning and Geo-localization in Multimodal Large Language Models
von: Xue, Kaiwen, et al.
Veröffentlicht: (2026)
von: Xue, Kaiwen, et al.
Veröffentlicht: (2026)
CFBenchmark-MM: Chinese Financial Assistant Benchmark for Multimodal Large Language Model
von: Li, Jiangtong, et al.
Veröffentlicht: (2025)
von: Li, Jiangtong, et al.
Veröffentlicht: (2025)
Face-Human-Bench: A Comprehensive Benchmark of Face and Human Understanding for Multi-modal Assistants
von: Qin, Lixiong, et al.
Veröffentlicht: (2025)
von: Qin, Lixiong, et al.
Veröffentlicht: (2025)
Language-Pretraining-Induced Bias: A Strong Foundation for General Vision Tasks
von: Luo, Yaxin, et al.
Veröffentlicht: (2026)
von: Luo, Yaxin, et al.
Veröffentlicht: (2026)
A Comprehensive Survey of Large Language Models and Multimodal Large Language Models in Medicine
von: Xiao, Hanguang, et al.
Veröffentlicht: (2024)
von: Xiao, Hanguang, et al.
Veröffentlicht: (2024)
DocRetriever: A Plug-and-Play Framework for Multimodal Document Retrieval with Comprehensive Benchmark
von: Hu, Ruofan, et al.
Veröffentlicht: (2026)
von: Hu, Ruofan, et al.
Veröffentlicht: (2026)
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models
von: Jiang, Jiachen, et al.
Veröffentlicht: (2025)
von: Jiang, Jiachen, et al.
Veröffentlicht: (2025)
Distilling Mathematical Reasoning Capabilities into Small Language Models
von: Zhu, Xunyu, et al.
Veröffentlicht: (2024)
von: Zhu, Xunyu, et al.
Veröffentlicht: (2024)
Language-Instructed Reasoning for Group Activity Detection via Multimodal Large Language Model
von: Peng, Jihua, et al.
Veröffentlicht: (2025)
von: Peng, Jihua, et al.
Veröffentlicht: (2025)
Mitigating Hallucinations in Large Vision-Language Models without Performance Degradation
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
CODIS: Benchmarking Context-Dependent Visual Comprehension for Multimodal Large Language Models
von: Luo, Fuwen, et al.
Veröffentlicht: (2024)
von: Luo, Fuwen, et al.
Veröffentlicht: (2024)
Towards Harmless Multimodal Assistants with Blind Preference Optimization
von: Li, Yongqi, et al.
Veröffentlicht: (2025)
von: Li, Yongqi, et al.
Veröffentlicht: (2025)
LLaSA: Large Language and E-Commerce Shopping Assistant
von: Zhang, Shuo, et al.
Veröffentlicht: (2024)
von: Zhang, Shuo, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model
von: Zhu, Yichen, et al.
Veröffentlicht: (2024) -
ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025) -
Language-Conditioned Robotic Manipulation with Fast and Slow Thinking
von: Zhu, Minjie, et al.
Veröffentlicht: (2024) -
MMRo: Are Multimodal LLMs Eligible as the Brain for In-Home Robotics?
von: Li, Jinming, et al.
Veröffentlicht: (2024) -
TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
von: Wen, Junjie, et al.
Veröffentlicht: (2024)