MM-Zero: Self-Evolving Multi-Model Vision Language Models From Zero Data
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Zongxia, Du, Hongyang, Huang, Chengsong, Wu, Xiyang, Yu, Lantao, He, Yicheng, Xie, Jing, Wu, Xiaomin, Liu, Zhichao, Zhang, Jiarui, Liu, Fuxiao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges
von: Li, Zongxia, et al.
Veröffentlicht: (2025)
von: Li, Zongxia, et al.
Veröffentlicht: (2025)
VisPlay: Self-Evolving Vision-Language Models from Images
von: He, Yicheng, et al.
Veröffentlicht: (2025)
von: He, Yicheng, et al.
Veröffentlicht: (2025)
R-Zero: Self-Evolving Reasoning LLM from Zero Data
von: Huang, Chengsong, et al.
Veröffentlicht: (2025)
von: Huang, Chengsong, et al.
Veröffentlicht: (2025)
G-Zero: Self-Play for Open-Ended Generation from Zero Data
von: Huang, Chengsong, et al.
Veröffentlicht: (2026)
von: Huang, Chengsong, et al.
Veröffentlicht: (2026)
Self-Rewarding Vision-Language Model via Reasoning Decomposition
von: Li, Zongxia, et al.
Veröffentlicht: (2025)
von: Li, Zongxia, et al.
Veröffentlicht: (2025)
VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding
von: Li, Zongxia, et al.
Veröffentlicht: (2025)
von: Li, Zongxia, et al.
Veröffentlicht: (2025)
First Frame Is the Place to Go for Video Content Customization
von: Chen, Jingxi, et al.
Veröffentlicht: (2025)
von: Chen, Jingxi, et al.
Veröffentlicht: (2025)
Graph-of-Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills
von: Liu, Dawei, et al.
Veröffentlicht: (2026)
von: Liu, Dawei, et al.
Veröffentlicht: (2026)
HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
von: Guan, Tianrui, et al.
Veröffentlicht: (2023)
von: Guan, Tianrui, et al.
Veröffentlicht: (2023)
Active Zero: Self-Evolving Vision-Language Models through Active Environment Exploration
von: He, Jinghan, et al.
Veröffentlicht: (2026)
von: He, Jinghan, et al.
Veröffentlicht: (2026)
SABER: A Stealthy Agentic Black-Box Attack Framework for Vision-Language-Action Models
von: Wu, Xiyang, et al.
Veröffentlicht: (2026)
von: Wu, Xiyang, et al.
Veröffentlicht: (2026)
Dr. Zero: Self-Evolving Search Agents without Training Data
von: Yue, Zhenrui, et al.
Veröffentlicht: (2026)
von: Yue, Zhenrui, et al.
Veröffentlicht: (2026)
StreamVoice+: Evolving into End-to-end Streaming Zero-shot Voice Conversion
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
Absolute Zero: Reinforced Self-play Reasoning with Zero Data
von: Zhao, Andrew, et al.
Veröffentlicht: (2025)
von: Zhao, Andrew, et al.
Veröffentlicht: (2025)
Agent0: Unleashing Self-Evolving Agents from Zero Data via Tool-Integrated Reasoning
von: Xia, Peng, et al.
Veröffentlicht: (2025)
von: Xia, Peng, et al.
Veröffentlicht: (2025)
Towards Understanding In-Context Learning with Contrastive Demonstrations and Saliency Maps
von: Liu, Fuxiao, et al.
Veröffentlicht: (2023)
von: Liu, Fuxiao, et al.
Veröffentlicht: (2023)
MASS: Motion-Aware Spatial-Temporal Grounding for Physics Reasoning and Comprehension in Vision-Language Models
von: Wu, Xiyang, et al.
Veröffentlicht: (2025)
von: Wu, Xiyang, et al.
Veröffentlicht: (2025)
Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks
von: Wu, Xiyang, et al.
Veröffentlicht: (2026)
von: Wu, Xiyang, et al.
Veröffentlicht: (2026)
Learning A Zero-shot Occupancy Network from Vision Foundation Models via Self-supervised Adaptation
von: Lin, Sihao, et al.
Veröffentlicht: (2025)
von: Lin, Sihao, et al.
Veröffentlicht: (2025)
Self-supervised Dynamic Heterogeneous Degradation Modeling for Unified Zero-Shot Image Restoration
von: Hu, XiaoWan, et al.
Veröffentlicht: (2026)
von: Hu, XiaoWan, et al.
Veröffentlicht: (2026)
zkLLM: Zero Knowledge Proofs for Large Language Models
von: Sun, Haochen, et al.
Veröffentlicht: (2024)
von: Sun, Haochen, et al.
Veröffentlicht: (2024)
Evolving Prompt Adaptation for Vision-Language Models
von: Zhang, Enming, et al.
Veröffentlicht: (2026)
von: Zhang, Enming, et al.
Veröffentlicht: (2026)
From Zero to Hero: Detecting Leaked Data through Synthetic Data Injection and Model Querying
von: Wu, Biao, et al.
Veröffentlicht: (2023)
von: Wu, Biao, et al.
Veröffentlicht: (2023)
MM-InstructEval: Zero-Shot Evaluation of (Multimodal) Large Language Models on Multimodal Reasoning Tasks
von: Yang, Xiaocui, et al.
Veröffentlicht: (2024)
von: Yang, Xiaocui, et al.
Veröffentlicht: (2024)
Hierarchical Micro-Segmentations for Zero-Trust Services via Large Language Model (LLM)-enhanced Graph Diffusion
von: Liu, Yinqiu, et al.
Veröffentlicht: (2024)
von: Liu, Yinqiu, et al.
Veröffentlicht: (2024)
Pretrained Optimization Model for Zero-Shot Black Box Optimization
von: Li, Xiaobin, et al.
Veröffentlicht: (2024)
von: Li, Xiaobin, et al.
Veröffentlicht: (2024)
Interleaved Current‐Fed Boost Converter With Output Voltage Self‐Balancing for Photovoltaics MVDC Integration
von: Xiaoquan Zhu, et al.
Veröffentlicht: (2025)
von: Xiaoquan Zhu, et al.
Veröffentlicht: (2025)
Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision
von: He, Yinghui, et al.
Veröffentlicht: (2026)
von: He, Yinghui, et al.
Veröffentlicht: (2026)
Self-Improving for Zero-Shot Named Entity Recognition with Large Language Models
von: Xie, Tingyu, et al.
Veröffentlicht: (2023)
von: Xie, Tingyu, et al.
Veröffentlicht: (2023)
TTCS: Test-Time Curriculum Synthesis for Self-Evolving
von: Yang, Chengyi, et al.
Veröffentlicht: (2026)
von: Yang, Chengyi, et al.
Veröffentlicht: (2026)
EvoVLA: Self-Evolving Vision-Language-Action Model
von: Liu, Zeting, et al.
Veröffentlicht: (2025)
von: Liu, Zeting, et al.
Veröffentlicht: (2025)
Cross-modal Proxy Evolving for OOD Detection with Vision-Language Models
von: Tang, Hao, et al.
Veröffentlicht: (2026)
von: Tang, Hao, et al.
Veröffentlicht: (2026)
Vision-Zero: Scalable VLM Self-Improvement via Strategic Gamified Self-Play
von: Wang, Qinsi, et al.
Veröffentlicht: (2025)
von: Wang, Qinsi, et al.
Veröffentlicht: (2025)
Pruner-Zero: Evolving Symbolic Pruning Metric from scratch for Large Language Models
von: Dong, Peijie, et al.
Veröffentlicht: (2024)
von: Dong, Peijie, et al.
Veröffentlicht: (2024)
Attention-guided Self-reflection for Zero-shot Hallucination Detection in Large Language Models
von: Liu, Qiang, et al.
Veröffentlicht: (2025)
von: Liu, Qiang, et al.
Veröffentlicht: (2025)
GOFA: A Generative One-For-All Model for Joint Graph Language Modeling
von: Kong, Lecheng, et al.
Veröffentlicht: (2024)
von: Kong, Lecheng, et al.
Veröffentlicht: (2024)
Open-Source Image Editing Models Are Zero-Shot Vision Learners
von: Liu, Wei, et al.
Veröffentlicht: (2026)
von: Liu, Wei, et al.
Veröffentlicht: (2026)
EDDA: A Encoder-Decoder Data Augmentation Framework for Zero-Shot Stance Detection
von: Ding, Daijun, et al.
Veröffentlicht: (2024)
von: Ding, Daijun, et al.
Veröffentlicht: (2024)
Tool-R0: Self-Evolving LLM Agents for Tool-Learning from Zero Data
von: Acikgoz, Emre Can, et al.
Veröffentlicht: (2026)
von: Acikgoz, Emre Can, et al.
Veröffentlicht: (2026)
Utilizing Large Language Models for Zero-Shot Medical Ontology Extension from Clinical Notes
von: Wu, Guanchen, et al.
Veröffentlicht: (2025)
von: Wu, Guanchen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges
von: Li, Zongxia, et al.
Veröffentlicht: (2025) -
VisPlay: Self-Evolving Vision-Language Models from Images
von: He, Yicheng, et al.
Veröffentlicht: (2025) -
R-Zero: Self-Evolving Reasoning LLM from Zero Data
von: Huang, Chengsong, et al.
Veröffentlicht: (2025) -
G-Zero: Self-Play for Open-Ended Generation from Zero Data
von: Huang, Chengsong, et al.
Veröffentlicht: (2026) -
Self-Rewarding Vision-Language Model via Reasoning Decomposition
von: Li, Zongxia, et al.
Veröffentlicht: (2025)