MiroFlow: Towards High-Performance and Robust Open-Source Agent Framework for General Deep Research Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Su, Shiqian, Xing, Sen, Dong, Xuan, Zhong, Muyan, Wang, Bin, Zhu, Xizhou, Chen, Yuntao, Wang, Wenhai, Deng, Yue, Zhu, Pengxiang, Liu, Ziyuan, Li, Tiantong, Yu, Jiaheng, Chen, Zhe, Bing, Lidong, Dai, Jifeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling
von: MiroMind Team, et al.
Veröffentlicht: (2025)
von: MiroMind Team, et al.
Veröffentlicht: (2025)
VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks
von: Wu, Jiannan, et al.
Veröffentlicht: (2024)
von: Wu, Jiannan, et al.
Veröffentlicht: (2024)
The All-Seeing Project V2: Towards General Relation Comprehension of the Open World
von: Wang, Weiyun, et al.
Veröffentlicht: (2024)
von: Wang, Weiyun, et al.
Veröffentlicht: (2024)
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
von: Chen, Zhe, et al.
Veröffentlicht: (2023)
von: Chen, Zhe, et al.
Veröffentlicht: (2023)
CoMemo: LVLMs Need Image Context with Image Memory
von: Liu, Shi, et al.
Veröffentlicht: (2025)
von: Liu, Shi, et al.
Veröffentlicht: (2025)
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding
von: Tao, Chenxin, et al.
Veröffentlicht: (2024)
von: Tao, Chenxin, et al.
Veröffentlicht: (2024)
PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models
von: Yang, Chenyu, et al.
Veröffentlicht: (2024)
von: Yang, Chenyu, et al.
Veröffentlicht: (2024)
MuLan: Adapting Multilingual Diffusion Models for Hundreds of Languages with Negligible Cost
von: Xing, Sen, et al.
Veröffentlicht: (2024)
von: Xing, Sen, et al.
Veröffentlicht: (2024)
MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
von: Tian, Changyao, et al.
Veröffentlicht: (2024)
von: Tian, Changyao, et al.
Veröffentlicht: (2024)
Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures
von: Duan, Yuchen, et al.
Veröffentlicht: (2024)
von: Duan, Yuchen, et al.
Veröffentlicht: (2024)
Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
von: Wang, Weiyun, et al.
Veröffentlicht: (2024)
von: Wang, Weiyun, et al.
Veröffentlicht: (2024)
ZeroGUI: Automating Online GUI Learning at Zero Human Cost
von: Yang, Chenyu, et al.
Veröffentlicht: (2025)
von: Yang, Chenyu, et al.
Veröffentlicht: (2025)
Bounding Box Stability against Feature Dropout Reflects Detector Generalization across Environments
von: Yang, Yang, et al.
Veröffentlicht: (2024)
von: Yang, Yang, et al.
Veröffentlicht: (2024)
VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models
von: Xu, Weiye, et al.
Veröffentlicht: (2025)
von: Xu, Weiye, et al.
Veröffentlicht: (2025)
Needle In A Multimodal Haystack
von: Wang, Weiyun, et al.
Veröffentlicht: (2024)
von: Wang, Weiyun, et al.
Veröffentlicht: (2024)
Efficient Deformable ConvNets: Rethinking Dynamic and Sparse Operator for Vision Applications
von: Xiong, Yuwen, et al.
Veröffentlicht: (2024)
von: Xiong, Yuwen, et al.
Veröffentlicht: (2024)
MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome
von: Ye, Fangda, et al.
Veröffentlicht: (2026)
von: Ye, Fangda, et al.
Veröffentlicht: (2026)
MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity
von: Liu, Yangzhou, et al.
Veröffentlicht: (2024)
von: Liu, Yangzhou, et al.
Veröffentlicht: (2024)
Vision Model Pre-training on Interleaved Image-Text Data via Latent Compression Learning
von: Yang, Chenyu, et al.
Veröffentlicht: (2024)
von: Yang, Chenyu, et al.
Veröffentlicht: (2024)
V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding
von: Ge, Junqi, et al.
Veröffentlicht: (2024)
von: Ge, Junqi, et al.
Veröffentlicht: (2024)
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization
von: Li, Xingxuan, et al.
Veröffentlicht: (2025)
von: Li, Xingxuan, et al.
Veröffentlicht: (2025)
Demystify Transformers & Convolutions in Modern Image Deep Networks
von: Hu, Xiaowei, et al.
Veröffentlicht: (2022)
von: Hu, Xiaowei, et al.
Veröffentlicht: (2022)
VisualPRM: An Effective Process Reward Model for Multimodal Reasoning
von: Wang, Weiyun, et al.
Veröffentlicht: (2025)
von: Wang, Weiyun, et al.
Veröffentlicht: (2025)
Learning 1D Causal Visual Representation with De-focus Attention Networks
von: Tao, Chenxin, et al.
Veröffentlicht: (2024)
von: Tao, Chenxin, et al.
Veröffentlicht: (2024)
Multi-Agent Tool-Integrated Policy Optimization
von: Mo, Zhanfeng, et al.
Veröffentlicht: (2025)
von: Mo, Zhanfeng, et al.
Veröffentlicht: (2025)
DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving
von: Cui, Erfei, et al.
Veröffentlicht: (2023)
von: Cui, Erfei, et al.
Veröffentlicht: (2023)
Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance
von: Gao, Zhangwei, et al.
Veröffentlicht: (2024)
von: Gao, Zhangwei, et al.
Veröffentlicht: (2024)
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
von: Luo, Gen, et al.
Veröffentlicht: (2025)
von: Luo, Gen, et al.
Veröffentlicht: (2025)
Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy
von: Hou, Zhi, et al.
Veröffentlicht: (2025)
von: Hou, Zhi, et al.
Veröffentlicht: (2025)
Evolving Prompts In-Context: An Open-ended, Self-replicating Perspective
von: Wang, Jianyu, et al.
Veröffentlicht: (2025)
von: Wang, Jianyu, et al.
Veröffentlicht: (2025)
NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
von: Tian, Changyao, et al.
Veröffentlicht: (2025)
von: Tian, Changyao, et al.
Veröffentlicht: (2025)
Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training
von: Luo, Gen, et al.
Veröffentlicht: (2024)
von: Luo, Gen, et al.
Veröffentlicht: (2024)
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
von: Chen, Zhe, et al.
Veröffentlicht: (2024)
von: Chen, Zhe, et al.
Veröffentlicht: (2024)
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning
von: Xing, Zhenghao, et al.
Veröffentlicht: (2025)
von: Xing, Zhenghao, et al.
Veröffentlicht: (2025)
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
von: Luo, Gen, et al.
Veröffentlicht: (2025)
von: Luo, Gen, et al.
Veröffentlicht: (2025)
Distributed Invariant Kalman Filter for Cooperative Localization using Matrix Lie Groups
von: Zhou, Yizhi, et al.
Veröffentlicht: (2024)
von: Zhou, Yizhi, et al.
Veröffentlicht: (2024)
M-Loss: Quantifying Model Merging Compatibility with Limited Unlabeled Data
von: Wang, Tiantong, et al.
Veröffentlicht: (2026)
von: Wang, Tiantong, et al.
Veröffentlicht: (2026)
Parameter-Inverted Image Pyramid Networks
von: Zhu, Xizhou, et al.
Veröffentlicht: (2024)
von: Zhu, Xizhou, et al.
Veröffentlicht: (2024)
A Unified Graph Language Model for Multi-Domain Multi-Task Graph Alignment Instruction Tuning
von: Chen, Haibo, et al.
Veröffentlicht: (2026)
von: Chen, Haibo, et al.
Veröffentlicht: (2026)
How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
von: Chen, Zhe, et al.
Veröffentlicht: (2024)
von: Chen, Zhe, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling
von: MiroMind Team, et al.
Veröffentlicht: (2025) -
VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks
von: Wu, Jiannan, et al.
Veröffentlicht: (2024) -
The All-Seeing Project V2: Towards General Relation Comprehension of the Open World
von: Wang, Weiyun, et al.
Veröffentlicht: (2024) -
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
von: Chen, Zhe, et al.
Veröffentlicht: (2023) -
CoMemo: LVLMs Need Image Context with Image Memory
von: Liu, Shi, et al.
Veröffentlicht: (2025)