MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome
Fuente:
arXiv
Saved in:
| Main Authors: | Ye, Fangda, Hu, Yuxin, Zhu, Pengxiang, Li, Yibo, Jin, Ziqi, Xiao, Yao, Wang, Yibo, Wang, Lei, Zhang, Zhen, Wang, Lu, Deng, Yue, Wang, Bin, Zhang, Yifan, Su, Liangcai, Wang, Xinyu, Zhao, He, Wei, Chen, Ren, Qiang, Hooi, Bryan, Bo, An, Yan, Shuicheng, Bing, Lidong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DeepResearchEval: An Automated Framework for Deep Research Task Construction and Agentic Evaluation
by: Wang, Yibo, et al.
Published: (2026)
by: Wang, Yibo, et al.
Published: (2026)
Learning to Learn-at-Test-Time: Language Agents with Learnable Adaptation Policies
by: Lou, Zhanzhi, et al.
Published: (2026)
by: Lou, Zhanzhi, et al.
Published: (2026)
MiroFlow: Towards High-Performance and Robust Open-Source Agent Framework for General Deep Research Tasks
by: Su, Shiqian, et al.
Published: (2026)
by: Su, Shiqian, et al.
Published: (2026)
ConfTuner: Training Large Language Models to Express Their Confidence Verbally
by: Li, Yibo, et al.
Published: (2025)
by: Li, Yibo, et al.
Published: (2025)
Towards Realistic Personalization: Evaluating Long-Horizon Preference Following in Personalized User-LLM Interactions
by: Guo, Qianyun, et al.
Published: (2026)
by: Guo, Qianyun, et al.
Published: (2026)
Automated Phishing Detection Using URLs and Webpages
by: Wang, Huilin, et al.
Published: (2024)
by: Wang, Huilin, et al.
Published: (2024)
IRNN: Innovation-driven Recurrent Neural Network for Time-Series Data Modeling and Prediction
by: Zhou, Yifan, et al.
Published: (2025)
by: Zhou, Yifan, et al.
Published: (2025)
On the Role of Discreteness in Diffusion LLMs
by: Jin, Ziqi, et al.
Published: (2025)
by: Jin, Ziqi, et al.
Published: (2025)
Argus: Evidence Assembly for Scalable Deep Research Agents
by: Zhang, Zhen, et al.
Published: (2026)
by: Zhang, Zhen, et al.
Published: (2026)
Ergodicity and invariant measure approximation of the stochastic Cahn-Hilliard equation via an explicit fully discrete scheme
by: Deng, Nan, et al.
Published: (2025)
by: Deng, Nan, et al.
Published: (2025)
MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs
by: Zhao, Chenchen, et al.
Published: (2025)
by: Zhao, Chenchen, et al.
Published: (2025)
Strong convergence of a fully discrete scheme for stochastic Burgers equation with fractional-type noise
by: Wang, Yibo, et al.
Published: (2024)
by: Wang, Yibo, et al.
Published: (2024)
Approximation of the invariant measure for stochastic Allen-Cahn equation via an explicit fully discrete scheme
by: Wang, Yibo, et al.
Published: (2024)
by: Wang, Yibo, et al.
Published: (2024)
Learning Off-policy with Model-based Intrinsic Motivation For Active Online Exploration
by: Wang, Yibo, et al.
Published: (2024)
by: Wang, Yibo, et al.
Published: (2024)
Revisiting Projection-Free Online Learning with Time-Varying Constraints
by: Wang, Yibo, et al.
Published: (2025)
by: Wang, Yibo, et al.
Published: (2025)
Deep-Reporter: Deep Research for Grounded Multimodal Long-Form Generation
by: Ye, Fangda, et al.
Published: (2026)
by: Ye, Fangda, et al.
Published: (2026)
Training-Free Multi-Style Fusion Through Reference-Based Adaptive Modulation
by: Liu, Xu, et al.
Published: (2025)
by: Liu, Xu, et al.
Published: (2025)
Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient Updates
by: Li, Yibo, et al.
Published: (2026)
by: Li, Yibo, et al.
Published: (2026)
Beyond 'Aha!': Toward Systematic Meta-Abilities Alignment in Large Reasoning Models
by: Hu, Zhiyuan, et al.
Published: (2025)
by: Hu, Zhiyuan, et al.
Published: (2025)
APEX: Autonomous Policy Exploration for Self-Evolving LLM Agents
by: Li, Yibo, et al.
Published: (2026)
by: Li, Yibo, et al.
Published: (2026)
Physics-Informed High-order Graph Dynamics Identification Learning for Predicting Complex Networks Long-term Dynamics
by: Wang, Bicheng, et al.
Published: (2025)
by: Wang, Bicheng, et al.
Published: (2025)
Physics-Inspired Spatial Temporal Graph Neural Networks for Predicting Industrial Chain Resilience
by: Wang, Bicheng, et al.
Published: (2025)
by: Wang, Bicheng, et al.
Published: (2025)
Asymmetric Coordination Modulating Co Spin State for Peroxymonosulfate Activation to Accelerate 1 O 2 Generation
by: Xiuze Li, et al.
Published: (2025)
by: Xiuze Li, et al.
Published: (2025)
Mining Agent OSS: AI and LangGraph for Natural Language Underground Tunnel Design
by: Yibo, Zhang
Published: (2026)
by: Yibo, Zhang
Published: (2026)
KM-UNet KAN Mamba UNet for medical image segmentation
by: Zhang, Yibo
Published: (2025)
by: Zhang, Yibo
Published: (2025)
Holomorphic Curves in Moduli Spaces Are Quasi-Isometrically Immersed
by: Zhang, Yibo
Published: (2024)
by: Zhang, Yibo
Published: (2024)
Quasi-isometric embeddings for shrinking maps from surfaces into the moduli space
by: Zhang, Yibo
Published: (2025)
by: Zhang, Yibo
Published: (2025)
Classification of Torus Fibrations Over $S^2$ Up to Fibre Sum Stabilisation
by: Zhang, Yibo
Published: (2023)
by: Zhang, Yibo
Published: (2023)
Tricking Retrievers with Influential Tokens: An Efficient Black-Box Corpus Poisoning Attack
by: Wang, Cheng, et al.
Published: (2025)
by: Wang, Cheng, et al.
Published: (2025)
Beyond Forgetting in Continual Medical Image Segmentation: A Comprehensive Benchmark Study
by: Wang, Bomin, et al.
Published: (2026)
by: Wang, Bomin, et al.
Published: (2026)
Automating Steering for Safe Multimodal Large Language Models
by: Wu, Lyucheng, et al.
Published: (2025)
by: Wu, Lyucheng, et al.
Published: (2025)
Enhanced Low‐Temperature Photothermal Combustion of C3H8 Using Surface‐Engineered Co3O4 Nanocatalysts
by: Yajun Wang, et al.
Published: (2025)
by: Yajun Wang, et al.
Published: (2025)
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization
by: Li, Xingxuan, et al.
Published: (2025)
by: Li, Xingxuan, et al.
Published: (2025)
Theoretical study of spin-dependent transport in WSe$_2$-based vertical spin valves
by: Wang, Yibo, et al.
Published: (2026)
by: Wang, Yibo, et al.
Published: (2026)
Recognition Mechanism of RNA by TLR13: Structural Insights and Implications for Immune Activation.
by: Wang, Yibo, et al.
Published: (2025)
by: Wang, Yibo, et al.
Published: (2025)
WebAgentGuard: A Reasoning-Driven Guard Model for Detecting Prompt Injection Attacks in Web Agents
by: Chen, Yulin, et al.
Published: (2026)
by: Chen, Yulin, et al.
Published: (2026)
High Tacrolimus Intra‐Patient Variability and Adverse Outcomes in Cardiac Transplant Recipients: A Single‐Center Study in China
by: Zhenzhen Wang, et al.
Published: (2025)
by: Zhenzhen Wang, et al.
Published: (2025)
AdaptEval: A Benchmark for Evaluating Large Language Models on Code Snippet Adaptation
by: Zhang, Tanghaoran, et al.
Published: (2026)
by: Zhang, Tanghaoran, et al.
Published: (2026)
MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling
by: MiroMind Team, et al.
Published: (2025)
by: MiroMind Team, et al.
Published: (2025)
Self-Rewarding Sequential Monte Carlo for Masked Diffusion Language Models
by: Luo, Ziwei, et al.
Published: (2026)
by: Luo, Ziwei, et al.
Published: (2026)
Similar Items
-
DeepResearchEval: An Automated Framework for Deep Research Task Construction and Agentic Evaluation
by: Wang, Yibo, et al.
Published: (2026) -
Learning to Learn-at-Test-Time: Language Agents with Learnable Adaptation Policies
by: Lou, Zhanzhi, et al.
Published: (2026) -
MiroFlow: Towards High-Performance and Robust Open-Source Agent Framework for General Deep Research Tasks
by: Su, Shiqian, et al.
Published: (2026) -
ConfTuner: Training Large Language Models to Express Their Confidence Verbally
by: Li, Yibo, et al.
Published: (2025) -
Towards Realistic Personalization: Evaluating Long-Horizon Preference Following in Personalized User-LLM Interactions
by: Guo, Qianyun, et al.
Published: (2026)