MTA-Agent: An Open Recipe for Multimodal Deep Search Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Peng, Xiangyu, Qin, Can, Yan, An, Yang, Xinyi, Chen, Zeyuan, Xu, Ran, Wu, Chien-Sheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
UNIDOC-BENCH: A Unified Benchmark for Document-Centric Multimodal RAG
von: Peng, Xiangyu, et al.
Veröffentlicht: (2025)
von: Peng, Xiangyu, et al.
Veröffentlicht: (2025)
OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents
von: Chen, Shuang, et al.
Veröffentlicht: (2026)
von: Chen, Shuang, et al.
Veröffentlicht: (2026)
DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents
von: Wang, Shengqin, et al.
Veröffentlicht: (2026)
von: Wang, Shengqin, et al.
Veröffentlicht: (2026)
Customize Your Visual Autoregressive Recipe with Set Autoregressive Modeling
von: Liu, Wenze, et al.
Veröffentlicht: (2024)
von: Liu, Wenze, et al.
Veröffentlicht: (2024)
ProMMSearchAgent: A Generalizable Multimodal Search Agent Trained with Process-Oriented Rewards
von: Yan, Wentao, et al.
Veröffentlicht: (2026)
von: Yan, Wentao, et al.
Veröffentlicht: (2026)
VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents
von: Meng, Rui, et al.
Veröffentlicht: (2025)
von: Meng, Rui, et al.
Veröffentlicht: (2025)
PresentAgent: Multimodal Agent for Presentation Video Generation
von: Shi, Jingwei, et al.
Veröffentlicht: (2025)
von: Shi, Jingwei, et al.
Veröffentlicht: (2025)
PresentAgent-2: Towards Generalist Multimodal Presentation Agents
von: Wu, Wei, et al.
Veröffentlicht: (2026)
von: Wu, Wei, et al.
Veröffentlicht: (2026)
CT-Agent: A Multimodal-LLM Agent for 3D CT Radiology Question Answering
von: Mao, Yuren, et al.
Veröffentlicht: (2025)
von: Mao, Yuren, et al.
Veröffentlicht: (2025)
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
von: Zhu, Jinguo, et al.
Veröffentlicht: (2025)
von: Zhu, Jinguo, et al.
Veröffentlicht: (2025)
Warm Diffusion: Recipe for Blur-Noise Mixture Diffusion Models
von: Hsueh, Hao-Chien, et al.
Veröffentlicht: (2025)
von: Hsueh, Hao-Chien, et al.
Veröffentlicht: (2025)
EEA: Exploration-Exploitation Agent for Long Video Understanding
von: Yang, Te, et al.
Veröffentlicht: (2025)
von: Yang, Te, et al.
Veröffentlicht: (2025)
MMDeepResearch-Bench: A Benchmark for Multimodal Deep Research Agents
von: Huang, Peizhou, et al.
Veröffentlicht: (2026)
von: Huang, Peizhou, et al.
Veröffentlicht: (2026)
OracleAgent: A Multimodal Reasoning Agent for Oracle Bone Script Research
von: Li, Caoshuo, et al.
Veröffentlicht: (2025)
von: Li, Caoshuo, et al.
Veröffentlicht: (2025)
From Web to Pixels: Bringing Agentic Search into Visual Perception
von: Yang, Bokang, et al.
Veröffentlicht: (2026)
von: Yang, Bokang, et al.
Veröffentlicht: (2026)
RecipeGen: A Step-Aligned Multimodal Benchmark for Real-World Recipe Generation
von: Zhang, Ruoxuan, et al.
Veröffentlicht: (2025)
von: Zhang, Ruoxuan, et al.
Veröffentlicht: (2025)
DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories
von: Deng, Chenlong, et al.
Veröffentlicht: (2026)
von: Deng, Chenlong, et al.
Veröffentlicht: (2026)
SQ-LLaVA: Self-Questioning for Large Vision-Language Assistant
von: Sun, Guohao, et al.
Veröffentlicht: (2024)
von: Sun, Guohao, et al.
Veröffentlicht: (2024)
VCA: Video Curious Agent for Long Video Understanding
von: Yang, Zeyuan, et al.
Veröffentlicht: (2024)
von: Yang, Zeyuan, et al.
Veröffentlicht: (2024)
VSearcher: Long-Horizon Multimodal Search Agent via Reinforcement Learning
von: Zhang, Ruiyang, et al.
Veröffentlicht: (2026)
von: Zhang, Ruiyang, et al.
Veröffentlicht: (2026)
LayoutDETR: Detection Transformer Is a Good Multimodal Layout Designer
von: Yu, Ning, et al.
Veröffentlicht: (2022)
von: Yu, Ning, et al.
Veröffentlicht: (2022)
EmbodiedEval: Evaluate Multimodal LLMs as Embodied Agents
von: Cheng, Zhili, et al.
Veröffentlicht: (2025)
von: Cheng, Zhili, et al.
Veröffentlicht: (2025)
Vero: An Open RL Recipe for General Visual Reasoning
von: Sarch, Gabriel, et al.
Veröffentlicht: (2026)
von: Sarch, Gabriel, et al.
Veröffentlicht: (2026)
VisBrowse-Bench: Benchmarking Visual-Native Search for Multimodal Browsing Agents
von: Zhang, Zhengbo, et al.
Veröffentlicht: (2026)
von: Zhang, Zhengbo, et al.
Veröffentlicht: (2026)
VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding
von: Fan, Yue, et al.
Veröffentlicht: (2024)
von: Fan, Yue, et al.
Veröffentlicht: (2024)
How Far Are Vision-Language Models from Constructing the Real World? A Benchmark for Physical Generative Reasoning
von: Yang, Luyu, et al.
Veröffentlicht: (2026)
von: Yang, Luyu, et al.
Veröffentlicht: (2026)
BLIP3o-NEXT: Next Frontier of Native Image Generation
von: Chen, Jiuhai, et al.
Veröffentlicht: (2025)
von: Chen, Jiuhai, et al.
Veröffentlicht: (2025)
CVP: Central-Peripheral Vision-Inspired Multimodal Model for Spatial Reasoning
von: Chen, Zeyuan, et al.
Veröffentlicht: (2025)
von: Chen, Zeyuan, et al.
Veröffentlicht: (2025)
MobileFlow: A Multimodal LLM For Mobile GUI Agent
von: Nong, Songqin, et al.
Veröffentlicht: (2024)
von: Nong, Songqin, et al.
Veröffentlicht: (2024)
AgentVista: Evaluating Multimodal Agents in Ultra-Challenging Realistic Visual Scenarios
von: Su, Zhaochen, et al.
Veröffentlicht: (2026)
von: Su, Zhaochen, et al.
Veröffentlicht: (2026)
Multimodal Classification via Total Correlation Maximization
von: Yu, Feng, et al.
Veröffentlicht: (2026)
von: Yu, Feng, et al.
Veröffentlicht: (2026)
ScaleCUA: Scaling Open-Source Computer Use Agents with Cross-Platform Data
von: Liu, Zhaoyang, et al.
Veröffentlicht: (2025)
von: Liu, Zhaoyang, et al.
Veröffentlicht: (2025)
Learning to Search: A Decision-Based Agent for Knowledge-Based Visual Question Answering
von: Chen, Zhuohong, et al.
Veröffentlicht: (2026)
von: Chen, Zhuohong, et al.
Veröffentlicht: (2026)
OpenCUA: Open Foundations for Computer-Use Agents
von: Wang, Xinyuan, et al.
Veröffentlicht: (2025)
von: Wang, Xinyuan, et al.
Veröffentlicht: (2025)
BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
von: Chen, Jiuhai, et al.
Veröffentlicht: (2025)
von: Chen, Jiuhai, et al.
Veröffentlicht: (2025)
GEMS: Agent-Native Multimodal Generation with Memory and Skills
von: He, Zefeng, et al.
Veröffentlicht: (2026)
von: He, Zefeng, et al.
Veröffentlicht: (2026)
MMA: Multimodal Memory Agent
von: Lu, Yihao, et al.
Veröffentlicht: (2026)
von: Lu, Yihao, et al.
Veröffentlicht: (2026)
Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis
von: Chen, Shuang, et al.
Veröffentlicht: (2026)
von: Chen, Shuang, et al.
Veröffentlicht: (2026)
MTA: Multimodal Task Alignment for BEV Perception and Captioning
von: Ma, Yunsheng, et al.
Veröffentlicht: (2024)
von: Ma, Yunsheng, et al.
Veröffentlicht: (2024)
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
von: Wu, Size, et al.
Veröffentlicht: (2025)
von: Wu, Size, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
UNIDOC-BENCH: A Unified Benchmark for Document-Centric Multimodal RAG
von: Peng, Xiangyu, et al.
Veröffentlicht: (2025) -
OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents
von: Chen, Shuang, et al.
Veröffentlicht: (2026) -
DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents
von: Wang, Shengqin, et al.
Veröffentlicht: (2026) -
Customize Your Visual Autoregressive Recipe with Set Autoregressive Modeling
von: Liu, Wenze, et al.
Veröffentlicht: (2024) -
ProMMSearchAgent: A Generalizable Multimodal Search Agent Trained with Process-Oriented Rewards
von: Yan, Wentao, et al.
Veröffentlicht: (2026)