Dynamic Orchestration of Multi-Agent System for Real-World Multi-Image Agricultural VQA
Fuente:
arXiv
Saved in:
| Main Authors: | Ke, Yan, Yu, Xin, Du, Heming, Chapman, Scott, Huang, Helen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DatasetAgent: A Novel Multi-Agent System for Auto-Constructing Datasets from Real-World Images
by: Sun, Haoran, et al.
Published: (2025)
by: Sun, Haoran, et al.
Published: (2025)
When One Moment Isn't Enough: Multi-Moment Retrieval with Cross-Moment Interactions
by: Cao, Zhuo, et al.
Published: (2025)
by: Cao, Zhuo, et al.
Published: (2025)
Maestro: Self-Improving Text-to-Image Generation via Agent Orchestration
by: Wan, Xingchen, et al.
Published: (2025)
by: Wan, Xingchen, et al.
Published: (2025)
Refine and Align: Confidence Calibration through Multi-Agent Interaction in VQA
by: Pandey, Ayush, et al.
Published: (2025)
by: Pandey, Ayush, et al.
Published: (2025)
SCOPE: Structured Decomposition and Conditional Skill Orchestration for Complex Image Generation
by: Ren, Tianfei, et al.
Published: (2026)
by: Ren, Tianfei, et al.
Published: (2026)
IMAGAgent: Orchestrating Multi-Turn Image Editing via Constraint-Aware Planning and Reflection
by: Shen, Fei, et al.
Published: (2026)
by: Shen, Fei, et al.
Published: (2026)
Sparse Mixture-of-Experts for Multi-Channel Imaging: Are All Channel Interactions Required?
by: Yun, Sukwon, et al.
Published: (2025)
by: Yun, Sukwon, et al.
Published: (2025)
M$^3$-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question Answering
by: Ma, Jiatong, et al.
Published: (2026)
by: Ma, Jiatong, et al.
Published: (2026)
MATA: A Trainable Hierarchical Automaton System for Multi-Agent Visual Reasoning
by: Cai, Zhixi, et al.
Published: (2026)
by: Cai, Zhixi, et al.
Published: (2026)
Multi-Agent Image Restoration
by: Jiang, Xu, et al.
Published: (2025)
by: Jiang, Xu, et al.
Published: (2025)
MTabVQA: Evaluating Multi-Tabular Reasoning of Language Models in Visual Space
by: Singh, Anshul, et al.
Published: (2025)
by: Singh, Anshul, et al.
Published: (2025)
AgroVG: A Large-Scale Multi-Source Benchmark for Agricultural Visual Grounding
by: Li, Haocheng, et al.
Published: (2026)
by: Li, Haocheng, et al.
Published: (2026)
UniShield: An Adaptive Multi-Agent Framework for Unified Forgery Image Detection and Localization
by: Huang, Qing, et al.
Published: (2025)
by: Huang, Qing, et al.
Published: (2025)
PulseMind: A Multi-Modal Medical Model for Real-World Clinical Diagnosis
by: Xu, Jiao, et al.
Published: (2026)
by: Xu, Jiao, et al.
Published: (2026)
Bird-SR: Bidirectional Reward-Guided Diffusion for Real-World Image Super-Resolution
by: Fan, Zihao, et al.
Published: (2026)
by: Fan, Zihao, et al.
Published: (2026)
MSRAMIE: Multimodal Structured Reasoning Agent for Multi-instruction Image Editing
by: Qiu, Zhaoyuan, et al.
Published: (2026)
by: Qiu, Zhaoyuan, et al.
Published: (2026)
CSAOT: Cooperative Multi-Agent System for Active Object Tracking
by: Nguyen, Hy, et al.
Published: (2025)
by: Nguyen, Hy, et al.
Published: (2025)
PASSION: Towards Effective Incomplete Multi-Modal Medical Image Segmentation with Imbalanced Missing Rates
by: Shi, Junjie, et al.
Published: (2024)
by: Shi, Junjie, et al.
Published: (2024)
Realism Control One-step Diffusion for Real-World Image Super-Resolution
by: Wu, Zongliang, et al.
Published: (2025)
by: Wu, Zongliang, et al.
Published: (2025)
Symphony: A Cognitively-Inspired Multi-Agent System for Long-Video Understanding
by: Yan, Haiyang, et al.
Published: (2026)
by: Yan, Haiyang, et al.
Published: (2026)
COMBO: Compositional World Models for Embodied Multi-Agent Cooperation
by: Zhang, Hongxin, et al.
Published: (2024)
by: Zhang, Hongxin, et al.
Published: (2024)
Kvasir-VQA: A Text-Image Pair GI Tract Dataset
by: Gautam, Sushant, et al.
Published: (2024)
by: Gautam, Sushant, et al.
Published: (2024)
GenPilot: A Multi-Agent System for Test-Time Prompt Optimization in Image Generation
by: Ye, Wen, et al.
Published: (2025)
by: Ye, Wen, et al.
Published: (2025)
Multi-Agent VQA: Exploring Multi-Agent Foundation Models in Zero-Shot Visual Question Answering
by: Jiang, Bowen, et al.
Published: (2024)
by: Jiang, Bowen, et al.
Published: (2024)
World-To-Image: Grounding Text-to-Image Generation with Agent-Driven World Knowledge
by: Son, Moo Hyun, et al.
Published: (2025)
by: Son, Moo Hyun, et al.
Published: (2025)
ShareVerse: Multi-Agent Consistent Video Generation for Shared World Modeling
by: Zhu, Jiayi, et al.
Published: (2026)
by: Zhu, Jiayi, et al.
Published: (2026)
Gather and Trace: Rethinking Video TextVQA from an Instance-oriented Perspective
by: Zhang, Yan, et al.
Published: (2025)
by: Zhang, Yan, et al.
Published: (2025)
MUSES: 3D-Controllable Image Generation via Multi-Modal Agent Collaboration
by: Ding, Yanbo, et al.
Published: (2024)
by: Ding, Yanbo, et al.
Published: (2024)
PathNavigate: A Training-Free Pathology Agent with Surprise-Guided Scan and Shared Slide Memory for Whole-Slide Image VQA
by: Yang, Chunze, et al.
Published: (2026)
by: Yang, Chunze, et al.
Published: (2026)
Dual Causal Inference: Integrating Backdoor Adjustment and Instrumental Variable Learning for Medical VQA
by: Xu, Zibo, et al.
Published: (2026)
by: Xu, Zibo, et al.
Published: (2026)
Box-QAymo: Box-Referring VQA Dataset for Autonomous Driving
by: Etchegaray, Djamahl, et al.
Published: (2025)
by: Etchegaray, Djamahl, et al.
Published: (2025)
ChromouVQA: Benchmarking Vision-Language Models under Chromatic Camouflaged Images
by: Zhang, Yunfei, et al.
Published: (2025)
by: Zhang, Yunfei, et al.
Published: (2025)
WSI-VQA: Interpreting Whole Slide Images by Generative Visual Question Answering
by: Chen, Pingyi, et al.
Published: (2024)
by: Chen, Pingyi, et al.
Published: (2024)
AgroNVILA: Perception-Reasoning Decoupling for Multi-view Agricultural Multimodal Large Language Models
by: Zhang, Jiarui, et al.
Published: (2026)
by: Zhang, Jiarui, et al.
Published: (2026)
LaRE$^2$: Latent Reconstruction Error Based Method for Diffusion-Generated Image Detection
by: Luo, Yunpeng, et al.
Published: (2024)
by: Luo, Yunpeng, et al.
Published: (2024)
OmniScience: A Large-scale Multi-modal Dataset for Scientific Image Understanding
by: Tao, Haoyi, et al.
Published: (2026)
by: Tao, Haoyi, et al.
Published: (2026)
A Denoising Framework for Real-World Ultra-Low-Dose Lung CT Images Based on an Image Purification Strategy
by: Gong, Guoliang, et al.
Published: (2025)
by: Gong, Guoliang, et al.
Published: (2025)
CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving
by: Huang, Minqing, et al.
Published: (2026)
by: Huang, Minqing, et al.
Published: (2026)
Multi-Knowledge-oriented Nighttime Haze Imaging Enhancer for Vision-driven Intelligent Systems
by: Chen, Ai, et al.
Published: (2025)
by: Chen, Ai, et al.
Published: (2025)
Advances and Innovations in the Multi-Agent Robotic System (MARS) Challenge
by: Kang, Li, et al.
Published: (2026)
by: Kang, Li, et al.
Published: (2026)
Similar Items
-
DatasetAgent: A Novel Multi-Agent System for Auto-Constructing Datasets from Real-World Images
by: Sun, Haoran, et al.
Published: (2025) -
When One Moment Isn't Enough: Multi-Moment Retrieval with Cross-Moment Interactions
by: Cao, Zhuo, et al.
Published: (2025) -
Maestro: Self-Improving Text-to-Image Generation via Agent Orchestration
by: Wan, Xingchen, et al.
Published: (2025) -
Refine and Align: Confidence Calibration through Multi-Agent Interaction in VQA
by: Pandey, Ayush, et al.
Published: (2025) -
SCOPE: Structured Decomposition and Conditional Skill Orchestration for Complex Image Generation
by: Ren, Tianfei, et al.
Published: (2026)