AMIGO: Agentic Multi-Image Grounding Oracle Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Min, Mahjoubfar, Ata |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CARE: Towards Clinical Accountability in Multi-Modal Medical Reasoning with an Evidence-Grounded Agentic Framework
by: Du, Yuexi, et al.
Published: (2026)
by: Du, Yuexi, et al.
Published: (2026)
Evaluating Self-Supervised Learning in Medical Imaging: A Benchmark for Robustness, Generalizability, and Multi-Domain Impact
by: Bundele, Valay, et al.
Published: (2024)
by: Bundele, Valay, et al.
Published: (2024)
Active Policy Improvement from Multiple Black-box Oracles
by: Liu, Xuefeng, et al.
Published: (2023)
by: Liu, Xuefeng, et al.
Published: (2023)
Characterization of Multi-Model Agentic AI Systems on General Tasks via Trace-Driven Simulation
by: Kim, Donghwan, et al.
Published: (2026)
by: Kim, Donghwan, et al.
Published: (2026)
Efficient Learning of Fuzzy Logic Systems for Large-Scale Data Using Deep Learning
by: Koklu, Ata, et al.
Published: (2024)
by: Koklu, Ata, et al.
Published: (2024)
Zadeh's Type-2 Fuzzy Logic Systems: Precision and High-Quality Prediction Intervals
by: Guven, Yusuf, et al.
Published: (2024)
by: Guven, Yusuf, et al.
Published: (2024)
Enhancing Interval Type-2 Fuzzy Logic Systems: Learning for Precision and Prediction Intervals
by: Koklu, Ata, et al.
Published: (2024)
by: Koklu, Ata, et al.
Published: (2024)
GCA Framework: A GCC Countries-Grounded Dataset and Agentic Pipeline for Climate Decision Support
by: Sheikh, Muhammad Umer, et al.
Published: (2026)
by: Sheikh, Muhammad Umer, et al.
Published: (2026)
Is Efficient PAC Learning Possible with an Oracle That Responds 'Yes' or 'No'?
by: Daskalakis, Constantinos, et al.
Published: (2024)
by: Daskalakis, Constantinos, et al.
Published: (2024)
Auto-bidding in real-time auctions via Oracle Imitation Learning (OIL)
by: Chiappa, Alberto Silvio, et al.
Published: (2024)
by: Chiappa, Alberto Silvio, et al.
Published: (2024)
Multi-Agent Debate: A Unified Agentic Framework for Tabular Anomaly Detection
by: Wang, Pinqiao, et al.
Published: (2026)
by: Wang, Pinqiao, et al.
Published: (2026)
Improving Autoregressive Training with Dynamic Oracles
by: Yang, Jianing, et al.
Published: (2024)
by: Yang, Jianing, et al.
Published: (2024)
Oracle-Guided Soft Shielding for Safe Move Prediction in Chess
by: Rajendran, Prajit T, et al.
Published: (2026)
by: Rajendran, Prajit T, et al.
Published: (2026)
PepBenchmark: A Standardized Benchmark for Peptide Machine Learning
by: Zhang, Jiahui, et al.
Published: (2026)
by: Zhang, Jiahui, et al.
Published: (2026)
MultiTab: A Comprehensive Benchmark Suite for Multi-Dimensional Evaluation in Tabular Domains
by: Lee, Kyungeun, et al.
Published: (2025)
by: Lee, Kyungeun, et al.
Published: (2025)
SLEA-RL: Step-Level Experience Augmented Reinforcement Learning for Multi-Turn Agentic Training
by: Wang, Prince Zizhuang, et al.
Published: (2026)
by: Wang, Prince Zizhuang, et al.
Published: (2026)
Contextual Linear Bandits under Noisy Features: Towards Bayesian Oracles
by: Kim, Jung-hun, et al.
Published: (2017)
by: Kim, Jung-hun, et al.
Published: (2017)
Agentic Uncertainty Reveals Agentic Overconfidence
by: Kaddour, Jean, et al.
Published: (2026)
by: Kaddour, Jean, et al.
Published: (2026)
BoSS: A Best-of-Strategies Selector as an Oracle for Deep Active Learning
by: Huseljic, Denis, et al.
Published: (2026)
by: Huseljic, Denis, et al.
Published: (2026)
Auto-Discovery-Bench: Diagnosing Structured State Tracking in Oracle-Guided Discovery
by: Chen, Tingting, et al.
Published: (2025)
by: Chen, Tingting, et al.
Published: (2025)
Sample and Oracle Efficient Reinforcement Learning for MDPs with Linearly-Realizable Value Functions
by: Mhammedi, Zakaria
Published: (2024)
by: Mhammedi, Zakaria
Published: (2024)
Human-Guided Agentic AI for Multimodal Clinical Prediction: Lessons from the AgentDS Healthcare Benchmark
by: Pulavarthy, Lalitha Pranathi, et al.
Published: (2026)
by: Pulavarthy, Lalitha Pranathi, et al.
Published: (2026)
Multi-View Encoders for Performance Prediction in LLM-Based Agentic Workflows
by: Trirat, Patara, et al.
Published: (2025)
by: Trirat, Patara, et al.
Published: (2025)
TDPNavigator-Placer: Thermal- and Wirelength-Aware Chiplet Placement in 2.5D Systems Through Multi-Agent Reinforcement Learning
by: Hou, Yubo, et al.
Published: (2026)
by: Hou, Yubo, et al.
Published: (2026)
HeuriGym: An Agentic Benchmark for LLM-Crafted Heuristics in Combinatorial Optimization
by: Chen, Hongzheng, et al.
Published: (2025)
by: Chen, Hongzheng, et al.
Published: (2025)
Zeroth-Order Optimization Meets Human Feedback: Provable Learning via Ranking Oracles
by: Tang, Zhiwei, et al.
Published: (2023)
by: Tang, Zhiwei, et al.
Published: (2023)
Harnessing Agentic Evolution
by: Zhang, Jiayi, et al.
Published: (2026)
by: Zhang, Jiayi, et al.
Published: (2026)
A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning
by: Wang, Ruiyi, et al.
Published: (2025)
by: Wang, Ruiyi, et al.
Published: (2025)
Towards Robust Agentic CUDA Kernel Benchmarking, Verification, and Optimization
by: Lange, Robert Tjarko, et al.
Published: (2025)
by: Lange, Robert Tjarko, et al.
Published: (2025)
Supplement Generation Training for Enhancing Agentic Task Performance
by: Cho, Young Min, et al.
Published: (2026)
by: Cho, Young Min, et al.
Published: (2026)
Understanding LLM-Driven Test Oracle Generation
by: Bodicoat, Adam, et al.
Published: (2026)
by: Bodicoat, Adam, et al.
Published: (2026)
EVGeoQA: Benchmarking LLMs on Dynamic, Multi-Objective Geo-Spatial Exploration
by: Wu, Jianfei, et al.
Published: (2026)
by: Wu, Jianfei, et al.
Published: (2026)
Fast Benchmarking of Asynchronous Multi-Fidelity Optimization on Zero-Cost Benchmarks
by: Watanabe, Shuhei, et al.
Published: (2024)
by: Watanabe, Shuhei, et al.
Published: (2024)
Data Curation Through the Lens of Spectral Dynamics: Static Limits, Dynamic Acceleration, and Practical Oracles
by: Zhang, Yizhou, et al.
Published: (2025)
by: Zhang, Yizhou, et al.
Published: (2025)
When Do Multi-Agent Systems Outperform? Analysing the Learning Efficiency of Agentic Systems
by: Su, Junwei, et al.
Published: (2026)
by: Su, Junwei, et al.
Published: (2026)
AgenticPay: A Multi-Agent LLM Negotiation System for Buyer-Seller Transactions
by: Liu, Xianyang, et al.
Published: (2026)
by: Liu, Xianyang, et al.
Published: (2026)
Agentics 2.0: Logical Transduction Algebra for Agentic Data Workflows
by: Gliozzo, Alfio Massimiliano, et al.
Published: (2026)
by: Gliozzo, Alfio Massimiliano, et al.
Published: (2026)
Membership Testing in Markov Equivalence Classes via Independence Query Oracles
by: Zhang, Jiaqi, et al.
Published: (2024)
by: Zhang, Jiaqi, et al.
Published: (2024)
Sanity Checks for Agentic Data Science
by: Rewolinski, Zachary T., et al.
Published: (2026)
by: Rewolinski, Zachary T., et al.
Published: (2026)
Agentic Unlearning: When LLM Agent Meets Machine Unlearning
by: Wang, Bin, et al.
Published: (2026)
by: Wang, Bin, et al.
Published: (2026)
Similar Items
-
CARE: Towards Clinical Accountability in Multi-Modal Medical Reasoning with an Evidence-Grounded Agentic Framework
by: Du, Yuexi, et al.
Published: (2026) -
Evaluating Self-Supervised Learning in Medical Imaging: A Benchmark for Robustness, Generalizability, and Multi-Domain Impact
by: Bundele, Valay, et al.
Published: (2024) -
Active Policy Improvement from Multiple Black-box Oracles
by: Liu, Xuefeng, et al.
Published: (2023) -
Characterization of Multi-Model Agentic AI Systems on General Tasks via Trace-Driven Simulation
by: Kim, Donghwan, et al.
Published: (2026) -
Efficient Learning of Fuzzy Logic Systems for Large-Scale Data Using Deep Learning
by: Koklu, Ata, et al.
Published: (2024)