EH-Benchmark Ophthalmic Hallucination Benchmark and Agent-Driven Top-Down Traceable Reasoning Workflow
Fuente:
arXiv
Saved in:
| Main Authors: | Pan, Xiaoyu, Bai, Yang, Zou, Ke, Zhou, Yang, Zhou, Jun, Fu, Huazhu, Tham, Yih-Chung, Liu, Yong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CataractSurg-80K: Knowledge-Driven Benchmarking for Structured Reasoning in Ophthalmic Surgery Planning
by: Meng, Yang, et al.
Published: (2025)
by: Meng, Yang, et al.
Published: (2025)
Benchmarking Agentic Workflow Generation
by: Qiao, Shuofei, et al.
Published: (2024)
by: Qiao, Shuofei, et al.
Published: (2024)
An Agentic System for Rare Disease Diagnosis with Traceable Reasoning
by: Zhao, Weike, et al.
Published: (2025)
by: Zhao, Weike, et al.
Published: (2025)
SEVADE: Self-Evolving Multi-Agent Analysis with Decoupled Evaluation for Hallucination-Resistant Irony Detection
by: Liu, Ziqi, et al.
Published: (2025)
by: Liu, Ziqi, et al.
Published: (2025)
MAS-Orchestra: Understanding and Improving Multi-Agent Reasoning Through Holistic Orchestration and Controlled Benchmarks
by: Ke, Zixuan, et al.
Published: (2026)
by: Ke, Zixuan, et al.
Published: (2026)
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind
by: Lupu, Andrei, et al.
Published: (2025)
by: Lupu, Andrei, et al.
Published: (2025)
Benchmarking LLMs' Swarm intelligence
by: Ruan, Kai, et al.
Published: (2025)
by: Ruan, Kai, et al.
Published: (2025)
ORCA: An Agentic Reasoning Framework for Hallucination and Adversarial Robustness in Vision-Language Models
by: Yu, Chung-En Johnny, et al.
Published: (2025)
by: Yu, Chung-En Johnny, et al.
Published: (2025)
Multi-agent Undercover Gaming: Hallucination Removal via Counterfactual Test for Multimodal Reasoning
by: Liang, Dayong, et al.
Published: (2025)
by: Liang, Dayong, et al.
Published: (2025)
DiscoVerse: Multi-Agent Pharmaceutical Co-Scientist for Traceable Drug Discovery and Reverse Translation
by: Zheng, Xiaochen, et al.
Published: (2025)
by: Zheng, Xiaochen, et al.
Published: (2025)
LingxiDiagBench: A Multi-Agent Framework for Benchmarking LLMs in Chinese Psychiatric Consultation and Diagnosis
by: Xu, Shihao, et al.
Published: (2026)
by: Xu, Shihao, et al.
Published: (2026)
Hydra: An Agentic Reasoning Approach for Enhancing Adversarial Robustness and Mitigating Hallucinations in Vision-Language Models
by: Chung-En, et al.
Published: (2025)
by: Chung-En, et al.
Published: (2025)
Finch: Benchmarking Finance & Accounting across Spreadsheet-Centric Enterprise Workflows
by: Dong, Haoyu, et al.
Published: (2025)
by: Dong, Haoyu, et al.
Published: (2025)
GNNs as Predictors of Agentic Workflow Performances
by: Zhang, Yuanshuo, et al.
Published: (2025)
by: Zhang, Yuanshuo, et al.
Published: (2025)
What Makes Good Collaborative Views? Contrastive Mutual Information Maximization for Multi-Agent Perception
by: Su, Wanfang, et al.
Published: (2024)
by: Su, Wanfang, et al.
Published: (2024)
RadAgents: Multimodal Agentic Reasoning for Chest X-ray Interpretation with Radiologist-like Workflows
by: Zhang, Kai, et al.
Published: (2025)
by: Zhang, Kai, et al.
Published: (2025)
Collab-Overcooked: Benchmarking and Evaluating Large Language Models as Collaborative Agents
by: Sun, Haochen, et al.
Published: (2025)
by: Sun, Haochen, et al.
Published: (2025)
Automating Structural Engineering Workflows with Large Language Model Agents
by: Liang, Haoran, et al.
Published: (2025)
by: Liang, Haoran, et al.
Published: (2025)
EvoFlow: Evolving Diverse Agentic Workflows On The Fly
by: Zhang, Guibin, et al.
Published: (2025)
by: Zhang, Guibin, et al.
Published: (2025)
PolicySimEval: A Benchmark for Evaluating Policy Outcomes through Agent-Based Simulation
by: Kang, Jiaju, et al.
Published: (2025)
by: Kang, Jiaju, et al.
Published: (2025)
Communication to Completion: Modeling Collaborative Workflows with Intelligent Multi-Agent Communication
by: Lu, Yiming, et al.
Published: (2025)
by: Lu, Yiming, et al.
Published: (2025)
Browsing Lost Unformed Recollections: A Benchmark for Tip-of-the-Tongue Search and Reasoning
by: CH-Wang, Sky, et al.
Published: (2025)
by: CH-Wang, Sky, et al.
Published: (2025)
How do Role Models Shape Collective Morality? Exemplar-Driven Moral Learning in Multi-Agent Simulation
by: Liao, Junjie, et al.
Published: (2026)
by: Liao, Junjie, et al.
Published: (2026)
FURINA: A Fully Customizable Role-Playing Benchmark via Scalable Multi-Agent Collaboration Pipeline
by: Wu, Haotian, et al.
Published: (2025)
by: Wu, Haotian, et al.
Published: (2025)
Beyond Static Tools: Test-Time Tool Evolution for Scientific Reasoning
by: Lu, Jiaxuan, et al.
Published: (2026)
by: Lu, Jiaxuan, et al.
Published: (2026)
Internet of Agentic AI: Incentive-Compatible Distributed Teaming and Workflow
by: Yang, Ya-Ting, et al.
Published: (2026)
by: Yang, Ya-Ting, et al.
Published: (2026)
HSCodeComp: A Realistic and Expert-level Benchmark for Deep Search Agents in Hierarchical Rule Application
by: Yang, Yiqian, et al.
Published: (2025)
by: Yang, Yiqian, et al.
Published: (2025)
Cooperative Profiles Predict Multi-Agent LLM Team Performance in AI for Science Workflows
by: Kumar, Shivani, et al.
Published: (2026)
by: Kumar, Shivani, et al.
Published: (2026)
Are LLM Agents the New RPA? A Comparative Study with RPA Across Enterprise Workflows
by: Průcha, Petr, et al.
Published: (2025)
by: Průcha, Petr, et al.
Published: (2025)
DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows
by: Gao, Yuxuan, et al.
Published: (2026)
by: Gao, Yuxuan, et al.
Published: (2026)
It's a TRAP! Task-Redirecting Agent Persuasion Benchmark for Web Agents
by: Korgul, Karolina, et al.
Published: (2025)
by: Korgul, Karolina, et al.
Published: (2025)
CONCAT: Consensus- and Confidence-Driven Ad Hoc Teaming for Efficient LLM-Based Multi-Agent Systems
by: Ma, Ziyang, et al.
Published: (2026)
by: Ma, Ziyang, et al.
Published: (2026)
KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows
by: Pan, Zaifeng, et al.
Published: (2025)
by: Pan, Zaifeng, et al.
Published: (2025)
DSBC : Data Science task Benchmarking with Context engineering
by: Kadiyala, Ram Mohan Rao, et al.
Published: (2025)
by: Kadiyala, Ram Mohan Rao, et al.
Published: (2025)
Beyond the All-in-One Agent: Benchmarking Role-Specialized Multi-Agent Collaboration in Enterprise Workflows
by: Yu, Tao, et al.
Published: (2026)
by: Yu, Tao, et al.
Published: (2026)
MedPriv-Bench: Benchmarking the Privacy-Utility Trade-off of Large Language Models in Medical Open-End Question Answering
by: Guan, Shaowei, et al.
Published: (2026)
by: Guan, Shaowei, et al.
Published: (2026)
EvoRoute: Experience-Driven Self-Routing LLM Agent Systems
by: Zhang, Guibin, et al.
Published: (2026)
by: Zhang, Guibin, et al.
Published: (2026)
Rethinking the Value of Multi-Agent Workflow: A Strong Single Agent Baseline
by: Xu, Jiawei, et al.
Published: (2026)
by: Xu, Jiawei, et al.
Published: (2026)
OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning
by: Lu, Pan, et al.
Published: (2025)
by: Lu, Pan, et al.
Published: (2025)
Empowering Scientific Workflows with Federated Agents
by: Kamatar, Alok, et al.
Published: (2025)
by: Kamatar, Alok, et al.
Published: (2025)
Similar Items
-
CataractSurg-80K: Knowledge-Driven Benchmarking for Structured Reasoning in Ophthalmic Surgery Planning
by: Meng, Yang, et al.
Published: (2025) -
Benchmarking Agentic Workflow Generation
by: Qiao, Shuofei, et al.
Published: (2024) -
An Agentic System for Rare Disease Diagnosis with Traceable Reasoning
by: Zhao, Weike, et al.
Published: (2025) -
SEVADE: Self-Evolving Multi-Agent Analysis with Decoupled Evaluation for Hallucination-Resistant Irony Detection
by: Liu, Ziqi, et al.
Published: (2025) -
MAS-Orchestra: Understanding and Improving Multi-Agent Reasoning Through Holistic Orchestration and Controlled Benchmarks
by: Ke, Zixuan, et al.
Published: (2026)