MIRAGE-Bench: LLM Agent is Hallucinating and Where to Find Them
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Weichen, Sun, Yiyou, Huang, Pohao, Pu, Jiayue, Lin, Heyue, Song, Dawn |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HALoGEN: Fantastic LLM Hallucinations and Where to Find Them
by: Ravichander, Abhilasha, et al.
Published: (2025)
by: Ravichander, Abhilasha, et al.
Published: (2025)
Fantastic Pretraining Optimizers and Where to Find Them
by: Wen, Kaiyue, et al.
Published: (2025)
by: Wen, Kaiyue, et al.
Published: (2025)
Low Rank Gradients and Where to Find Them
by: Sonthalia, Rishi, et al.
Published: (2025)
by: Sonthalia, Rishi, et al.
Published: (2025)
The Long-Horizon Task Mirage? Diagnosing Where and Why Agentic Systems Break
by: Wang, Xinyu Jessica, et al.
Published: (2026)
by: Wang, Xinyu Jessica, et al.
Published: (2026)
Fantastic Bugs and Where to Find Them in AI Benchmarks
by: Truong, Sang, et al.
Published: (2025)
by: Truong, Sang, et al.
Published: (2025)
HomeSafe-Bench: Evaluating Vision-Language Models on Unsafe Action Detection for Embodied Agents in Household Scenarios
by: Pu, Jiayue, et al.
Published: (2026)
by: Pu, Jiayue, et al.
Published: (2026)
MIRAGE-Bench: Automatic Multilingual Benchmark Arena for Retrieval-Augmented Generation Systems
by: Thakur, Nandan, et al.
Published: (2024)
by: Thakur, Nandan, et al.
Published: (2024)
Fantastic Reasoning Behaviors and Where to Find Them: Unsupervised Discovery of the Reasoning Process
by: Zhang, Zhenyu, et al.
Published: (2025)
by: Zhang, Zhenyu, et al.
Published: (2025)
Strategy Executability in Mathematical Reasoning: Leveraging Human-Model Differences for Effective Guidance
by: Liang, Weida, et al.
Published: (2026)
by: Liang, Weida, et al.
Published: (2026)
Fantastic Targets for Concept Erasure in Diffusion Models and Where To Find Them
by: Bui, Anh, et al.
Published: (2025)
by: Bui, Anh, et al.
Published: (2025)
Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents
by: Li, Xu, et al.
Published: (2026)
by: Li, Xu, et al.
Published: (2026)
MIRAGE: The Illusion of Visual Understanding
by: Asadi, Mohammad, et al.
Published: (2026)
by: Asadi, Mohammad, et al.
Published: (2026)
LifelongAgentBench: Evaluating LLM Agents as Lifelong Learners
by: Zheng, Junhao, et al.
Published: (2025)
by: Zheng, Junhao, et al.
Published: (2025)
OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization
by: Sun, Yiyou, et al.
Published: (2025)
by: Sun, Yiyou, et al.
Published: (2025)
Can LLMs Ask Good Questions?
by: Zhang, Yueheng, et al.
Published: (2025)
by: Zhang, Yueheng, et al.
Published: (2025)
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence
by: Liu, Chonghan, et al.
Published: (2025)
by: Liu, Chonghan, et al.
Published: (2025)
Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?
by: Sun, Yiyou, et al.
Published: (2025)
by: Sun, Yiyou, et al.
Published: (2025)
RL Grokking Recipe: How Does RL Unlock and Transfer New Algorithms in LLMs?
by: Sun, Yiyou, et al.
Published: (2025)
by: Sun, Yiyou, et al.
Published: (2025)
ESG-Bench: Benchmarking Long-Context ESG Reports for Hallucination Mitigation
by: Sun, Siqi, et al.
Published: (2026)
by: Sun, Siqi, et al.
Published: (2026)
Golden Layers and Where to Find Them: Improved Knowledge Editing for Large Language Models Via Layer Gradient Analysis
by: Datta, Shrestha, et al.
Published: (2026)
by: Datta, Shrestha, et al.
Published: (2026)
Calibrated Language Models and How to Find Them with Label Smoothing
by: Huang, Jerry, et al.
Published: (2025)
by: Huang, Jerry, et al.
Published: (2025)
PTCG-Bench: Can LLM Agents Master Pokémon Trading Card Game?
by: Hua, Dongdong, et al.
Published: (2026)
by: Hua, Dongdong, et al.
Published: (2026)
What Cohort INRs Encode and Where to Freeze Them
by: Sideri-Lampretsa, Vasiliki, et al.
Published: (2026)
by: Sideri-Lampretsa, Vasiliki, et al.
Published: (2026)
AgenticPay: A Multi-Agent LLM Negotiation System for Buyer-Seller Transactions
by: Liu, Xianyang, et al.
Published: (2026)
by: Liu, Xianyang, et al.
Published: (2026)
OptArgus: A Multi-Agent System to Detect Hallucinations in LLM-based Optimization Modeling
by: Li, Zhong, et al.
Published: (2026)
by: Li, Zhong, et al.
Published: (2026)
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
Where LLM Agents Fail and How They can Learn From Failures
by: Zhu, Kunlun, et al.
Published: (2025)
by: Zhu, Kunlun, et al.
Published: (2025)
InnovatorBench: Evaluating Agents' Ability to Conduct Innovative LLM Research
by: Wu, Yunze, et al.
Published: (2025)
by: Wu, Yunze, et al.
Published: (2025)
LLM-based Agents Suffer from Hallucinations: A Survey of Taxonomy, Methods, and Directions
by: Lin, Xixun, et al.
Published: (2025)
by: Lin, Xixun, et al.
Published: (2025)
Where to Search: Measure the Prior-Structured Search Space of LLM Agents
by: Song, Zhuo-Yang
Published: (2025)
by: Song, Zhuo-Yang
Published: (2025)
Multimodal Representation Learning using Adaptive Graph Construction
by: Huang, Weichen
Published: (2024)
by: Huang, Weichen
Published: (2024)
Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
by: Zhang, Hanrong, et al.
Published: (2024)
by: Zhang, Hanrong, et al.
Published: (2024)
MIRAGE: Towards AI-Generated Image Detection in the Wild
by: Xia, Cheng, et al.
Published: (2025)
by: Xia, Cheng, et al.
Published: (2025)
AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents
by: Guo, Zhengkang, et al.
Published: (2026)
by: Guo, Zhengkang, et al.
Published: (2026)
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
by: Wang, Yidong, et al.
Published: (2025)
by: Wang, Yidong, et al.
Published: (2025)
MIRAGE: Context-Aware Prompt Injection against Mobile GUI Agents via User-Generated Content
by: Guo, Ruoqi, et al.
Published: (2026)
by: Guo, Ruoqi, et al.
Published: (2026)
Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination
by: Zheng, Haojie, et al.
Published: (2024)
by: Zheng, Haojie, et al.
Published: (2024)
NewtonBench: Benchmarking Generalizable Scientific Law Discovery in LLM Agents
by: Zheng, Tianshi, et al.
Published: (2025)
by: Zheng, Tianshi, et al.
Published: (2025)
SkillGenBench: Benchmarking Skill Generation Pipelines for LLM Agents
by: Zhou, Yifan, et al.
Published: (2026)
by: Zhou, Yifan, et al.
Published: (2026)
ProcCtrlBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents
by: He, Jiawei, et al.
Published: (2026)
by: He, Jiawei, et al.
Published: (2026)
Similar Items
-
HALoGEN: Fantastic LLM Hallucinations and Where to Find Them
by: Ravichander, Abhilasha, et al.
Published: (2025) -
Fantastic Pretraining Optimizers and Where to Find Them
by: Wen, Kaiyue, et al.
Published: (2025) -
Low Rank Gradients and Where to Find Them
by: Sonthalia, Rishi, et al.
Published: (2025) -
The Long-Horizon Task Mirage? Diagnosing Where and Why Agentic Systems Break
by: Wang, Xinyu Jessica, et al.
Published: (2026) -
Fantastic Bugs and Where to Find Them in AI Benchmarks
by: Truong, Sang, et al.
Published: (2025)