AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Junyang, Wang, Yuhang, Xu, Guohai, Zhang, Jing, Gu, Yukai, Jia, Haitao, Wang, Jiaqi, Xu, Haiyang, Yan, Ming, Zhang, Ji, Sang, Jitao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration
di: Wang, Junyang, et al.
Pubblicazione: (2024)
di: Wang, Junyang, et al.
Pubblicazione: (2024)
Mobile-Agent-V: A Video-Guided Approach for Effortless and Efficient Operational Knowledge Injection in Mobile Automation
di: Wang, Junyang, et al.
Pubblicazione: (2025)
di: Wang, Junyang, et al.
Pubblicazione: (2025)
Mobile-Agent-V: A Video-Guided Approach for Effortless and Efficient Operational Knowledge Injection in Mobile Automation
di: Wang, Junyang, et al.
Pubblicazione: (2025)
di: Wang, Junyang, et al.
Pubblicazione: (2025)
STAMP: Training Explicit Memory for Mobile GUI Agents in Controllable and Scalable Virtual Environments
di: Wang, Junyang, et al.
Pubblicazione: (2026)
di: Wang, Junyang, et al.
Pubblicazione: (2026)
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
di: Wang, Junyang, et al.
Pubblicazione: (2024)
di: Wang, Junyang, et al.
Pubblicazione: (2024)
FairCLIP: Social Bias Elimination based on Attribute Prototype Learning and Representation Neutralization
di: Wang, Junyang, et al.
Pubblicazione: (2022)
di: Wang, Junyang, et al.
Pubblicazione: (2022)
VaLiD: Mitigating the Hallucination of Large Vision Language Models by Visual Layer Fusion Contrastive Decoding
di: Wang, Jiaqi, et al.
Pubblicazione: (2024)
di: Wang, Jiaqi, et al.
Pubblicazione: (2024)
KG-FPQ: Evaluating Factuality Hallucination in LLMs with Knowledge Graph-based False Premise Questions
di: Zhu, Yanxu, et al.
Pubblicazione: (2024)
di: Zhu, Yanxu, et al.
Pubblicazione: (2024)
AIGCs Confuse AI Too: Investigating and Explaining Synthetic Image-induced Hallucinations in Large Vision-Language Models
di: Gao, Yifei, et al.
Pubblicazione: (2024)
di: Gao, Yifei, et al.
Pubblicazione: (2024)
Mobile-Agent-E: Self-Evolving Mobile Assistant for Complex Tasks
di: Wang, Zhenhailong, et al.
Pubblicazione: (2025)
di: Wang, Zhenhailong, et al.
Pubblicazione: (2025)
ODE: Open-Set Evaluation of Hallucinations in Multimodal Large Language Models
di: Tu, Yahan, et al.
Pubblicazione: (2024)
di: Tu, Yahan, et al.
Pubblicazione: (2024)
FREAK: A Fine-grained Hallucination Evaluation Benchmark for Advanced MLLMs
di: Yin, Zhihan, et al.
Pubblicazione: (2026)
di: Yin, Zhihan, et al.
Pubblicazione: (2026)
Don't Command, Cultivate: An Exploratory Study of System-2 Alignment
di: Wang, Yuhang, et al.
Pubblicazione: (2024)
di: Wang, Yuhang, et al.
Pubblicazione: (2024)
CDEval: A Benchmark for Measuring the Cultural Dimensions of Large Language Models
di: Wang, Yuhang, et al.
Pubblicazione: (2023)
di: Wang, Yuhang, et al.
Pubblicazione: (2023)
Self-Guided Defense: Adaptive Safety Alignment for Reasoning Models via Synthesized Guidelines
di: Wang, Yuhang, et al.
Pubblicazione: (2025)
di: Wang, Yuhang, et al.
Pubblicazione: (2025)
Interpreting and Mitigating Hallucination in MLLMs through Multi-agent Debate
di: Lin, Zheng, et al.
Pubblicazione: (2024)
di: Lin, Zheng, et al.
Pubblicazione: (2024)
Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language Models
di: Jiang, Chaoya, et al.
Pubblicazione: (2024)
di: Jiang, Chaoya, et al.
Pubblicazione: (2024)
GOBench: Benchmarking Geometric Optics Generation and Understanding of MLLMs
di: Zhu, Xiaorong, et al.
Pubblicazione: (2025)
di: Zhu, Xiaorong, et al.
Pubblicazione: (2025)
DenoiseRep: Denoising Model for Representation Learning
di: Xu, Zhengrui, et al.
Pubblicazione: (2024)
di: Xu, Zhengrui, et al.
Pubblicazione: (2024)
Improving Weak-to-Strong Generalization with Scalable Oversight and Ensemble Learning
di: Sang, Jitao, et al.
Pubblicazione: (2024)
di: Sang, Jitao, et al.
Pubblicazione: (2024)
Linking Perception, Confidence and Accuracy in MLLMs
di: Du, Yuetian, et al.
Pubblicazione: (2026)
di: Du, Yuetian, et al.
Pubblicazione: (2026)
HalluLens: LLM Hallucination Benchmark
di: Bang, Yejin, et al.
Pubblicazione: (2025)
di: Bang, Yejin, et al.
Pubblicazione: (2025)
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
di: Ye, Qinghao, et al.
Pubblicazione: (2023)
di: Ye, Qinghao, et al.
Pubblicazione: (2023)
C-FAITH: A Chinese Fine-Grained Benchmark for Automated Hallucination Evaluation
di: Zhang, Xu, et al.
Pubblicazione: (2025)
di: Zhang, Xu, et al.
Pubblicazione: (2025)
Affordance Benchmark for MLLMs
di: Wang, Junying, et al.
Pubblicazione: (2025)
di: Wang, Junying, et al.
Pubblicazione: (2025)
Mitigating Object Hallucinations in MLLMs via Multi-Frequency Perturbations
di: Li, Shuo, et al.
Pubblicazione: (2025)
di: Li, Shuo, et al.
Pubblicazione: (2025)
Seek-and-Solve: Benchmarking MLLMs for Visual Clue-Driven Reasoning in Daily Scenarios
di: Li, Xiaomin, et al.
Pubblicazione: (2026)
di: Li, Xiaomin, et al.
Pubblicazione: (2026)
TCC-Bench: Benchmarking the Traditional Chinese Culture Understanding Capabilities of MLLMs
di: Xu, Pengju, et al.
Pubblicazione: (2025)
di: Xu, Pengju, et al.
Pubblicazione: (2025)
TARS: MinMax Token-Adaptive Preference Strategy for Hallucination Reduction in MLLMs
di: Zhang, Kejia, et al.
Pubblicazione: (2025)
di: Zhang, Kejia, et al.
Pubblicazione: (2025)
PC-Agent: A Hierarchical Multi-Agent Collaboration Framework for Complex Task Automation on PC
di: Liu, Haowei, et al.
Pubblicazione: (2025)
di: Liu, Haowei, et al.
Pubblicazione: (2025)
COPO: Causal-Oriented Policy Optimization for Hallucinations of MLLMs
di: Guo, Peizheng, et al.
Pubblicazione: (2025)
di: Guo, Peizheng, et al.
Pubblicazione: (2025)
Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture
di: Zhang, Wanyue, et al.
Pubblicazione: (2025)
di: Zhang, Wanyue, et al.
Pubblicazione: (2025)
Hallucination Augmented Contrastive Learning for Multimodal Large Language Model
di: Jiang, Chaoya, et al.
Pubblicazione: (2023)
di: Jiang, Chaoya, et al.
Pubblicazione: (2023)
PROST-LLM: Progressively Enhancing the Speech-to-Speech Translation Capability in LLMs
di: Xu, Jing, et al.
Pubblicazione: (2026)
di: Xu, Jing, et al.
Pubblicazione: (2026)
FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow
di: Sun, Haoyu, et al.
Pubblicazione: (2025)
di: Sun, Haoyu, et al.
Pubblicazione: (2025)
The Mirage of Performance Gains: Why Contrastive Decoding Fails to Mitigate Object Hallucinations in MLLMs?
di: Yin, Hao, et al.
Pubblicazione: (2025)
di: Yin, Hao, et al.
Pubblicazione: (2025)
A Probabilistic Inference Scaling Theory for LLM Self-Correction
di: Yang, Zhe, et al.
Pubblicazione: (2025)
di: Yang, Zhe, et al.
Pubblicazione: (2025)
SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs
di: Wang, Siting, et al.
Pubblicazione: (2025)
di: Wang, Siting, et al.
Pubblicazione: (2025)
Positional Failures in Long-Context LLMs: A Blind Spot in Reasoning Benchmarks
di: Zhang, Chuyifei, et al.
Pubblicazione: (2026)
di: Zhang, Chuyifei, et al.
Pubblicazione: (2026)
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs
di: Ding, Xuanwen, et al.
Pubblicazione: (2025)
di: Ding, Xuanwen, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration
di: Wang, Junyang, et al.
Pubblicazione: (2024) -
Mobile-Agent-V: A Video-Guided Approach for Effortless and Efficient Operational Knowledge Injection in Mobile Automation
di: Wang, Junyang, et al.
Pubblicazione: (2025) -
Mobile-Agent-V: A Video-Guided Approach for Effortless and Efficient Operational Knowledge Injection in Mobile Automation
di: Wang, Junyang, et al.
Pubblicazione: (2025) -
STAMP: Training Explicit Memory for Mobile GUI Agents in Controllable and Scalable Virtual Environments
di: Wang, Junyang, et al.
Pubblicazione: (2026) -
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
di: Wang, Junyang, et al.
Pubblicazione: (2024)