Salvato in:
| Autori principali: | Liu, Xuannan, Yang, Xiao, Li, Zekun, Li, Peipei, He, Ran |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2601.06818 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
T^2Agent A Tool-augmented Multimodal Misinformation Detection Agent with Monte Carlo Tree Search
di: Cui, Xing, et al.
Pubblicazione: (2025)
di: Cui, Xing, et al.
Pubblicazione: (2025)
Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs
di: Liu, Xuannan, et al.
Pubblicazione: (2025)
di: Liu, Xuannan, et al.
Pubblicazione: (2025)
HalluLens: LLM Hallucination Benchmark
di: Bang, Yejin, et al.
Pubblicazione: (2025)
di: Bang, Yejin, et al.
Pubblicazione: (2025)
MMFakeBench: A Mixed-Source Multimodal Misinformation Detection Benchmark for LVLMs
di: Liu, Xuannan, et al.
Pubblicazione: (2024)
di: Liu, Xuannan, et al.
Pubblicazione: (2024)
Jailbreak Attacks and Defenses against Multimodal Generative Models: A Survey
di: Liu, Xuannan, et al.
Pubblicazione: (2024)
di: Liu, Xuannan, et al.
Pubblicazione: (2024)
HalluEntity: Benchmarking and Understanding Entity-Level Hallucination Detection
di: Yeh, Min-Hsuan, et al.
Pubblicazione: (2025)
di: Yeh, Min-Hsuan, et al.
Pubblicazione: (2025)
HalluCana: Fixing LLM Hallucination with A Canary Lookahead
di: Li, Tianyi, et al.
Pubblicazione: (2024)
di: Li, Tianyi, et al.
Pubblicazione: (2024)
MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents
di: Zhang, Dongsen, et al.
Pubblicazione: (2025)
di: Zhang, Dongsen, et al.
Pubblicazione: (2025)
FKA-Owl: Advancing Multimodal Fake News Detection through Knowledge-Augmented LVLMs
di: Liu, Xuannan, et al.
Pubblicazione: (2024)
di: Liu, Xuannan, et al.
Pubblicazione: (2024)
HalluVerse25: Fine-grained Multilingual Benchmark Dataset for LLM Hallucinations
di: Abdaljalil, Samir, et al.
Pubblicazione: (2025)
di: Abdaljalil, Samir, et al.
Pubblicazione: (2025)
HalluDial: A Large-Scale Benchmark for Automatic Dialogue-Level Hallucination Evaluation
di: Luo, Wen, et al.
Pubblicazione: (2024)
di: Luo, Wen, et al.
Pubblicazione: (2024)
HalluCounter: Reference-free LLM Hallucination Detection in the Wild!
di: Urlana, Ashok, et al.
Pubblicazione: (2025)
di: Urlana, Ashok, et al.
Pubblicazione: (2025)
HalluHard: A Hard Multi-Turn Hallucination Benchmark
di: Fan, Dongyang, et al.
Pubblicazione: (2026)
di: Fan, Dongyang, et al.
Pubblicazione: (2026)
HalluScore: Large Language Model Hallucination Question Answering Benchmark
di: Alansari, Aisha, et al.
Pubblicazione: (2026)
di: Alansari, Aisha, et al.
Pubblicazione: (2026)
HalluDetect: Detecting, Mitigating, and Benchmarking Hallucinations in Conversational Systems in the Legal Domain
di: Anaokar, Spandan, et al.
Pubblicazione: (2025)
di: Anaokar, Spandan, et al.
Pubblicazione: (2025)
PerHalluEval: Persian Hallucination Evaluation Benchmark for Large Language Models
di: Hosseini, Mohammad, et al.
Pubblicazione: (2025)
di: Hosseini, Mohammad, et al.
Pubblicazione: (2025)
AVFakeBench: A Comprehensive Audio-Video Forgery Detection Benchmark for AV-LMMs
di: Xia, Shuhan, et al.
Pubblicazione: (2025)
di: Xia, Shuhan, et al.
Pubblicazione: (2025)
HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models
di: Liu, Emmy, et al.
Pubblicazione: (2026)
di: Liu, Emmy, et al.
Pubblicazione: (2026)
ID-Cloak: Crafting Identity-Specific Cloaks Against Personalized Text-to-Image Generation
di: Teng, Qianrui, et al.
Pubblicazione: (2025)
di: Teng, Qianrui, et al.
Pubblicazione: (2025)
FFE-Hallu:Hallucinations in Fixed Figurative Expressions:Benchmark of Idioms and Proverbs in the Persian Language
di: Hosseini, Faezeh, et al.
Pubblicazione: (2026)
di: Hosseini, Faezeh, et al.
Pubblicazione: (2026)
HalluZig: Hallucination Detection using Zigzag Persistence
di: Samaga, Shreyas N., et al.
Pubblicazione: (2026)
di: Samaga, Shreyas N., et al.
Pubblicazione: (2026)
Localize, Understand, Collaborate: Semantic-Aware Dragging via Intention Reasoner
di: Cui, Xing, et al.
Pubblicazione: (2024)
di: Cui, Xing, et al.
Pubblicazione: (2024)
Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems
di: Zhang, Shaokun, et al.
Pubblicazione: (2025)
di: Zhang, Shaokun, et al.
Pubblicazione: (2025)
AgentExpt: Automating AI Experiment Design with LLM-based Resource Retrieval Agent
di: Li, Yu, et al.
Pubblicazione: (2025)
di: Li, Yu, et al.
Pubblicazione: (2025)
HalluClean: A Unified Framework to Combat Hallucinations in LLMs
di: Zhao, Yaxin, et al.
Pubblicazione: (2025)
di: Zhao, Yaxin, et al.
Pubblicazione: (2025)
FinReflectKG -- HalluBench: GraphRAG Hallucination Benchmark for Financial Question Answering Systems
di: Kumar, Mahesh, et al.
Pubblicazione: (2026)
di: Kumar, Mahesh, et al.
Pubblicazione: (2026)
HalluMix: A Task-Agnostic, Multi-Domain Benchmark for Real-World Hallucination Detection
di: Emery, Deanna, et al.
Pubblicazione: (2025)
di: Emery, Deanna, et al.
Pubblicazione: (2025)
MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models
di: Pandit, Shrey, et al.
Pubblicazione: (2025)
di: Pandit, Shrey, et al.
Pubblicazione: (2025)
SafetyFlow: An Agent-Flow System for Automated LLM Safety Benchmarking
di: Zhu, Xiangyang, et al.
Pubblicazione: (2025)
di: Zhu, Xiangyang, et al.
Pubblicazione: (2025)
Survey on AI-Generated Media Detection: From Non-MLLM to MLLM
di: Zou, Yueying, et al.
Pubblicazione: (2025)
di: Zou, Yueying, et al.
Pubblicazione: (2025)
HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs
di: Cherif, Ahmed
Pubblicazione: (2026)
di: Cherif, Ahmed
Pubblicazione: (2026)
MARCH: Multi-Agent Reinforced Self-Check for LLM Hallucination
di: Li, Zhuo, et al.
Pubblicazione: (2026)
di: Li, Zhuo, et al.
Pubblicazione: (2026)
The HalluRAG Dataset: Detecting Closed-Domain Hallucinations in RAG Applications Using an LLM's Internal States
di: Ridder, Fabian, et al.
Pubblicazione: (2024)
di: Ridder, Fabian, et al.
Pubblicazione: (2024)
TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration
di: Ma, Zerun, et al.
Pubblicazione: (2026)
di: Ma, Zerun, et al.
Pubblicazione: (2026)
FlowBench: Revisiting and Benchmarking Workflow-Guided Planning for LLM-based Agents
di: Xiao, Ruixuan, et al.
Pubblicazione: (2024)
di: Xiao, Ruixuan, et al.
Pubblicazione: (2024)
DSC2025 -- ViHallu Challenge: Detecting Hallucination in Vietnamese LLMs
di: Nguyen, Anh Thi-Hoang, et al.
Pubblicazione: (2026)
di: Nguyen, Anh Thi-Hoang, et al.
Pubblicazione: (2026)
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
di: Deng, Shihan, et al.
Pubblicazione: (2024)
di: Deng, Shihan, et al.
Pubblicazione: (2024)
AraHalluEval: A Fine-grained Hallucination Evaluation Framework for Arabic LLMs
di: Alansari, Aisha, et al.
Pubblicazione: (2025)
di: Alansari, Aisha, et al.
Pubblicazione: (2025)
HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL Conferences
di: Sakai, Yusuke, et al.
Pubblicazione: (2026)
di: Sakai, Yusuke, et al.
Pubblicazione: (2026)
Preference-Aware Memory Update for Long-Term LLM Agents
di: Sun, Haoran, et al.
Pubblicazione: (2025)
di: Sun, Haoran, et al.
Pubblicazione: (2025)
Documenti analoghi
-
T^2Agent A Tool-augmented Multimodal Misinformation Detection Agent with Monte Carlo Tree Search
di: Cui, Xing, et al.
Pubblicazione: (2025) -
Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs
di: Liu, Xuannan, et al.
Pubblicazione: (2025) -
HalluLens: LLM Hallucination Benchmark
di: Bang, Yejin, et al.
Pubblicazione: (2025) -
MMFakeBench: A Mixed-Source Multimodal Misinformation Detection Benchmark for LVLMs
di: Liu, Xuannan, et al.
Pubblicazione: (2024) -
Jailbreak Attacks and Defenses against Multimodal Generative Models: A Survey
di: Liu, Xuannan, et al.
Pubblicazione: (2024)