HalluHard: A Hard Multi-Turn Hallucination Benchmark
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Fan, Dongyang, Delsad, Sebastien, Flammarion, Nicolas, Andriushchenko, Maksym |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Does Refusal Training in LLMs Generalize to the Past Tense?
par: Andriushchenko, Maksym, et autres
Publié: (2024)
par: Andriushchenko, Maksym, et autres
Publié: (2024)
Is In-Context Learning Sufficient for Instruction Following in LLMs?
par: Zhao, Hao, et autres
Publié: (2024)
par: Zhao, Hao, et autres
Publié: (2024)
HalluLens: LLM Hallucination Benchmark
par: Bang, Yejin, et autres
Publié: (2025)
par: Bang, Yejin, et autres
Publié: (2025)
HalluMix: A Task-Agnostic, Multi-Domain Benchmark for Real-World Hallucination Detection
par: Emery, Deanna, et autres
Publié: (2025)
par: Emery, Deanna, et autres
Publié: (2025)
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
par: Andriushchenko, Maksym, et autres
Publié: (2024)
par: Andriushchenko, Maksym, et autres
Publié: (2024)
HalluVerse25: Fine-grained Multilingual Benchmark Dataset for LLM Hallucinations
par: Abdaljalil, Samir, et autres
Publié: (2025)
par: Abdaljalil, Samir, et autres
Publié: (2025)
Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
par: Rando, Javier, et autres
Publié: (2024)
par: Rando, Javier, et autres
Publié: (2024)
HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models
par: Liu, Emmy, et autres
Publié: (2026)
par: Liu, Emmy, et autres
Publié: (2026)
Long Is More for Alignment: A Simple but Tough-to-Beat Baseline for Instruction Fine-Tuning
par: Zhao, Hao, et autres
Publié: (2024)
par: Zhao, Hao, et autres
Publié: (2024)
MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models
par: Pandit, Shrey, et autres
Publié: (2025)
par: Pandit, Shrey, et autres
Publié: (2025)
HalluSAE: Detecting Hallucinations in Large Language Models via Sparse Auto-Encoders
par: Chen, Boshui, et autres
Publié: (2026)
par: Chen, Boshui, et autres
Publié: (2026)
BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali
par: Adib, Shefayat E Shams, et autres
Publié: (2026)
par: Adib, Shefayat E Shams, et autres
Publié: (2026)
DSC2025 -- ViHallu Challenge: Detecting Hallucination in Vietnamese LLMs
par: Nguyen, Anh Thi-Hoang, et autres
Publié: (2026)
par: Nguyen, Anh Thi-Hoang, et autres
Publié: (2026)
TopoBench: Benchmarking LLMs on Hard Topological Reasoning
par: Maniparambil, Mayug, et autres
Publié: (2026)
par: Maniparambil, Mayug, et autres
Publié: (2026)
HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL Conferences
par: Sakai, Yusuke, et autres
Publié: (2026)
par: Sakai, Yusuke, et autres
Publié: (2026)
HalluShift: Measuring Distribution Shifts towards Hallucination Detection in LLMs
par: Dasgupta, Sharanya, et autres
Publié: (2025)
par: Dasgupta, Sharanya, et autres
Publié: (2025)
HalluSearch at SemEval-2025 Task 3: A Search-Enhanced RAG Pipeline for Hallucination Detection
par: Abdallah, Mohamed A., et autres
Publié: (2025)
par: Abdallah, Mohamed A., et autres
Publié: (2025)
HalluCiteChecker: A Lightweight Toolkit for Hallucinated Citation Detection and Verification in the Era of AI Scientists
par: Sakai, Yusuke, et autres
Publié: (2026)
par: Sakai, Yusuke, et autres
Publié: (2026)
HalluGraph: Auditable Hallucination Detection for Legal RAG Systems via Knowledge Graph Alignment
par: Noël, Valentin, et autres
Publié: (2025)
par: Noël, Valentin, et autres
Publié: (2025)
MARBLE: A Hard Benchmark for Multimodal Spatial Reasoning and Planning
par: Jiang, Yulun, et autres
Publié: (2025)
par: Jiang, Yulun, et autres
Publié: (2025)
Capability-Based Scaling Trends for LLM-Based Red-Teaming
par: Panfilov, Alexander, et autres
Publié: (2025)
par: Panfilov, Alexander, et autres
Publié: (2025)
AHA: Aligning Large Audio-Language Models for Reasoning Hallucinations via Counterfactual Hard Negatives
par: Chen, Yanxi, et autres
Publié: (2025)
par: Chen, Yanxi, et autres
Publié: (2025)
Residual Drift Dominates Contradiction in Multi-Turn Constraint Reasoning
par: Kawada, Sebastien
Publié: (2026)
par: Kawada, Sebastien
Publié: (2026)
Emergent Misalignment is Easy, Narrow Misalignment is Hard
par: Soligo, Anna, et autres
Publié: (2026)
par: Soligo, Anna, et autres
Publié: (2026)
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
par: Huang, Kaixuan, et autres
Publié: (2025)
par: Huang, Kaixuan, et autres
Publié: (2025)
Soft Tokens, Hard Truths
par: Butt, Natasha, et autres
Publié: (2025)
par: Butt, Natasha, et autres
Publié: (2025)
Hard Negatives, Hard Lessons: Revisiting Training Data Quality for Robust Information Retrieval with LLMs
par: Thakur, Nandan, et autres
Publié: (2025)
par: Thakur, Nandan, et autres
Publié: (2025)
Language Models Do Hard Arithmetic Tasks Easily and Hardly Do Easy Arithmetic Tasks
par: Gambardella, Andrew, et autres
Publié: (2024)
par: Gambardella, Andrew, et autres
Publié: (2024)
Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math
par: Pandit, Shrey, et autres
Publié: (2025)
par: Pandit, Shrey, et autres
Publié: (2025)
MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
par: Sirdeshmukh, Ved, et autres
Publié: (2025)
par: Sirdeshmukh, Ved, et autres
Publié: (2025)
Compositional Hardness of Code in Large Language Models -- A Probabilistic Perspective
par: Wolf, Yotam, et autres
Publié: (2024)
par: Wolf, Yotam, et autres
Publié: (2024)
From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
par: Li, Tianle, et autres
Publié: (2024)
par: Li, Tianle, et autres
Publié: (2024)
NP-Hard Lower Bound Complexity for Semantic Self-Verification
par: Young, Robin
Publié: (2025)
par: Young, Robin
Publié: (2025)
Decomposing and Measuring Evaluation Awareness
par: Li, Changling, et autres
Publié: (2026)
par: Li, Changling, et autres
Publié: (2026)
MTRAG: A Multi-Turn Conversational Benchmark for Evaluating Retrieval-Augmented Generation Systems
par: Katsis, Yannis, et autres
Publié: (2025)
par: Katsis, Yannis, et autres
Publié: (2025)
StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns
par: Wan, Luanbo, et autres
Publié: (2025)
par: Wan, Luanbo, et autres
Publié: (2025)
C-FAITH: A Chinese Fine-Grained Benchmark for Automated Hallucination Evaluation
par: Zhang, Xu, et autres
Publié: (2025)
par: Zhang, Xu, et autres
Publié: (2025)
Beyond Hard Masks: Progressive Token Evolution for Diffusion Language Models
par: Zhong, Linhao, et autres
Publié: (2026)
par: Zhong, Linhao, et autres
Publié: (2026)
TimeSage-MT: A Multi-Turn Benchmark for Evaluating Agentic Time Series Reasoning
par: Kong, Yaxuan, et autres
Publié: (2026)
par: Kong, Yaxuan, et autres
Publié: (2026)
FaithBench: A Diverse Hallucination Benchmark for Summarization by Modern LLMs
par: Bao, Forrest Sheng, et autres
Publié: (2024)
par: Bao, Forrest Sheng, et autres
Publié: (2024)
Documents similaires
-
Does Refusal Training in LLMs Generalize to the Past Tense?
par: Andriushchenko, Maksym, et autres
Publié: (2024) -
Is In-Context Learning Sufficient for Instruction Following in LLMs?
par: Zhao, Hao, et autres
Publié: (2024) -
HalluLens: LLM Hallucination Benchmark
par: Bang, Yejin, et autres
Publié: (2025) -
HalluMix: A Task-Agnostic, Multi-Domain Benchmark for Real-World Hallucination Detection
par: Emery, Deanna, et autres
Publié: (2025) -
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
par: Andriushchenko, Maksym, et autres
Publié: (2024)