ReEval: Automatic Hallucination Evaluation for Retrieval-Augmented Large Language Models via Transferable Adversarial Attacks
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Xiaodong, Cheng, Hao, Liu, Xiaodong, Roth, Dan, Gao, Jianfeng |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Model Tells Itself Where to Attend: Faithfulness Meets Automatic Attention Steering
by: Zhang, Qingru, et al.
Published: (2024)
by: Zhang, Qingru, et al.
Published: (2024)
Training Large Reasoning Models Efficiently via Progressive Thought Encoding
by: Zhang, Zeliang, et al.
Published: (2026)
by: Zhang, Zeliang, et al.
Published: (2026)
REFIND at SemEval-2025 Task 3: Retrieval-Augmented Factuality Hallucination Detection in Large Language Models
by: Lee, DongGeon, et al.
Published: (2025)
by: Lee, DongGeon, et al.
Published: (2025)
Alleviating Hallucination in Large Vision-Language Models with Active Retrieval Augmentation
by: Qu, Xiaoye, et al.
Published: (2024)
by: Qu, Xiaoye, et al.
Published: (2024)
TF-Attack: Transferable and Fast Adversarial Attacks on Large Language Models
by: Li, Zelin, et al.
Published: (2024)
by: Li, Zelin, et al.
Published: (2024)
PerHalluEval: Persian Hallucination Evaluation Benchmark for Large Language Models
by: Hosseini, Mohammad, et al.
Published: (2025)
by: Hosseini, Mohammad, et al.
Published: (2025)
R-Eval: A Unified Toolkit for Evaluating Domain Knowledge of Retrieval Augmented Large Language Models
by: Tu, Shangqing, et al.
Published: (2024)
by: Tu, Shangqing, et al.
Published: (2024)
AttackEval: How to Evaluate the Effectiveness of Jailbreak Attacking on Large Language Models
by: Shu, Dong, et al.
Published: (2024)
by: Shu, Dong, et al.
Published: (2024)
MARAGE: Transferable Multi-Model Adversarial Attack for Retrieval-Augmented Generation Data Extraction
by: Hu, Xiao, et al.
Published: (2025)
by: Hu, Xiao, et al.
Published: (2025)
OnionEval: An Unified Evaluation of Fact-conflicting Hallucination for Small-Large Language Models
by: Sun, Chongren, et al.
Published: (2025)
by: Sun, Chongren, et al.
Published: (2025)
Diversifying the Expert Knowledge for Task-Agnostic Pruning in Sparse Mixture-of-Experts
by: Zhang, Zeliang, et al.
Published: (2024)
by: Zhang, Zeliang, et al.
Published: (2024)
Language Models as Inductive Reasoners
by: Yang, Zonglin, et al.
Published: (2022)
by: Yang, Zonglin, et al.
Published: (2022)
HaluEval-Wild: Evaluating Hallucinations of Language Models in the Wild
by: Zhu, Zhiying, et al.
Published: (2024)
by: Zhu, Zhiying, et al.
Published: (2024)
ReDeEP: Detecting Hallucination in Retrieval-Augmented Generation via Mechanistic Interpretability
by: Sun, Zhongxiang, et al.
Published: (2024)
by: Sun, Zhongxiang, et al.
Published: (2024)
Mitigating Multilingual Hallucination in Large Vision-Language Models
by: Qu, Xiaoye, et al.
Published: (2024)
by: Qu, Xiaoye, et al.
Published: (2024)
PR-Attack: Coordinated Prompt-RAG Attacks on Retrieval-Augmented Generation in Large Language Models via Bilevel Optimization
by: Jiao, Yang, et al.
Published: (2025)
by: Jiao, Yang, et al.
Published: (2025)
Black-Box Opinion Manipulation Attacks to Retrieval-Augmented Generation of Large Language Models
by: Chen, Zhuo, et al.
Published: (2024)
by: Chen, Zhuo, et al.
Published: (2024)
Retrieval Augmented Generation Evaluation in the Era of Large Language Models: A Comprehensive Survey
by: Gan, Aoran, et al.
Published: (2025)
by: Gan, Aoran, et al.
Published: (2025)
Beyond Facts: Evaluating Intent Hallucination in Large Language Models
by: Hao, Yijie, et al.
Published: (2025)
by: Hao, Yijie, et al.
Published: (2025)
Generative Adapter: Contextualizing Language Models in Parameters with A Single Forward Pass
by: Chen, Tong, et al.
Published: (2024)
by: Chen, Tong, et al.
Published: (2024)
PAC: Pronunciation-Aware Contextualized Large Language Model-based Automatic Speech Recognition
by: Fu, Li, et al.
Published: (2025)
by: Fu, Li, et al.
Published: (2025)
RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models
by: Niu, Cheng, et al.
Published: (2023)
by: Niu, Cheng, et al.
Published: (2023)
SEMA: Simple yet Effective Learning for Multi-Turn Jailbreak Attacks
by: Feng, Mingqian, et al.
Published: (2026)
by: Feng, Mingqian, et al.
Published: (2026)
BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali
by: Adib, Shefayat E Shams, et al.
Published: (2026)
by: Adib, Shefayat E Shams, et al.
Published: (2026)
Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language Models
by: Jiang, Chaoya, et al.
Published: (2024)
by: Jiang, Chaoya, et al.
Published: (2024)
Towards Consistent Natural-Language Explanations via Explanation-Consistency Finetuning
by: Chen, Yanda, et al.
Published: (2024)
by: Chen, Yanda, et al.
Published: (2024)
Reducing Hallucinations of Medical Multimodal Large Language Models with Visual Retrieval-Augmented Generation
by: Chu, Yun-Wei, et al.
Published: (2025)
by: Chu, Yun-Wei, et al.
Published: (2025)
Judge as A Judge: Improving the Evaluation of Retrieval-Augmented Generation through the Judge-Consistency of Large Language Models
by: Liu, Shuliang, et al.
Published: (2025)
by: Liu, Shuliang, et al.
Published: (2025)
Token-Modification Adversarial Attacks for Natural Language Processing: A Survey
by: Roth, Tom, et al.
Published: (2021)
by: Roth, Tom, et al.
Published: (2021)
From Neurons to Semantics: Evaluating Cross-Linguistic Alignment Capabilities of Large Language Models via Neurons Alignment
by: Huang, Chongxuan, et al.
Published: (2025)
by: Huang, Chongxuan, et al.
Published: (2025)
CMoralEval: A Moral Evaluation Benchmark for Chinese Large Language Models
by: Yu, Linhao, et al.
Published: (2024)
by: Yu, Linhao, et al.
Published: (2024)
OpenHuEval: Evaluating Large Language Model on Hungarian Specifics
by: Yang, Haote, et al.
Published: (2025)
by: Yang, Haote, et al.
Published: (2025)
AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation
by: Zhang, Xiechi, et al.
Published: (2025)
by: Zhang, Xiechi, et al.
Published: (2025)
ReasonAgain: Using Extractable Symbolic Programs to Evaluate Mathematical Reasoning
by: Yu, Xiaodong, et al.
Published: (2024)
by: Yu, Xiaodong, et al.
Published: (2024)
Re-Initialization Token Learning for Tool-Augmented Large Language Models
by: Li, Chenghao, et al.
Published: (2025)
by: Li, Chenghao, et al.
Published: (2025)
SocREval: Large Language Models with the Socratic Method for Reference-Free Reasoning Evaluation
by: He, Hangfeng, et al.
Published: (2023)
by: He, Hangfeng, et al.
Published: (2023)
Topic-FlipRAG: Topic-Orientated Adversarial Opinion Manipulation Attacks to Retrieval-Augmented Generation Models
by: Gong, Yuyang, et al.
Published: (2025)
by: Gong, Yuyang, et al.
Published: (2025)
AttentionRAG: Attention-Guided Context Pruning in Retrieval-Augmented Generation
by: Fang, Yixiong, et al.
Published: (2025)
by: Fang, Yixiong, et al.
Published: (2025)
ConSiDERS-The-Human Evaluation Framework: Rethinking Human Evaluation for Generative Large Language Models
by: Elangovan, Aparna, et al.
Published: (2024)
by: Elangovan, Aparna, et al.
Published: (2024)
Mitigating Hallucinations in Large Language Models via Self-Refinement-Enhanced Knowledge Retrieval
by: Niu, Mengjia, et al.
Published: (2024)
by: Niu, Mengjia, et al.
Published: (2024)
Similar Items
-
Model Tells Itself Where to Attend: Faithfulness Meets Automatic Attention Steering
by: Zhang, Qingru, et al.
Published: (2024) -
Training Large Reasoning Models Efficiently via Progressive Thought Encoding
by: Zhang, Zeliang, et al.
Published: (2026) -
REFIND at SemEval-2025 Task 3: Retrieval-Augmented Factuality Hallucination Detection in Large Language Models
by: Lee, DongGeon, et al.
Published: (2025) -
Alleviating Hallucination in Large Vision-Language Models with Active Retrieval Augmentation
by: Qu, Xiaoye, et al.
Published: (2024) -
TF-Attack: Transferable and Fast Adversarial Attacks on Large Language Models
by: Li, Zelin, et al.
Published: (2024)