LLM Lies: Hallucinations are not Bugs, but Features as Adversarial Examples
Fuente:
arXiv
Saved in:
| Main Authors: | Yao, Jia-Yu, Ning, Kun-Peng, Liu, Zhen-Hui, Ning, Mu-Nan, Liu, Yu-Yang, Yuan, Li |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PiCO: Peer Review in LLMs based on the Consistency Optimization
by: Ning, Kun-Peng, et al.
Published: (2024)
by: Ning, Kun-Peng, et al.
Published: (2024)
GPT as a Monte Carlo Language Tree: A Probabilistic Perspective
by: Ning, Kun-Peng, et al.
Published: (2025)
by: Ning, Kun-Peng, et al.
Published: (2025)
Examples as the Prompt: A Scalable Approach for Efficient LLM Adaptation in E-Commerce
by: Zeng, Jingying, et al.
Published: (2025)
by: Zeng, Jingying, et al.
Published: (2025)
How to Detect and Defeat Molecular Mirage: A Metric-Driven Benchmark for Hallucination in LLM-based Molecular Comprehension
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
Unnatural Languages Are Not Bugs but Features for LLMs
by: Duan, Keyu, et al.
Published: (2025)
by: Duan, Keyu, et al.
Published: (2025)
Two Pathways to Truthfulness: On the Intrinsic Encoding of LLM Hallucinations
by: Luo, Wen, et al.
Published: (2026)
by: Luo, Wen, et al.
Published: (2026)
Fast Adversarial Training against Textual Adversarial Attacks
by: Yang, Yichen, et al.
Published: (2024)
by: Yang, Yichen, et al.
Published: (2024)
Teaching with Lies: Curriculum DPO on Synthetic Negatives for Hallucination Detection
by: Pandit, Shrey, et al.
Published: (2025)
by: Pandit, Shrey, et al.
Published: (2025)
PDC & DM-SFT: A Road for LLM SQL Bug-Fix Enhancing
by: Duan, Yiwen, et al.
Published: (2024)
by: Duan, Yiwen, et al.
Published: (2024)
Revision or Re-Solving? Decomposing Second-Pass Gains in Multi-LLM Pipelines
by: Ning, Jingjie, et al.
Published: (2026)
by: Ning, Jingjie, et al.
Published: (2026)
Measuring and Reducing LLM Hallucination without Gold-Standard Answers
by: Wei, Jiaheng, et al.
Published: (2024)
by: Wei, Jiaheng, et al.
Published: (2024)
Unlocking the Power of LLM Uncertainty for Active In-Context Example Selection
by: Huang, Hsiu-Yuan, et al.
Published: (2024)
by: Huang, Hsiu-Yuan, et al.
Published: (2024)
REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations
by: Liang, Buyun, et al.
Published: (2026)
by: Liang, Buyun, et al.
Published: (2026)
DiLA: Enhancing LLM Tool Learning with Differential Logic Layer
by: Zhang, Yu, et al.
Published: (2024)
by: Zhang, Yu, et al.
Published: (2024)
FASTTRACK: Fast and Accurate Fact Tracing for LLMs
by: Chen, Si, et al.
Published: (2024)
by: Chen, Si, et al.
Published: (2024)
FlexKBQA: A Flexible LLM-Powered Framework for Few-Shot Knowledge Base Question Answering
by: Li, Zhenyu, et al.
Published: (2023)
by: Li, Zhenyu, et al.
Published: (2023)
Generative Adversarial Reasoner: Enhancing LLM Reasoning with Adversarial Reinforcement Learning
by: Liu, Qihao, et al.
Published: (2025)
by: Liu, Qihao, et al.
Published: (2025)
ReEval: Automatic Hallucination Evaluation for Retrieval-Augmented Large Language Models via Transferable Adversarial Attacks
by: Yu, Xiaodong, et al.
Published: (2023)
by: Yu, Xiaodong, et al.
Published: (2023)
RigorLLM: Resilient Guardrails for Large Language Models against Undesired Content
by: Yuan, Zhuowen, et al.
Published: (2024)
by: Yuan, Zhuowen, et al.
Published: (2024)
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning
by: Ning, Yansong, et al.
Published: (2025)
by: Ning, Yansong, et al.
Published: (2025)
Dense SAE Latents Are Features, Not Bugs
by: Sun, Xiaoqing, et al.
Published: (2025)
by: Sun, Xiaoqing, et al.
Published: (2025)
Reliable and diverse evaluation of LLM medical knowledge mastery
by: Zhou, Yuxuan, et al.
Published: (2024)
by: Zhou, Yuxuan, et al.
Published: (2024)
SVDq: 1.25-bit and 410x Key Cache Compression for LLM Attention
by: Yankun, Hong, et al.
Published: (2025)
by: Yankun, Hong, et al.
Published: (2025)
Mitigating LLM Hallucinations via Conformal Abstention
by: Yadkori, Yasin Abbasi, et al.
Published: (2024)
by: Yadkori, Yasin Abbasi, et al.
Published: (2024)
Hallucination Detection and Hallucination Mitigation: An Investigation
by: Luo, Junliang, et al.
Published: (2024)
by: Luo, Junliang, et al.
Published: (2024)
Combating Adversarial Attacks with Multi-Agent Debate
by: Chern, Steffi, et al.
Published: (2024)
by: Chern, Steffi, et al.
Published: (2024)
Not the Example, but the Process: How Self-Generated Examples Enhance LLM Reasoning
by: Gwak, Daehoon, et al.
Published: (2026)
by: Gwak, Daehoon, et al.
Published: (2026)
Lie to Me: Knowledge Graphs for Robust Hallucination Self-Detection in LLMs
by: Kale, Sahil, et al.
Published: (2025)
by: Kale, Sahil, et al.
Published: (2025)
AI Security Beyond Core Domains: Resume Screening as a Case Study of Adversarial Vulnerabilities in Specialized LLM Applications
by: Mu, Honglin, et al.
Published: (2025)
by: Mu, Honglin, et al.
Published: (2025)
IAE: Irony-based Adversarial Examples for Sentiment Analysis Systems
by: Yi, Xiaoyin, et al.
Published: (2024)
by: Yi, Xiaoyin, et al.
Published: (2024)
Unfamiliar Finetuning Examples Control How Language Models Hallucinate
by: Kang, Katie, et al.
Published: (2024)
by: Kang, Katie, et al.
Published: (2024)
Tuning-Free Accountable Intervention for LLM Deployment -- A Metacognitive Approach
by: Tan, Zhen, et al.
Published: (2024)
by: Tan, Zhen, et al.
Published: (2024)
Delta -- Contrastive Decoding Mitigates Text Hallucinations in Large Language Models
by: Huang, Cheng Peng, et al.
Published: (2025)
by: Huang, Cheng Peng, et al.
Published: (2025)
BetterV: Controlled Verilog Generation with Discriminative Guidance
by: Pei, Zehua, et al.
Published: (2024)
by: Pei, Zehua, et al.
Published: (2024)
Look Within, Why LLMs Hallucinate: A Causal Perspective
by: Li, He, et al.
Published: (2024)
by: Li, He, et al.
Published: (2024)
PRISM: Probing Reasoning, Instruction, and Source Memory in LLM Hallucinations
by: Wu, Yuhe, et al.
Published: (2026)
by: Wu, Yuhe, et al.
Published: (2026)
Rethinking Text-based Protein Understanding: Retrieval or LLM?
by: Wu, Juntong, et al.
Published: (2025)
by: Wu, Juntong, et al.
Published: (2025)
STT-Arena: A More Realistic Environment for Tool-Using with Spatio-Temporal Dynamics
by: Hui, Tingfeng, et al.
Published: (2026)
by: Hui, Tingfeng, et al.
Published: (2026)
Enhancing LLM-Based Data Annotation with Error Decomposition
by: Xu, Zhen, et al.
Published: (2026)
by: Xu, Zhen, et al.
Published: (2026)
Banishing LLM Hallucinations Requires Rethinking Generalization
by: Li, Johnny, et al.
Published: (2024)
by: Li, Johnny, et al.
Published: (2024)
Similar Items
-
PiCO: Peer Review in LLMs based on the Consistency Optimization
by: Ning, Kun-Peng, et al.
Published: (2024) -
GPT as a Monte Carlo Language Tree: A Probabilistic Perspective
by: Ning, Kun-Peng, et al.
Published: (2025) -
Examples as the Prompt: A Scalable Approach for Efficient LLM Adaptation in E-Commerce
by: Zeng, Jingying, et al.
Published: (2025) -
How to Detect and Defeat Molecular Mirage: A Metric-Driven Benchmark for Hallucination in LLM-based Molecular Comprehension
by: Li, Hao, et al.
Published: (2025) -
Unnatural Languages Are Not Bugs but Features for LLMs
by: Duan, Keyu, et al.
Published: (2025)