Disentangling Deception and Hallucination Failures in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lu, Haolang, Peng, Hongrui, Fu, WeiYe, Nan, Guoshun, Cao, Xinye, Li, Xingrui, Guo, Hongcan, Wang, Kun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reallocating Attention Across Layers to Reduce Multimodal Hallucination
von: Lu, Haolang, et al.
Veröffentlicht: (2025)
von: Lu, Haolang, et al.
Veröffentlicht: (2025)
KGMark: A Diffusion Watermark for Knowledge Graphs
von: Peng, Hongrui, et al.
Veröffentlicht: (2025)
von: Peng, Hongrui, et al.
Veröffentlicht: (2025)
Streaming Hallucination Detection in Long Chain-of-Thought Reasoning
von: Lu, Haolang, et al.
Veröffentlicht: (2026)
von: Lu, Haolang, et al.
Veröffentlicht: (2026)
VideoMiner: Iteratively Grounding Key Frames of Hour-Long Videos via Tree-based Group Relative Policy Optimization
von: Cao, Xinye, et al.
Veröffentlicht: (2025)
von: Cao, Xinye, et al.
Veröffentlicht: (2025)
MALSIGHT: Exploring Malicious Source Code and Benign Pseudocode for Iterative Binary Malware Summarization
von: Lu, Haolang, et al.
Veröffentlicht: (2024)
von: Lu, Haolang, et al.
Veröffentlicht: (2024)
Diagnosing Knowledge Conflict in Multimodal Long-Chain Reasoning
von: Tang, Jing, et al.
Veröffentlicht: (2026)
von: Tang, Jing, et al.
Veröffentlicht: (2026)
ClueTracer: Question-to-Vision Clue Tracing for Training-Free Hallucination Suppression in Multimodal Reasoning
von: Xi, Gongli, et al.
Veröffentlicht: (2026)
von: Xi, Gongli, et al.
Veröffentlicht: (2026)
Advancing LLM-Based Security Automation with Customized Group Relative Policy Optimization for Zero-Touch Networks
von: Cao, Xinye, et al.
Veröffentlicht: (2025)
von: Cao, Xinye, et al.
Veröffentlicht: (2025)
Enhancing Cloud Network Resilience via a Robust LLM-Empowered Multi-Agent Reinforcement Learning Framework
von: Peng, Yixiao, et al.
Veröffentlicht: (2026)
von: Peng, Yixiao, et al.
Veröffentlicht: (2026)
Advancing Compositional LLM Reasoning with Structured Task Relations in Interactive Multimodal Communications
von: Cao, Xinye, et al.
Veröffentlicht: (2025)
von: Cao, Xinye, et al.
Veröffentlicht: (2025)
Auditing Meta-Cognitive Hallucinations in Reasoning Large Language Models
von: Lu, Haolang, et al.
Veröffentlicht: (2025)
von: Lu, Haolang, et al.
Veröffentlicht: (2025)
Advancing Expert Specialization for Better MoE
von: Guo, Hongcan, et al.
Veröffentlicht: (2025)
von: Guo, Hongcan, et al.
Veröffentlicht: (2025)
Deceptive Semantic Shortcuts on Reasoning Chains: How Far Can Models Go without Hallucination?
von: Li, Bangzheng, et al.
Veröffentlicht: (2023)
von: Li, Bangzheng, et al.
Veröffentlicht: (2023)
Helicase: Uncertainty-Guided Supply Chain Knowledge Graph Construction with Autonomous Multi-Agent LLMs
von: Long, Yunbo, et al.
Veröffentlicht: (2026)
von: Long, Yunbo, et al.
Veröffentlicht: (2026)
Quantized but Deceptive? A Multi-Dimensional Truthfulness Evaluation of Quantized LLMs
von: Fu, Yao, et al.
Veröffentlicht: (2025)
von: Fu, Yao, et al.
Veröffentlicht: (2025)
Causal-LLaVA: Causal Disentanglement for Mitigating Hallucination in Multimodal Large Language Models
von: Hu, Xinmiao, et al.
Veröffentlicht: (2025)
von: Hu, Xinmiao, et al.
Veröffentlicht: (2025)
A Rolling Stone Gathers No Moss: Adaptive Policy Optimization for Stable Self-Evaluation in Large Multimodal Models
von: Wang, Wenkai, et al.
Veröffentlicht: (2025)
von: Wang, Wenkai, et al.
Veröffentlicht: (2025)
3D-IDS: Doubly Disentangled Dynamic Intrusion Detection
von: Qiu, Chenyang, et al.
Veröffentlicht: (2023)
von: Qiu, Chenyang, et al.
Veröffentlicht: (2023)
Using Laplace Transform To Optimize the Hallucination of Generation Models
von: Kang, Cheng, et al.
Veröffentlicht: (2026)
von: Kang, Cheng, et al.
Veröffentlicht: (2026)
Adversarial Orthogonal Disentanglement for LVLM Hallucination Mitigation
von: Cheng, Ruoxi, et al.
Veröffentlicht: (2026)
von: Cheng, Ruoxi, et al.
Veröffentlicht: (2026)
LLM Lies: Hallucinations are not Bugs, but Features as Adversarial Examples
von: Yao, Jia-Yu, et al.
Veröffentlicht: (2023)
von: Yao, Jia-Yu, et al.
Veröffentlicht: (2023)
Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language Models
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024)
Cost-Effective Hallucination Detection for LLMs
von: Valentin, Simon, et al.
Veröffentlicht: (2024)
von: Valentin, Simon, et al.
Veröffentlicht: (2024)
Two Pathways to Truthfulness: On the Intrinsic Encoding of LLM Hallucinations
von: Luo, Wen, et al.
Veröffentlicht: (2026)
von: Luo, Wen, et al.
Veröffentlicht: (2026)
AI Deception: Risks, Dynamics, and Controls
von: Chen, Boyuan, et al.
Veröffentlicht: (2025)
von: Chen, Boyuan, et al.
Veröffentlicht: (2025)
VeriTrace: Evolving Mental Models for Deep Research Agents
von: Zhao, Haolang, et al.
Veröffentlicht: (2026)
von: Zhao, Haolang, et al.
Veröffentlicht: (2026)
Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-Evaluation
von: Zhang, Xiaoying, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaoying, et al.
Veröffentlicht: (2024)
Anchored Policy Optimization: Mitigating Exploration Collapse Via Support-Constrained Rectification
von: Wang, Tianyi, et al.
Veröffentlicht: (2026)
von: Wang, Tianyi, et al.
Veröffentlicht: (2026)
LLM-based Agents Suffer from Hallucinations: A Survey of Taxonomy, Methods, and Directions
von: Lin, Xixun, et al.
Veröffentlicht: (2025)
von: Lin, Xixun, et al.
Veröffentlicht: (2025)
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions
von: Xu, Jingxin, et al.
Veröffentlicht: (2025)
von: Xu, Jingxin, et al.
Veröffentlicht: (2025)
Revisiting Third-Party Library Detection: A Ground Truth Dataset and Its Implications Across Security Tasks
von: Gu, Jintao, et al.
Veröffentlicht: (2025)
von: Gu, Jintao, et al.
Veröffentlicht: (2025)
Distorted or Fabricated? A Survey on Hallucination in Video LLMs
von: Huang, Yiyang, et al.
Veröffentlicht: (2026)
von: Huang, Yiyang, et al.
Veröffentlicht: (2026)
DecepChain: Inducing Deceptive Reasoning in Large Language Models
von: Shen, Wei, et al.
Veröffentlicht: (2025)
von: Shen, Wei, et al.
Veröffentlicht: (2025)
TimeDRL: Disentangled Representation Learning for Multivariate Time-Series
von: Chang, Ching, et al.
Veröffentlicht: (2023)
von: Chang, Ching, et al.
Veröffentlicht: (2023)
Self-Supervised Learning of Disentangled Representations for Multivariate Time-Series
von: Chang, Ching, et al.
Veröffentlicht: (2024)
von: Chang, Ching, et al.
Veröffentlicht: (2024)
On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback
von: Williams, Marcus, et al.
Veröffentlicht: (2024)
von: Williams, Marcus, et al.
Veröffentlicht: (2024)
Time Series Stock Price Forecasting Based on Genetic Algorithm (GA)-Long Short-Term Memory Network (LSTM) Optimization
von: Sha, Xinye
Veröffentlicht: (2024)
von: Sha, Xinye
Veröffentlicht: (2024)
Multi-Objective Infeasibility Diagnosis for Routing Problems Using Large Language Models
von: Li, Kai, et al.
Veröffentlicht: (2025)
von: Li, Kai, et al.
Veröffentlicht: (2025)
Too Good to be Bad: On the Failure of LLMs to Role-Play Villains
von: Yi, Zihao, et al.
Veröffentlicht: (2025)
von: Yi, Zihao, et al.
Veröffentlicht: (2025)
Mitigating Deceptive Alignment via Self-Monitoring
von: Ji, Jiaming, et al.
Veröffentlicht: (2025)
von: Ji, Jiaming, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Reallocating Attention Across Layers to Reduce Multimodal Hallucination
von: Lu, Haolang, et al.
Veröffentlicht: (2025) -
KGMark: A Diffusion Watermark for Knowledge Graphs
von: Peng, Hongrui, et al.
Veröffentlicht: (2025) -
Streaming Hallucination Detection in Long Chain-of-Thought Reasoning
von: Lu, Haolang, et al.
Veröffentlicht: (2026) -
VideoMiner: Iteratively Grounding Key Frames of Hour-Long Videos via Tree-based Group Relative Policy Optimization
von: Cao, Xinye, et al.
Veröffentlicht: (2025) -
MALSIGHT: Exploring Malicious Source Code and Benign Pseudocode for Iterative Binary Malware Summarization
von: Lu, Haolang, et al.
Veröffentlicht: (2024)