Prompt-Guided Internal States for Hallucination Detection of Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Fujie, Yu, Peiqi, Yi, Biao, Zhang, Baolei, Li, Tong, Liu, Zheli |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CTRAP: Embedding Collapse Trap to Safeguard Large Language Models from Harmful Fine-Tuning
von: Yi, Biao, et al.
Veröffentlicht: (2025)
von: Yi, Biao, et al.
Veröffentlicht: (2025)
BadActs: A Universal Backdoor Defense in the Activation Space
von: Yi, Biao, et al.
Veröffentlicht: (2024)
von: Yi, Biao, et al.
Veröffentlicht: (2024)
Gradient Surgery for Safe LLM Fine-Tuning
von: Yi, Biao, et al.
Veröffentlicht: (2025)
von: Yi, Biao, et al.
Veröffentlicht: (2025)
Hallucination Detection via Internal States and Structured Reasoning Consistency in Large Language Models
von: Song, Yusheng, et al.
Veröffentlicht: (2025)
von: Song, Yusheng, et al.
Veröffentlicht: (2025)
Probe before You Talk: Towards Black-box Defense against Backdoor Unalignment for Large Language Models
von: Yi, Biao, et al.
Veröffentlicht: (2025)
von: Yi, Biao, et al.
Veröffentlicht: (2025)
BadReasoner: Planting Tunable Overthinking Backdoors into Large Reasoning Models for Fun or Profit
von: Yi, Biao, et al.
Veröffentlicht: (2025)
von: Yi, Biao, et al.
Veröffentlicht: (2025)
Detection Method for Prompt Injection by Integrating Pre-trained Model and Heuristic Feature Engineering
von: Ji, Yi, et al.
Veröffentlicht: (2025)
von: Ji, Yi, et al.
Veröffentlicht: (2025)
Unsupervised Real-Time Hallucination Detection based on the Internal States of Large Language Models
von: Su, Weihang, et al.
Veröffentlicht: (2024)
von: Su, Weihang, et al.
Veröffentlicht: (2024)
Confabulation: The Surprising Value of Large Language Model Hallucinations
von: Sui, Peiqi, et al.
Veröffentlicht: (2024)
von: Sui, Peiqi, et al.
Veröffentlicht: (2024)
INSIDE: LLMs' Internal States Retain the Power of Hallucination Detection
von: Chen, Chao, et al.
Veröffentlicht: (2024)
von: Chen, Chao, et al.
Veröffentlicht: (2024)
Detecting Hallucinations in Large Language Models via Internal Attention Divergence Signals
von: van Dijk, Gijs
Veröffentlicht: (2026)
von: van Dijk, Gijs
Veröffentlicht: (2026)
Attention Sinks as Internal Signals for Hallucination Detection in Large Language Models
von: Binkowski, Jakub, et al.
Veröffentlicht: (2026)
von: Binkowski, Jakub, et al.
Veröffentlicht: (2026)
Knowledge Overshadowing Causes Amalgamated Hallucination in Large Language Models
von: Zhang, Yuji, et al.
Veröffentlicht: (2024)
von: Zhang, Yuji, et al.
Veröffentlicht: (2024)
Bolster Hallucination Detection via Prompt-Guided Data Augmentation
von: Li, Wenyun, et al.
Veröffentlicht: (2025)
von: Li, Wenyun, et al.
Veröffentlicht: (2025)
Hallucination Detection and Evaluation of Large Language Model
von: Zhang, Chenggong, et al.
Veröffentlicht: (2025)
von: Zhang, Chenggong, et al.
Veröffentlicht: (2025)
Enhancing Robustness in Large Language Models: Prompting for Mitigating the Impact of Irrelevant Information
von: Jiang, Ming, et al.
Veröffentlicht: (2024)
von: Jiang, Ming, et al.
Veröffentlicht: (2024)
Active Prompting with Chain-of-Thought for Large Language Models
von: Diao, Shizhe, et al.
Veröffentlicht: (2023)
von: Diao, Shizhe, et al.
Veröffentlicht: (2023)
Chain-of-Thought Prompting Obscures Hallucination Cues in Large Language Models: An Empirical Evaluation
von: Cheng, Jiahao, et al.
Veröffentlicht: (2025)
von: Cheng, Jiahao, et al.
Veröffentlicht: (2025)
Mitigating Prompt-Induced Hallucinations in Large Language Models via Structured Reasoning
von: Hao, Jinbo, et al.
Veröffentlicht: (2026)
von: Hao, Jinbo, et al.
Veröffentlicht: (2026)
HaluNet: Learning Hallucination Risk from Internal Signals in LLM Question Answering
von: Tong, Chaodong, et al.
Veröffentlicht: (2025)
von: Tong, Chaodong, et al.
Veröffentlicht: (2025)
Hallucination Detection: Robustly Discerning Reliable Answers in Large Language Models
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
PretrainRL: Alleviating Factuality Hallucination of Large Language Models at the Beginning
von: Liu, Langming, et al.
Veröffentlicht: (2026)
von: Liu, Langming, et al.
Veröffentlicht: (2026)
Mitigating Hallucinations in Large Vision-Language Models by Self-Injecting Hallucinations
von: Lu, Yifan, et al.
Veröffentlicht: (2025)
von: Lu, Yifan, et al.
Veröffentlicht: (2025)
Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models
von: Li, Chaozhuo, et al.
Veröffentlicht: (2025)
von: Li, Chaozhuo, et al.
Veröffentlicht: (2025)
HIVE: Hidden-Evidence Verification for Hallucination Detection in Diffusion Large Language Models
von: Zhao, Guoshenghui, et al.
Veröffentlicht: (2026)
von: Zhao, Guoshenghui, et al.
Veröffentlicht: (2026)
PromptIntern: Saving Inference Costs by Internalizing Recurrent Prompt during Large Language Model Fine-tuning
von: Zou, Jiaru, et al.
Veröffentlicht: (2024)
von: Zou, Jiaru, et al.
Veröffentlicht: (2024)
Alleviating Hallucinations of Large Language Models through Induced Hallucinations
von: Zhang, Yue, et al.
Veröffentlicht: (2023)
von: Zhang, Yue, et al.
Veröffentlicht: (2023)
DecoPrompt : Decoding Prompts Reduces Hallucinations when Large Language Models Meet False Premises
von: Xu, Nan, et al.
Veröffentlicht: (2024)
von: Xu, Nan, et al.
Veröffentlicht: (2024)
Triggering Hallucinations in LLMs: A Quantitative Study of Prompt-Induced Hallucination in Large Language Models
von: Sato, Makoto
Veröffentlicht: (2025)
von: Sato, Makoto
Veröffentlicht: (2025)
Efficient Detection of Toxic Prompts in Large Language Models
von: Liu, Yi, et al.
Veröffentlicht: (2024)
von: Liu, Yi, et al.
Veröffentlicht: (2024)
HAD: HAllucination Detection Language Models Based on a Comprehensive Hallucination Taxonomy
von: Xu, Fan, et al.
Veröffentlicht: (2025)
von: Xu, Fan, et al.
Veröffentlicht: (2025)
Hallucination Detection with the Internal Layers of LLMs
von: Preiß, Martin
Veröffentlicht: (2025)
von: Preiß, Martin
Veröffentlicht: (2025)
Preference Orchestrator: Prompt-Aware Multi-Objective Alignment for Large Language Models
von: Liu, Biao, et al.
Veröffentlicht: (2025)
von: Liu, Biao, et al.
Veröffentlicht: (2025)
Dynamic Attention-Guided Context Decoding for Mitigating Context Faithfulness Hallucinations in Large Language Models
von: Huang, Yanwen, et al.
Veröffentlicht: (2025)
von: Huang, Yanwen, et al.
Veröffentlicht: (2025)
Mitigating Hallucinations in Large Vision-Language Models with Internal Fact-based Contrastive Decoding
von: Wang, Chao, et al.
Veröffentlicht: (2025)
von: Wang, Chao, et al.
Veröffentlicht: (2025)
CCNU at SemEval-2025 Task 3: Leveraging Internal and External Knowledge of Large Language Models for Multilingual Hallucination Annotation
von: Liu, Xu, et al.
Veröffentlicht: (2025)
von: Liu, Xu, et al.
Veröffentlicht: (2025)
Calibrating Reasoning in Language Models with Internal Consistency
von: Xie, Zhihui, et al.
Veröffentlicht: (2024)
von: Xie, Zhihui, et al.
Veröffentlicht: (2024)
Prompt-Response Semantic Divergence Metrics for Faithfulness Hallucination and Misalignment Detection in Large Language Models
von: Halperin, Igor
Veröffentlicht: (2025)
von: Halperin, Igor
Veröffentlicht: (2025)
Mitigating Hallucinations of Large Language Models in Medical Information Extraction via Contrastive Decoding
von: Xu, Derong, et al.
Veröffentlicht: (2024)
von: Xu, Derong, et al.
Veröffentlicht: (2024)
Practical Framework for Privacy-Preserving and Byzantine-robust Federated Learning
von: Zhang, Baolei, et al.
Veröffentlicht: (2025)
von: Zhang, Baolei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CTRAP: Embedding Collapse Trap to Safeguard Large Language Models from Harmful Fine-Tuning
von: Yi, Biao, et al.
Veröffentlicht: (2025) -
BadActs: A Universal Backdoor Defense in the Activation Space
von: Yi, Biao, et al.
Veröffentlicht: (2024) -
Gradient Surgery for Safe LLM Fine-Tuning
von: Yi, Biao, et al.
Veröffentlicht: (2025) -
Hallucination Detection via Internal States and Structured Reasoning Consistency in Large Language Models
von: Song, Yusheng, et al.
Veröffentlicht: (2025) -
Probe before You Talk: Towards Black-box Defense against Backdoor Unalignment for Large Language Models
von: Yi, Biao, et al.
Veröffentlicht: (2025)