Probing LLM Hallucination from Within: Perturbation-Driven Approach via Internal Knowledge
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Seongmin, Hsu, Hsiang, Chen, Chun-Fu, Chau, Duen Horng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Effective Guidance for Model Attention with Simple Yes-no Annotations
von: Lee, Seongmin, et al.
Veröffentlicht: (2024)
von: Lee, Seongmin, et al.
Veröffentlicht: (2024)
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety
von: Lee, Seongmin, et al.
Veröffentlicht: (2025)
von: Lee, Seongmin, et al.
Veröffentlicht: (2025)
LLM Attributor: Interactive Visual Attribution for LLM Generation
von: Lee, Seongmin, et al.
Veröffentlicht: (2024)
von: Lee, Seongmin, et al.
Veröffentlicht: (2024)
LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked
von: Phute, Mansi, et al.
Veröffentlicht: (2023)
von: Phute, Mansi, et al.
Veröffentlicht: (2023)
Wordflow: Social Prompt Engineering for Large Language Models
von: Wang, Zijie J., et al.
Veröffentlicht: (2024)
von: Wang, Zijie J., et al.
Veröffentlicht: (2024)
Transformer Explainer: Interactive Learning of Text-Generative Models
von: Cho, Aeree, et al.
Veröffentlicht: (2024)
von: Cho, Aeree, et al.
Veröffentlicht: (2024)
Mobile Fitting Room: On-device Virtual Try-on via Diffusion Models
von: Blalock, Justin, et al.
Veröffentlicht: (2024)
von: Blalock, Justin, et al.
Veröffentlicht: (2024)
Diffusion Explainer: Visual Explanation for Text-to-image Stable Diffusion
von: Lee, Seongmin, et al.
Veröffentlicht: (2023)
von: Lee, Seongmin, et al.
Veröffentlicht: (2023)
MeMemo: On-device Retrieval Augmentation for Private and Personalized Text Generation
von: Wang, Zijie J., et al.
Veröffentlicht: (2024)
von: Wang, Zijie J., et al.
Veröffentlicht: (2024)
PRISM: Probing Reasoning, Instruction, and Source Memory in LLM Hallucinations
von: Wu, Yuhe, et al.
Veröffentlicht: (2026)
von: Wu, Yuhe, et al.
Veröffentlicht: (2026)
Visualizing and Benchmarking LLM Factual Hallucination Tendencies via Internal State Analysis and Clustering
von: Mao, Nathan, et al.
Veröffentlicht: (2026)
von: Mao, Nathan, et al.
Veröffentlicht: (2026)
LLM-CAS: Dynamic Neuron Perturbation for Real-Time Hallucination Correction
von: Zhang, Jensen, et al.
Veröffentlicht: (2025)
von: Zhang, Jensen, et al.
Veröffentlicht: (2025)
Hallucination Detection with the Internal Layers of LLMs
von: Preiß, Martin
Veröffentlicht: (2025)
von: Preiß, Martin
Veröffentlicht: (2025)
Toxic HallucinAItions: Perturbing Prompts and Tracing LLM Circuits
von: Shimgekar, Soorya Ram, et al.
Veröffentlicht: (2026)
von: Shimgekar, Soorya Ram, et al.
Veröffentlicht: (2026)
Decoding by Perturbation: Mitigating MLLM Hallucinations via Dynamic Textual Perturbation
von: Jia, Sihang, et al.
Veröffentlicht: (2026)
von: Jia, Sihang, et al.
Veröffentlicht: (2026)
Look Within, Why LLMs Hallucinate: A Causal Perspective
von: Li, He, et al.
Veröffentlicht: (2024)
von: Li, He, et al.
Veröffentlicht: (2024)
Internal Knowledge Without External Expression: Probing the Generalization Boundary of a Classical Chinese Language Model
von: Chen, Jiuting, et al.
Veröffentlicht: (2026)
von: Chen, Jiuting, et al.
Veröffentlicht: (2026)
LLM-Enhanced Linear Autoencoders for Recommendation
von: Moon, Jaewan, et al.
Veröffentlicht: (2025)
von: Moon, Jaewan, et al.
Veröffentlicht: (2025)
TempPerturb-Eval: On the Joint Effects of Internal Temperature and External Perturbations in RAG Robustness
von: Zhou, Yongxin, et al.
Veröffentlicht: (2025)
von: Zhou, Yongxin, et al.
Veröffentlicht: (2025)
Internal Representation, Not Clinical Knowledge: Where Apparent LLM Triage Failures Originate
von: Navarro, David Fraile, et al.
Veröffentlicht: (2026)
von: Navarro, David Fraile, et al.
Veröffentlicht: (2026)
HuDEx: Integrating Hallucination Detection and Explainability for Enhancing the Reliability of LLM responses
von: Lee, Sujeong, et al.
Veröffentlicht: (2025)
von: Lee, Sujeong, et al.
Veröffentlicht: (2025)
Mitigating Object Hallucinations in MLLMs via Multi-Frequency Perturbations
von: Li, Shuo, et al.
Veröffentlicht: (2025)
von: Li, Shuo, et al.
Veröffentlicht: (2025)
Probing the Lack of Stable Internal Beliefs in LLMs
von: Luo, Yifan, et al.
Veröffentlicht: (2026)
von: Luo, Yifan, et al.
Veröffentlicht: (2026)
CHILL at SemEval-2025 Task 2: You Can't Just Throw Entities and Hope -- Make Your LLM to Get Them Right
von: Lee, Jaebok, et al.
Veröffentlicht: (2025)
von: Lee, Jaebok, et al.
Veröffentlicht: (2025)
From Internal Representations to Text Quality: A Geometric Approach to LLM Evaluation
von: Yusupov, Viacheslav, et al.
Veröffentlicht: (2025)
von: Yusupov, Viacheslav, et al.
Veröffentlicht: (2025)
Latent Paraphrasing: Perturbation on Layers Improves Knowledge Injection in Language Models
von: Kang, Minki, et al.
Veröffentlicht: (2024)
von: Kang, Minki, et al.
Veröffentlicht: (2024)
Question-to-Question Retrieval for Hallucination-Free Knowledge Access: An Approach for Wikipedia and Wikidata Question Answering
von: Thottingal, Santhosh
Veröffentlicht: (2025)
von: Thottingal, Santhosh
Veröffentlicht: (2025)
Neural Probe-Based Hallucination Detection for Large Language Models
von: Liang, Shize, et al.
Veröffentlicht: (2025)
von: Liang, Shize, et al.
Veröffentlicht: (2025)
Counterfactual Probing for Hallucination Detection and Mitigation in Large Language Models
von: Feng, Yijun
Veröffentlicht: (2025)
von: Feng, Yijun
Veröffentlicht: (2025)
Cross-Layer Attention Probing for Fine-Grained Hallucination Detection
von: Suresh, Malavika, et al.
Veröffentlicht: (2025)
von: Suresh, Malavika, et al.
Veröffentlicht: (2025)
Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and Refinement
von: Cao, Maosong, et al.
Veröffentlicht: (2025)
von: Cao, Maosong, et al.
Veröffentlicht: (2025)
HalluLens: LLM Hallucination Benchmark
von: Bang, Yejin, et al.
Veröffentlicht: (2025)
von: Bang, Yejin, et al.
Veröffentlicht: (2025)
Task-Specific Knowledge Distillation via Intermediate Probes
von: Brown, Ryan, et al.
Veröffentlicht: (2026)
von: Brown, Ryan, et al.
Veröffentlicht: (2026)
Unsupervised Extractive Dialogue Summarization in Hyperdimensional Space
von: Park, Seongmin, et al.
Veröffentlicht: (2024)
von: Park, Seongmin, et al.
Veröffentlicht: (2024)
Unsupervised Real-Time Hallucination Detection based on the Internal States of Large Language Models
von: Su, Weihang, et al.
Veröffentlicht: (2024)
von: Su, Weihang, et al.
Veröffentlicht: (2024)
Let's Fuse Step by Step: A Generative Fusion Decoding Algorithm with LLMs for Robust and Instruction-Aware ASR and OCR
von: Hsu, Chan-Jan, et al.
Veröffentlicht: (2024)
von: Hsu, Chan-Jan, et al.
Veröffentlicht: (2024)
Enhancing Hallucination Detection through Perturbation-Based Synthetic Data Generation in System Responses
von: Zhang, Dongxu, et al.
Veröffentlicht: (2024)
von: Zhang, Dongxu, et al.
Veröffentlicht: (2024)
Mitigating LLM Hallucinations via Conformal Abstention
von: Yadkori, Yasin Abbasi, et al.
Veröffentlicht: (2024)
von: Yadkori, Yasin Abbasi, et al.
Veröffentlicht: (2024)
Enhancing Hallucination Detection via Future Context
von: Lee, Joosung, et al.
Veröffentlicht: (2025)
von: Lee, Joosung, et al.
Veröffentlicht: (2025)
Does This Look Familiar to You? Knowledge Analysis via Model Internal Representations
von: Park, Sihyun
Veröffentlicht: (2025)
von: Park, Sihyun
Veröffentlicht: (2025)
Ähnliche Einträge
-
Effective Guidance for Model Attention with Simple Yes-no Annotations
von: Lee, Seongmin, et al.
Veröffentlicht: (2024) -
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety
von: Lee, Seongmin, et al.
Veröffentlicht: (2025) -
LLM Attributor: Interactive Visual Attribution for LLM Generation
von: Lee, Seongmin, et al.
Veröffentlicht: (2024) -
LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked
von: Phute, Mansi, et al.
Veröffentlicht: (2023) -
Wordflow: Social Prompt Engineering for Large Language Models
von: Wang, Zijie J., et al.
Veröffentlicht: (2024)