H-Neurons: On the Existence, Impact, and Origin of Hallucination-Associated Neurons in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gao, Cheng, Chen, Huimin, Xiao, Chaojun, Chen, Zhiyi, Liu, Zhiyuan, Sun, Maosong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
KARL: Mitigating Hallucinations in LLMs via Knowledge-Boundary-Aware Reinforcement Learning
von: Gao, Cheng, et al.
Veröffentlicht: (2026)
von: Gao, Cheng, et al.
Veröffentlicht: (2026)
Large Language Models' Complicit Responses to Illicit Instructions across Socio-Legal Contexts
von: Wang, Xing, et al.
Veröffentlicht: (2025)
von: Wang, Xing, et al.
Veröffentlicht: (2025)
PersLLM: A Personified Training Approach for Large Language Models
von: Zeng, Zheni, et al.
Veröffentlicht: (2024)
von: Zeng, Zheni, et al.
Veröffentlicht: (2024)
Where Fake Citations Are Made: Tracing Field-Level Hallucination to Specific Neurons in LLMs
von: Chen, Yuefei, et al.
Veröffentlicht: (2026)
von: Chen, Yuefei, et al.
Veröffentlicht: (2026)
Medical Hallucinations in Foundation Models and Their Impact on Healthcare
von: Kim, Yubin, et al.
Veröffentlicht: (2025)
von: Kim, Yubin, et al.
Veröffentlicht: (2025)
Densing Law of LLMs
von: Xiao, Chaojun, et al.
Veröffentlicht: (2024)
von: Xiao, Chaojun, et al.
Veröffentlicht: (2024)
Do Hallucination Neurons Generalize? Evidence from Cross-Domain Transfer in LLMs
von: Vaddi, Snehit, et al.
Veröffentlicht: (2026)
von: Vaddi, Snehit, et al.
Veröffentlicht: (2026)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
von: Xiao, Chaojun, et al.
Veröffentlicht: (2024)
von: Xiao, Chaojun, et al.
Veröffentlicht: (2024)
Detection and Mitigation of Hallucination in Large Reasoning Models: A Mechanistic Perspective
von: Sun, Zhongxiang, et al.
Veröffentlicht: (2025)
von: Sun, Zhongxiang, et al.
Veröffentlicht: (2025)
How Does Alignment Enhance LLMs' Multilingual Capabilities? A Language Neurons Perspective
von: Zhang, Shimao, et al.
Veröffentlicht: (2025)
von: Zhang, Shimao, et al.
Veröffentlicht: (2025)
LLM-CAS: Dynamic Neuron Perturbation for Real-Time Hallucination Correction
von: Zhang, Jensen, et al.
Veröffentlicht: (2025)
von: Zhang, Jensen, et al.
Veröffentlicht: (2025)
Growth First, Care Second? Tracing the Landscape of LLM Value Preferences in Everyday Dilemmas
von: Chen, Zhiyi, et al.
Veröffentlicht: (2026)
von: Chen, Zhiyi, et al.
Veröffentlicht: (2026)
Identifying Good and Bad Neurons for Task-Level Controllable LLMs
von: Li, Wenjie, et al.
Veröffentlicht: (2026)
von: Li, Wenjie, et al.
Veröffentlicht: (2026)
NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs
von: Pan, Birong, et al.
Veröffentlicht: (2025)
von: Pan, Birong, et al.
Veröffentlicht: (2025)
Dissecting Role Cognition in Medical LLMs via Neuronal Ablation
von: Liang, Xun, et al.
Veröffentlicht: (2025)
von: Liang, Xun, et al.
Veröffentlicht: (2025)
Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction
von: Li, Ming, et al.
Veröffentlicht: (2025)
von: Li, Ming, et al.
Veröffentlicht: (2025)
Exploring the Benefit of Activation Sparsity in Pre-training
von: Zhang, Zhengyan, et al.
Veröffentlicht: (2024)
von: Zhang, Zhengyan, et al.
Veröffentlicht: (2024)
Optima: Optimizing Effectiveness and Efficiency for LLM-Based Multi-Agent System
von: Chen, Weize, et al.
Veröffentlicht: (2024)
von: Chen, Weize, et al.
Veröffentlicht: (2024)
Beyond Natural Language: LLMs Leveraging Alternative Formats for Enhanced Reasoning and Communication
von: Chen, Weize, et al.
Veröffentlicht: (2024)
von: Chen, Weize, et al.
Veröffentlicht: (2024)
Evaluating the Capabilities of LLMs for Supporting Anticipatory Impact Assessment
von: Allaham, Mowafak, et al.
Veröffentlicht: (2024)
von: Allaham, Mowafak, et al.
Veröffentlicht: (2024)
How Large Language Models are Designed to Hallucinate
von: Ackermann, Richard, et al.
Veröffentlicht: (2025)
von: Ackermann, Richard, et al.
Veröffentlicht: (2025)
IROTE: Human-like Traits Elicitation of Large Language Model via In-Context Self-Reflective Optimization
von: Bai, Yuzhuo, et al.
Veröffentlicht: (2025)
von: Bai, Yuzhuo, et al.
Veröffentlicht: (2025)
An evaluation of LLMs for political bias in Western media: Israel-Hamas and Ukraine-Russia wars
von: Chandra, Rohitash, et al.
Veröffentlicht: (2026)
von: Chandra, Rohitash, et al.
Veröffentlicht: (2026)
Developmental trajectories of decision making and affective dynamics in large language models
von: Wang, Zhihao, et al.
Veröffentlicht: (2025)
von: Wang, Zhihao, et al.
Veröffentlicht: (2025)
HumT DumT: Measuring and controlling human-like language in LLMs
von: Cheng, Myra, et al.
Veröffentlicht: (2025)
von: Cheng, Myra, et al.
Veröffentlicht: (2025)
Sparse Neurons Carry Strong Signals of Question Ambiguity in LLMs
von: Zhang, Zhuoxuan, et al.
Veröffentlicht: (2025)
von: Zhang, Zhuoxuan, et al.
Veröffentlicht: (2025)
Are Rationales Necessary and Sufficient? Tuning LLMs for Explainable Misinformation Detection
von: Wang, Bing, et al.
Veröffentlicht: (2026)
von: Wang, Bing, et al.
Veröffentlicht: (2026)
Explore the Potential of LLMs in Misinformation Detection: An Empirical Study
von: Chen, Mengyang, et al.
Veröffentlicht: (2023)
von: Chen, Mengyang, et al.
Veröffentlicht: (2023)
Neuron Empirical Gradient: Discovering and Quantifying Neurons Global Linear Controllability
von: Zhao, Xin, et al.
Veröffentlicht: (2024)
von: Zhao, Xin, et al.
Veröffentlicht: (2024)
LocalBench: Benchmarking LLMs on County-Level Local Knowledge and Reasoning
von: Gao, Zihan, et al.
Veröffentlicht: (2025)
von: Gao, Zihan, et al.
Veröffentlicht: (2025)
Hallucination Detection: A Probabilistic Framework Using Embeddings Distance Analysis
von: Ricco, Emanuele, et al.
Veröffentlicht: (2025)
von: Ricco, Emanuele, et al.
Veröffentlicht: (2025)
From Text to Multimodality: Exploring the Evolution and Impact of Large Language Models in Medical Practice
von: Niu, Qian, et al.
Veröffentlicht: (2024)
von: Niu, Qian, et al.
Veröffentlicht: (2024)
Text Corpora as Concept Fields: Black-Box Hallucination and Novelty Measurement
von: Kersting, Nicholas S., et al.
Veröffentlicht: (2026)
von: Kersting, Nicholas S., et al.
Veröffentlicht: (2026)
Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models
von: Dahl, Matthew, et al.
Veröffentlicht: (2024)
von: Dahl, Matthew, et al.
Veröffentlicht: (2024)
ELEPHANT: Measuring and understanding social sycophancy in LLMs
von: Cheng, Myra, et al.
Veröffentlicht: (2025)
von: Cheng, Myra, et al.
Veröffentlicht: (2025)
Place Matters: Comparing LLM Hallucination Rates for Place-Based Legal Queries
von: Curran, Damian, et al.
Veröffentlicht: (2025)
von: Curran, Damian, et al.
Veröffentlicht: (2025)
Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models
von: Zhang, Yue, et al.
Veröffentlicht: (2023)
von: Zhang, Yue, et al.
Veröffentlicht: (2023)
Accommodation and Epistemic Vigilance: A Pragmatic Account of Why LLMs Fail to Challenge Harmful Beliefs
von: Cheng, Myra, et al.
Veröffentlicht: (2026)
von: Cheng, Myra, et al.
Veröffentlicht: (2026)
NEAT: Concept driven Neuron Attribution in LLMs
von: Kavuri, Vivek Hruday, et al.
Veröffentlicht: (2025)
von: Kavuri, Vivek Hruday, et al.
Veröffentlicht: (2025)
SteLLA: A Structured Grading System Using LLMs with RAG
von: Qiu, Hefei, et al.
Veröffentlicht: (2025)
von: Qiu, Hefei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
KARL: Mitigating Hallucinations in LLMs via Knowledge-Boundary-Aware Reinforcement Learning
von: Gao, Cheng, et al.
Veröffentlicht: (2026) -
Large Language Models' Complicit Responses to Illicit Instructions across Socio-Legal Contexts
von: Wang, Xing, et al.
Veröffentlicht: (2025) -
PersLLM: A Personified Training Approach for Large Language Models
von: Zeng, Zheni, et al.
Veröffentlicht: (2024) -
Where Fake Citations Are Made: Tracing Field-Level Hallucination to Specific Neurons in LLMs
von: Chen, Yuefei, et al.
Veröffentlicht: (2026) -
Medical Hallucinations in Foundation Models and Their Impact on Healthcare
von: Kim, Yubin, et al.
Veröffentlicht: (2025)