Psychological Concept Neurons: Can Neural Control Bias Probing and Shift Generation in LLMs?
Fuente:
arXiv
Saved in:
| Main Authors: | Harada, Yuto, Hamada, Hiro Taiyo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Measuring How LLMs Internalize Human Psychological Concepts: A preliminary analysis
by: Hamada, Hiro Taiyo, et al.
Published: (2025)
by: Hamada, Hiro Taiyo, et al.
Published: (2025)
Unsupervised Concept Vector Extraction for Bias Control in LLMs
by: Cyberey, Hannah, et al.
Published: (2025)
by: Cyberey, Hannah, et al.
Published: (2025)
Understanding and Mitigating Gender Bias in LLMs via Interpretable Neuron Editing
by: Yu, Zeping, et al.
Published: (2025)
by: Yu, Zeping, et al.
Published: (2025)
NEAT: Concept driven Neuron Attribution in LLMs
by: Kavuri, Vivek Hruday, et al.
Published: (2025)
by: Kavuri, Vivek Hruday, et al.
Published: (2025)
What are They Thinking? Delineation, Probing and Tracking of Concepts in LLMs
by: Abdelwahab, Mohamed, et al.
Published: (2026)
by: Abdelwahab, Mohamed, et al.
Published: (2026)
Can We Trust LLMs? Mitigate Overconfidence Bias in LLMs through Knowledge Transfer
by: Yang, Haoyan, et al.
Published: (2024)
by: Yang, Haoyan, et al.
Published: (2024)
Probing Social Identity Bias in Chinese LLMs with Gendered Pronouns and Social Groups
by: Liu, Geng, et al.
Published: (2025)
by: Liu, Geng, et al.
Published: (2025)
Probing Gender Bias in Multilingual LLMs: A Case Study of Stereotypes in Persian
by: Kalhor, Ghazal, et al.
Published: (2025)
by: Kalhor, Ghazal, et al.
Published: (2025)
Semantic Shifts of Psychological Concepts in Scientific and Popular Media Discourse: A Distributional Semantics Analysis of Russian-Language Corpora
by: Anastasia, Orlova
Published: (2026)
by: Anastasia, Orlova
Published: (2026)
Shifting Perspectives: Steering Vectors for Robust Bias Mitigation in LLMs
by: Siddique, Zara, et al.
Published: (2025)
by: Siddique, Zara, et al.
Published: (2025)
Can LLMs Learn New Concepts Incrementally without Forgetting?
by: Zheng, Junhao, et al.
Published: (2024)
by: Zheng, Junhao, et al.
Published: (2024)
Difficulty-Controllable Multiple-Choice Question Generation Using Large Language Models and Direct Preference Optimization
by: Tomikawa, Yuto, et al.
Published: (2025)
by: Tomikawa, Yuto, et al.
Published: (2025)
CoBia: Constructed Conversations Can Trigger Otherwise Concealed Societal Biases in LLMs
by: Nikeghbal, Nafiseh, et al.
Published: (2025)
by: Nikeghbal, Nafiseh, et al.
Published: (2025)
Understanding How Value Neurons Shape the Generation of Specified Values in LLMs
by: Su, Yi, et al.
Published: (2025)
by: Su, Yi, et al.
Published: (2025)
Identifying Good and Bad Neurons for Task-Level Controllable LLMs
by: Li, Wenjie, et al.
Published: (2026)
by: Li, Wenjie, et al.
Published: (2026)
Template-Based Probes Are Imperfect Lenses for Counterfactual Bias Evaluation in LLMs
by: Kohankhaki, Farnaz, et al.
Published: (2024)
by: Kohankhaki, Farnaz, et al.
Published: (2024)
How Programming Concepts and Neurons Are Shared in Code Language Models
by: Kargaran, Amir Hossein, et al.
Published: (2025)
by: Kargaran, Amir Hossein, et al.
Published: (2025)
Evaluate Bias without Manual Test Sets: A Concept Representation Perspective for LLMs
by: Gao, Lang, et al.
Published: (2025)
by: Gao, Lang, et al.
Published: (2025)
Collective Predictive Coding as Model of Science: Formalizing Scientific Activities Towards Generative Science
by: Taniguchi, Tadahiro, et al.
Published: (2024)
by: Taniguchi, Tadahiro, et al.
Published: (2024)
Investigating Gender Bias in LLM-Generated Stories via Psychological Stereotypes
by: Masoudian, Shahed, et al.
Published: (2025)
by: Masoudian, Shahed, et al.
Published: (2025)
Can Prompt Modifiers Control Bias? A Comparative Analysis of Text-to-Image Generative Models
by: Shin, Philip Wootaek, et al.
Published: (2024)
by: Shin, Philip Wootaek, et al.
Published: (2024)
The Bias is in the Details: An Assessment of Cognitive Bias in LLMs
by: Knipper, R. Alexander, et al.
Published: (2025)
by: Knipper, R. Alexander, et al.
Published: (2025)
H-Neurons: On the Existence, Impact, and Origin of Hallucination-Associated Neurons in LLMs
by: Gao, Cheng, et al.
Published: (2025)
by: Gao, Cheng, et al.
Published: (2025)
A Multilingual Perspective on Probing Gender Bias
by: Stańczak, Karolina
Published: (2024)
by: Stańczak, Karolina
Published: (2024)
Neural FOXP2 -- Language Specific Neuron Steering for Targeted Language Improvement in LLMs
by: Saha, Anusa, et al.
Published: (2026)
by: Saha, Anusa, et al.
Published: (2026)
FilBench: Can LLMs Understand and Generate Filipino?
by: Miranda, Lester James V., et al.
Published: (2025)
by: Miranda, Lester James V., et al.
Published: (2025)
Can LLMs Generate and Solve Linguistic Olympiad Puzzles?
by: Majmudar, Neh, et al.
Published: (2025)
by: Majmudar, Neh, et al.
Published: (2025)
BigO(Bench) -- Can LLMs Generate Code with Controlled Time and Space Complexity?
by: Chambon, Pierre, et al.
Published: (2025)
by: Chambon, Pierre, et al.
Published: (2025)
Mitigating Copy Bias in In-Context Learning through Neuron Pruning
by: Ali, Ameen, et al.
Published: (2024)
by: Ali, Ameen, et al.
Published: (2024)
Social Bias Probing: Fairness Benchmarking for Language Models
by: Manerba, Marta Marchiori, et al.
Published: (2023)
by: Manerba, Marta Marchiori, et al.
Published: (2023)
A Review of Incorporating Psychological Theories in LLMs
by: Liu, Zizhou, et al.
Published: (2025)
by: Liu, Zizhou, et al.
Published: (2025)
Massive Supervised Fine-tuning Experiments Reveal How Data, Layer, and Training Factors Shape LLM Alignment Quality
by: Harada, Yuto, et al.
Published: (2025)
by: Harada, Yuto, et al.
Published: (2025)
Do Hallucination Neurons Generalize? Evidence from Cross-Domain Transfer in LLMs
by: Vaddi, Snehit, et al.
Published: (2026)
by: Vaddi, Snehit, et al.
Published: (2026)
Counterfactual Probing for the Influence of Affect and Specificity on Intergroup Bias
by: Govindarajan, Venkata S, et al.
Published: (2023)
by: Govindarajan, Venkata S, et al.
Published: (2023)
Disclosure and Mitigation of Gender Bias in LLMs
by: Dong, Xiangjue, et al.
Published: (2024)
by: Dong, Xiangjue, et al.
Published: (2024)
Sharing Matters: Analysing Neurons Across Languages and Tasks in LLMs
by: Wang, Weixuan, et al.
Published: (2024)
by: Wang, Weixuan, et al.
Published: (2024)
Style-Specific Neurons for Steering LLMs in Text Style Transfer
by: Lai, Wen, et al.
Published: (2024)
by: Lai, Wen, et al.
Published: (2024)
Investigating Neurons and Heads in Transformer-based LLMs for Typographical Errors
by: Tsuji, Kohei, et al.
Published: (2025)
by: Tsuji, Kohei, et al.
Published: (2025)
Concept Space Alignment in Multilingual LLMs
by: Peng, Qiwei, et al.
Published: (2024)
by: Peng, Qiwei, et al.
Published: (2024)
FairCoder: Evaluating Social Bias of LLMs in Code Generation
by: Du, Yongkang, et al.
Published: (2025)
by: Du, Yongkang, et al.
Published: (2025)
Similar Items
-
Measuring How LLMs Internalize Human Psychological Concepts: A preliminary analysis
by: Hamada, Hiro Taiyo, et al.
Published: (2025) -
Unsupervised Concept Vector Extraction for Bias Control in LLMs
by: Cyberey, Hannah, et al.
Published: (2025) -
Understanding and Mitigating Gender Bias in LLMs via Interpretable Neuron Editing
by: Yu, Zeping, et al.
Published: (2025) -
NEAT: Concept driven Neuron Attribution in LLMs
by: Kavuri, Vivek Hruday, et al.
Published: (2025) -
What are They Thinking? Delineation, Probing and Tracking of Concepts in LLMs
by: Abdelwahab, Mohamed, et al.
Published: (2026)