NEAT: Concept driven Neuron Attribution in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Kavuri, Vivek Hruday, Shroff, Gargi, Mishra, Rahul |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KTCR: Improving Implicit Hate Detection with Knowledge Transfer driven Concept Refinement
by: Garg, Samarth, et al.
Published: (2024)
by: Garg, Samarth, et al.
Published: (2024)
Think$^{2}$: Grounded Metacognitive Reasoning in Large Language Models
by: Elenjical, Abraham Paul, et al.
Published: (2026)
by: Elenjical, Abraham Paul, et al.
Published: (2026)
Neurons Speak in Ranges: Breaking Free from Discrete Neuronal Attribution
by: Haider, Muhammad Umair, et al.
Published: (2025)
by: Haider, Muhammad Umair, et al.
Published: (2025)
NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs
by: Pan, Birong, et al.
Published: (2025)
by: Pan, Birong, et al.
Published: (2025)
Can LLMs Convert Graphs to Text-Attributed Graphs?
by: Wang, Zehong, et al.
Published: (2024)
by: Wang, Zehong, et al.
Published: (2024)
Beyond Single Concept Vector: Modeling Concept Subspace in LLMs with Gaussian Distribution
by: Zhao, Haiyan, et al.
Published: (2024)
by: Zhao, Haiyan, et al.
Published: (2024)
The Transfer Neurons Hypothesis: An Underlying Mechanism for Language Latent Space Transitions in Multilingual LLMs
by: Tezuka, Hinata, et al.
Published: (2025)
by: Tezuka, Hinata, et al.
Published: (2025)
Split, Unlearn, Merge: Leveraging Data Attributes for More Effective Unlearning in LLMs
by: Kadhe, Swanand Ravindra, et al.
Published: (2024)
by: Kadhe, Swanand Ravindra, et al.
Published: (2024)
DiscoGraMS: Enhancing Movie Screen-Play Summarization using Movie Character-Aware Discourse Graph
by: Chitale, Maitreya Prafulla, et al.
Published: (2024)
by: Chitale, Maitreya Prafulla, et al.
Published: (2024)
Attribution-Guided Pruning for Insight and Control: Circuit Discovery and Targeted Correction in Small-scale LLMs
by: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Published: (2025)
by: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Published: (2025)
Enhancing Temporal Understanding in LLMs for Semi-structured Tables
by: Deng, Irwin, et al.
Published: (2024)
by: Deng, Irwin, et al.
Published: (2024)
H-STAR: LLM-driven Hybrid SQL-Text Adaptive Reasoning on Tables
by: Abhyankar, Nikhil, et al.
Published: (2024)
by: Abhyankar, Nikhil, et al.
Published: (2024)
AttributionBench: How Hard is Automatic Attribution Evaluation?
by: Li, Yifei, et al.
Published: (2024)
by: Li, Yifei, et al.
Published: (2024)
One Example Shown, Many Concepts Known! Counterexample-Driven Conceptual Reasoning in Mathematical LLMs
by: Li, Yinghui, et al.
Published: (2025)
by: Li, Yinghui, et al.
Published: (2025)
Teaching LLMs How to Learn with Contextual Fine-Tuning
by: Choi, Younwoo, et al.
Published: (2025)
by: Choi, Younwoo, et al.
Published: (2025)
Confidence Regulation Neurons in Language Models
by: Stolfo, Alessandro, et al.
Published: (2024)
by: Stolfo, Alessandro, et al.
Published: (2024)
Mathematical Reasoning via Intervention-Based Time-Series Causal Discovery Using LLMs as Concept Mastery Simulators
by: Okita, Tsuyoshi
Published: (2026)
by: Okita, Tsuyoshi
Published: (2026)
Decomposing Attention To Find Context-Sensitive Neurons
by: Gibson, Alex
Published: (2025)
by: Gibson, Alex
Published: (2025)
CoSy: Evaluating Textual Explanations of Neurons
by: Kopf, Laura, et al.
Published: (2024)
by: Kopf, Laura, et al.
Published: (2024)
Universal Neurons in GPT2 Language Models
by: Gurnee, Wes, et al.
Published: (2024)
by: Gurnee, Wes, et al.
Published: (2024)
Concept Tokens: Learning Behavioral Embeddings Through Concept Definitions
by: Sastre, Ignacio, et al.
Published: (2026)
by: Sastre, Ignacio, et al.
Published: (2026)
Finding Culture-Sensitive Neurons in Vision-Language Models
by: Zhao, Xiutian, et al.
Published: (2025)
by: Zhao, Xiutian, et al.
Published: (2025)
The Compliance Trap: How Structural Constraints Degrade Frontier AI Metacognition Under Adversarial Pressure
by: Kumar, Rahul
Published: (2026)
by: Kumar, Rahul
Published: (2026)
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers
by: Su, Jingtong, et al.
Published: (2025)
by: Su, Jingtong, et al.
Published: (2025)
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
by: Sheshadri, Abhay, et al.
Published: (2024)
by: Sheshadri, Abhay, et al.
Published: (2024)
Mechanistic Interpretability of GPT-like Models on Summarization Tasks
by: Mishra, Anurag
Published: (2025)
by: Mishra, Anurag
Published: (2025)
Dissecting Chronos: Sparse Autoencoders Reveal Causal Feature Hierarchies in Time Series Foundation Models
by: Mishra, Anurag
Published: (2026)
by: Mishra, Anurag
Published: (2026)
Reverse Thinking Makes LLMs Stronger Reasoners
by: Chen, Justin Chih-Yao, et al.
Published: (2024)
by: Chen, Justin Chih-Yao, et al.
Published: (2024)
Language-specific Neurons Do Not Facilitate Cross-Lingual Transfer
by: Mondal, Soumen Kumar, et al.
Published: (2025)
by: Mondal, Soumen Kumar, et al.
Published: (2025)
Attributing Response to Context: A Jensen-Shannon Divergence Driven Mechanistic Study of Context Attribution in Retrieval-Augmented Generation
by: Li, Ruizhe, et al.
Published: (2025)
by: Li, Ruizhe, et al.
Published: (2025)
Incorporating Attribution Importance for Improving Faithfulness Metrics
by: Zhao, Zhixue, et al.
Published: (2023)
by: Zhao, Zhixue, et al.
Published: (2023)
A Synthetic Dataset for Personal Attribute Inference
by: Yukhymenko, Hanna, et al.
Published: (2024)
by: Yukhymenko, Hanna, et al.
Published: (2024)
Reliable, Adaptable, and Attributable Language Models with Retrieval
by: Asai, Akari, et al.
Published: (2024)
by: Asai, Akari, et al.
Published: (2024)
Fast Byte Latent Transformer
by: Kallini, Julie, et al.
Published: (2026)
by: Kallini, Julie, et al.
Published: (2026)
Exploring Concept Depth: How Large Language Models Acquire Knowledge and Concept at Different Layers?
by: Jin, Mingyu, et al.
Published: (2024)
by: Jin, Mingyu, et al.
Published: (2024)
Fast Training Dataset Attribution via In-Context Learning
by: Fotouhi, Milad, et al.
Published: (2024)
by: Fotouhi, Milad, et al.
Published: (2024)
Efficient Estimation of Kernel Surrogate Models for Task Attribution
by: Zhang, Zhenshuo, et al.
Published: (2026)
by: Zhang, Zhenshuo, et al.
Published: (2026)
NeuroAda: Activating Each Neuron's Potential for Parameter-Efficient Fine-Tuning
by: Zhang, Zhi, et al.
Published: (2025)
by: Zhang, Zhi, et al.
Published: (2025)
Zeroth-Order Adaptive Neuron Alignment Based Pruning without Re-Training
by: Cunegatti, Elia, et al.
Published: (2024)
by: Cunegatti, Elia, et al.
Published: (2024)
A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models
by: Kazemi, Hamid, et al.
Published: (2026)
by: Kazemi, Hamid, et al.
Published: (2026)
Similar Items
-
KTCR: Improving Implicit Hate Detection with Knowledge Transfer driven Concept Refinement
by: Garg, Samarth, et al.
Published: (2024) -
Think$^{2}$: Grounded Metacognitive Reasoning in Large Language Models
by: Elenjical, Abraham Paul, et al.
Published: (2026) -
Neurons Speak in Ranges: Breaking Free from Discrete Neuronal Attribution
by: Haider, Muhammad Umair, et al.
Published: (2025) -
NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs
by: Pan, Birong, et al.
Published: (2025) -
Can LLMs Convert Graphs to Text-Attributed Graphs?
by: Wang, Zehong, et al.
Published: (2024)