Disentangling MLP Neuron Weights in Vocabulary Space
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Avrahamy, Asaf, Gur-Arieh, Yoav, Geva, Mor |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2026)
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2026)
Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2025)
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2025)
Enhancing Automated Interpretability with Output-Centric Feature Descriptions
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2025)
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2025)
Precise In-Parameter Concept Erasure in Large Language Models
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2025)
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2025)
LMEnt: A Suite for Analyzing Knowledge in Language Models from Pretraining Data to Representations
von: Gottesman, Daniela, et al.
Veröffentlicht: (2025)
von: Gottesman, Daniela, et al.
Veröffentlicht: (2025)
Constructing Interpretable Features from Compositional Neuron Groups
von: Shafran, Or, et al.
Veröffentlicht: (2025)
von: Shafran, Or, et al.
Veröffentlicht: (2025)
Backward Lens: Projecting Language Model Gradients into the Vocabulary Space
von: Katz, Shahar, et al.
Veröffentlicht: (2024)
von: Katz, Shahar, et al.
Veröffentlicht: (2024)
Estimating Knowledge in Large Language Models Without Generating a Single Token
von: Gottesman, Daniela, et al.
Veröffentlicht: (2024)
von: Gottesman, Daniela, et al.
Veröffentlicht: (2024)
Inferring Functionality of Attention Heads from their Parameters
von: Elhelo, Amit, et al.
Veröffentlicht: (2024)
von: Elhelo, Amit, et al.
Veröffentlicht: (2024)
RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations
von: Huang, Jing, et al.
Veröffentlicht: (2024)
von: Huang, Jing, et al.
Veröffentlicht: (2024)
Hallucinations Undermine Trust; Metacognition is a Way Forward
von: Yona, Gal, et al.
Veröffentlicht: (2026)
von: Yona, Gal, et al.
Veröffentlicht: (2026)
Eliciting Textual Descriptions from Representations of Continuous Prompts
von: Ramati, Dana, et al.
Veröffentlicht: (2024)
von: Ramati, Dana, et al.
Veröffentlicht: (2024)
Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers
von: Yona, Gal, et al.
Veröffentlicht: (2024)
von: Yona, Gal, et al.
Veröffentlicht: (2024)
Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?
von: Yona, Gal, et al.
Veröffentlicht: (2024)
von: Yona, Gal, et al.
Veröffentlicht: (2024)
The Hidden Space of Transformer Language Adapters
von: Alabi, Jesujoba O., et al.
Veröffentlicht: (2024)
von: Alabi, Jesujoba O., et al.
Veröffentlicht: (2024)
Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts
von: Ahrac, Sagi, et al.
Veröffentlicht: (2026)
von: Ahrac, Sagi, et al.
Veröffentlicht: (2026)
Preventing Rogue Agents Improves Multi-Agent Collaboration
von: Barbi, Ohav, et al.
Veröffentlicht: (2025)
von: Barbi, Ohav, et al.
Veröffentlicht: (2025)
Detecting (Un)answerability in Large Language Models with Linear Directions
von: Lavi, Maor Juliet, et al.
Veröffentlicht: (2025)
von: Lavi, Maor Juliet, et al.
Veröffentlicht: (2025)
Rethinking Selective Knowledge Distillation
von: Tavor, Almog, et al.
Veröffentlicht: (2026)
von: Tavor, Almog, et al.
Veröffentlicht: (2026)
Performance Gap in Entity Knowledge Extraction Across Modalities in Vision Language Models
von: Cohen, Ido, et al.
Veröffentlicht: (2024)
von: Cohen, Ido, et al.
Veröffentlicht: (2024)
Indications of Belief-Guided Agency and Meta-Cognitive Monitoring in Large Language Models
von: Yalon, Noam Steinmetz, et al.
Veröffentlicht: (2026)
von: Yalon, Noam Steinmetz, et al.
Veröffentlicht: (2026)
Jump to Conclusions: Short-Cutting Transformers With Linear Transformations
von: Din, Alexander Yom, et al.
Veröffentlicht: (2023)
von: Din, Alexander Yom, et al.
Veröffentlicht: (2023)
Friends and Grandmothers in Silico: Localizing Entity Cells in Language Models
von: Yona, Itay, et al.
Veröffentlicht: (2026)
von: Yona, Itay, et al.
Veröffentlicht: (2026)
From Loops to Oops: Fallback Behaviors of Language Models Under Uncertainty
von: Ivgi, Maor, et al.
Veröffentlicht: (2024)
von: Ivgi, Maor, et al.
Veröffentlicht: (2024)
MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of Documents
von: Wolfson, Tomer, et al.
Veröffentlicht: (2025)
von: Wolfson, Tomer, et al.
Veröffentlicht: (2025)
NoviCode: Generating Programs from Natural Language Utterances by Novices
von: Mordechai, Asaf Achi, et al.
Veröffentlicht: (2024)
von: Mordechai, Asaf Achi, et al.
Veröffentlicht: (2024)
Do Large Language Models Perform Latent Multi-Hop Reasoning without Exploiting Shortcuts?
von: Yang, Sohee, et al.
Veröffentlicht: (2024)
von: Yang, Sohee, et al.
Veröffentlicht: (2024)
Hopping Too Late: Exploring the Limitations of Large Language Models on Multi-Hop Queries
von: Biran, Eden, et al.
Veröffentlicht: (2024)
von: Biran, Eden, et al.
Veröffentlicht: (2024)
Do Large Language Models Latently Perform Multi-Hop Reasoning?
von: Yang, Sohee, et al.
Veröffentlicht: (2024)
von: Yang, Sohee, et al.
Veröffentlicht: (2024)
From Directions to Regions: Decomposing Activations in Language Models via Local Geometry
von: Shafran, Or, et al.
Veröffentlicht: (2026)
von: Shafran, Or, et al.
Veröffentlicht: (2026)
Don't Blame the Annotator: Bias Already Starts in the Annotation Instructions
von: Parmar, Mihir, et al.
Veröffentlicht: (2022)
von: Parmar, Mihir, et al.
Veröffentlicht: (2022)
From Insights to Actions: The Impact of Interpretability and Analysis Research on NLP
von: Mosbach, Marius, et al.
Veröffentlicht: (2024)
von: Mosbach, Marius, et al.
Veröffentlicht: (2024)
Intrinsic Test of Unlearning Using Parametric Knowledge Traces
von: Hong, Yihuai, et al.
Veröffentlicht: (2024)
von: Hong, Yihuai, et al.
Veröffentlicht: (2024)
Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs
von: Gekhman, Zorik, et al.
Veröffentlicht: (2026)
von: Gekhman, Zorik, et al.
Veröffentlicht: (2026)
NeuronMLP: Efficient LLM Inference via Singular Value Decomposition Compression and Tiling on AWS Trainium
von: Song, Dinghong, et al.
Veröffentlicht: (2025)
von: Song, Dinghong, et al.
Veröffentlicht: (2025)
How Well Can Reasoning Models Identify and Recover from Unhelpful Thoughts?
von: Yang, Sohee, et al.
Veröffentlicht: (2025)
von: Yang, Sohee, et al.
Veröffentlicht: (2025)
Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models
von: Ghandeharioun, Asma, et al.
Veröffentlicht: (2024)
von: Ghandeharioun, Asma, et al.
Veröffentlicht: (2024)
Mechanistically Interpretable Neural Encoding Reveals Fine-Grained Functional Selectivity in Human Visual Cortex
von: Grosbard, Idan Daniel, et al.
Veröffentlicht: (2026)
von: Grosbard, Idan Daniel, et al.
Veröffentlicht: (2026)
Latent Reasoning with Supervised Thinking States
von: Amos, Ido, et al.
Veröffentlicht: (2026)
von: Amos, Ido, et al.
Veröffentlicht: (2026)
A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning Chains
von: Jacovi, Alon, et al.
Veröffentlicht: (2024)
von: Jacovi, Alon, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2026) -
Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2025) -
Enhancing Automated Interpretability with Output-Centric Feature Descriptions
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2025) -
Precise In-Parameter Concept Erasure in Large Language Models
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2025) -
LMEnt: A Suite for Analyzing Knowledge in Language Models from Pretraining Data to Representations
von: Gottesman, Daniela, et al.
Veröffentlicht: (2025)