Data-driven Circuit Discovery for Interpretability of Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Rai, Daking, Geva, Mor, Yao, Ziyu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
An Investigation of Neuron Activation as a Unified Lens to Explain Chain-of-Thought Eliciting Arithmetic Reasoning of LLMs
di: Rai, Daking, et al.
Pubblicazione: (2024)
di: Rai, Daking, et al.
Pubblicazione: (2024)
A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models
di: Rai, Daking, et al.
Pubblicazione: (2024)
di: Rai, Daking, et al.
Pubblicazione: (2024)
Mechanistic Understanding of Language Models in Syntactic Code Completion
di: Miller, Samuel, et al.
Pubblicazione: (2025)
di: Miller, Samuel, et al.
Pubblicazione: (2025)
NeuronSpark: A Spiking Neural Network Language Model with Selective State Space Dynamics
di: Tang, Zhengzheng
Pubblicazione: (2026)
di: Tang, Zhengzheng
Pubblicazione: (2026)
AIPsy-Affect: A Keyword-Free Clinical Stimulus Battery for Mechanistic Interpretability of Emotion in Language Models
di: Keeman, Michael
Pubblicazione: (2026)
di: Keeman, Michael
Pubblicazione: (2026)
CLMN: Concept based Language Models via Neural Symbolic Reasoning
di: Yang, Yibo
Pubblicazione: (2025)
di: Yang, Yibo
Pubblicazione: (2025)
The Origins of Representation Manifolds in Large Language Models
di: Modell, Alexander, et al.
Pubblicazione: (2025)
di: Modell, Alexander, et al.
Pubblicazione: (2025)
ExpliCa: Evaluating Explicit Causal Reasoning in Large Language Models
di: Miliani, Martina, et al.
Pubblicazione: (2025)
di: Miliani, Martina, et al.
Pubblicazione: (2025)
Failure by Interference: Language Models Make Balanced Parentheses Errors When Faulty Mechanisms Overshadow Sound Ones
di: Rai, Daking, et al.
Pubblicazione: (2025)
di: Rai, Daking, et al.
Pubblicazione: (2025)
MedMemoryBench: Benchmarking Agent Memory in Personalized Healthcare
di: Wang, Yihao, et al.
Pubblicazione: (2026)
di: Wang, Yihao, et al.
Pubblicazione: (2026)
Large Language Models Report Subjective Experience Under Self-Referential Processing
di: Berg, Cameron, et al.
Pubblicazione: (2025)
di: Berg, Cameron, et al.
Pubblicazione: (2025)
How Pruning Reshapes Features: Sparse Autoencoder Analysis of Weight-Pruned Language Models
di: Borobia, Hector, et al.
Pubblicazione: (2026)
di: Borobia, Hector, et al.
Pubblicazione: (2026)
Anonymization-Enhanced Privacy Protection for Mobile GUI Agents: Available but Invisible
di: Zhao, Lepeng, et al.
Pubblicazione: (2026)
di: Zhao, Lepeng, et al.
Pubblicazione: (2026)
Product-of-Experts Training Reduces Dataset Artifacts in Natural Language Inference
di: Mathew, Aby Mammen
Pubblicazione: (2026)
di: Mathew, Aby Mammen
Pubblicazione: (2026)
Prompting Encoder Models for Zero-Shot Classification: A Cross-Domain Study in Italian
di: Auriemma, Serena, et al.
Pubblicazione: (2024)
di: Auriemma, Serena, et al.
Pubblicazione: (2024)
Harnessing non-adversarial robustness in large language models
di: Zhou, Qinghua, et al.
Pubblicazione: (2026)
di: Zhou, Qinghua, et al.
Pubblicazione: (2026)
The GPT-4o Shock Emotional Attachment to AI Models and Its Impact on Regulatory Acceptance: A Cross-Cultural Analysis of the Immediate Transition from GPT-4o to GPT-5
di: Naito, Hiroki
Pubblicazione: (2025)
di: Naito, Hiroki
Pubblicazione: (2025)
Accelerating Suffix Jailbreak attacks with Prefix-Shared KV-cache
di: Wang, Xinhai, et al.
Pubblicazione: (2026)
di: Wang, Xinhai, et al.
Pubblicazione: (2026)
Probing for Representation Manifolds in Superposition
di: Modell, Alexander
Pubblicazione: (2026)
di: Modell, Alexander
Pubblicazione: (2026)
HR-Agent: A Task-Oriented Dialogue (TOD) LLM Agent Tailored for HR Applications
di: Xu, Weijie, et al.
Pubblicazione: (2024)
di: Xu, Weijie, et al.
Pubblicazione: (2024)
ProactBench: Beyond What The User Asked For
di: Harfi, Sepehr, et al.
Pubblicazione: (2026)
di: Harfi, Sepehr, et al.
Pubblicazione: (2026)
Thinking Machines: Mathematical Reasoning in the Age of LLMs
di: Asperti, Andrea, et al.
Pubblicazione: (2025)
di: Asperti, Andrea, et al.
Pubblicazione: (2025)
Context Aware Lemmatization and Morphological Tagging Method in Turkish
di: Sayallar, Cagri
Pubblicazione: (2025)
di: Sayallar, Cagri
Pubblicazione: (2025)
MetaCheckGPT -- A Multi-task Hallucination Detector Using LLM Uncertainty and Meta-models
di: Mehta, Rahul, et al.
Pubblicazione: (2024)
di: Mehta, Rahul, et al.
Pubblicazione: (2024)
Stick to your Role! Stability of Personal Values Expressed in Large Language Models
di: Kovač, Grgur, et al.
Pubblicazione: (2024)
di: Kovač, Grgur, et al.
Pubblicazione: (2024)
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models
di: Guo, Dongxin, et al.
Pubblicazione: (2026)
di: Guo, Dongxin, et al.
Pubblicazione: (2026)
The Concept Allocation Zone: Tracking How Concepts Form Across Transformer Depth
di: Henry, James
Pubblicazione: (2026)
di: Henry, James
Pubblicazione: (2026)
A Practical Guide to Streaming Continual Learning
di: Cossu, Andrea, et al.
Pubblicazione: (2026)
di: Cossu, Andrea, et al.
Pubblicazione: (2026)
Don't Look Back in Anger: MAGIC Net for Streaming Continual Learning with Temporal Dependence
di: Giannini, Federico, et al.
Pubblicazione: (2026)
di: Giannini, Federico, et al.
Pubblicazione: (2026)
cPNN: Continuous Progressive Neural Networks for Evolving Streaming Time Series
di: Giannini, Federico, et al.
Pubblicazione: (2026)
di: Giannini, Federico, et al.
Pubblicazione: (2026)
Geometric Evolution Maps: Extracting Stable Concept Probes from Transformer Residual Streams
di: Henry, James
Pubblicazione: (2026)
di: Henry, James
Pubblicazione: (2026)
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
di: Viveiros, André G., et al.
Pubblicazione: (2025)
di: Viveiros, André G., et al.
Pubblicazione: (2025)
A Survey on Collaborating Small and Large Language Models for Performance, Cost-effectiveness, Cloud-edge Privacy, and Trustworthiness
di: Wang, Fali, et al.
Pubblicazione: (2025)
di: Wang, Fali, et al.
Pubblicazione: (2025)
ProbeScale: Probing Analysis to Optimize Neural Scaling Laws for Efficient Small Language Model Inference
di: Das, Sourav
Pubblicazione: (2026)
di: Das, Sourav
Pubblicazione: (2026)
DeepPersona: A Generative Engine for Scaling Deep Synthetic Personas
di: Wang, Zhen, et al.
Pubblicazione: (2025)
di: Wang, Zhen, et al.
Pubblicazione: (2025)
PerkwE_COQA: Enhanced Persian Conversational Question Answering by combining contextual keyword extraction with Large Language Models
di: Moradbeiki, Pardis, et al.
Pubblicazione: (2024)
di: Moradbeiki, Pardis, et al.
Pubblicazione: (2024)
SwiftDossier: Tailored Automatic Dossier for Drug Discovery with LLMs and Agents
di: Fossi, Gabriele, et al.
Pubblicazione: (2024)
di: Fossi, Gabriele, et al.
Pubblicazione: (2024)
Unlocking UML Class Diagram Understanding in Vision Language Models
di: Naboichenko, Artem, et al.
Pubblicazione: (2026)
di: Naboichenko, Artem, et al.
Pubblicazione: (2026)
Do Reasoning Models Enhance Embedding Models?
di: Chan, Wun Yu, et al.
Pubblicazione: (2026)
di: Chan, Wun Yu, et al.
Pubblicazione: (2026)
Project Synapse: A Hierarchical Multi-Agent Framework with Hybrid Memory for Autonomous Resolution of Last-Mile Delivery Disruptions
di: Yadav, Arin Gopalan, et al.
Pubblicazione: (2026)
di: Yadav, Arin Gopalan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
An Investigation of Neuron Activation as a Unified Lens to Explain Chain-of-Thought Eliciting Arithmetic Reasoning of LLMs
di: Rai, Daking, et al.
Pubblicazione: (2024) -
A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models
di: Rai, Daking, et al.
Pubblicazione: (2024) -
Mechanistic Understanding of Language Models in Syntactic Code Completion
di: Miller, Samuel, et al.
Pubblicazione: (2025) -
NeuronSpark: A Spiking Neural Network Language Model with Selective State Space Dynamics
di: Tang, Zhengzheng
Pubblicazione: (2026) -
AIPsy-Affect: A Keyword-Free Clinical Stimulus Battery for Mechanistic Interpretability of Emotion in Language Models
di: Keeman, Michael
Pubblicazione: (2026)