Time Course MechInterp: Analyzing the Evolution of Components and Knowledge in Large Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Hakimi, Ahmad Dawar, Modarressi, Ali, Wicke, Philipp, Schütze, Hinrich |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
RET-LLM: Towards a General Read-Write Memory for Large Language Models
par: Modarressi, Ali, et autres
Publié: (2023)
par: Modarressi, Ali, et autres
Publié: (2023)
Do We Still Need Humans in the Loop? Comparing Human and LLM Annotation in Active Learning for Hostility Detection
par: Hakimi, Ahmad Dawar, et autres
Publié: (2026)
par: Hakimi, Ahmad Dawar, et autres
Publié: (2026)
Consistent Document-Level Relation Extraction via Counterfactuals
par: Modarressi, Ali, et autres
Publié: (2024)
par: Modarressi, Ali, et autres
Publié: (2024)
SLAyiNG: Towards Queer Language Processing
par: Veloso, Leonor, et autres
Publié: (2025)
par: Veloso, Leonor, et autres
Publié: (2025)
On Relation-Specific Neurons in Large Language Models
par: Liu, Yihong, et autres
Publié: (2025)
par: Liu, Yihong, et autres
Publié: (2025)
Do We Know What LLMs Don't Know? A Study of Consistency in Knowledge Probing
par: Zhao, Raoyuan, et autres
Publié: (2025)
par: Zhao, Raoyuan, et autres
Publié: (2025)
ImpliRet: Benchmarking the Implicit Fact Retrieval Challenge
par: Taghavi, Zeinab Sadat, et autres
Publié: (2025)
par: Taghavi, Zeinab Sadat, et autres
Publié: (2025)
GLUScope: A Tool for Analyzing GLU Neurons in Transformer Language Models
par: Gerstner, Sebastian, et autres
Publié: (2026)
par: Gerstner, Sebastian, et autres
Publié: (2026)
MemLLM: Finetuning LLMs to Use An Explicit Read-Write Memory
par: Modarressi, Ali, et autres
Publié: (2024)
par: Modarressi, Ali, et autres
Publié: (2024)
BlackboxNLP-2025 MIB Shared Task: Exploring Ensemble Strategies for Circuit Localization Methods
par: Mondorf, Philipp, et autres
Publié: (2025)
par: Mondorf, Philipp, et autres
Publié: (2025)
Probing Language Models' Gesture Understanding for Enhanced Human-AI Interaction
par: Wicke, Philipp
Publié: (2024)
par: Wicke, Philipp
Publié: (2024)
Exploring Spatial Schema Intuitions in Large Language and Vision Models
par: Wicke, Philipp, et autres
Publié: (2024)
par: Wicke, Philipp, et autres
Publié: (2024)
Mechanistic Understanding and Mitigation of Language Confusion in English-Centric Large Language Models
par: Nie, Ercong, et autres
Publié: (2025)
par: Nie, Ercong, et autres
Publié: (2025)
Evaluating Contextually Mediated Factual Recall in Multilingual Large Language Models
par: Liu, Yihong, et autres
Publié: (2026)
par: Liu, Yihong, et autres
Publié: (2026)
MEXA: Multilingual Evaluation of English-Centric LLMs via Cross-Lingual Alignment
par: Kargaran, Amir Hossein, et autres
Publié: (2024)
par: Kargaran, Amir Hossein, et autres
Publié: (2024)
A Recipe of Parallel Corpora Exploitation for Multilingual Large Language Models
par: Lin, Peiqin, et autres
Publié: (2024)
par: Lin, Peiqin, et autres
Publié: (2024)
NoLiMa: Long-Context Evaluation Beyond Literal Matching
par: Modarressi, Ali, et autres
Publié: (2025)
par: Modarressi, Ali, et autres
Publié: (2025)
Decomposed Prompting: Probing Multilingual Linguistic Structure Knowledge in Large Language Models
par: Nie, Ercong, et autres
Publié: (2024)
par: Nie, Ercong, et autres
Publié: (2024)
MechELK: A Mechanistic Interpretability Framework for Eliciting Latent Knowledge in Large Language Models
par: Park, Ji-jun, et autres
Publié: (2026)
par: Park, Ji-jun, et autres
Publié: (2026)
The Anatomy of an Edit: Mechanism-Guided Activation Steering for Knowledge Editing
par: Cao, Yuan, et autres
Publié: (2026)
par: Cao, Yuan, et autres
Publié: (2026)
Collapse of Dense Retrievers: Short, Early, and Literal Biases Outranking Factual Evidence
par: Fayyaz, Mohsen, et autres
Publié: (2025)
par: Fayyaz, Mohsen, et autres
Publié: (2025)
Breaking the Script Barrier in Multilingual Pre-Trained Language Models with Transliteration-Based Post-Training Alignment
par: Xhelili, Orgest, et autres
Publié: (2024)
par: Xhelili, Orgest, et autres
Publié: (2024)
MaLA-500: Massive Language Adaptation of Large Language Models
par: Lin, Peiqin, et autres
Publié: (2024)
par: Lin, Peiqin, et autres
Publié: (2024)
Steering MoE LLMs via Expert (De)Activation
par: Fayyaz, Mohsen, et autres
Publié: (2025)
par: Fayyaz, Mohsen, et autres
Publié: (2025)
GKnow: Measuring the Entanglement of Gender Bias and Factual Gender
par: Veloso, Leonor, et autres
Publié: (2026)
par: Veloso, Leonor, et autres
Publié: (2026)
MenuCraft: Interactive Menu System Design with Large Language Models
par: Kargaran, Amir Hossein, et autres
Publié: (2023)
par: Kargaran, Amir Hossein, et autres
Publié: (2023)
Evaluating Robustness of Large Language Models Against Multilingual Typographical Errors
par: Zhao, Raoyuan, et autres
Publié: (2025)
par: Zhao, Raoyuan, et autres
Publié: (2025)
Large Reasoning Models Are (Not Yet) Multilingual Latent Reasoners
par: Liu, Yihong, et autres
Publié: (2026)
par: Liu, Yihong, et autres
Publié: (2026)
Large Language Models as Neurolinguistic Subjects: Discrepancy between Performance and Competence
par: He, Linyang, et autres
Publié: (2024)
par: He, Linyang, et autres
Publié: (2024)
GNNavi: Navigating the Information Flow in Large Language Models by Graph Neural Network
par: Yuan, Shuzhou, et autres
Publié: (2024)
par: Yuan, Shuzhou, et autres
Publié: (2024)
TransliCo: A Contrastive Learning Framework to Address the Script Barrier in Multilingual Pretrained Language Models
par: Liu, Yihong, et autres
Publié: (2024)
par: Liu, Yihong, et autres
Publié: (2024)
Hateful Person or Hateful Model? Investigating the Role of Personas in Hate Speech Detection by Large Language Models
par: Yuan, Shuzhou, et autres
Publié: (2025)
par: Yuan, Shuzhou, et autres
Publié: (2025)
Understanding Gated Neurons in Transformers from Their Input-Output Functionality
par: Gerstner, Sebastian, et autres
Publié: (2025)
par: Gerstner, Sebastian, et autres
Publié: (2025)
TransMI: A Framework to Create Strong Baselines from Multilingual Pretrained Language Models for Transliterated Data
par: Liu, Yihong, et autres
Publié: (2024)
par: Liu, Yihong, et autres
Publié: (2024)
MaskLID: Code-Switching Language Identification through Iterative Masking
par: Kargaran, Amir Hossein, et autres
Publié: (2024)
par: Kargaran, Amir Hossein, et autres
Publié: (2024)
How Programming Concepts and Neurons Are Shared in Code Language Models
par: Kargaran, Amir Hossein, et autres
Publié: (2025)
par: Kargaran, Amir Hossein, et autres
Publié: (2025)
Derivational Morphology Reveals Analogical Generalization in Large Language Models
par: Hofmann, Valentin, et autres
Publié: (2024)
par: Hofmann, Valentin, et autres
Publié: (2024)
Left, Right, or Center? Evaluating LLM Framing in News Classification and Generation
par: Kennedy, Molly, et autres
Publié: (2026)
par: Kennedy, Molly, et autres
Publié: (2026)
GlotLID: Language Identification for Low-Resource Languages
par: Kargaran, Amir Hossein, et autres
Publié: (2023)
par: Kargaran, Amir Hossein, et autres
Publié: (2023)
Constructions Are So Difficult That Even Large Language Models Get Them Right for the Wrong Reasons
par: Zhou, Shijia, et autres
Publié: (2024)
par: Zhou, Shijia, et autres
Publié: (2024)
Documents similaires
-
RET-LLM: Towards a General Read-Write Memory for Large Language Models
par: Modarressi, Ali, et autres
Publié: (2023) -
Do We Still Need Humans in the Loop? Comparing Human and LLM Annotation in Active Learning for Hostility Detection
par: Hakimi, Ahmad Dawar, et autres
Publié: (2026) -
Consistent Document-Level Relation Extraction via Counterfactuals
par: Modarressi, Ali, et autres
Publié: (2024) -
SLAyiNG: Towards Queer Language Processing
par: Veloso, Leonor, et autres
Publié: (2025) -
On Relation-Specific Neurons in Large Language Models
par: Liu, Yihong, et autres
Publié: (2025)