Time Course MechInterp: Analyzing the Evolution of Components and Knowledge in Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Hakimi, Ahmad Dawar, Modarressi, Ali, Wicke, Philipp, Schütze, Hinrich |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
RET-LLM: Towards a General Read-Write Memory for Large Language Models
di: Modarressi, Ali, et al.
Pubblicazione: (2023)
di: Modarressi, Ali, et al.
Pubblicazione: (2023)
Do We Still Need Humans in the Loop? Comparing Human and LLM Annotation in Active Learning for Hostility Detection
di: Hakimi, Ahmad Dawar, et al.
Pubblicazione: (2026)
di: Hakimi, Ahmad Dawar, et al.
Pubblicazione: (2026)
Consistent Document-Level Relation Extraction via Counterfactuals
di: Modarressi, Ali, et al.
Pubblicazione: (2024)
di: Modarressi, Ali, et al.
Pubblicazione: (2024)
SLAyiNG: Towards Queer Language Processing
di: Veloso, Leonor, et al.
Pubblicazione: (2025)
di: Veloso, Leonor, et al.
Pubblicazione: (2025)
On Relation-Specific Neurons in Large Language Models
di: Liu, Yihong, et al.
Pubblicazione: (2025)
di: Liu, Yihong, et al.
Pubblicazione: (2025)
Do We Know What LLMs Don't Know? A Study of Consistency in Knowledge Probing
di: Zhao, Raoyuan, et al.
Pubblicazione: (2025)
di: Zhao, Raoyuan, et al.
Pubblicazione: (2025)
ImpliRet: Benchmarking the Implicit Fact Retrieval Challenge
di: Taghavi, Zeinab Sadat, et al.
Pubblicazione: (2025)
di: Taghavi, Zeinab Sadat, et al.
Pubblicazione: (2025)
GLUScope: A Tool for Analyzing GLU Neurons in Transformer Language Models
di: Gerstner, Sebastian, et al.
Pubblicazione: (2026)
di: Gerstner, Sebastian, et al.
Pubblicazione: (2026)
MemLLM: Finetuning LLMs to Use An Explicit Read-Write Memory
di: Modarressi, Ali, et al.
Pubblicazione: (2024)
di: Modarressi, Ali, et al.
Pubblicazione: (2024)
BlackboxNLP-2025 MIB Shared Task: Exploring Ensemble Strategies for Circuit Localization Methods
di: Mondorf, Philipp, et al.
Pubblicazione: (2025)
di: Mondorf, Philipp, et al.
Pubblicazione: (2025)
Probing Language Models' Gesture Understanding for Enhanced Human-AI Interaction
di: Wicke, Philipp
Pubblicazione: (2024)
di: Wicke, Philipp
Pubblicazione: (2024)
Exploring Spatial Schema Intuitions in Large Language and Vision Models
di: Wicke, Philipp, et al.
Pubblicazione: (2024)
di: Wicke, Philipp, et al.
Pubblicazione: (2024)
Mechanistic Understanding and Mitigation of Language Confusion in English-Centric Large Language Models
di: Nie, Ercong, et al.
Pubblicazione: (2025)
di: Nie, Ercong, et al.
Pubblicazione: (2025)
Evaluating Contextually Mediated Factual Recall in Multilingual Large Language Models
di: Liu, Yihong, et al.
Pubblicazione: (2026)
di: Liu, Yihong, et al.
Pubblicazione: (2026)
MEXA: Multilingual Evaluation of English-Centric LLMs via Cross-Lingual Alignment
di: Kargaran, Amir Hossein, et al.
Pubblicazione: (2024)
di: Kargaran, Amir Hossein, et al.
Pubblicazione: (2024)
A Recipe of Parallel Corpora Exploitation for Multilingual Large Language Models
di: Lin, Peiqin, et al.
Pubblicazione: (2024)
di: Lin, Peiqin, et al.
Pubblicazione: (2024)
NoLiMa: Long-Context Evaluation Beyond Literal Matching
di: Modarressi, Ali, et al.
Pubblicazione: (2025)
di: Modarressi, Ali, et al.
Pubblicazione: (2025)
Decomposed Prompting: Probing Multilingual Linguistic Structure Knowledge in Large Language Models
di: Nie, Ercong, et al.
Pubblicazione: (2024)
di: Nie, Ercong, et al.
Pubblicazione: (2024)
MechELK: A Mechanistic Interpretability Framework for Eliciting Latent Knowledge in Large Language Models
di: Park, Ji-jun, et al.
Pubblicazione: (2026)
di: Park, Ji-jun, et al.
Pubblicazione: (2026)
The Anatomy of an Edit: Mechanism-Guided Activation Steering for Knowledge Editing
di: Cao, Yuan, et al.
Pubblicazione: (2026)
di: Cao, Yuan, et al.
Pubblicazione: (2026)
Collapse of Dense Retrievers: Short, Early, and Literal Biases Outranking Factual Evidence
di: Fayyaz, Mohsen, et al.
Pubblicazione: (2025)
di: Fayyaz, Mohsen, et al.
Pubblicazione: (2025)
Breaking the Script Barrier in Multilingual Pre-Trained Language Models with Transliteration-Based Post-Training Alignment
di: Xhelili, Orgest, et al.
Pubblicazione: (2024)
di: Xhelili, Orgest, et al.
Pubblicazione: (2024)
MaLA-500: Massive Language Adaptation of Large Language Models
di: Lin, Peiqin, et al.
Pubblicazione: (2024)
di: Lin, Peiqin, et al.
Pubblicazione: (2024)
Steering MoE LLMs via Expert (De)Activation
di: Fayyaz, Mohsen, et al.
Pubblicazione: (2025)
di: Fayyaz, Mohsen, et al.
Pubblicazione: (2025)
GKnow: Measuring the Entanglement of Gender Bias and Factual Gender
di: Veloso, Leonor, et al.
Pubblicazione: (2026)
di: Veloso, Leonor, et al.
Pubblicazione: (2026)
MenuCraft: Interactive Menu System Design with Large Language Models
di: Kargaran, Amir Hossein, et al.
Pubblicazione: (2023)
di: Kargaran, Amir Hossein, et al.
Pubblicazione: (2023)
Evaluating Robustness of Large Language Models Against Multilingual Typographical Errors
di: Zhao, Raoyuan, et al.
Pubblicazione: (2025)
di: Zhao, Raoyuan, et al.
Pubblicazione: (2025)
Large Reasoning Models Are (Not Yet) Multilingual Latent Reasoners
di: Liu, Yihong, et al.
Pubblicazione: (2026)
di: Liu, Yihong, et al.
Pubblicazione: (2026)
Large Language Models as Neurolinguistic Subjects: Discrepancy between Performance and Competence
di: He, Linyang, et al.
Pubblicazione: (2024)
di: He, Linyang, et al.
Pubblicazione: (2024)
GNNavi: Navigating the Information Flow in Large Language Models by Graph Neural Network
di: Yuan, Shuzhou, et al.
Pubblicazione: (2024)
di: Yuan, Shuzhou, et al.
Pubblicazione: (2024)
TransliCo: A Contrastive Learning Framework to Address the Script Barrier in Multilingual Pretrained Language Models
di: Liu, Yihong, et al.
Pubblicazione: (2024)
di: Liu, Yihong, et al.
Pubblicazione: (2024)
Hateful Person or Hateful Model? Investigating the Role of Personas in Hate Speech Detection by Large Language Models
di: Yuan, Shuzhou, et al.
Pubblicazione: (2025)
di: Yuan, Shuzhou, et al.
Pubblicazione: (2025)
Understanding Gated Neurons in Transformers from Their Input-Output Functionality
di: Gerstner, Sebastian, et al.
Pubblicazione: (2025)
di: Gerstner, Sebastian, et al.
Pubblicazione: (2025)
TransMI: A Framework to Create Strong Baselines from Multilingual Pretrained Language Models for Transliterated Data
di: Liu, Yihong, et al.
Pubblicazione: (2024)
di: Liu, Yihong, et al.
Pubblicazione: (2024)
MaskLID: Code-Switching Language Identification through Iterative Masking
di: Kargaran, Amir Hossein, et al.
Pubblicazione: (2024)
di: Kargaran, Amir Hossein, et al.
Pubblicazione: (2024)
How Programming Concepts and Neurons Are Shared in Code Language Models
di: Kargaran, Amir Hossein, et al.
Pubblicazione: (2025)
di: Kargaran, Amir Hossein, et al.
Pubblicazione: (2025)
Derivational Morphology Reveals Analogical Generalization in Large Language Models
di: Hofmann, Valentin, et al.
Pubblicazione: (2024)
di: Hofmann, Valentin, et al.
Pubblicazione: (2024)
Left, Right, or Center? Evaluating LLM Framing in News Classification and Generation
di: Kennedy, Molly, et al.
Pubblicazione: (2026)
di: Kennedy, Molly, et al.
Pubblicazione: (2026)
GlotLID: Language Identification for Low-Resource Languages
di: Kargaran, Amir Hossein, et al.
Pubblicazione: (2023)
di: Kargaran, Amir Hossein, et al.
Pubblicazione: (2023)
Constructions Are So Difficult That Even Large Language Models Get Them Right for the Wrong Reasons
di: Zhou, Shijia, et al.
Pubblicazione: (2024)
di: Zhou, Shijia, et al.
Pubblicazione: (2024)
Documenti analoghi
-
RET-LLM: Towards a General Read-Write Memory for Large Language Models
di: Modarressi, Ali, et al.
Pubblicazione: (2023) -
Do We Still Need Humans in the Loop? Comparing Human and LLM Annotation in Active Learning for Hostility Detection
di: Hakimi, Ahmad Dawar, et al.
Pubblicazione: (2026) -
Consistent Document-Level Relation Extraction via Counterfactuals
di: Modarressi, Ali, et al.
Pubblicazione: (2024) -
SLAyiNG: Towards Queer Language Processing
di: Veloso, Leonor, et al.
Pubblicazione: (2025) -
On Relation-Specific Neurons in Large Language Models
di: Liu, Yihong, et al.
Pubblicazione: (2025)