Triangulation as an Acceptance Rule for Multilingual Mechanistic Interpretability
Fuente:
arXiv
Salvato in:
| Autore principale: | Long, Yanan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Mechanistic Interpretability of Binary and Ternary Transformers
di: Li, Jason
Pubblicazione: (2024)
di: Li, Jason
Pubblicazione: (2024)
Tracking Equivalent Mechanistic Interpretations Across Neural Networks
di: Sun, Alan, et al.
Pubblicazione: (2026)
di: Sun, Alan, et al.
Pubblicazione: (2026)
MIB: A Mechanistic Interpretability Benchmark
di: Mueller, Aaron, et al.
Pubblicazione: (2025)
di: Mueller, Aaron, et al.
Pubblicazione: (2025)
HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks
di: Sun, Jiuding, et al.
Pubblicazione: (2025)
di: Sun, Jiuding, et al.
Pubblicazione: (2025)
Mechanistic Interpretability of GPT-like Models on Summarization Tasks
di: Mishra, Anurag
Pubblicazione: (2025)
di: Mishra, Anurag
Pubblicazione: (2025)
Binary Autoencoder for Mechanistic Interpretability of Large Language Models
di: Cho, Hakaze, et al.
Pubblicazione: (2025)
di: Cho, Hakaze, et al.
Pubblicazione: (2025)
Mechanistic Interpretability as Statistical Estimation: A Variance Analysis
di: Méloux, Maxime, et al.
Pubblicazione: (2025)
di: Méloux, Maxime, et al.
Pubblicazione: (2025)
Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable?
di: Méloux, Maxime, et al.
Pubblicazione: (2025)
di: Méloux, Maxime, et al.
Pubblicazione: (2025)
RSD: A Local Triangulation Audit Primitive for Learned Vector Blocks
di: Jin, Seungmin
Pubblicazione: (2026)
di: Jin, Seungmin
Pubblicazione: (2026)
How Do Transformers Learn to Associate Tokens: Gradient Leading Terms Bring Mechanistic Interpretability
di: Im, Shawn, et al.
Pubblicazione: (2026)
di: Im, Shawn, et al.
Pubblicazione: (2026)
LAS: Loss-less ANN-SNN Conversion for Fully Spike-Driven Large Language Models
di: Chen, Long, et al.
Pubblicazione: (2025)
di: Chen, Long, et al.
Pubblicazione: (2025)
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs
di: Song, Xiangchen, et al.
Pubblicazione: (2025)
di: Song, Xiangchen, et al.
Pubblicazione: (2025)
Detecting and Understanding Vulnerabilities in Language Models via Mechanistic Interpretability
di: García-Carrasco, Jorge, et al.
Pubblicazione: (2024)
di: García-Carrasco, Jorge, et al.
Pubblicazione: (2024)
Beyond Transcription: Mechanistic Interpretability in ASR
di: Glazer, Neta, et al.
Pubblicazione: (2025)
di: Glazer, Neta, et al.
Pubblicazione: (2025)
Mechanistic Data Attribution: Tracing the Training Origins of Interpretable LLM Units
di: Chen, Jianhui, et al.
Pubblicazione: (2026)
di: Chen, Jianhui, et al.
Pubblicazione: (2026)
Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference
di: Kim, Geonhee, et al.
Pubblicazione: (2024)
di: Kim, Geonhee, et al.
Pubblicazione: (2024)
The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research
di: Bai, Xiaoyan, et al.
Pubblicazione: (2026)
di: Bai, Xiaoyan, et al.
Pubblicazione: (2026)
One Sample to Rule Them All: Extreme Data Efficiency in Multidiscipline Reasoning with Reinforcement Learning
di: Li, Yiyuan, et al.
Pubblicazione: (2026)
di: Li, Yiyuan, et al.
Pubblicazione: (2026)
Comparing Feature Importance and Rule Extraction for Interpretability on Text Data
di: Lopardo, Gianluigi, et al.
Pubblicazione: (2022)
di: Lopardo, Gianluigi, et al.
Pubblicazione: (2022)
Intrinsic Self-Correction in LLMs: Towards Explainable Prompting via Mechanistic Interpretability
di: Lee, Yu-Ting, et al.
Pubblicazione: (2025)
di: Lee, Yu-Ting, et al.
Pubblicazione: (2025)
DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders
di: Wang, Xu, et al.
Pubblicazione: (2026)
di: Wang, Xu, et al.
Pubblicazione: (2026)
Mechanistic Unlearning: Robust Knowledge Unlearning and Editing via Mechanistic Localization
di: Guo, Phillip, et al.
Pubblicazione: (2024)
di: Guo, Phillip, et al.
Pubblicazione: (2024)
Beyond Accuracy: Introducing a Symbolic-Mechanistic Approach to Interpretable Evaluation
di: Habibi, Reza, et al.
Pubblicazione: (2026)
di: Habibi, Reza, et al.
Pubblicazione: (2026)
LK Losses: Direct Acceptance Rate Optimization for Speculative Decoding
di: Samarin, Alexander, et al.
Pubblicazione: (2026)
di: Samarin, Alexander, et al.
Pubblicazione: (2026)
Multilingual Amnesia: On the Transferability of Unlearning in Multilingual LLMs
di: Farashah, Alireza Dehghanpour, et al.
Pubblicazione: (2026)
di: Farashah, Alireza Dehghanpour, et al.
Pubblicazione: (2026)
How do Large Language Models Understand Relevance? A Mechanistic Interpretability Perspective
di: Liu, Qi, et al.
Pubblicazione: (2025)
di: Liu, Qi, et al.
Pubblicazione: (2025)
A Mathematical Philosophy of Explanations in Mechanistic Interpretability -- The Strange Science Part I.i
di: Ayonrinde, Kola, et al.
Pubblicazione: (2025)
di: Ayonrinde, Kola, et al.
Pubblicazione: (2025)
Adaptive Circuit Behavior and Generalization in Mechanistic Interpretability
di: Nainani, Jatin, et al.
Pubblicazione: (2024)
di: Nainani, Jatin, et al.
Pubblicazione: (2024)
On Mechanistic Circuits for Extractive Question-Answering
di: Basu, Samyadeep, et al.
Pubblicazione: (2025)
di: Basu, Samyadeep, et al.
Pubblicazione: (2025)
Mechanistic?
di: Saphra, Naomi, et al.
Pubblicazione: (2024)
di: Saphra, Naomi, et al.
Pubblicazione: (2024)
Mechanistic Insights into Grokking from the Embedding Layer
di: AlquBoj, H. V., et al.
Pubblicazione: (2025)
di: AlquBoj, H. V., et al.
Pubblicazione: (2025)
Mechanistic Anomaly Detection for "Quirky" Language Models
di: Johnston, David O., et al.
Pubblicazione: (2025)
di: Johnston, David O., et al.
Pubblicazione: (2025)
Simple Mechanistic Explanations for Out-Of-Context Reasoning
di: Wang, Atticus, et al.
Pubblicazione: (2025)
di: Wang, Atticus, et al.
Pubblicazione: (2025)
ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of Multilinguality
di: Longpre, Shayne, et al.
Pubblicazione: (2025)
di: Longpre, Shayne, et al.
Pubblicazione: (2025)
Stochastic Attention: Connectome-Inspired Randomized Routing for Expressive Linear-Time Attention
di: Jin, Zehao, et al.
Pubblicazione: (2026)
di: Jin, Zehao, et al.
Pubblicazione: (2026)
Multilingual Needle in a Haystack: Investigating Long-Context Behavior of Multilingual Large Language Models
di: Hengle, Amey, et al.
Pubblicazione: (2024)
di: Hengle, Amey, et al.
Pubblicazione: (2024)
Adapting Multilingual LLMs to Low-Resource Languages using Continued Pre-training and Synthetic Corpus
di: Joshi, Raviraj, et al.
Pubblicazione: (2024)
di: Joshi, Raviraj, et al.
Pubblicazione: (2024)
CultureGuard: Towards Culturally-Aware Dataset and Guard Model for Multilingual Safety Applications
di: Joshi, Raviraj, et al.
Pubblicazione: (2025)
di: Joshi, Raviraj, et al.
Pubblicazione: (2025)
Multilingual Self-Taught Faithfulness Evaluators
di: Alfano, Carlo, et al.
Pubblicazione: (2025)
di: Alfano, Carlo, et al.
Pubblicazione: (2025)
On the Calibration of Multilingual Question Answering LLMs
di: Yang, Yahan, et al.
Pubblicazione: (2023)
di: Yang, Yahan, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Mechanistic Interpretability of Binary and Ternary Transformers
di: Li, Jason
Pubblicazione: (2024) -
Tracking Equivalent Mechanistic Interpretations Across Neural Networks
di: Sun, Alan, et al.
Pubblicazione: (2026) -
MIB: A Mechanistic Interpretability Benchmark
di: Mueller, Aaron, et al.
Pubblicazione: (2025) -
HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks
di: Sun, Jiuding, et al.
Pubblicazione: (2025) -
Mechanistic Interpretability of GPT-like Models on Summarization Tasks
di: Mishra, Anurag
Pubblicazione: (2025)