NNsight and NDIF: Democratizing Access to Open-Weight Foundation Model Internals
Fuente:
arXiv
Salvato in:
| Autori principali: | Fiotto-Kaufman, Jaden, Loftus, Alexander R., Todd, Eric, Brinkmann, Jannik, Pal, Koyena, Troitskii, Dmitrii, Ripa, Michael, Belfki, Adam, Rager, Can, Juang, Caden, Mueller, Aaron, Marks, Samuel, Sharma, Arnab Sen, Lucchetti, Francesca, Prakash, Nikhil, Brodley, Carla, Guha, Arjun, Bell, Jonathan, Wallace, Byron C., Bau, David |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Quest for the Right Mediator: Surveying Mechanistic Interpretability Through the Lens of Causal Mediation Analysis
di: Mueller, Aaron, et al.
Pubblicazione: (2024)
di: Mueller, Aaron, et al.
Pubblicazione: (2024)
In-Context Learning Without Copying
di: Sahin, Kerem, et al.
Pubblicazione: (2025)
di: Sahin, Kerem, et al.
Pubblicazione: (2025)
Understanding How CodeLLMs (Mis)Predict Types with Activation Steering
di: Lucchetti, Francesca, et al.
Pubblicazione: (2024)
di: Lucchetti, Francesca, et al.
Pubblicazione: (2024)
Internal states before wait modulate reasoning patterns
di: Troitskii, Dmitrii, et al.
Pubblicazione: (2025)
di: Troitskii, Dmitrii, et al.
Pubblicazione: (2025)
Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models
di: Karvonen, Adam, et al.
Pubblicazione: (2024)
di: Karvonen, Adam, et al.
Pubblicazione: (2024)
In-Context Algebra
di: Todd, Eric, et al.
Pubblicazione: (2025)
di: Todd, Eric, et al.
Pubblicazione: (2025)
Agents of Chaos
di: Shapira, Natalie, et al.
Pubblicazione: (2026)
di: Shapira, Natalie, et al.
Pubblicazione: (2026)
Future Lens: Anticipating Subsequent Tokens from a Single Hidden State
di: Pal, Koyena, et al.
Pubblicazione: (2023)
di: Pal, Koyena, et al.
Pubblicazione: (2023)
Poser: Unmasking Alignment Faking LLMs by Manipulating Their Internals
di: Clymer, Joshua, et al.
Pubblicazione: (2024)
di: Clymer, Joshua, et al.
Pubblicazione: (2024)
Do explanations generalize across large reasoning models?
di: Pal, Koyena, et al.
Pubblicazione: (2026)
di: Pal, Koyena, et al.
Pubblicazione: (2026)
Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
di: Casademunt, Helena, et al.
Pubblicazione: (2025)
di: Casademunt, Helena, et al.
Pubblicazione: (2025)
Lithology of sediment core KARGA, Russia
di: Troitskii, S L
Pubblicazione: (2014)
di: Troitskii, S L
Pubblicazione: (2014)
Pollen profile KARGA, Russia
di: Troitskii, S L
Pubblicazione: (2014)
di: Troitskii, S L
Pubblicazione: (2014)
Model Lakes
di: Pal, Koyena, et al.
Pubblicazione: (2024)
di: Pal, Koyena, et al.
Pubblicazione: (2024)
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
di: Marks, Samuel, et al.
Pubblicazione: (2024)
di: Marks, Samuel, et al.
Pubblicazione: (2024)
Automatically Interpreting Millions of Features in Large Language Models
di: Paulo, Gonçalo, et al.
Pubblicazione: (2024)
di: Paulo, Gonçalo, et al.
Pubblicazione: (2024)
Discovering Forbidden Topics in Language Models
di: Rager, Can, et al.
Pubblicazione: (2025)
di: Rager, Can, et al.
Pubblicazione: (2025)
Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning
di: Minder, Julian, et al.
Pubblicazione: (2025)
di: Minder, Julian, et al.
Pubblicazione: (2025)
Substance Beats Style: Why Beginning Students Fail to Code with LLMs
di: Lucchetti, Francesca, et al.
Pubblicazione: (2024)
di: Lucchetti, Francesca, et al.
Pubblicazione: (2024)
Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks
di: Karvonen, Adam, et al.
Pubblicazione: (2024)
di: Karvonen, Adam, et al.
Pubblicazione: (2024)
Non-Abelian fusion and braiding in many-body parton states
di: Bose, Koyena
Pubblicazione: (2026)
di: Bose, Koyena
Pubblicazione: (2026)
Mechanisms of AI Protein Folding in ESMFold
di: Lu, Kevin, et al.
Pubblicazione: (2026)
di: Lu, Kevin, et al.
Pubblicazione: (2026)
MIB: A Mechanistic Interpretability Benchmark
di: Mueller, Aaron, et al.
Pubblicazione: (2025)
di: Mueller, Aaron, et al.
Pubblicazione: (2025)
Mitigating Adaptive Attacks against Reasoning Models with Activation Consistency Training
di: Shah, Avidan, et al.
Pubblicazione: (2026)
di: Shah, Avidan, et al.
Pubblicazione: (2026)
GOV-REK: Governed Reward Engineering Kernels for Designing Robust Multi-Agent Reinforcement Learning Systems
di: Rana, Ashish, et al.
Pubblicazione: (2024)
di: Rana, Ashish, et al.
Pubblicazione: (2024)
Jailbreak Transferability Emerges from Shared Representations
di: Angell, Rico, et al.
Pubblicazione: (2025)
di: Angell, Rico, et al.
Pubblicazione: (2025)
NSA: Neuro-symbolic ARC Challenge
di: Batorski, Paweł, et al.
Pubblicazione: (2025)
di: Batorski, Paweł, et al.
Pubblicazione: (2025)
Elucidating Mechanisms of Demographic Bias in LLMs for Healthcare
di: Ahsan, Hiba, et al.
Pubblicazione: (2025)
di: Ahsan, Hiba, et al.
Pubblicazione: (2025)
Les diverses formes de colonisation pour les chômeurs en Autriche
di: Fritz Rager
Pubblicazione: (1934)
di: Fritz Rager
Pubblicazione: (1934)
Apprentice training in the Austrian metal industry
di: Fritz Rager
Pubblicazione: (1923)
di: Fritz Rager
Pubblicazione: (1923)
The settlement of the unemployed on the land in Austria
di: Fritz Rager
Pubblicazione: (1934)
di: Fritz Rager
Pubblicazione: (1934)
Provision for prolonged unemployment in certain industrial states
di: Fritz Rager
Pubblicazione: (1927)
di: Fritz Rager
Pubblicazione: (1927)
L'assistance extraordinaire en cas de chômage prolongé: étude internationale
di: Fritz Rager
Pubblicazione: (1927)
di: Fritz Rager
Pubblicazione: (1927)
Vector Arithmetic in Concept and Token Subspaces
di: Feucht, Sheridan, et al.
Pubblicazione: (2025)
di: Feucht, Sheridan, et al.
Pubblicazione: (2025)
Dyna-5G: Dynamic Role Switching for Self-Organizing 5G M2M Networks
di: Bitsikas, Evangelos, et al.
Pubblicazione: (2024)
di: Bitsikas, Evangelos, et al.
Pubblicazione: (2024)
Transfer Learning from Foundational Optimization Embeddings to Unsupervised SAT Representations
di: Pal, Koyena, et al.
Pubblicazione: (2026)
di: Pal, Koyena, et al.
Pubblicazione: (2026)
Deploying and Evaluating LLMs to Program Service Mobile Robots
di: Hu, Zichao, et al.
Pubblicazione: (2023)
di: Hu, Zichao, et al.
Pubblicazione: (2023)
Large Language Models Share Representations of Latent Grammatical Concepts Across Typologically Diverse Languages
di: Brinkmann, Jannik, et al.
Pubblicazione: (2025)
di: Brinkmann, Jannik, et al.
Pubblicazione: (2025)
Steering Language Models in Multi-Token Generation: A Case Study on Tense and Aspect
di: Klerings, Alina, et al.
Pubblicazione: (2025)
di: Klerings, Alina, et al.
Pubblicazione: (2025)
Fed-NDIF: A Noise-Embedded Federated Diffusion Model For Low-Count Whole-Body PET Denoising
di: Zhou, Yinchi, et al.
Pubblicazione: (2025)
di: Zhou, Yinchi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
The Quest for the Right Mediator: Surveying Mechanistic Interpretability Through the Lens of Causal Mediation Analysis
di: Mueller, Aaron, et al.
Pubblicazione: (2024) -
In-Context Learning Without Copying
di: Sahin, Kerem, et al.
Pubblicazione: (2025) -
Understanding How CodeLLMs (Mis)Predict Types with Activation Steering
di: Lucchetti, Francesca, et al.
Pubblicazione: (2024) -
Internal states before wait modulate reasoning patterns
di: Troitskii, Dmitrii, et al.
Pubblicazione: (2025) -
Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models
di: Karvonen, Adam, et al.
Pubblicazione: (2024)