pyvene: A Library for Understanding and Improving PyTorch Models via Interventions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Zhengxuan, Geiger, Atticus, Arora, Aryaman, Huang, Jing, Wang, Zheng, Goodman, Noah D., Manning, Christopher D., Potts, Christopher |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2024)
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2024)
ReFT: Representation Finetuning for Language Models
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2024)
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2024)
AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2025)
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2025)
Interpretability at Scale: Identifying Causal Mechanisms in Alpaca
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2023)
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2023)
Improved Representation Steering for Language Models
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2025)
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2025)
RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations
von: Huang, Jing, et al.
Veröffentlicht: (2024)
von: Huang, Jing, et al.
Veröffentlicht: (2024)
Bayesian scaling laws for in-context learning
von: Arora, Aryaman, et al.
Veröffentlicht: (2024)
von: Arora, Aryaman, et al.
Veröffentlicht: (2024)
Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations
von: Geiger, Atticus, et al.
Veröffentlicht: (2023)
von: Geiger, Atticus, et al.
Veröffentlicht: (2023)
ReCOGS: How Incidental Details of a Logical Form Overshadow an Evaluation of Semantic Interpretation
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2023)
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2023)
HyperSteer: Activation Steering at Scale with Hypernetworks
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop Questions
von: Zhong, Zexuan, et al.
Veröffentlicht: (2023)
von: Zhong, Zexuan, et al.
Veröffentlicht: (2023)
Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability
von: Geiger, Atticus, et al.
Veröffentlicht: (2023)
von: Geiger, Atticus, et al.
Veröffentlicht: (2023)
Recurrent Neural Networks Learn to Store and Generate Sequences using Non-Linear Representations
von: Csordás, Róbert, et al.
Veröffentlicht: (2024)
von: Csordás, Róbert, et al.
Veröffentlicht: (2024)
CausalGym: Benchmarking causal interpretability methods on linguistic tasks
von: Arora, Aryaman, et al.
Veröffentlicht: (2024)
von: Arora, Aryaman, et al.
Veröffentlicht: (2024)
PyTorch-IE: Fast and Reproducible Prototyping for Information Extraction
von: Binder, Arne, et al.
Veröffentlicht: (2024)
von: Binder, Arne, et al.
Veröffentlicht: (2024)
PreFT: Prefill-only finetuning for efficient inference
von: Lanpouthakoun, Andrew, et al.
Veröffentlicht: (2026)
von: Lanpouthakoun, Andrew, et al.
Veröffentlicht: (2026)
DMFF in PyTorch backend
von: Zhu, Jia-Xin, et al.
Veröffentlicht: (2026)
von: Zhu, Jia-Xin, et al.
Veröffentlicht: (2026)
ADMP in PyTorch backend
von: Zhu, Jia-Xin, et al.
Veröffentlicht: (2026)
von: Zhu, Jia-Xin, et al.
Veröffentlicht: (2026)
TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training
von: Liang, Wanchao, et al.
Veröffentlicht: (2024)
von: Liang, Wanchao, et al.
Veröffentlicht: (2024)
ADAG: Automatically Describing Attribution Graphs
von: Arora, Aryaman, et al.
Veröffentlicht: (2026)
von: Arora, Aryaman, et al.
Veröffentlicht: (2026)
CrypTorch: PyTorch-based Auto-tuning Compiler for Machine Learning with Multi-party Computation
von: Liu, Jinyu, et al.
Veröffentlicht: (2025)
von: Liu, Jinyu, et al.
Veröffentlicht: (2025)
ScoNe: Benchmarking Negation Reasoning in Language Models With Fine-Tuning and In-Context Learning
von: She, Jingyuan Selena, et al.
Veröffentlicht: (2023)
von: She, Jingyuan Selena, et al.
Veröffentlicht: (2023)
Updating CLIP to Prefer Descriptions Over Captions
von: Zur, Amir, et al.
Veröffentlicht: (2024)
von: Zur, Amir, et al.
Veröffentlicht: (2024)
Language Model Circuits Are Sparse in the Neuron Basis
von: Arora, Aryaman, et al.
Veröffentlicht: (2026)
von: Arora, Aryaman, et al.
Veröffentlicht: (2026)
depyf: Open the Opaque Box of PyTorch Compiler for Machine Learning Researchers
von: You, Kaichao, et al.
Veröffentlicht: (2024)
von: You, Kaichao, et al.
Veröffentlicht: (2024)
GraphMend: Code Transformations for Fixing Graph Breaks in PyTorch 2
von: Kashmira, Savini, et al.
Veröffentlicht: (2025)
von: Kashmira, Savini, et al.
Veröffentlicht: (2025)
HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
LeetDecoding: A PyTorch Library for Exponentially Decaying Causal Linear Attention with CUDA Implementations
von: Wang, Jiaping, et al.
Veröffentlicht: (2025)
von: Wang, Jiaping, et al.
Veröffentlicht: (2025)
Hessian of Perplexity for Large Language Models by PyTorch autograd (Open Source)
von: Ilin, Ivan
Veröffentlicht: (2025)
von: Ilin, Ivan
Veröffentlicht: (2025)
PyLO: Towards Accessible Learned Optimizers in PyTorch
von: Janson, Paul, et al.
Veröffentlicht: (2025)
von: Janson, Paul, et al.
Veröffentlicht: (2025)
torchgfn: A PyTorch GFlowNet library
von: Viviano, Joseph D., et al.
Veröffentlicht: (2023)
von: Viviano, Joseph D., et al.
Veröffentlicht: (2023)
TorchAO: PyTorch-Native Training-to-Serving Model Optimization
von: Or, Andrew, et al.
Veröffentlicht: (2025)
von: Or, Andrew, et al.
Veröffentlicht: (2025)
TorchSim: An efficient atomistic simulation engine in PyTorch
von: Cohen, Orion, et al.
Veröffentlicht: (2025)
von: Cohen, Orion, et al.
Veröffentlicht: (2025)
A PyTorch Library of Turing-Complete Neural Networks
von: Bates, Jonathan
Veröffentlicht: (2026)
von: Bates, Jonathan
Veröffentlicht: (2026)
torchsom: The Reference PyTorch Library for Self-Organizing Maps
von: Berthier, Louis, et al.
Veröffentlicht: (2025)
von: Berthier, Louis, et al.
Veröffentlicht: (2025)
NewsTorch: A PyTorch-based Toolkit for Learner-oriented News Recommendation
von: Wang, Rongyao, et al.
Veröffentlicht: (2026)
von: Wang, Rongyao, et al.
Veröffentlicht: (2026)
MRpro - open PyTorch-based MR reconstruction and processing package
von: Zimmermann, Felix Frederik, et al.
Veröffentlicht: (2025)
von: Zimmermann, Felix Frederik, et al.
Veröffentlicht: (2025)
GRAFX: An Open-Source Library for Audio Processing Graphs in PyTorch
von: Lee, Sungho, et al.
Veröffentlicht: (2024)
von: Lee, Sungho, et al.
Veröffentlicht: (2024)
PyGraph: Robust Compiler Support for CUDA Graphs in PyTorch
von: Ghosh, Abhishek, et al.
Veröffentlicht: (2025)
von: Ghosh, Abhishek, et al.
Veröffentlicht: (2025)
Venom: A PyTorch Generative Modeling Toolkit
von: Yan, Liang
Veröffentlicht: (2026)
von: Yan, Liang
Veröffentlicht: (2026)
Ähnliche Einträge
-
A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2024) -
ReFT: Representation Finetuning for Language Models
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2024) -
AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2025) -
Interpretability at Scale: Identifying Causal Mechanisms in Alpaca
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2023) -
Improved Representation Steering for Language Models
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2025)