Training Neural Networks for Modularity aids Interpretability
Fuente:
arXiv
Salvato in:
| Autori principali: | Golechha, Satvik, Cope, Dylan, Schoots, Nandi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Studying Cross-cluster Modularity in Neural Networks
di: Golechha, Satvik, et al.
Pubblicazione: (2025)
di: Golechha, Satvik, et al.
Pubblicazione: (2025)
Challenges in Mechanistically Interpreting Model Representations
di: Golechha, Satvik, et al.
Pubblicazione: (2024)
di: Golechha, Satvik, et al.
Pubblicazione: (2024)
Progress Measures for Grokking on Real-world Tasks
di: Golechha, Satvik
Pubblicazione: (2024)
di: Golechha, Satvik
Pubblicazione: (2024)
Among Us: A Sandbox for Measuring and Detecting Agentic Deception
di: Golechha, Satvik, et al.
Pubblicazione: (2025)
di: Golechha, Satvik, et al.
Pubblicazione: (2025)
Relating Piecewise Linear Kolmogorov Arnold Networks to ReLU Networks
di: Schoots, Nandi, et al.
Pubblicazione: (2025)
di: Schoots, Nandi, et al.
Pubblicazione: (2025)
NICE: To Optimize In-Context Examples or Not?
di: Srivastava, Pragya, et al.
Pubblicazione: (2024)
di: Srivastava, Pragya, et al.
Pubblicazione: (2024)
Building Better Deception Probes Using Targeted Instruction Pairs
di: Natarajan, Vikram, et al.
Pubblicazione: (2026)
di: Natarajan, Vikram, et al.
Pubblicazione: (2026)
ABBEL: LLM Agents Acting through Belief Bottlenecks Expressed in Language
di: Lidayan, Aly, et al.
Pubblicazione: (2025)
di: Lidayan, Aly, et al.
Pubblicazione: (2025)
Who's the Evil Twin? Differential Auditing for Undesired Behavior
di: Balappanawar, Ishwar, et al.
Pubblicazione: (2025)
di: Balappanawar, Ishwar, et al.
Pubblicazione: (2025)
The Propensity for Density in Feed-forward Models
di: Schoots, Nandi, et al.
Pubblicazione: (2024)
di: Schoots, Nandi, et al.
Pubblicazione: (2024)
Extending Activation Steering to Broad Skills and Multiple Behaviours
di: van der Weij, Teun, et al.
Pubblicazione: (2024)
di: van der Weij, Teun, et al.
Pubblicazione: (2024)
Logic Sketch Prompting (LSP): A Deterministic and Interpretable Prompting Method
di: Tripathi, Satvik
Pubblicazione: (2025)
di: Tripathi, Satvik
Pubblicazione: (2025)
FocusLearn: Fully-Interpretable, High-Performance Modular Neural Networks for Time Series
di: Su, Qiqi, et al.
Pubblicazione: (2023)
di: Su, Qiqi, et al.
Pubblicazione: (2023)
Soft Contamination Means Benchmarks Test Shallow Generalization
di: Spiesberger, Ari, et al.
Pubblicazione: (2026)
di: Spiesberger, Ari, et al.
Pubblicazione: (2026)
Human-like Forgetting Curves in Deep Neural Networks
di: Kline, Dylan
Pubblicazione: (2025)
di: Kline, Dylan
Pubblicazione: (2025)
DiffusionBlocks: Block-wise Neural Network Training via Diffusion Interpretation
di: Shing, Makoto, et al.
Pubblicazione: (2025)
di: Shing, Makoto, et al.
Pubblicazione: (2025)
Faithful Interpretation for Graph Neural Networks
di: Hu, Lijie, et al.
Pubblicazione: (2024)
di: Hu, Lijie, et al.
Pubblicazione: (2024)
Modular Boundaries in Recurrent Neural Networks
di: Tanner, Jacob, et al.
Pubblicazione: (2023)
di: Tanner, Jacob, et al.
Pubblicazione: (2023)
Interpreting Neural Networks through Mahalanobis Distance
di: Oursland, Alan
Pubblicazione: (2024)
di: Oursland, Alan
Pubblicazione: (2024)
The Interpretable and Effective Graph Neural Additive Networks
di: Bechler-Speicher, Maya, et al.
Pubblicazione: (2024)
di: Bechler-Speicher, Maya, et al.
Pubblicazione: (2024)
Interpretable Neural Networks with Random Constructive Algorithm
di: Nan, Jing, et al.
Pubblicazione: (2023)
di: Nan, Jing, et al.
Pubblicazione: (2023)
Interpretable Graph Neural Networks for Tabular Data
di: Alkhatib, Amr, et al.
Pubblicazione: (2023)
di: Alkhatib, Amr, et al.
Pubblicazione: (2023)
Factor Graph-based Interpretable Neural Networks
di: Li, Yicong, et al.
Pubblicazione: (2025)
di: Li, Yicong, et al.
Pubblicazione: (2025)
Uncovering a Universal Abstract Algorithm for Modular Addition in Neural Networks
di: McCracken, Gavin, et al.
Pubblicazione: (2025)
di: McCracken, Gavin, et al.
Pubblicazione: (2025)
A Comprehensive Survey on Self-Interpretable Neural Networks
di: Ji, Yang, et al.
Pubblicazione: (2025)
di: Ji, Yang, et al.
Pubblicazione: (2025)
Spark: Modular Spiking Neural Networks
di: Franco, Mario, et al.
Pubblicazione: (2026)
di: Franco, Mario, et al.
Pubblicazione: (2026)
Closed-Form Interpretation of Neural Network Classifiers with Symbolic Gradients
di: Wetzel, Sebastian Johann
Pubblicazione: (2024)
di: Wetzel, Sebastian Johann
Pubblicazione: (2024)
Attention Consistency Regularization for Interpretable Early-Exit Neural Networks
di: Zhao, Yanhua
Pubblicazione: (2026)
di: Zhao, Yanhua
Pubblicazione: (2026)
Efficient and Interpretable Neural Networks Using Complex Lehmer Transform
di: Ataei, Masoud, et al.
Pubblicazione: (2025)
di: Ataei, Masoud, et al.
Pubblicazione: (2025)
Dissecting Language Models: Machine Unlearning via Selective Pruning
di: Pochinkov, Nicholas, et al.
Pubblicazione: (2024)
di: Pochinkov, Nicholas, et al.
Pubblicazione: (2024)
A Unified Theory of Compositionality, Modularity, and Interpretability in Markov Decision Processes
di: Ringstrom, Thomas J., et al.
Pubblicazione: (2025)
di: Ringstrom, Thomas J., et al.
Pubblicazione: (2025)
Gradient-Free Training of Quantized Neural Networks
di: Cohen, Noa, et al.
Pubblicazione: (2024)
di: Cohen, Noa, et al.
Pubblicazione: (2024)
Z-Error Loss for Training Neural Networks
di: Godin, Guillaume
Pubblicazione: (2025)
di: Godin, Guillaume
Pubblicazione: (2025)
Automatic Stability and Recovery for Neural Network Training
di: Or, Barak
Pubblicazione: (2026)
di: Or, Barak
Pubblicazione: (2026)
Energy Consumption in Parallel Neural Network Training
di: Huber, Philipp, et al.
Pubblicazione: (2025)
di: Huber, Philipp, et al.
Pubblicazione: (2025)
Train Faster, Perform Better: Modular Adaptive Training in Over-Parameterized Models
di: Shi, Yubin, et al.
Pubblicazione: (2024)
di: Shi, Yubin, et al.
Pubblicazione: (2024)
GINN-KAN: Interpretability pipelining with applications in Physics Informed Neural Networks
di: Ranasinghe, Nisal, et al.
Pubblicazione: (2024)
di: Ranasinghe, Nisal, et al.
Pubblicazione: (2024)
Deep Model Merging: The Sister of Neural Network Interpretability -- A Survey
di: Khan, Arham, et al.
Pubblicazione: (2024)
di: Khan, Arham, et al.
Pubblicazione: (2024)
Closed-Form Interpretation of Neural Network Latent Spaces with Symbolic Gradients
di: Wetzel, Sebastian J., et al.
Pubblicazione: (2024)
di: Wetzel, Sebastian J., et al.
Pubblicazione: (2024)
Framework GNN-AID: Graph Neural Network Analysis Interpretation and Defense
di: Lukyanov, Kirill, et al.
Pubblicazione: (2025)
di: Lukyanov, Kirill, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Studying Cross-cluster Modularity in Neural Networks
di: Golechha, Satvik, et al.
Pubblicazione: (2025) -
Challenges in Mechanistically Interpreting Model Representations
di: Golechha, Satvik, et al.
Pubblicazione: (2024) -
Progress Measures for Grokking on Real-world Tasks
di: Golechha, Satvik
Pubblicazione: (2024) -
Among Us: A Sandbox for Measuring and Detecting Agentic Deception
di: Golechha, Satvik, et al.
Pubblicazione: (2025) -
Relating Piecewise Linear Kolmogorov Arnold Networks to ReLU Networks
di: Schoots, Nandi, et al.
Pubblicazione: (2025)