Estimating the Probability of Sampling a Trained Neural Network at Random
Fuente:
arXiv
Saved in:
| Main Authors: | Scherlis, Adam, Belrose, Nora |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding Gradient Descent through the Training Jacobian
by: Belrose, Nora, et al.
Published: (2024)
by: Belrose, Nora, et al.
Published: (2024)
Refusal in LLMs is an Affine Function
by: Marshall, Thomas, et al.
Published: (2024)
by: Marshall, Thomas, et al.
Published: (2024)
Sparse Autoencoders Trained on the Same Data Learn Different Features
by: Paulo, Gonçalo, et al.
Published: (2025)
by: Paulo, Gonçalo, et al.
Published: (2025)
Evaluating SAE interpretability without explanations
by: Paulo, Gonçalo, et al.
Published: (2025)
by: Paulo, Gonçalo, et al.
Published: (2025)
Slowing Learning by Erasing Simple Features
by: Quirke, Lucia, et al.
Published: (2025)
by: Quirke, Lucia, et al.
Published: (2025)
Converting MLPs into Polynomials in Closed Form
by: Belrose, Nora, et al.
Published: (2025)
by: Belrose, Nora, et al.
Published: (2025)
Balancing Label Quantity and Quality for Scalable Elicitation
by: Mallen, Alex, et al.
Published: (2024)
by: Mallen, Alex, et al.
Published: (2024)
Partially Rewriting a Transformer in Natural Language
by: Paulo, Gonçalo, et al.
Published: (2025)
by: Paulo, Gonçalo, et al.
Published: (2025)
Neural Networks Learn Statistics of Increasing Complexity
by: Belrose, Nora, et al.
Published: (2024)
by: Belrose, Nora, et al.
Published: (2024)
Examining Two Hop Reasoning Through Information Content Scaling
by: Johnston, David, et al.
Published: (2025)
by: Johnston, David, et al.
Published: (2025)
Polysemanticity and Capacity in Neural Networks
by: Scherlis, Adam, et al.
Published: (2022)
by: Scherlis, Adam, et al.
Published: (2022)
Binary Sparse Coding for Interpretability
by: Quirke, Lucia, et al.
Published: (2025)
by: Quirke, Lucia, et al.
Published: (2025)
Transcoders Beat Sparse Autoencoders for Interpretability
by: Paulo, Gonçalo, et al.
Published: (2025)
by: Paulo, Gonçalo, et al.
Published: (2025)
Does Transformer Interpretability Transfer to RNNs?
by: Paulo, Gonçalo, et al.
Published: (2024)
by: Paulo, Gonçalo, et al.
Published: (2024)
Mechanistic Anomaly Detection for "Quirky" Language Models
by: Johnston, David O., et al.
Published: (2025)
by: Johnston, David O., et al.
Published: (2025)
Automatically Interpreting Millions of Features in Large Language Models
by: Paulo, Gonçalo, et al.
Published: (2024)
by: Paulo, Gonçalo, et al.
Published: (2024)
Eliciting Latent Knowledge from Quirky Language Models
by: Mallen, Alex, et al.
Published: (2023)
by: Mallen, Alex, et al.
Published: (2023)
Efficient Training of Deep Neural Operator Networks via Randomized Sampling
by: Karumuri, Sharmila, et al.
Published: (2024)
by: Karumuri, Sharmila, et al.
Published: (2024)
Sample Complexity Bounds for Estimating Probability Divergences under Invariances
by: Tahmasebi, Behrooz, et al.
Published: (2023)
by: Tahmasebi, Behrooz, et al.
Published: (2023)
Hierarchic Flows to Estimate and Sample High-dimensional Probabilities
by: Lempereur, Etienne, et al.
Published: (2024)
by: Lempereur, Etienne, et al.
Published: (2024)
ReLU Networks as Random Functions: Their Distribution in Probability Space
by: Chaudhari, Shreyas, et al.
Published: (2025)
by: Chaudhari, Shreyas, et al.
Published: (2025)
Deep Neural Network Training as Random Effects: An Optimization-Inference Duality
by: Yao, Minhao, et al.
Published: (2026)
by: Yao, Minhao, et al.
Published: (2026)
Gradient-Free Training of Recurrent Neural Networks using Random Perturbations
by: Fernandez, Jesus Garcia, et al.
Published: (2024)
by: Fernandez, Jesus Garcia, et al.
Published: (2024)
Finding Probably Approximate Optimal Solutions by Training to Estimate the Optimal Values of Subproblems
by: Megiddo, Nimrod, et al.
Published: (2025)
by: Megiddo, Nimrod, et al.
Published: (2025)
Training Guarantees of Neural Network Classification Two-Sample Tests by Kernel Analysis
by: Khurana, Varun, et al.
Published: (2024)
by: Khurana, Varun, et al.
Published: (2024)
LEACE: Perfect linear concept erasure in closed form
by: Belrose, Nora, et al.
Published: (2023)
by: Belrose, Nora, et al.
Published: (2023)
Eliciting Latent Predictions from Transformers with the Tuned Lens
by: Belrose, Nora, et al.
Published: (2023)
by: Belrose, Nora, et al.
Published: (2023)
SGM-PINN: Sampling Graphical Models for Faster Training of Physics-Informed Neural Networks
by: Anticev, John, et al.
Published: (2024)
by: Anticev, John, et al.
Published: (2024)
Improving the Finite Sample Estimation of Average Treatment Effects using Double/Debiased Machine Learning with Propensity Score Calibration
by: Ballinari, Daniele, et al.
Published: (2024)
by: Ballinari, Daniele, et al.
Published: (2024)
YOSO: You-Only-Sample-Once via Compressed Sensing for Graph Neural Network Training
by: Li, Yi, et al.
Published: (2024)
by: Li, Yi, et al.
Published: (2024)
How Many Training Samples Are Needed for the Inverse Kinematics Solutions by Artificial Neural Networks
by: Lim, Dong-Won
Published: (2026)
by: Lim, Dong-Won
Published: (2026)
Neighbor-Sampling Based Momentum Stochastic Methods for Training Graph Neural Networks
by: Noel, Molly, et al.
Published: (2025)
by: Noel, Molly, et al.
Published: (2025)
On the Dataless Training of Neural Networks
by: Velasquez, Alvaro, et al.
Published: (2025)
by: Velasquez, Alvaro, et al.
Published: (2025)
Probability Passing for Graph Neural Networks: Graph Structure and Representations Joint Learning
by: Wang, Ziyan, et al.
Published: (2024)
by: Wang, Ziyan, et al.
Published: (2024)
Forecasting Probability Distributions of Financial Returns with Deep Neural Networks
by: Michańków, Jakub
Published: (2025)
by: Michańków, Jakub
Published: (2025)
Gibbs Sampling the Posterior of Neural Networks
by: Piccioli, Giovanni, et al.
Published: (2023)
by: Piccioli, Giovanni, et al.
Published: (2023)
Distributed Matrix-Based Sampling for Graph Neural Network Training
by: Tripathy, Alok, et al.
Published: (2023)
by: Tripathy, Alok, et al.
Published: (2023)
Sample-Free Safety Assessment of Neural Network Controllers via Taylor Methods
by: Evans, Adam, et al.
Published: (2026)
by: Evans, Adam, et al.
Published: (2026)
Controlled Langevin Dynamics for Sampling of Feedforward Neural Networks Trained with Minibatches
by: Zambon, Alessandro, et al.
Published: (2026)
by: Zambon, Alessandro, et al.
Published: (2026)
Density Ratio Estimation with Conditional Probability Paths
by: Yu, Hanlin, et al.
Published: (2025)
by: Yu, Hanlin, et al.
Published: (2025)
Similar Items
-
Understanding Gradient Descent through the Training Jacobian
by: Belrose, Nora, et al.
Published: (2024) -
Refusal in LLMs is an Affine Function
by: Marshall, Thomas, et al.
Published: (2024) -
Sparse Autoencoders Trained on the Same Data Learn Different Features
by: Paulo, Gonçalo, et al.
Published: (2025) -
Evaluating SAE interpretability without explanations
by: Paulo, Gonçalo, et al.
Published: (2025) -
Slowing Learning by Erasing Simple Features
by: Quirke, Lucia, et al.
Published: (2025)