Neural Networks Learn Statistics of Increasing Complexity
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Belrose, Nora, Pope, Quintin, Quirke, Lucia, Mallen, Alex, Fern, Xiaoli |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Slowing Learning by Erasing Simple Features
von: Quirke, Lucia, et al.
Veröffentlicht: (2025)
von: Quirke, Lucia, et al.
Veröffentlicht: (2025)
Balancing Label Quantity and Quality for Scalable Elicitation
von: Mallen, Alex, et al.
Veröffentlicht: (2024)
von: Mallen, Alex, et al.
Veröffentlicht: (2024)
Binary Sparse Coding for Interpretability
von: Quirke, Lucia, et al.
Veröffentlicht: (2025)
von: Quirke, Lucia, et al.
Veröffentlicht: (2025)
Automatically Interpreting Millions of Features in Large Language Models
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2024)
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2024)
Eliciting Latent Knowledge from Quirky Language Models
von: Mallen, Alex, et al.
Veröffentlicht: (2023)
von: Mallen, Alex, et al.
Veröffentlicht: (2023)
Estimating the Probability of Sampling a Trained Neural Network at Random
von: Scherlis, Adam, et al.
Veröffentlicht: (2025)
von: Scherlis, Adam, et al.
Veröffentlicht: (2025)
Sparse Autoencoders Trained on the Same Data Learn Different Features
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025)
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025)
Understanding Gradient Descent through the Training Jacobian
von: Belrose, Nora, et al.
Veröffentlicht: (2024)
von: Belrose, Nora, et al.
Veröffentlicht: (2024)
Evaluating SAE interpretability without explanations
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025)
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025)
Converting MLPs into Polynomials in Closed Form
von: Belrose, Nora, et al.
Veröffentlicht: (2025)
von: Belrose, Nora, et al.
Veröffentlicht: (2025)
Partially Rewriting a Transformer in Natural Language
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025)
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025)
Examining Two Hop Reasoning Through Information Content Scaling
von: Johnston, David, et al.
Veröffentlicht: (2025)
von: Johnston, David, et al.
Veröffentlicht: (2025)
Transcoders Beat Sparse Autoencoders for Interpretability
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025)
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025)
Refusal in LLMs is an Affine Function
von: Marshall, Thomas, et al.
Veröffentlicht: (2024)
von: Marshall, Thomas, et al.
Veröffentlicht: (2024)
Does Transformer Interpretability Transfer to RNNs?
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2024)
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2024)
Mechanistic Anomaly Detection for "Quirky" Language Models
von: Johnston, David O., et al.
Veröffentlicht: (2025)
von: Johnston, David O., et al.
Veröffentlicht: (2025)
Automatically Finding and Validating Unexpected Side-Effects of Interventions on Language Models
von: Pope, Quintin, et al.
Veröffentlicht: (2026)
von: Pope, Quintin, et al.
Veröffentlicht: (2026)
Understanding Addition in Transformers
von: Quirke, Philip, et al.
Veröffentlicht: (2023)
von: Quirke, Philip, et al.
Veröffentlicht: (2023)
ComplexityNet: Increasing LLM Inference Efficiency by Learning Task Complexity
von: Bae, Henry, et al.
Veröffentlicht: (2023)
von: Bae, Henry, et al.
Veröffentlicht: (2023)
Understanding Addition and Subtraction in Transformers
von: Quirke, Philip, et al.
Veröffentlicht: (2024)
von: Quirke, Philip, et al.
Veröffentlicht: (2024)
Graph Neural Network Based Action Ranking for Planning
von: Mangannavar, Rajesh, et al.
Veröffentlicht: (2024)
von: Mangannavar, Rajesh, et al.
Veröffentlicht: (2024)
Active Learning with Neural Networks: Insights from Nonparametric Statistics
von: Zhu, Yinglun, et al.
Veröffentlicht: (2022)
von: Zhu, Yinglun, et al.
Veröffentlicht: (2022)
Hybrid Real- and Complex-valued Neural Network Architecture
von: Young, Alex, et al.
Veröffentlicht: (2025)
von: Young, Alex, et al.
Veröffentlicht: (2025)
LEACE: Perfect linear concept erasure in closed form
von: Belrose, Nora, et al.
Veröffentlicht: (2023)
von: Belrose, Nora, et al.
Veröffentlicht: (2023)
Eliciting Latent Predictions from Transformers with the Tuned Lens
von: Belrose, Nora, et al.
Veröffentlicht: (2023)
von: Belrose, Nora, et al.
Veröffentlicht: (2023)
Statistical Learning Analysis of Physics-Informed Neural Networks
von: Barajas-Solano, David A.
Veröffentlicht: (2026)
von: Barajas-Solano, David A.
Veröffentlicht: (2026)
Subversion Strategy Eval: Can language models statelessly strategize to subvert control protocols?
von: Mallen, Alex, et al.
Veröffentlicht: (2024)
von: Mallen, Alex, et al.
Veröffentlicht: (2024)
Transfer Learning via Auxiliary Labels with Application to Cold-Hardiness Prediction
von: Goebel, Kristen, et al.
Veröffentlicht: (2025)
von: Goebel, Kristen, et al.
Veröffentlicht: (2025)
Why Do Some Language Models Fake Alignment While Others Don't?
von: Sheshadri, Abhay, et al.
Veröffentlicht: (2025)
von: Sheshadri, Abhay, et al.
Veröffentlicht: (2025)
Rethinking Deep Learning: Propagating Information in Neural Networks without Backpropagation and Statistical Optimization
von: Itoh, Kei
Veröffentlicht: (2024)
von: Itoh, Kei
Veröffentlicht: (2024)
Learning to Compile Programs to Neural Networks
von: Weber, Logan, et al.
Veröffentlicht: (2024)
von: Weber, Logan, et al.
Veröffentlicht: (2024)
Understanding Scaling Laws with Statistical and Approximation Theory for Transformer Neural Networks on Intrinsically Low-dimensional Data
von: Havrilla, Alex, et al.
Veröffentlicht: (2024)
von: Havrilla, Alex, et al.
Veröffentlicht: (2024)
Learned Random Label Predictions as a Neural Network Complexity Metric
von: Becker, Marlon, et al.
Veröffentlicht: (2024)
von: Becker, Marlon, et al.
Veröffentlicht: (2024)
Provable Guarantees for Nonlinear Feature Learning in Three-Layer Neural Networks
von: Nichani, Eshaan, et al.
Veröffentlicht: (2023)
von: Nichani, Eshaan, et al.
Veröffentlicht: (2023)
Increasing the Confidence of Deep Neural Networks by Coverage Analysis
von: Rossolini, Giulio, et al.
Veröffentlicht: (2021)
von: Rossolini, Giulio, et al.
Veröffentlicht: (2021)
Characterizing Learning in Deep Neural Networks using Tractable Algorithmic Complexity Analysis
von: Bakhtiarifard, Pedram, et al.
Veröffentlicht: (2026)
von: Bakhtiarifard, Pedram, et al.
Veröffentlicht: (2026)
DeepAveragers: Offline Reinforcement Learning by Solving Derived Non-Parametric MDPs
von: Shrestha, Aayam, et al.
Veröffentlicht: (2020)
von: Shrestha, Aayam, et al.
Veröffentlicht: (2020)
A Manifold Perspective on the Statistical Generalization of Graph Neural Networks
von: Wang, Zhiyang, et al.
Veröffentlicht: (2024)
von: Wang, Zhiyang, et al.
Veröffentlicht: (2024)
Budgeted Online Active Learning with Expert Advice and Episodic Priors
von: Goebel, Kristen, et al.
Veröffentlicht: (2025)
von: Goebel, Kristen, et al.
Veröffentlicht: (2025)
Steinmetz Neural Networks for Complex-Valued Data
von: Venkatasubramanian, Shyam, et al.
Veröffentlicht: (2024)
von: Venkatasubramanian, Shyam, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Slowing Learning by Erasing Simple Features
von: Quirke, Lucia, et al.
Veröffentlicht: (2025) -
Balancing Label Quantity and Quality for Scalable Elicitation
von: Mallen, Alex, et al.
Veröffentlicht: (2024) -
Binary Sparse Coding for Interpretability
von: Quirke, Lucia, et al.
Veröffentlicht: (2025) -
Automatically Interpreting Millions of Features in Large Language Models
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2024) -
Eliciting Latent Knowledge from Quirky Language Models
von: Mallen, Alex, et al.
Veröffentlicht: (2023)