Hypothesis Testing the Circuit Hypothesis in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Shi, Claudia, Beltran-Velez, Nicolas, Nazaret, Achille, Zheng, Carolina, Garriga-Alonso, Adrià, Jesson, Andrew, Makar, Maggie, Blei, David M. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can Generative AI Solve Your In-Context Learning Problem? A Martingale Perspective
by: Jesson, Andrew, et al.
Published: (2024)
by: Jesson, Andrew, et al.
Published: (2024)
Extremely Greedy Equivalence Search
by: Nazaret, Achille, et al.
Published: (2025)
by: Nazaret, Achille, et al.
Published: (2025)
SynthSAEBench: Evaluating Sparse Autoencoders on Scalable Realistic Synthetic Data
by: Chanin, David, et al.
Published: (2026)
by: Chanin, David, et al.
Published: (2026)
Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
by: Chanin, David, et al.
Published: (2025)
by: Chanin, David, et al.
Published: (2025)
Among Us: A Sandbox for Measuring and Detecting Agentic Deception
by: Golechha, Satvik, et al.
Published: (2025)
by: Golechha, Satvik, et al.
Published: (2025)
Biases in the Blind Spot: Detecting What LLMs Fail to Mention
by: Arcuschin, Iván, et al.
Published: (2026)
by: Arcuschin, Iván, et al.
Published: (2026)
Treeffuser: Probabilistic Predictions via Conditional Diffusions with Gradient-Boosted Trees
by: Beltran-Velez, Nicolas, et al.
Published: (2024)
by: Beltran-Velez, Nicolas, et al.
Published: (2024)
Duel-Evolve: Reward-Free Test-Time Scaling via LLM Self-Preferences
by: Karlekar, Sweta, et al.
Published: (2026)
by: Karlekar, Sweta, et al.
Published: (2026)
Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders
by: Chanin, David, et al.
Published: (2025)
by: Chanin, David, et al.
Published: (2025)
Model Directions, Not Words: Mechanistic Topic Models Using Sparse Autoencoders
by: Zheng, Carolina, et al.
Published: (2025)
by: Zheng, Carolina, et al.
Published: (2025)
Investigating the Lottery Ticket Hypothesis for Variational Quantum Circuits
by: Kölle, Michael, et al.
Published: (2025)
by: Kölle, Michael, et al.
Published: (2025)
Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNN
by: Taufeeque, Mohammad, et al.
Published: (2025)
by: Taufeeque, Mohammad, et al.
Published: (2025)
Interpreting Emergent Planning in Model-Free Reinforcement Learning
by: Bush, Thomas, et al.
Published: (2025)
by: Bush, Thomas, et al.
Published: (2025)
Size-adaptive Hypothesis Testing for Fairness
by: Ferrara, Antonio, et al.
Published: (2025)
by: Ferrara, Antonio, et al.
Published: (2025)
Wrist Photoplethysmography Predicts Dietary Information
by: Verrier, Kyle, et al.
Published: (2025)
by: Verrier, Kyle, et al.
Published: (2025)
Stable Differentiable Causal Discovery
by: Nazaret, Achille, et al.
Published: (2023)
by: Nazaret, Achille, et al.
Published: (2023)
On the (In)feasibility of ML Backdoor Detection as an Hypothesis Testing Problem
by: Pichler, Georg, et al.
Published: (2024)
by: Pichler, Georg, et al.
Published: (2024)
E-valuator: Reliable Agent Verifiers with Sequential Hypothesis Testing
by: Sadhuka, Shuvom, et al.
Published: (2025)
by: Sadhuka, Shuvom, et al.
Published: (2025)
The Neural Pruning Law Hypothesis
by: Barbulescu, Eugen, et al.
Published: (2025)
by: Barbulescu, Eugen, et al.
Published: (2025)
On the Sparsity of the Strong Lottery Ticket Hypothesis
by: Natale, Emanuele, et al.
Published: (2024)
by: Natale, Emanuele, et al.
Published: (2024)
Scaling Law Hypothesis for Multimodal Model
by: Sun, Qingyun, et al.
Published: (2024)
by: Sun, Qingyun, et al.
Published: (2024)
DiFR: Inference Verification Despite Nondeterminism
by: Karvonen, Adam, et al.
Published: (2025)
by: Karvonen, Adam, et al.
Published: (2025)
Hypothesis Hunting with Evolving Networks of Autonomous Scientific Agents
by: Liu, Tennison, et al.
Published: (2025)
by: Liu, Tennison, et al.
Published: (2025)
Testing the Machine Consciousness Hypothesis
by: Fitz, Stephen
Published: (2025)
by: Fitz, Stephen
Published: (2025)
AutoSciDACT: Automated Scientific Discovery through Contrastive Embedding and Hypothesis Testing
by: Bright-Thonney, Samuel, et al.
Published: (2025)
by: Bright-Thonney, Samuel, et al.
Published: (2025)
f-INE: A Hypothesis Testing Framework for Estimating Influence under Training Randomness
by: Panda, Subhodip, et al.
Published: (2025)
by: Panda, Subhodip, et al.
Published: (2025)
Decoding Answers Before Chain-of-Thought: Evidence from Pre-CoT Probes and Activation Steering
by: Cox, Kyle, et al.
Published: (2026)
by: Cox, Kyle, et al.
Published: (2026)
The Two-Stage Decision-Sampling Hypothesis: Understanding the Emergence of Self-Reflection in RL-Trained LLMs
by: Zhao, Zibo, et al.
Published: (2026)
by: Zhao, Zibo, et al.
Published: (2026)
Planning in a recurrent neural network that plays Sokoban
by: Taufeeque, Mohammad, et al.
Published: (2024)
by: Taufeeque, Mohammad, et al.
Published: (2024)
Revisiting the Superficial Alignment Hypothesis
by: Raghavendra, Mohit, et al.
Published: (2024)
by: Raghavendra, Mohit, et al.
Published: (2024)
The Optimal Choice of Hypothesis Is the Weakest, Not the Shortest
by: Bennett, Michael Timothy
Published: (2023)
by: Bennett, Michael Timothy
Published: (2023)
A Black Swan Hypothesis: The Role of Human Irrationality in AI Safety
by: Lee, Hyunin, et al.
Published: (2024)
by: Lee, Hyunin, et al.
Published: (2024)
The Transfer Neurons Hypothesis: An Underlying Mechanism for Language Latent Space Transitions in Multilingual LLMs
by: Tezuka, Hinata, et al.
Published: (2025)
by: Tezuka, Hinata, et al.
Published: (2025)
Adversarial Circuit Evaluation
by: de Bos, Niels uit, et al.
Published: (2024)
by: de Bos, Niels uit, et al.
Published: (2024)
The Universal Weight Subspace Hypothesis
by: Kaushik, Prakhar, et al.
Published: (2025)
by: Kaushik, Prakhar, et al.
Published: (2025)
The Heterophilic Snowflake Hypothesis: Training and Empowering GNNs for Heterophilic Graphs
by: Wang, Kun, et al.
Published: (2024)
by: Wang, Kun, et al.
Published: (2024)
Risk Horizons: Structured Hypothesis Spaces for Longitudinal Clinical Prediction
by: Qu, Zhan, et al.
Published: (2026)
by: Qu, Zhan, et al.
Published: (2026)
Unveiling the Magic of Code Reasoning through Hypothesis Decomposition and Amendment
by: Zhao, Yuze, et al.
Published: (2025)
by: Zhao, Yuze, et al.
Published: (2025)
The Multiple Ticket Hypothesis: Random Sparse Subnetworks Suffice for RLVR
by: Adewuyi, Israel, et al.
Published: (2026)
by: Adewuyi, Israel, et al.
Published: (2026)
The Strong Lottery Ticket Hypothesis for Multi-Head Attention Mechanisms
by: Otsuka, Hikari, et al.
Published: (2025)
by: Otsuka, Hikari, et al.
Published: (2025)
Similar Items
-
Can Generative AI Solve Your In-Context Learning Problem? A Martingale Perspective
by: Jesson, Andrew, et al.
Published: (2024) -
Extremely Greedy Equivalence Search
by: Nazaret, Achille, et al.
Published: (2025) -
SynthSAEBench: Evaluating Sparse Autoencoders on Scalable Realistic Synthetic Data
by: Chanin, David, et al.
Published: (2026) -
Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
by: Chanin, David, et al.
Published: (2025) -
Among Us: A Sandbox for Measuring and Detecting Agentic Deception
by: Golechha, Satvik, et al.
Published: (2025)