Toward universal steering and monitoring of AI models
Fuente:
arXiv
Saved in:
| Main Authors: | Beaglehole, Daniel, Radhakrishnan, Adityanarayanan, Boix-Adserà, Enric, Belkin, Mikhail |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards a theory of model distillation
by: Boix-Adsera, Enric
Published: (2024)
by: Boix-Adsera, Enric
Published: (2024)
Secret mixtures of experts inside your LLM
by: Boix-Adsera, Enric
Published: (2025)
by: Boix-Adsera, Enric
Published: (2025)
On the inductive bias of infinite-depth ResNets and the bottleneck rank
by: Boix-Adsera, Enric
Published: (2025)
by: Boix-Adsera, Enric
Published: (2025)
The Features at Convergence Theorem: a first-principles alternative to the Neural Feature Ansatz for how networks learn representations
by: Boix-Adsera, Enric, et al.
Published: (2025)
by: Boix-Adsera, Enric, et al.
Published: (2025)
xRFM: Accurate, scalable, and interpretable feature learning models for tabular data
by: Beaglehole, Daniel, et al.
Published: (2025)
by: Beaglehole, Daniel, et al.
Published: (2025)
Let Me Think! A Long Chain-of-Thought Can Be Worth Exponentially Many Short Ones
by: Mirtaheri, Parsa, et al.
Published: (2025)
by: Mirtaheri, Parsa, et al.
Published: (2025)
The Weight Gram Matrix Captures Sequential Feature Linearization in Deep Networks
by: Cha, Taehun, et al.
Published: (2026)
by: Cha, Taehun, et al.
Published: (2026)
The power of fine-grained experts: Granularity boosts expressivity in Mixture of Experts
by: Boix-Adsera, Enric, et al.
Published: (2025)
by: Boix-Adsera, Enric, et al.
Published: (2025)
Contextual Linear Activation Steering of Language Models
by: Hsu, Brandon, et al.
Published: (2026)
by: Hsu, Brandon, et al.
Published: (2026)
When can transformers reason with abstract symbols?
by: Boix-Adsera, Enric, et al.
Published: (2023)
by: Boix-Adsera, Enric, et al.
Published: (2023)
Emergence in non-neural models: grokking modular arithmetic via average gradient outer product
by: Mallinar, Neil, et al.
Published: (2024)
by: Mallinar, Neil, et al.
Published: (2024)
Catching rationalization in the act: detecting motivated reasoning before and after CoT via activation probing
by: Mirtaheri, Parsa, et al.
Published: (2026)
by: Mirtaheri, Parsa, et al.
Published: (2026)
Linear Recursive Feature Machines provably recover low-rank matrices
by: Radhakrishnan, Adityanarayanan, et al.
Published: (2024)
by: Radhakrishnan, Adityanarayanan, et al.
Published: (2024)
Convergent Evolution: How Different Language Models Learn Similar Number Representations
by: Fu, Deqing, et al.
Published: (2026)
by: Fu, Deqing, et al.
Published: (2026)
Efficient semantic uncertainty quantification in language models via diversity-steered sampling
by: Park, Ji Won, et al.
Published: (2025)
by: Park, Ji Won, et al.
Published: (2025)
Quadratic models for understanding catapult dynamics of neural networks
by: Zhu, Libin, et al.
Published: (2022)
by: Zhu, Libin, et al.
Published: (2022)
Can sparse autoencoders be used to decompose and interpret steering vectors?
by: Mayne, Harry, et al.
Published: (2024)
by: Mayne, Harry, et al.
Published: (2024)
Context-Scaling versus Task-Scaling in In-Context Learning
by: Abedsoltan, Amirhesam, et al.
Published: (2024)
by: Abedsoltan, Amirhesam, et al.
Published: (2024)
Catapults in SGD: spikes in the training loss and their impact on generalization through feature learning
by: Zhu, Libin, et al.
Published: (2023)
by: Zhu, Libin, et al.
Published: (2023)
Towards Conversational Diagnostic AI
by: Tu, Tao, et al.
Published: (2024)
by: Tu, Tao, et al.
Published: (2024)
Efficient and accurate steering of Large Language Models through attention-guided feature learning
by: Davarmanesh, Parmida, et al.
Published: (2026)
by: Davarmanesh, Parmida, et al.
Published: (2026)
Merge to Mix: Mixing Datasets via Model Merging
by: Tao, Zhixu Silvia, et al.
Published: (2025)
by: Tao, Zhixu Silvia, et al.
Published: (2025)
Domain-level metacognitive monitoring in frontier LLMs: A 33-model atlas
by: Cacioli, Jon-Paul
Published: (2026)
by: Cacioli, Jon-Paul
Published: (2026)
Towards Conversational AI for Disease Management
by: Palepu, Anil, et al.
Published: (2025)
by: Palepu, Anil, et al.
Published: (2025)
Towards Execution-Grounded Automated AI Research
by: Si, Chenglei, et al.
Published: (2026)
by: Si, Chenglei, et al.
Published: (2026)
ClimateGPT: Towards AI Synthesizing Interdisciplinary Research on Climate Change
by: Thulke, David, et al.
Published: (2024)
by: Thulke, David, et al.
Published: (2024)
The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
by: Lu, Chris, et al.
Published: (2024)
by: Lu, Chris, et al.
Published: (2024)
Large Language Models in the Task of Automatic Validation of Text Classifier Predictions
by: Tsymbalov, Aleksandr, et al.
Published: (2025)
by: Tsymbalov, Aleksandr, et al.
Published: (2025)
RaguTeam at SemEval-2026 Task 8: Meno and Friends in a Judge-Orchestrated LLM Ensemble for Faithful Multi-Turn Response Generation
by: Bondarenko, Ivan, et al.
Published: (2026)
by: Bondarenko, Ivan, et al.
Published: (2026)
Towards Automated Patent Workflows: AI-Orchestrated Multi-Agent Framework for Intellectual Property Management and Analysis
by: Srinivas, Sakhinana Sagar, et al.
Published: (2024)
by: Srinivas, Sakhinana Sagar, et al.
Published: (2024)
AI-AI Bias: large language models favor communications generated by large language models
by: Laurito, Walter, et al.
Published: (2024)
by: Laurito, Walter, et al.
Published: (2024)
The merged-staircase property: a necessary and nearly sufficient condition for SGD learning of sparse functions on two-layer neural networks
by: Abbe, Emmanuel, et al.
Published: (2022)
by: Abbe, Emmanuel, et al.
Published: (2022)
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI
by: Yang, Chao, et al.
Published: (2024)
by: Yang, Chao, et al.
Published: (2024)
Feature learning as alignment: a structural property of gradient descent in non-linear neural networks
by: Beaglehole, Daniel, et al.
Published: (2024)
by: Beaglehole, Daniel, et al.
Published: (2024)
Out-of-Distribution Detection using Synthetic Data Generation
by: Abbas, Momin, et al.
Published: (2025)
by: Abbas, Momin, et al.
Published: (2025)
Wikontic: Constructing Wikidata-Aligned, Ontology-Aware Knowledge Graphs with Large Language Models
by: Chepurova, Alla, et al.
Published: (2025)
by: Chepurova, Alla, et al.
Published: (2025)
Limitations of Normalization in Attention Mechanism
by: Mudarisov, Timur, et al.
Published: (2025)
by: Mudarisov, Timur, et al.
Published: (2025)
Scaling Transformer to 1M tokens and beyond with RMT
by: Bulatov, Aydar, et al.
Published: (2023)
by: Bulatov, Aydar, et al.
Published: (2023)
AI and Generative AI for Research Discovery and Summarization
by: Glickman, Mark, et al.
Published: (2024)
by: Glickman, Mark, et al.
Published: (2024)
BEExAI: Benchmark to Evaluate Explainable AI
by: Sithakoul, Samuel, et al.
Published: (2024)
by: Sithakoul, Samuel, et al.
Published: (2024)
Similar Items
-
Towards a theory of model distillation
by: Boix-Adsera, Enric
Published: (2024) -
Secret mixtures of experts inside your LLM
by: Boix-Adsera, Enric
Published: (2025) -
On the inductive bias of infinite-depth ResNets and the bottleneck rank
by: Boix-Adsera, Enric
Published: (2025) -
The Features at Convergence Theorem: a first-principles alternative to the Neural Feature Ansatz for how networks learn representations
by: Boix-Adsera, Enric, et al.
Published: (2025) -
xRFM: Accurate, scalable, and interpretable feature learning models for tabular data
by: Beaglehole, Daniel, et al.
Published: (2025)