Position: Capability Control Should be a Separate Goal From Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Siddiqui, Shoaib Ahmed, Triantafillou, Eleni, Krueger, David, Weller, Adrian |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Dormant to Deleted: Tamper-Resistant Unlearning Through Weight-Space Regularization
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2025)
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2025)
On Evaluating LLMs' Capabilities as Functional Approximators: A Bayesian Perspective
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2024)
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2024)
Blockwise Self-Supervised Learning at Scale
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2023)
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2023)
A deeper look at depth pruning of LLMs
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2024)
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2024)
Exploring the design space of deep-learning-based weather forecasting systems
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2024)
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2024)
The Topological Trouble With Transformers
by: Mozer, Michael C., et al.
Published: (2026)
by: Mozer, Michael C., et al.
Published: (2026)
Permissive Information-Flow Analysis for Large Language Models
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2024)
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2024)
Position: AI Evaluations Should be Grounded on a Theory of Capability
by: Jo, Nathanael, et al.
Published: (2025)
by: Jo, Nathanael, et al.
Published: (2025)
Step-resolved data attribution for looped transformers
by: Kaissis, Georgios, et al.
Published: (2026)
by: Kaissis, Georgios, et al.
Published: (2026)
Your Privacy Depends on Others: Collusion Vulnerabilities in Individual Differential Privacy
by: Kaiser, Johannes, et al.
Published: (2026)
by: Kaiser, Johannes, et al.
Published: (2026)
Estimation of Concept Explanations Should be Uncertainty Aware
by: Piratla, Vihari, et al.
Published: (2023)
by: Piratla, Vihari, et al.
Published: (2023)
Protecting against simultaneous data poisoning attacks
by: Alex, Neel, et al.
Published: (2024)
by: Alex, Neel, et al.
Published: (2024)
Position: Weight Space Should Be a First-Class Generative AI Modality
by: Wang, Zhangyang, et al.
Published: (2026)
by: Wang, Zhangyang, et al.
Published: (2026)
To Each (Textual Sequence) Its Own: Improving Memorized-Data Unlearning in Large Language Models
by: Barbulescu, George-Octavian, et al.
Published: (2024)
by: Barbulescu, George-Octavian, et al.
Published: (2024)
On the Limitations and Capabilities of Position Embeddings for Length Generalization
by: Chen, Yang, et al.
Published: (2025)
by: Chen, Yang, et al.
Published: (2025)
Rotary Position Encodings for Graphs
by: Reid, Isaac, et al.
Published: (2025)
by: Reid, Isaac, et al.
Published: (2025)
Position: Beyond Euclidean -- Foundation Models Should Embrace Non-Euclidean Geometries
by: He, Neil, et al.
Published: (2025)
by: He, Neil, et al.
Published: (2025)
GRAIL: Goal Recognition Alignment through Imitation Learning
by: Elhadad, Osher, et al.
Published: (2026)
by: Elhadad, Osher, et al.
Published: (2026)
ALVIN: Active Learning Via INterpolation
by: Korakakis, Michalis, et al.
Published: (2024)
by: Korakakis, Michalis, et al.
Published: (2024)
Mitigating Shortcut Learning with InterpoLated Learning
by: Korakakis, Michalis, et al.
Published: (2025)
by: Korakakis, Michalis, et al.
Published: (2025)
Position: AI Security Policy Should Target Systems, Not Models
by: Riegler, Michael A., et al.
Published: (2026)
by: Riegler, Michael A., et al.
Published: (2026)
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs
by: Song, Xiangchen, et al.
Published: (2025)
by: Song, Xiangchen, et al.
Published: (2025)
Getting By Goal Misgeneralization With a Little Help From a Mentor
by: Trinh, Tu, et al.
Published: (2024)
by: Trinh, Tu, et al.
Published: (2024)
1000 Layer Networks for Self-Supervised RL: Scaling Depth Can Enable New Goal-Reaching Capabilities
by: Wang, Kevin, et al.
Published: (2025)
by: Wang, Kevin, et al.
Published: (2025)
From Tokenizer Bias to Backbone Capability: A Controlled Study of LLMs for Time Series Forecasting
by: Zhang, Xinyu, et al.
Published: (2025)
by: Zhang, Xinyu, et al.
Published: (2025)
You Don't Need All That Attention: Surgical Memorization Mitigation in Text-to-Image Diffusion Models
by: Zhao, Kairan, et al.
Published: (2026)
by: Zhao, Kairan, et al.
Published: (2026)
Controlling Grokking with Nonlinearity and Data Symmetry
by: Salah, Ahmed, et al.
Published: (2024)
by: Salah, Ahmed, et al.
Published: (2024)
Noisy-Pair Robust Representation Alignment for Positive-Unlabeled Learning
by: Zhao, Hengwei, et al.
Published: (2025)
by: Zhao, Hengwei, et al.
Published: (2025)
The Master Key Hypothesis: Unlocking Cross-Model Capability Transfer via Linear Subspace Alignment
by: Balasubramanian, Rishab, et al.
Published: (2026)
by: Balasubramanian, Rishab, et al.
Published: (2026)
Why Deep Jacobian Spectra Separate: Depth-Induced Scaling and Singular-Vector Alignment
by: Haas, Nathanaël, et al.
Published: (2026)
by: Haas, Nathanaël, et al.
Published: (2026)
Graph Alignment for Benchmarking Graph Neural Networks and Learning Positional Encodings
by: Lagesse, Adrien, et al.
Published: (2025)
by: Lagesse, Adrien, et al.
Published: (2025)
Why Goal-Conditioned Reinforcement Learning Works: Relation to Dual Control
by: Lawrence, Nathan P., et al.
Published: (2025)
by: Lawrence, Nathan P., et al.
Published: (2025)
Preference Goal Tuning: Post-Training as Latent Control for Frozen Policies
by: Zhao, Guangyu, et al.
Published: (2024)
by: Zhao, Guangyu, et al.
Published: (2024)
Position: Machine Learning Conferences Should Establish a "Refutations and Critiques" Track
by: Schaeffer, Rylan, et al.
Published: (2025)
by: Schaeffer, Rylan, et al.
Published: (2025)
Token Reduction Should Go Beyond Efficiency in Generative Models -- From Vision, Language to Multimodality
by: Kong, Zhenglun, et al.
Published: (2025)
by: Kong, Zhenglun, et al.
Published: (2025)
How Should We Meta-Learn Reinforcement Learning Algorithms?
by: Goldie, Alexander David, et al.
Published: (2025)
by: Goldie, Alexander David, et al.
Published: (2025)
Learning to Decide with AI Assistance under Human-Alignment
by: Benz, Nina Corvelo, et al.
Published: (2026)
by: Benz, Nina Corvelo, et al.
Published: (2026)
Dense and Diverse Goal Coverage in Multi Goal Reinforcement Learning
by: Singh, Sagalpreet, et al.
Published: (2025)
by: Singh, Sagalpreet, et al.
Published: (2025)
Implicit meta-learning may lead language models to trust more reliable sources
by: Krasheninnikov, Dmitrii, et al.
Published: (2023)
by: Krasheninnikov, Dmitrii, et al.
Published: (2023)
Autonomous Learning From Success and Failure: Goal-Conditioned Supervised Learning with Negative Feedback
by: Zhang, Zeqiang, et al.
Published: (2025)
by: Zhang, Zeqiang, et al.
Published: (2025)
Similar Items
-
From Dormant to Deleted: Tamper-Resistant Unlearning Through Weight-Space Regularization
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2025) -
On Evaluating LLMs' Capabilities as Functional Approximators: A Bayesian Perspective
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2024) -
Blockwise Self-Supervised Learning at Scale
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2023) -
A deeper look at depth pruning of LLMs
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2024) -
Exploring the design space of deep-learning-based weather forecasting systems
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2024)