Out of Control -- Why Alignment Needs Formal Control Theory (and an Alignment Control Stack)
Fuente:
arXiv
Guardado en:
| Autor principal: | Perrier, Elija |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Operationalising Extended Cognition: Formal Metrics for Corporate Knowledge and Legal Accountability
por: Perrier, Elija
Publicado: (2025)
por: Perrier, Elija
Publicado: (2025)
Hamiltonian Formalism for Comparing Quantum and Classical Intelligence
por: Perrier, Elija
Publicado: (2025)
por: Perrier, Elija
Publicado: (2025)
Towards Measurement Theory for Artificial Intelligence
por: Perrier, Elija
Publicado: (2025)
por: Perrier, Elija
Publicado: (2025)
Typed Chain-of-Thought: A Curry-Howard Framework for Verifying LLM Reasoning
por: Perrier, Elija
Publicado: (2025)
por: Perrier, Elija
Publicado: (2025)
Statistical Scenario Modelling and Lookalike Distributions for Multi-Variate AI Risk
por: Perrier, Elija
Publicado: (2025)
por: Perrier, Elija
Publicado: (2025)
Deconstructing Superintelligence: Identity, Self-Modification and Différance
por: Perrier, Elija
Publicado: (2026)
por: Perrier, Elija
Publicado: (2026)
K-P Quantum Neural Networks
por: Perrier, Elija
Publicado: (2025)
por: Perrier, Elija
Publicado: (2025)
Quantum AIXI: Universal Intelligence via Quantum Information
por: Perrier, Elija
Publicado: (2025)
por: Perrier, Elija
Publicado: (2025)
Threshold Crossings as Tail Events for Catastrophic AI Risk
por: Perrier, Elija
Publicado: (2025)
por: Perrier, Elija
Publicado: (2025)
Watts-per-Intelligence Part II: Algorithmic Catalysis
por: Perrier, Elija
Publicado: (2026)
por: Perrier, Elija
Publicado: (2026)
Post-AGI Economies: Autonomy and the First Fundamental Theorem of Welfare Economics
por: Perrier, Elija
Publicado: (2026)
por: Perrier, Elija
Publicado: (2026)
Time, Identity and Consciousness in Language Model Agents
por: Perrier, Elija, et al.
Publicado: (2026)
por: Perrier, Elija, et al.
Publicado: (2026)
Quantum AGI: Ontological Foundations
por: Perrier, Elija, et al.
Publicado: (2025)
por: Perrier, Elija, et al.
Publicado: (2025)
Position: Stop Acting Like Language Model Agents Are Normal Agents
por: Perrier, Elija, et al.
Publicado: (2025)
por: Perrier, Elija, et al.
Publicado: (2025)
Agent Identity Evals: Measuring Agentic Identity
por: Perrier, Elija, et al.
Publicado: (2025)
por: Perrier, Elija, et al.
Publicado: (2025)
Beyond Ordinal Preferences: Why Alignment Needs Cardinal Human Feedback
por: Whitfill, Parker, et al.
Publicado: (2025)
por: Whitfill, Parker, et al.
Publicado: (2025)
MOSAIC: Composable Safety Alignment with Modular Control Tokens
por: Peng, Jingyu, et al.
Publicado: (2026)
por: Peng, Jingyu, et al.
Publicado: (2026)
Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment
por: Guo, Yiju, et al.
Publicado: (2024)
por: Guo, Yiju, et al.
Publicado: (2024)
CAP: Controllable Alignment Prompting for Unlearning in LLMs
por: Wang, Zhaokun, et al.
Publicado: (2026)
por: Wang, Zhaokun, et al.
Publicado: (2026)
CBF-LLM: Safe Control for LLM Alignment
por: Miyaoka, Yuya, et al.
Publicado: (2024)
por: Miyaoka, Yuya, et al.
Publicado: (2024)
Collab: Controlled Decoding using Mixture of Agents for LLM Alignment
por: Chakraborty, Souradip, et al.
Publicado: (2025)
por: Chakraborty, Souradip, et al.
Publicado: (2025)
Safe Transformer: An Explicit Safety Bit For Interpretable And Controllable Alignment
por: Feng, Jingyuan, et al.
Publicado: (2026)
por: Feng, Jingyuan, et al.
Publicado: (2026)
Universe Routing: Why Self-Evolving Agents Need Epistemic Control
por: Wang, Zhaohui Geoffrey
Publicado: (2026)
por: Wang, Zhaohui Geoffrey
Publicado: (2026)
Watts-Per-Intelligence: Part I (Energy Efficiency)
por: Perrier, Elija
Publicado: (2025)
por: Perrier, Elija
Publicado: (2025)
Quantum Geometric Machine Learning
por: Perrier, Elija
Publicado: (2024)
por: Perrier, Elija
Publicado: (2024)
Position: Assistive Agents Need Accessibility Alignment
por: Hu, Jie, et al.
Publicado: (2026)
por: Hu, Jie, et al.
Publicado: (2026)
Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance
por: Jin, Luozhijie, et al.
Publicado: (2025)
por: Jin, Luozhijie, et al.
Publicado: (2025)
Position: Capability Control Should be a Separate Goal From Alignment
por: Siddiqui, Shoaib Ahmed, et al.
Publicado: (2026)
por: Siddiqui, Shoaib Ahmed, et al.
Publicado: (2026)
SudoLM: Learning Access Control of Parametric Knowledge with Authorization Alignment
por: Liu, Qin, et al.
Publicado: (2024)
por: Liu, Qin, et al.
Publicado: (2024)
Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements
por: Zhang, Jingyu, et al.
Publicado: (2024)
por: Zhang, Jingyu, et al.
Publicado: (2024)
Analogous Alignments: Digital "Formally" meets Analog
por: Mohanty, Hansa, et al.
Publicado: (2024)
por: Mohanty, Hansa, et al.
Publicado: (2024)
GAC: Stabilizing Asynchronous RL Training for LLMs via Gradient Alignment Control
por: Xu, Haofeng, et al.
Publicado: (2026)
por: Xu, Haofeng, et al.
Publicado: (2026)
C-MORAL: Controllable Multi-Objective Molecular Optimization with Reinforcement Alignment for LLMs
por: Gao, Rui, et al.
Publicado: (2026)
por: Gao, Rui, et al.
Publicado: (2026)
The Geometry of Compromise: Unlocking Generative Capabilities via Controllable Modality Alignment
por: Liu, Hongyuan, et al.
Publicado: (2026)
por: Liu, Hongyuan, et al.
Publicado: (2026)
Expected Value Alignment for Generative Reward Modeling in Formal Mathematics Verification
por: Ji, Shihao, et al.
Publicado: (2026)
por: Ji, Shihao, et al.
Publicado: (2026)
JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment
por: Liu, Renhang, et al.
Publicado: (2025)
por: Liu, Renhang, et al.
Publicado: (2025)
Maximize Your Diffusion: A Study into Reward Maximization and Alignment for Diffusion-based Control
por: Huh, Dom, et al.
Publicado: (2025)
por: Huh, Dom, et al.
Publicado: (2025)
Medoid Prototype Alignment for Cross-Plant Unknown Attack Detection in Industrial Control Systems
por: Wang, Luyao
Publicado: (2026)
por: Wang, Luyao
Publicado: (2026)
Infrastructure for AI Agents
por: Chan, Alan, et al.
Publicado: (2025)
por: Chan, Alan, et al.
Publicado: (2025)
How Alignment Routes: Localizing, Scaling, and Controlling Policy Circuits in Language Models
por: Frank, Gregory N.
Publicado: (2026)
por: Frank, Gregory N.
Publicado: (2026)
Ejemplares similares
-
Operationalising Extended Cognition: Formal Metrics for Corporate Knowledge and Legal Accountability
por: Perrier, Elija
Publicado: (2025) -
Hamiltonian Formalism for Comparing Quantum and Classical Intelligence
por: Perrier, Elija
Publicado: (2025) -
Towards Measurement Theory for Artificial Intelligence
por: Perrier, Elija
Publicado: (2025) -
Typed Chain-of-Thought: A Curry-Howard Framework for Verifying LLM Reasoning
por: Perrier, Elija
Publicado: (2025) -
Statistical Scenario Modelling and Lookalike Distributions for Multi-Variate AI Risk
por: Perrier, Elija
Publicado: (2025)