The Specification Trap: Why Static Value Alignment Alone Is Insufficient for Robust Alignment
Fuente:
arXiv
Saved in:
| Main Author: | Spizzirri, Austin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Machine Learning as Performative Materialist Practice: Thirteen Theses on the Epistemology, Methodology, and Politics of Applied ML
by: De Unánue, Adolfo, et al.
Published: (2026)
by: De Unánue, Adolfo, et al.
Published: (2026)
Revenue-Sharing as Infrastructure: A Distributed Business Model for Generative AI Platforms
by: Mondjo, Ghislain Dorian Tchuente
Published: (2026)
by: Mondjo, Ghislain Dorian Tchuente
Published: (2026)
Alpay Algebra III: Observer-Coupled Collapse and the Temporal Drift of Identity
by: Alpay, Faruk
Published: (2025)
by: Alpay, Faruk
Published: (2025)
Cognitive Castes: Artificial Intelligence, Epistemic Stratification, and the Dissolution of Democratic Discourse
by: Wright, Craig S
Published: (2025)
by: Wright, Craig S
Published: (2025)
Modeling the Path of Structural Strategic Deterrence: A Sand Table Simulation and Research Report on China's Military-Industrial Capability System against the United States Based on Rare Earth Supply Disconnection
by: Meng, Wei
Published: (2025)
by: Meng, Wei
Published: (2025)
Expert Insight-Based Modeling of Non-Kinetic Strategic Deterrence of Rare Earth Supply Disruption:A Simulation-Driven Systematic Framework
by: Meng, Wei
Published: (2025)
by: Meng, Wei
Published: (2025)
Memory-Augmented State Machine Prompting: A Novel LLM Agent Framework for Real-Time Strategy Games
by: Qi, Runnan, et al.
Published: (2025)
by: Qi, Runnan, et al.
Published: (2025)
Cooperation After the Algorithm: Designing Human-AI Coexistence Beyond the Illusion of Collaboration
by: Codreanu, Tatia
Published: (2026)
by: Codreanu, Tatia
Published: (2026)
Intersymbolic AI: Interlinking Symbolic AI and Subsymbolic AI
by: Platzer, André
Published: (2024)
by: Platzer, André
Published: (2024)
Formal Proofs as Structured Explanations: Proposing Several Tasks on Explainable Natural Language Inference
by: Abzianidze, Lasha
Published: (2023)
by: Abzianidze, Lasha
Published: (2023)
Fixed-Point Traps and Identity Emergence in Educational Feedback Systems
by: Alpay, Faruk
Published: (2025)
by: Alpay, Faruk
Published: (2025)
Civilizational Metamaterials: Engineering Coordination Under Capability Gradients and Structural Turbulence
by: Orban, David
Published: (2026)
by: Orban, David
Published: (2026)
Strategic Counterfactual Modeling of Deep-Target Airstrike Systems via Intervention-Aware Spatio-Causal Graph Networks
by: Meng, Wei
Published: (2025)
by: Meng, Wei
Published: (2025)
Online Decision Making with Generative Action Sets
by: Xu, Jianyu, et al.
Published: (2025)
by: Xu, Jianyu, et al.
Published: (2025)
Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback
by: Conitzer, Vincent, et al.
Published: (2024)
by: Conitzer, Vincent, et al.
Published: (2024)
Coopetition-Gym v1: A Formally Grounded Platform for Mixed-Motive Multi-Agent Reinforcement Learning under Strategic Coopetition
by: Pant, Vik, et al.
Published: (2026)
by: Pant, Vik, et al.
Published: (2026)
Hilbert-Geo: Solving Solid Geometric Problems by Neural-Symbolic Reasoning
by: Xu, Ruoran, et al.
Published: (2026)
by: Xu, Ruoran, et al.
Published: (2026)
Modeling and Visualization Reasoning for Stakeholders in Education and Industry Integration Systems: Research on Structured Synthetic Dialogue Data Generation Based on NIST Standards
by: Meng, Wei
Published: (2025)
by: Meng, Wei
Published: (2025)
Standardized Threat Taxonomy for AI Security, Governance, and Regulatory Compliance
by: Huwyler, Hernan
Published: (2025)
by: Huwyler, Hernan
Published: (2025)
Informed AI Regulation: Comparing the Ethical Frameworks of Leading LLM Chatbots Using an Ethics-Based Audit to Assess Moral Reasoning and Normative Values
by: Chun, Jon, et al.
Published: (2024)
by: Chun, Jon, et al.
Published: (2024)
Bridging Voting and Deliberation with Algorithms: Field Insights from vTaiwan and Kultur Komitee
by: Yang, Joshua C., et al.
Published: (2025)
by: Yang, Joshua C., et al.
Published: (2025)
Yanasse: Finding New Proofs from Deep Vision's Analogies, Part 1
by: Linhares, Alexandre
Published: (2026)
by: Linhares, Alexandre
Published: (2026)
Fixed-Point Theorems and the Ethics of Radical Transparency: A Logic-First Treatment
by: Alpay, Faruk, et al.
Published: (2025)
by: Alpay, Faruk, et al.
Published: (2025)
Value-Aware Multiagent Systems
by: Osman, Nardine
Published: (2025)
by: Osman, Nardine
Published: (2025)
Approximating Discrimination Within Models When Faced With Several Non-Binary Sensitive Attributes
by: Bian, Yijun, et al.
Published: (2024)
by: Bian, Yijun, et al.
Published: (2024)
Does Machine Bring in Extra Bias in Learning? Approximating Fairness in Models Promptly
by: Bian, Yijun, et al.
Published: (2024)
by: Bian, Yijun, et al.
Published: (2024)
Interpretation, Learning, and Empathy as One Constraint: A Residual-Adequacy Architecture with Accountable Abstention
by: Amornbunchornvej, Chainarong
Published: (2026)
by: Amornbunchornvej, Chainarong
Published: (2026)
Logical Modalities within the European AI Act: An Analysis
by: Lawniczak, Lara, et al.
Published: (2025)
by: Lawniczak, Lara, et al.
Published: (2025)
Prices, Bids, Values: One ML-Powered Combinatorial Auction to Rule Them All
by: Soumalias, Ermis, et al.
Published: (2024)
by: Soumalias, Ermis, et al.
Published: (2024)
A Monad-Based Clause Architecture for Artificial Age Score (AAS) in Large Language Models
by: Kayadibi, Seyma Yaman
Published: (2025)
by: Kayadibi, Seyma Yaman
Published: (2025)
How Worrying Are Privacy Attacks Against Machine Learning?
by: Domingo-Ferrer, Josep
Published: (2025)
by: Domingo-Ferrer, Josep
Published: (2025)
Beyond Reward Suppression: Reshaping Steganographic Communication Protocols in MARL via Dynamic Representational Circuit Breaking
by: Ming, Liu Hung
Published: (2026)
by: Ming, Liu Hung
Published: (2026)
$ϕ^{\infty}$: Clause Purification, Embedding Realignment, and the Total Suppression of the Em Dash in Autoregressive Language Models
by: Kilictas, Bugra, et al.
Published: (2025)
by: Kilictas, Bugra, et al.
Published: (2025)
Reinsuring AI: Energy, Agriculture, Finance & Medicine as Precedents for Scalable Governance of Frontier Artificial Intelligence
by: Stetler, Nicholas
Published: (2025)
by: Stetler, Nicholas
Published: (2025)
Transforming Business with Generative AI: Research, Innovation, Market Deployment and Future Shifts in Business Models
by: Singh, Narotam, et al.
Published: (2024)
by: Singh, Narotam, et al.
Published: (2024)
Interpretability Guarantees with Merlin-Arthur Classifiers
by: Wäldchen, Stephan, et al.
Published: (2022)
by: Wäldchen, Stephan, et al.
Published: (2022)
Emergent Coordination in Multi-Agent Systems via Pressure Fields and Temporal Decay
by: Rodriguez, Roland
Published: (2026)
by: Rodriguez, Roland
Published: (2026)
Black Box Deployed -- Functional Criteria for Artificial Moral Agents in the LLM Era
by: Brophy, Matthew E.
Published: (2025)
by: Brophy, Matthew E.
Published: (2025)
The Axiom of Consent: Friction Dynamics in Multi-Agent Coordination
by: Farzulla, Murad
Published: (2026)
by: Farzulla, Murad
Published: (2026)
Solving Zebra Puzzles Using Constraint-Guided Multi-Agent Systems
by: Berman, Shmuel, et al.
Published: (2024)
by: Berman, Shmuel, et al.
Published: (2024)
Similar Items
-
Machine Learning as Performative Materialist Practice: Thirteen Theses on the Epistemology, Methodology, and Politics of Applied ML
by: De Unánue, Adolfo, et al.
Published: (2026) -
Revenue-Sharing as Infrastructure: A Distributed Business Model for Generative AI Platforms
by: Mondjo, Ghislain Dorian Tchuente
Published: (2026) -
Alpay Algebra III: Observer-Coupled Collapse and the Temporal Drift of Identity
by: Alpay, Faruk
Published: (2025) -
Cognitive Castes: Artificial Intelligence, Epistemic Stratification, and the Dissolution of Democratic Discourse
by: Wright, Craig S
Published: (2025) -
Modeling the Path of Structural Strategic Deterrence: A Sand Table Simulation and Research Report on China's Military-Industrial Capability System against the United States Based on Rare Earth Supply Disconnection
by: Meng, Wei
Published: (2025)