The Specification Trap: Why Static Value Alignment Alone Is Insufficient for Robust Alignment
Fuente:
arXiv
Guardado en:
| Autor principal: | Spizzirri, Austin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Machine Learning as Performative Materialist Practice: Thirteen Theses on the Epistemology, Methodology, and Politics of Applied ML
por: De Unánue, Adolfo, et al.
Publicado: (2026)
por: De Unánue, Adolfo, et al.
Publicado: (2026)
Revenue-Sharing as Infrastructure: A Distributed Business Model for Generative AI Platforms
por: Mondjo, Ghislain Dorian Tchuente
Publicado: (2026)
por: Mondjo, Ghislain Dorian Tchuente
Publicado: (2026)
Alpay Algebra III: Observer-Coupled Collapse and the Temporal Drift of Identity
por: Alpay, Faruk
Publicado: (2025)
por: Alpay, Faruk
Publicado: (2025)
Cognitive Castes: Artificial Intelligence, Epistemic Stratification, and the Dissolution of Democratic Discourse
por: Wright, Craig S
Publicado: (2025)
por: Wright, Craig S
Publicado: (2025)
Modeling the Path of Structural Strategic Deterrence: A Sand Table Simulation and Research Report on China's Military-Industrial Capability System against the United States Based on Rare Earth Supply Disconnection
por: Meng, Wei
Publicado: (2025)
por: Meng, Wei
Publicado: (2025)
Expert Insight-Based Modeling of Non-Kinetic Strategic Deterrence of Rare Earth Supply Disruption:A Simulation-Driven Systematic Framework
por: Meng, Wei
Publicado: (2025)
por: Meng, Wei
Publicado: (2025)
Memory-Augmented State Machine Prompting: A Novel LLM Agent Framework for Real-Time Strategy Games
por: Qi, Runnan, et al.
Publicado: (2025)
por: Qi, Runnan, et al.
Publicado: (2025)
Cooperation After the Algorithm: Designing Human-AI Coexistence Beyond the Illusion of Collaboration
por: Codreanu, Tatia
Publicado: (2026)
por: Codreanu, Tatia
Publicado: (2026)
Intersymbolic AI: Interlinking Symbolic AI and Subsymbolic AI
por: Platzer, André
Publicado: (2024)
por: Platzer, André
Publicado: (2024)
Formal Proofs as Structured Explanations: Proposing Several Tasks on Explainable Natural Language Inference
por: Abzianidze, Lasha
Publicado: (2023)
por: Abzianidze, Lasha
Publicado: (2023)
Fixed-Point Traps and Identity Emergence in Educational Feedback Systems
por: Alpay, Faruk
Publicado: (2025)
por: Alpay, Faruk
Publicado: (2025)
Civilizational Metamaterials: Engineering Coordination Under Capability Gradients and Structural Turbulence
por: Orban, David
Publicado: (2026)
por: Orban, David
Publicado: (2026)
Strategic Counterfactual Modeling of Deep-Target Airstrike Systems via Intervention-Aware Spatio-Causal Graph Networks
por: Meng, Wei
Publicado: (2025)
por: Meng, Wei
Publicado: (2025)
Online Decision Making with Generative Action Sets
por: Xu, Jianyu, et al.
Publicado: (2025)
por: Xu, Jianyu, et al.
Publicado: (2025)
Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback
por: Conitzer, Vincent, et al.
Publicado: (2024)
por: Conitzer, Vincent, et al.
Publicado: (2024)
Coopetition-Gym v1: A Formally Grounded Platform for Mixed-Motive Multi-Agent Reinforcement Learning under Strategic Coopetition
por: Pant, Vik, et al.
Publicado: (2026)
por: Pant, Vik, et al.
Publicado: (2026)
Hilbert-Geo: Solving Solid Geometric Problems by Neural-Symbolic Reasoning
por: Xu, Ruoran, et al.
Publicado: (2026)
por: Xu, Ruoran, et al.
Publicado: (2026)
Modeling and Visualization Reasoning for Stakeholders in Education and Industry Integration Systems: Research on Structured Synthetic Dialogue Data Generation Based on NIST Standards
por: Meng, Wei
Publicado: (2025)
por: Meng, Wei
Publicado: (2025)
Standardized Threat Taxonomy for AI Security, Governance, and Regulatory Compliance
por: Huwyler, Hernan
Publicado: (2025)
por: Huwyler, Hernan
Publicado: (2025)
Informed AI Regulation: Comparing the Ethical Frameworks of Leading LLM Chatbots Using an Ethics-Based Audit to Assess Moral Reasoning and Normative Values
por: Chun, Jon, et al.
Publicado: (2024)
por: Chun, Jon, et al.
Publicado: (2024)
Bridging Voting and Deliberation with Algorithms: Field Insights from vTaiwan and Kultur Komitee
por: Yang, Joshua C., et al.
Publicado: (2025)
por: Yang, Joshua C., et al.
Publicado: (2025)
Yanasse: Finding New Proofs from Deep Vision's Analogies, Part 1
por: Linhares, Alexandre
Publicado: (2026)
por: Linhares, Alexandre
Publicado: (2026)
Fixed-Point Theorems and the Ethics of Radical Transparency: A Logic-First Treatment
por: Alpay, Faruk, et al.
Publicado: (2025)
por: Alpay, Faruk, et al.
Publicado: (2025)
Value-Aware Multiagent Systems
por: Osman, Nardine
Publicado: (2025)
por: Osman, Nardine
Publicado: (2025)
Approximating Discrimination Within Models When Faced With Several Non-Binary Sensitive Attributes
por: Bian, Yijun, et al.
Publicado: (2024)
por: Bian, Yijun, et al.
Publicado: (2024)
Does Machine Bring in Extra Bias in Learning? Approximating Fairness in Models Promptly
por: Bian, Yijun, et al.
Publicado: (2024)
por: Bian, Yijun, et al.
Publicado: (2024)
Interpretation, Learning, and Empathy as One Constraint: A Residual-Adequacy Architecture with Accountable Abstention
por: Amornbunchornvej, Chainarong
Publicado: (2026)
por: Amornbunchornvej, Chainarong
Publicado: (2026)
Logical Modalities within the European AI Act: An Analysis
por: Lawniczak, Lara, et al.
Publicado: (2025)
por: Lawniczak, Lara, et al.
Publicado: (2025)
Prices, Bids, Values: One ML-Powered Combinatorial Auction to Rule Them All
por: Soumalias, Ermis, et al.
Publicado: (2024)
por: Soumalias, Ermis, et al.
Publicado: (2024)
A Monad-Based Clause Architecture for Artificial Age Score (AAS) in Large Language Models
por: Kayadibi, Seyma Yaman
Publicado: (2025)
por: Kayadibi, Seyma Yaman
Publicado: (2025)
How Worrying Are Privacy Attacks Against Machine Learning?
por: Domingo-Ferrer, Josep
Publicado: (2025)
por: Domingo-Ferrer, Josep
Publicado: (2025)
Beyond Reward Suppression: Reshaping Steganographic Communication Protocols in MARL via Dynamic Representational Circuit Breaking
por: Ming, Liu Hung
Publicado: (2026)
por: Ming, Liu Hung
Publicado: (2026)
$ϕ^{\infty}$: Clause Purification, Embedding Realignment, and the Total Suppression of the Em Dash in Autoregressive Language Models
por: Kilictas, Bugra, et al.
Publicado: (2025)
por: Kilictas, Bugra, et al.
Publicado: (2025)
Reinsuring AI: Energy, Agriculture, Finance & Medicine as Precedents for Scalable Governance of Frontier Artificial Intelligence
por: Stetler, Nicholas
Publicado: (2025)
por: Stetler, Nicholas
Publicado: (2025)
Transforming Business with Generative AI: Research, Innovation, Market Deployment and Future Shifts in Business Models
por: Singh, Narotam, et al.
Publicado: (2024)
por: Singh, Narotam, et al.
Publicado: (2024)
Interpretability Guarantees with Merlin-Arthur Classifiers
por: Wäldchen, Stephan, et al.
Publicado: (2022)
por: Wäldchen, Stephan, et al.
Publicado: (2022)
Emergent Coordination in Multi-Agent Systems via Pressure Fields and Temporal Decay
por: Rodriguez, Roland
Publicado: (2026)
por: Rodriguez, Roland
Publicado: (2026)
Black Box Deployed -- Functional Criteria for Artificial Moral Agents in the LLM Era
por: Brophy, Matthew E.
Publicado: (2025)
por: Brophy, Matthew E.
Publicado: (2025)
The Axiom of Consent: Friction Dynamics in Multi-Agent Coordination
por: Farzulla, Murad
Publicado: (2026)
por: Farzulla, Murad
Publicado: (2026)
Solving Zebra Puzzles Using Constraint-Guided Multi-Agent Systems
por: Berman, Shmuel, et al.
Publicado: (2024)
por: Berman, Shmuel, et al.
Publicado: (2024)
Ejemplares similares
-
Machine Learning as Performative Materialist Practice: Thirteen Theses on the Epistemology, Methodology, and Politics of Applied ML
por: De Unánue, Adolfo, et al.
Publicado: (2026) -
Revenue-Sharing as Infrastructure: A Distributed Business Model for Generative AI Platforms
por: Mondjo, Ghislain Dorian Tchuente
Publicado: (2026) -
Alpay Algebra III: Observer-Coupled Collapse and the Temporal Drift of Identity
por: Alpay, Faruk
Publicado: (2025) -
Cognitive Castes: Artificial Intelligence, Epistemic Stratification, and the Dissolution of Democratic Discourse
por: Wright, Craig S
Publicado: (2025) -
Modeling the Path of Structural Strategic Deterrence: A Sand Table Simulation and Research Report on China's Military-Industrial Capability System against the United States Based on Rare Earth Supply Disconnection
por: Meng, Wei
Publicado: (2025)