The AI Alignment Paradox
Fuente:
arXiv
Salvato in:
| Autori principali: | West, Robert, Aydin, Roland |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
From Model Training to Model Raising
di: Aydin, Roland, et al.
Pubblicazione: (2025)
di: Aydin, Roland, et al.
Pubblicazione: (2025)
Assessing Social Alignment: Do Personality-Prompted Large Language Models Behave Like Humans?
di: Zakazov, Ivan, et al.
Pubblicazione: (2024)
di: Zakazov, Ivan, et al.
Pubblicazione: (2024)
The Psychology of Learning from Machines: Anthropomorphic AI and the Paradox of Automation in Education
di: Qadir, Junaid, et al.
Pubblicazione: (2026)
di: Qadir, Junaid, et al.
Pubblicazione: (2026)
Rethinking AI Cultural Alignment
di: Bravansky, Michal, et al.
Pubblicazione: (2025)
di: Bravansky, Michal, et al.
Pubblicazione: (2025)
Justifications for Democratizing AI Alignment and Their Prospects
di: Steingrüber, André, et al.
Pubblicazione: (2025)
di: Steingrüber, André, et al.
Pubblicazione: (2025)
"Think First, Verify Always": Training Humans to Face AI Risks
di: Aydin, Yuksel
Pubblicazione: (2025)
di: Aydin, Yuksel
Pubblicazione: (2025)
Understanding the Process of Human-AI Value Alignment
di: McKinlay, Jack, et al.
Pubblicazione: (2025)
di: McKinlay, Jack, et al.
Pubblicazione: (2025)
The Generative AI Paradox: GenAI and the Erosion of Trust, the Corrosion of Information Verification, and the Demise of Truth
di: Ferrara, Emilio
Pubblicazione: (2026)
di: Ferrara, Emilio
Pubblicazione: (2026)
Privacy Ethics Alignment in AI: A Stakeholder-Centric Framework for Ethical AI
di: Barthwal, Ankur, et al.
Pubblicazione: (2025)
di: Barthwal, Ankur, et al.
Pubblicazione: (2025)
Whose Truth? Pluralistic Geo-Alignment for (Agentic) AI
di: Janowicz, Krzysztof, et al.
Pubblicazione: (2025)
di: Janowicz, Krzysztof, et al.
Pubblicazione: (2025)
The Value of Gen-AI Conversations: A bottom-up Framework for AI Value Alignment
di: Motnikar, Lenart, et al.
Pubblicazione: (2025)
di: Motnikar, Lenart, et al.
Pubblicazione: (2025)
AI and Human Oversight: A Risk-Based Framework for Alignment
di: Kandikatla, Laxmiraju, et al.
Pubblicazione: (2025)
di: Kandikatla, Laxmiraju, et al.
Pubblicazione: (2025)
Scaling Truth: The Confidence Paradox in AI Fact-Checking
di: Qazi, Ihsan A., et al.
Pubblicazione: (2025)
di: Qazi, Ihsan A., et al.
Pubblicazione: (2025)
Using AI Alignment Theory to understand the potential pitfalls of regulatory frameworks
di: Tlaie, Alejandro
Pubblicazione: (2024)
di: Tlaie, Alejandro
Pubblicazione: (2024)
Alignment, Agency and Autonomy in Frontier AI: A Systems Engineering Perspective
di: Tallam, Krti
Pubblicazione: (2025)
di: Tallam, Krti
Pubblicazione: (2025)
BioBlue: Systematic runaway-optimiser-like LLM failure modes on biologically and economically aligned AI safety benchmarks for LLMs with simplified observation format
di: Pihlakas, Roland, et al.
Pubblicazione: (2025)
di: Pihlakas, Roland, et al.
Pubblicazione: (2025)
Alignment Is Not Enough: A Relational Framework for Moral Standing in Human-AI Interaction
di: Pasandi, Faezeh B., et al.
Pubblicazione: (2026)
di: Pasandi, Faezeh B., et al.
Pubblicazione: (2026)
Wide Reflective Equilibrium in LLM Alignment: Bridging Moral Epistemology and AI Safety
di: Brophy, Matthew
Pubblicazione: (2025)
di: Brophy, Matthew
Pubblicazione: (2025)
AI Alignment at Your Discretion
di: Buyl, Maarten, et al.
Pubblicazione: (2025)
di: Buyl, Maarten, et al.
Pubblicazione: (2025)
Rules, Cases, and Reasoning: Positivist Legal Theory as a Framework for Pluralistic AI Alignment
di: Caputo, Nicholas A.
Pubblicazione: (2024)
di: Caputo, Nicholas A.
Pubblicazione: (2024)
Characterizing AI Agents for Alignment and Governance
di: Kasirzadeh, Atoosa, et al.
Pubblicazione: (2025)
di: Kasirzadeh, Atoosa, et al.
Pubblicazione: (2025)
The Coming Crisis of Multi-Agent Misalignment: AI Alignment Must Be a Dynamic and Social Process
di: Carichon, Florian, et al.
Pubblicazione: (2025)
di: Carichon, Florian, et al.
Pubblicazione: (2025)
The Memory Paradox: Why Our Brains Need Knowledge in an Age of AI
di: Oakley, Barbara, et al.
Pubblicazione: (2025)
di: Oakley, Barbara, et al.
Pubblicazione: (2025)
The Adoption Paradox for Veterinary Professionals in China: High Use of Artificial Intelligence Despite Low Familiarity
di: Li, Shumin, et al.
Pubblicazione: (2025)
di: Li, Shumin, et al.
Pubblicazione: (2025)
AI From the Margins (AIM): Rethinking Participatory AI Design Through the Lived Experience of Minoritized Communities
di: Portegies, Tijs, et al.
Pubblicazione: (2026)
di: Portegies, Tijs, et al.
Pubblicazione: (2026)
Why AI Alignment Failure Is Structural: Learned Human Interaction Structures and AGI as an Endogenous Evolutionary Shock
di: Sornette, Didier, et al.
Pubblicazione: (2026)
di: Sornette, Didier, et al.
Pubblicazione: (2026)
A Tutorial on Teaching Data Analytics with Generative AI
di: Bray, Robert L.
Pubblicazione: (2024)
di: Bray, Robert L.
Pubblicazione: (2024)
Alignment Debt: The Hidden Work of Making AI Usable
di: Oyemike, Cumi, et al.
Pubblicazione: (2025)
di: Oyemike, Cumi, et al.
Pubblicazione: (2025)
Balancing Safety and Helpfulness in Healthcare AI Assistants through Iterative Preference Alignment
di: Nghiem, Huy, et al.
Pubblicazione: (2025)
di: Nghiem, Huy, et al.
Pubblicazione: (2025)
Distributed agency in second language learning and teaching through generative AI
di: Godwin-Jones, Robert
Pubblicazione: (2024)
di: Godwin-Jones, Robert
Pubblicazione: (2024)
Towards Integrated Alignment
di: Reis, Ben Y., et al.
Pubblicazione: (2025)
di: Reis, Ben Y., et al.
Pubblicazione: (2025)
Bidirectional Human-AI Alignment in Education for Trustworthy Learning Environments
di: Shen, Hua
Pubblicazione: (2025)
di: Shen, Hua
Pubblicazione: (2025)
Breaking Up with Normatively Monolithic Agency with GRACE: A Reason-Based Neuro-Symbolic Architecture for Safe and Ethical AI Alignment
di: Jahn, Felix, et al.
Pubblicazione: (2026)
di: Jahn, Felix, et al.
Pubblicazione: (2026)
Risk Alignment in Agentic AI Systems
di: Clatterbuck, Hayley, et al.
Pubblicazione: (2024)
di: Clatterbuck, Hayley, et al.
Pubblicazione: (2024)
The Law-Following AI Framework: Legal Foundations and Technical Constraints. Legal Analogues for AI Actorship and technical feasibility of Law Alignment
di: Delgado, Katalina Hernandez
Pubblicazione: (2025)
di: Delgado, Katalina Hernandez
Pubblicazione: (2025)
Antisocial Analagous Behavior, Alignment and Human Impact of Google AI Systems: Evaluating through the lens of modified Antisocial Behavior Criteria by Human Interaction, Independent LLM Analysis, and AI Self-Reflection
di: Ogilvie, Alan D.
Pubblicazione: (2024)
di: Ogilvie, Alan D.
Pubblicazione: (2024)
The Emotional Alignment Design Policy
di: Schwitzgebel, Eric, et al.
Pubblicazione: (2025)
di: Schwitzgebel, Eric, et al.
Pubblicazione: (2025)
Disentangling AI Alignment: A Structured Taxonomy Beyond Safety and Ethics
di: Baum, Kevin
Pubblicazione: (2025)
di: Baum, Kevin
Pubblicazione: (2025)
Lightweight Prompt Engineering for Cognitive Alignment in Educational AI: A OneClickQuiz Case Study
di: Yaacoub, Antoun, et al.
Pubblicazione: (2025)
di: Yaacoub, Antoun, et al.
Pubblicazione: (2025)
Perceptions of AI Across Sectors: A Comparative Review of Public Attitudes
di: Bialy, Filip, et al.
Pubblicazione: (2025)
di: Bialy, Filip, et al.
Pubblicazione: (2025)
Documenti analoghi
-
From Model Training to Model Raising
di: Aydin, Roland, et al.
Pubblicazione: (2025) -
Assessing Social Alignment: Do Personality-Prompted Large Language Models Behave Like Humans?
di: Zakazov, Ivan, et al.
Pubblicazione: (2024) -
The Psychology of Learning from Machines: Anthropomorphic AI and the Paradox of Automation in Education
di: Qadir, Junaid, et al.
Pubblicazione: (2026) -
Rethinking AI Cultural Alignment
di: Bravansky, Michal, et al.
Pubblicazione: (2025) -
Justifications for Democratizing AI Alignment and Their Prospects
di: Steingrüber, André, et al.
Pubblicazione: (2025)