Using AI Alignment Theory to understand the potential pitfalls of regulatory frameworks
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Tlaie, Alejandro |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Large language models in medicine: the potentials and pitfalls
von: Omiye, Jesutofunmi A., et al.
Veröffentlicht: (2023)
von: Omiye, Jesutofunmi A., et al.
Veröffentlicht: (2023)
Exploring and steering the moral compass of Large Language Models
von: Tlaie, Alejandro
Veröffentlicht: (2024)
von: Tlaie, Alejandro
Veröffentlicht: (2024)
Promises and pitfalls of artificial intelligence for legal applications
von: Kapoor, Sayash, et al.
Veröffentlicht: (2024)
von: Kapoor, Sayash, et al.
Veröffentlicht: (2024)
Rules, Cases, and Reasoning: Positivist Legal Theory as a Framework for Pluralistic AI Alignment
von: Caputo, Nicholas A.
Veröffentlicht: (2024)
von: Caputo, Nicholas A.
Veröffentlicht: (2024)
The AI Alignment Paradox
von: West, Robert, et al.
Veröffentlicht: (2024)
von: West, Robert, et al.
Veröffentlicht: (2024)
Rethinking AI Cultural Alignment
von: Bravansky, Michal, et al.
Veröffentlicht: (2025)
von: Bravansky, Michal, et al.
Veröffentlicht: (2025)
Justifications for Democratizing AI Alignment and Their Prospects
von: Steingrüber, André, et al.
Veröffentlicht: (2025)
von: Steingrüber, André, et al.
Veröffentlicht: (2025)
AI and Social Theory
von: Mokander, Jakob, et al.
Veröffentlicht: (2024)
von: Mokander, Jakob, et al.
Veröffentlicht: (2024)
Understanding the Process of Human-AI Value Alignment
von: McKinlay, Jack, et al.
Veröffentlicht: (2025)
von: McKinlay, Jack, et al.
Veröffentlicht: (2025)
Privacy Ethics Alignment in AI: A Stakeholder-Centric Framework for Ethical AI
von: Barthwal, Ankur, et al.
Veröffentlicht: (2025)
von: Barthwal, Ankur, et al.
Veröffentlicht: (2025)
Securing External Deeper-than-black-box GPAI Evaluations
von: Tlaie, Alejandro, et al.
Veröffentlicht: (2025)
von: Tlaie, Alejandro, et al.
Veröffentlicht: (2025)
Whose Truth? Pluralistic Geo-Alignment for (Agentic) AI
von: Janowicz, Krzysztof, et al.
Veröffentlicht: (2025)
von: Janowicz, Krzysztof, et al.
Veröffentlicht: (2025)
The Value of Gen-AI Conversations: A bottom-up Framework for AI Value Alignment
von: Motnikar, Lenart, et al.
Veröffentlicht: (2025)
von: Motnikar, Lenart, et al.
Veröffentlicht: (2025)
Building the ethical AI framework of the future: from philosophy to practice
von: Catapang, Jasper Kyle
Veröffentlicht: (2026)
von: Catapang, Jasper Kyle
Veröffentlicht: (2026)
AI and Human Oversight: A Risk-Based Framework for Alignment
von: Kandikatla, Laxmiraju, et al.
Veröffentlicht: (2025)
von: Kandikatla, Laxmiraju, et al.
Veröffentlicht: (2025)
Addressing the regulatory gap: moving towards an EU AI audit ecosystem beyond the AI Act by including civil society
von: Hartmann, David, et al.
Veröffentlicht: (2024)
von: Hartmann, David, et al.
Veröffentlicht: (2024)
Alignment, Agency and Autonomy in Frontier AI: A Systems Engineering Perspective
von: Tallam, Krti
Veröffentlicht: (2025)
von: Tallam, Krti
Veröffentlicht: (2025)
Security, privacy, and agentic AI in a regulatory view: From definitions and distinctions to provisions and reflections
von: Zhang, Shiliang, et al.
Veröffentlicht: (2026)
von: Zhang, Shiliang, et al.
Veröffentlicht: (2026)
Alignment Is Not Enough: A Relational Framework for Moral Standing in Human-AI Interaction
von: Pasandi, Faezeh B., et al.
Veröffentlicht: (2026)
von: Pasandi, Faezeh B., et al.
Veröffentlicht: (2026)
Wide Reflective Equilibrium in LLM Alignment: Bridging Moral Epistemology and AI Safety
von: Brophy, Matthew
Veröffentlicht: (2025)
von: Brophy, Matthew
Veröffentlicht: (2025)
AI Alignment at Your Discretion
von: Buyl, Maarten, et al.
Veröffentlicht: (2025)
von: Buyl, Maarten, et al.
Veröffentlicht: (2025)
AI for bureaucratic productivity: Measuring the potential of AI to help automate 143 million UK government transactions
von: Straub, Vincent J., et al.
Veröffentlicht: (2024)
von: Straub, Vincent J., et al.
Veröffentlicht: (2024)
AI threats to national security can be countered through an incident regime
von: Ortega, Alejandro
Veröffentlicht: (2025)
von: Ortega, Alejandro
Veröffentlicht: (2025)
AI-Educational Development Loop (AI-EDL): A Conceptual Framework to Bridge AI Capabilities with Classical Educational Theories
von: Yu, Ning, et al.
Veröffentlicht: (2025)
von: Yu, Ning, et al.
Veröffentlicht: (2025)
Teacher agency in the age of generative AI: towards a framework of hybrid intelligence for learning design
von: Frøsig, Thomas B, et al.
Veröffentlicht: (2024)
von: Frøsig, Thomas B, et al.
Veröffentlicht: (2024)
Large language models eroding science understanding: an experimental study
von: Collins, Harry, et al.
Veröffentlicht: (2026)
von: Collins, Harry, et al.
Veröffentlicht: (2026)
Characterizing AI Agents for Alignment and Governance
von: Kasirzadeh, Atoosa, et al.
Veröffentlicht: (2025)
von: Kasirzadeh, Atoosa, et al.
Veröffentlicht: (2025)
The Coming Crisis of Multi-Agent Misalignment: AI Alignment Must Be a Dynamic and Social Process
von: Carichon, Florian, et al.
Veröffentlicht: (2025)
von: Carichon, Florian, et al.
Veröffentlicht: (2025)
Standing on FURM ground -- A framework for evaluating Fair, Useful, and Reliable AI Models in healthcare systems
von: Callahan, Alison, et al.
Veröffentlicht: (2024)
von: Callahan, Alison, et al.
Veröffentlicht: (2024)
Deconstructing Student Perceptions of Generative AI (GenAI) through an Expectancy Value Theory (EVT)-based Instrument
von: Chan, Cecilia Ka Yuk, et al.
Veröffentlicht: (2023)
von: Chan, Cecilia Ka Yuk, et al.
Veröffentlicht: (2023)
Why AI Alignment Failure Is Structural: Learned Human Interaction Structures and AGI as an Endogenous Evolutionary Shock
von: Sornette, Didier, et al.
Veröffentlicht: (2026)
von: Sornette, Didier, et al.
Veröffentlicht: (2026)
ELEPHANT: Measuring and understanding social sycophancy in LLMs
von: Cheng, Myra, et al.
Veröffentlicht: (2025)
von: Cheng, Myra, et al.
Veröffentlicht: (2025)
Representative Social Choice: From Learning Theory to AI Alignment
von: Qiu, Tianyi
Veröffentlicht: (2024)
von: Qiu, Tianyi
Veröffentlicht: (2024)
Exploring Public Opinion on Responsible AI Through The Lens of Cultural Consensus Theory
von: Gurkan, Necdet, et al.
Veröffentlicht: (2024)
von: Gurkan, Necdet, et al.
Veröffentlicht: (2024)
Alignment Debt: The Hidden Work of Making AI Usable
von: Oyemike, Cumi, et al.
Veröffentlicht: (2025)
von: Oyemike, Cumi, et al.
Veröffentlicht: (2025)
Balancing Safety and Helpfulness in Healthcare AI Assistants through Iterative Preference Alignment
von: Nghiem, Huy, et al.
Veröffentlicht: (2025)
von: Nghiem, Huy, et al.
Veröffentlicht: (2025)
The potential functions of an international institution for AI safety. Insights from adjacent policy areas and recent trends
von: De Castris, A. Leone, et al.
Veröffentlicht: (2024)
von: De Castris, A. Leone, et al.
Veröffentlicht: (2024)
Enhanced Interpretable Knowledge Tracing for Students Performance Prediction with Human understandable Feature Space
von: Minn, Sein, et al.
Veröffentlicht: (2025)
von: Minn, Sein, et al.
Veröffentlicht: (2025)
AI Thinking: A framework for rethinking artificial intelligence in practice
von: Newman-Griffis, Denis
Veröffentlicht: (2024)
von: Newman-Griffis, Denis
Veröffentlicht: (2024)
Towards Integrated Alignment
von: Reis, Ben Y., et al.
Veröffentlicht: (2025)
von: Reis, Ben Y., et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Large language models in medicine: the potentials and pitfalls
von: Omiye, Jesutofunmi A., et al.
Veröffentlicht: (2023) -
Exploring and steering the moral compass of Large Language Models
von: Tlaie, Alejandro
Veröffentlicht: (2024) -
Promises and pitfalls of artificial intelligence for legal applications
von: Kapoor, Sayash, et al.
Veröffentlicht: (2024) -
Rules, Cases, and Reasoning: Positivist Legal Theory as a Framework for Pluralistic AI Alignment
von: Caputo, Nicholas A.
Veröffentlicht: (2024) -
The AI Alignment Paradox
von: West, Robert, et al.
Veröffentlicht: (2024)