AI Alignment at Your Discretion
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Buyl, Maarten, Khalaf, Hadi, Verdun, Claudio Mayrink, Paes, Lucas Monteiro, Machado, Caio C. Vieira, Calmon, Flavio du Pin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Soft Best-of-n Sampling for Model Alignment
von: Verdun, Claudio Mayrink, et al.
Veröffentlicht: (2025)
von: Verdun, Claudio Mayrink, et al.
Veröffentlicht: (2025)
Inference-Time Reward Hacking in Large Language Models
von: Khalaf, Hadi, et al.
Veröffentlicht: (2025)
von: Khalaf, Hadi, et al.
Veröffentlicht: (2025)
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
von: Bhalla, Usha, et al.
Veröffentlicht: (2025)
von: Bhalla, Usha, et al.
Veröffentlicht: (2025)
Multi-Group Proportional Representation in Retrieval
von: Oesterling, Alex, et al.
Veröffentlicht: (2024)
von: Oesterling, Alex, et al.
Veröffentlicht: (2024)
Algorithmic Arbitrariness in Content Moderation
von: Gomez, Juan Felipe, et al.
Veröffentlicht: (2024)
von: Gomez, Juan Felipe, et al.
Veröffentlicht: (2024)
HeavyWater and SimplexWater: Distortion-Free LLM Watermarks for Low-Entropy Next-Token Predictions
von: Tsur, Dor, et al.
Veröffentlicht: (2025)
von: Tsur, Dor, et al.
Veröffentlicht: (2025)
Multi-Group Proportional Representation for Text-to-Image Models
von: Jung, Sangwon, et al.
Veröffentlicht: (2025)
von: Jung, Sangwon, et al.
Veröffentlicht: (2025)
Optimized Couplings for Watermarking Large Language Models
von: Tsur, Dor, et al.
Veröffentlicht: (2025)
von: Tsur, Dor, et al.
Veröffentlicht: (2025)
Rigor in AI: Doing Rigorous AI Work Requires a Broader, Responsible AI-Informed Conception of Rigor
von: Olteanu, Alexandra, et al.
Veröffentlicht: (2025)
von: Olteanu, Alexandra, et al.
Veröffentlicht: (2025)
ProofCompass: Enhancing Specialized Provers with LLM Guidance
von: Wischermann, Nicolas, et al.
Veröffentlicht: (2025)
von: Wischermann, Nicolas, et al.
Veröffentlicht: (2025)
Selective Explanations
von: Paes, Lucas Monteiro, et al.
Veröffentlicht: (2024)
von: Paes, Lucas Monteiro, et al.
Veröffentlicht: (2024)
The Pedagogy of AI Mistakes: Fostering Higher-Order Thinking
von: Hosseini, Hadi
Veröffentlicht: (2026)
von: Hosseini, Hadi
Veröffentlicht: (2026)
Designing escalation criteria for international AI incident response: criteria, triggers, and thresholds
von: Gomez, Francesca, et al.
Veröffentlicht: (2026)
von: Gomez, Francesca, et al.
Veröffentlicht: (2026)
Why AI Alignment Failure Is Structural: Learned Human Interaction Structures and AGI as an Endogenous Evolutionary Shock
von: Sornette, Didier, et al.
Veröffentlicht: (2026)
von: Sornette, Didier, et al.
Veröffentlicht: (2026)
AI Education in a Mirror: Challenges Faced by Academic and Industry Experts
von: Akgun, Mahir, et al.
Veröffentlicht: (2025)
von: Akgun, Mahir, et al.
Veröffentlicht: (2025)
Forecasting Open-Weight AI Model Growth on HuggingFace
von: Bhandari, Kushal Raj, et al.
Veröffentlicht: (2025)
von: Bhandari, Kushal Raj, et al.
Veröffentlicht: (2025)
The AI Alignment Paradox
von: West, Robert, et al.
Veröffentlicht: (2024)
von: West, Robert, et al.
Veröffentlicht: (2024)
Attack-Aware Noise Calibration for Differential Privacy
von: Kulynych, Bogdan, et al.
Veröffentlicht: (2024)
von: Kulynych, Bogdan, et al.
Veröffentlicht: (2024)
Rethinking AI Cultural Alignment
von: Bravansky, Michal, et al.
Veröffentlicht: (2025)
von: Bravansky, Michal, et al.
Veröffentlicht: (2025)
Justifications for Democratizing AI Alignment and Their Prospects
von: Steingrüber, André, et al.
Veröffentlicht: (2025)
von: Steingrüber, André, et al.
Veröffentlicht: (2025)
Multi-Group Fairness Evaluation via Conditional Value-at-Risk Testing
von: Paes, Lucas Monteiro, et al.
Veröffentlicht: (2023)
von: Paes, Lucas Monteiro, et al.
Veröffentlicht: (2023)
Is Your AI Truly Yours? Leveraging Blockchain for Copyrights, Provenance, and Lineage
von: Wang, Qin, et al.
Veröffentlicht: (2024)
von: Wang, Qin, et al.
Veröffentlicht: (2024)
Understanding the Process of Human-AI Value Alignment
von: McKinlay, Jack, et al.
Veröffentlicht: (2025)
von: McKinlay, Jack, et al.
Veröffentlicht: (2025)
Toward a Public and Secure Generative AI: A Comparative Analysis of Open and Closed LLMs
von: Machado, Jorge
Veröffentlicht: (2025)
von: Machado, Jorge
Veröffentlicht: (2025)
Privacy Ethics Alignment in AI: A Stakeholder-Centric Framework for Ethical AI
von: Barthwal, Ankur, et al.
Veröffentlicht: (2025)
von: Barthwal, Ankur, et al.
Veröffentlicht: (2025)
Whose Truth? Pluralistic Geo-Alignment for (Agentic) AI
von: Janowicz, Krzysztof, et al.
Veröffentlicht: (2025)
von: Janowicz, Krzysztof, et al.
Veröffentlicht: (2025)
The Value of Gen-AI Conversations: A bottom-up Framework for AI Value Alignment
von: Motnikar, Lenart, et al.
Veröffentlicht: (2025)
von: Motnikar, Lenart, et al.
Veröffentlicht: (2025)
AI and Human Oversight: A Risk-Based Framework for Alignment
von: Kandikatla, Laxmiraju, et al.
Veröffentlicht: (2025)
von: Kandikatla, Laxmiraju, et al.
Veröffentlicht: (2025)
Robust AI Evaluation through Maximal Lotteries
von: Khalaf, Hadi, et al.
Veröffentlicht: (2026)
von: Khalaf, Hadi, et al.
Veröffentlicht: (2026)
Using AI Alignment Theory to understand the potential pitfalls of regulatory frameworks
von: Tlaie, Alejandro
Veröffentlicht: (2024)
von: Tlaie, Alejandro
Veröffentlicht: (2024)
Alignment, Agency and Autonomy in Frontier AI: A Systems Engineering Perspective
von: Tallam, Krti
Veröffentlicht: (2025)
von: Tallam, Krti
Veröffentlicht: (2025)
A Robust Governance for the AI Act: AI Office, AI Board, Scientific Panel, and National Authorities
von: Novelli, Claudio, et al.
Veröffentlicht: (2024)
von: Novelli, Claudio, et al.
Veröffentlicht: (2024)
Beware! The AI Act Can Also Apply to Your AI Research Practices
von: Wernick, Alina, et al.
Veröffentlicht: (2025)
von: Wernick, Alina, et al.
Veröffentlicht: (2025)
Particip-AI: A Democratic Surveying Framework for Anticipating Future AI Use Cases, Harms and Benefits
von: Mun, Jimin, et al.
Veröffentlicht: (2024)
von: Mun, Jimin, et al.
Veröffentlicht: (2024)
Alignment Is Not Enough: A Relational Framework for Moral Standing in Human-AI Interaction
von: Pasandi, Faezeh B., et al.
Veröffentlicht: (2026)
von: Pasandi, Faezeh B., et al.
Veröffentlicht: (2026)
Wide Reflective Equilibrium in LLM Alignment: Bridging Moral Epistemology and AI Safety
von: Brophy, Matthew
Veröffentlicht: (2025)
von: Brophy, Matthew
Veröffentlicht: (2025)
Characterizing AI Agents for Alignment and Governance
von: Kasirzadeh, Atoosa, et al.
Veröffentlicht: (2025)
von: Kasirzadeh, Atoosa, et al.
Veröffentlicht: (2025)
Unifying Re-Identification, Attribute Inference, and Data Reconstruction Risks in Differential Privacy
von: Kulynych, Bogdan, et al.
Veröffentlicht: (2025)
von: Kulynych, Bogdan, et al.
Veröffentlicht: (2025)
Rules, Cases, and Reasoning: Positivist Legal Theory as a Framework for Pluralistic AI Alignment
von: Caputo, Nicholas A.
Veröffentlicht: (2024)
von: Caputo, Nicholas A.
Veröffentlicht: (2024)
Desk-AId: Humanitarian Aid Desk Assessment with Geospatial AI for Predicting Landmine Areas
von: Cirillo, Flavio, et al.
Veröffentlicht: (2024)
von: Cirillo, Flavio, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Soft Best-of-n Sampling for Model Alignment
von: Verdun, Claudio Mayrink, et al.
Veröffentlicht: (2025) -
Inference-Time Reward Hacking in Large Language Models
von: Khalaf, Hadi, et al.
Veröffentlicht: (2025) -
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
von: Bhalla, Usha, et al.
Veröffentlicht: (2025) -
Multi-Group Proportional Representation in Retrieval
von: Oesterling, Alex, et al.
Veröffentlicht: (2024) -
Algorithmic Arbitrariness in Content Moderation
von: Gomez, Juan Felipe, et al.
Veröffentlicht: (2024)