Rethinking AI Cultural Alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bravansky, Michal, Trhlik, Filip, Barez, Fazl |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Embodied AI: Emerging Risks and Opportunities for Policy Action
von: Perlo, Jared, et al.
Veröffentlicht: (2025)
von: Perlo, Jared, et al.
Veröffentlicht: (2025)
VAL-Bench: Belief Consistency as a measure for Value Alignment in Language Models
von: Gupta, Aman, et al.
Veröffentlicht: (2025)
von: Gupta, Aman, et al.
Veröffentlicht: (2025)
Quantifying Generative Media Bias with a Corpus of Real-world and Generated News Articles
von: Trhlik, Filip, et al.
Veröffentlicht: (2024)
von: Trhlik, Filip, et al.
Veröffentlicht: (2024)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2025)
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2025)
Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness
von: Fu, Tingchen, et al.
Veröffentlicht: (2025)
von: Fu, Tingchen, et al.
Veröffentlicht: (2025)
Bias Dynamics in BabyLMs: Towards a Compute-Efficient Sandbox for Democratising Pre-Training Debiasing
von: Trhlik, Filip, et al.
Veröffentlicht: (2026)
von: Trhlik, Filip, et al.
Veröffentlicht: (2026)
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
von: Lan, Michael, et al.
Veröffentlicht: (2023)
von: Lan, Michael, et al.
Veröffentlicht: (2023)
Understanding Addition in Transformers
von: Quirke, Philip, et al.
Veröffentlicht: (2023)
von: Quirke, Philip, et al.
Veröffentlicht: (2023)
Query Circuits: Explaining How Language Models Answer User Prompts
von: Wu, Tung-Yu, et al.
Veröffentlicht: (2025)
von: Wu, Tung-Yu, et al.
Veröffentlicht: (2025)
An Evaluation of Cultural Value Alignment in LLM
von: Sukiennik, Nicholas, et al.
Veröffentlicht: (2025)
von: Sukiennik, Nicholas, et al.
Veröffentlicht: (2025)
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
von: Neo, Clement, et al.
Veröffentlicht: (2024)
von: Neo, Clement, et al.
Veröffentlicht: (2024)
The AI Alignment Paradox
von: West, Robert, et al.
Veröffentlicht: (2024)
von: West, Robert, et al.
Veröffentlicht: (2024)
Token Taxes: mitigating AGI's economic risks
von: Irwin, Lucas, et al.
Veröffentlicht: (2026)
von: Irwin, Lucas, et al.
Veröffentlicht: (2026)
Open Problems in Machine Unlearning for AI Safety
von: Barez, Fazl, et al.
Veröffentlicht: (2025)
von: Barez, Fazl, et al.
Veröffentlicht: (2025)
Dataset Featurization: Uncovering Natural Language Features through Unsupervised Data Reconstruction
von: Bravansky, Michal, et al.
Veröffentlicht: (2025)
von: Bravansky, Michal, et al.
Veröffentlicht: (2025)
Perceptions of AI Across Sectors: A Comparative Review of Public Attitudes
von: Bialy, Filip, et al.
Veröffentlicht: (2025)
von: Bialy, Filip, et al.
Veröffentlicht: (2025)
Fluent but Foreign: Even Regional LLMs Lack Cultural Alignment
von: Agarwal, Dhruv, et al.
Veröffentlicht: (2025)
von: Agarwal, Dhruv, et al.
Veröffentlicht: (2025)
Justifications for Democratizing AI Alignment and Their Prospects
von: Steingrüber, André, et al.
Veröffentlicht: (2025)
von: Steingrüber, André, et al.
Veröffentlicht: (2025)
Beyond Agreement: Rethinking Ground Truth in Educational AI Annotation
von: Thomas, Danielle R., et al.
Veröffentlicht: (2025)
von: Thomas, Danielle R., et al.
Veröffentlicht: (2025)
Urban Comfort Assessment in the Era of Digital Planning: A Multidimensional, Data-driven, and AI-assisted Framework
von: Yang, Sijie, et al.
Veröffentlicht: (2025)
von: Yang, Sijie, et al.
Veröffentlicht: (2025)
Understanding the Process of Human-AI Value Alignment
von: McKinlay, Jack, et al.
Veröffentlicht: (2025)
von: McKinlay, Jack, et al.
Veröffentlicht: (2025)
Beyond Automation: Rethinking Work, Creativity, and Governance in the Age of Generative AI
von: Lin, Haocheng
Veröffentlicht: (2025)
von: Lin, Haocheng
Veröffentlicht: (2025)
AI From the Margins (AIM): Rethinking Participatory AI Design Through the Lived Experience of Minoritized Communities
von: Portegies, Tijs, et al.
Veröffentlicht: (2026)
von: Portegies, Tijs, et al.
Veröffentlicht: (2026)
Privacy Ethics Alignment in AI: A Stakeholder-Centric Framework for Ethical AI
von: Barthwal, Ankur, et al.
Veröffentlicht: (2025)
von: Barthwal, Ankur, et al.
Veröffentlicht: (2025)
Ask before you Build: Rethinking AI-for-Good in Human Trafficking Interventions
von: Nair, Pratheeksha, et al.
Veröffentlicht: (2025)
von: Nair, Pratheeksha, et al.
Veröffentlicht: (2025)
From Cloud to Edge: Rethinking Generative AI for Low-Resource Design Challenges
von: Vuruma, Sai Krishna Revanth, et al.
Veröffentlicht: (2024)
von: Vuruma, Sai Krishna Revanth, et al.
Veröffentlicht: (2024)
Whose Truth? Pluralistic Geo-Alignment for (Agentic) AI
von: Janowicz, Krzysztof, et al.
Veröffentlicht: (2025)
von: Janowicz, Krzysztof, et al.
Veröffentlicht: (2025)
The Value of Gen-AI Conversations: A bottom-up Framework for AI Value Alignment
von: Motnikar, Lenart, et al.
Veröffentlicht: (2025)
von: Motnikar, Lenart, et al.
Veröffentlicht: (2025)
Rethinking AI Literacy Education in Higher Education: Bridging Risk Perception and Responsible Adoption
von: Yu, Shasha, et al.
Veröffentlicht: (2026)
von: Yu, Shasha, et al.
Veröffentlicht: (2026)
Evidence of a Cognitive Shift in AI Education: How Students Are Rethinking Human Intelligence?
von: Rekik, Islem
Veröffentlicht: (2026)
von: Rekik, Islem
Veröffentlicht: (2026)
Large Language Models Relearn Removed Concepts
von: Lo, Michelle, et al.
Veröffentlicht: (2024)
von: Lo, Michelle, et al.
Veröffentlicht: (2024)
AI and Human Oversight: A Risk-Based Framework for Alignment
von: Kandikatla, Laxmiraju, et al.
Veröffentlicht: (2025)
von: Kandikatla, Laxmiraju, et al.
Veröffentlicht: (2025)
Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMs
von: Kabir, Mohsinul, et al.
Veröffentlicht: (2025)
von: Kabir, Mohsinul, et al.
Veröffentlicht: (2025)
Alignment, Agency and Autonomy in Frontier AI: A Systems Engineering Perspective
von: Tallam, Krti
Veröffentlicht: (2025)
von: Tallam, Krti
Veröffentlicht: (2025)
Using AI Alignment Theory to understand the potential pitfalls of regulatory frameworks
von: Tlaie, Alejandro
Veröffentlicht: (2024)
von: Tlaie, Alejandro
Veröffentlicht: (2024)
Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs
von: Rystrøm, Jonathan, et al.
Veröffentlicht: (2025)
von: Rystrøm, Jonathan, et al.
Veröffentlicht: (2025)
Training-Free Cultural Alignment of Large Language Models via Persona Disagreement
von: Kiet, Huynh Trung, et al.
Veröffentlicht: (2026)
von: Kiet, Huynh Trung, et al.
Veröffentlicht: (2026)
Not someone, but something: Rethinking trust in the age of medical AI
von: Beger, Jan
Veröffentlicht: (2025)
von: Beger, Jan
Veröffentlicht: (2025)
Wide Reflective Equilibrium in LLM Alignment: Bridging Moral Epistemology and AI Safety
von: Brophy, Matthew
Veröffentlicht: (2025)
von: Brophy, Matthew
Veröffentlicht: (2025)
AI Alignment at Your Discretion
von: Buyl, Maarten, et al.
Veröffentlicht: (2025)
von: Buyl, Maarten, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Embodied AI: Emerging Risks and Opportunities for Policy Action
von: Perlo, Jared, et al.
Veröffentlicht: (2025) -
VAL-Bench: Belief Consistency as a measure for Value Alignment in Language Models
von: Gupta, Aman, et al.
Veröffentlicht: (2025) -
Quantifying Generative Media Bias with a Corpus of Real-world and Generated News Articles
von: Trhlik, Filip, et al.
Veröffentlicht: (2024) -
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2025) -
Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness
von: Fu, Tingchen, et al.
Veröffentlicht: (2025)