Normative Conflicts and Shallow AI Alignment
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | Millière, Raphaël |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Philosophy of Cognitive Science in the Age of Deep Learning
par: Millière, Raphaël
Publié: (2024)
par: Millière, Raphaël
Publié: (2024)
Language Models as Models of Language
par: Millière, Raphaël
Publié: (2024)
par: Millière, Raphaël
Publié: (2024)
A Philosophical Introduction to Language Models - Part II: The Way Forward
par: Millière, Raphaël, et autres
Publié: (2024)
par: Millière, Raphaël, et autres
Publié: (2024)
Anthropocentric bias in language model evaluation
par: Millière, Raphaël, et autres
Publié: (2024)
par: Millière, Raphaël, et autres
Publié: (2024)
The Vector Grounding Problem
par: Mollo, Dimitri Coelho, et autres
Publié: (2023)
par: Mollo, Dimitri Coelho, et autres
Publié: (2023)
A Philosophical Introduction to Language Models -- Part I: Continuity With Classic Debates
par: Millière, Raphaël, et autres
Publié: (2024)
par: Millière, Raphaël, et autres
Publié: (2024)
How Do Transformers Learn Variable Binding in Symbolic Programs?
par: Wu, Yiwei, et autres
Publié: (2025)
par: Wu, Yiwei, et autres
Publié: (2025)
LLMs as Models for Analogical Reasoning
par: Musker, Sam, et autres
Publié: (2024)
par: Musker, Sam, et autres
Publié: (2024)
Decoding In-Context Learning: Neuroscience-inspired Analysis of Representations in Large Language Models
par: Yousefi, Safoora, et autres
Publié: (2023)
par: Yousefi, Safoora, et autres
Publié: (2023)
Why Is RLHF Alignment Shallow? A Gradient Analysis
par: Young, Robin
Publié: (2026)
par: Young, Robin
Publié: (2026)
Decoupled Proxy Alignment: Mitigating Language Prior Conflict for Multimodal Alignment in MLLM
par: Tan, Chenkun, et autres
Publié: (2025)
par: Tan, Chenkun, et autres
Publié: (2025)
Philosophy of cognitive science in the age of deep learning
par: Raphaël Millière
Publié: (2024)
par: Raphaël Millière
Publié: (2024)
CONGRAD:Conflicting Gradient Filtering for Multilingual Preference Alignment
par: Li, Jiangnan, et autres
Publié: (2025)
par: Li, Jiangnan, et autres
Publié: (2025)
Reward-free Alignment for Conflicting Objectives
par: Chen, Peter, et autres
Publié: (2026)
par: Chen, Peter, et autres
Publié: (2026)
ConflictBench: Evaluating Human-AI Conflict via Interactive and Visually Grounded Environments
par: Zhao, Weixiang, et autres
Publié: (2026)
par: Zhao, Weixiang, et autres
Publié: (2026)
Targeting Misalignment: A Conflict-Aware Framework for Reward-Model-based LLM Alignment
par: Liu, Zixuan, et autres
Publié: (2025)
par: Liu, Zixuan, et autres
Publié: (2025)
AI Alignment Breaks at the Edge
par: Bao, Han, et autres
Publié: (2026)
par: Bao, Han, et autres
Publié: (2026)
Self-Improvement Towards Pareto Optimality: Mitigating Preference Conflicts in Multi-Objective Alignment
par: Li, Moxin, et autres
Publié: (2025)
par: Li, Moxin, et autres
Publié: (2025)
Deontological Keyword Bias: The Impact of Modal Expressions on Normative Judgments of Language Models
par: Park, Bumjin, et autres
Publié: (2025)
par: Park, Bumjin, et autres
Publié: (2025)
Value-Conflict Diagnostics Reveal Widespread Alignment Faking in Language Models
par: Nair, Inderjeet, et autres
Publié: (2026)
par: Nair, Inderjeet, et autres
Publié: (2026)
Beyond Shallow Heuristics: Leveraging Human Intuition for Curriculum Learning
par: Toborek, Vanessa, et autres
Publié: (2025)
par: Toborek, Vanessa, et autres
Publié: (2025)
Consensus or Conflict? Fine-Grained Evaluation of Conflicting Answers in Question-Answering
par: Nachshoni, Eviatar, et autres
Publié: (2025)
par: Nachshoni, Eviatar, et autres
Publié: (2025)
Evaluating Behavioral Alignment in Conflict Dialogue: A Multi-Dimensional Comparison of LLM Agents and Humans
par: Kwon, Deuksin, et autres
Publié: (2025)
par: Kwon, Deuksin, et autres
Publié: (2025)
Challenges and Future Directions of Data-Centric AI Alignment
par: Yeh, Min-Hsuan, et autres
Publié: (2024)
par: Yeh, Min-Hsuan, et autres
Publié: (2024)
Value-Action Alignment in Large Language Models under Privacy-Prosocial Conflict
par: Chen, Guanyu, et autres
Publié: (2026)
par: Chen, Guanyu, et autres
Publié: (2026)
Shallow Cross-Encoders for Low-Latency Retrieval
par: Petrov, Aleksandr V., et autres
Publié: (2024)
par: Petrov, Aleksandr V., et autres
Publié: (2024)
Omni-Safety under Cross-Modality Conflict: Vulnerabilities, Dynamics Mechanisms and Efficient Alignment
par: Wang, Kun, et autres
Publié: (2026)
par: Wang, Kun, et autres
Publié: (2026)
Freeze Deep, Train Shallow: Interpretable Layer Allocation for Continued Pre-Training
par: Wu, Yu-Hang, et autres
Publié: (2026)
par: Wu, Yu-Hang, et autres
Publié: (2026)
Llama SLayer 8B: Shallow Layers Hold the Key to Knowledge Injection
par: Chen, Tianxiang, et autres
Publié: (2024)
par: Chen, Tianxiang, et autres
Publié: (2024)
Rehearsal: Simulating Conflict to Teach Conflict Resolution
par: Shaikh, Omar, et autres
Publié: (2023)
par: Shaikh, Omar, et autres
Publié: (2023)
Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest
par: Wu, Addison J., et autres
Publié: (2026)
par: Wu, Addison J., et autres
Publié: (2026)
Shallow Preference Signals: Large Language Model Aligns Even Better with Truncated Data?
par: Qi, Xuan, et autres
Publié: (2025)
par: Qi, Xuan, et autres
Publié: (2025)
The Reader is the Metric: How Textual Features and Reader Profiles Explain Conflicting Evaluations of AI Creative Writing
par: Marco, Guillermo, et autres
Publié: (2025)
par: Marco, Guillermo, et autres
Publié: (2025)
Conflicts in Texts: Data, Implications and Challenges
par: Liu, Siyi, et autres
Publié: (2025)
par: Liu, Siyi, et autres
Publié: (2025)
Where Norms and References Collide: Evaluating LLMs on Normative Reasoning
par: Abrams, Mitchell, et autres
Publié: (2026)
par: Abrams, Mitchell, et autres
Publié: (2026)
The Unintended Trade-off of AI Alignment:Balancing Hallucination Mitigation and Safety in LLMs
par: Mahmoud, Omar, et autres
Publié: (2025)
par: Mahmoud, Omar, et autres
Publié: (2025)
Theoretical Understanding of In-Context Learning in Shallow Transformers with Unstructured Data
par: Xing, Yue, et autres
Publié: (2024)
par: Xing, Yue, et autres
Publié: (2024)
Great Memory, Shallow Reasoning: Limits of $k$NN-LMs
par: Geng, Shangyi, et autres
Publié: (2024)
par: Geng, Shangyi, et autres
Publié: (2024)
Pluralism in AI Governance: Toward Sociotechnical Alignment and Normative Coherence
par: Nkongolo, Mike Wa
Publié: (2026)
par: Nkongolo, Mike Wa
Publié: (2026)
Position: Towards Bidirectional Human-AI Alignment
par: Shen, Hua, et autres
Publié: (2024)
par: Shen, Hua, et autres
Publié: (2024)
Documents similaires
-
Philosophy of Cognitive Science in the Age of Deep Learning
par: Millière, Raphaël
Publié: (2024) -
Language Models as Models of Language
par: Millière, Raphaël
Publié: (2024) -
A Philosophical Introduction to Language Models - Part II: The Way Forward
par: Millière, Raphaël, et autres
Publié: (2024) -
Anthropocentric bias in language model evaluation
par: Millière, Raphaël, et autres
Publié: (2024) -
The Vector Grounding Problem
par: Mollo, Dimitri Coelho, et autres
Publié: (2023)