Normative Conflicts and Shallow AI Alignment
Fuente:
arXiv
Saved in:
| Main Author: | Millière, Raphaël |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Philosophy of Cognitive Science in the Age of Deep Learning
by: Millière, Raphaël
Published: (2024)
by: Millière, Raphaël
Published: (2024)
Language Models as Models of Language
by: Millière, Raphaël
Published: (2024)
by: Millière, Raphaël
Published: (2024)
A Philosophical Introduction to Language Models - Part II: The Way Forward
by: Millière, Raphaël, et al.
Published: (2024)
by: Millière, Raphaël, et al.
Published: (2024)
Anthropocentric bias in language model evaluation
by: Millière, Raphaël, et al.
Published: (2024)
by: Millière, Raphaël, et al.
Published: (2024)
The Vector Grounding Problem
by: Mollo, Dimitri Coelho, et al.
Published: (2023)
by: Mollo, Dimitri Coelho, et al.
Published: (2023)
A Philosophical Introduction to Language Models -- Part I: Continuity With Classic Debates
by: Millière, Raphaël, et al.
Published: (2024)
by: Millière, Raphaël, et al.
Published: (2024)
How Do Transformers Learn Variable Binding in Symbolic Programs?
by: Wu, Yiwei, et al.
Published: (2025)
by: Wu, Yiwei, et al.
Published: (2025)
LLMs as Models for Analogical Reasoning
by: Musker, Sam, et al.
Published: (2024)
by: Musker, Sam, et al.
Published: (2024)
Decoding In-Context Learning: Neuroscience-inspired Analysis of Representations in Large Language Models
by: Yousefi, Safoora, et al.
Published: (2023)
by: Yousefi, Safoora, et al.
Published: (2023)
Why Is RLHF Alignment Shallow? A Gradient Analysis
by: Young, Robin
Published: (2026)
by: Young, Robin
Published: (2026)
Decoupled Proxy Alignment: Mitigating Language Prior Conflict for Multimodal Alignment in MLLM
by: Tan, Chenkun, et al.
Published: (2025)
by: Tan, Chenkun, et al.
Published: (2025)
Philosophy of cognitive science in the age of deep learning
by: Raphaël Millière
Published: (2024)
by: Raphaël Millière
Published: (2024)
CONGRAD:Conflicting Gradient Filtering for Multilingual Preference Alignment
by: Li, Jiangnan, et al.
Published: (2025)
by: Li, Jiangnan, et al.
Published: (2025)
Reward-free Alignment for Conflicting Objectives
by: Chen, Peter, et al.
Published: (2026)
by: Chen, Peter, et al.
Published: (2026)
ConflictBench: Evaluating Human-AI Conflict via Interactive and Visually Grounded Environments
by: Zhao, Weixiang, et al.
Published: (2026)
by: Zhao, Weixiang, et al.
Published: (2026)
Targeting Misalignment: A Conflict-Aware Framework for Reward-Model-based LLM Alignment
by: Liu, Zixuan, et al.
Published: (2025)
by: Liu, Zixuan, et al.
Published: (2025)
AI Alignment Breaks at the Edge
by: Bao, Han, et al.
Published: (2026)
by: Bao, Han, et al.
Published: (2026)
Self-Improvement Towards Pareto Optimality: Mitigating Preference Conflicts in Multi-Objective Alignment
by: Li, Moxin, et al.
Published: (2025)
by: Li, Moxin, et al.
Published: (2025)
Deontological Keyword Bias: The Impact of Modal Expressions on Normative Judgments of Language Models
by: Park, Bumjin, et al.
Published: (2025)
by: Park, Bumjin, et al.
Published: (2025)
Value-Conflict Diagnostics Reveal Widespread Alignment Faking in Language Models
by: Nair, Inderjeet, et al.
Published: (2026)
by: Nair, Inderjeet, et al.
Published: (2026)
Beyond Shallow Heuristics: Leveraging Human Intuition for Curriculum Learning
by: Toborek, Vanessa, et al.
Published: (2025)
by: Toborek, Vanessa, et al.
Published: (2025)
Consensus or Conflict? Fine-Grained Evaluation of Conflicting Answers in Question-Answering
by: Nachshoni, Eviatar, et al.
Published: (2025)
by: Nachshoni, Eviatar, et al.
Published: (2025)
Evaluating Behavioral Alignment in Conflict Dialogue: A Multi-Dimensional Comparison of LLM Agents and Humans
by: Kwon, Deuksin, et al.
Published: (2025)
by: Kwon, Deuksin, et al.
Published: (2025)
Value-Action Alignment in Large Language Models under Privacy-Prosocial Conflict
by: Chen, Guanyu, et al.
Published: (2026)
by: Chen, Guanyu, et al.
Published: (2026)
Challenges and Future Directions of Data-Centric AI Alignment
by: Yeh, Min-Hsuan, et al.
Published: (2024)
by: Yeh, Min-Hsuan, et al.
Published: (2024)
Shallow Cross-Encoders for Low-Latency Retrieval
by: Petrov, Aleksandr V., et al.
Published: (2024)
by: Petrov, Aleksandr V., et al.
Published: (2024)
Omni-Safety under Cross-Modality Conflict: Vulnerabilities, Dynamics Mechanisms and Efficient Alignment
by: Wang, Kun, et al.
Published: (2026)
by: Wang, Kun, et al.
Published: (2026)
Freeze Deep, Train Shallow: Interpretable Layer Allocation for Continued Pre-Training
by: Wu, Yu-Hang, et al.
Published: (2026)
by: Wu, Yu-Hang, et al.
Published: (2026)
Llama SLayer 8B: Shallow Layers Hold the Key to Knowledge Injection
by: Chen, Tianxiang, et al.
Published: (2024)
by: Chen, Tianxiang, et al.
Published: (2024)
Rehearsal: Simulating Conflict to Teach Conflict Resolution
by: Shaikh, Omar, et al.
Published: (2023)
by: Shaikh, Omar, et al.
Published: (2023)
Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest
by: Wu, Addison J., et al.
Published: (2026)
by: Wu, Addison J., et al.
Published: (2026)
Shallow Preference Signals: Large Language Model Aligns Even Better with Truncated Data?
by: Qi, Xuan, et al.
Published: (2025)
by: Qi, Xuan, et al.
Published: (2025)
The Reader is the Metric: How Textual Features and Reader Profiles Explain Conflicting Evaluations of AI Creative Writing
by: Marco, Guillermo, et al.
Published: (2025)
by: Marco, Guillermo, et al.
Published: (2025)
Where Norms and References Collide: Evaluating LLMs on Normative Reasoning
by: Abrams, Mitchell, et al.
Published: (2026)
by: Abrams, Mitchell, et al.
Published: (2026)
Conflicts in Texts: Data, Implications and Challenges
by: Liu, Siyi, et al.
Published: (2025)
by: Liu, Siyi, et al.
Published: (2025)
The Unintended Trade-off of AI Alignment:Balancing Hallucination Mitigation and Safety in LLMs
by: Mahmoud, Omar, et al.
Published: (2025)
by: Mahmoud, Omar, et al.
Published: (2025)
Pluralism in AI Governance: Toward Sociotechnical Alignment and Normative Coherence
by: Nkongolo, Mike Wa
Published: (2026)
by: Nkongolo, Mike Wa
Published: (2026)
Theoretical Understanding of In-Context Learning in Shallow Transformers with Unstructured Data
by: Xing, Yue, et al.
Published: (2024)
by: Xing, Yue, et al.
Published: (2024)
Great Memory, Shallow Reasoning: Limits of $k$NN-LMs
by: Geng, Shangyi, et al.
Published: (2024)
by: Geng, Shangyi, et al.
Published: (2024)
Position: Towards Bidirectional Human-AI Alignment
by: Shen, Hua, et al.
Published: (2024)
by: Shen, Hua, et al.
Published: (2024)
Similar Items
-
Philosophy of Cognitive Science in the Age of Deep Learning
by: Millière, Raphaël
Published: (2024) -
Language Models as Models of Language
by: Millière, Raphaël
Published: (2024) -
A Philosophical Introduction to Language Models - Part II: The Way Forward
by: Millière, Raphaël, et al.
Published: (2024) -
Anthropocentric bias in language model evaluation
by: Millière, Raphaël, et al.
Published: (2024) -
The Vector Grounding Problem
by: Mollo, Dimitri Coelho, et al.
Published: (2023)