Can AI Make Conflicts Worse? An Alignment Failure in LLM Deployment Across Conflict Contexts
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | Kryshtal, Andrii |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Can LLMs Understand What We Cannot Say? Measuring Multilevel Alignment Through Abortion Stigma Across Cognitive, Interpersonal, and Structural Levels
par: Sharma, Anika, et autres
Publié: (2025)
par: Sharma, Anika, et autres
Publié: (2025)
CogNarr Ecosystem: Facilitating Group Cognition at Scale
par: Boik, John C.
Publié: (2024)
par: Boik, John C.
Publié: (2024)
The Epistemic Suite: A Post-Foundational Diagnostic Methodology for Assessing AI Knowledge Claims
par: Kelly, Matthew
Publié: (2025)
par: Kelly, Matthew
Publié: (2025)
The Fragility Of Moral Judgment In Large Language Models
par: van Nuenen, Tom, et autres
Publié: (2026)
par: van Nuenen, Tom, et autres
Publié: (2026)
AI and My Values: User Perceptions of LLMs' Ability to Extract, Embody, and Explain Human Values from Casual Conversations
par: Yun, Bhada, et autres
Publié: (2026)
par: Yun, Bhada, et autres
Publié: (2026)
Aurora: Neuro-Symbolic AI Driven Advising Agent
par: Lugones, Lorena Amanda Quincoso, et autres
Publié: (2026)
par: Lugones, Lorena Amanda Quincoso, et autres
Publié: (2026)
Not Everything That Counts Can Be Counted: A Case for Safe Qualitative AI
par: Beltoft, Stine, et autres
Publié: (2025)
par: Beltoft, Stine, et autres
Publié: (2025)
Revealing Hidden Bias in AI: Lessons from Large Language Models
par: Beatty, Django, et autres
Publié: (2024)
par: Beatty, Django, et autres
Publié: (2024)
LLM-FACETS: A Privacy-Preserving Framework for Evaluating LLM Transparency and Accountability
par: Lucas, Tom, et autres
Publié: (2026)
par: Lucas, Tom, et autres
Publié: (2026)
A Socratic RAG Approach to Connect Natural Language Queries on Research Topics with Knowledge Organization Systems
par: Lefton, Lew, et autres
Publié: (2025)
par: Lefton, Lew, et autres
Publié: (2025)
Eroding the Truth-Default: A Causal Analysis of Human Susceptibility to Foundation Model Hallucinations and Disinformation in the Wild
par: Loth, Alexander, et autres
Publié: (2026)
par: Loth, Alexander, et autres
Publié: (2026)
Tatemae: Detecting Alignment Faking via Tool Selection in LLMs
par: Leonesi, Matteo, et autres
Publié: (2026)
par: Leonesi, Matteo, et autres
Publié: (2026)
Superhuman Game AI Disclosure: Expertise and Context Moderate Effects on Trust and Fairness
par: Chua, Jaymari, et autres
Publié: (2025)
par: Chua, Jaymari, et autres
Publié: (2025)
AI to Learn 2.0: A Deliverable-Oriented Governance Framework and Maturity Rubric for Opaque AI in Learning-Intensive Domains
par: Shintani, Seine A.
Publié: (2026)
par: Shintani, Seine A.
Publié: (2026)
From Helpfulness to Toxic Proactivity: Diagnosing Behavioral Misalignment in LLM Agents
par: Wang, Xinyue, et autres
Publié: (2026)
par: Wang, Xinyue, et autres
Publié: (2026)
AgentLeak: A Full-Stack Benchmark for Privacy Leakage in Multi-Agent LLM Systems
par: Yagoubi, Faouzi El, et autres
Publié: (2026)
par: Yagoubi, Faouzi El, et autres
Publié: (2026)
Domain-Specific Constitutional AI: Enhancing Safety in LLM-Powered Mental Health Chatbots
par: Lyu, Chenhan, et autres
Publié: (2025)
par: Lyu, Chenhan, et autres
Publié: (2025)
MixAT: Combining Continuous and Discrete Adversarial Training for LLMs
par: Dékány, Csaba, et autres
Publié: (2025)
par: Dékány, Csaba, et autres
Publié: (2025)
APPSI-139: A Parallel Corpus of English Application Privacy Policy Summarization and Interpretation
par: Zhu, Pengyun, et autres
Publié: (2026)
par: Zhu, Pengyun, et autres
Publié: (2026)
Cultural Encoding in Large Language Models: The Existence Gap in AI-Mediated Brand Discovery
par: Junyao, Huang, et autres
Publié: (2025)
par: Junyao, Huang, et autres
Publié: (2025)
The Hall of AI Fears and Hopes: Comparing the Views of AI Influencers and those of Members of the U.S. Public Through an Interactive Platform
par: Moreira, Gustavo, et autres
Publié: (2025)
par: Moreira, Gustavo, et autres
Publié: (2025)
The Reliance Negotiation Framework: A Dynamic Process Model of Student LLM Engagement in Academic Writing
par: Hossain, Shahin
Publié: (2026)
par: Hossain, Shahin
Publié: (2026)
Developing Critical Thinking in Second Language Learners: Exploring Generative AI like ChatGPT as a Tool for Argumentative Essay Writing
par: Suh, Simon, et autres
Publié: (2025)
par: Suh, Simon, et autres
Publié: (2025)
Collective Constitutional AI: Aligning a Language Model with Public Input
par: Huang, Saffron, et autres
Publié: (2024)
par: Huang, Saffron, et autres
Publié: (2024)
AI-induced sexual harassment: Investigating Contextual Characteristics and User Reactions of Sexual Harassment by a Companion Chatbot
par: Mohammad, et autres
Publié: (2025)
par: Mohammad, et autres
Publié: (2025)
Implicit Geographic Inference in LLM Medical Triage: Language-Driven Disparities in Emergency Recommendations
par: Wong, Qi Han
Publié: (2026)
par: Wong, Qi Han
Publié: (2026)
Human Values in a Single Sentence: Moral Presence, Hierarchies, and Transformer Ensembles on the Schwartz Continuum
par: Yeste, Víctor, et autres
Publié: (2026)
par: Yeste, Víctor, et autres
Publié: (2026)
Auditing Preferences for Brands and Cultures in LLMs
par: Rienecker, Jasmine, et autres
Publié: (2026)
par: Rienecker, Jasmine, et autres
Publié: (2026)
Bye Bye Perspective API: Lessons for Measurement Infrastructure in NLP, CSS and LLM Evaluation
par: Hartmann, David, et autres
Publié: (2026)
par: Hartmann, David, et autres
Publié: (2026)
Controlled Territory and Conflict Tracking (CONTACT): (Geo-)Mapping Occupied Territory from Open Source Intelligence
par: Mandal, Paul K., et autres
Publié: (2025)
par: Mandal, Paul K., et autres
Publié: (2025)
More Context, Larger Models, or Moral Knowledge? A Systematic Study of Schwartz Value Detection in Political Texts
par: Yeste, Víctor, et autres
Publié: (2026)
par: Yeste, Víctor, et autres
Publié: (2026)
Ideology as a Problem: Lightweight Logit Steering for Annotator-Specific Alignment in Social Media Analysis
par: Xia, Wei, et autres
Publié: (2025)
par: Xia, Wei, et autres
Publié: (2025)
Personalities at Play: Probing Alignment in AI Teammates
par: Samadi, Mohammad Amin, et autres
Publié: (2026)
par: Samadi, Mohammad Amin, et autres
Publié: (2026)
Qwerty AI: Explainable Automated Age Rating and Content Safety Assessment for Russian-Language Screenplays
par: Zmanovskii, Nikita
Publié: (2025)
par: Zmanovskii, Nikita
Publié: (2025)
Beyond Imperfect Alternatives with Rulemapping: A Neuro-Symbolic Case Study on Online Hate Speech
par: von Cossel, Oskar
Publié: (2026)
par: von Cossel, Oskar
Publié: (2026)
CAMP: Cumulative Agentic Masking and Pruning for Privacy Protection in Multi-Turn LLM Conversations
par: Panjwani, Aman
Publié: (2026)
par: Panjwani, Aman
Publié: (2026)
Exploring and Mitigating Gender Bias in Encoder-Based Transformer Models
par: Hossain, Ariyan, et autres
Publié: (2025)
par: Hossain, Ariyan, et autres
Publié: (2025)
Seeing Is No Longer Believing: Frontier Image Generation Models, Synthetic Visual Evidence, and Real-World Risk
par: Wu, Shuai, et autres
Publié: (2026)
par: Wu, Shuai, et autres
Publié: (2026)
Benchmarking Deception Probes via Black-to-White Performance Boosts
par: Parrack, Avi, et autres
Publié: (2025)
par: Parrack, Avi, et autres
Publié: (2025)
Powerful Training-Free Membership Inference Against Autoregressive Language Models
par: Ilić, David, et autres
Publié: (2026)
par: Ilić, David, et autres
Publié: (2026)
Documents similaires
-
Can LLMs Understand What We Cannot Say? Measuring Multilevel Alignment Through Abortion Stigma Across Cognitive, Interpersonal, and Structural Levels
par: Sharma, Anika, et autres
Publié: (2025) -
CogNarr Ecosystem: Facilitating Group Cognition at Scale
par: Boik, John C.
Publié: (2024) -
The Epistemic Suite: A Post-Foundational Diagnostic Methodology for Assessing AI Knowledge Claims
par: Kelly, Matthew
Publié: (2025) -
The Fragility Of Moral Judgment In Large Language Models
par: van Nuenen, Tom, et autres
Publié: (2026) -
AI and My Values: User Perceptions of LLMs' Ability to Extract, Embody, and Explain Human Values from Casual Conversations
par: Yun, Bhada, et autres
Publié: (2026)