Are Dilemmas and Conflicts in LLM Alignment Solvable? A View from Priority Graph
Fuente:
arXiv
Salvato in:
| Autori principali: | Tang, Zhenheng, Liu, Xiang, Wang, Qian, Choi, Eunsol, Li, Bo, Chu, Xiaowen |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Lottery LLM Hypothesis, Rethinking What Abilities Should LLM Compression Preserve?
di: Tang, Zhenheng, et al.
Pubblicazione: (2025)
di: Tang, Zhenheng, et al.
Pubblicazione: (2025)
Mediator: Memory-efficient LLM Merging with Less Parameter Conflicts and Uncertainty Based Routing
di: Lai, Kunfeng, et al.
Pubblicazione: (2025)
di: Lai, Kunfeng, et al.
Pubblicazione: (2025)
Is Your LLM-as-a-Recommender Agent Trustable? LLMs' Recommendation is Easily Hacked by Biases (Preferences)
di: Tang, Zichen, et al.
Pubblicazione: (2026)
di: Tang, Zichen, et al.
Pubblicazione: (2026)
Growth First, Care Second? Tracing the Landscape of LLM Value Preferences in Everyday Dilemmas
di: Chen, Zhiyi, et al.
Pubblicazione: (2026)
di: Chen, Zhiyi, et al.
Pubblicazione: (2026)
An Evaluation of Cultural Value Alignment in LLM
di: Sukiennik, Nicholas, et al.
Pubblicazione: (2025)
di: Sukiennik, Nicholas, et al.
Pubblicazione: (2025)
When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas
di: Backmann, Steffen, et al.
Pubblicazione: (2025)
di: Backmann, Steffen, et al.
Pubblicazione: (2025)
Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression
di: Liu, Xiang, et al.
Pubblicazione: (2025)
di: Liu, Xiang, et al.
Pubblicazione: (2025)
Societal Alignment Frameworks Can Improve LLM Alignment
di: Stańczak, Karolina, et al.
Pubblicazione: (2025)
di: Stańczak, Karolina, et al.
Pubblicazione: (2025)
The Singapore Consensus on Global AI Safety Research Priorities
di: Bengio, Yoshua, et al.
Pubblicazione: (2025)
di: Bengio, Yoshua, et al.
Pubblicazione: (2025)
Copyleft for Alleviating AIGC Copyright Dilemma: What-if Analysis, Public Perception and Implications
di: Guo, Xinwei, et al.
Pubblicazione: (2024)
di: Guo, Xinwei, et al.
Pubblicazione: (2024)
Expert Survey: AI Reliability & Security Research Priorities
di: O'Brien, Joe, et al.
Pubblicazione: (2025)
di: O'Brien, Joe, et al.
Pubblicazione: (2025)
Normative Evaluation of Large Language Models with Everyday Moral Dilemmas
di: Sachdeva, Pratik S., et al.
Pubblicazione: (2025)
di: Sachdeva, Pratik S., et al.
Pubblicazione: (2025)
Beyond Static Question Banks: Dynamic Knowledge Expansion via LLM-Automated Graph Construction and Adaptive Generation
di: Wang, Yingquan, et al.
Pubblicazione: (2026)
di: Wang, Yingquan, et al.
Pubblicazione: (2026)
Thinking in Graphs with CoMAP: A Shared Visual Workspace for Designing Project-Based Learning
di: Li, Ruijia, et al.
Pubblicazione: (2026)
di: Li, Ruijia, et al.
Pubblicazione: (2026)
The Dilemma of Uncertainty Estimation for General Purpose AI in the EU AI Act
di: Valdenegro-Toro, Matias, et al.
Pubblicazione: (2024)
di: Valdenegro-Toro, Matias, et al.
Pubblicazione: (2024)
LSSF: Safety Alignment for Large Language Models through Low-Rank Safety Subspace Fusion
di: Zhou, Guanghao, et al.
Pubblicazione: (2026)
di: Zhou, Guanghao, et al.
Pubblicazione: (2026)
Generative Value Conflicts Reveal LLM Priorities
di: Liu, Andy, et al.
Pubblicazione: (2025)
di: Liu, Andy, et al.
Pubblicazione: (2025)
LLM Safety Alignment is Divergence Estimation in Disguise
di: Haldar, Rajdeep, et al.
Pubblicazione: (2025)
di: Haldar, Rajdeep, et al.
Pubblicazione: (2025)
Moral Alignment for LLM Agents
di: Tennant, Elizaveta, et al.
Pubblicazione: (2024)
di: Tennant, Elizaveta, et al.
Pubblicazione: (2024)
Social Catalysts, Not Moral Agents: The Illusion of Alignment in LLM Societies
di: Hu, Yueqing, et al.
Pubblicazione: (2026)
di: Hu, Yueqing, et al.
Pubblicazione: (2026)
Wide Reflective Equilibrium in LLM Alignment: Bridging Moral Epistemology and AI Safety
di: Brophy, Matthew
Pubblicazione: (2025)
di: Brophy, Matthew
Pubblicazione: (2025)
Dropouts in Confidence: Moral Uncertainty in Human-LLM Alignment
di: Kwon, Jea, et al.
Pubblicazione: (2025)
di: Kwon, Jea, et al.
Pubblicazione: (2025)
OmniReview: A Large-scale Benchmark and LLM-enhanced Framework for Realistic Reviewer Recommendation
di: Huang, Yehua, et al.
Pubblicazione: (2026)
di: Huang, Yehua, et al.
Pubblicazione: (2026)
A More Advanced Group Polarization Measurement Approach Based on LLM-Based Agents and Graphs
di: Liu, Zixin, et al.
Pubblicazione: (2024)
di: Liu, Zixin, et al.
Pubblicazione: (2024)
Exploring Persona-dependent LLM Alignment for the Moral Machine Experiment
di: Kim, Jiseon, et al.
Pubblicazione: (2025)
di: Kim, Jiseon, et al.
Pubblicazione: (2025)
Bias in Decision-Making for AI's Ethical Dilemmas: A Comparative Study of ChatGPT and Claude
di: Xu, Wentao, et al.
Pubblicazione: (2025)
di: Xu, Wentao, et al.
Pubblicazione: (2025)
From Delegates to Trustees: How Optimizing for Long-Term Interests Shapes Bias and Alignment in LLM
di: Fulay, Suyash, et al.
Pubblicazione: (2025)
di: Fulay, Suyash, et al.
Pubblicazione: (2025)
"Pull or Not to Pull?'': Investigating Moral Biases in Leading Large Language Models Across Ethical Dilemmas
di: Ding, Junchen, et al.
Pubblicazione: (2025)
di: Ding, Junchen, et al.
Pubblicazione: (2025)
Real-World AI Evaluation: How FRAME Generates Systematic Evidence to Resolve the Decision-Maker's Dilemma
di: Schwartz, Reva, et al.
Pubblicazione: (2026)
di: Schwartz, Reva, et al.
Pubblicazione: (2026)
Leveraging Pedagogical Theories to Understand Student Learning Process with Graph-based Reasonable Knowledge Tracing
di: Cui, Jiajun, et al.
Pubblicazione: (2024)
di: Cui, Jiajun, et al.
Pubblicazione: (2024)
COMPASS: A Framework for Evaluating Organization-Specific Policy Alignment in LLMs
di: Choi, Dasol, et al.
Pubblicazione: (2026)
di: Choi, Dasol, et al.
Pubblicazione: (2026)
FedImpro: Measuring and Improving Client Update in Federated Learning
di: Tang, Zhenheng, et al.
Pubblicazione: (2024)
di: Tang, Zhenheng, et al.
Pubblicazione: (2024)
Should We Really Edit Language Models? On the Evaluation of Edited Language Models
di: Li, Qi, et al.
Pubblicazione: (2024)
di: Li, Qi, et al.
Pubblicazione: (2024)
Trustless Autonomy: Understanding Motivations, Benefits, and Governance Dilemmas in Self-Sovereign Decentralized AI Agents
di: Hu, Botao Amber, et al.
Pubblicazione: (2025)
di: Hu, Botao Amber, et al.
Pubblicazione: (2025)
RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
di: Shin, Jisu, et al.
Pubblicazione: (2025)
di: Shin, Jisu, et al.
Pubblicazione: (2025)
Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest
di: Wu, Addison J., et al.
Pubblicazione: (2026)
di: Wu, Addison J., et al.
Pubblicazione: (2026)
From Descriptive to Prescriptive: Uncover the Social Value Alignment of LLM-based Agents
di: Qu, Jinxian, et al.
Pubblicazione: (2026)
di: Qu, Jinxian, et al.
Pubblicazione: (2026)
The Convergent Ethics of AI? Analyzing Moral Foundation Priorities in Large Language Models with a Multi-Framework Approach
di: Coleman, Chad, et al.
Pubblicazione: (2025)
di: Coleman, Chad, et al.
Pubblicazione: (2025)
Whose Truth? Pluralistic Geo-Alignment for (Agentic) AI
di: Janowicz, Krzysztof, et al.
Pubblicazione: (2025)
di: Janowicz, Krzysztof, et al.
Pubblicazione: (2025)
The Fake Friend Dilemma: Trust and the Political Economy of Conversational AI
di: Erickson, Jacob
Pubblicazione: (2026)
di: Erickson, Jacob
Pubblicazione: (2026)
Documenti analoghi
-
The Lottery LLM Hypothesis, Rethinking What Abilities Should LLM Compression Preserve?
di: Tang, Zhenheng, et al.
Pubblicazione: (2025) -
Mediator: Memory-efficient LLM Merging with Less Parameter Conflicts and Uncertainty Based Routing
di: Lai, Kunfeng, et al.
Pubblicazione: (2025) -
Is Your LLM-as-a-Recommender Agent Trustable? LLMs' Recommendation is Easily Hacked by Biases (Preferences)
di: Tang, Zichen, et al.
Pubblicazione: (2026) -
Growth First, Care Second? Tracing the Landscape of LLM Value Preferences in Everyday Dilemmas
di: Chen, Zhiyi, et al.
Pubblicazione: (2026) -
An Evaluation of Cultural Value Alignment in LLM
di: Sukiennik, Nicholas, et al.
Pubblicazione: (2025)