Scopes of Alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Varshney, Kush R., Ashktorab, Zahra, Bouneffouf, Djallel, Riemer, Matthew, Weisz, Justin D. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Ultimate Test of Superintelligent AI Agents: Can an AI Balance Care and Control in Asymmetric Relationships?
von: Bouneffouf, Djallel, et al.
Veröffentlicht: (2025)
von: Bouneffouf, Djallel, et al.
Veröffentlicht: (2025)
Contextual Moral Value Alignment Through Context-Based Aggregation
von: Dognin, Pierre, et al.
Veröffentlicht: (2024)
von: Dognin, Pierre, et al.
Veröffentlicht: (2024)
Position: Theory of Mind Benchmarks are Broken for Large Language Models
von: Riemer, Matthew, et al.
Veröffentlicht: (2024)
von: Riemer, Matthew, et al.
Veröffentlicht: (2024)
Assessing AI Utility: The Random Guesser Test for Sequential Decision-Making Systems
von: Ide, Shun, et al.
Veröffentlicht: (2024)
von: Ide, Shun, et al.
Veröffentlicht: (2024)
Empathy and the Right to Be an Exception: What LLMs Can and Cannot Do
von: Kidder, William, et al.
Veröffentlicht: (2024)
von: Kidder, William, et al.
Veröffentlicht: (2024)
Decolonial AI Alignment: Openness, Viśe\d{s}a-Dharma, and Including Excluded Knowledges
von: Varshney, Kush R.
Veröffentlicht: (2023)
von: Varshney, Kush R.
Veröffentlicht: (2023)
An Annotated Reading of 'The Singer of Tales' in the LLM Era
von: Varshney, Kush R.
Veröffentlicht: (2025)
von: Varshney, Kush R.
Veröffentlicht: (2025)
Mitigating Misalignment Contagion by Steering with Implicit Traits
von: Chang, Maria, et al.
Veröffentlicht: (2026)
von: Chang, Maria, et al.
Veröffentlicht: (2026)
Are LLMs Court-Ready? Evaluating Frontier Models on Indian Legal Reasoning
von: Juvekar, Kush, et al.
Veröffentlicht: (2025)
von: Juvekar, Kush, et al.
Veröffentlicht: (2025)
Enhancing Value Alignment of LLMs with Multi-agent system and Combinatorial Fusion
von: Wu, Yuanhong, et al.
Veröffentlicht: (2026)
von: Wu, Yuanhong, et al.
Veröffentlicht: (2026)
Alignment Studio: Aligning Large Language Models to Particular Contextual Regulations
von: Achintalwar, Swapnaja, et al.
Veröffentlicht: (2024)
von: Achintalwar, Swapnaja, et al.
Veröffentlicht: (2024)
Societal Alignment Frameworks Can Improve LLM Alignment
von: Stańczak, Karolina, et al.
Veröffentlicht: (2025)
von: Stańczak, Karolina, et al.
Veröffentlicht: (2025)
Which Contributions Deserve Credit? Perceptions of Attribution in Human-AI Co-Creation
von: He, Jessica, et al.
Veröffentlicht: (2025)
von: He, Jessica, et al.
Veröffentlicht: (2025)
COMPASS: Computational Mapping of Patient-Therapist Alliance Strategies with Language Modeling
von: Lin, Baihan, et al.
Veröffentlicht: (2024)
von: Lin, Baihan, et al.
Veröffentlicht: (2024)
A Scoping Review of LLM-as-a-Judge in Healthcare and the MedJUDGE Framework
von: Li, Chenyu, et al.
Veröffentlicht: (2026)
von: Li, Chenyu, et al.
Veröffentlicht: (2026)
Dropouts in Confidence: Moral Uncertainty in Human-LLM Alignment
von: Kwon, Jea, et al.
Veröffentlicht: (2025)
von: Kwon, Jea, et al.
Veröffentlicht: (2025)
Fluent but Foreign: Even Regional LLMs Lack Cultural Alignment
von: Agarwal, Dhruv, et al.
Veröffentlicht: (2025)
von: Agarwal, Dhruv, et al.
Veröffentlicht: (2025)
Evaluating the Prompt Steerability of Large Language Models
von: Miehling, Erik, et al.
Veröffentlicht: (2024)
von: Miehling, Erik, et al.
Veröffentlicht: (2024)
Exploring Persona-dependent LLM Alignment for the Moral Machine Experiment
von: Kim, Jiseon, et al.
Veröffentlicht: (2025)
von: Kim, Jiseon, et al.
Veröffentlicht: (2025)
The Effectiveness of Approximate Regularized Replay for Efficient Supervised Fine-Tuning of Large Language Models
von: Riemer, Matthew, et al.
Veröffentlicht: (2025)
von: Riemer, Matthew, et al.
Veröffentlicht: (2025)
Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMs
von: Kabir, Mohsinul, et al.
Veröffentlicht: (2025)
von: Kabir, Mohsinul, et al.
Veröffentlicht: (2025)
Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race
von: Sun, Lihao, et al.
Veröffentlicht: (2025)
von: Sun, Lihao, et al.
Veröffentlicht: (2025)
Implicit Behavioral Alignment of Language Agents in High-Stakes Crowd Simulations
von: Wang, Yunzhe, et al.
Veröffentlicht: (2025)
von: Wang, Yunzhe, et al.
Veröffentlicht: (2025)
Balancing Safety and Helpfulness in Healthcare AI Assistants through Iterative Preference Alignment
von: Nghiem, Huy, et al.
Veröffentlicht: (2025)
von: Nghiem, Huy, et al.
Veröffentlicht: (2025)
PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization
von: Jiang, Han, et al.
Veröffentlicht: (2025)
von: Jiang, Han, et al.
Veröffentlicht: (2025)
Beyond Single-Sentence Prompts: Upgrading Value Alignment Benchmarks with Dialogues and Stories
von: Zhang, Yazhou, et al.
Veröffentlicht: (2025)
von: Zhang, Yazhou, et al.
Veröffentlicht: (2025)
Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs
von: Rystrøm, Jonathan, et al.
Veröffentlicht: (2025)
von: Rystrøm, Jonathan, et al.
Veröffentlicht: (2025)
From Descriptive to Prescriptive: Uncover the Social Value Alignment of LLM-based Agents
von: Qu, Jinxian, et al.
Veröffentlicht: (2026)
von: Qu, Jinxian, et al.
Veröffentlicht: (2026)
Training-Free Cultural Alignment of Large Language Models via Persona Disagreement
von: Kiet, Huynh Trung, et al.
Veröffentlicht: (2026)
von: Kiet, Huynh Trung, et al.
Veröffentlicht: (2026)
Bye-bye, Bluebook? Automating Legal Procedure with Large Language Models
von: Dahl, Matthew
Veröffentlicht: (2025)
von: Dahl, Matthew
Veröffentlicht: (2025)
Conversational Topic Recommendation in Counseling and Psychotherapy with Decision Transformer and Large Language Models
von: Gunal, Aylin, et al.
Veröffentlicht: (2024)
von: Gunal, Aylin, et al.
Veröffentlicht: (2024)
Empirical Evidence for Alignment Faking in a Small LLM and Prompt-Based Mitigation Techniques
von: Koorndijk, Jeanice
Veröffentlicht: (2025)
von: Koorndijk, Jeanice
Veröffentlicht: (2025)
Self-Alignment of Large Language Models via Monopolylogue-based Social Scene Simulation
von: Pang, Xianghe, et al.
Veröffentlicht: (2024)
von: Pang, Xianghe, et al.
Veröffentlicht: (2024)
LLMs as Strategic Actors: Behavioral Alignment, Risk Calibration, and Argumentation Framing in Geopolitical Simulations
von: Solopova, Veronika, et al.
Veröffentlicht: (2026)
von: Solopova, Veronika, et al.
Veröffentlicht: (2026)
Political Alignment in Large Language Models: A Multidimensional Audit of Psychometric Identity and Behavioral Bias
von: Sakhawat, Adib, et al.
Veröffentlicht: (2026)
von: Sakhawat, Adib, et al.
Veröffentlicht: (2026)
AI Steerability 360: A Toolkit for Steering Large Language Models
von: Miehling, Erik, et al.
Veröffentlicht: (2026)
von: Miehling, Erik, et al.
Veröffentlicht: (2026)
Lightweight Prompt Engineering for Cognitive Alignment in Educational AI: A OneClickQuiz Case Study
von: Yaacoub, Antoun, et al.
Veröffentlicht: (2025)
von: Yaacoub, Antoun, et al.
Veröffentlicht: (2025)
Are LLMs (Really) Ideological? An IRT-based Analysis and Alignment Tool for Perceived Socio-Economic Bias in LLMs
von: Wachter, Jasmin, et al.
Veröffentlicht: (2025)
von: Wachter, Jasmin, et al.
Veröffentlicht: (2025)
Alignment Whack-a-Mole : Finetuning Activates Verbatim Recall of Copyrighted Books in Large Language Models
von: Liu, Xinyue, et al.
Veröffentlicht: (2026)
von: Liu, Xinyue, et al.
Veröffentlicht: (2026)
Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction
von: Li, Ming, et al.
Veröffentlicht: (2025)
von: Li, Ming, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
The Ultimate Test of Superintelligent AI Agents: Can an AI Balance Care and Control in Asymmetric Relationships?
von: Bouneffouf, Djallel, et al.
Veröffentlicht: (2025) -
Contextual Moral Value Alignment Through Context-Based Aggregation
von: Dognin, Pierre, et al.
Veröffentlicht: (2024) -
Position: Theory of Mind Benchmarks are Broken for Large Language Models
von: Riemer, Matthew, et al.
Veröffentlicht: (2024) -
Assessing AI Utility: The Random Guesser Test for Sequential Decision-Making Systems
von: Ide, Shun, et al.
Veröffentlicht: (2024) -
Empathy and the Right to Be an Exception: What LLMs Can and Cannot Do
von: Kidder, William, et al.
Veröffentlicht: (2024)