Scopes of Alignment
Fuente:
arXiv
Salvato in:
| Autori principali: | Varshney, Kush R., Ashktorab, Zahra, Bouneffouf, Djallel, Riemer, Matthew, Weisz, Justin D. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Ultimate Test of Superintelligent AI Agents: Can an AI Balance Care and Control in Asymmetric Relationships?
di: Bouneffouf, Djallel, et al.
Pubblicazione: (2025)
di: Bouneffouf, Djallel, et al.
Pubblicazione: (2025)
Contextual Moral Value Alignment Through Context-Based Aggregation
di: Dognin, Pierre, et al.
Pubblicazione: (2024)
di: Dognin, Pierre, et al.
Pubblicazione: (2024)
Position: Theory of Mind Benchmarks are Broken for Large Language Models
di: Riemer, Matthew, et al.
Pubblicazione: (2024)
di: Riemer, Matthew, et al.
Pubblicazione: (2024)
Assessing AI Utility: The Random Guesser Test for Sequential Decision-Making Systems
di: Ide, Shun, et al.
Pubblicazione: (2024)
di: Ide, Shun, et al.
Pubblicazione: (2024)
Empathy and the Right to Be an Exception: What LLMs Can and Cannot Do
di: Kidder, William, et al.
Pubblicazione: (2024)
di: Kidder, William, et al.
Pubblicazione: (2024)
Decolonial AI Alignment: Openness, Viśe\d{s}a-Dharma, and Including Excluded Knowledges
di: Varshney, Kush R.
Pubblicazione: (2023)
di: Varshney, Kush R.
Pubblicazione: (2023)
An Annotated Reading of 'The Singer of Tales' in the LLM Era
di: Varshney, Kush R.
Pubblicazione: (2025)
di: Varshney, Kush R.
Pubblicazione: (2025)
Mitigating Misalignment Contagion by Steering with Implicit Traits
di: Chang, Maria, et al.
Pubblicazione: (2026)
di: Chang, Maria, et al.
Pubblicazione: (2026)
Are LLMs Court-Ready? Evaluating Frontier Models on Indian Legal Reasoning
di: Juvekar, Kush, et al.
Pubblicazione: (2025)
di: Juvekar, Kush, et al.
Pubblicazione: (2025)
Enhancing Value Alignment of LLMs with Multi-agent system and Combinatorial Fusion
di: Wu, Yuanhong, et al.
Pubblicazione: (2026)
di: Wu, Yuanhong, et al.
Pubblicazione: (2026)
Alignment Studio: Aligning Large Language Models to Particular Contextual Regulations
di: Achintalwar, Swapnaja, et al.
Pubblicazione: (2024)
di: Achintalwar, Swapnaja, et al.
Pubblicazione: (2024)
Societal Alignment Frameworks Can Improve LLM Alignment
di: Stańczak, Karolina, et al.
Pubblicazione: (2025)
di: Stańczak, Karolina, et al.
Pubblicazione: (2025)
Which Contributions Deserve Credit? Perceptions of Attribution in Human-AI Co-Creation
di: He, Jessica, et al.
Pubblicazione: (2025)
di: He, Jessica, et al.
Pubblicazione: (2025)
COMPASS: Computational Mapping of Patient-Therapist Alliance Strategies with Language Modeling
di: Lin, Baihan, et al.
Pubblicazione: (2024)
di: Lin, Baihan, et al.
Pubblicazione: (2024)
A Scoping Review of LLM-as-a-Judge in Healthcare and the MedJUDGE Framework
di: Li, Chenyu, et al.
Pubblicazione: (2026)
di: Li, Chenyu, et al.
Pubblicazione: (2026)
Dropouts in Confidence: Moral Uncertainty in Human-LLM Alignment
di: Kwon, Jea, et al.
Pubblicazione: (2025)
di: Kwon, Jea, et al.
Pubblicazione: (2025)
Fluent but Foreign: Even Regional LLMs Lack Cultural Alignment
di: Agarwal, Dhruv, et al.
Pubblicazione: (2025)
di: Agarwal, Dhruv, et al.
Pubblicazione: (2025)
Evaluating the Prompt Steerability of Large Language Models
di: Miehling, Erik, et al.
Pubblicazione: (2024)
di: Miehling, Erik, et al.
Pubblicazione: (2024)
Exploring Persona-dependent LLM Alignment for the Moral Machine Experiment
di: Kim, Jiseon, et al.
Pubblicazione: (2025)
di: Kim, Jiseon, et al.
Pubblicazione: (2025)
The Effectiveness of Approximate Regularized Replay for Efficient Supervised Fine-Tuning of Large Language Models
di: Riemer, Matthew, et al.
Pubblicazione: (2025)
di: Riemer, Matthew, et al.
Pubblicazione: (2025)
Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMs
di: Kabir, Mohsinul, et al.
Pubblicazione: (2025)
di: Kabir, Mohsinul, et al.
Pubblicazione: (2025)
Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race
di: Sun, Lihao, et al.
Pubblicazione: (2025)
di: Sun, Lihao, et al.
Pubblicazione: (2025)
Implicit Behavioral Alignment of Language Agents in High-Stakes Crowd Simulations
di: Wang, Yunzhe, et al.
Pubblicazione: (2025)
di: Wang, Yunzhe, et al.
Pubblicazione: (2025)
Balancing Safety and Helpfulness in Healthcare AI Assistants through Iterative Preference Alignment
di: Nghiem, Huy, et al.
Pubblicazione: (2025)
di: Nghiem, Huy, et al.
Pubblicazione: (2025)
PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization
di: Jiang, Han, et al.
Pubblicazione: (2025)
di: Jiang, Han, et al.
Pubblicazione: (2025)
Beyond Single-Sentence Prompts: Upgrading Value Alignment Benchmarks with Dialogues and Stories
di: Zhang, Yazhou, et al.
Pubblicazione: (2025)
di: Zhang, Yazhou, et al.
Pubblicazione: (2025)
Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs
di: Rystrøm, Jonathan, et al.
Pubblicazione: (2025)
di: Rystrøm, Jonathan, et al.
Pubblicazione: (2025)
From Descriptive to Prescriptive: Uncover the Social Value Alignment of LLM-based Agents
di: Qu, Jinxian, et al.
Pubblicazione: (2026)
di: Qu, Jinxian, et al.
Pubblicazione: (2026)
Training-Free Cultural Alignment of Large Language Models via Persona Disagreement
di: Kiet, Huynh Trung, et al.
Pubblicazione: (2026)
di: Kiet, Huynh Trung, et al.
Pubblicazione: (2026)
Bye-bye, Bluebook? Automating Legal Procedure with Large Language Models
di: Dahl, Matthew
Pubblicazione: (2025)
di: Dahl, Matthew
Pubblicazione: (2025)
Conversational Topic Recommendation in Counseling and Psychotherapy with Decision Transformer and Large Language Models
di: Gunal, Aylin, et al.
Pubblicazione: (2024)
di: Gunal, Aylin, et al.
Pubblicazione: (2024)
Empirical Evidence for Alignment Faking in a Small LLM and Prompt-Based Mitigation Techniques
di: Koorndijk, Jeanice
Pubblicazione: (2025)
di: Koorndijk, Jeanice
Pubblicazione: (2025)
Self-Alignment of Large Language Models via Monopolylogue-based Social Scene Simulation
di: Pang, Xianghe, et al.
Pubblicazione: (2024)
di: Pang, Xianghe, et al.
Pubblicazione: (2024)
LLMs as Strategic Actors: Behavioral Alignment, Risk Calibration, and Argumentation Framing in Geopolitical Simulations
di: Solopova, Veronika, et al.
Pubblicazione: (2026)
di: Solopova, Veronika, et al.
Pubblicazione: (2026)
Political Alignment in Large Language Models: A Multidimensional Audit of Psychometric Identity and Behavioral Bias
di: Sakhawat, Adib, et al.
Pubblicazione: (2026)
di: Sakhawat, Adib, et al.
Pubblicazione: (2026)
AI Steerability 360: A Toolkit for Steering Large Language Models
di: Miehling, Erik, et al.
Pubblicazione: (2026)
di: Miehling, Erik, et al.
Pubblicazione: (2026)
Lightweight Prompt Engineering for Cognitive Alignment in Educational AI: A OneClickQuiz Case Study
di: Yaacoub, Antoun, et al.
Pubblicazione: (2025)
di: Yaacoub, Antoun, et al.
Pubblicazione: (2025)
Are LLMs (Really) Ideological? An IRT-based Analysis and Alignment Tool for Perceived Socio-Economic Bias in LLMs
di: Wachter, Jasmin, et al.
Pubblicazione: (2025)
di: Wachter, Jasmin, et al.
Pubblicazione: (2025)
Alignment Whack-a-Mole : Finetuning Activates Verbatim Recall of Copyrighted Books in Large Language Models
di: Liu, Xinyue, et al.
Pubblicazione: (2026)
di: Liu, Xinyue, et al.
Pubblicazione: (2026)
Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction
di: Li, Ming, et al.
Pubblicazione: (2025)
di: Li, Ming, et al.
Pubblicazione: (2025)
Documenti analoghi
-
The Ultimate Test of Superintelligent AI Agents: Can an AI Balance Care and Control in Asymmetric Relationships?
di: Bouneffouf, Djallel, et al.
Pubblicazione: (2025) -
Contextual Moral Value Alignment Through Context-Based Aggregation
di: Dognin, Pierre, et al.
Pubblicazione: (2024) -
Position: Theory of Mind Benchmarks are Broken for Large Language Models
di: Riemer, Matthew, et al.
Pubblicazione: (2024) -
Assessing AI Utility: The Random Guesser Test for Sequential Decision-Making Systems
di: Ide, Shun, et al.
Pubblicazione: (2024) -
Empathy and the Right to Be an Exception: What LLMs Can and Cannot Do
di: Kidder, William, et al.
Pubblicazione: (2024)