Contextual Moral Value Alignment Through Context-Based Aggregation
Fuente:
arXiv
Saved in:
| Main Authors: | Dognin, Pierre, Rios, Jesus, Luss, Ronny, Padhi, Inkit, Riemer, Matthew D, Liu, Miao, Sattigeri, Prasanna, Nagireddy, Manish, Varshney, Kush R., Bouneffouf, Djallel |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Value Alignment from Unstructured Text
by: Padhi, Inkit, et al.
Published: (2024)
by: Padhi, Inkit, et al.
Published: (2024)
The Ultimate Test of Superintelligent AI Agents: Can an AI Balance Care and Control in Asymmetric Relationships?
by: Bouneffouf, Djallel, et al.
Published: (2025)
by: Bouneffouf, Djallel, et al.
Published: (2025)
When in Doubt, Cascade: Towards Building Efficient and Capable Guardrails
by: Nagireddy, Manish, et al.
Published: (2024)
by: Nagireddy, Manish, et al.
Published: (2024)
Answering the Wrong Question: Reasoning Trace Inversion for Abstention in LLMs
by: Gourabathina, Abinitha, et al.
Published: (2026)
by: Gourabathina, Abinitha, et al.
Published: (2026)
Alignment Studio: Aligning Large Language Models to Particular Contextual Regulations
by: Achintalwar, Swapnaja, et al.
Published: (2024)
by: Achintalwar, Swapnaja, et al.
Published: (2024)
Scopes of Alignment
by: Varshney, Kush R., et al.
Published: (2025)
by: Varshney, Kush R., et al.
Published: (2025)
Programming Refusal with Conditional Activation Steering
by: Lee, Bruce W., et al.
Published: (2024)
by: Lee, Bruce W., et al.
Published: (2024)
Mitigating Misalignment Contagion by Steering with Implicit Traits
by: Chang, Maria, et al.
Published: (2026)
by: Chang, Maria, et al.
Published: (2026)
Sparsity May Be All You Need: Sparse Random Parameter Adaptation
by: Rios, Jesus, et al.
Published: (2025)
by: Rios, Jesus, et al.
Published: (2025)
The Effectiveness of Approximate Regularized Replay for Efficient Supervised Fine-Tuning of Large Language Models
by: Riemer, Matthew, et al.
Published: (2025)
by: Riemer, Matthew, et al.
Published: (2025)
Enhancing Value Alignment of LLMs with Multi-agent system and Combinatorial Fusion
by: Wu, Yuanhong, et al.
Published: (2026)
by: Wu, Yuanhong, et al.
Published: (2026)
Agentic AI Needs a Systems Theory
by: Miehling, Erik, et al.
Published: (2025)
by: Miehling, Erik, et al.
Published: (2025)
Evaluating the Prompt Steerability of Large Language Models
by: Miehling, Erik, et al.
Published: (2024)
by: Miehling, Erik, et al.
Published: (2024)
Survey: Multi-Armed Bandits Meet Large Language Models
by: Bouneffouf, Djallel, et al.
Published: (2025)
by: Bouneffouf, Djallel, et al.
Published: (2025)
Position: Theory of Mind Benchmarks are Broken for Large Language Models
by: Riemer, Matthew, et al.
Published: (2024)
by: Riemer, Matthew, et al.
Published: (2024)
Keeping Up with the Language Models: Systematic Benchmark Extension for Bias Auditing
by: Baldini, Ioana, et al.
Published: (2023)
by: Baldini, Ioana, et al.
Published: (2023)
Multi-Level Explanations for Generative Language Models
by: Paes, Lucas Monteiro, et al.
Published: (2024)
by: Paes, Lucas Monteiro, et al.
Published: (2024)
An Algebraic Exposition of the Theory of Dyadic Morality
by: Varshney, Kush R.
Published: (2026)
by: Varshney, Kush R.
Published: (2026)
Language Models in Dialogue: Conversational Maxims for Human-AI Interactions
by: Miehling, Erik, et al.
Published: (2024)
by: Miehling, Erik, et al.
Published: (2024)
Granite Guardian
by: Padhi, Inkit, et al.
Published: (2024)
by: Padhi, Inkit, et al.
Published: (2024)
Building a Foundational Guardrail for General Agentic Systems via Synthetic Data
by: Huang, Yue, et al.
Published: (2025)
by: Huang, Yue, et al.
Published: (2025)
WikiContradict: A Benchmark for Evaluating LLMs on Real-World Knowledge Conflicts from Wikipedia
by: Hou, Yufang, et al.
Published: (2024)
by: Hou, Yufang, et al.
Published: (2024)
Assessing AI Utility: The Random Guesser Test for Sequential Decision-Making Systems
by: Ide, Shun, et al.
Published: (2024)
by: Ide, Shun, et al.
Published: (2024)
Conversational Topic Recommendation in Counseling and Psychotherapy with Decision Transformer and Large Language Models
by: Gunal, Aylin, et al.
Published: (2024)
by: Gunal, Aylin, et al.
Published: (2024)
Decolonial AI Alignment: Openness, Viśe\d{s}a-Dharma, and Including Excluded Knowledges
by: Varshney, Kush R.
Published: (2023)
by: Varshney, Kush R.
Published: (2023)
When Stability meets Sufficiency: Informative Explanations that do not Overwhelm
by: Luss, Ronny, et al.
Published: (2021)
by: Luss, Ronny, et al.
Published: (2021)
Targeted Advertising on Social Networks Using Online Variational Tensor Regression
by: Idé, Tsuyoshi, et al.
Published: (2022)
by: Idé, Tsuyoshi, et al.
Published: (2022)
Image Captioning as an Assistive Technology: Lessons Learned from VizWiz 2020 Challenge
by: Dognin, Pierre, et al.
Published: (2020)
by: Dognin, Pierre, et al.
Published: (2020)
Detectors for Safe and Reliable LLMs: Implementations, Uses, and Limitations
by: Achintalwar, Swapnaja, et al.
Published: (2024)
by: Achintalwar, Swapnaja, et al.
Published: (2024)
Interpolating Item and User Fairness in Multi-Sided Recommendations
by: Chen, Qinyi, et al.
Published: (2023)
by: Chen, Qinyi, et al.
Published: (2023)
An Annotated Reading of 'The Singer of Tales' in the LLM Era
by: Varshney, Kush R.
Published: (2025)
by: Varshney, Kush R.
Published: (2025)
CELL your Model: Contrastive Explanations for Large Language Models
by: Luss, Ronny, et al.
Published: (2024)
by: Luss, Ronny, et al.
Published: (2024)
The RealHumanEval: Evaluating Large Language Models' Abilities to Support Programmers
by: Mozannar, Hussein, et al.
Published: (2024)
by: Mozannar, Hussein, et al.
Published: (2024)
AI Steerability 360: A Toolkit for Steering Large Language Models
by: Miehling, Erik, et al.
Published: (2026)
by: Miehling, Erik, et al.
Published: (2026)
COMPASS: Computational Mapping of Patient-Therapist Alliance Strategies with Language Modeling
by: Lin, Baihan, et al.
Published: (2024)
by: Lin, Baihan, et al.
Published: (2024)
Split, Unlearn, Merge: Leveraging Data Attributes for More Effective Unlearning in LLMs
by: Kadhe, Swanand Ravindra, et al.
Published: (2024)
by: Kadhe, Swanand Ravindra, et al.
Published: (2024)
Adversarial Prompt Evaluation: Systematic Benchmarking of Guardrails Against Prompt Input Attacks on LLMs
by: Zizzo, Giulio, et al.
Published: (2025)
by: Zizzo, Giulio, et al.
Published: (2025)
Hétérocères nouveaux de l'Amérique du Sud
by: Dognin, Paul
Published: (1901)
by: Dognin, Paul
Published: (1901)
Heterocores nouveaux de l'Amerique du Sud
by: Dognin, Paul
Published: (1913)
by: Dognin, Paul
Published: (1913)
Auditing and Generating Synthetic Data with Controllable Trust Trade-offs
by: Belgodere, Brian, et al.
Published: (2023)
by: Belgodere, Brian, et al.
Published: (2023)
Similar Items
-
Value Alignment from Unstructured Text
by: Padhi, Inkit, et al.
Published: (2024) -
The Ultimate Test of Superintelligent AI Agents: Can an AI Balance Care and Control in Asymmetric Relationships?
by: Bouneffouf, Djallel, et al.
Published: (2025) -
When in Doubt, Cascade: Towards Building Efficient and Capable Guardrails
by: Nagireddy, Manish, et al.
Published: (2024) -
Answering the Wrong Question: Reasoning Trace Inversion for Abstention in LLMs
by: Gourabathina, Abinitha, et al.
Published: (2026) -
Alignment Studio: Aligning Large Language Models to Particular Contextual Regulations
by: Achintalwar, Swapnaja, et al.
Published: (2024)