Collective Constitutional AI: Aligning a Language Model with Public Input
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Saffron, Siddarth, Divya, Lovitt, Liane, Liao, Thomas I., Durmus, Esin, Tamkin, Alex, Ganguli, Deep |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AI-induced sexual harassment: Investigating Contextual Characteristics and User Reactions of Sexual Harassment by a Companion Chatbot
by: Mohammad, et al.
Published: (2025)
by: Mohammad, et al.
Published: (2025)
Understanding Gen Alpha Digital Language: Evaluation of LLM Safety Systems for Content Moderation
by: Mehta, Manisha, et al.
Published: (2025)
by: Mehta, Manisha, et al.
Published: (2025)
"I followed what felt right, not what I was told": Autonomy, Coaching, and Recognizing Bias Through AI-Mediated Dialogue
by: Taheri, Atieh, et al.
Published: (2026)
by: Taheri, Atieh, et al.
Published: (2026)
Resisting AI Solutionism through Workplace Collective Action
by: Zheng, Kevin, et al.
Published: (2025)
by: Zheng, Kevin, et al.
Published: (2025)
Benchmarking Educational LLMs with Analytics: A Case Study on Gender Bias in Feedback
by: Du, Yishan, et al.
Published: (2025)
by: Du, Yishan, et al.
Published: (2025)
SectEval: Evaluating the Latent Sectarian Preferences of Large Language Models
by: Maheshwari, Aditya, et al.
Published: (2026)
by: Maheshwari, Aditya, et al.
Published: (2026)
Chatbot Deployment Considerations for Application-Agnostic Human-Machine Dialogues
by: Rivas, Pablo, et al.
Published: (2025)
by: Rivas, Pablo, et al.
Published: (2025)
Generative UI as an Accessibility Bridge: Lessons from C2C E-Commerce
by: Ryskeldiev, Bektur
Published: (2026)
by: Ryskeldiev, Bektur
Published: (2026)
How Frontier LLMs Adapt to Neurodivergence Context: A Measurement Framework for Surface vs. Structural Change in System-Prompted Responses
by: Gupta, Ishan, et al.
Published: (2026)
by: Gupta, Ishan, et al.
Published: (2026)
What Would GPT Click: Practical Effects of Human-AI Behavioral Misalignment and the Cost of Synthetic Participants in User Experience
by: Kuric, Eduard, et al.
Published: (2026)
by: Kuric, Eduard, et al.
Published: (2026)
The Human-AI Delegation Dilemma: Individual Strategies, Collective Equilibria and Sociotechnical Lock-in
by: Hila, Angjelin
Published: (2026)
by: Hila, Angjelin
Published: (2026)
CURATe: Benchmarking Personalised Alignment of Conversational AI Assistants
by: Alberts, Lize, et al.
Published: (2024)
by: Alberts, Lize, et al.
Published: (2024)
Textual Entailment is not a Better Bias Metric than Token Probability
by: Felkner, Virginia K., et al.
Published: (2025)
by: Felkner, Virginia K., et al.
Published: (2025)
GPT is Not an Annotator: The Necessity of Human Annotation in Fairness Benchmark Construction
by: Felkner, Virginia K., et al.
Published: (2024)
by: Felkner, Virginia K., et al.
Published: (2024)
Using a cognitive architecture to consider antiBlackness in design and development of AI systems
by: Dancy, Christopher L.
Published: (2022)
by: Dancy, Christopher L.
Published: (2022)
Can Humans Tell? A Dual-Axis Study of Human Perception of LLM-Generated News
by: Loth, Alexander, et al.
Published: (2026)
by: Loth, Alexander, et al.
Published: (2026)
Generative Confidants: How do People Experience Trust in Emotional Support from Generative AI?
by: Volpato, Riccardo, et al.
Published: (2026)
by: Volpato, Riccardo, et al.
Published: (2026)
A Contextual Help Browser Extension to Assist Digital Illiterate Internet Users
by: Koutsiaris, Christos
Published: (2026)
by: Koutsiaris, Christos
Published: (2026)
How Utilitarian Are OpenAI's Models Really? Replicating and Reinterpreting Pfeffer, Krügel, and Uhl (2025)
by: Himmelreich, Johannes
Published: (2026)
by: Himmelreich, Johannes
Published: (2026)
LLM-Driven Accessible Interface: A Model-Based Approach
by: Jerry, Blessing, et al.
Published: (2026)
by: Jerry, Blessing, et al.
Published: (2026)
Will AI shape the way we speak? The emerging sociolinguistic influence of synthetic voices
by: Székely, Éva, et al.
Published: (2025)
by: Székely, Éva, et al.
Published: (2025)
The Hall of AI Fears and Hopes: Comparing the Views of AI Influencers and those of Members of the U.S. Public Through an Interactive Platform
by: Moreira, Gustavo, et al.
Published: (2025)
by: Moreira, Gustavo, et al.
Published: (2025)
Evaluation of Hate Speech Detection Using Large Language Models and Geographical Contextualization
by: Zahid, Anwar Hossain, et al.
Published: (2025)
by: Zahid, Anwar Hossain, et al.
Published: (2025)
The Company You Keep: How LLMs Respond to Dark Triad Traits
by: Lu, Zeyi, et al.
Published: (2026)
by: Lu, Zeyi, et al.
Published: (2026)
Balancing Innovation and Integrity: AI Integration in Liberal Arts College Administration
by: Read, Ian Olivo
Published: (2025)
by: Read, Ian Olivo
Published: (2025)
Rejected Dialects: Biases Against African American Language in Reward Models
by: Mire, Joel, et al.
Published: (2025)
by: Mire, Joel, et al.
Published: (2025)
Benchmarking Bengali Dialectal Bias: A Multi-Stage Framework Integrating RAG-Based Translation and Human-Augmented RLAIF
by: Sami, K. M. Jubair, et al.
Published: (2026)
by: Sami, K. M. Jubair, et al.
Published: (2026)
Playing Games with My Heart: An Evaluation of AI Companion Apps
by: Rauh, Maribeth, et al.
Published: (2026)
by: Rauh, Maribeth, et al.
Published: (2026)
AI for Accessible Education: Personalized Audio-Based Learning for Blind Students
by: Yang, Crystal, et al.
Published: (2025)
by: Yang, Crystal, et al.
Published: (2025)
Choreographing Trash Cans: On Speculative Futures of Weak Robots in Public Spaces
by: Axelsson, Minja, et al.
Published: (2025)
by: Axelsson, Minja, et al.
Published: (2025)
Closing Africa's Early Warning Gap: AI Weather Forecasting for Disaster Prevention
by: Ndlovu, Qness
Published: (2026)
by: Ndlovu, Qness
Published: (2026)
Extreme Self-Preference in Language Models
by: Lehr, Steven A., et al.
Published: (2025)
by: Lehr, Steven A., et al.
Published: (2025)
Aurora: Neuro-Symbolic AI Driven Advising Agent
by: Lugones, Lorena Amanda Quincoso, et al.
Published: (2026)
by: Lugones, Lorena Amanda Quincoso, et al.
Published: (2026)
Personalities at Play: Probing Alignment in AI Teammates
by: Samadi, Mohammad Amin, et al.
Published: (2026)
by: Samadi, Mohammad Amin, et al.
Published: (2026)
Disaster Question Answering with LoRA Efficiency and Accurate End Position
by: Yasuno, Takato
Published: (2026)
by: Yasuno, Takato
Published: (2026)
VEAT Quantifies Implicit Associations in Text-to-Video Generator Sora and Reveals Challenges in Bias Mitigation
by: Sun, Yongxu, et al.
Published: (2026)
by: Sun, Yongxu, et al.
Published: (2026)
LLM Wardens: Mitigating Adversarial Persuasion with Third-Party Conversational Oversight
by: Wachowiak, Lennart, et al.
Published: (2026)
by: Wachowiak, Lennart, et al.
Published: (2026)
Characterizing Resource Sharing Practices on Underground Internet Forum Synthetic Non-Consensual Intimate Image Content Creation Communities
by: Medeiros, Bernardo B. P., et al.
Published: (2026)
by: Medeiros, Bernardo B. P., et al.
Published: (2026)
Can LLMs Understand What We Cannot Say? Measuring Multilevel Alignment Through Abortion Stigma Across Cognitive, Interpersonal, and Structural Levels
by: Sharma, Anika, et al.
Published: (2025)
by: Sharma, Anika, et al.
Published: (2025)
AI-based Classification of Customer Support Tickets: State of the Art and Implementation with AutoML
by: Truss, Mario, et al.
Published: (2024)
by: Truss, Mario, et al.
Published: (2024)
Similar Items
-
AI-induced sexual harassment: Investigating Contextual Characteristics and User Reactions of Sexual Harassment by a Companion Chatbot
by: Mohammad, et al.
Published: (2025) -
Understanding Gen Alpha Digital Language: Evaluation of LLM Safety Systems for Content Moderation
by: Mehta, Manisha, et al.
Published: (2025) -
"I followed what felt right, not what I was told": Autonomy, Coaching, and Recognizing Bias Through AI-Mediated Dialogue
by: Taheri, Atieh, et al.
Published: (2026) -
Resisting AI Solutionism through Workplace Collective Action
by: Zheng, Kevin, et al.
Published: (2025) -
Benchmarking Educational LLMs with Analytics: A Case Study on Gender Bias in Feedback
by: Du, Yishan, et al.
Published: (2025)