Human Preferences for Constructive Interactions in Language Model Alignment
Fuente:
arXiv
Salvato in:
| Autori principali: | Kyrychenko, Yara, Roozenbeek, Jon, Davidson, Brandon, van der Linden, Sander, Debnath, Ramit |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Addressing Climate Action Misperceptions with Generative AI
di: Remshard, Miriam, et al.
Pubblicazione: (2026)
di: Remshard, Miriam, et al.
Pubblicazione: (2026)
Generative Language Models Exhibit Social Identity Biases
di: Hu, Tiancheng, et al.
Pubblicazione: (2023)
di: Hu, Tiancheng, et al.
Pubblicazione: (2023)
Alignment Drift in Long-Term Human-LLM Interaction: A Mechanism-Oriented Framework
di: Yao, Xintong
Pubblicazione: (2026)
di: Yao, Xintong
Pubblicazione: (2026)
Oyster-I: Beyond Refusal -- Constructive Safety Alignment for Responsible Language Models
di: Duan, Ranjie, et al.
Pubblicazione: (2025)
di: Duan, Ranjie, et al.
Pubblicazione: (2025)
Culturally Grounded Personas in Large Language Models: Characterization and Alignment with Socio-Psychological Value Frameworks
di: Greco, Candida M., et al.
Pubblicazione: (2026)
di: Greco, Candida M., et al.
Pubblicazione: (2026)
Mind the Style: Impact of Communication Style on Human-Chatbot Interaction
di: Derner, Erik, et al.
Pubblicazione: (2026)
di: Derner, Erik, et al.
Pubblicazione: (2026)
Imperfectly Cooperative Human-AI Interactions: Comparing the Impacts of Human and AI Attributes in Simulated and User Studies
di: Cohen, Myke C., et al.
Pubblicazione: (2026)
di: Cohen, Myke C., et al.
Pubblicazione: (2026)
Large Language Models Show Human-like Social Desirability Biases in Survey Responses
di: Salecha, Aadesh, et al.
Pubblicazione: (2024)
di: Salecha, Aadesh, et al.
Pubblicazione: (2024)
GenAI Against Humanity: Nefarious Applications of Generative Artificial Intelligence and Large Language Models
di: Ferrara, Emilio
Pubblicazione: (2023)
di: Ferrara, Emilio
Pubblicazione: (2023)
MONAL: Model Autophagy Analysis for Modeling Human-AI Interactions
di: Yang, Shu, et al.
Pubblicazione: (2024)
di: Yang, Shu, et al.
Pubblicazione: (2024)
Open Models, Closed Minds? On Agents Capabilities in Mimicking Human Personalities through Open Large Language Models
di: La Cava, Lucio, et al.
Pubblicazione: (2024)
di: La Cava, Lucio, et al.
Pubblicazione: (2024)
Beyond One-Way Influence: Bidirectional Opinion Dynamics in Multi-Turn Human-LLM Interactions
di: Jiang, Yuyang, et al.
Pubblicazione: (2025)
di: Jiang, Yuyang, et al.
Pubblicazione: (2025)
DaKultur: Evaluating the Cultural Awareness of Language Models for Danish with Native Speakers
di: Müller-Eberstein, Max, et al.
Pubblicazione: (2025)
di: Müller-Eberstein, Max, et al.
Pubblicazione: (2025)
Conversational DNA: A New Visual Language for Understanding Dialogue Structure in Human and AI
di: Lin, Baihan
Pubblicazione: (2025)
di: Lin, Baihan
Pubblicazione: (2025)
A Metasemantic-Metapragmatic Framework for Taxonomizing Multimodal Communicative Alignment
di: Ji, Eugene Yu
Pubblicazione: (2025)
di: Ji, Eugene Yu
Pubblicazione: (2025)
Who Would Chatbots Vote For? Political Preferences of ChatGPT and Gemini in the 2024 European Union Elections
di: Haman, Michael, et al.
Pubblicazione: (2024)
di: Haman, Michael, et al.
Pubblicazione: (2024)
Generating the Modal Worker: A Cross-Model Audit of Race and Gender in LLM-Generated Personas Across 41 Occupations
di: van der Linden, Ilona, et al.
Pubblicazione: (2025)
di: van der Linden, Ilona, et al.
Pubblicazione: (2025)
More is More: Addition Bias in Large Language Models
di: Santagata, Luca, et al.
Pubblicazione: (2024)
di: Santagata, Luca, et al.
Pubblicazione: (2024)
Impacts of Anthropomorphizing Large Language Models in Learning Environments
di: Schaaff, Kristina, et al.
Pubblicazione: (2024)
di: Schaaff, Kristina, et al.
Pubblicazione: (2024)
From Prompts to Constructs: A Dual-Validity Framework for LLM Research in Psychology
di: Lin, Zhicheng
Pubblicazione: (2025)
di: Lin, Zhicheng
Pubblicazione: (2025)
Small but Significant: On the Promise of Small Language Models for Accessible AIED
di: Wei, Yumou, et al.
Pubblicazione: (2025)
di: Wei, Yumou, et al.
Pubblicazione: (2025)
Large Language Models as Psychological Simulators: A Methodological Guide
di: Lin, Zhicheng
Pubblicazione: (2025)
di: Lin, Zhicheng
Pubblicazione: (2025)
Evidence of conceptual mastery in the application of rules by Large Language Models
di: Nunes, José Luiz, et al.
Pubblicazione: (2025)
di: Nunes, José Luiz, et al.
Pubblicazione: (2025)
On the Reliability of Large Language Models to Misinformed and Demographically-Informed Prompts
di: Aremu, Toluwani, et al.
Pubblicazione: (2024)
di: Aremu, Toluwani, et al.
Pubblicazione: (2024)
STAR: SocioTechnical Approach to Red Teaming Language Models
di: Weidinger, Laura, et al.
Pubblicazione: (2024)
di: Weidinger, Laura, et al.
Pubblicazione: (2024)
Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence
di: McIntosh, Timothy R., et al.
Pubblicazione: (2024)
di: McIntosh, Timothy R., et al.
Pubblicazione: (2024)
Evaluating the Application of Large Language Models to Generate Feedback in Programming Education
di: Jacobs, Sven, et al.
Pubblicazione: (2024)
di: Jacobs, Sven, et al.
Pubblicazione: (2024)
Empirical evidence of Large Language Model's influence on human spoken communication
di: Yakura, Hiromu, et al.
Pubblicazione: (2024)
di: Yakura, Hiromu, et al.
Pubblicazione: (2024)
Chain of Alignment: Integrating Public Will with Expert Intelligence for Language Model Alignment
di: Konya, Andrew, et al.
Pubblicazione: (2024)
di: Konya, Andrew, et al.
Pubblicazione: (2024)
An Empirical Investigation of Gender Stereotype Representation in Large Language Models: The Italian Case
di: Giachino, Gioele, et al.
Pubblicazione: (2025)
di: Giachino, Gioele, et al.
Pubblicazione: (2025)
Gender Trouble in Language Models: An Empirical Audit Guided by Gender Performativity Theory
di: Hafner, Franziska Sofia, et al.
Pubblicazione: (2025)
di: Hafner, Franziska Sofia, et al.
Pubblicazione: (2025)
Decoding the Mind of Large Language Models: A Quantitative Evaluation of Ideology and Biases
di: Hirose, Manari, et al.
Pubblicazione: (2025)
di: Hirose, Manari, et al.
Pubblicazione: (2025)
KidLM: Advancing Language Models for Children -- Early Insights and Future Directions
di: Nayeem, Mir Tafseer, et al.
Pubblicazione: (2024)
di: Nayeem, Mir Tafseer, et al.
Pubblicazione: (2024)
Getting in the Door: Streamlining Intake in Civil Legal Services with Large Language Models
di: Steenhuis, Quinten, et al.
Pubblicazione: (2024)
di: Steenhuis, Quinten, et al.
Pubblicazione: (2024)
TUX: Measuring Human--AI Tacit Understanding
di: Li, Yueshen, et al.
Pubblicazione: (2026)
di: Li, Yueshen, et al.
Pubblicazione: (2026)
Creativity Support in the Age of Large Language Models: An Empirical Study Involving Emerging Writers
di: Chakrabarty, Tuhin, et al.
Pubblicazione: (2023)
di: Chakrabarty, Tuhin, et al.
Pubblicazione: (2023)
Conversational AI Powered by Large Language Models Amplifies False Memories in Witness Interviews
di: Chan, Samantha, et al.
Pubblicazione: (2024)
di: Chan, Samantha, et al.
Pubblicazione: (2024)
Can AI writing be salvaged? Mitigating Idiosyncrasies and Improving Human-AI Alignment in the Writing Process through Edits
di: Chakrabarty, Tuhin, et al.
Pubblicazione: (2024)
di: Chakrabarty, Tuhin, et al.
Pubblicazione: (2024)
Human-Centred LLM Privacy Audits: Findings and Frictions
di: Staufer, Dimitri, et al.
Pubblicazione: (2026)
di: Staufer, Dimitri, et al.
Pubblicazione: (2026)
Human Decision-making is Susceptible to AI-driven Manipulation
di: Sabour, Sahand, et al.
Pubblicazione: (2025)
di: Sabour, Sahand, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Addressing Climate Action Misperceptions with Generative AI
di: Remshard, Miriam, et al.
Pubblicazione: (2026) -
Generative Language Models Exhibit Social Identity Biases
di: Hu, Tiancheng, et al.
Pubblicazione: (2023) -
Alignment Drift in Long-Term Human-LLM Interaction: A Mechanism-Oriented Framework
di: Yao, Xintong
Pubblicazione: (2026) -
Oyster-I: Beyond Refusal -- Constructive Safety Alignment for Responsible Language Models
di: Duan, Ranjie, et al.
Pubblicazione: (2025) -
Culturally Grounded Personas in Large Language Models: Characterization and Alignment with Socio-Psychological Value Frameworks
di: Greco, Candida M., et al.
Pubblicazione: (2026)