Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)
Fuente:
arXiv
Salvato in:
| Autori principali: | Jiang, Liwei, Chai, Yuanjun, Li, Margaret, Liu, Mickel, Fok, Raymond, Dziri, Nouha, Tsvetkov, Yulia, Sap, Maarten, Albalak, Alon, Choi, Yejin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
di: Jiang, Liwei, et al.
Pubblicazione: (2024)
di: Jiang, Liwei, et al.
Pubblicazione: (2024)
Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
di: Mireshghallah, Niloofar, et al.
Pubblicazione: (2023)
di: Mireshghallah, Niloofar, et al.
Pubblicazione: (2023)
Multi-Attribute Constraint Satisfaction via Language Model Rewriting
di: Baheti, Ashutosh, et al.
Pubblicazione: (2024)
di: Baheti, Ashutosh, et al.
Pubblicazione: (2024)
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
di: Han, Seungju, et al.
Pubblicazione: (2024)
di: Han, Seungju, et al.
Pubblicazione: (2024)
SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior
di: Li, Jing-Jing, et al.
Pubblicazione: (2024)
di: Li, Jing-Jing, et al.
Pubblicazione: (2024)
OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
di: Vijayvargiya, Sanidhya, et al.
Pubblicazione: (2025)
di: Vijayvargiya, Sanidhya, et al.
Pubblicazione: (2025)
Rel-A.I.: An Interaction-Centered Approach To Measuring Human-LM Reliance
di: Zhou, Kaitlyn, et al.
Pubblicazione: (2024)
di: Zhou, Kaitlyn, et al.
Pubblicazione: (2024)
What Makes it Ok to Set a Fire? Iterative Self-distillation of Contexts and Rationales for Disambiguating Defeasible Social and Moral Situations
di: Rao, Kavel, et al.
Pubblicazione: (2023)
di: Rao, Kavel, et al.
Pubblicazione: (2023)
Value Kaleidoscope: Engaging AI with Pluralistic Human Values, Rights, and Duties
di: Sorensen, Taylor, et al.
Pubblicazione: (2023)
di: Sorensen, Taylor, et al.
Pubblicazione: (2023)
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
di: Sclar, Melanie, et al.
Pubblicazione: (2023)
di: Sclar, Melanie, et al.
Pubblicazione: (2023)
From Dogwhistles to Bullhorns: Unveiling Coded Rhetoric with Language Models
di: Mendelsohn, Julia, et al.
Pubblicazione: (2023)
di: Mendelsohn, Julia, et al.
Pubblicazione: (2023)
CULTURE-GEN: Revealing Global Cultural Perception in Language Models through Natural Language Prompting
di: Li, Huihan, et al.
Pubblicazione: (2024)
di: Li, Huihan, et al.
Pubblicazione: (2024)
Phenomenal Yet Puzzling: Testing Inductive Reasoning Capabilities of Language Models with Hypothesis Refinement
di: Qiu, Linlu, et al.
Pubblicazione: (2023)
di: Qiu, Linlu, et al.
Pubblicazione: (2023)
AI as Humanity's Salieri: Quantifying Linguistic Creativity of Language Models via Systematic Attribution of Machine Text against Web Text
di: Lu, Ximing, et al.
Pubblicazione: (2024)
di: Lu, Ximing, et al.
Pubblicazione: (2024)
The Art of Saying No: Contextual Noncompliance in Language Models
di: Brahman, Faeze, et al.
Pubblicazione: (2024)
di: Brahman, Faeze, et al.
Pubblicazione: (2024)
Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models
di: Liu, Mickel, et al.
Pubblicazione: (2025)
di: Liu, Mickel, et al.
Pubblicazione: (2025)
Tuning Language Models by Proxy
di: Liu, Alisa, et al.
Pubblicazione: (2024)
di: Liu, Alisa, et al.
Pubblicazione: (2024)
Particip-AI: A Democratic Surveying Framework for Anticipating Future AI Use Cases, Harms and Benefits
di: Mun, Jimin, et al.
Pubblicazione: (2024)
di: Mun, Jimin, et al.
Pubblicazione: (2024)
TurnWise: The Gap between Single- and Multi-turn Language Model Capabilities
di: Graf, Victoria, et al.
Pubblicazione: (2026)
di: Graf, Victoria, et al.
Pubblicazione: (2026)
Surfacing Semantic Orthogonality Across Model Safety Benchmarks: A Multi-Dimensional Analysis
di: Bennion, Jonathan, et al.
Pubblicazione: (2025)
di: Bennion, Jonathan, et al.
Pubblicazione: (2025)
Can Language Models Reason about Individualistic Human Values and Preferences?
di: Jiang, Liwei, et al.
Pubblicazione: (2024)
di: Jiang, Liwei, et al.
Pubblicazione: (2024)
A Roadmap to Pluralistic Alignment
di: Sorensen, Taylor, et al.
Pubblicazione: (2024)
di: Sorensen, Taylor, et al.
Pubblicazione: (2024)
Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations
di: Zheng, Mingqian, et al.
Pubblicazione: (2026)
di: Zheng, Mingqian, et al.
Pubblicazione: (2026)
DailyDilemmas: Revealing Value Preferences of LLMs with Quandaries of Daily Life
di: Chiu, Yu Ying, et al.
Pubblicazione: (2024)
di: Chiu, Yu Ying, et al.
Pubblicazione: (2024)
PolyGuard: A Multilingual Safety Moderation Tool for 17 Languages
di: Kumar, Priyanshu, et al.
Pubblicazione: (2025)
di: Kumar, Priyanshu, et al.
Pubblicazione: (2025)
SoK: Privacy-Enhancing Technologies in Artificial Intelligence
di: Oualha, Nouha
Pubblicazione: (2025)
di: Oualha, Nouha
Pubblicazione: (2025)
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
di: Lin, Bill Yuchen, et al.
Pubblicazione: (2024)
di: Lin, Bill Yuchen, et al.
Pubblicazione: (2024)
Finding Flawed Fictions: Evaluating Complex Reasoning in Language Models via Plot Hole Detection
di: Ahuja, Kabir, et al.
Pubblicazione: (2025)
di: Ahuja, Kabir, et al.
Pubblicazione: (2025)
A False Sense of Privacy: Evaluating Textual Data Sanitization Beyond Surface-level Privacy Leakage
di: Xin, Rui, et al.
Pubblicazione: (2025)
di: Xin, Rui, et al.
Pubblicazione: (2025)
CulturalTeaming: AI-Assisted Interactive Red-Teaming for Challenging LLMs' (Lack of) Multicultural Knowledge
di: Chiu, Yu Ying, et al.
Pubblicazione: (2024)
di: Chiu, Yu Ying, et al.
Pubblicazione: (2024)
Political Neutrality in AI Is Impossible- But Here Is How to Approximate It
di: Fisher, Jillian, et al.
Pubblicazione: (2025)
di: Fisher, Jillian, et al.
Pubblicazione: (2025)
Quality Control in Open-Ended Crowdsourcing: A Survey
di: Chai, Lei, et al.
Pubblicazione: (2024)
di: Chai, Lei, et al.
Pubblicazione: (2024)
Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning
di: Sclar, Melanie, et al.
Pubblicazione: (2024)
di: Sclar, Melanie, et al.
Pubblicazione: (2024)
Modular Pluralism: Pluralistic Alignment via Multi-LLM Collaboration
di: Feng, Shangbin, et al.
Pubblicazione: (2024)
di: Feng, Shangbin, et al.
Pubblicazione: (2024)
Open-Endedness is Essential for Artificial Superhuman Intelligence
di: Hughes, Edward, et al.
Pubblicazione: (2024)
di: Hughes, Edward, et al.
Pubblicazione: (2024)
RewardBench: Evaluating Reward Models for Language Modeling
di: Lambert, Nathan, et al.
Pubblicazione: (2024)
di: Lambert, Nathan, et al.
Pubblicazione: (2024)
In Search of Verifiability: Explanations Rarely Enable Complementary Performance in AI-Advised Decision Making
di: Fok, Raymond, et al.
Pubblicazione: (2023)
di: Fok, Raymond, et al.
Pubblicazione: (2023)
Augmenting Expert Cognition in the Age of Generative AI: Insights from Document-Centric Knowledge Work
di: Siu, Alexa, et al.
Pubblicazione: (2025)
di: Siu, Alexa, et al.
Pubblicazione: (2025)
In search of verifiability: Explanations rarely enable complementary performance in AI‐advised decision making
di: Raymond Fok, et al.
Pubblicazione: (2024)
di: Raymond Fok, et al.
Pubblicazione: (2024)
MAGPIE: A dataset for Multi-AGent contextual PrIvacy Evaluation
di: Juneja, Gurusha, et al.
Pubblicazione: (2025)
di: Juneja, Gurusha, et al.
Pubblicazione: (2025)
Documenti analoghi
-
WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
di: Jiang, Liwei, et al.
Pubblicazione: (2024) -
Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
di: Mireshghallah, Niloofar, et al.
Pubblicazione: (2023) -
Multi-Attribute Constraint Satisfaction via Language Model Rewriting
di: Baheti, Ashutosh, et al.
Pubblicazione: (2024) -
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
di: Han, Seungju, et al.
Pubblicazione: (2024) -
SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior
di: Li, Jing-Jing, et al.
Pubblicazione: (2024)