The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhou, Jiaxu, Huang, Jen-tse, Zhou, Xuhui, Lam, Man Ho, Wang, Xintao, Zhu, Hao, Wang, Wenxuan, Sap, Maarten |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SOTOPIA-S4: a user-friendly system for flexible, customizable, and large-scale social simulation
di: Zhou, Xuhui, et al.
Pubblicazione: (2025)
di: Zhou, Xuhui, et al.
Pubblicazione: (2025)
Minion: A Technology Probe to Explore How Users Negotiate Harmful Value Conflicts with AI Companions
di: Fan, Xianzhe, et al.
Pubblicazione: (2024)
di: Fan, Xianzhe, et al.
Pubblicazione: (2024)
On the Resilience of LLM-Based Multi-Agent Collaboration with Faulty Agents
di: Huang, Jen-tse, et al.
Pubblicazione: (2024)
di: Huang, Jen-tse, et al.
Pubblicazione: (2024)
When Should AI Read the Room? Public Perceptions of Social Intelligence in AI Agents
di: Mathur, Leena, et al.
Pubblicazione: (2026)
di: Mathur, Leena, et al.
Pubblicazione: (2026)
From Dogwhistles to Bullhorns: Unveiling Coded Rhetoric with Language Models
di: Mendelsohn, Julia, et al.
Pubblicazione: (2023)
di: Mendelsohn, Julia, et al.
Pubblicazione: (2023)
Imperfectly Cooperative Human-AI Interactions: Comparing the Impacts of Human and AI Attributes in Simulated and User Studies
di: Cohen, Myke C., et al.
Pubblicazione: (2026)
di: Cohen, Myke C., et al.
Pubblicazione: (2026)
Revisiting the Reliability of Psychological Scales on Large Language Models
di: Huang, Jen-tse, et al.
Pubblicazione: (2023)
di: Huang, Jen-tse, et al.
Pubblicazione: (2023)
Social Welfare Function Leaderboard: When LLM Agents Allocate Social Welfare
di: Shi, Zhengliang, et al.
Pubblicazione: (2025)
di: Shi, Zhengliang, et al.
Pubblicazione: (2025)
SoMi-ToM: Evaluating Multi-Perspective Theory of Mind in Embodied Social Interactions
di: Fan, Xianzhe, et al.
Pubblicazione: (2025)
di: Fan, Xianzhe, et al.
Pubblicazione: (2025)
BIG5-CHAT: Shaping LLM Personalities Through Training on Human-Grounded Data
di: Li, Wenkai, et al.
Pubblicazione: (2024)
di: Li, Wenkai, et al.
Pubblicazione: (2024)
Rethinking Theory of Mind Benchmarks for LLMs: Towards A User-Centered Perspective
di: Wang, Qiaosi, et al.
Pubblicazione: (2025)
di: Wang, Qiaosi, et al.
Pubblicazione: (2025)
Cognitive Chain-of-Thought (CoCoT): Structured Multimodal Reasoning about Social Situations
di: Park, Eunkyu, et al.
Pubblicazione: (2025)
di: Park, Eunkyu, et al.
Pubblicazione: (2025)
Data Defenses Against Large Language Models
di: Agnew, William, et al.
Pubblicazione: (2024)
di: Agnew, William, et al.
Pubblicazione: (2024)
SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior
di: Li, Jing-Jing, et al.
Pubblicazione: (2024)
di: Li, Jing-Jing, et al.
Pubblicazione: (2024)
Why (not) use AI? Analyzing People's Reasoning and Conditions for AI Acceptability
di: Mun, Jimin, et al.
Pubblicazione: (2025)
di: Mun, Jimin, et al.
Pubblicazione: (2025)
Translating With Feeling: Centering Translator Perspectives within Translation Technologies
di: Chechelnitsky, Daniel, et al.
Pubblicazione: (2026)
di: Chechelnitsky, Daniel, et al.
Pubblicazione: (2026)
Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs
di: Zhou, Xuhui, et al.
Pubblicazione: (2024)
di: Zhou, Xuhui, et al.
Pubblicazione: (2024)
GoodPoint: Learning Constructive Scientific Paper Feedback from Author Responses
di: Mun, Jimin, et al.
Pubblicazione: (2026)
di: Mun, Jimin, et al.
Pubblicazione: (2026)
Emotionally Numb or Empathetic? Evaluating How LLMs Feel Using EmotionBench
di: Huang, Jen-tse, et al.
Pubblicazione: (2023)
di: Huang, Jen-tse, et al.
Pubblicazione: (2023)
Prioritization First, Principles Second: An Adaptive Interpretation of Helpful, Honest, and Harmless Principles
di: Huang, Yue, et al.
Pubblicazione: (2025)
di: Huang, Yue, et al.
Pubblicazione: (2025)
Social Catalysts, Not Moral Agents: The Illusion of Alignment in LLM Societies
di: Hu, Yueqing, et al.
Pubblicazione: (2026)
di: Hu, Yueqing, et al.
Pubblicazione: (2026)
On the Failure of Latent State Persistence in Large Language Models
di: Huang, Jen-tse, et al.
Pubblicazione: (2025)
di: Huang, Jen-tse, et al.
Pubblicazione: (2025)
The Generative AI Ethics Playbook
di: Smith, Jessie J., et al.
Pubblicazione: (2024)
di: Smith, Jessie J., et al.
Pubblicazione: (2024)
MCTSr-Zero: Self-Reflective Psychological Counseling Dialogues Generation via Principles and Adaptive Exploration
di: Lu, Hao, et al.
Pubblicazione: (2025)
di: Lu, Hao, et al.
Pubblicazione: (2025)
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
di: Li, Jing-Jing, et al.
Pubblicazione: (2026)
di: Li, Jing-Jing, et al.
Pubblicazione: (2026)
Training Proactive and Personalized LLM Agents
di: Sun, Weiwei, et al.
Pubblicazione: (2025)
di: Sun, Weiwei, et al.
Pubblicazione: (2025)
AI-LieDar: Examine the Trade-off Between Utility and Truthfulness in LLM Agents
di: Su, Zhe, et al.
Pubblicazione: (2024)
di: Su, Zhe, et al.
Pubblicazione: (2024)
Emergence of Social Norms in Generative Agent Societies: Principles and Architecture
di: Ren, Siyue, et al.
Pubblicazione: (2024)
di: Ren, Siyue, et al.
Pubblicazione: (2024)
Ensuring User-side Fairness in Dynamic Recommender Systems
di: Yoo, Hyunsik, et al.
Pubblicazione: (2023)
di: Yoo, Hyunsik, et al.
Pubblicazione: (2023)
Ensuring Fairness with Transparent Auditing of Quantitative Bias in AI Systems
di: Yuan, Chih-Cheng Rex, et al.
Pubblicazione: (2024)
di: Yuan, Chih-Cheng Rex, et al.
Pubblicazione: (2024)
PolygloToxicityPrompts: Multilingual Evaluation of Neural Toxic Degeneration in Large Language Models
di: Jain, Devansh, et al.
Pubblicazione: (2024)
di: Jain, Devansh, et al.
Pubblicazione: (2024)
Who is ChatGPT? Benchmarking LLMs' Psychological Portrayal Using PsychoBench
di: Huang, Jen-tse, et al.
Pubblicazione: (2023)
di: Huang, Jen-tse, et al.
Pubblicazione: (2023)
Human or LLM as Standardized Patients? A Comparative Study for Medical Education
di: Zhang, Bingquan, et al.
Pubblicazione: (2025)
di: Zhang, Bingquan, et al.
Pubblicazione: (2025)
Navigating Towards Fairness with Data Selection
di: Zhang, Yixuan, et al.
Pubblicazione: (2024)
di: Zhang, Yixuan, et al.
Pubblicazione: (2024)
Law in Silico: Simulating Legal Society with LLM-Based Agents
di: Wang, Yiding, et al.
Pubblicazione: (2025)
di: Wang, Yiding, et al.
Pubblicazione: (2025)
Measuring Validity in LLM-based Resume Screening
di: Castleman, Jane, et al.
Pubblicazione: (2026)
di: Castleman, Jane, et al.
Pubblicazione: (2026)
Particip-AI: A Democratic Surveying Framework for Anticipating Future AI Use Cases, Harms and Benefits
di: Mun, Jimin, et al.
Pubblicazione: (2024)
di: Mun, Jimin, et al.
Pubblicazione: (2024)
Words Like Knives: Backstory-Personalized Modeling and Detection of Violent Communication
di: Shen, Jocelyn, et al.
Pubblicazione: (2025)
di: Shen, Jocelyn, et al.
Pubblicazione: (2025)
Does Socialization Emerge in AI Agent Society? A Case Study of Moltbook
di: Li, Ming, et al.
Pubblicazione: (2026)
di: Li, Ming, et al.
Pubblicazione: (2026)
Bridging the Regulatory Divide: Ensuring Safety and Equity in Wearable Health Technologies
di: Kelshiker, Akshay, et al.
Pubblicazione: (2025)
di: Kelshiker, Akshay, et al.
Pubblicazione: (2025)
Documenti analoghi
-
SOTOPIA-S4: a user-friendly system for flexible, customizable, and large-scale social simulation
di: Zhou, Xuhui, et al.
Pubblicazione: (2025) -
Minion: A Technology Probe to Explore How Users Negotiate Harmful Value Conflicts with AI Companions
di: Fan, Xianzhe, et al.
Pubblicazione: (2024) -
On the Resilience of LLM-Based Multi-Agent Collaboration with Faulty Agents
di: Huang, Jen-tse, et al.
Pubblicazione: (2024) -
When Should AI Read the Room? Public Perceptions of Social Intelligence in AI Agents
di: Mathur, Leena, et al.
Pubblicazione: (2026) -
From Dogwhistles to Bullhorns: Unveiling Coded Rhetoric with Language Models
di: Mendelsohn, Julia, et al.
Pubblicazione: (2023)