Evaluating Cooperation in LLM Social Groups through Elected Leadership
Fuente:
arXiv
Salvato in:
| Autori principali: | Faulkner, Ryan, Deshpande, Anushka, Piedrahita, David Guzman, Leibo, Joel Z., Jin, Zhijing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)
SocialJax: An Evaluation Suite for Multi-agent Reinforcement Learning in Sequential Social Dilemmas
di: Guo, Zihao, et al.
Pubblicazione: (2025)
di: Guo, Zihao, et al.
Pubblicazione: (2025)
Causality for Natural Language Processing
di: Jin, Zhijing
Pubblicazione: (2025)
di: Jin, Zhijing
Pubblicazione: (2025)
CORE: Measuring Multi-Agent LLM Interaction Quality under Game-Theoretic Pressures
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)
BinaryPPO: Efficient Policy Optimization for Binary Classification
di: Pandey, Punya Syon, et al.
Pubblicazione: (2026)
di: Pandey, Punya Syon, et al.
Pubblicazione: (2026)
When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas
di: Backmann, Steffen, et al.
Pubblicazione: (2025)
di: Backmann, Steffen, et al.
Pubblicazione: (2025)
Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap
di: Qi, Xuan, et al.
Pubblicazione: (2025)
di: Qi, Xuan, et al.
Pubblicazione: (2025)
Decomposing and Measuring Evaluation Awareness
di: Li, Changling, et al.
Pubblicazione: (2026)
di: Li, Changling, et al.
Pubblicazione: (2026)
Improving Large Language Model Safety with Contrastive Representation Learning
di: Simko, Samuel, et al.
Pubblicazione: (2025)
di: Simko, Samuel, et al.
Pubblicazione: (2025)
Accidental Vulnerability: Factors in Fine-Tuning that Shift Model Safeguards
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)
Quriosity: Analyzing Human Questioning Behavior and Causal Inquiry through Curiosity-Driven Queries
di: Ceraolo, Roberto, et al.
Pubblicazione: (2024)
di: Ceraolo, Roberto, et al.
Pubblicazione: (2024)
CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas
di: Tewolde, Emanuel, et al.
Pubblicazione: (2026)
di: Tewolde, Emanuel, et al.
Pubblicazione: (2026)
Corrupted by Reasoning: Reasoning Language Models Become Free-Riders in Public Goods Games
di: Piedrahita, David Guzman, et al.
Pubblicazione: (2025)
di: Piedrahita, David Guzman, et al.
Pubblicazione: (2025)
Aligned at the Start: Conceptual Groupings in LLM Embeddings
di: Khatir, Mehrdad, et al.
Pubblicazione: (2024)
di: Khatir, Mehrdad, et al.
Pubblicazione: (2024)
QualEval: Qualitative Evaluation for Model Improvement
di: Murahari, Vishvak, et al.
Pubblicazione: (2023)
di: Murahari, Vishvak, et al.
Pubblicazione: (2023)
Voices of Her: Analyzing Gender Differences in the AI Publication World
di: Ding, Yiwen, et al.
Pubblicazione: (2023)
di: Ding, Yiwen, et al.
Pubblicazione: (2023)
Code-Mix Sentiment Analysis on Hinglish Tweets
di: Garg, Aashi, et al.
Pubblicazione: (2026)
di: Garg, Aashi, et al.
Pubblicazione: (2026)
Visualizing Uncertainty in Translation Tasks: An Evaluation of LLM Performance and Confidence Metrics
di: Park, Jin Hyun, et al.
Pubblicazione: (2025)
di: Park, Jin Hyun, et al.
Pubblicazione: (2025)
Balanced Accuracy: The Right Metric for Evaluating LLM Judges -- Explained through Youden's J statistic
di: Collot, Stephane, et al.
Pubblicazione: (2025)
di: Collot, Stephane, et al.
Pubblicazione: (2025)
Peering Inside the Black Box: Uncovering LLM Errors in Optimization Modelling through Component-Level Evaluation
di: Refai, Dania, et al.
Pubblicazione: (2025)
di: Refai, Dania, et al.
Pubblicazione: (2025)
Can Large Language Models Infer Causation from Correlation?
di: Jin, Zhijing, et al.
Pubblicazione: (2023)
di: Jin, Zhijing, et al.
Pubblicazione: (2023)
Analyzing the Role of Semantic Representations in the Era of Large Language Models
di: Jin, Zhijing, et al.
Pubblicazione: (2024)
di: Jin, Zhijing, et al.
Pubblicazione: (2024)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
di: Liu, Yixin, et al.
Pubblicazione: (2025)
di: Liu, Yixin, et al.
Pubblicazione: (2025)
AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
di: Ma, Chang, et al.
Pubblicazione: (2024)
di: Ma, Chang, et al.
Pubblicazione: (2024)
Think Globally, Group Locally: Evaluating LLMs Using Multi-Lingual Word Grouping Games
di: Guerra-Solano, César, et al.
Pubblicazione: (2025)
di: Guerra-Solano, César, et al.
Pubblicazione: (2025)
PersonaGym: Evaluating Persona Agents and LLMs
di: Samuel, Vinay, et al.
Pubblicazione: (2024)
di: Samuel, Vinay, et al.
Pubblicazione: (2024)
Group Distributionally Robust Optimization-Driven Reinforcement Learning for LLM Reasoning
di: Panaganti, Kishan, et al.
Pubblicazione: (2026)
di: Panaganti, Kishan, et al.
Pubblicazione: (2026)
Group Reasoning Emission Estimation Networks
di: Guo, Yanming, et al.
Pubblicazione: (2025)
di: Guo, Yanming, et al.
Pubblicazione: (2025)
Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning
di: Zhang, Xichen, et al.
Pubblicazione: (2025)
di: Zhang, Xichen, et al.
Pubblicazione: (2025)
Enhancing LLM Evaluations: The Garbling Trick
di: Bradley, William F.
Pubblicazione: (2024)
di: Bradley, William F.
Pubblicazione: (2024)
Survey on Evaluation of LLM-based Agents
di: Yehudai, Asaf, et al.
Pubblicazione: (2025)
di: Yehudai, Asaf, et al.
Pubblicazione: (2025)
Enhancing NLP Robustness and Generalization through LLM-Generated Contrast Sets: A Scalable Framework for Systematic Evaluation and Adversarial Training
di: Lin, Hender
Pubblicazione: (2025)
di: Lin, Hender
Pubblicazione: (2025)
NICE: To Optimize In-Context Examples or Not?
di: Srivastava, Pragya, et al.
Pubblicazione: (2024)
di: Srivastava, Pragya, et al.
Pubblicazione: (2024)
Beyond Perplexity: A Geometric and Spectral Study of Low-Rank Pre-Training
di: Shivagunde, Namrata, et al.
Pubblicazione: (2026)
di: Shivagunde, Namrata, et al.
Pubblicazione: (2026)
Enhancing LLM Problem Solving with REAP: Reflection, Explicit Problem Deconstruction, and Advanced Prompting
di: Lingo, Ryan, et al.
Pubblicazione: (2024)
di: Lingo, Ryan, et al.
Pubblicazione: (2024)
Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations
di: Roytburg, Dani, et al.
Pubblicazione: (2026)
di: Roytburg, Dani, et al.
Pubblicazione: (2026)
LLM generation novelty through the lens of semantic similarity
di: Davydov, Philipp, et al.
Pubblicazione: (2025)
di: Davydov, Philipp, et al.
Pubblicazione: (2025)
Reinforce LLM Reasoning through Multi-Agent Reflection
di: Yuan, Yurun, et al.
Pubblicazione: (2025)
di: Yuan, Yurun, et al.
Pubblicazione: (2025)
KALAVAI: Predicting When Independent Specialist Fusion Works -- A Quantitative Model for Post-Hoc Cooperative LLM Training
di: Kumaresan, Ramchand
Pubblicazione: (2026)
di: Kumaresan, Ramchand
Pubblicazione: (2026)
Towards Multilingual LLM Evaluation for European Languages
di: Thellmann, Klaudia, et al.
Pubblicazione: (2024)
di: Thellmann, Klaudia, et al.
Pubblicazione: (2024)
Documenti analoghi
-
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025) -
SocialJax: An Evaluation Suite for Multi-agent Reinforcement Learning in Sequential Social Dilemmas
di: Guo, Zihao, et al.
Pubblicazione: (2025) -
Causality for Natural Language Processing
di: Jin, Zhijing
Pubblicazione: (2025) -
CORE: Measuring Multi-Agent LLM Interaction Quality under Game-Theoretic Pressures
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025) -
BinaryPPO: Efficient Policy Optimization for Binary Classification
di: Pandey, Punya Syon, et al.
Pubblicazione: (2026)