CORE: Measuring Multi-Agent LLM Interaction Quality under Game-Theoretic Pressures
Fuente:
arXiv
Saved in:
| Main Authors: | Pandey, Punya Syon, Yang, Yongjin, Liu, Jiarui, Jin, Zhijing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BinaryPPO: Efficient Policy Optimization for Binary Classification
by: Pandey, Punya Syon, et al.
Published: (2026)
by: Pandey, Punya Syon, et al.
Published: (2026)
Accidental Vulnerability: Factors in Fine-Tuning that Shift Model Safeguards
by: Pandey, Punya Syon, et al.
Published: (2025)
by: Pandey, Punya Syon, et al.
Published: (2025)
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
by: Pandey, Punya Syon, et al.
Published: (2025)
by: Pandey, Punya Syon, et al.
Published: (2025)
Quriosity: Analyzing Human Questioning Behavior and Causal Inquiry through Curiosity-Driven Queries
by: Ceraolo, Roberto, et al.
Published: (2024)
by: Ceraolo, Roberto, et al.
Published: (2024)
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness
by: Yang, Yongjin, et al.
Published: (2025)
by: Yang, Yongjin, et al.
Published: (2025)
CLT-Forge: A Scalable Library for Cross-Layer Transcoders and Attribution Graphs
by: Draye, Florent, et al.
Published: (2026)
by: Draye, Florent, et al.
Published: (2026)
Causality for Natural Language Processing
by: Jin, Zhijing
Published: (2025)
by: Jin, Zhijing
Published: (2025)
Voices of Her: Analyzing Gender Differences in the AI Publication World
by: Ding, Yiwen, et al.
Published: (2023)
by: Ding, Yiwen, et al.
Published: (2023)
AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
by: Ma, Chang, et al.
Published: (2024)
by: Ma, Chang, et al.
Published: (2024)
Decomposing and Measuring Evaluation Awareness
by: Li, Changling, et al.
Published: (2026)
by: Li, Changling, et al.
Published: (2026)
Can Large Language Models Infer Causation from Correlation?
by: Jin, Zhijing, et al.
Published: (2023)
by: Jin, Zhijing, et al.
Published: (2023)
Analyzing the Role of Semantic Representations in the Era of Large Language Models
by: Jin, Zhijing, et al.
Published: (2024)
by: Jin, Zhijing, et al.
Published: (2024)
Evaluating Cooperation in LLM Social Groups through Elected Leadership
by: Faulkner, Ryan, et al.
Published: (2026)
by: Faulkner, Ryan, et al.
Published: (2026)
Objective Matters: Fine-Tuning Objectives Shape Safety, Robustness, and Persona Drift
by: Vennemeyer, Daniel, et al.
Published: (2026)
by: Vennemeyer, Daniel, et al.
Published: (2026)
CORE-KG: An LLM-Driven Knowledge Graph Construction Framework for Human Smuggling Networks
by: Meher, Dipak, et al.
Published: (2025)
by: Meher, Dipak, et al.
Published: (2025)
Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap
by: Qi, Xuan, et al.
Published: (2025)
by: Qi, Xuan, et al.
Published: (2025)
Benevolent Dictators? On LLM Agent Behavior in Dictator Games
by: Einwiller, Andreas, et al.
Published: (2025)
by: Einwiller, Andreas, et al.
Published: (2025)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Models
by: Choi, Younwoo, et al.
Published: (2025)
by: Choi, Younwoo, et al.
Published: (2025)
Improving Large Language Model Safety with Contrastive Representation Learning
by: Simko, Samuel, et al.
Published: (2025)
by: Simko, Samuel, et al.
Published: (2025)
SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning
by: Liu, Bo, et al.
Published: (2025)
by: Liu, Bo, et al.
Published: (2025)
Probing the Multi-turn Planning Capabilities of LLMs via 20 Question Games
by: Zhang, Yizhe, et al.
Published: (2023)
by: Zhang, Yizhe, et al.
Published: (2023)
CORE: Comprehensive Ontological Relation Evaluation for Large Language Models
by: Dwivedi, Satyam, et al.
Published: (2026)
by: Dwivedi, Satyam, et al.
Published: (2026)
ExpeL: LLM Agents Are Experiential Learners
by: Zhao, Andrew, et al.
Published: (2023)
by: Zhao, Andrew, et al.
Published: (2023)
RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
Simple Yet Effective: An Information-Theoretic Approach to Multi-LLM Uncertainty Quantification
by: Kruse, Maya, et al.
Published: (2025)
by: Kruse, Maya, et al.
Published: (2025)
Tracing Multilingual Representations in LLMs with Cross-Layer Transcoders
by: Harrasse, Abir, et al.
Published: (2025)
by: Harrasse, Abir, et al.
Published: (2025)
The Reasoning Trap: An Information-Theoretic Bound on Closed-System Multi-Step LLM Reasoning
by: Shin, Kwan Soo
Published: (2026)
by: Shin, Kwan Soo
Published: (2026)
SE-Bench: Benchmarking Self-Evolution with Knowledge Internalization
by: Yuan, Jiarui, et al.
Published: (2026)
by: Yuan, Jiarui, et al.
Published: (2026)
Corrupted by Reasoning: Reasoning Language Models Become Free-Riders in Public Goods Games
by: Piedrahita, David Guzman, et al.
Published: (2025)
by: Piedrahita, David Guzman, et al.
Published: (2025)
MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent
by: Yu, Hongli, et al.
Published: (2025)
by: Yu, Hongli, et al.
Published: (2025)
BAGEN: Are LLM Agents Budget-Aware?
by: Lin, Yuxiang, et al.
Published: (2026)
by: Lin, Yuxiang, et al.
Published: (2026)
Reinforce LLM Reasoning through Multi-Agent Reflection
by: Yuan, Yurun, et al.
Published: (2025)
by: Yuan, Yurun, et al.
Published: (2025)
Game-theoretic LLM: Agent Workflow for Negotiation Games
by: Hua, Wenyue, et al.
Published: (2024)
by: Hua, Wenyue, et al.
Published: (2024)
Stargazer: A Scalable Model-Fitting Benchmark Environment for AI Agents under Astrophysical Constraints
by: Liu, Xinge, et al.
Published: (2026)
by: Liu, Xinge, et al.
Published: (2026)
Calibrated Surprise: An Information-Theoretic Account of Creative Quality
by: Zou, Bo, et al.
Published: (2026)
by: Zou, Bo, et al.
Published: (2026)
From Words to Actions: Unveiling the Theoretical Underpinnings of LLM-Driven Autonomous Systems
by: He, Jianliang, et al.
Published: (2024)
by: He, Jianliang, et al.
Published: (2024)
KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation
by: Shi, Jiajun, et al.
Published: (2025)
by: Shi, Jiajun, et al.
Published: (2025)
ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
by: Lu, Jiarui, et al.
Published: (2024)
by: Lu, Jiarui, et al.
Published: (2024)
Measuring and Reducing LLM Hallucination without Gold-Standard Answers
by: Wei, Jiaheng, et al.
Published: (2024)
by: Wei, Jiaheng, et al.
Published: (2024)
Similar Items
-
BinaryPPO: Efficient Policy Optimization for Binary Classification
by: Pandey, Punya Syon, et al.
Published: (2026) -
Accidental Vulnerability: Factors in Fine-Tuning that Shift Model Safeguards
by: Pandey, Punya Syon, et al.
Published: (2025) -
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
by: Pandey, Punya Syon, et al.
Published: (2025) -
Quriosity: Analyzing Human Questioning Behavior and Causal Inquiry through Curiosity-Driven Queries
by: Ceraolo, Roberto, et al.
Published: (2024) -
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness
by: Yang, Yongjin, et al.
Published: (2025)