Interpretable Risk Mitigation in LLM Agent Systems
Fuente:
arXiv
Saved in:
| Main Author: | Chojnacki, Jan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-Agent Strategic Games with LLMs
by: Chupilkin, Maxim
Published: (2026)
by: Chupilkin, Maxim
Published: (2026)
Using deep reinforcement learning to promote sustainable human behaviour on a common pool resource problem
by: Koster, Raphael, et al.
Published: (2024)
by: Koster, Raphael, et al.
Published: (2024)
Modeling the Economic Impacts of AI Openness Regulation
by: Qiu, Tori, et al.
Published: (2025)
by: Qiu, Tori, et al.
Published: (2025)
Designing DSIC Mechanisms for Data Sharing in the Era of Large Language Models
by: Ayyoubzadeh, Seyed Moein, et al.
Published: (2025)
by: Ayyoubzadeh, Seyed Moein, et al.
Published: (2025)
Designing Algorithmic Delegates: The Role of Indistinguishability in Human-AI Handoff
by: Greenwood, Sophie, et al.
Published: (2025)
by: Greenwood, Sophie, et al.
Published: (2025)
Towards Strategic Persuasion with Language Models
by: Cheng, Zirui, et al.
Published: (2025)
by: Cheng, Zirui, et al.
Published: (2025)
A Game-Theoretic Negotiation Framework for Cross-Cultural Consensus in LLMs
by: Zhang, Guoxi, et al.
Published: (2025)
by: Zhang, Guoxi, et al.
Published: (2025)
If It's Nice, Do It Twice: We Should Try Iterative Corpus Curation
by: Young, Robin
Published: (2025)
by: Young, Robin
Published: (2025)
Path Dependence under Adaptive AI Delegation
by: Huang, Lingxiao, et al.
Published: (2026)
by: Huang, Lingxiao, et al.
Published: (2026)
AI of the People, by the People, for the People: A Social Choice Approach to Collective Control of Artificial Intelligence
by: Bachmann, Paul Anton, et al.
Published: (2026)
by: Bachmann, Paul Anton, et al.
Published: (2026)
Competition and Diversity in Generative AI
by: Raghavan, Manish
Published: (2024)
by: Raghavan, Manish
Published: (2024)
AI Arms and Influence: Frontier Models Exhibit Sophisticated Reasoning in Simulated Nuclear Crises
by: Payne, Kenneth
Published: (2026)
by: Payne, Kenneth
Published: (2026)
CERN for AI: A Theoretical Framework for Autonomous Simulation-Based Artificial Intelligence Testing and Alignment
by: Bojic, Ljubisa, et al.
Published: (2023)
by: Bojic, Ljubisa, et al.
Published: (2023)
Assessing Group Fairness with Social Welfare Optimization
by: Chen, Violet, et al.
Published: (2024)
by: Chen, Violet, et al.
Published: (2024)
AgentSociety: Incentivizing Agentic Social Intelligence
by: Kesari, Aditya Vema Reddy, et al.
Published: (2026)
by: Kesari, Aditya Vema Reddy, et al.
Published: (2026)
Is Your LLM Overcharging You? Tokenization, Transparency, and Incentives
by: Velasco, Ander Artola, et al.
Published: (2025)
by: Velasco, Ander Artola, et al.
Published: (2025)
Integrated Design and Governance of Agentic AI Systems through Adaptive Information Modulation
by: Chen, Qiliang, et al.
Published: (2024)
by: Chen, Qiliang, et al.
Published: (2024)
LLM Strategic Reasoning: Agentic Study through Behavioral Game Theory
by: Jia, Jingru, et al.
Published: (2025)
by: Jia, Jingru, et al.
Published: (2025)
Do LLMs trust AI regulation? Emerging behaviour of game-theoretic LLM agents
by: Buscemi, Alessio, et al.
Published: (2025)
by: Buscemi, Alessio, et al.
Published: (2025)
Game Theory Meets LLM and Agentic AI: Reimagining Cybersecurity for the Age of Intelligent Threats
by: Zhu, Quanyan
Published: (2025)
by: Zhu, Quanyan
Published: (2025)
Two Tickets are Better than One: Fair and Accurate Hiring Under Strategic LLM Manipulations
by: Cohen, Lee, et al.
Published: (2025)
by: Cohen, Lee, et al.
Published: (2025)
From Outcome-Based to Language-Based Preferences
by: Capraro, Valerio, et al.
Published: (2022)
by: Capraro, Valerio, et al.
Published: (2022)
Nicer Than Humans: How do Large Language Models Behave in the Prisoner's Dilemma?
by: Fontana, Nicoló, et al.
Published: (2024)
by: Fontana, Nicoló, et al.
Published: (2024)
Alignment as Institutional Design: From Behavioral Correction to Transaction Structure in Intelligent Systems
by: Chai, Rui
Published: (2026)
by: Chai, Rui
Published: (2026)
TERMS-Bench: Diagnosing LLM Negotiation Agents Beyond Deal Rate
by: Zhang, Erica, et al.
Published: (2026)
by: Zhang, Erica, et al.
Published: (2026)
The Price of Paranoia: Robust Risk-Sensitive Cooperation in Non-Stationary Multi-Agent Reinforcement Learning
by: Ganguly, Deep Kumar, et al.
Published: (2026)
by: Ganguly, Deep Kumar, et al.
Published: (2026)
GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory
by: Cobben, Pepijn, et al.
Published: (2026)
by: Cobben, Pepijn, et al.
Published: (2026)
Institutional AI: Governing LLM Collusion in Multi-Agent Cournot Markets via Public Governance Graphs
by: Syrnikov, Marcantonio Bracale, et al.
Published: (2026)
by: Syrnikov, Marcantonio Bracale, et al.
Published: (2026)
ALYMPICS: LLM Agents Meet Game Theory -- Exploring Strategic Decision-Making with AI Agents
by: Mao, Shaoguang, et al.
Published: (2023)
by: Mao, Shaoguang, et al.
Published: (2023)
FinEvo: From Isolated Backtests to Ecological Market Games for Multi-Agent Financial Strategy Evolution
by: Zou, Mingxi, et al.
Published: (2026)
by: Zou, Mingxi, et al.
Published: (2026)
Cooperation Dynamics in Multi-Agent Systems: Exploring Game-Theoretic Scenarios with Mean-Field Equilibria
by: Sathi, Vaigarai, et al.
Published: (2023)
by: Sathi, Vaigarai, et al.
Published: (2023)
EconEvals: Benchmarks and Litmus Tests for Economic Decision-Making by LLM Agents
by: Fish, Sara, et al.
Published: (2025)
by: Fish, Sara, et al.
Published: (2025)
Everyone Contributes! Incentivizing Strategic Cooperation in Multi-LLM Systems via Sequential Public Goods Games
by: Liang, Yunhao, et al.
Published: (2025)
by: Liang, Yunhao, et al.
Published: (2025)
The Backfiring Effect of Weak AI Safety Regulation
by: Laufer, Benjamin, et al.
Published: (2025)
by: Laufer, Benjamin, et al.
Published: (2025)
The Fair Game: Auditing & Debiasing AI Algorithms Over Time
by: Basu, Debabrota, et al.
Published: (2025)
by: Basu, Debabrota, et al.
Published: (2025)
Can LLMs effectively provide game-theoretic-based scenarios for cybersecurity?
by: Proverbio, Daniele, et al.
Published: (2025)
by: Proverbio, Daniele, et al.
Published: (2025)
Towards Affordable Energy: A Gymnasium Environment for Electric Utility Demand-Response Programs
by: Escamilla, Jose E. Aguilar, et al.
Published: (2026)
by: Escamilla, Jose E. Aguilar, et al.
Published: (2026)
Agreement, Diversity, and Polarization Indices for Approval Elections
by: Faliszewski, Piotr, et al.
Published: (2026)
by: Faliszewski, Piotr, et al.
Published: (2026)
Learning Aggregation Rules in Participatory Budgeting: A Data-Driven Approach
by: Fairstein, Roy, et al.
Published: (2024)
by: Fairstein, Roy, et al.
Published: (2024)
Test-Time Compute Games
by: Velasco, Ander Artola, et al.
Published: (2026)
by: Velasco, Ander Artola, et al.
Published: (2026)
Similar Items
-
Multi-Agent Strategic Games with LLMs
by: Chupilkin, Maxim
Published: (2026) -
Using deep reinforcement learning to promote sustainable human behaviour on a common pool resource problem
by: Koster, Raphael, et al.
Published: (2024) -
Modeling the Economic Impacts of AI Openness Regulation
by: Qiu, Tori, et al.
Published: (2025) -
Designing DSIC Mechanisms for Data Sharing in the Era of Large Language Models
by: Ayyoubzadeh, Seyed Moein, et al.
Published: (2025) -
Designing Algorithmic Delegates: The Role of Indistinguishability in Human-AI Handoff
by: Greenwood, Sophie, et al.
Published: (2025)