Self-Improving AI Agents through Self-Play
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | Chojecki, Przemyslaw |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Psychometric Tests for AI Agents and Their Moduli Space
par: Chojecki, Przemyslaw
Publié: (2025)
par: Chojecki, Przemyslaw
Publié: (2025)
Mathematics and Coding are Universal AI Benchmarks
par: Chojecki, Przemyslaw
Publié: (2025)
par: Chojecki, Przemyslaw
Publié: (2025)
The Geometry of Benchmarks: A New Path Toward AGI
par: Chojecki, Przemyslaw
Publié: (2025)
par: Chojecki, Przemyslaw
Publié: (2025)
An Operational Kardashev-Style Scale for Autonomous AI - Towards AGI and Superintelligence
par: Chojecki, Przemyslaw
Publié: (2025)
par: Chojecki, Przemyslaw
Publié: (2025)
Learning Robust Reasoning through Guided Adversarial Self-Play
par: Li, Shuozhe, et autres
Publié: (2026)
par: Li, Shuozhe, et autres
Publié: (2026)
On The Statistical Limits of Self-Improving Agents
par: Wang, Charles L., et autres
Publié: (2025)
par: Wang, Charles L., et autres
Publié: (2025)
VideoAgent: Self-Improving Video Generation
par: Soni, Achint, et autres
Publié: (2024)
par: Soni, Achint, et autres
Publié: (2024)
Toward Training Superintelligent Software Agents through Self-Play SWE-RL
par: Wei, Yuxiang, et autres
Publié: (2025)
par: Wei, Yuxiang, et autres
Publié: (2025)
Superhuman AI for Stratego Using Self-Play Reinforcement Learning and Test-Time Search
par: Sokota, Samuel, et autres
Publié: (2025)
par: Sokota, Samuel, et autres
Publié: (2025)
Experiential Reflective Learning for Self-Improving LLM Agents
par: Allard, Marc-Antoine, et autres
Publié: (2026)
par: Allard, Marc-Antoine, et autres
Publié: (2026)
Continual Harness: Online Adaptation for Self-Improving Foundation Agents
par: Karten, Seth, et autres
Publié: (2026)
par: Karten, Seth, et autres
Publié: (2026)
TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment
par: Tan, Zhewen, et autres
Publié: (2026)
par: Tan, Zhewen, et autres
Publié: (2026)
Robust Autonomy Emerges from Self-Play
par: Cusumano-Towner, Marco, et autres
Publié: (2025)
par: Cusumano-Towner, Marco, et autres
Publié: (2025)
RSPO: Regularized Self-Play Alignment of Large Language Models
par: Tang, Xiaohang, et autres
Publié: (2025)
par: Tang, Xiaohang, et autres
Publié: (2025)
Memory Self-Regeneration: Uncovering Hidden Knowledge in Unlearned Models
par: Polowczyk, Agnieszka, et autres
Publié: (2025)
par: Polowczyk, Agnieszka, et autres
Publié: (2025)
Self-Improving LLM Agents at Test-Time
par: Acikgoz, Emre Can, et autres
Publié: (2025)
par: Acikgoz, Emre Can, et autres
Publié: (2025)
Self-Harmony: Learning to Harmonize Self-Supervision and Self-Play in Test-Time Reinforcement Learning
par: Wang, Ru, et autres
Publié: (2025)
par: Wang, Ru, et autres
Publié: (2025)
Enabling Self-Improving Agents to Learn at Test Time With Human-In-The-Loop Guidance
par: He, Yufei, et autres
Publié: (2025)
par: He, Yufei, et autres
Publié: (2025)
Self-Play Reinforcement Learning under Imperfect Information in Big 2
par: Patwa, Aalok
Publié: (2026)
par: Patwa, Aalok
Publié: (2026)
Soft Self-Consistency Improves Language Model Agents
par: Wang, Han, et autres
Publié: (2024)
par: Wang, Han, et autres
Publié: (2024)
Model Science: getting serious about verification, explanation and control of AI systems
par: Biecek, Przemyslaw, et autres
Publié: (2025)
par: Biecek, Przemyslaw, et autres
Publié: (2025)
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models
par: Cheng, Jiale, et autres
Publié: (2024)
par: Cheng, Jiale, et autres
Publié: (2024)
Self-Play Preference Optimization for Language Model Alignment
par: Wu, Yue, et autres
Publié: (2024)
par: Wu, Yue, et autres
Publié: (2024)
Foundation Model Self-Play: Open-Ended Strategy Innovation via Foundation Models
par: Dharna, Aaron, et autres
Publié: (2025)
par: Dharna, Aaron, et autres
Publié: (2025)
WIST: Web-Grounded Iterative Self-Play Tree for Domain-Targeted Reasoning Improvement
par: Li, Fangyuan, et autres
Publié: (2026)
par: Li, Fangyuan, et autres
Publié: (2026)
Self-Improving Robust Preference Optimization
par: Choi, Eugene, et autres
Publié: (2024)
par: Choi, Eugene, et autres
Publié: (2024)
CoSPlay: Cooperative Self-Play at Test-Time with Self-Generated Code and Unit Test
par: Hu, Zhangyi, et autres
Publié: (2026)
par: Hu, Zhangyi, et autres
Publié: (2026)
Training Agents to Self-Report Misbehavior
par: Lee, Bruce W., et autres
Publié: (2026)
par: Lee, Bruce W., et autres
Publié: (2026)
AlphaVerus: Bootstrapping Formally Verified Code Generation through Self-Improving Translation and Treefinement
par: Aggarwal, Pranjal, et autres
Publié: (2024)
par: Aggarwal, Pranjal, et autres
Publié: (2024)
Improving LLM Code Reasoning via Semantic Equivalence Self-Play with Formal Verification
par: Barone, Antonio Valerio Miceli, et autres
Publié: (2026)
par: Barone, Antonio Valerio Miceli, et autres
Publié: (2026)
Combee: Scaling Prompt Learning for Self-Improving Language Model Agents
par: Li, Hanchen, et autres
Publié: (2026)
par: Li, Hanchen, et autres
Publié: (2026)
Large Language Models Can Self-Improve At Web Agent Tasks
par: Patel, Ajay, et autres
Publié: (2024)
par: Patel, Ajay, et autres
Publié: (2024)
Recursive Introspection: Teaching Language Model Agents How to Self-Improve
par: Qu, Yuxiao, et autres
Publié: (2024)
par: Qu, Yuxiao, et autres
Publié: (2024)
Self-Improving Diffusion Models with Synthetic Data
par: Alemohammad, Sina, et autres
Publié: (2024)
par: Alemohammad, Sina, et autres
Publié: (2024)
SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning
par: Liu, Bo, et autres
Publié: (2025)
par: Liu, Bo, et autres
Publié: (2025)
Heterogeneous Self-Play for Realistic Highway Traffic Simulation
par: Qiu, Jinkai, et autres
Publié: (2026)
par: Qiu, Jinkai, et autres
Publié: (2026)
AgentOCR: Reimagining Agent History via Optical Self-Compression
par: Feng, Lang, et autres
Publié: (2026)
par: Feng, Lang, et autres
Publié: (2026)
Self-Improved Learning for Scalable Neural Combinatorial Optimization
par: Luo, Fu, et autres
Publié: (2024)
par: Luo, Fu, et autres
Publié: (2024)
Annealing Self-Distillation Rectification Improves Adversarial Training
par: Wu, Yu-Yu, et autres
Publié: (2023)
par: Wu, Yu-Yu, et autres
Publié: (2023)
Near-Optimal Reinforcement Learning with Self-Play under Adaptivity Constraints
par: Qiao, Dan, et autres
Publié: (2024)
par: Qiao, Dan, et autres
Publié: (2024)
Documents similaires
-
Psychometric Tests for AI Agents and Their Moduli Space
par: Chojecki, Przemyslaw
Publié: (2025) -
Mathematics and Coding are Universal AI Benchmarks
par: Chojecki, Przemyslaw
Publié: (2025) -
The Geometry of Benchmarks: A New Path Toward AGI
par: Chojecki, Przemyslaw
Publié: (2025) -
An Operational Kardashev-Style Scale for Autonomous AI - Towards AGI and Superintelligence
par: Chojecki, Przemyslaw
Publié: (2025) -
Learning Robust Reasoning through Guided Adversarial Self-Play
par: Li, Shuozhe, et autres
Publié: (2026)