Learning Robust Reasoning through Guided Adversarial Self-Play
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Shuozhe, Tadiparthi, Vaishnav, Lee, Kwonjoon, Agarwal, Nakul, Mahjoub, Hossein Nourkhiz, Pari, Ehsan Moradi, Chen, Lizhang, Zhang, Amy, Leqi, Liu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Language Grounded Multi-agent Reinforcement Learning with Human-interpretable Communication
by: Li, Huao, et al.
Published: (2024)
by: Li, Huao, et al.
Published: (2024)
Metacognitive Self-Correction for Multi-Agent System via Prototype-Guided Next-Execution Reconstruction
by: Shen, Xu, et al.
Published: (2025)
by: Shen, Xu, et al.
Published: (2025)
In Search of a Lost Metric: Human Empowerment as a Pillar of Socially Conscious Navigation
by: Baddam, Vasanth Reddy, et al.
Published: (2025)
by: Baddam, Vasanth Reddy, et al.
Published: (2025)
Understanding the Role of Hallucination in Reinforcement Post-Training of Multimodal Reasoning Models
by: Zhang, Gengwei, et al.
Published: (2026)
by: Zhang, Gengwei, et al.
Published: (2026)
SSR: A Generic Framework for Text-Aided Map Compression for Localization
by: Omama, Mohammad, et al.
Published: (2026)
by: Omama, Mohammad, et al.
Published: (2026)
SMART-Merge Planner: A Safe Merging and Real-Time Motion Planner for Autonomous Highway On-Ramp Merging
by: Mohammadnejad, Toktam, et al.
Published: (2025)
by: Mohammadnejad, Toktam, et al.
Published: (2025)
ExPO: Unlocking Hard Reasoning with Self-Explanation-Guided Reinforcement Learning
by: Zhou, Ruiyang, et al.
Published: (2025)
by: Zhou, Ruiyang, et al.
Published: (2025)
R3DM: Enabling Role Discovery and Diversity Through Dynamics Models in Multi-agent Reinforcement Learning
by: Goel, Harsh, et al.
Published: (2025)
by: Goel, Harsh, et al.
Published: (2025)
Dual Control for Interactive Autonomous Merging with Model Predictive Diffusion
by: Knaup, Jacob, et al.
Published: (2025)
by: Knaup, Jacob, et al.
Published: (2025)
Modeling the Lane-Change Reactions to Merging Vehicles for Highway On-Ramp Simulations
by: Holley, Dustin, et al.
Published: (2024)
by: Holley, Dustin, et al.
Published: (2024)
Speaking the Language of Teamwork: LLM-Guided Credit Assignment in Multi-Agent Reinforcement Learning
by: Lin, Muhan, et al.
Published: (2025)
by: Lin, Muhan, et al.
Published: (2025)
Symbolic Graph Inference for Compound Scene Understanding
by: Aryan, FNU, et al.
Published: (2024)
by: Aryan, FNU, et al.
Published: (2024)
Task-aware Distributed Source Coding under Dynamic Bandwidth
by: Li, Po-han, et al.
Published: (2023)
by: Li, Po-han, et al.
Published: (2023)
Filtered Reasoning Score: Evaluating Reasoning Quality on a Model's Most-Confident Traces
by: Pathak, Manas, et al.
Published: (2026)
by: Pathak, Manas, et al.
Published: (2026)
Navigating Noisy Feedback: Enhancing Reinforcement Learning with Error-Prone Language Models
by: Lin, Muhan, et al.
Published: (2024)
by: Lin, Muhan, et al.
Published: (2024)
Can't make an Omelette without Breaking some Eggs: Plausible Action Anticipation using Large Video-Language Models
by: Mittal, Himangi, et al.
Published: (2024)
by: Mittal, Himangi, et al.
Published: (2024)
A Learnable Wavelet Transformer for Long-Short Equity Trading and Risk-Adjusted Return Optimization
by: Li, Shuozhe, et al.
Published: (2026)
by: Li, Shuozhe, et al.
Published: (2026)
CARE-RFT: Confidence-Anchored Reinforcement Finetuning for Reliable Reasoning in Large Language Models
by: Li, Shuozhe, et al.
Published: (2026)
by: Li, Shuozhe, et al.
Published: (2026)
MR-LDM -- The Merge-Reactive Longitudinal Decision Model: Game Theoretic Human Decision Modeling for Interactive Sim Agents
by: Holley, Dustin, et al.
Published: (2025)
by: Holley, Dustin, et al.
Published: (2025)
Vamos: Versatile Action Models for Video Understanding
by: Wang, Shijie, et al.
Published: (2023)
by: Wang, Shijie, et al.
Published: (2023)
Overcoming Multi-step Complexity in Multimodal Theory-of-Mind Reasoning: A Scalable Bayesian Planner
by: Zhang, Chunhui, et al.
Published: (2025)
by: Zhang, Chunhui, et al.
Published: (2025)
M2D2M: Multi-Motion Generation from Text with Discrete Diffusion Models
by: Chi, Seunggeun, et al.
Published: (2024)
by: Chi, Seunggeun, et al.
Published: (2024)
Enhancing Language Agent Strategic Reasoning through Self-Play in Adversarial Games
by: Zhang, Yikai, et al.
Published: (2025)
by: Zhang, Yikai, et al.
Published: (2025)
AROW: V2X-based Automated Right-of-Way Algorithm for Cooperative Intersection Management
by: Shah, Ghayoor, et al.
Published: (2023)
by: Shah, Ghayoor, et al.
Published: (2023)
AntGPT: Can Large Language Models Help Long-term Action Anticipation from Videos?
by: Zhao, Qi, et al.
Published: (2023)
by: Zhao, Qi, et al.
Published: (2023)
$ϕ$-Balancing for Mixture-of-Experts Training
by: Chen, Lizhang, et al.
Published: (2026)
by: Chen, Lizhang, et al.
Published: (2026)
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning
by: Xu, Haoran, et al.
Published: (2025)
by: Xu, Haoran, et al.
Published: (2025)
Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning
by: Li, Shangzhe, et al.
Published: (2026)
by: Li, Shangzhe, et al.
Published: (2026)
Community Detection on Model Explanation Graphs for Explainable AI
by: Moradi, Ehsan
Published: (2025)
by: Moradi, Ehsan
Published: (2025)
Seirênes: Adversarial Self-Play with Evolving Distractions for LLM Reasoning
by: Zhang, Chi, et al.
Published: (2026)
by: Zhang, Chi, et al.
Published: (2026)
ATLS: Automated Trailer Loading for Surface Vessels
by: Abughaida, Amer, et al.
Published: (2024)
by: Abughaida, Amer, et al.
Published: (2024)
Momentum Guidance: Plug-and-Play Guidance for Flow Models
by: Liao, Runlong, et al.
Published: (2026)
by: Liao, Runlong, et al.
Published: (2026)
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
by: Chen, Jiaqi, et al.
Published: (2025)
by: Chen, Jiaqi, et al.
Published: (2025)
DiSCTT: Consensus-Guided Self-Curriculum for Efficient Test-Time Adaptation in Reasoning
by: Moradi, Mohammad Mahdi, et al.
Published: (2026)
by: Moradi, Mohammad Mahdi, et al.
Published: (2026)
Towards Robust 3D Pose Transfer with Adversarial Learning
by: Chen, Haoyu, et al.
Published: (2024)
by: Chen, Haoyu, et al.
Published: (2024)
Vibration Simulation of the cylindrical reservoir shell containing fluid vortex with the help of Vib-Shape software
by: Hossein Moradi
Published: (2015)
by: Hossein Moradi
Published: (2015)
ViCor: Bridging Visual Understanding and Commonsense Reasoning with Large Language Models
by: Zhou, Kaiwen, et al.
Published: (2023)
by: Zhou, Kaiwen, et al.
Published: (2023)
Adversarial Robustness of Nonparametric Regression
by: Moradi, Parsa, et al.
Published: (2025)
by: Moradi, Parsa, et al.
Published: (2025)
CGES: Confidence-Guided Early Stopping for Efficient and Accurate Self-Consistency
by: Aghazadeh, Ehsan, et al.
Published: (2025)
by: Aghazadeh, Ehsan, et al.
Published: (2025)
Efficient Self‐Guided One‐Pass Multi‐View Subspace Clustering
by: Haoyu Duan, et al.
Published: (2026)
by: Haoyu Duan, et al.
Published: (2026)
Similar Items
-
Language Grounded Multi-agent Reinforcement Learning with Human-interpretable Communication
by: Li, Huao, et al.
Published: (2024) -
Metacognitive Self-Correction for Multi-Agent System via Prototype-Guided Next-Execution Reconstruction
by: Shen, Xu, et al.
Published: (2025) -
In Search of a Lost Metric: Human Empowerment as a Pillar of Socially Conscious Navigation
by: Baddam, Vasanth Reddy, et al.
Published: (2025) -
Understanding the Role of Hallucination in Reinforcement Post-Training of Multimodal Reasoning Models
by: Zhang, Gengwei, et al.
Published: (2026) -
SSR: A Generic Framework for Text-Aided Map Compression for Localization
by: Omama, Mohammad, et al.
Published: (2026)