Search Self-play: Pushing the Frontier of Agent Capability without Supervision
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lu, Hongliang, Wen, Yuhang, Cheng, Pengyu, Ding, Ruijin, Guo, Jiaqi, Xu, Haotian, Wang, Chutian, Chen, Haonan, Jiang, Xiaoxi, Jiang, Guanjun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MARCH: Multi-Agent Reinforced Self-Check for LLM Hallucination
von: Li, Zhuo, et al.
Veröffentlicht: (2026)
von: Li, Zhuo, et al.
Veröffentlicht: (2026)
Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills
von: Ni, Jingwei, et al.
Veröffentlicht: (2026)
von: Ni, Jingwei, et al.
Veröffentlicht: (2026)
Pushing the Frontiers of Self-Distillation Prototypes Network with Dimension Regularization and Score Normalization
von: Chen, Yafeng, et al.
Veröffentlicht: (2025)
von: Chen, Yafeng, et al.
Veröffentlicht: (2025)
CLIPO: Contrastive Learning in Policy Optimization Generalizes RLVR
von: Cui, Sijia, et al.
Veröffentlicht: (2026)
von: Cui, Sijia, et al.
Veröffentlicht: (2026)
Bringing Stability to Diffusion: Decomposing and Reducing Variance of Training Masked Diffusion Models
von: Jia, Mengni, et al.
Veröffentlicht: (2025)
von: Jia, Mengni, et al.
Veröffentlicht: (2025)
ZeroSearch: Incentivize the Search Capability of LLMs without Searching
von: Sun, Hao, et al.
Veröffentlicht: (2025)
von: Sun, Hao, et al.
Veröffentlicht: (2025)
AgentFrontier: Expanding the Capability Frontier of LLM Agents with ZPD-Guided Data Synthesis
von: Chen, Xuanzhong, et al.
Veröffentlicht: (2025)
von: Chen, Xuanzhong, et al.
Veröffentlicht: (2025)
Eliminating Inductive Bias in Reward Models with Information-Theoretic Guidance
von: Li, Zhuo, et al.
Veröffentlicht: (2025)
von: Li, Zhuo, et al.
Veröffentlicht: (2025)
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization
von: Yao, Yihang, et al.
Veröffentlicht: (2026)
von: Yao, Yihang, et al.
Veröffentlicht: (2026)
S-Agents: Self-organizing Agents in Open-ended Environments
von: Chen, Jiaqi, et al.
Veröffentlicht: (2024)
von: Chen, Jiaqi, et al.
Veröffentlicht: (2024)
Pushing Frontiers for Proteoglycans
von: Marissa L. Maciej‐Hulme
Veröffentlicht: (2026)
von: Marissa L. Maciej‐Hulme
Veröffentlicht: (2026)
IndustryEQA: Pushing the Frontiers of Embodied Question Answering in Industrial Scenarios
von: Li, Yifan, et al.
Veröffentlicht: (2025)
von: Li, Yifan, et al.
Veröffentlicht: (2025)
How Can Haptic Feedback Assist People with Blind and Low Vision (BLV): A Systematic Literature Review
von: Jiang, Chutian, et al.
Veröffentlicht: (2024)
von: Jiang, Chutian, et al.
Veröffentlicht: (2024)
Iterative Data-Consistent Inversion with Multiple Push-forward Constraints
von: Jiang, Tianyi, et al.
Veröffentlicht: (2026)
von: Jiang, Tianyi, et al.
Veröffentlicht: (2026)
Self-playing Adversarial Language Game Enhances LLM Reasoning
von: Cheng, Pengyu, et al.
Veröffentlicht: (2024)
von: Cheng, Pengyu, et al.
Veröffentlicht: (2024)
Pushing Radar Odometry Beyond the Pavement: Current Capabilities and Challenges
von: Kolhe, Shaunak, et al.
Veröffentlicht: (2026)
von: Kolhe, Shaunak, et al.
Veröffentlicht: (2026)
Alignment Tipping Process: How Self-Evolution Pushes LLM Agents Off the Rails
von: Han, Siwei, et al.
Veröffentlicht: (2025)
von: Han, Siwei, et al.
Veröffentlicht: (2025)
Forecasting Frontier Language Model Agent Capabilities
von: Pimpale, Govind, et al.
Veröffentlicht: (2025)
von: Pimpale, Govind, et al.
Veröffentlicht: (2025)
Ola: Pushing the Frontiers of Omni-Modal Language Model
von: Liu, Zuyan, et al.
Veröffentlicht: (2025)
von: Liu, Zuyan, et al.
Veröffentlicht: (2025)
Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
von: Comanici, Gheorghe, et al.
Veröffentlicht: (2025)
von: Comanici, Gheorghe, et al.
Veröffentlicht: (2025)
Pushing the Frontier on Approximate EFX Allocations
von: Amanatidis, Georgios, et al.
Veröffentlicht: (2024)
von: Amanatidis, Georgios, et al.
Veröffentlicht: (2024)
Multi-Modality Spatio-Temporal Forecasting via Self-Supervised Learning
von: Deng, Jiewen, et al.
Veröffentlicht: (2024)
von: Deng, Jiewen, et al.
Veröffentlicht: (2024)
DOCTOR: Dynamic On-Chip Temporal Variation Remediation Toward Self-Corrected Photonic Tensor Accelerators
von: Lu, Haotian, et al.
Veröffentlicht: (2024)
von: Lu, Haotian, et al.
Veröffentlicht: (2024)
QuarkMedSearch: A Long-Horizon Deep Search Agent for Exploring Medical Intelligence
von: Lin, Zhichao, et al.
Veröffentlicht: (2026)
von: Lin, Zhichao, et al.
Veröffentlicht: (2026)
Almost sharp global wellposedness and scattering for the defocusing conformal wave equation on the hyperbolic space
von: Ma, Chutian
Veröffentlicht: (2023)
von: Ma, Chutian
Veröffentlicht: (2023)
Writing-Zero: Bridge the Gap Between Non-verifiable Tasks and Verifiable Rewards
von: Jia, Ruipeng, et al.
Veröffentlicht: (2025)
von: Jia, Ruipeng, et al.
Veröffentlicht: (2025)
Physical Neural Networks with Self-Learning Capabilities
von: Yu, Weichao, et al.
Veröffentlicht: (2024)
von: Yu, Weichao, et al.
Veröffentlicht: (2024)
SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL
von: Wang, Junke, et al.
Veröffentlicht: (2025)
von: Wang, Junke, et al.
Veröffentlicht: (2025)
Answer First, Reason Later: Aligning Search Relevance via Mode-Balanced Reinforcement Learning
von: Zhang, Shijie, et al.
Veröffentlicht: (2026)
von: Zhang, Shijie, et al.
Veröffentlicht: (2026)
Probing the Mid-level Vision Capabilities of Self-Supervised Learning
von: Chen, Xuweiyi, et al.
Veröffentlicht: (2024)
von: Chen, Xuweiyi, et al.
Veröffentlicht: (2024)
Advancing the Search Frontier with AI Agents
von: White, Ryen W.
Veröffentlicht: (2023)
von: White, Ryen W.
Veröffentlicht: (2023)
MiLorE-SSL: Scaling Multilingual Capabilities in Self-Supervised Models without Forgetting
von: Xu, Jing, et al.
Veröffentlicht: (2026)
von: Xu, Jing, et al.
Veröffentlicht: (2026)
Pushing the Frontiers of Non-equilibrium Dynamics of Collisionless and Weakly Collisional Self-gravitating Systems
von: Banik, Uddipan
Veröffentlicht: (2024)
von: Banik, Uddipan
Veröffentlicht: (2024)
Modularity in Argyres-Douglas Theories with $a=c$
von: Jiang, Hongliang
Veröffentlicht: (2024)
von: Jiang, Hongliang
Veröffentlicht: (2024)
Time-reversal invariant TQFTs from self-mirror symmetric SCFTs
von: Jiang, Hongliang
Veröffentlicht: (2024)
von: Jiang, Hongliang
Veröffentlicht: (2024)
D1-D5 CFT data from $AdS_3 \times S^3$ Virasoro-Shapiro amplitude
von: Jiang, Hongliang
Veröffentlicht: (2026)
von: Jiang, Hongliang
Veröffentlicht: (2026)
Macdonald Index from VOA and Graded Unitarity
von: Jiang, Hongliang
Veröffentlicht: (2026)
von: Jiang, Hongliang
Veröffentlicht: (2026)
Mem$^2$Evolve: Towards Self-Evolving Agents via Co-Evolutionary Capability Expansion and Experience Distillation
von: Cheng, Zihao, et al.
Veröffentlicht: (2026)
von: Cheng, Zihao, et al.
Veröffentlicht: (2026)
SCC/PVA Hydrogel with Self‐Healing Capability for Electrochromic Device
von: Weixing Song, et al.
Veröffentlicht: (2024)
von: Weixing Song, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MARCH: Multi-Agent Reinforced Self-Check for LLM Hallucination
von: Li, Zhuo, et al.
Veröffentlicht: (2026) -
Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills
von: Ni, Jingwei, et al.
Veröffentlicht: (2026) -
Pushing the Frontiers of Self-Distillation Prototypes Network with Dimension Regularization and Score Normalization
von: Chen, Yafeng, et al.
Veröffentlicht: (2025) -
CLIPO: Contrastive Learning in Policy Optimization Generalizes RLVR
von: Cui, Sijia, et al.
Veröffentlicht: (2026) -
Bringing Stability to Diffusion: Decomposing and Reducing Variance of Training Masked Diffusion Models
von: Jia, Mengni, et al.
Veröffentlicht: (2025)