Absolute Zero: Reinforced Self-play Reasoning with Zero Data
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Andrew, Wu, Yiran, Yue, Yang, Wu, Tong, Xu, Quentin, Lin, Matthieu, Wang, Shenzhi, Wu, Qingyun, Zheng, Zilong, Huang, Gao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DiveR-CT: Diversity-enhanced Red Teaming Large Language Model Assistants with Relaxing Constraints
von: Zhao, Andrew, et al.
Veröffentlicht: (2024)
von: Zhao, Andrew, et al.
Veröffentlicht: (2024)
From Reasoning Chains to Verifiable Subproblems: Curriculum Reinforcement Learning Enables Credit Assignment for LLM Reasoning
von: Jiang, Xitai, et al.
Veröffentlicht: (2026)
von: Jiang, Xitai, et al.
Veröffentlicht: (2026)
Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning
von: Wu, Tong, et al.
Veröffentlicht: (2025)
von: Wu, Tong, et al.
Veröffentlicht: (2025)
StateFlow: Enhancing LLM Task-Solving through State-Driven Workflows
von: Wu, Yiran, et al.
Veröffentlicht: (2024)
von: Wu, Yiran, et al.
Veröffentlicht: (2024)
An Efficient Recipe for Long Context Extension via Middle-Focused Positional Encoding
von: Wu, Tong, et al.
Veröffentlicht: (2024)
von: Wu, Tong, et al.
Veröffentlicht: (2024)
R-Zero: Self-Evolving Reasoning LLM from Zero Data
von: Huang, Chengsong, et al.
Veröffentlicht: (2025)
von: Huang, Chengsong, et al.
Veröffentlicht: (2025)
G-Zero: Self-Play for Open-Ended Generation from Zero Data
von: Huang, Chengsong, et al.
Veröffentlicht: (2026)
von: Huang, Chengsong, et al.
Veröffentlicht: (2026)
Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
von: Wang, Shenzhi, et al.
Veröffentlicht: (2025)
von: Wang, Shenzhi, et al.
Veröffentlicht: (2025)
Scaffolded Language Models with Language Supervision for Mixed-Autonomy: A Survey
von: Lin, Matthieu, et al.
Veröffentlicht: (2024)
von: Lin, Matthieu, et al.
Veröffentlicht: (2024)
Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
von: Yue, Yang, et al.
Veröffentlicht: (2025)
von: Yue, Yang, et al.
Veröffentlicht: (2025)
Beyond English-Centric Training: How Reinforcement Learning Improves Cross-Lingual Reasoning in LLMs
von: Huang, Shulin, et al.
Veröffentlicht: (2025)
von: Huang, Shulin, et al.
Veröffentlicht: (2025)
SSR-Zero: Simple Self-Rewarding Reinforcement Learning for Machine Translation
von: Yang, Wenjie, et al.
Veröffentlicht: (2025)
von: Yang, Wenjie, et al.
Veröffentlicht: (2025)
RuleReasoner: Reinforced Rule-based Reasoning via Domain-aware Dynamic Sampling
von: Liu, Yang, et al.
Veröffentlicht: (2025)
von: Liu, Yang, et al.
Veröffentlicht: (2025)
Persona is a Double-edged Sword: Mitigating the Negative Impact of Role-playing Prompts in Zero-shot Reasoning Tasks
von: Kim, Junseok, et al.
Veröffentlicht: (2024)
von: Kim, Junseok, et al.
Veröffentlicht: (2024)
Explain Less, Understand More: Jargon Detection via Personalized Parameter-Efficient Fine-tuning
von: Wu, Bohao, et al.
Veröffentlicht: (2025)
von: Wu, Bohao, et al.
Veröffentlicht: (2025)
Attention-guided Self-reflection for Zero-shot Hallucination Detection in Large Language Models
von: Liu, Qiang, et al.
Veröffentlicht: (2025)
von: Liu, Qiang, et al.
Veröffentlicht: (2025)
The Flexibility Trap: Why Arbitrary Order Limits Reasoning Potential in Diffusion Language Models
von: Ni, Zanlin, et al.
Veröffentlicht: (2026)
von: Ni, Zanlin, et al.
Veröffentlicht: (2026)
ZeQR: Zero-shot Query Reformulation for Conversational Search
von: Yang, Dayu, et al.
Veröffentlicht: (2023)
von: Yang, Dayu, et al.
Veröffentlicht: (2023)
CEC-Zero: Zero-Supervision Character Error Correction with Self-Generated Rewards
von: Lin, Zhiming, et al.
Veröffentlicht: (2025)
von: Lin, Zhiming, et al.
Veröffentlicht: (2025)
ExpeL: LLM Agents Are Experiential Learners
von: Zhao, Andrew, et al.
Veröffentlicht: (2023)
von: Zhao, Andrew, et al.
Veröffentlicht: (2023)
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
von: Zeng, Yifan, et al.
Veröffentlicht: (2024)
von: Zeng, Yifan, et al.
Veröffentlicht: (2024)
The Flip Side of RLHF: On-Policy Feedback for Reward Model Self-Supervised Improvement
von: Wang, Xiaobo, et al.
Veröffentlicht: (2026)
von: Wang, Xiaobo, et al.
Veröffentlicht: (2026)
Socratic-Zero : Bootstrapping Reasoning via Data-Free Agent Co-evolution
von: Wang, Shaobo, et al.
Veröffentlicht: (2025)
von: Wang, Shaobo, et al.
Veröffentlicht: (2025)
Fast-DetectGPT: Efficient Zero-Shot Detection of Machine-Generated Text via Conditional Probability Curvature
von: Bao, Guangsheng, et al.
Veröffentlicht: (2023)
von: Bao, Guangsheng, et al.
Veröffentlicht: (2023)
Self-playing Adversarial Language Game Enhances LLM Reasoning
von: Cheng, Pengyu, et al.
Veröffentlicht: (2024)
von: Cheng, Pengyu, et al.
Veröffentlicht: (2024)
HopChain: Multi-Hop Data Synthesis for Generalizable Vision-Language Reasoning
von: Wang, Shenzhi, et al.
Veröffentlicht: (2026)
von: Wang, Shenzhi, et al.
Veröffentlicht: (2026)
GCOF: Self-iterative Text Generation for Copywriting Using Large Language Model
von: Zhou, Jianghui, et al.
Veröffentlicht: (2024)
von: Zhou, Jianghui, et al.
Veröffentlicht: (2024)
Label Set Optimization via Activation Distribution Kurtosis for Zero-shot Classification with Generative Models
von: Li, Yue, et al.
Veröffentlicht: (2024)
von: Li, Yue, et al.
Veröffentlicht: (2024)
GRLO: Towards Generalizable Reinforcement Learning in Open-Ended Environments from Zero
von: Yin, Shangjian, et al.
Veröffentlicht: (2026)
von: Yin, Shangjian, et al.
Veröffentlicht: (2026)
Trans-Zero: Self-Play Incentivizes Large Language Models for Multilingual Translation Without Parallel Data
von: Zou, Wei, et al.
Veröffentlicht: (2025)
von: Zou, Wei, et al.
Veröffentlicht: (2025)
HARGPT: Are LLMs Zero-Shot Human Activity Recognizers?
von: Ji, Sijie, et al.
Veröffentlicht: (2024)
von: Ji, Sijie, et al.
Veröffentlicht: (2024)
Reasoning-Oriented and Analogy-Based Methods for Locating and Editing in Zero-Shot Event-Relational Reasoning
von: Tang, Jingyao, et al.
Veröffentlicht: (2025)
von: Tang, Jingyao, et al.
Veröffentlicht: (2025)
Are You Being Tracked? Discover the Power of Zero-Shot Trajectory Tracing with LLMs!
von: Yang, Huanqi, et al.
Veröffentlicht: (2024)
von: Yang, Huanqi, et al.
Veröffentlicht: (2024)
Detecting RLVR Training Data via Structural Convergence of Reasoning
von: Zhang, Hongbo, et al.
Veröffentlicht: (2026)
von: Zhang, Hongbo, et al.
Veröffentlicht: (2026)
LLMs are not Zero-Shot Reasoners for Biomedical Information Extraction
von: Nagar, Aishik, et al.
Veröffentlicht: (2024)
von: Nagar, Aishik, et al.
Veröffentlicht: (2024)
Retrieval-Reasoning Large Language Model-based Synthetic Clinical Trial Generation
von: Xu, Zerui, et al.
Veröffentlicht: (2024)
von: Xu, Zerui, et al.
Veröffentlicht: (2024)
TokenSwift: Lossless Acceleration of Ultra Long Sequence Generation
von: Wu, Tong, et al.
Veröffentlicht: (2025)
von: Wu, Tong, et al.
Veröffentlicht: (2025)
MT-R1-Zero: Advancing LLM-based Machine Translation via R1-Zero-like Reinforcement Learning
von: Feng, Zhaopeng, et al.
Veröffentlicht: (2025)
von: Feng, Zhaopeng, et al.
Veröffentlicht: (2025)
Towards Zero-Shot Multimodal Machine Translation
von: Futeral, Matthieu, et al.
Veröffentlicht: (2024)
von: Futeral, Matthieu, et al.
Veröffentlicht: (2024)
Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning
von: Zhang, Shaokun, et al.
Veröffentlicht: (2025)
von: Zhang, Shaokun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DiveR-CT: Diversity-enhanced Red Teaming Large Language Model Assistants with Relaxing Constraints
von: Zhao, Andrew, et al.
Veröffentlicht: (2024) -
From Reasoning Chains to Verifiable Subproblems: Curriculum Reinforcement Learning Enables Credit Assignment for LLM Reasoning
von: Jiang, Xitai, et al.
Veröffentlicht: (2026) -
Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning
von: Wu, Tong, et al.
Veröffentlicht: (2025) -
StateFlow: Enhancing LLM Task-Solving through State-Driven Workflows
von: Wu, Yiran, et al.
Veröffentlicht: (2024) -
An Efficient Recipe for Long Context Extension via Middle-Focused Positional Encoding
von: Wu, Tong, et al.
Veröffentlicht: (2024)