Better LLM Reasoning via Dual-Play
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Zhengxin, Huang, Chengyu, Li, Aochong Oliver, Cardie, Claire |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HAPO: Training Language Models to Reason Concisely via History-Aware Policy Optimization
von: Huang, Chengyu, et al.
Veröffentlicht: (2025)
von: Huang, Chengyu, et al.
Veröffentlicht: (2025)
Memorization vs. Reasoning: Updating LLMs with New Knowledge
von: Li, Aochong Oliver, et al.
Veröffentlicht: (2025)
von: Li, Aochong Oliver, et al.
Veröffentlicht: (2025)
Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text
von: Huang, Chengyu, et al.
Veröffentlicht: (2026)
von: Huang, Chengyu, et al.
Veröffentlicht: (2026)
How Far Are We From True Auto-Research?
von: Zhang, Zhengxin, et al.
Veröffentlicht: (2026)
von: Zhang, Zhengxin, et al.
Veröffentlicht: (2026)
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
von: Chen, Jiaqi, et al.
Veröffentlicht: (2025)
von: Chen, Jiaqi, et al.
Veröffentlicht: (2025)
Reasoning Court: Combining Reasoning, Action, and Judgment for Multi-Hop Reasoning
von: Wu, Jingtian, et al.
Veröffentlicht: (2025)
von: Wu, Jingtian, et al.
Veröffentlicht: (2025)
Improving LLM Code Reasoning via Semantic Equivalence Self-Play with Formal Verification
von: Barone, Antonio Valerio Miceli, et al.
Veröffentlicht: (2026)
von: Barone, Antonio Valerio Miceli, et al.
Veröffentlicht: (2026)
RelayLLM: Efficient Reasoning via Collaborative Decoding
von: Huang, Chengsong, et al.
Veröffentlicht: (2026)
von: Huang, Chengsong, et al.
Veröffentlicht: (2026)
Green Prompting: Characterizing Prompt-driven Energy Costs of LLM Inference
von: Adamska, Marta, et al.
Veröffentlicht: (2025)
von: Adamska, Marta, et al.
Veröffentlicht: (2025)
Policy-Gradient Training of Language Models for Ranking
von: Gao, Ge, et al.
Veröffentlicht: (2023)
von: Gao, Ge, et al.
Veröffentlicht: (2023)
Understanding Generalization in Role-Playing Models via Information Theory
von: Li, Yongqi, et al.
Veröffentlicht: (2025)
von: Li, Yongqi, et al.
Veröffentlicht: (2025)
BPO: Staying Close to the Behavior LLM Creates Better Online LLM Alignment
von: Xu, Wenda, et al.
Veröffentlicht: (2024)
von: Xu, Wenda, et al.
Veröffentlicht: (2024)
DCRM: A Heuristic to Measure Response Pair Quality in Preference Optimization
von: Huang, Chengyu, et al.
Veröffentlicht: (2025)
von: Huang, Chengyu, et al.
Veröffentlicht: (2025)
PLAY2PROMPT: Zero-shot Tool Instruction Optimization for LLM Agents via Tool Play
von: Fang, Wei, et al.
Veröffentlicht: (2025)
von: Fang, Wei, et al.
Veröffentlicht: (2025)
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment
von: Zhang, Jiazheng, et al.
Veröffentlicht: (2025)
von: Zhang, Jiazheng, et al.
Veröffentlicht: (2025)
R-Zero: Self-Evolving Reasoning LLM from Zero Data
von: Huang, Chengsong, et al.
Veröffentlicht: (2025)
von: Huang, Chengsong, et al.
Veröffentlicht: (2025)
Universal Reasoner: A Single, Composable Plug-and-Play Reasoner for Frozen LLMs
von: Kim, Jaemin, et al.
Veröffentlicht: (2025)
von: Kim, Jaemin, et al.
Veröffentlicht: (2025)
Internalizing LLM Reasoning via Discovery and Replay of Latent Actions
von: Shi, Zhenning, et al.
Veröffentlicht: (2026)
von: Shi, Zhenning, et al.
Veröffentlicht: (2026)
RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?
von: Xu, Haotian, et al.
Veröffentlicht: (2025)
von: Xu, Haotian, et al.
Veröffentlicht: (2025)
Beyond Markovian: Reflective Exploration via Bayes-Adaptive RL for LLM Reasoning
von: Zhang, Shenao, et al.
Veröffentlicht: (2025)
von: Zhang, Shenao, et al.
Veröffentlicht: (2025)
ReasonFlux: Hierarchical LLM Reasoning via Scaling Thought Templates
von: Yang, Ling, et al.
Veröffentlicht: (2025)
von: Yang, Ling, et al.
Veröffentlicht: (2025)
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision
von: Xi, Zhiheng, et al.
Veröffentlicht: (2024)
von: Xi, Zhiheng, et al.
Veröffentlicht: (2024)
SABER: Switchable and Balanced Training for Efficient LLM Reasoning
von: Zhao, Kai, et al.
Veröffentlicht: (2025)
von: Zhao, Kai, et al.
Veröffentlicht: (2025)
SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning
von: Liu, Bo, et al.
Veröffentlicht: (2025)
von: Liu, Bo, et al.
Veröffentlicht: (2025)
Improving Reasoning Performance in Large Language Models via Representation Engineering
von: Højer, Bertram, et al.
Veröffentlicht: (2025)
von: Højer, Bertram, et al.
Veröffentlicht: (2025)
AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-Play
von: Xu, Ran, et al.
Veröffentlicht: (2025)
von: Xu, Ran, et al.
Veröffentlicht: (2025)
Interpreting and Controlling LLM Reasoning through Integrated Policy Gradient
von: Li, Changming, et al.
Veröffentlicht: (2026)
von: Li, Changming, et al.
Veröffentlicht: (2026)
Assessing LLM Reasoning Steps via Principal Knowledge Grounding
von: Hwang, Hyeon, et al.
Veröffentlicht: (2025)
von: Hwang, Hyeon, et al.
Veröffentlicht: (2025)
POSS: Position Specialist Generates Better Draft for Speculative Decoding
von: Huang, Langlin, et al.
Veröffentlicht: (2025)
von: Huang, Langlin, et al.
Veröffentlicht: (2025)
Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs
von: Su, Jinyan, et al.
Veröffentlicht: (2025)
von: Su, Jinyan, et al.
Veröffentlicht: (2025)
Policy Split: Incentivizing Dual-Mode Exploration in LLM Reinforcement with Dual-Mode Entropy Regularization
von: Yao, Jiashu, et al.
Veröffentlicht: (2026)
von: Yao, Jiashu, et al.
Veröffentlicht: (2026)
Adaptive Test-Time Reasoning via Reward-Guided Dual-Phase Search
von: Cui, Yingqian, et al.
Veröffentlicht: (2025)
von: Cui, Yingqian, et al.
Veröffentlicht: (2025)
KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation
von: Shi, Jiajun, et al.
Veröffentlicht: (2025)
von: Shi, Jiajun, et al.
Veröffentlicht: (2025)
WildVis: Open Source Visualizer for Million-Scale Chat Logs in the Wild
von: Deng, Yuntian, et al.
Veröffentlicht: (2024)
von: Deng, Yuntian, et al.
Veröffentlicht: (2024)
ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning
von: Potamitis, Nearchos, et al.
Veröffentlicht: (2025)
von: Potamitis, Nearchos, et al.
Veröffentlicht: (2025)
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
von: Dong, Guanting, et al.
Veröffentlicht: (2025)
von: Dong, Guanting, et al.
Veröffentlicht: (2025)
DRAGON: Guard LLM Unlearning in Context via Negative Detection and Reasoning
von: Wang, Yaxuan, et al.
Veröffentlicht: (2025)
von: Wang, Yaxuan, et al.
Veröffentlicht: (2025)
Temporal Consistency for LLM Reasoning Process Error Identification
von: Guo, Jiacheng, et al.
Veröffentlicht: (2025)
von: Guo, Jiacheng, et al.
Veröffentlicht: (2025)
How Is LLM Reasoning Distracted by Irrelevant Context? An Analysis Using a Controlled Benchmark
von: Yang, Minglai, et al.
Veröffentlicht: (2025)
von: Yang, Minglai, et al.
Veröffentlicht: (2025)
QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
von: Tseng, Albert, et al.
Veröffentlicht: (2024)
von: Tseng, Albert, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
HAPO: Training Language Models to Reason Concisely via History-Aware Policy Optimization
von: Huang, Chengyu, et al.
Veröffentlicht: (2025) -
Memorization vs. Reasoning: Updating LLMs with New Knowledge
von: Li, Aochong Oliver, et al.
Veröffentlicht: (2025) -
Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text
von: Huang, Chengyu, et al.
Veröffentlicht: (2026) -
How Far Are We From True Auto-Research?
von: Zhang, Zhengxin, et al.
Veröffentlicht: (2026) -
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
von: Chen, Jiaqi, et al.
Veröffentlicht: (2025)