Gespeichert in:
| Hauptverfasser: | Zhang, Haoran, Li, Yafu, Wang, Zhi, Wang, Zhilin, Zhang, Shunkai, Qu, Xiaoye, Cheng, Yu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2602.08498 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
New Skills or Sharper Primitives? A Probabilistic Perspective on the Emergence of Reasoning in RLVR
von: Wang, Zhilin, et al.
Veröffentlicht: (2026)
von: Wang, Zhilin, et al.
Veröffentlicht: (2026)
SEE: Continual Fine-tuning with Sequential Ensemble of Experts
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models
von: Fu, Tingchen, et al.
Veröffentlicht: (2025)
von: Fu, Tingchen, et al.
Veröffentlicht: (2025)
Draft-OPD: On-Policy Distillation for Speculative Draft Models
von: Lei, Haodi, et al.
Veröffentlicht: (2026)
von: Lei, Haodi, et al.
Veröffentlicht: (2026)
Towards an AI Musician: Synthesizing Sheet Music Problems for Musical Reasoning
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual Feedback
von: Li, Yafu, et al.
Veröffentlicht: (2025)
von: Li, Yafu, et al.
Veröffentlicht: (2025)
Reasoning over Boundaries: Enhancing Specification Alignment via Test-time Deliberation
von: Zhang, Haoran, et al.
Veröffentlicht: (2025)
von: Zhang, Haoran, et al.
Veröffentlicht: (2025)
Learning to Reason under Off-Policy Guidance
von: Yan, Jianhao, et al.
Veröffentlicht: (2025)
von: Yan, Jianhao, et al.
Veröffentlicht: (2025)
FaithRL: Learning to Reason Faithfully through Step-Level Faithfulness Maximization
von: Gui, Runquan, et al.
Veröffentlicht: (2026)
von: Gui, Runquan, et al.
Veröffentlicht: (2026)
ExGRPO: Learning to Reason from Experience
von: Zhan, Runzhe, et al.
Veröffentlicht: (2025)
von: Zhan, Runzhe, et al.
Veröffentlicht: (2025)
Unveiling Attractor Cycles in Large Language Models: A Dynamical Systems View of Successive Paraphrasing
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
Advancing Multimodal Reasoning: From Optimized Cold Start to Staged Reinforcement Learning
von: Chen, Shuang, et al.
Veröffentlicht: (2025)
von: Chen, Shuang, et al.
Veröffentlicht: (2025)
Teaching Thinking Models to Reason with Tools: A Full-Pipeline Recipe for Tool-Integrated Reasoning
von: Cheng, Qianjia, et al.
Veröffentlicht: (2026)
von: Cheng, Qianjia, et al.
Veröffentlicht: (2026)
Achieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified Scaling
von: Li, Yafu, et al.
Veröffentlicht: (2026)
von: Li, Yafu, et al.
Veröffentlicht: (2026)
From Drafts to Answers: Unlocking LLM Potential via Aggregation Fine-Tuning
von: Li, Yafu, et al.
Veröffentlicht: (2025)
von: Li, Yafu, et al.
Veröffentlicht: (2025)
Spotting AI's Touch: Identifying LLM-Paraphrased Spans in Text
von: Li, Yafu, et al.
Veröffentlicht: (2024)
von: Li, Yafu, et al.
Veröffentlicht: (2024)
LatentMem: Customizing Latent Memory for Multi-Agent Systems
von: Fu, Muxin, et al.
Veröffentlicht: (2026)
von: Fu, Muxin, et al.
Veröffentlicht: (2026)
Multi-LLM Collaborative Search for Complex Problem Solving
von: Yang, Sen, et al.
Veröffentlicht: (2025)
von: Yang, Sen, et al.
Veröffentlicht: (2025)
Timo: Towards Better Temporal Reasoning for Language Models
von: Su, Zhaochen, et al.
Veröffentlicht: (2024)
von: Su, Zhaochen, et al.
Veröffentlicht: (2024)
Lost in Literalism: How Supervised Training Shapes Translationese in LLMs
von: Li, Yafu, et al.
Veröffentlicht: (2025)
von: Li, Yafu, et al.
Veröffentlicht: (2025)
$π$-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows
von: Zhang, Haoran, et al.
Veröffentlicht: (2026)
von: Zhang, Haoran, et al.
Veröffentlicht: (2026)
DiffThinker: Towards Generative Multimodal Reasoning with Diffusion Models
von: He, Zefeng, et al.
Veröffentlicht: (2025)
von: He, Zefeng, et al.
Veröffentlicht: (2025)
MAGE: Machine-generated Text Detection in the Wild
von: Li, Yafu, et al.
Veröffentlicht: (2023)
von: Li, Yafu, et al.
Veröffentlicht: (2023)
A Survey of Efficient Reasoning for Large Reasoning Models: Language, Multimodality, and Beyond
von: Qu, Xiaoye, et al.
Veröffentlicht: (2025)
von: Qu, Xiaoye, et al.
Veröffentlicht: (2025)
Think Longer to Explore Deeper: Learn to Explore In-Context via Length-Incentivized Reinforcement Learning
von: Wang, Futing, et al.
Veröffentlicht: (2026)
von: Wang, Futing, et al.
Veröffentlicht: (2026)
Living in the Moment: Can Large Language Models Grasp Co-Temporal Reasoning?
von: Su, Zhaochen, et al.
Veröffentlicht: (2024)
von: Su, Zhaochen, et al.
Veröffentlicht: (2024)
ConflictBank: A Benchmark for Evaluating the Influence of Knowledge Conflicts in LLM
von: Su, Zhaochen, et al.
Veröffentlicht: (2024)
von: Su, Zhaochen, et al.
Veröffentlicht: (2024)
Joint Multi-Facts Reasoning Network For Complex Temporal Question Answering Over Knowledge Graph
von: Huang, Rikui, et al.
Veröffentlicht: (2024)
von: Huang, Rikui, et al.
Veröffentlicht: (2024)
Rethinking Entropy Regularization in Large Reasoning Models
von: Jiang, Yuxian, et al.
Veröffentlicht: (2025)
von: Jiang, Yuxian, et al.
Veröffentlicht: (2025)
Potential and Challenges of Model Editing for Social Debiasing
von: Yan, Jianhao, et al.
Veröffentlicht: (2024)
von: Yan, Jianhao, et al.
Veröffentlicht: (2024)
From Conflict to Consensus: Boosting Medical Reasoning via Multi-Round Agentic RAG
von: Wu, Wenhao, et al.
Veröffentlicht: (2026)
von: Wu, Wenhao, et al.
Veröffentlicht: (2026)
Make LoRA Great Again: Boosting LoRA with Adaptive Singular Values and Mixture-of-Experts Optimization Alignment
von: Fan, Chenghao, et al.
Veröffentlicht: (2025)
von: Fan, Chenghao, et al.
Veröffentlicht: (2025)
Keys to Robust Edits: from Theoretical Insights to Practical Advances
von: Yan, Jianhao, et al.
Veröffentlicht: (2024)
von: Yan, Jianhao, et al.
Veröffentlicht: (2024)
VideoSSR: Video Self-Supervised Reinforcement Learning
von: He, Zefeng, et al.
Veröffentlicht: (2025)
von: He, Zefeng, et al.
Veröffentlicht: (2025)
FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame Spotlighting
von: He, Zefeng, et al.
Veröffentlicht: (2025)
von: He, Zefeng, et al.
Veröffentlicht: (2025)
Mitigating Multilingual Hallucination in Large Vision-Language Models
von: Qu, Xiaoye, et al.
Veröffentlicht: (2024)
von: Qu, Xiaoye, et al.
Veröffentlicht: (2024)
What Have We Achieved on Non-autoregressive Translation?
von: Li, Yafu, et al.
Veröffentlicht: (2024)
von: Li, Yafu, et al.
Veröffentlicht: (2024)
Dynamic Data Mixing Maximizes Instruction Tuning for Mixture-of-Experts
von: Zhu, Tong, et al.
Veröffentlicht: (2024)
von: Zhu, Tong, et al.
Veröffentlicht: (2024)
LLaMA-MoE v2: Exploring Sparsity of LLaMA from Perspective of Mixture-of-Experts with Post-Training
von: Qu, Xiaoye, et al.
Veröffentlicht: (2024)
von: Qu, Xiaoye, et al.
Veröffentlicht: (2024)
SentiXRL: An advanced large language Model Framework for Multilingual Fine-Grained Emotion Classification in Complex Text Environment
von: Wang, Jie, et al.
Veröffentlicht: (2024)
von: Wang, Jie, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
New Skills or Sharper Primitives? A Probabilistic Perspective on the Emergence of Reasoning in RLVR
von: Wang, Zhilin, et al.
Veröffentlicht: (2026) -
SEE: Continual Fine-tuning with Sequential Ensemble of Experts
von: Wang, Zhilin, et al.
Veröffentlicht: (2025) -
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models
von: Fu, Tingchen, et al.
Veröffentlicht: (2025) -
Draft-OPD: On-Policy Distillation for Speculative Draft Models
von: Lei, Haodi, et al.
Veröffentlicht: (2026) -
Towards an AI Musician: Synthesizing Sheet Music Problems for Musical Reasoning
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)