Gespeichert in:
| Hauptverfasser: | Liu, Wanlong, Xu, Junxiao, Yu, Fei, Lin, Yukang, Ji, Ke, Chen, Wenyu, Xu, Yan, Wang, Yasheng, Shang, Lifeng, Wang, Benyou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2506.12860 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RAG-Instruct: Boosting LLMs with Diverse Retrieval-Augmented Instructions
von: Liu, Wanlong, et al.
Veröffentlicht: (2024)
von: Liu, Wanlong, et al.
Veröffentlicht: (2024)
Towards Understanding Fine-Tuning Mechanisms of LLMs via Circuit Analysis
von: Wang, Xu, et al.
Veröffentlicht: (2025)
von: Wang, Xu, et al.
Veröffentlicht: (2025)
Teaching LLMs According to Their Aptitude: Adaptive Reasoning for Mathematical Problem Solving
von: Xu, Xin, et al.
Veröffentlicht: (2025)
von: Xu, Xin, et al.
Veröffentlicht: (2025)
HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs
von: Chen, Junying, et al.
Veröffentlicht: (2024)
von: Chen, Junying, et al.
Veröffentlicht: (2024)
The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning Models
von: Ji, Ke, et al.
Veröffentlicht: (2025)
von: Ji, Ke, et al.
Veröffentlicht: (2025)
KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
von: Xu, Hongling, et al.
Veröffentlicht: (2025)
von: Xu, Hongling, et al.
Veröffentlicht: (2025)
Instruction-Tuning Data Synthesis from Scratch via Web Reconstruction
von: Jiang, Yuxin, et al.
Veröffentlicht: (2025)
von: Jiang, Yuxin, et al.
Veröffentlicht: (2025)
Scaling Flaws of Verifier-Guided Search in Mathematical Reasoning
von: Yu, Fei, et al.
Veröffentlicht: (2025)
von: Yu, Fei, et al.
Veröffentlicht: (2025)
Safe: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal Verification
von: Liu, Chengwu, et al.
Veröffentlicht: (2025)
von: Liu, Chengwu, et al.
Veröffentlicht: (2025)
OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning
von: Yu, Fei, et al.
Veröffentlicht: (2023)
von: Yu, Fei, et al.
Veröffentlicht: (2023)
Robust Search with Uncertainty-Aware Value Models for Language Model Reasoning
von: Yu, Fei, et al.
Veröffentlicht: (2025)
von: Yu, Fei, et al.
Veröffentlicht: (2025)
EssayBench: Evaluating Large Language Models in Multi-Genre Chinese Essay Writing
von: Gao, Fan, et al.
Veröffentlicht: (2025)
von: Gao, Fan, et al.
Veröffentlicht: (2025)
Double-Checker: Enhancing Reasoning of Slow-Thinking LLMs via Self-Critical Fine-Tuning
von: Xu, Xin, et al.
Veröffentlicht: (2025)
von: Xu, Xin, et al.
Veröffentlicht: (2025)
ReliableMath: Benchmark of Reliable Mathematical Reasoning on Large Language Models
von: Xue, Boyang, et al.
Veröffentlicht: (2025)
von: Xue, Boyang, et al.
Veröffentlicht: (2025)
Stepwise Reasoning Checkpoint Analysis: A Test Time Scaling Method to Enhance LLMs' Reasoning
von: Wang, Zezhong, et al.
Veröffentlicht: (2025)
von: Wang, Zezhong, et al.
Veröffentlicht: (2025)
Cut Your Losses! Learning to Prune Paths Early for Efficient Parallel Reasoning
von: Bi, Jiaxi, et al.
Veröffentlicht: (2026)
von: Bi, Jiaxi, et al.
Veröffentlicht: (2026)
ToolACE-R: Model-aware Iterative Training and Adaptive Refinement for Tool Learning
von: Zeng, Xingshan, et al.
Veröffentlicht: (2025)
von: Zeng, Xingshan, et al.
Veröffentlicht: (2025)
MathScale: Scaling Instruction Tuning for Mathematical Reasoning
von: Tang, Zhengyang, et al.
Veröffentlicht: (2024)
von: Tang, Zhengyang, et al.
Veröffentlicht: (2024)
DeepDiver: Adaptive Search Intensity Scaling via Open-Web Reinforcement Learning
von: Shi, Wenxuan, et al.
Veröffentlicht: (2025)
von: Shi, Wenxuan, et al.
Veröffentlicht: (2025)
DAST: Difficulty-Aware Self-Training on Large Language Models
von: Xue, Boyang, et al.
Veröffentlicht: (2025)
von: Xue, Boyang, et al.
Veröffentlicht: (2025)
Crowd Comparative Reasoning: Unlocking Comprehensive Evaluations for LLM-as-a-Judge
von: Zhang, Qiyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Qiyuan, et al.
Veröffentlicht: (2025)
How Should We Enhance the Safety of Large Reasoning Models: An Empirical Study
von: Zhang, Zhexin, et al.
Veröffentlicht: (2025)
von: Zhang, Zhexin, et al.
Veröffentlicht: (2025)
ToolACE-MT: Non-Autoregressive Generation for Agentic Multi-Turn Interaction
von: Zeng, Xingshan, et al.
Veröffentlicht: (2025)
von: Zeng, Xingshan, et al.
Veröffentlicht: (2025)
ARTIS: Agentic Risk-Aware Test-Time Scaling via Iterative Simulation
von: Zeng, Xingshan, et al.
Veröffentlicht: (2026)
von: Zeng, Xingshan, et al.
Veröffentlicht: (2026)
Improving Language Model Reasoning with Self-motivated Learning
von: Feng, Yunlong, et al.
Veröffentlicht: (2024)
von: Feng, Yunlong, et al.
Veröffentlicht: (2024)
Learning from Peers in Reasoning Models
von: Luo, Tongxu, et al.
Veröffentlicht: (2025)
von: Luo, Tongxu, et al.
Veröffentlicht: (2025)
Does higher interpretability imply better utility? A Pairwise Analysis on Sparse Autoencoders
von: Wang, Xu, et al.
Veröffentlicht: (2025)
von: Wang, Xu, et al.
Veröffentlicht: (2025)
LLMs Could Autonomously Learn Without External Supervision
von: Ji, Ke, et al.
Veröffentlicht: (2024)
von: Ji, Ke, et al.
Veröffentlicht: (2024)
Position: The Real Barrier to LLM Agent Usability is Agentic ROI
von: Liu, Weiwen, et al.
Veröffentlicht: (2025)
von: Liu, Weiwen, et al.
Veröffentlicht: (2025)
Model Unlearning via Sparse Autoencoder Subspace Guided Projections
von: Wang, Xu, et al.
Veröffentlicht: (2025)
von: Wang, Xu, et al.
Veröffentlicht: (2025)
Recall, Retrieve and Reason: Towards Better In-Context Relation Extraction
von: Li, Guozheng, et al.
Veröffentlicht: (2024)
von: Li, Guozheng, et al.
Veröffentlicht: (2024)
Self-Error-Instruct: Generalizing from Errors for LLMs Mathematical Reasoning
von: Yu, Erxin, et al.
Veröffentlicht: (2025)
von: Yu, Erxin, et al.
Veröffentlicht: (2025)
ToolFlow: Boosting LLM Tool-Calling Through Natural and Coherent Dialogue Synthesis
von: Wang, Zezhong, et al.
Veröffentlicht: (2024)
von: Wang, Zezhong, et al.
Veröffentlicht: (2024)
Adapting Large Language Models to Low-Resource Tibetan: A Two-Stage Continual and Supervised Fine-Tuning Study
von: Chen, Lifeng, et al.
Veröffentlicht: (2025)
von: Chen, Lifeng, et al.
Veröffentlicht: (2025)
Bridging and Modeling Correlations in Pairwise Data for Direct Preference Optimization
von: Jiang, Yuxin, et al.
Veröffentlicht: (2024)
von: Jiang, Yuxin, et al.
Veröffentlicht: (2024)
Long-Short Chain-of-Thought Mixture Supervised Fine-Tuning Eliciting Efficient Reasoning in Large Language Models
von: Yu, Bin, et al.
Veröffentlicht: (2025)
von: Yu, Bin, et al.
Veröffentlicht: (2025)
MLPs Compass: What is learned when MLPs are combined with PLMs?
von: Zhou, Li, et al.
Veröffentlicht: (2024)
von: Zhou, Li, et al.
Veröffentlicht: (2024)
Mixed-Precision Graph Neural Quantization for Low Bit Large Language Models
von: Liu, Wanlong, et al.
Veröffentlicht: (2025)
von: Liu, Wanlong, et al.
Veröffentlicht: (2025)
MindGYM: What Matters in Question Synthesis for Thinking-Centric Fine-Tuning?
von: Xu, Zhe, et al.
Veröffentlicht: (2025)
von: Xu, Zhe, et al.
Veröffentlicht: (2025)
Chain-of-Probe: Examining the Necessity and Accuracy of CoT Step-by-Step
von: Wang, Zezhong, et al.
Veröffentlicht: (2024)
von: Wang, Zezhong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
RAG-Instruct: Boosting LLMs with Diverse Retrieval-Augmented Instructions
von: Liu, Wanlong, et al.
Veröffentlicht: (2024) -
Towards Understanding Fine-Tuning Mechanisms of LLMs via Circuit Analysis
von: Wang, Xu, et al.
Veröffentlicht: (2025) -
Teaching LLMs According to Their Aptitude: Adaptive Reasoning for Mathematical Problem Solving
von: Xu, Xin, et al.
Veröffentlicht: (2025) -
HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs
von: Chen, Junying, et al.
Veröffentlicht: (2024) -
The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning Models
von: Ji, Ke, et al.
Veröffentlicht: (2025)