DAST: Difficulty-Aware Self-Training on Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xue, Boyang, Zhu, Qi, Wang, Hongru, Wang, Rui, Wang, Sheng, Xu, Hongling, Mi, Fei, Wang, Yasheng, Shang, Lifeng, Liu, Qun, Wong, Kam-Fai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ReliableMath: Benchmark of Reliable Mathematical Reasoning on Large Language Models
von: Xue, Boyang, et al.
Veröffentlicht: (2025)
von: Xue, Boyang, et al.
Veröffentlicht: (2025)
UAlign: Leveraging Uncertainty Estimations for Factuality Alignment on Large Language Models
von: Xue, Boyang, et al.
Veröffentlicht: (2024)
von: Xue, Boyang, et al.
Veröffentlicht: (2024)
UniRetriever: Multi-task Candidates Selection for Various Context-Adaptive Conversational Retrieval
von: Wang, Hongru, et al.
Veröffentlicht: (2024)
von: Wang, Hongru, et al.
Veröffentlicht: (2024)
Role Prompting Guided Domain Adaptation with General Capability Preserve for Large Language Models
von: Wang, Rui, et al.
Veröffentlicht: (2024)
von: Wang, Rui, et al.
Veröffentlicht: (2024)
Enhancing Large Language Models Against Inductive Instructions with Dual-critique Prompting
von: Wang, Rui, et al.
Veröffentlicht: (2023)
von: Wang, Rui, et al.
Veröffentlicht: (2023)
MlingConf: A Comprehensive Study of Multilingual Confidence Estimation on Large Language Models
von: Xue, Boyang, et al.
Veröffentlicht: (2024)
von: Xue, Boyang, et al.
Veröffentlicht: (2024)
MlingConf: A Comprehensive Study of Multilingual Confidence Estimation on Large Language Models
von: Xue, Boyang, et al.
Veröffentlicht: (2024)
von: Xue, Boyang, et al.
Veröffentlicht: (2024)
Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models
von: Wang, Rui, et al.
Veröffentlicht: (2025)
von: Wang, Rui, et al.
Veröffentlicht: (2025)
KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
von: Xu, Hongling, et al.
Veröffentlicht: (2025)
von: Xu, Hongling, et al.
Veröffentlicht: (2025)
Data Management For Training Large Language Models: A Survey
von: Wang, Zige, et al.
Veröffentlicht: (2023)
von: Wang, Zige, et al.
Veröffentlicht: (2023)
ToolFlow: Boosting LLM Tool-Calling Through Natural and Coherent Dialogue Synthesis
von: Wang, Zezhong, et al.
Veröffentlicht: (2024)
von: Wang, Zezhong, et al.
Veröffentlicht: (2024)
Stepwise Reasoning Checkpoint Analysis: A Test Time Scaling Method to Enhance LLMs' Reasoning
von: Wang, Zezhong, et al.
Veröffentlicht: (2025)
von: Wang, Zezhong, et al.
Veröffentlicht: (2025)
Chain-of-Probe: Examining the Necessity and Accuracy of CoT Step-by-Step
von: Wang, Zezhong, et al.
Veröffentlicht: (2024)
von: Wang, Zezhong, et al.
Veröffentlicht: (2024)
EssayBench: Evaluating Large Language Models in Multi-Genre Chinese Essay Writing
von: Gao, Fan, et al.
Veröffentlicht: (2025)
von: Gao, Fan, et al.
Veröffentlicht: (2025)
Teaching Large Reasoning Models Effective Reflection
von: Wang, Hanbin, et al.
Veröffentlicht: (2026)
von: Wang, Hanbin, et al.
Veröffentlicht: (2026)
OSPC: Detecting Harmful Memes with Large Language Model as a Catalyst
von: Cao, Jingtao, et al.
Veröffentlicht: (2024)
von: Cao, Jingtao, et al.
Veröffentlicht: (2024)
DAST: Difficulty-Adaptive Slow-Thinking for Large Reasoning Models
von: Shen, Yi, et al.
Veröffentlicht: (2025)
von: Shen, Yi, et al.
Veröffentlicht: (2025)
AppBench: Planning of Multiple APIs from Various APPs for Complex User Instruction
von: Wang, Hongru, et al.
Veröffentlicht: (2024)
von: Wang, Hongru, et al.
Veröffentlicht: (2024)
VLEU: a Method for Automatic Evaluation for Generalizability of Text-to-Image Models
von: Cao, Jingtao, et al.
Veröffentlicht: (2024)
von: Cao, Jingtao, et al.
Veröffentlicht: (2024)
M4LE: A Multi-Ability Multi-Range Multi-Task Multi-Domain Long-Context Evaluation Benchmark for Large Language Models
von: Kwan, Wai-Chung, et al.
Veröffentlicht: (2023)
von: Kwan, Wai-Chung, et al.
Veröffentlicht: (2023)
Self-DC: When to Reason and When to Act? Self Divide-and-Conquer for Compositional Unknown Questions
von: Wang, Hongru, et al.
Veröffentlicht: (2024)
von: Wang, Hongru, et al.
Veröffentlicht: (2024)
UniMS-RAG: A Unified Multi-source Retrieval-Augmented Generation for Personalized Dialogue Systems
von: Wang, Hongru, et al.
Veröffentlicht: (2024)
von: Wang, Hongru, et al.
Veröffentlicht: (2024)
MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models
von: Kwan, Wai-Chung, et al.
Veröffentlicht: (2024)
von: Kwan, Wai-Chung, et al.
Veröffentlicht: (2024)
YODA: Teacher-Student Progressive Learning for Language Models
von: Lu, Jianqiao, et al.
Veröffentlicht: (2024)
von: Lu, Jianqiao, et al.
Veröffentlicht: (2024)
SELF: Self-Evolution with Language Feedback
von: Lu, Jianqiao, et al.
Veröffentlicht: (2023)
von: Lu, Jianqiao, et al.
Veröffentlicht: (2023)
A Survey of the Evolution of Language Model-Based Dialogue Systems: Data, Task and Models
von: Wang, Hongru, et al.
Veröffentlicht: (2023)
von: Wang, Hongru, et al.
Veröffentlicht: (2023)
Position: Agent Should Invoke External Tools ONLY When Epistemically Necessary
von: Wang, Hongru, et al.
Veröffentlicht: (2025)
von: Wang, Hongru, et al.
Veröffentlicht: (2025)
ToolACE-MT: Non-Autoregressive Generation for Agentic Multi-Turn Interaction
von: Zeng, Xingshan, et al.
Veröffentlicht: (2025)
von: Zeng, Xingshan, et al.
Veröffentlicht: (2025)
Self-Guard: Empower the LLM to Safeguard Itself
von: Wang, Zezhong, et al.
Veröffentlicht: (2023)
von: Wang, Zezhong, et al.
Veröffentlicht: (2023)
Self-Reasoning Language Models: Unfold Hidden Reasoning Chains with Few Reasoning Catalyst
von: Wang, Hongru, et al.
Veröffentlicht: (2025)
von: Wang, Hongru, et al.
Veröffentlicht: (2025)
ARTIS: Agentic Risk-Aware Test-Time Scaling via Iterative Simulation
von: Zeng, Xingshan, et al.
Veröffentlicht: (2026)
von: Zeng, Xingshan, et al.
Veröffentlicht: (2026)
Self-Error-Instruct: Generalizing from Errors for LLMs Mathematical Reasoning
von: Yu, Erxin, et al.
Veröffentlicht: (2025)
von: Yu, Erxin, et al.
Veröffentlicht: (2025)
Memory-T1: Reinforcement Learning for Temporal Reasoning in Multi-session Agents
von: Du, Yiming, et al.
Veröffentlicht: (2025)
von: Du, Yiming, et al.
Veröffentlicht: (2025)
WebAggregator: Enhancing Compositional Reasoning Capabilities of Deep Research Agent Foundation Models
von: Wang, Rui, et al.
Veröffentlicht: (2025)
von: Wang, Rui, et al.
Veröffentlicht: (2025)
FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models
von: Jiang, Yuxin, et al.
Veröffentlicht: (2023)
von: Jiang, Yuxin, et al.
Veröffentlicht: (2023)
IndiVec: An Exploration of Leveraging Large Language Models for Media Bias Detection with Fine-Grained Bias Indicators
von: Lin, Luyang, et al.
Veröffentlicht: (2024)
von: Lin, Luyang, et al.
Veröffentlicht: (2024)
PerLTQA: A Personal Long-Term Memory Dataset for Memory Classification, Retrieval, and Synthesis in Question Answering
von: Du, Yiming, et al.
Veröffentlicht: (2024)
von: Du, Yiming, et al.
Veröffentlicht: (2024)
ToolACE-R: Model-aware Iterative Training and Adaptive Refinement for Tool Learning
von: Zeng, Xingshan, et al.
Veröffentlicht: (2025)
von: Zeng, Xingshan, et al.
Veröffentlicht: (2025)
Safe: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal Verification
von: Liu, Chengwu, et al.
Veröffentlicht: (2025)
von: Liu, Chengwu, et al.
Veröffentlicht: (2025)
Evaluating the External and Parametric Knowledge Fusion of Large Language Models
von: Zhang, Hao, et al.
Veröffentlicht: (2024)
von: Zhang, Hao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ReliableMath: Benchmark of Reliable Mathematical Reasoning on Large Language Models
von: Xue, Boyang, et al.
Veröffentlicht: (2025) -
UAlign: Leveraging Uncertainty Estimations for Factuality Alignment on Large Language Models
von: Xue, Boyang, et al.
Veröffentlicht: (2024) -
UniRetriever: Multi-task Candidates Selection for Various Context-Adaptive Conversational Retrieval
von: Wang, Hongru, et al.
Veröffentlicht: (2024) -
Role Prompting Guided Domain Adaptation with General Capability Preserve for Large Language Models
von: Wang, Rui, et al.
Veröffentlicht: (2024) -
Enhancing Large Language Models Against Inductive Instructions with Dual-critique Prompting
von: Wang, Rui, et al.
Veröffentlicht: (2023)