GRACE: Gradient-aligned Reasoning Data Curation for Efficient Post-training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Junjie, Wang, Ziao, Ma, NingXuan, Ma, Jianghong, Zhang, Xiaofeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Uncovering Intrinsic Capabilities: A Paradigm for Data Curation in Vision-Language Models
von: Li, Junjie, et al.
Veröffentlicht: (2025)
von: Li, Junjie, et al.
Veröffentlicht: (2025)
GiVE: Guiding Visual Encoder to Perceive Overlooked Information
von: Li, Junjie, et al.
Veröffentlicht: (2024)
von: Li, Junjie, et al.
Veröffentlicht: (2024)
MLGIB: Multi-Label Graph Information Bottleneck for Expressive and Robust Message Passing
von: Wu, Chaokai, et al.
Veröffentlicht: (2026)
von: Wu, Chaokai, et al.
Veröffentlicht: (2026)
Does Faithfulness Conflict with Plausibility? An Empirical Study in Explainable AI across NLP Tasks
von: Lu, Xiaolei, et al.
Veröffentlicht: (2024)
von: Lu, Xiaolei, et al.
Veröffentlicht: (2024)
GRACE: Discriminator-Guided Chain-of-Thought Reasoning
von: Khalifa, Muhammad, et al.
Veröffentlicht: (2023)
von: Khalifa, Muhammad, et al.
Veröffentlicht: (2023)
On Data Synthesis and Post-training for Visual Abstract Reasoning
von: Zhu, Ke, et al.
Veröffentlicht: (2025)
von: Zhu, Ke, et al.
Veröffentlicht: (2025)
Dual-Phase Playtime-guided Recommendation: Interest Intensity Exploration and Multimodal Random Walks
von: Zhang, Jingmao, et al.
Veröffentlicht: (2025)
von: Zhang, Jingmao, et al.
Veröffentlicht: (2025)
Diversity Recommendation via Causal Deconfounding of Co-purchase Relations and Counterfactual Exposure
von: Zhang, Jingmao, et al.
Veröffentlicht: (2025)
von: Zhang, Jingmao, et al.
Veröffentlicht: (2025)
STARec: An Efficient Agent Framework for Recommender Systems via Autonomous Deliberate Reasoning
von: Wu, Chenghao, et al.
Veröffentlicht: (2025)
von: Wu, Chenghao, et al.
Veröffentlicht: (2025)
IDVT: Interest-aware Denoising and View-guided Tuning for Social Recommendation
von: Yang, Dezhao, et al.
Veröffentlicht: (2023)
von: Yang, Dezhao, et al.
Veröffentlicht: (2023)
CircuitSeer: Mining High-Quality Data by Probing Mathematical Reasoning Circuits in LLMs
von: Wang, Shaobo, et al.
Veröffentlicht: (2025)
von: Wang, Shaobo, et al.
Veröffentlicht: (2025)
DCA-Bench: A Benchmark for Dataset Curation Agents
von: Huang, Benhao, et al.
Veröffentlicht: (2024)
von: Huang, Benhao, et al.
Veröffentlicht: (2024)
DaGRPO: Rectifying Gradient Conflict in Reasoning via Distinctiveness-Aware Group Relative Policy Optimization
von: Xie, Xuan, et al.
Veröffentlicht: (2025)
von: Xie, Xuan, et al.
Veröffentlicht: (2025)
What Matters in Data Curation for Multimodal Reasoning? Insights from the DCVLR Challenge
von: Shin, Yosub, et al.
Veröffentlicht: (2026)
von: Shin, Yosub, et al.
Veröffentlicht: (2026)
Layer-Aware Influence for Online Data Valuation Estimation
von: Yang, Ziao, et al.
Veröffentlicht: (2025)
von: Yang, Ziao, et al.
Veröffentlicht: (2025)
RADAR: Accelerating Large Language Model Inference With RL-Based Dynamic Draft Trees
von: Ma, Junjie, et al.
Veröffentlicht: (2025)
von: Ma, Junjie, et al.
Veröffentlicht: (2025)
How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise Gradients
von: Li, Ming, et al.
Veröffentlicht: (2025)
von: Li, Ming, et al.
Veröffentlicht: (2025)
Data to Defense: The Role of Curation in Customizing LLMs Against Jailbreaking Attacks
von: Liu, Xiaoqun, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoqun, et al.
Veröffentlicht: (2024)
The Impact of Post-training on Data Contamination
von: Kocyigit, Muhammed Yusuf, et al.
Veröffentlicht: (2026)
von: Kocyigit, Muhammed Yusuf, et al.
Veröffentlicht: (2026)
Memory as Action: Autonomous Context Curation for Long-Horizon Agentic Tasks
von: Zhang, Yuxiang, et al.
Veröffentlicht: (2025)
von: Zhang, Yuxiang, et al.
Veröffentlicht: (2025)
Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost
von: Xuan, Richmond Sin Jing, et al.
Veröffentlicht: (2026)
von: Xuan, Richmond Sin Jing, et al.
Veröffentlicht: (2026)
Efficient Post-training Quantization with FP8 Formats
von: Shen, Haihao, et al.
Veröffentlicht: (2023)
von: Shen, Haihao, et al.
Veröffentlicht: (2023)
Leash: Adaptive Length Penalty and Reward Shaping for Efficient Large Reasoning Model
von: Li, Yanhao, et al.
Veröffentlicht: (2025)
von: Li, Yanhao, et al.
Veröffentlicht: (2025)
Not All Preferences Are Created Equal: Stability-Aware and Gradient-Efficient Alignment for Reasoning Models
von: Wu, Hui, et al.
Veröffentlicht: (2026)
von: Wu, Hui, et al.
Veröffentlicht: (2026)
Post-training for Efficient Communication via Convention Formation
von: Hua, Yilun, et al.
Veröffentlicht: (2025)
von: Hua, Yilun, et al.
Veröffentlicht: (2025)
Efficient Post-training of LLMs for Code Generation With Offline Reinforcement Learning
von: Wu, Mingze, et al.
Veröffentlicht: (2026)
von: Wu, Mingze, et al.
Veröffentlicht: (2026)
Efficient Reasoning Models: A Survey
von: Feng, Sicheng, et al.
Veröffentlicht: (2025)
von: Feng, Sicheng, et al.
Veröffentlicht: (2025)
CuraLight: Debate-Guided Data Curation for LLM-Centered Traffic Signal Control
von: Guo, Qing, et al.
Veröffentlicht: (2026)
von: Guo, Qing, et al.
Veröffentlicht: (2026)
Neuro-Symbolic Data Generation for Math Reasoning
von: Li, Zenan, et al.
Veröffentlicht: (2024)
von: Li, Zenan, et al.
Veröffentlicht: (2024)
Mitigating Overthinking in Large Reasoning Models via Difficulty-aware Reinforcement Learning
von: Wan, Qian, et al.
Veröffentlicht: (2026)
von: Wan, Qian, et al.
Veröffentlicht: (2026)
Revisiting the Data Sampling in Multimodal Post-training from a Difficulty-Distinguish View
von: Qi, Jianyu, et al.
Veröffentlicht: (2025)
von: Qi, Jianyu, et al.
Veröffentlicht: (2025)
GRACE: A Dynamic Coreset Selection Framework for Large Language Model Optimization
von: Tang, Tianhao, et al.
Veröffentlicht: (2026)
von: Tang, Tianhao, et al.
Veröffentlicht: (2026)
Efficient Reasoning with Balanced Thinking
von: Li, Yulin, et al.
Veröffentlicht: (2026)
von: Li, Yulin, et al.
Veröffentlicht: (2026)
MPS-Prover: Advancing Stepwise Theorem Proving by Multi-Perspective Search and Data Curation
von: Liang, Zhenwen, et al.
Veröffentlicht: (2025)
von: Liang, Zhenwen, et al.
Veröffentlicht: (2025)
Resource-Efficient Reinforcement for Reasoning Large Language Models via Dynamic One-Shot Policy Refinement
von: Zhang, Yunjian, et al.
Veröffentlicht: (2026)
von: Zhang, Yunjian, et al.
Veröffentlicht: (2026)
Training Data Selection with Gradient Orthogonality for Efficient Domain Adaptation
von: Zhang, Xiyang, et al.
Veröffentlicht: (2026)
von: Zhang, Xiyang, et al.
Veröffentlicht: (2026)
DocMamba: Efficient Document Pre-training with State Space Model
von: Hu, Pengfei, et al.
Veröffentlicht: (2024)
von: Hu, Pengfei, et al.
Veröffentlicht: (2024)
On-Policy Supervised Fine-Tuning for Efficient Reasoning
von: Zhao, Anhao, et al.
Veröffentlicht: (2026)
von: Zhao, Anhao, et al.
Veröffentlicht: (2026)
Curated LLM: Synergy of LLMs and Data Curation for tabular augmentation in low-data regimes
von: Seedat, Nabeel, et al.
Veröffentlicht: (2023)
von: Seedat, Nabeel, et al.
Veröffentlicht: (2023)
CPGRec+: A Balance-oriented Framework for Personalized Video Game Recommendations
von: Li, Xiping, et al.
Veröffentlicht: (2026)
von: Li, Xiping, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Uncovering Intrinsic Capabilities: A Paradigm for Data Curation in Vision-Language Models
von: Li, Junjie, et al.
Veröffentlicht: (2025) -
GiVE: Guiding Visual Encoder to Perceive Overlooked Information
von: Li, Junjie, et al.
Veröffentlicht: (2024) -
MLGIB: Multi-Label Graph Information Bottleneck for Expressive and Robust Message Passing
von: Wu, Chaokai, et al.
Veröffentlicht: (2026) -
Does Faithfulness Conflict with Plausibility? An Empirical Study in Explainable AI across NLP Tasks
von: Lu, Xiaolei, et al.
Veröffentlicht: (2024) -
GRACE: Discriminator-Guided Chain-of-Thought Reasoning
von: Khalifa, Muhammad, et al.
Veröffentlicht: (2023)