Curriculum Learning for Efficient Chain-of-Thought Distillation via Structure-Aware Masking and GRPO
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Bowen, Wang, Maolin, Zhang, Sheng, Wang, Binhao, Wen, Yi, Gao, Jingtong, Liu, Bowen, Zhao, Zimo, Wang, Wanyu, Zhao, Xiangyu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Data Efficient Adaptation in Large Language Models via Continuous Low-Rank Fine-Tuning
by: Han, Xiao, et al.
Published: (2025)
by: Han, Xiao, et al.
Published: (2025)
Embedding in Recommender Systems: A Survey
by: Wang, Maolin, et al.
Published: (2023)
by: Wang, Maolin, et al.
Published: (2023)
GLINT-RU: Gated Lightweight Intelligent Recurrent Units for Sequential Recommender Systems
by: Zhang, Sheng, et al.
Published: (2024)
by: Zhang, Sheng, et al.
Published: (2024)
SEARCH-R: Structured Entity-Aware Retrieval with Chain-of-Reasoning Navigator for Multi-hop Question Answering
by: Fu, Yuqing, et al.
Published: (2026)
by: Fu, Yuqing, et al.
Published: (2026)
Renormalization Group Guided Tensor Network Structure Search
by: Wang, Maolin, et al.
Published: (2025)
by: Wang, Maolin, et al.
Published: (2025)
Learning to Maximize Mutual Information for Chain-of-Thought Distillation
by: Chen, Xin, et al.
Published: (2024)
by: Chen, Xin, et al.
Published: (2024)
DANCE: Resource-Efficient Neural Architecture Search with Data-Aware and Continuous Adaptation
by: Wang, Maolin, et al.
Published: (2025)
by: Wang, Maolin, et al.
Published: (2025)
Improving Generalization in Intent Detection: GRPO with Reward-Based Curriculum Sampling
by: Feng, Zihao, et al.
Published: (2025)
by: Feng, Zihao, et al.
Published: (2025)
SkyLink: A Large Vision-Language Model Driven Re-ranking Framework for Cross-View UAV geolocalization
by: Liu, Bowen, et al.
Published: (2026)
by: Liu, Bowen, et al.
Published: (2026)
MetaLoRA: Tensor-Enhanced Adaptive Low-Rank Fine-tuning
by: Wang, Maolin, et al.
Published: (2025)
by: Wang, Maolin, et al.
Published: (2025)
SPARK: Adaptive Low-Rank Knowledge Graph Modeling in Hybrid Geometric Spaces for Recommendation
by: Wang, Binhao, et al.
Published: (2025)
by: Wang, Binhao, et al.
Published: (2025)
FindRec: Stein-Guided Entropic Flow for Multi-Modal Sequential Recommendation
by: Wang, Maolin, et al.
Published: (2025)
by: Wang, Maolin, et al.
Published: (2025)
Large Multimodal Model Compression via Efficient Pruning and Distillation at AntGroup
by: Wang, Maolin, et al.
Published: (2023)
by: Wang, Maolin, et al.
Published: (2023)
Efficient Reasoning for LLMs through Speculative Chain-of-Thought
by: Wang, Jikai, et al.
Published: (2025)
by: Wang, Jikai, et al.
Published: (2025)
Less is More Tokens: Efficient Math Reasoning via Difficulty-Aware Chain-of-Thought Distillation
by: Waheed, Abdul, et al.
Published: (2025)
by: Waheed, Abdul, et al.
Published: (2025)
Scenario-Wise Rec: A Multi-Scenario Recommendation Benchmark
by: Li, Xiaopeng, et al.
Published: (2024)
by: Li, Xiaopeng, et al.
Published: (2024)
STAR-Rec: Making Peace with Length Variance and Pattern Diversity in Sequential Recommendation
by: Wang, Maolin, et al.
Published: (2025)
by: Wang, Maolin, et al.
Published: (2025)
RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward
by: Wang, Zongsheng, et al.
Published: (2025)
by: Wang, Zongsheng, et al.
Published: (2025)
BranchGRPO: Stable and Efficient GRPO with Structured Branching in Diffusion Models
by: Li, Yuming, et al.
Published: (2025)
by: Li, Yuming, et al.
Published: (2025)
LaF-GRPO: In-Situ Navigation Instruction Generation for the Visually Impaired via GRPO with LLM-as-Follower Reward
by: Zhao, Yi, et al.
Published: (2025)
by: Zhao, Yi, et al.
Published: (2025)
Multimodal Recommender Systems: A Survey
by: Liu, Qidong, et al.
Published: (2023)
by: Liu, Qidong, et al.
Published: (2023)
LLM4Rerank: LLM-based Auto-Reranking Framework for Recommendations
by: Gao, Jingtong, et al.
Published: (2024)
by: Gao, Jingtong, et al.
Published: (2024)
PRISM: Purified Representation and Integrated Semantic Modeling for Generative Sequential Recommendation
by: Fang, Dengzhao, et al.
Published: (2026)
by: Fang, Dengzhao, et al.
Published: (2026)
HiD-VAE: Interpretable Generative Recommendation via Hierarchical and Disentangled Semantic IDs
by: Fang, Dengzhao, et al.
Published: (2025)
by: Fang, Dengzhao, et al.
Published: (2025)
ToolSample: Dual Dynamic Sampling Methods with Curriculum Learning for RL-based Tool Learning
by: Feng, Zihao, et al.
Published: (2025)
by: Feng, Zihao, et al.
Published: (2025)
SIGMA: Selective Gated Mamba for Sequential Recommendation
by: Liu, Ziwei, et al.
Published: (2024)
by: Liu, Ziwei, et al.
Published: (2024)
Navigate the Unknown: Enhancing LLM Reasoning with Intrinsic Motivation Guided Exploration
by: Gao, Jingtong, et al.
Published: (2025)
by: Gao, Jingtong, et al.
Published: (2025)
Logic Jailbreak: Efficiently Unlocking LLM Safety Restrictions Through Formal Logical Expression
by: Peng, Jingyu, et al.
Published: (2025)
by: Peng, Jingyu, et al.
Published: (2025)
AdaSwitch: Balancing Exploration and Guidance in Knowledge Distillation via Adaptive Switching
by: Peng, Jingyu, et al.
Published: (2025)
by: Peng, Jingyu, et al.
Published: (2025)
Cumulative Distribution Function based General Temporal Point Processes
by: Wang, Maolin, et al.
Published: (2024)
by: Wang, Maolin, et al.
Published: (2024)
Tailored Teaching with Balanced Difficulty: Elevating Reasoning in Multimodal Chain-of-Thought via Prompt Curriculum
by: Yang, Xinglong, et al.
Published: (2025)
by: Yang, Xinglong, et al.
Published: (2025)
Job Skill Extraction via LLM-Centric Multi-Module Framework
by: Li, Guojing, et al.
Published: (2026)
by: Li, Guojing, et al.
Published: (2026)
Efficient Reasoning via Chain of Unconscious Thought
by: Gong, Ruihan, et al.
Published: (2025)
by: Gong, Ruihan, et al.
Published: (2025)
Put Teacher in Student's Shoes: Cross-Distillation for Ultra-compact Model Compression Framework
by: Wang, Maolin, et al.
Published: (2025)
by: Wang, Maolin, et al.
Published: (2025)
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
by: Chen, Minghan, et al.
Published: (2025)
by: Chen, Minghan, et al.
Published: (2025)
Slow Tuning and Low-Entropy Masking for Safe Chain-of-Thought Distillation
by: Ma, Ziyang, et al.
Published: (2025)
by: Ma, Ziyang, et al.
Published: (2025)
LinRec: Linear Attention Mechanism for Long-term Sequential Recommender Systems
by: Liu, Langming, et al.
Published: (2024)
by: Liu, Langming, et al.
Published: (2024)
EchoSR: Efficient Context Harnessing for Lightweight Image Super-Resolution
by: Zhao, Hanli, et al.
Published: (2026)
by: Zhao, Hanli, et al.
Published: (2026)
CoT-BERT: Enhancing Unsupervised Sentence Representation through Chain-of-Thought
by: Zhang, Bowen, et al.
Published: (2023)
by: Zhang, Bowen, et al.
Published: (2023)
OP-GRPO: Efficient Off-Policy GRPO for Flow-Matching Models
by: Zhang, Liyu, et al.
Published: (2026)
by: Zhang, Liyu, et al.
Published: (2026)
Similar Items
-
Data Efficient Adaptation in Large Language Models via Continuous Low-Rank Fine-Tuning
by: Han, Xiao, et al.
Published: (2025) -
Embedding in Recommender Systems: A Survey
by: Wang, Maolin, et al.
Published: (2023) -
GLINT-RU: Gated Lightweight Intelligent Recurrent Units for Sequential Recommender Systems
by: Zhang, Sheng, et al.
Published: (2024) -
SEARCH-R: Structured Entity-Aware Retrieval with Chain-of-Reasoning Navigator for Multi-hop Question Answering
by: Fu, Yuqing, et al.
Published: (2026) -
Renormalization Group Guided Tensor Network Structure Search
by: Wang, Maolin, et al.
Published: (2025)