Compress the Easy, Explore the Hard: Difficulty-Aware Entropy Regularization for Efficient LLM Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Luo, Qin-Wen, Ren, Sheng, Chen, Xiang, Liu, Rui, Fang, Jun, Tan, Naiqiang, Huang, Sheng-Jun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Agent-Omit: Adaptive Context Omission for Efficient LLM Agents
von: Ning, Yansong, et al.
Veröffentlicht: (2026)
von: Ning, Yansong, et al.
Veröffentlicht: (2026)
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning
von: Ning, Yansong, et al.
Veröffentlicht: (2025)
von: Ning, Yansong, et al.
Veröffentlicht: (2025)
SafeThinker: Reasoning about Risk to Deepen Safety Beyond Shallow Alignment
von: Fang, Xianya, et al.
Veröffentlicht: (2026)
von: Fang, Xianya, et al.
Veröffentlicht: (2026)
PACE: Prefix-Protected and Difficulty-Aware Compression for Efficient Reasoning
von: Feng, Ruixiang, et al.
Veröffentlicht: (2026)
von: Feng, Ruixiang, et al.
Veröffentlicht: (2026)
Think Smart, Not Hard: Difficulty Adaptive Reasoning for Large Audio Language Models
von: Sheng, Zhichao, et al.
Veröffentlicht: (2025)
von: Sheng, Zhichao, et al.
Veröffentlicht: (2025)
Bag of Tricks for Inference-time Computation of LLM Reasoning
von: Liu, Fan, et al.
Veröffentlicht: (2025)
von: Liu, Fan, et al.
Veröffentlicht: (2025)
Easy2Hard-Bench: Standardized Difficulty Labels for Profiling LLM Performance and Generalization
von: Ding, Mucong, et al.
Veröffentlicht: (2024)
von: Ding, Mucong, et al.
Veröffentlicht: (2024)
Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL
von: Luo, Qin-Wen, et al.
Veröffentlicht: (2025)
von: Luo, Qin-Wen, et al.
Veröffentlicht: (2025)
Efficient Code LLM Training via Distribution-Consistent and Diversity-Aware Data Selection
von: Lyu, Weijie, et al.
Veröffentlicht: (2025)
von: Lyu, Weijie, et al.
Veröffentlicht: (2025)
Relative Difficulty Distillation for Semantic Segmentation
von: Liang, Dong, et al.
Veröffentlicht: (2024)
von: Liang, Dong, et al.
Veröffentlicht: (2024)
DiMA: An LLM-Powered Ride-Hailing Assistant at DiDi
von: Ning, Yansong, et al.
Veröffentlicht: (2025)
von: Ning, Yansong, et al.
Veröffentlicht: (2025)
Reasoning Aware Self-Consistency: Leveraging Reasoning Paths for Efficient LLM Sampling
von: Wan, Guangya, et al.
Veröffentlicht: (2024)
von: Wan, Guangya, et al.
Veröffentlicht: (2024)
RIMO: An Easy-to-Evaluate, Hard-to-Solve Olympiad Benchmark for Advanced Mathematical Reasoning
von: Chen, Ziye, et al.
Veröffentlicht: (2025)
von: Chen, Ziye, et al.
Veröffentlicht: (2025)
DiffAdapt: Difficulty-Adaptive Reasoning for Token-Efficient LLM Inference
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
CameraBench: Benchmarking Visual Reasoning in MLLMs via Photography
von: Fang, I-Sheng, et al.
Veröffentlicht: (2025)
von: Fang, I-Sheng, et al.
Veröffentlicht: (2025)
HardSATGEN: Understanding the Difficulty of Hard SAT Formula Generation and A Strong Structure-Hardness-Aware Baseline
von: Li, Yang, et al.
Veröffentlicht: (2023)
von: Li, Yang, et al.
Veröffentlicht: (2023)
Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization
von: Hua, Xingyuan, et al.
Veröffentlicht: (2026)
von: Hua, Xingyuan, et al.
Veröffentlicht: (2026)
R1-Compress: Long Chain-of-Thought Compression via Chunk Compression and Search
von: Wang, Yibo, et al.
Veröffentlicht: (2025)
von: Wang, Yibo, et al.
Veröffentlicht: (2025)
DeepTravel: An End-to-End Agentic Reinforcement Learning Framework for Autonomous Travel Planning Agents
von: Ning, Yansong, et al.
Veröffentlicht: (2025)
von: Ning, Yansong, et al.
Veröffentlicht: (2025)
ARES: Multimodal Adaptive Reasoning via Difficulty-Aware Token-Level Entropy Shaping
von: Chen, Shuang, et al.
Veröffentlicht: (2025)
von: Chen, Shuang, et al.
Veröffentlicht: (2025)
Tailored Teaching with Balanced Difficulty: Elevating Reasoning in Multimodal Chain-of-Thought via Prompt Curriculum
von: Yang, Xinglong, et al.
Veröffentlicht: (2025)
von: Yang, Xinglong, et al.
Veröffentlicht: (2025)
Fast on the Easy, Deep on the Hard: Efficient Reasoning via Powered Length Penalty
von: Ling, Zehui, et al.
Veröffentlicht: (2025)
von: Ling, Zehui, et al.
Veröffentlicht: (2025)
Beyond Superficial Unlearning: Sharpness-Aware Robust Erasure of Hallucinations in Multimodal LLMs
von: Fang, Xianya, et al.
Veröffentlicht: (2026)
von: Fang, Xianya, et al.
Veröffentlicht: (2026)
FourierCompress: Layer-Aware Spectral Activation Compression for Efficient and Accurate Collaborative LLM Inference
von: Ma, Jian, et al.
Veröffentlicht: (2025)
von: Ma, Jian, et al.
Veröffentlicht: (2025)
Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM Reasoning
von: Parashar, Shubham, et al.
Veröffentlicht: (2025)
von: Parashar, Shubham, et al.
Veröffentlicht: (2025)
PEAR: Phase Entropy Aware Reward for Efficient Reasoning
von: Huang, Chen, et al.
Veröffentlicht: (2025)
von: Huang, Chen, et al.
Veröffentlicht: (2025)
Reasoning with OmniThought: A Large CoT Dataset with Verbosity and Cognitive Difficulty Annotations
von: Cai, Wenrui, et al.
Veröffentlicht: (2025)
von: Cai, Wenrui, et al.
Veröffentlicht: (2025)
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL
von: Luo, Qin-Wen, et al.
Veröffentlicht: (2024)
von: Luo, Qin-Wen, et al.
Veröffentlicht: (2024)
From Easy to Hard: The MIR Benchmark for Progressive Interleaved Multi-Image Reasoning
von: Du, Hang, et al.
Veröffentlicht: (2025)
von: Du, Hang, et al.
Veröffentlicht: (2025)
D3: Diversity, Difficulty, and Dependability-Aware Data Selection for Sample-Efficient LLM Instruction Tuning
von: Zhang, Jia, et al.
Veröffentlicht: (2025)
von: Zhang, Jia, et al.
Veröffentlicht: (2025)
KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
von: Liu, Zedong, et al.
Veröffentlicht: (2026)
von: Liu, Zedong, et al.
Veröffentlicht: (2026)
Beyond Easy Wins: A Text Hardness-Aware Benchmark for LLM-generated Text Detection
von: Ayoobi, Navid, et al.
Veröffentlicht: (2025)
von: Ayoobi, Navid, et al.
Veröffentlicht: (2025)
Enhancing Multi-Modal LLMs Reasoning via Difficulty-Aware Group Normalization
von: Li, Jinghan, et al.
Veröffentlicht: (2026)
von: Li, Jinghan, et al.
Veröffentlicht: (2026)
Molecular mechanism analyses of post‐traumatic epilepsy and hereditary epilepsy based on 10× single‐cell transcriptome sequencing technology
von: Fang Wen, et al.
Veröffentlicht: (2024)
von: Fang Wen, et al.
Veröffentlicht: (2024)
Rethinking Easy-to-Hard: Limits of Curriculum Learning in Post-Training for Deductive Reasoning
von: Mordig, Maximilian, et al.
Veröffentlicht: (2026)
von: Mordig, Maximilian, et al.
Veröffentlicht: (2026)
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning
von: Li, Ang, et al.
Veröffentlicht: (2025)
von: Li, Ang, et al.
Veröffentlicht: (2025)
The LLM Already Knows: Estimating LLM-Perceived Question Difficulty via Hidden Representations
von: Zhu, Yubo, et al.
Veröffentlicht: (2025)
von: Zhu, Yubo, et al.
Veröffentlicht: (2025)
DAST: Difficulty-Aware Self-Training on Large Language Models
von: Xue, Boyang, et al.
Veröffentlicht: (2025)
von: Xue, Boyang, et al.
Veröffentlicht: (2025)
Data-efficient LLM Fine-tuning for Code Generation
von: Lv, Weijie, et al.
Veröffentlicht: (2025)
von: Lv, Weijie, et al.
Veröffentlicht: (2025)
Ada-R1: Hybrid-CoT via Bi-Level Adaptive Reasoning Optimization
von: Luo, Haotian, et al.
Veröffentlicht: (2025)
von: Luo, Haotian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Agent-Omit: Adaptive Context Omission for Efficient LLM Agents
von: Ning, Yansong, et al.
Veröffentlicht: (2026) -
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning
von: Ning, Yansong, et al.
Veröffentlicht: (2025) -
SafeThinker: Reasoning about Risk to Deepen Safety Beyond Shallow Alignment
von: Fang, Xianya, et al.
Veröffentlicht: (2026) -
PACE: Prefix-Protected and Difficulty-Aware Compression for Efficient Reasoning
von: Feng, Ruixiang, et al.
Veröffentlicht: (2026) -
Think Smart, Not Hard: Difficulty Adaptive Reasoning for Large Audio Language Models
von: Sheng, Zhichao, et al.
Veröffentlicht: (2025)