DiffAdapt: Difficulty-Adaptive Reasoning for Token-Efficient LLM Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Xiang, Hu, Xuming, Chu, Xiaowen, Choi, Eunsol |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
by: Liu, Xiang, et al.
Published: (2025)
by: Liu, Xiang, et al.
Published: (2025)
FlowKV: Enhancing Multi-Turn Conversational Coherence in LLMs via Isolated Key-Value Cache Management
by: Liu, Xiang, et al.
Published: (2025)
by: Liu, Xiang, et al.
Published: (2025)
LongGenBench: Long-context Generation Benchmark
by: Liu, Xiang, et al.
Published: (2024)
by: Liu, Xiang, et al.
Published: (2024)
Improving LLM-as-a-Judge Inference with the Judgment Distribution
by: Wang, Victor, et al.
Published: (2025)
by: Wang, Victor, et al.
Published: (2025)
Learning to Reason Across Parallel Samples for LLM Reasoning
by: Qi, Jianing, et al.
Published: (2025)
by: Qi, Jianing, et al.
Published: (2025)
ARES: Multimodal Adaptive Reasoning via Difficulty-Aware Token-Level Entropy Shaping
by: Chen, Shuang, et al.
Published: (2025)
by: Chen, Shuang, et al.
Published: (2025)
AmbigDocs: Reasoning across Documents on Different Entities under the Same Name
by: Lee, Yoonsang, et al.
Published: (2024)
by: Lee, Yoonsang, et al.
Published: (2024)
SONIC: Segmented Optimized Nexus for Information Compression in Key-Value Caching
by: Chen, Hong, et al.
Published: (2026)
by: Chen, Hong, et al.
Published: (2026)
User Feedback in Human-LLM Dialogues: A Lens to Understand Users But Noisy as a Learning Signal
by: Liu, Yuhan, et al.
Published: (2025)
by: Liu, Yuhan, et al.
Published: (2025)
Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression
by: Liu, Xiang, et al.
Published: (2025)
by: Liu, Xiang, et al.
Published: (2025)
To Diff or Not to Diff? Structure-Aware and Adaptive Output Formats for Efficient LLM-based Code Editing
by: Cheng, Wei, et al.
Published: (2026)
by: Cheng, Wei, et al.
Published: (2026)
SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning
by: Li, Zheng, et al.
Published: (2025)
by: Li, Zheng, et al.
Published: (2025)
Less is More Tokens: Efficient Math Reasoning via Difficulty-Aware Chain-of-Thought Distillation
by: Waheed, Abdul, et al.
Published: (2025)
by: Waheed, Abdul, et al.
Published: (2025)
The Multi-Round Diagnostic RAG Framework for Emulating Clinical Reasoning
by: Sun, Penglei, et al.
Published: (2025)
by: Sun, Penglei, et al.
Published: (2025)
SDSAT: Accelerating LLM Inference through Speculative Decoding with Semantic Adaptive Tokens
by: Liu, Chengbo, et al.
Published: (2024)
by: Liu, Chengbo, et al.
Published: (2024)
Make Every Penny Count: Difficulty-Adaptive Self-Consistency for Cost-Efficient Reasoning
by: Wang, Xinglin, et al.
Published: (2024)
by: Wang, Xinglin, et al.
Published: (2024)
ResAdapt: Adaptive Resolution for Efficient Multimodal Reasoning
by: Liao, Huanxuan, et al.
Published: (2026)
by: Liao, Huanxuan, et al.
Published: (2026)
Are Dilemmas and Conflicts in LLM Alignment Solvable? A View from Priority Graph
by: Tang, Zhenheng, et al.
Published: (2026)
by: Tang, Zhenheng, et al.
Published: (2026)
Efficient Vision-Language Reasoning via Adaptive Token Pruning
by: Li, Xue, et al.
Published: (2025)
by: Li, Xue, et al.
Published: (2025)
Beyond Single Embeddings: Capturing Diverse Targets with Multi-Query Retrieval
by: Chen, Hung-Ting, et al.
Published: (2025)
by: Chen, Hung-Ting, et al.
Published: (2025)
CODA: Difficulty-Aware Compute Allocation for Adaptive Reasoning
by: Wu, Siye, et al.
Published: (2026)
by: Wu, Siye, et al.
Published: (2026)
Bridging Latent Reasoning and Target-Language Generation via Retrieval-Transition Heads
by: Patel, Shaswat, et al.
Published: (2026)
by: Patel, Shaswat, et al.
Published: (2026)
Adaptive Group Policy Optimization: Towards Stable Training and Token-Efficient Reasoning
by: Li, Chen, et al.
Published: (2025)
by: Li, Chen, et al.
Published: (2025)
Mitigating Temporal Misalignment by Discarding Outdated Facts
by: Zhang, Michael J. Q., et al.
Published: (2023)
by: Zhang, Michael J. Q., et al.
Published: (2023)
RefreshKV: Updating Small KV Cache During Long-form Generation
by: Xu, Fangyuan, et al.
Published: (2024)
by: Xu, Fangyuan, et al.
Published: (2024)
RedWhale: An Adapted Korean LLM Through Efficient Continual Pretraining
by: Vo, Anh-Dung, et al.
Published: (2024)
by: Vo, Anh-Dung, et al.
Published: (2024)
Open-World Evaluation for Retrieving Diverse Perspectives
by: Chen, Hung-Ting, et al.
Published: (2024)
by: Chen, Hung-Ting, et al.
Published: (2024)
ExpertFlow: Efficient Mixture-of-Experts Inference via Predictive Expert Caching and Token Scheduling
by: He, Xin, et al.
Published: (2024)
by: He, Xin, et al.
Published: (2024)
InstructDiff: Domain-Adaptive Data Selection via Differential Entropy for Efficient LLM Fine-Tuning
by: Su, Junyou, et al.
Published: (2026)
by: Su, Junyou, et al.
Published: (2026)
HAMburger: Accelerating LLM Inference via Token Smashing
by: Liu, Jingyu, et al.
Published: (2025)
by: Liu, Jingyu, et al.
Published: (2025)
Mask Tokens as Prophet: Fine-Grained Cache Eviction for Efficient dLLM Inference
by: Huang, Jianuo, et al.
Published: (2025)
by: Huang, Jianuo, et al.
Published: (2025)
Adapt to Thrive! Adaptive Power-Mean Policy Optimization for Improved LLM Reasoning
by: Huang, Yiming, et al.
Published: (2026)
by: Huang, Yiming, et al.
Published: (2026)
LLM-Oriented Token-Adaptive Knowledge Distillation
by: Xie, Xurong, et al.
Published: (2025)
by: Xie, Xurong, et al.
Published: (2025)
ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference
by: Liang, Yesheng, et al.
Published: (2025)
by: Liang, Yesheng, et al.
Published: (2025)
PACE: Prefix-Protected and Difficulty-Aware Compression for Efficient Reasoning
by: Feng, Ruixiang, et al.
Published: (2026)
by: Feng, Ruixiang, et al.
Published: (2026)
Not All Tokens Matter: Towards Efficient LLM Reasoning via Token Significance in Reinforcement Learning
by: Liu, Hanbing, et al.
Published: (2025)
by: Liu, Hanbing, et al.
Published: (2025)
From Distributional to Overton Pluralism: Investigating Large Language Model Alignment
by: Lake, Thom, et al.
Published: (2024)
by: Lake, Thom, et al.
Published: (2024)
PropMEND: Hypernetworks for Knowledge Propagation in LLMs
by: Liu, Zeyu Leo, et al.
Published: (2025)
by: Liu, Zeyu Leo, et al.
Published: (2025)
Rhapsody: A Dataset for Highlight Detection in Podcasts
by: Park, Younghan, et al.
Published: (2025)
by: Park, Younghan, et al.
Published: (2025)
On Language Models' Sensitivity to Suspicious Coincidences
by: Padmanabhan, Sriram, et al.
Published: (2025)
by: Padmanabhan, Sriram, et al.
Published: (2025)
Similar Items
-
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
by: Liu, Xiang, et al.
Published: (2025) -
FlowKV: Enhancing Multi-Turn Conversational Coherence in LLMs via Isolated Key-Value Cache Management
by: Liu, Xiang, et al.
Published: (2025) -
LongGenBench: Long-context Generation Benchmark
by: Liu, Xiang, et al.
Published: (2024) -
Improving LLM-as-a-Judge Inference with the Judgment Distribution
by: Wang, Victor, et al.
Published: (2025) -
Learning to Reason Across Parallel Samples for LLM Reasoning
by: Qi, Jianing, et al.
Published: (2025)