DiffAdapt: Difficulty-Adaptive Reasoning for Token-Efficient LLM Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Xiang, Hu, Xuming, Chu, Xiaowen, Choi, Eunsol |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
FlowKV: Enhancing Multi-Turn Conversational Coherence in LLMs via Isolated Key-Value Cache Management
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
LongGenBench: Long-context Generation Benchmark
von: Liu, Xiang, et al.
Veröffentlicht: (2024)
von: Liu, Xiang, et al.
Veröffentlicht: (2024)
Improving LLM-as-a-Judge Inference with the Judgment Distribution
von: Wang, Victor, et al.
Veröffentlicht: (2025)
von: Wang, Victor, et al.
Veröffentlicht: (2025)
Learning to Reason Across Parallel Samples for LLM Reasoning
von: Qi, Jianing, et al.
Veröffentlicht: (2025)
von: Qi, Jianing, et al.
Veröffentlicht: (2025)
ARES: Multimodal Adaptive Reasoning via Difficulty-Aware Token-Level Entropy Shaping
von: Chen, Shuang, et al.
Veröffentlicht: (2025)
von: Chen, Shuang, et al.
Veröffentlicht: (2025)
AmbigDocs: Reasoning across Documents on Different Entities under the Same Name
von: Lee, Yoonsang, et al.
Veröffentlicht: (2024)
von: Lee, Yoonsang, et al.
Veröffentlicht: (2024)
SONIC: Segmented Optimized Nexus for Information Compression in Key-Value Caching
von: Chen, Hong, et al.
Veröffentlicht: (2026)
von: Chen, Hong, et al.
Veröffentlicht: (2026)
User Feedback in Human-LLM Dialogues: A Lens to Understand Users But Noisy as a Learning Signal
von: Liu, Yuhan, et al.
Veröffentlicht: (2025)
von: Liu, Yuhan, et al.
Veröffentlicht: (2025)
Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
To Diff or Not to Diff? Structure-Aware and Adaptive Output Formats for Efficient LLM-based Code Editing
von: Cheng, Wei, et al.
Veröffentlicht: (2026)
von: Cheng, Wei, et al.
Veröffentlicht: (2026)
SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning
von: Li, Zheng, et al.
Veröffentlicht: (2025)
von: Li, Zheng, et al.
Veröffentlicht: (2025)
Less is More Tokens: Efficient Math Reasoning via Difficulty-Aware Chain-of-Thought Distillation
von: Waheed, Abdul, et al.
Veröffentlicht: (2025)
von: Waheed, Abdul, et al.
Veröffentlicht: (2025)
The Multi-Round Diagnostic RAG Framework for Emulating Clinical Reasoning
von: Sun, Penglei, et al.
Veröffentlicht: (2025)
von: Sun, Penglei, et al.
Veröffentlicht: (2025)
SDSAT: Accelerating LLM Inference through Speculative Decoding with Semantic Adaptive Tokens
von: Liu, Chengbo, et al.
Veröffentlicht: (2024)
von: Liu, Chengbo, et al.
Veröffentlicht: (2024)
Make Every Penny Count: Difficulty-Adaptive Self-Consistency for Cost-Efficient Reasoning
von: Wang, Xinglin, et al.
Veröffentlicht: (2024)
von: Wang, Xinglin, et al.
Veröffentlicht: (2024)
ResAdapt: Adaptive Resolution for Efficient Multimodal Reasoning
von: Liao, Huanxuan, et al.
Veröffentlicht: (2026)
von: Liao, Huanxuan, et al.
Veröffentlicht: (2026)
Are Dilemmas and Conflicts in LLM Alignment Solvable? A View from Priority Graph
von: Tang, Zhenheng, et al.
Veröffentlicht: (2026)
von: Tang, Zhenheng, et al.
Veröffentlicht: (2026)
Efficient Vision-Language Reasoning via Adaptive Token Pruning
von: Li, Xue, et al.
Veröffentlicht: (2025)
von: Li, Xue, et al.
Veröffentlicht: (2025)
Beyond Single Embeddings: Capturing Diverse Targets with Multi-Query Retrieval
von: Chen, Hung-Ting, et al.
Veröffentlicht: (2025)
von: Chen, Hung-Ting, et al.
Veröffentlicht: (2025)
CODA: Difficulty-Aware Compute Allocation for Adaptive Reasoning
von: Wu, Siye, et al.
Veröffentlicht: (2026)
von: Wu, Siye, et al.
Veröffentlicht: (2026)
Bridging Latent Reasoning and Target-Language Generation via Retrieval-Transition Heads
von: Patel, Shaswat, et al.
Veröffentlicht: (2026)
von: Patel, Shaswat, et al.
Veröffentlicht: (2026)
Adaptive Group Policy Optimization: Towards Stable Training and Token-Efficient Reasoning
von: Li, Chen, et al.
Veröffentlicht: (2025)
von: Li, Chen, et al.
Veröffentlicht: (2025)
Mitigating Temporal Misalignment by Discarding Outdated Facts
von: Zhang, Michael J. Q., et al.
Veröffentlicht: (2023)
von: Zhang, Michael J. Q., et al.
Veröffentlicht: (2023)
RefreshKV: Updating Small KV Cache During Long-form Generation
von: Xu, Fangyuan, et al.
Veröffentlicht: (2024)
von: Xu, Fangyuan, et al.
Veröffentlicht: (2024)
RedWhale: An Adapted Korean LLM Through Efficient Continual Pretraining
von: Vo, Anh-Dung, et al.
Veröffentlicht: (2024)
von: Vo, Anh-Dung, et al.
Veröffentlicht: (2024)
Open-World Evaluation for Retrieving Diverse Perspectives
von: Chen, Hung-Ting, et al.
Veröffentlicht: (2024)
von: Chen, Hung-Ting, et al.
Veröffentlicht: (2024)
ExpertFlow: Efficient Mixture-of-Experts Inference via Predictive Expert Caching and Token Scheduling
von: He, Xin, et al.
Veröffentlicht: (2024)
von: He, Xin, et al.
Veröffentlicht: (2024)
InstructDiff: Domain-Adaptive Data Selection via Differential Entropy for Efficient LLM Fine-Tuning
von: Su, Junyou, et al.
Veröffentlicht: (2026)
von: Su, Junyou, et al.
Veröffentlicht: (2026)
HAMburger: Accelerating LLM Inference via Token Smashing
von: Liu, Jingyu, et al.
Veröffentlicht: (2025)
von: Liu, Jingyu, et al.
Veröffentlicht: (2025)
Mask Tokens as Prophet: Fine-Grained Cache Eviction for Efficient dLLM Inference
von: Huang, Jianuo, et al.
Veröffentlicht: (2025)
von: Huang, Jianuo, et al.
Veröffentlicht: (2025)
Adapt to Thrive! Adaptive Power-Mean Policy Optimization for Improved LLM Reasoning
von: Huang, Yiming, et al.
Veröffentlicht: (2026)
von: Huang, Yiming, et al.
Veröffentlicht: (2026)
LLM-Oriented Token-Adaptive Knowledge Distillation
von: Xie, Xurong, et al.
Veröffentlicht: (2025)
von: Xie, Xurong, et al.
Veröffentlicht: (2025)
ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference
von: Liang, Yesheng, et al.
Veröffentlicht: (2025)
von: Liang, Yesheng, et al.
Veröffentlicht: (2025)
PACE: Prefix-Protected and Difficulty-Aware Compression for Efficient Reasoning
von: Feng, Ruixiang, et al.
Veröffentlicht: (2026)
von: Feng, Ruixiang, et al.
Veröffentlicht: (2026)
Not All Tokens Matter: Towards Efficient LLM Reasoning via Token Significance in Reinforcement Learning
von: Liu, Hanbing, et al.
Veröffentlicht: (2025)
von: Liu, Hanbing, et al.
Veröffentlicht: (2025)
From Distributional to Overton Pluralism: Investigating Large Language Model Alignment
von: Lake, Thom, et al.
Veröffentlicht: (2024)
von: Lake, Thom, et al.
Veröffentlicht: (2024)
PropMEND: Hypernetworks for Knowledge Propagation in LLMs
von: Liu, Zeyu Leo, et al.
Veröffentlicht: (2025)
von: Liu, Zeyu Leo, et al.
Veröffentlicht: (2025)
Rhapsody: A Dataset for Highlight Detection in Podcasts
von: Park, Younghan, et al.
Veröffentlicht: (2025)
von: Park, Younghan, et al.
Veröffentlicht: (2025)
On Language Models' Sensitivity to Suspicious Coincidences
von: Padmanabhan, Sriram, et al.
Veröffentlicht: (2025)
von: Padmanabhan, Sriram, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
von: Liu, Xiang, et al.
Veröffentlicht: (2025) -
FlowKV: Enhancing Multi-Turn Conversational Coherence in LLMs via Isolated Key-Value Cache Management
von: Liu, Xiang, et al.
Veröffentlicht: (2025) -
LongGenBench: Long-context Generation Benchmark
von: Liu, Xiang, et al.
Veröffentlicht: (2024) -
Improving LLM-as-a-Judge Inference with the Judgment Distribution
von: Wang, Victor, et al.
Veröffentlicht: (2025) -
Learning to Reason Across Parallel Samples for LLM Reasoning
von: Qi, Jianing, et al.
Veröffentlicht: (2025)