Scaling LLM Inference with Optimized Sample Compute Allocation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhang, Kexun, Zhou, Shang, Wang, Danqing, Wang, William Yang, Li, Lei |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Don't Fine-Tune, Decode: Syntax Error-Free Tool Use via Constrained Decoding
par: Zhang, Kexun, et autres
Publié: (2023)
par: Zhang, Kexun, et autres
Publié: (2023)
DeFT: Decoding with Flash Tree-attention for Efficient Tree-structured LLM Inference
par: Yao, Jinwei, et autres
Publié: (2024)
par: Yao, Jinwei, et autres
Publié: (2024)
Human Bias in the Face of AI: Examining Human Judgment Against Text Labeled as AI Generated
par: Zhu, Tiffany, et autres
Publié: (2024)
par: Zhu, Tiffany, et autres
Publié: (2024)
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
par: Feng, Yuan, et autres
Publié: (2024)
par: Feng, Yuan, et autres
Publié: (2024)
Cooperative Strategic Planning Enhances Reasoning Capabilities in Large Language Models
par: Wang, Danqing, et autres
Publié: (2024)
par: Wang, Danqing, et autres
Publié: (2024)
Adaptive Rectification Sampling for Test-Time Compute Scaling
par: Tan, Zhendong, et autres
Publié: (2025)
par: Tan, Zhendong, et autres
Publié: (2025)
Understanding Reasoning Ability of Language Models From the Perspective of Reasoning Paths Aggregation
par: Wang, Xinyi, et autres
Publié: (2024)
par: Wang, Xinyi, et autres
Publié: (2024)
Thinking Slow, Fast: Scaling Inference Compute with Distilled Reasoners
par: Paliotta, Daniele, et autres
Publié: (2025)
par: Paliotta, Daniele, et autres
Publié: (2025)
When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMs
par: Khairi, Ammar, et autres
Publié: (2025)
par: Khairi, Ammar, et autres
Publié: (2025)
Self-Resource Allocation in Multi-Agent LLM Systems
par: Amayuelas, Alfonso, et autres
Publié: (2025)
par: Amayuelas, Alfonso, et autres
Publié: (2025)
Sleep-time Compute: Beyond Inference Scaling at Test-time
par: Lin, Kevin, et autres
Publié: (2025)
par: Lin, Kevin, et autres
Publié: (2025)
Speculative Decoding for Multi-Sample Inference
par: Li, Yiwei, et autres
Publié: (2025)
par: Li, Yiwei, et autres
Publié: (2025)
LLM Inference Unveiled: Survey and Roofline Model Insights
par: Yuan, Zhihang, et autres
Publié: (2024)
par: Yuan, Zhihang, et autres
Publié: (2024)
Sampling-Efficient Test-Time Scaling: Self-Estimating the Best-of-N Sampling in Early Decoding
par: Wang, Yiming, et autres
Publié: (2025)
par: Wang, Yiming, et autres
Publié: (2025)
Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data
par: Wang, Xinyi, et autres
Publié: (2024)
par: Wang, Xinyi, et autres
Publié: (2024)
Revealing the Barriers of Language Agents in Planning
par: Xie, Jian, et autres
Publié: (2024)
par: Xie, Jian, et autres
Publié: (2024)
SCALE: Selective Resource Allocation for Overcoming Performance Bottlenecks in Mathematical Test-time Scaling
par: Xiao, Yang, et autres
Publié: (2025)
par: Xiao, Yang, et autres
Publié: (2025)
BPO: Staying Close to the Behavior LLM Creates Better Online LLM Alignment
par: Xu, Wenda, et autres
Publié: (2024)
par: Xu, Wenda, et autres
Publié: (2024)
Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning
par: Bi, Zhenni, et autres
Publié: (2024)
par: Bi, Zhenni, et autres
Publié: (2024)
Examining False Positives under Inference Scaling for Mathematical Reasoning
par: Wang, Yu, et autres
Publié: (2025)
par: Wang, Yu, et autres
Publié: (2025)
A Survey of Frontiers in LLM Reasoning: Inference Scaling, Learning to Reason, and Agentic Systems
par: Ke, Zixuan, et autres
Publié: (2025)
par: Ke, Zixuan, et autres
Publié: (2025)
UniScale: Adaptive Unified Inference Scaling via Online Joint Optimization of Model Routing and Test-Time Scaling
par: Huang, Kaiyu, et autres
Publié: (2026)
par: Huang, Kaiyu, et autres
Publié: (2026)
DynScaling: Efficient Verifier-free Inference Scaling via Dynamic and Integrated Sampling
par: Wang, Fei, et autres
Publié: (2025)
par: Wang, Fei, et autres
Publié: (2025)
Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning
par: Yang, Wenkai, et autres
Publié: (2025)
par: Yang, Wenkai, et autres
Publié: (2025)
Optimizing Temperature for Language Models with Multi-Sample Inference
par: Du, Weihua, et autres
Publié: (2025)
par: Du, Weihua, et autres
Publié: (2025)
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference
par: Ouyang, Haojie, et autres
Publié: (2025)
par: Ouyang, Haojie, et autres
Publié: (2025)
SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning
par: Li, Zheng, et autres
Publié: (2025)
par: Li, Zheng, et autres
Publié: (2025)
Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement
par: Xu, Wenda, et autres
Publié: (2024)
par: Xu, Wenda, et autres
Publié: (2024)
Budgeted LoRA: Distillation as Structured Compute Allocation for Efficient Inference
par: Sabry, Mohammed, et autres
Publié: (2026)
par: Sabry, Mohammed, et autres
Publié: (2026)
Human-Instruction-Free LLM Self-Alignment with Limited Samples
par: Guo, Hongyi, et autres
Publié: (2024)
par: Guo, Hongyi, et autres
Publié: (2024)
CA-SQL: Complexity-Aware Inference Time Reasoning for Text-to-SQL via Exploration and Compute Budget Allocation
par: Petullo, James, et autres
Publié: (2026)
par: Petullo, James, et autres
Publié: (2026)
Chunks as Arms: Multi-Armed Bandit-Guided Sampling for Long-Context LLM Preference Optimization
par: Duan, Shaohua, et autres
Publié: (2025)
par: Duan, Shaohua, et autres
Publié: (2025)
RankAdaptor: Hierarchical Rank Allocation for Efficient Fine-Tuning Pruned LLMs via Performance Model
par: Zhou, Changhai, et autres
Publié: (2024)
par: Zhou, Changhai, et autres
Publié: (2024)
Scaling Textual Gradients via Sampling-Based Momentum
par: Ding, Zixin, et autres
Publié: (2025)
par: Ding, Zixin, et autres
Publié: (2025)
ParaThinker: Native Parallel Thinking as a New Paradigm to Scale LLM Test-time Compute
par: Wen, Hao, et autres
Publié: (2025)
par: Wen, Hao, et autres
Publié: (2025)
LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues
par: Li, Haoyang, et autres
Publié: (2025)
par: Li, Haoyang, et autres
Publié: (2025)
XL3M: A Training-free Framework for LLM Length Extension Based on Segment-wise Inference
par: Wang, Shengnan, et autres
Publié: (2024)
par: Wang, Shengnan, et autres
Publié: (2024)
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
par: Liu, Baihui, et autres
Publié: (2026)
par: Liu, Baihui, et autres
Publié: (2026)
MarkovScale: Towards Optimal Sequential Scaling at Inference Time
par: Wang, Youkang, et autres
Publié: (2026)
par: Wang, Youkang, et autres
Publié: (2026)
Understanding Dynamic Compute Allocation in Recurrent Transformers
par: Moosa, Ibraheem Muhammad, et autres
Publié: (2026)
par: Moosa, Ibraheem Muhammad, et autres
Publié: (2026)
Documents similaires
-
Don't Fine-Tune, Decode: Syntax Error-Free Tool Use via Constrained Decoding
par: Zhang, Kexun, et autres
Publié: (2023) -
DeFT: Decoding with Flash Tree-attention for Efficient Tree-structured LLM Inference
par: Yao, Jinwei, et autres
Publié: (2024) -
Human Bias in the Face of AI: Examining Human Judgment Against Text Labeled as AI Generated
par: Zhu, Tiffany, et autres
Publié: (2024) -
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
par: Feng, Yuan, et autres
Publié: (2024) -
Cooperative Strategic Planning Enhances Reasoning Capabilities in Large Language Models
par: Wang, Danqing, et autres
Publié: (2024)