SpecReason: Fast and Accurate Inference-Time Compute via Speculative Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pan, Rui, Dai, Yinwei, Zhang, Zhihao, Oliaro, Gabriele, Jia, Zhihao, Netravali, Ravi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SuffixDecoding: Extreme Speculative Decoding for Emerging AI Applications
von: Oliaro, Gabriele, et al.
Veröffentlicht: (2024)
von: Oliaro, Gabriele, et al.
Veröffentlicht: (2024)
Fail Fast, Win Big: Rethinking the Drafting Strategy in Speculative Decoding via Diffusion LLMs
von: Pan, Rui, et al.
Veröffentlicht: (2025)
von: Pan, Rui, et al.
Veröffentlicht: (2025)
Less Is More: Fast and Accurate Reasoning with Cross-Head Unified Sparse Attention
von: Yang, Lijie, et al.
Veröffentlicht: (2025)
von: Yang, Lijie, et al.
Veröffentlicht: (2025)
Quantized Side Tuning: Fast and Memory-Efficient Tuning of Quantized Large Language Models
von: Zhang, Zhengxin, et al.
Veröffentlicht: (2024)
von: Zhang, Zhengxin, et al.
Veröffentlicht: (2024)
TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention
von: Yang, Lijie, et al.
Veröffentlicht: (2024)
von: Yang, Lijie, et al.
Veröffentlicht: (2024)
FastKernels: Benchmarking GPU Kernel Generation in Production
von: Oliaro, Gabriele, et al.
Veröffentlicht: (2026)
von: Oliaro, Gabriele, et al.
Veröffentlicht: (2026)
SpecExit: Accelerating Large Reasoning Model via Speculative Exit
von: Yang, Rubing, et al.
Veröffentlicht: (2025)
von: Yang, Rubing, et al.
Veröffentlicht: (2025)
Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems
von: Miao, Xupeng, et al.
Veröffentlicht: (2023)
von: Miao, Xupeng, et al.
Veröffentlicht: (2023)
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding
von: Bang, Jehyeon, et al.
Veröffentlicht: (2026)
von: Bang, Jehyeon, et al.
Veröffentlicht: (2026)
BanditSpec: Adaptive Speculative Decoding via Bandit Algorithms
von: Hou, Yunlong, et al.
Veröffentlicht: (2025)
von: Hou, Yunlong, et al.
Veröffentlicht: (2025)
Slipstream: Trajectory-Grounded Compaction Validation for Long-Horizon Agents
von: Chen, Zhuofu, et al.
Veröffentlicht: (2026)
von: Chen, Zhuofu, et al.
Veröffentlicht: (2026)
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
von: Li, Zikun, et al.
Veröffentlicht: (2025)
von: Li, Zikun, et al.
Veröffentlicht: (2025)
CAS-Spec: Cascade Adaptive Self-Speculative Decoding for On-the-Fly Lossless Inference Acceleration of LLMs
von: Ning, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Ning, Zhiyuan, et al.
Veröffentlicht: (2025)
SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification
von: Miao, Xupeng, et al.
Veröffentlicht: (2023)
von: Miao, Xupeng, et al.
Veröffentlicht: (2023)
Geometry Guided Self-Consistency for Physical AI
von: Dai, Yinwei, et al.
Veröffentlicht: (2026)
von: Dai, Yinwei, et al.
Veröffentlicht: (2026)
Accelerating Large-Scale Reasoning Model Inference with Sparse Self-Speculative Decoding
von: Zhao, Yilong, et al.
Veröffentlicht: (2025)
von: Zhao, Yilong, et al.
Veröffentlicht: (2025)
Apparate: Rethinking Early Exits to Tame Latency-Throughput Tensions in ML Serving
von: Dai, Yinwei, et al.
Veröffentlicht: (2023)
von: Dai, Yinwei, et al.
Veröffentlicht: (2023)
DistillSpec: Improving Speculative Decoding via Knowledge Distillation
von: Zhou, Yongchao, et al.
Veröffentlicht: (2023)
von: Zhou, Yongchao, et al.
Veröffentlicht: (2023)
SpecMD: A Comprehensive Study On Speculative Expert Prefetching
von: Hoang, Duc, et al.
Veröffentlicht: (2026)
von: Hoang, Duc, et al.
Veröffentlicht: (2026)
FR-Spec: Accelerating Large-Vocabulary Language Models via Frequency-Ranked Speculative Sampling
von: Zhao, Weilin, et al.
Veröffentlicht: (2025)
von: Zhao, Weilin, et al.
Veröffentlicht: (2025)
Inference-Time Computations for LLM Reasoning and Planning: A Benchmark and Insights
von: Parashar, Shubham, et al.
Veröffentlicht: (2025)
von: Parashar, Shubham, et al.
Veröffentlicht: (2025)
A Survey of Test-Time Compute: From Intuitive Inference to Deliberate Reasoning
von: Ji, Yixin, et al.
Veröffentlicht: (2025)
von: Ji, Yixin, et al.
Veröffentlicht: (2025)
Fractional Reasoning via Latent Steering Vectors Improves Inference Time Compute
von: Liu, Sheng, et al.
Veröffentlicht: (2025)
von: Liu, Sheng, et al.
Veröffentlicht: (2025)
HiSpec: Hierarchical Speculative Decoding for LLMs
von: Kumar, Avinash, et al.
Veröffentlicht: (2025)
von: Kumar, Avinash, et al.
Veröffentlicht: (2025)
Skim: Speculative Execution for Fast and Efficient Web Agents
von: Wong, Mike, et al.
Veröffentlicht: (2026)
von: Wong, Mike, et al.
Veröffentlicht: (2026)
SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths
von: Huang, Kaixuan, et al.
Veröffentlicht: (2024)
von: Huang, Kaixuan, et al.
Veröffentlicht: (2024)
KnapSpec: Self-Speculative Decoding via Adaptive Layer Selection as a Knapsack Problem
von: Cha, Seongjin, et al.
Veröffentlicht: (2026)
von: Cha, Seongjin, et al.
Veröffentlicht: (2026)
Accurate and Scalable Graph Neural Networks via Message Invariance
von: Shi, Zhihao, et al.
Veröffentlicht: (2025)
von: Shi, Zhihao, et al.
Veröffentlicht: (2025)
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2025)
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2025)
FastVLM: Self-Speculative Decoding for Fast Vision-Language Model Inference
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2025)
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2025)
TS-Reasoner: Domain-Oriented Time Series Inference Agents for Reasoning and Automated Analysis
von: Ye, Wen, et al.
Veröffentlicht: (2024)
von: Ye, Wen, et al.
Veröffentlicht: (2024)
Arbitrage: Efficient Reasoning via Advantage-Aware Speculation
von: Maheswaran, Monishwaran, et al.
Veröffentlicht: (2025)
von: Maheswaran, Monishwaran, et al.
Veröffentlicht: (2025)
Spec-VLA: Speculative Decoding for Vision-Language-Action Models with Relaxed Acceptance
von: Wang, Songsheng, et al.
Veröffentlicht: (2025)
von: Wang, Songsheng, et al.
Veröffentlicht: (2025)
Clover-2: Accurate Inference for Regressive Lightweight Speculative Decoding
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
Legilimens: Performant Video Analytics on the System-on-Chip Edge
von: Ramanujam, Murali, et al.
Veröffentlicht: (2025)
von: Ramanujam, Murali, et al.
Veröffentlicht: (2025)
SpecForge: A Flexible and Efficient Open-Source Training Framework for Speculative Decoding
von: Li, Shenggui, et al.
Veröffentlicht: (2026)
von: Li, Shenggui, et al.
Veröffentlicht: (2026)
Improving LLM Reasoning through Scaling Inference Computation with Collaborative Verification
von: Liang, Zhenwen, et al.
Veröffentlicht: (2024)
von: Liang, Zhenwen, et al.
Veröffentlicht: (2024)
SpecMemo: Speculative Decoding is in Your Pocket
von: Yildirim, Selin, et al.
Veröffentlicht: (2025)
von: Yildirim, Selin, et al.
Veröffentlicht: (2025)
Experience-Guided Adaptation of Inference-Time Reasoning Strategies
von: Stein, Adam, et al.
Veröffentlicht: (2025)
von: Stein, Adam, et al.
Veröffentlicht: (2025)
LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification
von: Yang, Penghui, et al.
Veröffentlicht: (2025)
von: Yang, Penghui, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SuffixDecoding: Extreme Speculative Decoding for Emerging AI Applications
von: Oliaro, Gabriele, et al.
Veröffentlicht: (2024) -
Fail Fast, Win Big: Rethinking the Drafting Strategy in Speculative Decoding via Diffusion LLMs
von: Pan, Rui, et al.
Veröffentlicht: (2025) -
Less Is More: Fast and Accurate Reasoning with Cross-Head Unified Sparse Attention
von: Yang, Lijie, et al.
Veröffentlicht: (2025) -
Quantized Side Tuning: Fast and Memory-Efficient Tuning of Quantized Large Language Models
von: Zhang, Zhengxin, et al.
Veröffentlicht: (2024) -
TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention
von: Yang, Lijie, et al.
Veröffentlicht: (2024)