$\texttt{SPECS}$: Faster Test-Time Scaling through Speculative Drafts
Fuente:
arXiv
Saved in:
| Main Authors: | Cemri, Mert, Rajaraman, Nived, Tiwari, Rishabh, Liu, Xiaoxuan, Keutzer, Kurt, Stoica, Ion, Ramchandran, Kannan, Beirami, Ahmad, Sun, Ziteng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Toward a Theory of Tokenization in LLMs
by: Rajaraman, Nived, et al.
Published: (2024)
by: Rajaraman, Nived, et al.
Published: (2024)
Statistical Complexity and Optimal Algorithms for Non-linear Ridge Bandits
by: Rajaraman, Nived, et al.
Published: (2023)
by: Rajaraman, Nived, et al.
Published: (2023)
Why Do Multi-Agent LLM Systems Fail?
by: Cemri, Mert, et al.
Published: (2025)
by: Cemri, Mert, et al.
Published: (2025)
Transformers on Markov Data: Constant Depth Suffices
by: Rajaraman, Nived, et al.
Published: (2024)
by: Rajaraman, Nived, et al.
Published: (2024)
Scaling Test-Time Compute Without Verification or RL is Suboptimal
by: Setlur, Amrith, et al.
Published: (2025)
by: Setlur, Amrith, et al.
Published: (2025)
S*: Test Time Scaling for Code Generation
by: Li, Dacheng, et al.
Published: (2025)
by: Li, Dacheng, et al.
Published: (2025)
Speculative Decoding: Performance or Illusion?
by: Liu, Xiaoxuan, et al.
Published: (2025)
by: Liu, Xiaoxuan, et al.
Published: (2025)
From Markov to Laplace: How Mamba In-Context Learns Markov Chains
by: Bondaschi, Marco, et al.
Published: (2025)
by: Bondaschi, Marco, et al.
Published: (2025)
Online Speculative Decoding
by: Liu, Xiaoxuan, et al.
Published: (2023)
by: Liu, Xiaoxuan, et al.
Published: (2023)
Interactive Learning of Single-Index Models via Stochastic Gradient Descent
by: Rajaraman, Nived, et al.
Published: (2026)
by: Rajaraman, Nived, et al.
Published: (2026)
SpecTr: Fast Speculative Decoding via Optimal Transport
by: Sun, Ziteng, et al.
Published: (2023)
by: Sun, Ziteng, et al.
Published: (2023)
Block Verification Accelerates Speculative Decoding
by: Sun, Ziteng, et al.
Published: (2024)
by: Sun, Ziteng, et al.
Published: (2024)
Cascade Speculative Drafting for Even Faster LLM Inference
by: Chen, Ziyi, et al.
Published: (2023)
by: Chen, Ziyi, et al.
Published: (2023)
Revisiting Cache Freshness for Emerging Real-Time Applications
by: Mao, Ziming, et al.
Published: (2024)
by: Mao, Ziming, et al.
Published: (2024)
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
by: Tiwari, Rishabh, et al.
Published: (2025)
by: Tiwari, Rishabh, et al.
Published: (2025)
Computational Intractability of Strategizing against Online Learners
by: Assos, Angelos, et al.
Published: (2025)
by: Assos, Angelos, et al.
Published: (2025)
Ouroboros: Generating Longer Drafts Phrase by Phrase for Faster Speculative Decoding
by: Zhao, Weilin, et al.
Published: (2024)
by: Zhao, Weilin, et al.
Published: (2024)
Balancing Coverage and Draft Latency in Vocabulary Trimming for Faster Speculative Decoding
by: Shoham, Ofir Ben
Published: (2026)
by: Shoham, Ofir Ben
Published: (2026)
The Time is Here for Just-in-Time Systems: Challenges and Opportunities
by: Liu, Shu, et al.
Published: (2026)
by: Liu, Shu, et al.
Published: (2026)
Arbitrage: Efficient Reasoning via Advantage-Aware Speculation
by: Maheswaran, Monishwaran, et al.
Published: (2025)
by: Maheswaran, Monishwaran, et al.
Published: (2025)
ODE$_t$(ODE$_l$): Shortcutting the Time and the Length in Diffusion and Flow Models for Faster Sampling
by: Gudovskiy, Denis, et al.
Published: (2025)
by: Gudovskiy, Denis, et al.
Published: (2025)
Some Present-Day Problems of Romanian Library Science
by: Stoica, Ion
Published: (1973)
by: Stoica, Ion
Published: (1973)
The Central University Library, Bucharest. Over Seventy-five Years in the History of a Collection
by: Stoica, Ion
Published: (1972)
by: Stoica, Ion
Published: (1972)
TABED: Test-Time Adaptive Ensemble Drafting for Robust Speculative Decoding in LVLMs
by: Lee, Minjae, et al.
Published: (2026)
by: Lee, Minjae, et al.
Published: (2026)
MITRA: A Large-Scale Parallel Corpus and Multilingual Pretrained Language Model for Machine Translation and Semantic Retrieval for Pāli, Sanskrit, Buddhist Chinese, and Tibetan
by: Nehrdich, Sebastian, et al.
Published: (2026)
by: Nehrdich, Sebastian, et al.
Published: (2026)
Agentic Test-Time Scaling for WebAgents
by: Lee, Nicholas, et al.
Published: (2026)
by: Lee, Nicholas, et al.
Published: (2026)
The Fair Value of Data Under Heterogeneous Privacy Constraints in Federated Learning
by: Kang, Justin, et al.
Published: (2023)
by: Kang, Justin, et al.
Published: (2023)
Draft-OPD: On-Policy Distillation for Speculative Draft Models
by: Lei, Haodi, et al.
Published: (2026)
by: Lei, Haodi, et al.
Published: (2026)
Speculative RAG: Enhancing Retrieval Augmented Generation through Drafting
by: Wang, Zilong, et al.
Published: (2024)
by: Wang, Zilong, et al.
Published: (2024)
Learned Best-Effort LLM Serving
by: Jha, Siddharth, et al.
Published: (2024)
by: Jha, Siddharth, et al.
Published: (2024)
EvoX: Meta-Evolution for Automated Discovery
by: Liu, Shu, et al.
Published: (2026)
by: Liu, Shu, et al.
Published: (2026)
Speculative Interaction Agents: Building Real-Time Agents with Asynchronous I/O and Speculative Tool Calling
by: Hooper, Coleman, et al.
Published: (2026)
by: Hooper, Coleman, et al.
Published: (2026)
HashAttention: Semantic Sparsity for Faster Inference
by: Desai, Aditya, et al.
Published: (2024)
by: Desai, Aditya, et al.
Published: (2024)
SPEED: Speculative Pipelined Execution for Efficient Decoding
by: Hooper, Coleman, et al.
Published: (2023)
by: Hooper, Coleman, et al.
Published: (2023)
Sample Complexity and Representation Ability of Test-time Scaling Paradigms
by: Huang, Baihe, et al.
Published: (2025)
by: Huang, Baihe, et al.
Published: (2025)
SAGE: A Realistic Benchmark for Semantic Understanding
by: Goel, Samarth, et al.
Published: (2025)
by: Goel, Samarth, et al.
Published: (2025)
Quantifying Positional Biases in Text Embedding Models
by: Lee, Reagan J., et al.
Published: (2024)
by: Lee, Reagan J., et al.
Published: (2024)
Asymptotics of Language Model Alignment
by: Yang, Joy Qiping, et al.
Published: (2024)
by: Yang, Joy Qiping, et al.
Published: (2024)
DFlare: Scaling Up Draft Capacity for Block Diffusion Speculative Decoding
by: Zhang, Jiebin, et al.
Published: (2026)
by: Zhang, Jiebin, et al.
Published: (2026)
RoboMonkey: Scaling Test-Time Sampling and Verification for Vision-Language-Action Models
by: Kwok, Jacky, et al.
Published: (2025)
by: Kwok, Jacky, et al.
Published: (2025)
Similar Items
-
Toward a Theory of Tokenization in LLMs
by: Rajaraman, Nived, et al.
Published: (2024) -
Statistical Complexity and Optimal Algorithms for Non-linear Ridge Bandits
by: Rajaraman, Nived, et al.
Published: (2023) -
Why Do Multi-Agent LLM Systems Fail?
by: Cemri, Mert, et al.
Published: (2025) -
Transformers on Markov Data: Constant Depth Suffices
by: Rajaraman, Nived, et al.
Published: (2024) -
Scaling Test-Time Compute Without Verification or RL is Suboptimal
by: Setlur, Amrith, et al.
Published: (2025)