SSR: Speculative Parallel Scaling Reasoning in Test-time
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chu, Yuanlin, Wang, Bo, Liu, Xiang, Chen, Hong, Liu, Aiwei, Hu, Xuming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Deconstructing Pre-training: Knowledge Attribution Analysis in MoE and Dense Models
von: Wang, Bo, et al.
Veröffentlicht: (2026)
von: Wang, Bo, et al.
Veröffentlicht: (2026)
You only need 4 extra tokens: Synergistic Test-time Adaptation for LLMs
von: Xu, Yijie, et al.
Veröffentlicht: (2025)
von: Xu, Yijie, et al.
Veröffentlicht: (2025)
Parallel Test-Time Scaling for Latent Reasoning Models
von: You, Runyang, et al.
Veröffentlicht: (2025)
von: You, Runyang, et al.
Veröffentlicht: (2025)
SSR: Socratic Self-Refine for Large Language Model Reasoning
von: Shi, Haizhou, et al.
Veröffentlicht: (2025)
von: Shi, Haizhou, et al.
Veröffentlicht: (2025)
Mark Your LLM: Detecting the Misuse of Open-Source Large Language Models via Watermarking
von: Xu, Yijie, et al.
Veröffentlicht: (2025)
von: Xu, Yijie, et al.
Veröffentlicht: (2025)
Scaling Speculative Decoding with Lookahead Reasoning
von: Fu, Yichao, et al.
Veröffentlicht: (2025)
von: Fu, Yichao, et al.
Veröffentlicht: (2025)
PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning
von: Hu, Jingcheng, et al.
Veröffentlicht: (2026)
von: Hu, Jingcheng, et al.
Veröffentlicht: (2026)
TEMPO: Scaling Test-time Training for Large Reasoning Models
von: Zhang, Qingyang, et al.
Veröffentlicht: (2026)
von: Zhang, Qingyang, et al.
Veröffentlicht: (2026)
ParallelSpec: Parallel Drafter for Efficient Speculative Decoding
von: Xiao, Zilin, et al.
Veröffentlicht: (2024)
von: Xiao, Zilin, et al.
Veröffentlicht: (2024)
$\texttt{SPECS}$: Faster Test-Time Scaling through Speculative Drafts
von: Cemri, Mert, et al.
Veröffentlicht: (2025)
von: Cemri, Mert, et al.
Veröffentlicht: (2025)
Gumbel Distillation for Parallel Text Generation
von: Zhang, Chi, et al.
Veröffentlicht: (2026)
von: Zhang, Chi, et al.
Veröffentlicht: (2026)
Test-Time Speculation
von: Kumar, Avinash, et al.
Veröffentlicht: (2026)
von: Kumar, Avinash, et al.
Veröffentlicht: (2026)
Not-a-Bandit: Provably No-Regret Drafter Selection in Speculative Decoding for LLMs
von: Liu, Hongyi, et al.
Veröffentlicht: (2025)
von: Liu, Hongyi, et al.
Veröffentlicht: (2025)
The Virtues of Brevity: Avoid Overthinking in Parallel Test-Time Reasoning
von: Dinardi, Raul Cavalcante, et al.
Veröffentlicht: (2025)
von: Dinardi, Raul Cavalcante, et al.
Veröffentlicht: (2025)
DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation
von: Liu, Zining, et al.
Veröffentlicht: (2026)
von: Liu, Zining, et al.
Veröffentlicht: (2026)
Parallel Scaling Law for Language Models
von: Chen, Mouxiang, et al.
Veröffentlicht: (2025)
von: Chen, Mouxiang, et al.
Veröffentlicht: (2025)
Accelerating Large-Scale Reasoning Model Inference with Sparse Self-Speculative Decoding
von: Zhao, Yilong, et al.
Veröffentlicht: (2025)
von: Zhao, Yilong, et al.
Veröffentlicht: (2025)
RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time
von: Wang, Haozhe, et al.
Veröffentlicht: (2026)
von: Wang, Haozhe, et al.
Veröffentlicht: (2026)
SONIC: Segmented Optimized Nexus for Information Compression in Key-Value Caching
von: Chen, Hong, et al.
Veröffentlicht: (2026)
von: Chen, Hong, et al.
Veröffentlicht: (2026)
SOLARIS: Speculative Offloading of Latent-bAsed Representation for Inference Scaling
von: Liu, Zikun, et al.
Veröffentlicht: (2026)
von: Liu, Zikun, et al.
Veröffentlicht: (2026)
SpecPipe: Accelerating Pipeline Parallelism-based LLM Inference with Speculative Decoding
von: Yin, Haofei, et al.
Veröffentlicht: (2025)
von: Yin, Haofei, et al.
Veröffentlicht: (2025)
SpecExit: Accelerating Large Reasoning Model via Speculative Exit
von: Yang, Rubing, et al.
Veröffentlicht: (2025)
von: Yang, Rubing, et al.
Veröffentlicht: (2025)
Exploring Response Uncertainty in MLLMs: An Empirical Evaluation under Misleading Scenarios
von: Dang, Yunkai, et al.
Veröffentlicht: (2024)
von: Dang, Yunkai, et al.
Veröffentlicht: (2024)
KERV: Kinematic-Rectified Speculative Decoding for Embodied VLA Models
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
Exploring Test-time Scaling via Prediction Merging on Large-Scale Recommendation
von: Lyu, Fuyuan, et al.
Veröffentlicht: (2025)
von: Lyu, Fuyuan, et al.
Veröffentlicht: (2025)
PRISM: Parametrically Refactoring Inference for Speculative Sampling Draft Models
von: Wang, Xuliang, et al.
Veröffentlicht: (2026)
von: Wang, Xuliang, et al.
Veröffentlicht: (2026)
Speculative Speculative Decoding
von: Kumar, Tanishq, et al.
Veröffentlicht: (2026)
von: Kumar, Tanishq, et al.
Veröffentlicht: (2026)
D-PACE: Dynamic Position-Aware Cross-Entropy for Parallel Speculative Drafting
von: Wu, Tianyu, et al.
Veröffentlicht: (2026)
von: Wu, Tianyu, et al.
Veröffentlicht: (2026)
Dynamic Expert Specialization: Towards Catastrophic Forgetting-Free Multi-Domain MoE Adaptation
von: Li, Junzhuo, et al.
Veröffentlicht: (2025)
von: Li, Junzhuo, et al.
Veröffentlicht: (2025)
MDN: Parallelizing Stepwise Momentum for Delta Linear Attention
von: Huang, Yulong, et al.
Veröffentlicht: (2026)
von: Huang, Yulong, et al.
Veröffentlicht: (2026)
EAGLE-Pangu: Accelerator-Safe Tree Speculative Decoding on Ascend NPUs
von: Han, Chang, et al.
Veröffentlicht: (2026)
von: Han, Chang, et al.
Veröffentlicht: (2026)
CARD: A Cache-Assisted Parallel Speculative Decoding Framework via Query-and-Correct Paradigm for Accelerating LLM Inference
von: Zhou, Enyu, et al.
Veröffentlicht: (2025)
von: Zhou, Enyu, et al.
Veröffentlicht: (2025)
Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration
von: Zhong, Shuzhang, et al.
Veröffentlicht: (2026)
von: Zhong, Shuzhang, et al.
Veröffentlicht: (2026)
Decoding Knowledge Attribution in Mixture-of-Experts: A Framework of Basic-Refinement Collaboration and Efficiency Analysis
von: Li, Junzhuo, et al.
Veröffentlicht: (2025)
von: Li, Junzhuo, et al.
Veröffentlicht: (2025)
HeiSD: Hybrid Speculative Decoding for Embodied Vision-Language-Action Models with Kinematic Awareness
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
Online Speculative Decoding
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
TriSpec: Ternary Speculative Decoding via Lightweight Proxy Verification
von: Jiang, Haoyun, et al.
Veröffentlicht: (2026)
von: Jiang, Haoyun, et al.
Veröffentlicht: (2026)
FastTTS: Accelerating Test-Time Scaling for Edge LLM Reasoning
von: Chen, Hao Mark, et al.
Veröffentlicht: (2025)
von: Chen, Hao Mark, et al.
Veröffentlicht: (2025)
Scaling Laws for Many-Shot In-Context Learning with Self-Generated Annotations
von: Gu, Zhengyao, et al.
Veröffentlicht: (2025)
von: Gu, Zhengyao, et al.
Veröffentlicht: (2025)
Reviving Any-Subset Autoregressive Models with Principled Parallel Sampling and Speculative Decoding
von: Guo, Gabe, et al.
Veröffentlicht: (2025)
von: Guo, Gabe, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Deconstructing Pre-training: Knowledge Attribution Analysis in MoE and Dense Models
von: Wang, Bo, et al.
Veröffentlicht: (2026) -
You only need 4 extra tokens: Synergistic Test-time Adaptation for LLMs
von: Xu, Yijie, et al.
Veröffentlicht: (2025) -
Parallel Test-Time Scaling for Latent Reasoning Models
von: You, Runyang, et al.
Veröffentlicht: (2025) -
SSR: Socratic Self-Refine for Large Language Model Reasoning
von: Shi, Haizhou, et al.
Veröffentlicht: (2025) -
Mark Your LLM: Detecting the Misuse of Open-Source Large Language Models via Watermarking
von: Xu, Yijie, et al.
Veröffentlicht: (2025)