SpecBound: Adaptive Bounded Self-Speculation with Layer-wise Confidence Calibration
Fuente:
arXiv
Guardado en:
| Autores principales: | Wen, Zhuofan, Feng, Yang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration
por: Wen, Zhuofan, et al.
Publicado: (2024)
por: Wen, Zhuofan, et al.
Publicado: (2024)
SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths
por: Huang, Kaixuan, et al.
Publicado: (2024)
por: Huang, Kaixuan, et al.
Publicado: (2024)
HiSpec: Hierarchical Speculative Decoding for LLMs
por: Kumar, Avinash, et al.
Publicado: (2025)
por: Kumar, Avinash, et al.
Publicado: (2025)
SpecExit: Accelerating Large Reasoning Model via Speculative Exit
por: Yang, Rubing, et al.
Publicado: (2025)
por: Yang, Rubing, et al.
Publicado: (2025)
DistillSpec: Improving Speculative Decoding via Knowledge Distillation
por: Zhou, Yongchao, et al.
Publicado: (2023)
por: Zhou, Yongchao, et al.
Publicado: (2023)
LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification
por: Yang, Penghui, et al.
Publicado: (2025)
por: Yang, Penghui, et al.
Publicado: (2025)
ML-SpecQD: Multi-Level Speculative Decoding with Quantized Drafts
por: Georganas, Evangelos, et al.
Publicado: (2025)
por: Georganas, Evangelos, et al.
Publicado: (2025)
SpecForge: A Flexible and Efficient Open-Source Training Framework for Speculative Decoding
por: Li, Shenggui, et al.
Publicado: (2026)
por: Li, Shenggui, et al.
Publicado: (2026)
Token-Driven GammaTune: Adaptive Calibration for Enhanced Speculative Decoding
por: Gautam, Aayush, et al.
Publicado: (2025)
por: Gautam, Aayush, et al.
Publicado: (2025)
DynaSpec: Context-aware Dynamic Speculative Sampling for Large-Vocabulary Language Models
por: Zhang, Jinbin, et al.
Publicado: (2025)
por: Zhang, Jinbin, et al.
Publicado: (2025)
Confidence-Modulated Speculative Decoding for Large Language Models
por: Sen, Jaydip, et al.
Publicado: (2025)
por: Sen, Jaydip, et al.
Publicado: (2025)
FR-Spec: Accelerating Large-Vocabulary Language Models via Frequency-Ranked Speculative Sampling
por: Zhao, Weilin, et al.
Publicado: (2025)
por: Zhao, Weilin, et al.
Publicado: (2025)
LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding
por: Elhoushi, Mostafa, et al.
Publicado: (2024)
por: Elhoushi, Mostafa, et al.
Publicado: (2024)
SpecKV: Adaptive Speculative Decoding with Compression-Aware Gamma Selection
por: Shukla, Shikhar
Publicado: (2026)
por: Shukla, Shikhar
Publicado: (2026)
KnapSpec: Self-Speculative Decoding via Adaptive Layer Selection as a Knapsack Problem
por: Cha, Seongjin, et al.
Publicado: (2026)
por: Cha, Seongjin, et al.
Publicado: (2026)
Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable Rewards
por: Ma, Zhengzhao, et al.
Publicado: (2026)
por: Ma, Zhengzhao, et al.
Publicado: (2026)
ConfLayers: Adaptive Confidence-based Layer Skipping for Self-Speculative Decoding
por: Amer, Walaa, et al.
Publicado: (2026)
por: Amer, Walaa, et al.
Publicado: (2026)
Adaptive-Boundary-Clipping GRPO: Ensuring Bounded Ratios for Stable and Generalizable Training
por: Liu, Chi, et al.
Publicado: (2026)
por: Liu, Chi, et al.
Publicado: (2026)
Calibration Across Layers: Understanding Calibration Evolution in LLMs
por: Joshi, Abhinav, et al.
Publicado: (2025)
por: Joshi, Abhinav, et al.
Publicado: (2025)
Confidence-aware Self-Semantic Distillation on Knowledge Graph Embedding
por: Liu, Yichen, et al.
Publicado: (2022)
por: Liu, Yichen, et al.
Publicado: (2022)
MICE for CATs: Model-Internal Confidence Estimation for Calibrating Agents with Tools
por: Subramani, Nishant, et al.
Publicado: (2025)
por: Subramani, Nishant, et al.
Publicado: (2025)
Self-Speculative Biased Decoding for Faster Re-Translation
por: Zeng, Linxiao, et al.
Publicado: (2025)
por: Zeng, Linxiao, et al.
Publicado: (2025)
LEAP: Layer-wise Exit-Aware Pretraining for Efficient Transformer Inference
por: Kapadia, Shashank, et al.
Publicado: (2026)
por: Kapadia, Shashank, et al.
Publicado: (2026)
MaskPrune: Mask-based LLM Pruning for Layer-wise Uniform Structures
por: Qin, Jiayu, et al.
Publicado: (2025)
por: Qin, Jiayu, et al.
Publicado: (2025)
SpecExtend: A Drop-in Enhancement for Speculative Decoding of Long Sequences
por: Cha, Jungyoub, et al.
Publicado: (2025)
por: Cha, Jungyoub, et al.
Publicado: (2025)
Entropy-Adaptive Fine-Tuning: Resolving Confident Conflicts to Mitigate Forgetting
por: Diao, Muxi, et al.
Publicado: (2026)
por: Diao, Muxi, et al.
Publicado: (2026)
Self-Guided Process Reward Optimization with Redefined Step-wise Advantage for Process Reinforcement Learning
por: Fei, Wu, et al.
Publicado: (2025)
por: Fei, Wu, et al.
Publicado: (2025)
CopySpec: Accelerating LLMs with Speculative Copy-and-Paste Without Compromising Quality
por: Dumitru, Razvan-Gabriel, et al.
Publicado: (2025)
por: Dumitru, Razvan-Gabriel, et al.
Publicado: (2025)
ConfSpec: Efficient Step-Level Speculative Reasoning via Confidence-Gated Verification
por: Liu, Siran, et al.
Publicado: (2026)
por: Liu, Siran, et al.
Publicado: (2026)
SLaNC: Static LayerNorm Calibration
por: Salmani, Mahsa, et al.
Publicado: (2024)
por: Salmani, Mahsa, et al.
Publicado: (2024)
BanditSpec: Adaptive Speculative Decoding via Bandit Algorithms
por: Hou, Yunlong, et al.
Publicado: (2025)
por: Hou, Yunlong, et al.
Publicado: (2025)
SaySelf: Teaching LLMs to Express Confidence with Self-Reflective Rationales
por: Xu, Tianyang, et al.
Publicado: (2024)
por: Xu, Tianyang, et al.
Publicado: (2024)
Certified Robustness Under Bounded Levenshtein Distance
por: Rocamora, Elias Abad, et al.
Publicado: (2025)
por: Rocamora, Elias Abad, et al.
Publicado: (2025)
AdaSPEC: Selective Knowledge Distillation for Efficient Speculative Decoders
por: Hu, Yuezhou, et al.
Publicado: (2025)
por: Hu, Yuezhou, et al.
Publicado: (2025)
Graph-based Confidence Calibration for Large Language Models
por: Li, Yukun, et al.
Publicado: (2024)
por: Li, Yukun, et al.
Publicado: (2024)
Beyond Confidence: Rethinking Self-Assessments for Performance Prediction in LLMs
por: Bhattacharyya, Sree, et al.
Publicado: (2026)
por: Bhattacharyya, Sree, et al.
Publicado: (2026)
AdaptiveStep: Automatically Dividing Reasoning Step through Model Confidence
por: Liu, Yuliang, et al.
Publicado: (2025)
por: Liu, Yuliang, et al.
Publicado: (2025)
Calibrating Language Models with Adaptive Temperature Scaling
por: Xie, Johnathan, et al.
Publicado: (2024)
por: Xie, Johnathan, et al.
Publicado: (2024)
Online Speculative Decoding
por: Liu, Xiaoxuan, et al.
Publicado: (2023)
por: Liu, Xiaoxuan, et al.
Publicado: (2023)
CAS-Spec: Cascade Adaptive Self-Speculative Decoding for On-the-Fly Lossless Inference Acceleration of LLMs
por: Ning, Zhiyuan, et al.
Publicado: (2025)
por: Ning, Zhiyuan, et al.
Publicado: (2025)
Ejemplares similares
-
Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration
por: Wen, Zhuofan, et al.
Publicado: (2024) -
SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths
por: Huang, Kaixuan, et al.
Publicado: (2024) -
HiSpec: Hierarchical Speculative Decoding for LLMs
por: Kumar, Avinash, et al.
Publicado: (2025) -
SpecExit: Accelerating Large Reasoning Model via Speculative Exit
por: Yang, Rubing, et al.
Publicado: (2025) -
DistillSpec: Improving Speculative Decoding via Knowledge Distillation
por: Zhou, Yongchao, et al.
Publicado: (2023)