When RL Meets Adaptive Speculative Training: A Unified Training-Serving System
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Junxiong, Bie, Fengxiang, Li, Jisen, Zhou, Zhongzhu, Shao, Zelei, Wang, Yubo, Liu, Yinghui, Wu, Qingyang, May, Avner, Yanamandra, Sri, Zhang, Ce, Dao, Tri, Liang, Percy, Athiwaratkun, Ben, Song, Shuaiwen Leon, Xu, Chenfeng, Wu, Xiaoxia |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Beat the long tail: Distribution-Aware Speculative Decoding for RL Training
di: Shao, Zelei, et al.
Pubblicazione: (2025)
di: Shao, Zelei, et al.
Pubblicazione: (2025)
SAW-INT4: System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving
di: Jia, Jinda, et al.
Pubblicazione: (2026)
di: Jia, Jinda, et al.
Pubblicazione: (2026)
Introspective Diffusion Language Models
di: Yu, Yifan, et al.
Pubblicazione: (2026)
di: Yu, Yifan, et al.
Pubblicazione: (2026)
CARE: Covariance-Aware and Rank-Enhanced Decomposition for Enabling Multi-Head Latent Attention
di: Zhou, Zhongzhu, et al.
Pubblicazione: (2026)
di: Zhou, Zhongzhu, et al.
Pubblicazione: (2026)
Speculative Speculative Decoding
di: Kumar, Tanishq, et al.
Pubblicazione: (2026)
di: Kumar, Tanishq, et al.
Pubblicazione: (2026)
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization
di: Zhou, Zhongzhu, et al.
Pubblicazione: (2026)
di: Zhou, Zhongzhu, et al.
Pubblicazione: (2026)
Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient
di: Zhou, Zhongzhu, et al.
Pubblicazione: (2025)
di: Zhou, Zhongzhu, et al.
Pubblicazione: (2025)
The Mamba in the Llama: Distilling and Accelerating Hybrid Models
di: Wang, Junxiong, et al.
Pubblicazione: (2024)
di: Wang, Junxiong, et al.
Pubblicazione: (2024)
Data Diversification Methods In Alignment Enhance Math Performance In LLMs
di: Dokmeci, Berkan, et al.
Pubblicazione: (2025)
di: Dokmeci, Berkan, et al.
Pubblicazione: (2025)
Understanding and Steering the Cognitive Behaviors of Reasoning Models at Test-Time
di: Zhang, Zhenyu, et al.
Pubblicazione: (2025)
di: Zhang, Zhenyu, et al.
Pubblicazione: (2025)
Squeeze Evolve: Unified Multi-Model Orchestration for Verifier-Free Evolution
di: Maheswaran, Monishwaran, et al.
Pubblicazione: (2026)
di: Maheswaran, Monishwaran, et al.
Pubblicazione: (2026)
Opportunistic Expert Activation: Batch-Aware Expert Routing for Faster Decode Without Retraining
di: Oncescu, Costin-Andrei, et al.
Pubblicazione: (2025)
di: Oncescu, Costin-Andrei, et al.
Pubblicazione: (2025)
Kitty: Accurate and Efficient 2-bit KV Cache Quantization with Dynamic Channel-wise Precision Boost
di: Xia, Haojun, et al.
Pubblicazione: (2025)
di: Xia, Haojun, et al.
Pubblicazione: (2025)
Ladder-residual: parallelism-aware architecture for accelerating large model inference with communication overlapping
di: Zhang, Muru, et al.
Pubblicazione: (2025)
di: Zhang, Muru, et al.
Pubblicazione: (2025)
Think Deep, Think Fast: Investigating Efficiency of Verifier-free Inference-time-scaling Methods
di: Wang, Junlin, et al.
Pubblicazione: (2025)
di: Wang, Junlin, et al.
Pubblicazione: (2025)
RedPajama: an Open Dataset for Training Large Language Models
di: Weber, Maurice, et al.
Pubblicazione: (2024)
di: Weber, Maurice, et al.
Pubblicazione: (2024)
FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design
di: Xia, Haojun, et al.
Pubblicazione: (2024)
di: Xia, Haojun, et al.
Pubblicazione: (2024)
Search Your Block Floating Point Scales!
di: Gupta, Tanmaey, et al.
Pubblicazione: (2026)
di: Gupta, Tanmaey, et al.
Pubblicazione: (2026)
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation
di: Liu, Jingyu, et al.
Pubblicazione: (2025)
di: Liu, Jingyu, et al.
Pubblicazione: (2025)
RICE: Breaking Through the Training Bottlenecks of Reinforcement Learning with Explanation
di: Cheng, Zelei, et al.
Pubblicazione: (2024)
di: Cheng, Zelei, et al.
Pubblicazione: (2024)
Scaling Speculative Decoding with Lookahead Reasoning
di: Fu, Yichao, et al.
Pubblicazione: (2025)
di: Fu, Yichao, et al.
Pubblicazione: (2025)
Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding
di: Chen, Zhuoming, et al.
Pubblicazione: (2024)
di: Chen, Zhuoming, et al.
Pubblicazione: (2024)
SpecExec: Massively Parallel Speculative Decoding for Interactive LLM Inference on Consumer Devices
di: Svirschevski, Ruslan, et al.
Pubblicazione: (2024)
di: Svirschevski, Ruslan, et al.
Pubblicazione: (2024)
CorDA: Context-Oriented Decomposition Adaptation of Large Language Models for Task-Aware Parameter-Efficient Fine-tuning
di: Yang, Yibo, et al.
Pubblicazione: (2024)
di: Yang, Yibo, et al.
Pubblicazione: (2024)
How Well Can General Vision-Language Models Learn Medicine By Watching Public Educational Videos?
di: Thapa, Rahul, et al.
Pubblicazione: (2025)
di: Thapa, Rahul, et al.
Pubblicazione: (2025)
CDLM: Consistency Diffusion Language Models For Faster Sampling
di: Kim, Minseo, et al.
Pubblicazione: (2025)
di: Kim, Minseo, et al.
Pubblicazione: (2025)
Improving Model Alignment Through Collective Intelligence of Open-Source LLMS
di: Wang, Junlin, et al.
Pubblicazione: (2025)
di: Wang, Junlin, et al.
Pubblicazione: (2025)
Securing Recommender System via Cooperative Training
di: Wang, Qingyang, et al.
Pubblicazione: (2024)
di: Wang, Qingyang, et al.
Pubblicazione: (2024)
Disentangling Reasoning and Knowledge in Medical Large Language Models
di: Thapa, Rahul, et al.
Pubblicazione: (2025)
di: Thapa, Rahul, et al.
Pubblicazione: (2025)
Efficient Training-Free Online Routing for High-Volume Multi-LLM Serving
di: Wu, Fangzhou, et al.
Pubblicazione: (2025)
di: Wu, Fangzhou, et al.
Pubblicazione: (2025)
Goose: Anisotropic Speculation Trees for Training-Free Speculative Decoding
di: Jin, Tao, et al.
Pubblicazione: (2026)
di: Jin, Tao, et al.
Pubblicazione: (2026)
$V_1$: Unifying Generation and Self-Verification for Parallel Reasoners
di: Singh, Harman, et al.
Pubblicazione: (2026)
di: Singh, Harman, et al.
Pubblicazione: (2026)
M1: Towards Scalable Test-Time Compute with Mamba Reasoning Models
di: Wang, Junxiong, et al.
Pubblicazione: (2025)
di: Wang, Junxiong, et al.
Pubblicazione: (2025)
Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models
di: Thapa, Rahul, et al.
Pubblicazione: (2024)
di: Thapa, Rahul, et al.
Pubblicazione: (2024)
Speculate Deep and Accurate: Lossless and Training-Free Acceleration for Offloaded LLMs via Substitute Speculative Decoding
di: Wang, Pei-Shuo, et al.
Pubblicazione: (2025)
di: Wang, Pei-Shuo, et al.
Pubblicazione: (2025)
Training-Free Activation Sparsity in Large Language Models
di: Liu, James, et al.
Pubblicazione: (2024)
di: Liu, James, et al.
Pubblicazione: (2024)
Mamba: Linear-Time Sequence Modeling with Selective State Spaces
di: Gu, Albert, et al.
Pubblicazione: (2023)
di: Gu, Albert, et al.
Pubblicazione: (2023)
Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
di: Dao, Tri, et al.
Pubblicazione: (2024)
di: Dao, Tri, et al.
Pubblicazione: (2024)
Chiplet Cloud: Building AI Supercomputers for Serving Large Generative Language Models
di: Peng, Huwan, et al.
Pubblicazione: (2023)
di: Peng, Huwan, et al.
Pubblicazione: (2023)
Investigating the Synergistic Effects of Dropout and Residual Connections on Language Model Training
di: Li, Qingyang, et al.
Pubblicazione: (2024)
di: Li, Qingyang, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Beat the long tail: Distribution-Aware Speculative Decoding for RL Training
di: Shao, Zelei, et al.
Pubblicazione: (2025) -
SAW-INT4: System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving
di: Jia, Jinda, et al.
Pubblicazione: (2026) -
Introspective Diffusion Language Models
di: Yu, Yifan, et al.
Pubblicazione: (2026) -
CARE: Covariance-Aware and Rank-Enhanced Decomposition for Enabling Multi-Head Latent Attention
di: Zhou, Zhongzhu, et al.
Pubblicazione: (2026) -
Speculative Speculative Decoding
di: Kumar, Tanishq, et al.
Pubblicazione: (2026)