Accelerating RL Post-Training Rollouts via System-Integrated Speculative Decoding
Fuente:
arXiv
Saved in:
| Main Authors: | Iso, Hayate, Mitra, Tiyasa, Mondal, Sudipta, Shafipour, Rasoul, Elango, Venmugil, Kong, Terry, Huang, Yuki, Na, Seonjin, Putterman, Izzy, Chislett, Benjamin, Ashkenazi, Maor, Guman, Joseph, Shen, Gerald, Konuk, Tugrul, Aithal, Ashwath, Borkar, Ritika, Zilberstein, Ran, Rouhani, Bita |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding
by: Abramovich, Talor, et al.
Published: (2026)
by: Abramovich, Talor, et al.
Published: (2026)
LatentMoE: Toward Optimal Accuracy per FLOP and Parameter in Mixture of Experts
by: Elango, Venmugil, et al.
Published: (2026)
by: Elango, Venmugil, et al.
Published: (2026)
ATTENTION2D: Communication Efficient Distributed Self-Attention Mechanism
by: Elango, Venmugil
Published: (2025)
by: Elango, Venmugil
Published: (2025)
PaSE: Parallelization Strategies for Efficient DNN Training
by: Elango, Venmugil
Published: (2024)
by: Elango, Venmugil
Published: (2024)
AutoTemplate: A Simple Recipe for Lexically Constrained Text Generation
by: Iso, Hayate
Published: (2022)
by: Iso, Hayate
Published: (2022)
Helix Parallelism: Rethinking Sharding Strategies for Interactive Multi-Million-Token LLM Decoding
by: Bhatia, Nidhi, et al.
Published: (2025)
by: Bhatia, Nidhi, et al.
Published: (2025)
Elastic least‐squares reverse time migration from topography through anisotropic tensorial elastodynamics
by: Tugrul Konuk, et al.
Published: (2024)
by: Tugrul Konuk, et al.
Published: (2024)
AmbigNLG: Addressing Task Ambiguity in Instruction for NLG
by: Niwa, Ayana, et al.
Published: (2024)
by: Niwa, Ayana, et al.
Published: (2024)
Beyond the Buzz: A Pragmatic Take on Inference Disaggregation
by: Mitra, Tiyasa, et al.
Published: (2025)
by: Mitra, Tiyasa, et al.
Published: (2025)
Towards Croppable Implicit Neural Representations
by: Ashkenazi, Maor, et al.
Published: (2024)
by: Ashkenazi, Maor, et al.
Published: (2024)
Tied-Lora: Enhancing parameter efficiency of LoRA with weight tying
by: Renduchintala, Adithya, et al.
Published: (2023)
by: Renduchintala, Adithya, et al.
Published: (2023)
Noisy Pairing and Partial Supervision for Stylized Opinion Summarization
by: Iso, Hayate, et al.
Published: (2022)
by: Iso, Hayate, et al.
Published: (2022)
Holistic Reasoning with Long-Context LMs: A Benchmark for Database Operations on Massive Textual Data
by: Maekawa, Seiji, et al.
Published: (2024)
by: Maekawa, Seiji, et al.
Published: (2024)
The Rarity Blind Spot: A Framework for Evaluating Statistical Reasoning in LLMs
by: Maekawa, Seiji, et al.
Published: (2025)
by: Maekawa, Seiji, et al.
Published: (2025)
XATU: A Fine-grained Instruction-based Benchmark for Explainable Text Updates
by: Zhang, Haopeng, et al.
Published: (2023)
by: Zhang, Haopeng, et al.
Published: (2023)
Evaluating Bias in LLMs for Job-Resume Matching: Gender, Race, and Education
by: Iso, Hayate, et al.
Published: (2025)
by: Iso, Hayate, et al.
Published: (2025)
Retrieval Helps or Hurts? A Deeper Dive into the Efficacy of Retrieval Augmentation to Language Models
by: Maekawa, Seiji, et al.
Published: (2024)
by: Maekawa, Seiji, et al.
Published: (2024)
Key, Value, Compress: A Systematic Exploration of KV Cache Compression Techniques
by: Javidnia, Neusha, et al.
Published: (2025)
by: Javidnia, Neusha, et al.
Published: (2025)
Emission Distribution for the quantas of Maxwell-Chern-Simon Gauge Field coupled to External Current
by: Kar, Tiyasa
Published: (2021)
by: Kar, Tiyasa
Published: (2021)
Less is More for Long Document Summary Evaluation by LLMs
by: Wu, Yunshu, et al.
Published: (2023)
by: Wu, Yunshu, et al.
Published: (2023)
Efficient MoE Serving in the Memory-Bound Regime: Balance Activated Experts, Not Tokens
by: Yu, Yanpeng, et al.
Published: (2025)
by: Yu, Yanpeng, et al.
Published: (2025)
Zero-Shot Detection of LLM-Generated Code via Approximated Task Conditioning
by: Ashkenazi, Maor, et al.
Published: (2025)
by: Ashkenazi, Maor, et al.
Published: (2025)
La Psicología y La Psicoterapia en Otros Países "La Psicología Tiene un Largo Pasado Pero una Historia Corta": Turquía, como un ejemplo
by: Emre Konuk
Published: (2011)
by: Emre Konuk
Published: (2011)
From Single to Multi: How LLMs Hallucinate in Multi-Document Summarization
by: Belem, Catarina G., et al.
Published: (2024)
by: Belem, Catarina G., et al.
Published: (2024)
Llama 3 Meets MoE: Efficient Upcycling
by: Vavre, Aditya, et al.
Published: (2024)
by: Vavre, Aditya, et al.
Published: (2024)
Asymptotic expansion of the weighted power variation with second order differences of a stochastic differential equation driven by fBm
by: Yamagishi, Hayate
Published: (2024)
by: Yamagishi, Hayate
Published: (2024)
Seat number configuration of the box-ball system, and its relation to the 10-elimination and invariant measures
by: Suda, Hayate
Published: (2023)
by: Suda, Hayate
Published: (2023)
Asymptotic expansion of a Hurst index estimator for a stochastic differential equation driven by fBm
by: Yamagishi, Hayate
Published: (2024)
by: Yamagishi, Hayate
Published: (2024)
The MacWilliams Identity for Krawtchouk Association Schemes
by: Friedlander, Izzy
Published: (2024)
by: Friedlander, Izzy
Published: (2024)
The MacWilliams Identity for the Hermitian Rank Metric
by: Friedlander, Izzy
Published: (2023)
by: Friedlander, Izzy
Published: (2023)
An Investigation into Emotional Intelligence, Foreign Language Anxiety and Empathy through a Cognitive-Affective Course in an EFL Context
by: Ali Rouhani
Published: (2008)
by: Ali Rouhani
Published: (2008)
AI-driven GPT Prompts for Industry Analysis Scholarly Article
by: Aithal, Sreeramana
Published: (2025)
by: Aithal, Sreeramana
Published: (2025)
Ethical Leadership in the Light of the Bhagavad Gita: Exploring Dharma-Centered Management Practices for the 21st Century
by: Aithal, Sreeramana
Published: (2025)
by: Aithal, Sreeramana
Published: (2025)
Outcome Logic: A Unified Approach to the Metatheory of Program Logics with Branching Effects
by: Zilberstein, Noam
Published: (2024)
by: Zilberstein, Noam
Published: (2024)
SeedLM: Compressing LLM Weights into Seeds of Pseudo-Random Generators
by: Shafipour, Rasoul, et al.
Published: (2024)
by: Shafipour, Rasoul, et al.
Published: (2024)
A New Dataset and Methodology for Malicious URL Classification
by: Schvartzman, Ilan, et al.
Published: (2024)
by: Schvartzman, Ilan, et al.
Published: (2024)
Thermodynamic Characteristics of a Fermi Gas with an Invariant Energy Scale and its Astrophysical Implications
by: Kar, Tiyasa, et al.
Published: (2026)
by: Kar, Tiyasa, et al.
Published: (2026)
X-Token: Projection-Guided Cross-Tokenizer Knowledge Distillation
by: Sreenivas, Sharath Turuvekere, et al.
Published: (2026)
by: Sreenivas, Sharath Turuvekere, et al.
Published: (2026)
Localization and corruption : panacea or pandora's box? / Tugrul Gurgur, Anwar Shah
by: Gurgur, Tugrul
Published: (2005)
by: Gurgur, Tugrul
Published: (2005)
Stochastic Approximation with Two Time Scales: The General Case
by: Borkar, Vivek S
Published: (2024)
by: Borkar, Vivek S
Published: (2024)
Similar Items
-
SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding
by: Abramovich, Talor, et al.
Published: (2026) -
LatentMoE: Toward Optimal Accuracy per FLOP and Parameter in Mixture of Experts
by: Elango, Venmugil, et al.
Published: (2026) -
ATTENTION2D: Communication Efficient Distributed Self-Attention Mechanism
by: Elango, Venmugil
Published: (2025) -
PaSE: Parallelization Strategies for Efficient DNN Training
by: Elango, Venmugil
Published: (2024) -
AutoTemplate: A Simple Recipe for Lexically Constrained Text Generation
by: Iso, Hayate
Published: (2022)