Arbitrage: Efficient Reasoning via Advantage-Aware Speculation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Maheswaran, Monishwaran, Tiwari, Rishabh, Hu, Yuezhou, Dilmen, Kerem, Hooper, Coleman, Xi, Haocheng, Lee, Nicholas, Farajtabar, Mehrdad, Mahoney, Michael W., Keutzer, Kurt, Gholami, Amir |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ETS: Efficient Tree Search for Inference-Time Scaling
von: Hooper, Coleman, et al.
Veröffentlicht: (2025)
von: Hooper, Coleman, et al.
Veröffentlicht: (2025)
Residual Context Diffusion Language Models
von: Hu, Yuezhou, et al.
Veröffentlicht: (2026)
von: Hu, Yuezhou, et al.
Veröffentlicht: (2026)
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2025)
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2025)
Reward Under Attack: Analyzing the Robustness and Hackability of Process Reward Models
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2026)
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2026)
LoSA: Locality Aware Sparse Attention for Block-Wise Diffusion Language Models
von: Xi, Haocheng, et al.
Veröffentlicht: (2026)
von: Xi, Haocheng, et al.
Veröffentlicht: (2026)
Squeezed Attention: Accelerating Long Context Length LLM Inference
von: Hooper, Coleman, et al.
Veröffentlicht: (2024)
von: Hooper, Coleman, et al.
Veröffentlicht: (2024)
Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models
von: Kim, Minseo, et al.
Veröffentlicht: (2025)
von: Kim, Minseo, et al.
Veröffentlicht: (2025)
SPEED: Speculative Pipelined Execution for Efficient Decoding
von: Hooper, Coleman, et al.
Veröffentlicht: (2023)
von: Hooper, Coleman, et al.
Veröffentlicht: (2023)
XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization
von: Tomar, Aditya, et al.
Veröffentlicht: (2025)
von: Tomar, Aditya, et al.
Veröffentlicht: (2025)
SciML Agents: Write the Solver, Not the Solution
von: Gaonkar, Saarth, et al.
Veröffentlicht: (2025)
von: Gaonkar, Saarth, et al.
Veröffentlicht: (2025)
Multipole Attention for Efficient Long Context Reasoning
von: Hooper, Coleman, et al.
Veröffentlicht: (2025)
von: Hooper, Coleman, et al.
Veröffentlicht: (2025)
Speculative Interaction Agents: Building Real-Time Agents with Asynchronous I/O and Speculative Tool Calling
von: Hooper, Coleman, et al.
Veröffentlicht: (2026)
von: Hooper, Coleman, et al.
Veröffentlicht: (2026)
AI and Memory Wall
von: Gholami, Amir, et al.
Veröffentlicht: (2024)
von: Gholami, Amir, et al.
Veröffentlicht: (2024)
KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
von: Hooper, Coleman, et al.
Veröffentlicht: (2024)
von: Hooper, Coleman, et al.
Veröffentlicht: (2024)
SqueezeLLM: Dense-and-Sparse Quantization
von: Kim, Sehoon, et al.
Veröffentlicht: (2023)
von: Kim, Sehoon, et al.
Veröffentlicht: (2023)
Squeeze Evolve: Unified Multi-Model Orchestration for Verifier-Free Evolution
von: Maheswaran, Monishwaran, et al.
Veröffentlicht: (2026)
von: Maheswaran, Monishwaran, et al.
Veröffentlicht: (2026)
CDLM: Consistency Diffusion Language Models For Faster Sampling
von: Kim, Minseo, et al.
Veröffentlicht: (2025)
von: Kim, Minseo, et al.
Veröffentlicht: (2025)
TASER: Translation Assessment via Systematic Evaluation and Reasoning
von: Maheswaran, Monishwaran, et al.
Veröffentlicht: (2025)
von: Maheswaran, Monishwaran, et al.
Veröffentlicht: (2025)
An LLM Compiler for Parallel Function Calling
von: Kim, Sehoon, et al.
Veröffentlicht: (2023)
von: Kim, Sehoon, et al.
Veröffentlicht: (2023)
Learned Best-Effort LLM Serving
von: Jha, Siddharth, et al.
Veröffentlicht: (2024)
von: Jha, Siddharth, et al.
Veröffentlicht: (2024)
TinyAgent: Function Calling at the Edge
von: Erdogan, Lutfi Eren, et al.
Veröffentlicht: (2024)
von: Erdogan, Lutfi Eren, et al.
Veröffentlicht: (2024)
Towards Foundation Models for Scientific Machine Learning: Characterizing Scaling and Transfer Behavior
von: Subramanian, Shashank, et al.
Veröffentlicht: (2023)
von: Subramanian, Shashank, et al.
Veröffentlicht: (2023)
Agentic Test-Time Scaling for WebAgents
von: Lee, Nicholas, et al.
Veröffentlicht: (2026)
von: Lee, Nicholas, et al.
Veröffentlicht: (2026)
$V_1$: Unifying Generation and Self-Verification for Parallel Reasoners
von: Singh, Harman, et al.
Veröffentlicht: (2026)
von: Singh, Harman, et al.
Veröffentlicht: (2026)
$\texttt{SPECS}$: Faster Test-Time Scaling through Speculative Drafts
von: Cemri, Mert, et al.
Veröffentlicht: (2025)
von: Cemri, Mert, et al.
Veröffentlicht: (2025)
LLM2LLM: Boosting LLMs with Novel Iterative Data Enhancement
von: Lee, Nicholas, et al.
Veröffentlicht: (2024)
von: Lee, Nicholas, et al.
Veröffentlicht: (2024)
Characterizing Prompt Compression Methods for Long Context Inference
von: Jha, Siddharth, et al.
Veröffentlicht: (2024)
von: Jha, Siddharth, et al.
Veröffentlicht: (2024)
Efficient and Scalable Estimation of Tool Representations in Vector Space
von: Moon, Suhong, et al.
Veröffentlicht: (2024)
von: Moon, Suhong, et al.
Veröffentlicht: (2024)
Speculative Decoding for Autoregressive Video Generation
von: Hu, Yuezhou, et al.
Veröffentlicht: (2026)
von: Hu, Yuezhou, et al.
Veröffentlicht: (2026)
COAT: Compressing Optimizer states and Activation for Memory-Efficient FP8 Training
von: Xi, Haocheng, et al.
Veröffentlicht: (2024)
von: Xi, Haocheng, et al.
Veröffentlicht: (2024)
FGMP: Fine-Grained Mixed-Precision Weight and Activation Quantization for Hardware-Accelerated LLM Inference
von: Hooper, Coleman, et al.
Veröffentlicht: (2025)
von: Hooper, Coleman, et al.
Veröffentlicht: (2025)
AdaSPEC: Selective Knowledge Distillation for Efficient Speculative Decoders
von: Hu, Yuezhou, et al.
Veröffentlicht: (2025)
von: Hu, Yuezhou, et al.
Veröffentlicht: (2025)
Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks
von: Erdogan, Lutfi Eren, et al.
Veröffentlicht: (2025)
von: Erdogan, Lutfi Eren, et al.
Veröffentlicht: (2025)
MEV Capture Through Time-Advantaged Arbitrage
von: Fritsch, Robin, et al.
Veröffentlicht: (2024)
von: Fritsch, Robin, et al.
Veröffentlicht: (2024)
SAFE-RL: Saliency-Aware Counterfactual Explainer for Deep Reinforcement Learning Policies
von: Samadi, Amir, et al.
Veröffentlicht: (2024)
von: Samadi, Amir, et al.
Veröffentlicht: (2024)
Repeated Auctions with Speculators: Arbitrage Incentives and Forks in DAOs
von: Eschenbaum, Nicolas, et al.
Veröffentlicht: (2025)
von: Eschenbaum, Nicolas, et al.
Veröffentlicht: (2025)
MITRA: A Large-Scale Parallel Corpus and Multilingual Pretrained Language Model for Machine Translation and Semantic Retrieval for Pāli, Sanskrit, Buddhist Chinese, and Tibetan
von: Nehrdich, Sebastian, et al.
Veröffentlicht: (2026)
von: Nehrdich, Sebastian, et al.
Veröffentlicht: (2026)
Learning, Fast and Slow: Towards LLMs That Adapt Continually
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2026)
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2026)
GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
von: Mirzadeh, Iman, et al.
Veröffentlicht: (2024)
von: Mirzadeh, Iman, et al.
Veröffentlicht: (2024)
Recursive Language Models Meet Uncertainty: The Surprising Effectiveness of Self-Reflective Program Search for Long Context
von: Alizadeh, Keivan, et al.
Veröffentlicht: (2026)
von: Alizadeh, Keivan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
ETS: Efficient Tree Search for Inference-Time Scaling
von: Hooper, Coleman, et al.
Veröffentlicht: (2025) -
Residual Context Diffusion Language Models
von: Hu, Yuezhou, et al.
Veröffentlicht: (2026) -
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2025) -
Reward Under Attack: Analyzing the Robustness and Hackability of Process Reward Models
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2026) -
LoSA: Locality Aware Sparse Attention for Block-Wise Diffusion Language Models
von: Xi, Haocheng, et al.
Veröffentlicht: (2026)