MarginGate: Sparse Margin-Triggered Verification for Batch-Invariant LLM Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chu, Kexin, Zhou, Yang, Zhang, Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
von: Chu, Kexin, et al.
Veröffentlicht: (2025)
von: Chu, Kexin, et al.
Veröffentlicht: (2025)
An Inquiry into Datacenter TCO for LLM Inference with FP8
von: Kim, Jiwoo, et al.
Veröffentlicht: (2025)
von: Kim, Jiwoo, et al.
Veröffentlicht: (2025)
PARD: Accelerating LLM Inference with Low-Cost PARallel Draft Model Adaptation
von: An, Zihao, et al.
Veröffentlicht: (2025)
von: An, Zihao, et al.
Veröffentlicht: (2025)
One Size Does Not Fit All: Architecture-Aware Adaptive Batch Scheduling with DEBA
von: Belias, François, et al.
Veröffentlicht: (2025)
von: Belias, François, et al.
Veröffentlicht: (2025)
Light Differentiable Logic Gate Networks
von: Rüttgers, Lukas, et al.
Veröffentlicht: (2025)
von: Rüttgers, Lukas, et al.
Veröffentlicht: (2025)
GreenServ: Energy-Efficient Context-Aware Dynamic Routing for Multi-Model LLM Inference
von: Ziller, Thomas, et al.
Veröffentlicht: (2026)
von: Ziller, Thomas, et al.
Veröffentlicht: (2026)
An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference
von: Yao, Feiyu, et al.
Veröffentlicht: (2026)
von: Yao, Feiyu, et al.
Veröffentlicht: (2026)
Mind the Gap: Removing the Discretization Gap in Differentiable Logic Gate Networks
von: Yousefi, Shakir, et al.
Veröffentlicht: (2025)
von: Yousefi, Shakir, et al.
Veröffentlicht: (2025)
Risk-Aware Batch Testing for Performance Regression Detection
von: Sayedsalehi, Ali, et al.
Veröffentlicht: (2026)
von: Sayedsalehi, Ali, et al.
Veröffentlicht: (2026)
Ragged Paged Attention: A High-Performance and Flexible LLM Inference Kernel for TPU
von: Jiang, Jevin, et al.
Veröffentlicht: (2026)
von: Jiang, Jevin, et al.
Veröffentlicht: (2026)
KVDirect: Distributed Disaggregated LLM Inference
von: Chen, Shiyang, et al.
Veröffentlicht: (2024)
von: Chen, Shiyang, et al.
Veröffentlicht: (2024)
Pushing the Envelope of LLM Inference on AI-PC and Intel GPUs
von: Georganas, Evangelos, et al.
Veröffentlicht: (2025)
von: Georganas, Evangelos, et al.
Veröffentlicht: (2025)
Accelerating Sparse Ternary GEMM for Quantized ML on Apple Silicon
von: Lipshitz, Baraq, et al.
Veröffentlicht: (2025)
von: Lipshitz, Baraq, et al.
Veröffentlicht: (2025)
Anatomizing Deep Learning Inference in Web Browsers
von: Wang, Qipeng, et al.
Veröffentlicht: (2024)
von: Wang, Qipeng, et al.
Veröffentlicht: (2024)
ModeSwitch-LLM: A Lightweight Phase-Aware Controller for Cross-Mode LLM Inference on a Single GPU
von: Sunesh, Aman, et al.
Veröffentlicht: (2026)
von: Sunesh, Aman, et al.
Veröffentlicht: (2026)
Multi-DNN Inference of Sparse Models on Edge SoCs
von: Luo, Jiawei, et al.
Veröffentlicht: (2026)
von: Luo, Jiawei, et al.
Veröffentlicht: (2026)
Forecasting GPU Performance for Deep Learning Training and Inference
von: Lee, Seonho, et al.
Veröffentlicht: (2024)
von: Lee, Seonho, et al.
Veröffentlicht: (2024)
ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference
von: Yin, Wangsong, et al.
Veröffentlicht: (2025)
von: Yin, Wangsong, et al.
Veröffentlicht: (2025)
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
von: Fan, Ruibo, et al.
Veröffentlicht: (2026)
von: Fan, Ruibo, et al.
Veröffentlicht: (2026)
lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
von: Wang, Haoxin, et al.
Veröffentlicht: (2025)
von: Wang, Haoxin, et al.
Veröffentlicht: (2025)
Plug-and-Play Performance Estimation for LLM Services without Relying on Labeled Data
von: Wang, Can, et al.
Veröffentlicht: (2024)
von: Wang, Can, et al.
Veröffentlicht: (2024)
AutoSAGE: Input-Aware CUDA Scheduling for Sparse GNN Aggregation (SpMM/SDDMM) and CSR Attention
von: Stankovic, Aleksandar
Veröffentlicht: (2025)
von: Stankovic, Aleksandar
Veröffentlicht: (2025)
KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels
von: Wang, Han, et al.
Veröffentlicht: (2026)
von: Wang, Han, et al.
Veröffentlicht: (2026)
SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models
von: Chitty-Venkata, Krishna Teja, et al.
Veröffentlicht: (2025)
von: Chitty-Venkata, Krishna Teja, et al.
Veröffentlicht: (2025)
Private LLM Inference on Consumer Blackwell GPUs: A Practical Guide for Cost-Effective Local Deployment in SMEs
von: Knoop, Jonathan, et al.
Veröffentlicht: (2026)
von: Knoop, Jonathan, et al.
Veröffentlicht: (2026)
An Interpretable Latency Model for Speculative Decoding in LLM Serving
von: Kong, Linghao, et al.
Veröffentlicht: (2026)
von: Kong, Linghao, et al.
Veröffentlicht: (2026)
MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
von: Xue, Leyang, et al.
Veröffentlicht: (2024)
von: Xue, Leyang, et al.
Veröffentlicht: (2024)
DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
von: Holmes, Connor, et al.
Veröffentlicht: (2024)
von: Holmes, Connor, et al.
Veröffentlicht: (2024)
CEBench: A Benchmarking Toolkit for the Cost-Effectiveness of LLM Pipelines
von: Sun, Wenbo, et al.
Veröffentlicht: (2024)
von: Sun, Wenbo, et al.
Veröffentlicht: (2024)
Block Sparse Flash Attention
von: Ohayon, Daniel, et al.
Veröffentlicht: (2025)
von: Ohayon, Daniel, et al.
Veröffentlicht: (2025)
Neuralink: Fast LLM Inference on Smartphones with Neuron Co-Activation Linking
von: Wang, Tuowei, et al.
Veröffentlicht: (2024)
von: Wang, Tuowei, et al.
Veröffentlicht: (2024)
V-Seek: Accelerating LLM Reasoning on Open-hardware Server-class RISC-V Platforms
von: Rodrigo, Javier J. Poveda, et al.
Veröffentlicht: (2025)
von: Rodrigo, Javier J. Poveda, et al.
Veröffentlicht: (2025)
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference
von: Hamadanian, Pouya, et al.
Veröffentlicht: (2025)
von: Hamadanian, Pouya, et al.
Veröffentlicht: (2025)
Enabling Performant and Flexible Model-Internal Observability for LLM Inference
von: Yu, Nengneng, et al.
Veröffentlicht: (2026)
von: Yu, Nengneng, et al.
Veröffentlicht: (2026)
Forecasting LLM Inference Performance via Hardware-Agnostic Analytical Modeling
von: Patwari, Rajeev, et al.
Veröffentlicht: (2025)
von: Patwari, Rajeev, et al.
Veröffentlicht: (2025)
Bench360: Benchmarking Local LLM Inference from 360 Degrees
von: Stuhlmann, Linus, et al.
Veröffentlicht: (2025)
von: Stuhlmann, Linus, et al.
Veröffentlicht: (2025)
Confidential LLM Inference: Performance and Cost Across CPU and GPU TEEs
von: Chrapek, Marcin, et al.
Veröffentlicht: (2025)
von: Chrapek, Marcin, et al.
Veröffentlicht: (2025)
On the Sustainability of AI Inferences in the Edge
von: Sobhani, Ghazal, et al.
Veröffentlicht: (2025)
von: Sobhani, Ghazal, et al.
Veröffentlicht: (2025)
Generalizing Scaling Laws for Dense and Sparse Large Language Models
von: Hossain, Md Arafat, et al.
Veröffentlicht: (2025)
von: Hossain, Md Arafat, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
von: Chu, Kexin, et al.
Veröffentlicht: (2025) -
An Inquiry into Datacenter TCO for LLM Inference with FP8
von: Kim, Jiwoo, et al.
Veröffentlicht: (2025) -
PARD: Accelerating LLM Inference with Low-Cost PARallel Draft Model Adaptation
von: An, Zihao, et al.
Veröffentlicht: (2025) -
One Size Does Not Fit All: Architecture-Aware Adaptive Batch Scheduling with DEBA
von: Belias, François, et al.
Veröffentlicht: (2025) -
Light Differentiable Logic Gate Networks
von: Rüttgers, Lukas, et al.
Veröffentlicht: (2025)