Saved in:
| Main Authors: | Johnson, Benjamin K., Goralski, Thomas, Semwal, Ayush, Shen, Hui, Jang, H. Josh |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.18780 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
by: Shao, Zishan, et al.
Published: (2025)
by: Shao, Zishan, et al.
Published: (2025)
FlashFormer: Whole-Model Kernels for Efficient Low-Batch Inference
by: Nrusimha, Aniruddha, et al.
Published: (2025)
by: Nrusimha, Aniruddha, et al.
Published: (2025)
PathCRF: Ball-Free Soccer Event Detection via Possession Path Inference from Player Trajectories
by: Kim, Hyunsung, et al.
Published: (2026)
by: Kim, Hyunsung, et al.
Published: (2026)
FlashDecoding++: Faster Large Language Model Inference on GPUs
by: Hong, Ke, et al.
Published: (2023)
by: Hong, Ke, et al.
Published: (2023)
Flash Inference: Near Linear Time Inference for Long Convolution Sequence Models and Beyond
by: Oncescu, Costin-Andrei, et al.
Published: (2024)
by: Oncescu, Costin-Andrei, et al.
Published: (2024)
MatKV: Trading Compute for Flash Storage in LLM Inference
by: Shin, Kun-Woo, et al.
Published: (2025)
by: Shin, Kun-Woo, et al.
Published: (2025)
Statistical Inference and Learning for Shapley Additive Explanations (SHAP)
by: Whitehouse, Justin, et al.
Published: (2026)
by: Whitehouse, Justin, et al.
Published: (2026)
FlashMask: Efficient and Rich Mask Extension of FlashAttention
by: Wang, Guoxia, et al.
Published: (2024)
by: Wang, Guoxia, et al.
Published: (2024)
Is Flash Attention Stable?
by: Golden, Alicia, et al.
Published: (2024)
by: Golden, Alicia, et al.
Published: (2024)
Flash PD-SSM: Memory-Optimized Structured Sparse State-Space Models
by: Terzić, Aleksandar, et al.
Published: (2026)
by: Terzić, Aleksandar, et al.
Published: (2026)
ISO-Bench: Can Coding Agents Optimize Real-World Inference Workloads?
by: Nangia, Ayush, et al.
Published: (2026)
by: Nangia, Ayush, et al.
Published: (2026)
Robust Simulation-Based Inference under Missing Data via Neural Processes
by: Verma, Yogesh, et al.
Published: (2025)
by: Verma, Yogesh, et al.
Published: (2025)
Particle Semi-Implicit Variational Inference
by: Lim, Jen Ning, et al.
Published: (2024)
by: Lim, Jen Ning, et al.
Published: (2024)
Bayesian Semi-structured Subspace Inference
by: Dold, Daniel, et al.
Published: (2024)
by: Dold, Daniel, et al.
Published: (2024)
Kernel Semi-Implicit Variational Inference
by: Cheng, Ziheng, et al.
Published: (2024)
by: Cheng, Ziheng, et al.
Published: (2024)
Dataset Mention Extraction in Scientific Articles Using Bi-LSTM-CRF Model
by: Zeng, Tong, et al.
Published: (2024)
by: Zeng, Tong, et al.
Published: (2024)
HREB-CRF: Hierarchical Reduced-bias EMA for Chinese Named Entity Recognition
by: Sun, Sijin, et al.
Published: (2025)
by: Sun, Sijin, et al.
Published: (2025)
Autoregressive Synthesis of Sparse and Semi-Structured Mixed-Type Data
by: Rückstieß, Thomas, et al.
Published: (2026)
by: Rückstieß, Thomas, et al.
Published: (2026)
Spatial Deconfounder: Interference-Aware Deconfounding for Spatial Causal Inference
by: Khot, Ayush, et al.
Published: (2025)
by: Khot, Ayush, et al.
Published: (2025)
ALINE: Joint Amortization for Bayesian Inference and Active Data Acquisition
by: Huang, Daolang, et al.
Published: (2025)
by: Huang, Daolang, et al.
Published: (2025)
Beyond Distribution Estimation: Simplex Anchored Structural Inference Towards Universal Semi-Supervised Learning
by: Hou, Yaxin, et al.
Published: (2026)
by: Hou, Yaxin, et al.
Published: (2026)
VFA: Relieving Vector Operations in Flash Attention with Global Maximum Pre-computation
by: Sun, Yupeng, et al.
Published: (2026)
by: Sun, Yupeng, et al.
Published: (2026)
Realistic Image-to-Image Machine Unlearning via Decoupling and Knowledge Retention
by: Varshney, Ayush K., et al.
Published: (2025)
by: Varshney, Ayush K., et al.
Published: (2025)
FlashHead: Efficient Drop-In Replacement for the Classification Head in Language Model Inference
by: Tranheden, Wilhelm, et al.
Published: (2026)
by: Tranheden, Wilhelm, et al.
Published: (2026)
INT-FlashAttention: Enabling Flash Attention for INT8 Quantization
by: Chen, Shimao, et al.
Published: (2024)
by: Chen, Shimao, et al.
Published: (2024)
StreamPhy: Streaming Inference of High-Dimensional Physical Dynamics via State Space Models
by: Chen, Panqi, et al.
Published: (2026)
by: Chen, Panqi, et al.
Published: (2026)
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
by: Ye, Zihao, et al.
Published: (2025)
by: Ye, Zihao, et al.
Published: (2025)
A Kernel Approach for Semi-implicit Variational Inference
by: Yu, Longlin, et al.
Published: (2026)
by: Yu, Longlin, et al.
Published: (2026)
FlashNorm: Fast Normalization for Transformers
by: Graef, Nils, et al.
Published: (2024)
by: Graef, Nils, et al.
Published: (2024)
Structure Maintained Representation Learning Neural Network for Causal Inference
by: Sun, Yang, et al.
Published: (2025)
by: Sun, Yang, et al.
Published: (2025)
Explore BiLSTM-CRF-Based Models for Open Relation Extraction
by: Ni, Tao, et al.
Published: (2021)
by: Ni, Tao, et al.
Published: (2021)
FlashSVD v1.5: Making Low-Rank Transformers Inference Actually Fast
by: Wu, Wenhao, et al.
Published: (2026)
by: Wu, Wenhao, et al.
Published: (2026)
FlashSampling: Fast and Memory-Efficient Exact Sampling
by: Ruiz, Tomas, et al.
Published: (2026)
by: Ruiz, Tomas, et al.
Published: (2026)
Additive Large Language Models for Semi-Structured Text
by: K, Karthikeyan, et al.
Published: (2025)
by: K, Karthikeyan, et al.
Published: (2025)
Semi-Supervised Risk Control via Prediction-Powered Inference
by: Einbinder, Bat-Sheva, et al.
Published: (2024)
by: Einbinder, Bat-Sheva, et al.
Published: (2024)
End-to-end Sequence Labeling via Bi-directional LSTM-CNNs-CRF: A Reproducibility Study
by: Ganesh, Anirudh, et al.
Published: (2025)
by: Ganesh, Anirudh, et al.
Published: (2025)
Efficient Federated Unlearning under Plausible Deniability
by: Varshney, Ayush K., et al.
Published: (2024)
by: Varshney, Ayush K., et al.
Published: (2024)
Concept Drift Detection using Ensemble of Integrally Private Models
by: Varshney, Ayush K., et al.
Published: (2024)
by: Varshney, Ayush K., et al.
Published: (2024)
Causal-INSIGHT: Probing Temporal Models to Extract Causal Structure
by: Redden, Benjamin, et al.
Published: (2026)
by: Redden, Benjamin, et al.
Published: (2026)
FlashBias: Fast Computation of Attention with Bias
by: Wu, Haixu, et al.
Published: (2025)
by: Wu, Haixu, et al.
Published: (2025)
Similar Items
-
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
by: Shao, Zishan, et al.
Published: (2025) -
FlashFormer: Whole-Model Kernels for Efficient Low-Batch Inference
by: Nrusimha, Aniruddha, et al.
Published: (2025) -
PathCRF: Ball-Free Soccer Event Detection via Possession Path Inference from Player Trajectories
by: Kim, Hyunsung, et al.
Published: (2026) -
FlashDecoding++: Faster Large Language Model Inference on GPUs
by: Hong, Ke, et al.
Published: (2023) -
Flash Inference: Near Linear Time Inference for Long Convolution Sequence Models and Beyond
by: Oncescu, Costin-Andrei, et al.
Published: (2024)