FluidML: Fast and Memory Efficient Inference Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Jinjie, Qiu, Hang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Data-Efficient Inference of Neural Fluid Fields via SciML Foundation Model
by: Liu, Yuqiu, et al.
Published: (2024)
by: Liu, Yuqiu, et al.
Published: (2024)
Fast Compute for ML Optimization
by: Polson, Nick, et al.
Published: (2026)
by: Polson, Nick, et al.
Published: (2026)
FAST: An Optimization Framework for Fast Additive Segmentation in Transparent ML
by: Liu, Brian, et al.
Published: (2024)
by: Liu, Brian, et al.
Published: (2024)
Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints
by: Ao, Ruicheng, et al.
Published: (2025)
by: Ao, Ruicheng, et al.
Published: (2025)
Toward Robust and Efficient ML-Based GPU Caching for Modern Inference
by: Chen, Peng, et al.
Published: (2025)
by: Chen, Peng, et al.
Published: (2025)
Memory-Efficient Optimization with Factorized Hamiltonian Descent
by: Nguyen, Son, et al.
Published: (2024)
by: Nguyen, Son, et al.
Published: (2024)
The ML.ENERGY Benchmark: Toward Automated Inference Energy Measurement and Optimization
by: Chung, Jae-Won, et al.
Published: (2025)
by: Chung, Jae-Won, et al.
Published: (2025)
ML Inference Scheduling with Predictable Latency
by: Zhao, Haidong, et al.
Published: (2025)
by: Zhao, Haidong, et al.
Published: (2025)
MicroFlow: An Efficient Rust-Based Inference Engine for TinyML
by: Carnelos, Matteo, et al.
Published: (2024)
by: Carnelos, Matteo, et al.
Published: (2024)
Fast and Robust Simulation-Based Inference With Optimization Monte Carlo
by: Gkolemis, Vasilis, et al.
Published: (2025)
by: Gkolemis, Vasilis, et al.
Published: (2025)
On Simplifying Large-Scale Spatial Vectors: Fast, Memory-Efficient, and Cost-Predictable k-means
by: Ji, Yushuai, et al.
Published: (2024)
by: Ji, Yushuai, et al.
Published: (2024)
MLonMCU: TinyML Benchmarking with Fast Retargeting
by: van Kempen, Philipp, et al.
Published: (2023)
by: van Kempen, Philipp, et al.
Published: (2023)
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
by: Fan, Ruibo, et al.
Published: (2026)
by: Fan, Ruibo, et al.
Published: (2026)
Strait: Perceiving Priority and Interference in ML Inference Serving
by: Zhao, Haidong, et al.
Published: (2026)
by: Zhao, Haidong, et al.
Published: (2026)
MCUBERT: Memory-Efficient BERT Inference on Commodity Microcontrollers
by: Yang, Zebin, et al.
Published: (2024)
by: Yang, Zebin, et al.
Published: (2024)
Fast and Compact Tsetlin Machine Inference on CPUs Using Instruction-Level Optimization
by: Zeng, Yefan, et al.
Published: (2025)
by: Zeng, Yefan, et al.
Published: (2025)
Rectified Flows for Fast Multiscale Fluid Flow Modeling
by: Armegioiu, Victor, et al.
Published: (2025)
by: Armegioiu, Victor, et al.
Published: (2025)
COSMOS: A Hybrid Adaptive Optimizer for Memory-Efficient Training of LLMs
by: Liu, Liming, et al.
Published: (2025)
by: Liu, Liming, et al.
Published: (2025)
Accelerating TinyML Inference on Microcontrollers through Approximate Kernels
by: Armeniakos, Giorgos, et al.
Published: (2024)
by: Armeniakos, Giorgos, et al.
Published: (2024)
Scalable and Cost-Efficient ML Inference: Parallel Batch Processing with Serverless Functions
by: Barrak, Amine, et al.
Published: (2025)
by: Barrak, Amine, et al.
Published: (2025)
Applied Causal Inference Powered by ML and AI
by: Chernozhukov, Victor, et al.
Published: (2024)
by: Chernozhukov, Victor, et al.
Published: (2024)
FastAV: Efficient Token Pruning for Audio-Visual Large Language Model Inference
by: Jung, Chaeyoung, et al.
Published: (2026)
by: Jung, Chaeyoung, et al.
Published: (2026)
A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models
by: Chen, Yiming, et al.
Published: (2025)
by: Chen, Yiming, et al.
Published: (2025)
FlashSampling: Fast and Memory-Efficient Exact Sampling
by: Ruiz, Tomas, et al.
Published: (2026)
by: Ruiz, Tomas, et al.
Published: (2026)
FERMI-ML: A Flexible and Resource-Efficient Memory-In-Situ SRAM Macro for TinyML acceleration
by: Lokhande, Mukul, et al.
Published: (2025)
by: Lokhande, Mukul, et al.
Published: (2025)
Fast-PGM: Fast Probabilistic Graphical Model Learning and Inference
by: Jiang, Jiantong, et al.
Published: (2024)
by: Jiang, Jiantong, et al.
Published: (2024)
Fast KVzip: Efficient and Accurate LLM Inference with Gated KV Eviction
by: Kim, Jang-Hyun, et al.
Published: (2026)
by: Kim, Jang-Hyun, et al.
Published: (2026)
FlashOptim: Optimizers for Memory-Efficient Training
by: Ortiz, Jose Javier Gonzalez, et al.
Published: (2026)
by: Ortiz, Jose Javier Gonzalez, et al.
Published: (2026)
Biathlon: Harnessing Model Resilience for Accelerating ML Inference Pipelines
by: Chang, Chaokun, et al.
Published: (2024)
by: Chang, Chaokun, et al.
Published: (2024)
HeadInfer: Memory-Efficient LLM Inference by Head-wise Offloading
by: Luo, Cheng, et al.
Published: (2025)
by: Luo, Cheng, et al.
Published: (2025)
Fast Inference with Kronecker-Sparse Matrices
by: Gonon, Antoine, et al.
Published: (2024)
by: Gonon, Antoine, et al.
Published: (2024)
Impact of ML Optimization Tactics on Greener Pre-Trained ML Models
by: Álvarez, Alexandra González, et al.
Published: (2024)
by: Álvarez, Alexandra González, et al.
Published: (2024)
FRUGAL: Memory-Efficient Optimization by Reducing State Overhead for Scalable Training
by: Zmushko, Philip, et al.
Published: (2024)
by: Zmushko, Philip, et al.
Published: (2024)
Alada: Alternating Adaptation of Momentum Method for Memory-Efficient Matrix Optimization
by: He, Xiaoyu, et al.
Published: (2025)
by: He, Xiaoyu, et al.
Published: (2025)
MISA: Memory-Efficient LLMs Optimization with Module-wise Importance Sampling
by: Liu, Yuxi, et al.
Published: (2025)
by: Liu, Yuxi, et al.
Published: (2025)
Computing Within Limits: An Empirical Study of Energy Consumption in ML Training and Inference
by: Mavromatis, Ioannis, et al.
Published: (2024)
by: Mavromatis, Ioannis, et al.
Published: (2024)
vMCU: Coordinated Memory Management and Kernel Optimization for DNN Inference on MCUs
by: Zheng, Size, et al.
Published: (2024)
by: Zheng, Size, et al.
Published: (2024)
SERFLOW: A Cross-Service Cost Optimization Framework for SLO-Aware Dynamic ML Inference
by: Zhang, Zongshun, et al.
Published: (2025)
by: Zhang, Zongshun, et al.
Published: (2025)
InferF: Declarative Factorization of AI/ML Inferences over Joins
by: Chowdhury, Kanchan, et al.
Published: (2025)
by: Chowdhury, Kanchan, et al.
Published: (2025)
FDC: Fast KV Dimensionality Compression for Efficient LLM Inference
by: Zhang, Zeyu, et al.
Published: (2024)
by: Zhang, Zeyu, et al.
Published: (2024)
Similar Items
-
Data-Efficient Inference of Neural Fluid Fields via SciML Foundation Model
by: Liu, Yuqiu, et al.
Published: (2024) -
Fast Compute for ML Optimization
by: Polson, Nick, et al.
Published: (2026) -
FAST: An Optimization Framework for Fast Additive Segmentation in Transparent ML
by: Liu, Brian, et al.
Published: (2024) -
Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints
by: Ao, Ruicheng, et al.
Published: (2025) -
Toward Robust and Efficient ML-Based GPU Caching for Modern Inference
by: Chen, Peng, et al.
Published: (2025)