Inf-MLLM: Efficient Streaming Inference of Multimodal Large Language Models on a Single GPU
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ning, Zhenyu, Zhao, Jieru, Jin, Qihao, Ding, Wenchao, Guo, Minyi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
von: Lin, Mao, et al.
Veröffentlicht: (2026)
von: Lin, Mao, et al.
Veröffentlicht: (2026)
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
Efficient allocation of image recognition and LLM tasks on multi-GPU system
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
Efficient GPU-Centered Singular Value Decomposition Using the Divide-and-Conquer Method
von: Liu, Shifang, et al.
Veröffentlicht: (2025)
von: Liu, Shifang, et al.
Veröffentlicht: (2025)
CUTHERMO: Understanding GPU Memory Inefficiencies with Heat Map Profiling
von: Zhao, Yanbo, et al.
Veröffentlicht: (2025)
von: Zhao, Yanbo, et al.
Veröffentlicht: (2025)
ADELIA: Automatic Differentiation for Efficient Laplace Inference Approximations
von: Boudaoud, Afif, et al.
Veröffentlicht: (2026)
von: Boudaoud, Afif, et al.
Veröffentlicht: (2026)
The Energy Cost of Execution-Idle in GPU Clusters
von: Lei, Yiran, et al.
Veröffentlicht: (2026)
von: Lei, Yiran, et al.
Veröffentlicht: (2026)
Scalable GPU Performance Variability Analysis framework
von: Lahiry, Ankur, et al.
Veröffentlicht: (2025)
von: Lahiry, Ankur, et al.
Veröffentlicht: (2025)
On the Partitioning of GPU Power among Multi-Instances
von: Vamja, Tirth, et al.
Veröffentlicht: (2025)
von: Vamja, Tirth, et al.
Veröffentlicht: (2025)
Disaggregated Design for GPU-Based Volumetric Data Structures
von: Meneghin, Massimiliano, et al.
Veröffentlicht: (2025)
von: Meneghin, Massimiliano, et al.
Veröffentlicht: (2025)
Taking GPU Programming Models to Task for Performance Portability
von: Davis, Joshua H., et al.
Veröffentlicht: (2024)
von: Davis, Joshua H., et al.
Veröffentlicht: (2024)
Performance Optimization in Stream Processing Systems: Experiment-Driven Configuration Tuning for Kafka Streams
von: Chen, David, et al.
Veröffentlicht: (2026)
von: Chen, David, et al.
Veröffentlicht: (2026)
Profiling and optimization of multi-card GPU machine learning jobs
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
KEET: Explaining Performance of GPU Kernels Using LLM Agents
von: Davis, Joshua H., et al.
Veröffentlicht: (2026)
von: Davis, Joshua H., et al.
Veröffentlicht: (2026)
High-Performance Portable GPU Primitives for Arbitrary Types and Operators in Julia
von: Pilliat, Emmanuel
Veröffentlicht: (2026)
von: Pilliat, Emmanuel
Veröffentlicht: (2026)
Data-Driven Analysis to Understand GPU Hardware Resource Usage of Optimizations
von: Islam, Tanzima Z., et al.
Veröffentlicht: (2024)
von: Islam, Tanzima Z., et al.
Veröffentlicht: (2024)
LLMPerf: GPU Performance Modeling meets Large Language Models
von: Nguyen, Khoi N. M., et al.
Veröffentlicht: (2025)
von: Nguyen, Khoi N. M., et al.
Veröffentlicht: (2025)
Minos: Systematically Classifying Performance and Power Characteristics of GPU Workloads on HPC Clusters
von: Jain, Rutwik, et al.
Veröffentlicht: (2026)
von: Jain, Rutwik, et al.
Veröffentlicht: (2026)
Unleashing the Power of Preemptive Priority-based Scheduling for Real-Time GPU Tasks
von: Wang, Yidi, et al.
Veröffentlicht: (2024)
von: Wang, Yidi, et al.
Veröffentlicht: (2024)
Towards Stream-Based Monitoring for EVM Networks
von: Onica, Emanuel, et al.
Veröffentlicht: (2025)
von: Onica, Emanuel, et al.
Veröffentlicht: (2025)
Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2025)
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2025)
LEO: Tracing GPU Stall Root Causes via Cross-Vendor Backward Slicing
von: Xia, Yuning, et al.
Veröffentlicht: (2026)
von: Xia, Yuning, et al.
Veröffentlicht: (2026)
Fast and Scalable Mixed Precision Euclidean Distance Calculations Using GPU Tensor Cores
von: Curless, Brian, et al.
Veröffentlicht: (2025)
von: Curless, Brian, et al.
Veröffentlicht: (2025)
Optimizing Near Field Computation in the MLFMA Algorithm with Data Redundancy and Performance Modeling on a Single GPU
von: Sadeghi, Morteza, et al.
Veröffentlicht: (2024)
von: Sadeghi, Morteza, et al.
Veröffentlicht: (2024)
SProBench: Stream Processing Benchmark for High Performance Computing Infrastructure
von: Kulkarni, Apurv Deepak, et al.
Veröffentlicht: (2025)
von: Kulkarni, Apurv Deepak, et al.
Veröffentlicht: (2025)
Towards Portability at Scale: A Cross-Architecture Performance Evaluation of a GPU-enabled Shallow Water Solver
von: Villalobos, Johansell, et al.
Veröffentlicht: (2025)
von: Villalobos, Johansell, et al.
Veröffentlicht: (2025)
Cyclic Data Streaming on GPUs for Short Range Stencils Applied to Molecular Dynamics
von: Rose, Martin, et al.
Veröffentlicht: (2025)
von: Rose, Martin, et al.
Veröffentlicht: (2025)
FalconFS: Distributed File System for Large-Scale Deep Learning Pipeline
von: Xu, Jingwei, et al.
Veröffentlicht: (2025)
von: Xu, Jingwei, et al.
Veröffentlicht: (2025)
Characterizing WebGPU Dispatch Overhead for LLM Inference Across Four GPU Vendors, Three Backends, and Three Browsers
von: Maczan, Jędrzej
Veröffentlicht: (2026)
von: Maczan, Jędrzej
Veröffentlicht: (2026)
InkStream: Real-time GNN Inference on Streaming Graphs via Incremental Update
von: Wu, Dan, et al.
Veröffentlicht: (2023)
von: Wu, Dan, et al.
Veröffentlicht: (2023)
An Empirical Characterization of Outages and Incidents in Public Services for Large Language Models
von: Chu, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Chu, Xiaoyu, et al.
Veröffentlicht: (2025)
Accelerating Mobile Inference through Fine-Grained CPU-GPU Co-Execution
von: Li, Zhuojin, et al.
Veröffentlicht: (2025)
von: Li, Zhuojin, et al.
Veröffentlicht: (2025)
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
von: Zhang, Li, et al.
Veröffentlicht: (2025)
von: Zhang, Li, et al.
Veröffentlicht: (2025)
Collaborative Processing for Multi-Tenant Inference on Memory-Constrained Edge TPUs
von: Ng, Nathan, et al.
Veröffentlicht: (2026)
von: Ng, Nathan, et al.
Veröffentlicht: (2026)
Fine-Grained Energy Prediction For Parallellized LLM Inference With PIE-P
von: Dutt, Anurag, et al.
Veröffentlicht: (2025)
von: Dutt, Anurag, et al.
Veröffentlicht: (2025)
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles
von: Arif, Moiz, et al.
Veröffentlicht: (2026)
von: Arif, Moiz, et al.
Veröffentlicht: (2026)
Kairos: Efficient Temporal Graph Analytics on a Single Machine
von: da Trindade, Joana M. F., et al.
Veröffentlicht: (2024)
von: da Trindade, Joana M. F., et al.
Veröffentlicht: (2024)
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
von: Karfakis, George, et al.
Veröffentlicht: (2025)
von: Karfakis, George, et al.
Veröffentlicht: (2025)
Towards Universal Performance Modeling for Machine Learning Training on Multi-GPU Platforms
von: Lin, Zhongyi, et al.
Veröffentlicht: (2024)
von: Lin, Zhongyi, et al.
Veröffentlicht: (2024)
Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
von: Zhang, Yaozheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yaozheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
von: Lin, Mao, et al.
Veröffentlicht: (2026) -
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024) -
Efficient allocation of image recognition and LLM tasks on multi-GPU system
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025) -
Efficient GPU-Centered Singular Value Decomposition Using the Divide-and-Conquer Method
von: Liu, Shifang, et al.
Veröffentlicht: (2025) -
CUTHERMO: Understanding GPU Memory Inefficiencies with Heat Map Profiling
von: Zhao, Yanbo, et al.
Veröffentlicht: (2025)