Performance Modeling of Data Storage Systems using Generative Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Al-Maeeni, Abdalaziz Rashid, Temirkhanov, Aziz, Ryzhikov, Artem, Hushchyn, Mikhail |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Accelerating AI Performance using Anderson Extrapolation on GPUs
von: Dajani, Saleem Abdul Fattah Ahmed Al, et al.
Veröffentlicht: (2024)
von: Dajani, Saleem Abdul Fattah Ahmed Al, et al.
Veröffentlicht: (2024)
Generalizing Scaling Laws for Dense and Sparse Large Language Models
von: Hossain, Md Arafat, et al.
Veröffentlicht: (2025)
von: Hossain, Md Arafat, et al.
Veröffentlicht: (2025)
Quantum Neural Networks for Wind Energy Forecasting: A Comparative Study of Performance and Scalability with Classical Models
von: Hangun, Batuhan, et al.
Veröffentlicht: (2025)
von: Hangun, Batuhan, et al.
Veröffentlicht: (2025)
Automatic Generation of Fast and Accurate Performance Models for Deep Neural Network Accelerators
von: Lübeck, Konstantin, et al.
Veröffentlicht: (2024)
von: Lübeck, Konstantin, et al.
Veröffentlicht: (2024)
LLM Swiss Round: Aggregating Multi-Benchmark Performance via Competitive Swiss-System Dynamics
von: Liu, Jiashuo, et al.
Veröffentlicht: (2025)
von: Liu, Jiashuo, et al.
Veröffentlicht: (2025)
Research on Low-Latency Inference and Training Efficiency Optimization for Graph Neural Network and Large Language Model-Based Recommendation Systems
von: Zhao, Yushang, et al.
Veröffentlicht: (2025)
von: Zhao, Yushang, et al.
Veröffentlicht: (2025)
Leveraging Speculative Sampling and KV-Cache Optimizations Together for Generative AI using OpenVINO
von: Barad, Haim, et al.
Veröffentlicht: (2023)
von: Barad, Haim, et al.
Veröffentlicht: (2023)
Fairness in Serving Large Language Models
von: Sheng, Ying, et al.
Veröffentlicht: (2023)
von: Sheng, Ying, et al.
Veröffentlicht: (2023)
Data Efficacy for Language Model Training
von: Dai, Yalun, et al.
Veröffentlicht: (2025)
von: Dai, Yalun, et al.
Veröffentlicht: (2025)
APOLLO: SGD-like Memory, AdamW-level Performance
von: Zhu, Hanqing, et al.
Veröffentlicht: (2024)
von: Zhu, Hanqing, et al.
Veröffentlicht: (2024)
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
von: Shao, Zishan, et al.
Veröffentlicht: (2025)
von: Shao, Zishan, et al.
Veröffentlicht: (2025)
Deploying Open-Source Large Language Models: A performance Analysis
von: Bendi-Ouis, Yannis, et al.
Veröffentlicht: (2024)
von: Bendi-Ouis, Yannis, et al.
Veröffentlicht: (2024)
Applied Federated Model Personalisation in the Industrial Domain: A Comparative Study
von: Siniosoglou, Ilias, et al.
Veröffentlicht: (2024)
von: Siniosoglou, Ilias, et al.
Veröffentlicht: (2024)
Forecasting LLM Inference Performance via Hardware-Agnostic Analytical Modeling
von: Patwari, Rajeev, et al.
Veröffentlicht: (2025)
von: Patwari, Rajeev, et al.
Veröffentlicht: (2025)
Knowledge Grafting: A Mechanism for Optimizing AI Model Deployment in Resource-Constrained Environments
von: Almurshed, Osama, et al.
Veröffentlicht: (2025)
von: Almurshed, Osama, et al.
Veröffentlicht: (2025)
ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching
von: Zhao, Youpeng, et al.
Veröffentlicht: (2024)
von: Zhao, Youpeng, et al.
Veröffentlicht: (2024)
Ensuring Reliability of Curated EHR-Derived Data: The Validation of Accuracy for LLM/ML-Extracted Information and Data (VALID) Framework
von: Estevez, Melissa, et al.
Veröffentlicht: (2025)
von: Estevez, Melissa, et al.
Veröffentlicht: (2025)
Ragged Paged Attention: A High-Performance and Flexible LLM Inference Kernel for TPU
von: Jiang, Jevin, et al.
Veröffentlicht: (2026)
von: Jiang, Jevin, et al.
Veröffentlicht: (2026)
Rapid Augmentations for Time Series (RATS): A High-Performance Library for Time Series Augmentation
von: Skaf, Wadie, et al.
Veröffentlicht: (2026)
von: Skaf, Wadie, et al.
Veröffentlicht: (2026)
Energy per Successful Goal: Goal-Level Energy Accounting for Agentic AI Systems
von: Panigrahy, Deepak, et al.
Veröffentlicht: (2026)
von: Panigrahy, Deepak, et al.
Veröffentlicht: (2026)
ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference
von: Yin, Wangsong, et al.
Veröffentlicht: (2025)
von: Yin, Wangsong, et al.
Veröffentlicht: (2025)
Reducing Latency of LLM Search Agent via Speculation-based Algorithm-System Co-Design
von: Huang, Zixiao, et al.
Veröffentlicht: (2025)
von: Huang, Zixiao, et al.
Veröffentlicht: (2025)
Towards Generalized Parameter Tuning in Coherent Ising Machines: A Portfolio-Based Approach
von: Hanyu, Tatsuro, et al.
Veröffentlicht: (2025)
von: Hanyu, Tatsuro, et al.
Veröffentlicht: (2025)
It's all about PR -- Smart Benchmarking AI Accelerators using Performance Representatives
von: Jung, Alexander Louis-Ferdinand, et al.
Veröffentlicht: (2024)
von: Jung, Alexander Louis-Ferdinand, et al.
Veröffentlicht: (2024)
Throughput Optimization as a Strategic Lever in Large-Scale AI Systems: Evidence from Dataloader and Memory Profiling Innovations
von: Jha, Mayank
Veröffentlicht: (2026)
von: Jha, Mayank
Veröffentlicht: (2026)
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management
von: Zhang, Hang, et al.
Veröffentlicht: (2025)
von: Zhang, Hang, et al.
Veröffentlicht: (2025)
A Performance Evaluation of a Quantized Large Language Model on Various Smartphones
von: Çöplü, Tolga, et al.
Veröffentlicht: (2023)
von: Çöplü, Tolga, et al.
Veröffentlicht: (2023)
Model Compression and Efficient Inference for Large Language Models: A Survey
von: Wang, Wenxiao, et al.
Veröffentlicht: (2024)
von: Wang, Wenxiao, et al.
Veröffentlicht: (2024)
Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems
von: Miao, Xupeng, et al.
Veröffentlicht: (2023)
von: Miao, Xupeng, et al.
Veröffentlicht: (2023)
Predictive Modeling of I/O Performance for Machine Learning Training Pipelines: A Data-Driven Approach to Storage Optimization
von: Prabhakar, Karthik, et al.
Veröffentlicht: (2025)
von: Prabhakar, Karthik, et al.
Veröffentlicht: (2025)
Enabling Performant and Flexible Model-Internal Observability for LLM Inference
von: Yu, Nengneng, et al.
Veröffentlicht: (2026)
von: Yu, Nengneng, et al.
Veröffentlicht: (2026)
Learning Performance-Improving Code Edits
von: Shypula, Alexander, et al.
Veröffentlicht: (2023)
von: Shypula, Alexander, et al.
Veröffentlicht: (2023)
The Hidden Power of Pure 16-bit Floating-Point Neural Networks
von: Yun, Juyoung, et al.
Veröffentlicht: (2023)
von: Yun, Juyoung, et al.
Veröffentlicht: (2023)
Machine Learning Methods for Evaluating Public Crisis: Meta-Analysis
von: Okpala, Izunna, et al.
Veröffentlicht: (2023)
von: Okpala, Izunna, et al.
Veröffentlicht: (2023)
SLiM: One-shot Quantization and Sparsity with Low-rank Approximation for LLM Weight Compression
von: Mozaffari, Mohammad, et al.
Veröffentlicht: (2024)
von: Mozaffari, Mohammad, et al.
Veröffentlicht: (2024)
Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs
von: Yun, Vincent-Daniel, et al.
Veröffentlicht: (2026)
von: Yun, Vincent-Daniel, et al.
Veröffentlicht: (2026)
GreedySnake: Accelerating SSD-Offloaded LLM Training with Efficient Scheduling and Optimizer Step Overlapping
von: Yin, Yishu, et al.
Veröffentlicht: (2025)
von: Yin, Yishu, et al.
Veröffentlicht: (2025)
EXAQ: Exponent Aware Quantization For LLMs Acceleration
von: Shkolnik, Moran, et al.
Veröffentlicht: (2024)
von: Shkolnik, Moran, et al.
Veröffentlicht: (2024)
Energy-Efficient Transformer Inference: Optimization Strategies for Time Series Classification
von: Kermani, Arshia, et al.
Veröffentlicht: (2025)
von: Kermani, Arshia, et al.
Veröffentlicht: (2025)
Private LLM Inference on Consumer Blackwell GPUs: A Practical Guide for Cost-Effective Local Deployment in SMEs
von: Knoop, Jonathan, et al.
Veröffentlicht: (2026)
von: Knoop, Jonathan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Accelerating AI Performance using Anderson Extrapolation on GPUs
von: Dajani, Saleem Abdul Fattah Ahmed Al, et al.
Veröffentlicht: (2024) -
Generalizing Scaling Laws for Dense and Sparse Large Language Models
von: Hossain, Md Arafat, et al.
Veröffentlicht: (2025) -
Quantum Neural Networks for Wind Energy Forecasting: A Comparative Study of Performance and Scalability with Classical Models
von: Hangun, Batuhan, et al.
Veröffentlicht: (2025) -
Automatic Generation of Fast and Accurate Performance Models for Deep Neural Network Accelerators
von: Lübeck, Konstantin, et al.
Veröffentlicht: (2024) -
LLM Swiss Round: Aggregating Multi-Benchmark Performance via Competitive Swiss-System Dynamics
von: Liu, Jiashuo, et al.
Veröffentlicht: (2025)