Enregistré dans:
| Auteurs principaux: | Rychkov, Valentin, Picoco, Claudia, Caleca, Emilie |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2406.01133 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Modeling Common Cause Failure in Dynamic PRA
par: Picoco, Claudia, et autres
Publié: (2024)
par: Picoco, Claudia, et autres
Publié: (2024)
Green AI: Exploring Carbon Footprints, Mitigation Strategies, and Trade Offs in Large Language Model Training
par: Liu, Vivian, et autres
Publié: (2024)
par: Liu, Vivian, et autres
Publié: (2024)
AI Application Benchmarking: Power-Aware Performance Analysis for Vision and Language Models
par: Mayr, Martin, et autres
Publié: (2026)
par: Mayr, Martin, et autres
Publié: (2026)
Profiling Large Language Model Inference on Apple Silicon: A Quantization Perspective
par: Benazir, Afsara, et autres
Publié: (2025)
par: Benazir, Afsara, et autres
Publié: (2025)
Impact of AI-Triage on Radiologist Report Turnaround Time: Real-World Time-Savings and Insights from Model Predictions
par: Thompson, Yee Lam Elim, et autres
Publié: (2025)
par: Thompson, Yee Lam Elim, et autres
Publié: (2025)
Generalizing Scaling Laws for Dense and Sparse Large Language Models
par: Hossain, Md Arafat, et autres
Publié: (2025)
par: Hossain, Md Arafat, et autres
Publié: (2025)
A Review on Proprietary Accelerators for Large Language Models
par: Park, Sihyeong, et autres
Publié: (2025)
par: Park, Sihyeong, et autres
Publié: (2025)
ALISE: Accelerating Large Language Model Serving with Speculative Scheduling
par: Zhao, Youpeng, et autres
Publié: (2024)
par: Zhao, Youpeng, et autres
Publié: (2024)
SysLLMatic: Large Language Models are Software System Optimizers
par: Peng, Huiyun, et autres
Publié: (2025)
par: Peng, Huiyun, et autres
Publié: (2025)
LFED: A Literary Fiction Evaluation Dataset for Large Language Models
par: Yu, Linhao, et autres
Publié: (2024)
par: Yu, Linhao, et autres
Publié: (2024)
Fairness in Serving Large Language Models
par: Sheng, Ying, et autres
Publié: (2023)
par: Sheng, Ying, et autres
Publié: (2023)
LoPace: A Lossless Optimized Prompt Accurate Compression Engine for Large Language Model Applications
par: Ulla, Aman
Publié: (2026)
par: Ulla, Aman
Publié: (2026)
Priority Sampling of Large Language Models for Compilers
par: Grubisic, Dejan, et autres
Publié: (2024)
par: Grubisic, Dejan, et autres
Publié: (2024)
Should AI Optimize Your Code? A Comparative Study of Classical Optimizing Compilers Versus Current Large Language Models
par: Rosas, Miguel Romero, et autres
Publié: (2024)
par: Rosas, Miguel Romero, et autres
Publié: (2024)
GPU-Accelerated INT8 Quantization for KV Cache Compression in Large Language Models
par: Taneja, Maanas, et autres
Publié: (2026)
par: Taneja, Maanas, et autres
Publié: (2026)
Hardware optimization on Android for inference of AI models
par: Gherasim, Iulius, et autres
Publié: (2025)
par: Gherasim, Iulius, et autres
Publié: (2025)
Information Retrieval in the Age of Generative AI: The RGB Model
par: Garetto, Michele, et autres
Publié: (2025)
par: Garetto, Michele, et autres
Publié: (2025)
Layer Importance and Hallucination Analysis in Large Language Models via Enhanced Activation Variance-Sparsity
par: Song, Zichen, et autres
Publié: (2024)
par: Song, Zichen, et autres
Publié: (2024)
MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models
par: Chitty-Venkata, Krishna Teja, et autres
Publié: (2025)
par: Chitty-Venkata, Krishna Teja, et autres
Publié: (2025)
SimLens for Early Exit in Large Language Models: Eliciting Accurate Latent Predictions with One More Token
par: Ma, Ming, et autres
Publié: (2025)
par: Ma, Ming, et autres
Publié: (2025)
ZO2: Scalable Zeroth-Order Fine-Tuning for Extremely Large Language Models with Limited GPU Memory
par: Wang, Liangyu, et autres
Publié: (2025)
par: Wang, Liangyu, et autres
Publié: (2025)
Modeling the Impact of Fiber Latency on Compute-Communication Overlap in Geo-Distributed Multi-Datacenter AI Training
par: Papavasileiou, Ioannis, et autres
Publié: (2026)
par: Papavasileiou, Ioannis, et autres
Publié: (2026)
What are the key determinants of maintenance performance?
par: Soroush Avakh Darestani
Publié: (2020)
par: Soroush Avakh Darestani
Publié: (2020)
SysOM-AI: Continuous Cross-Layer Performance Diagnosis for Production AI Training
par: Zheng, Yusheng, et autres
Publié: (2026)
par: Zheng, Yusheng, et autres
Publié: (2026)
Deploying Open-Source Large Language Models: A performance Analysis
par: Bendi-Ouis, Yannis, et autres
Publié: (2024)
par: Bendi-Ouis, Yannis, et autres
Publié: (2024)
An Empirical Characterization of Outages and Incidents in Public Services for Large Language Models
par: Chu, Xiaoyu, et autres
Publié: (2025)
par: Chu, Xiaoyu, et autres
Publié: (2025)
Optimas: An Intelligent Analytics-Informed Generative AI Framework for Performance Optimization
par: Zaeed, Mohammad, et autres
Publié: (2026)
par: Zaeed, Mohammad, et autres
Publié: (2026)
SLM-Bench: A Comprehensive Benchmark of Small Language Models on Environmental Impacts--Extended Version
par: Pham, Nghiem Thanh, et autres
Publié: (2025)
par: Pham, Nghiem Thanh, et autres
Publié: (2025)
Approximations to Study the Impact of the Service Discipline in Systems with Redundancy
par: Gast, Nicolas, et autres
Publié: (2024)
par: Gast, Nicolas, et autres
Publié: (2024)
SemaTune: Semantic-Aware Online OS Tuning with Large Language Models
par: Liargkovas, Georgios, et autres
Publié: (2026)
par: Liargkovas, Georgios, et autres
Publié: (2026)
Breaking the Loop: Detecting and Mitigating Denial-of-Service Vulnerabilities in Large Language Models
par: Yu, Junzhe, et autres
Publié: (2025)
par: Yu, Junzhe, et autres
Publié: (2025)
Adaptive Cache Pollution Control for Large Language Model Inference Workloads Using Temporal CNN-Based Prediction and Priority-Aware Replacement
par: Liu, Songze, et autres
Publié: (2025)
par: Liu, Songze, et autres
Publié: (2025)
Benchmarking Quantum Annealers with Near-Optimal Minor-Embedded Instances
par: Gilbert, Valentin, et autres
Publié: (2024)
par: Gilbert, Valentin, et autres
Publié: (2024)
Enhancing Instruction Prefetching via Cache and TLB Management
par: Jamet, Alexandre Valentin, et autres
Publié: (2026)
par: Jamet, Alexandre Valentin, et autres
Publié: (2026)
LLMPerf: GPU Performance Modeling meets Large Language Models
par: Nguyen, Khoi N. M., et autres
Publié: (2025)
par: Nguyen, Khoi N. M., et autres
Publié: (2025)
ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching
par: Zhao, Youpeng, et autres
Publié: (2024)
par: Zhao, Youpeng, et autres
Publié: (2024)
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
par: Zhao, Xuanlei, et autres
Publié: (2024)
par: Zhao, Xuanlei, et autres
Publié: (2024)
CarbonCall: Sustainability-Aware Function Calling for Large Language Models on Edge Devices
par: Paramanayakam, Varatheepan, et autres
Publié: (2025)
par: Paramanayakam, Varatheepan, et autres
Publié: (2025)
Accuracy and Consumption analysis from a compressed model by CompactifAI from Multiverse Computing
par: Fovet, Damien, et autres
Publié: (2025)
par: Fovet, Damien, et autres
Publié: (2025)
Dawn of the Dead(line Misses): Impact of Job Dismiss on the Deadline Miss Rate
par: Chen, Jian-Jia, et autres
Publié: (2024)
par: Chen, Jian-Jia, et autres
Publié: (2024)
Documents similaires
-
Modeling Common Cause Failure in Dynamic PRA
par: Picoco, Claudia, et autres
Publié: (2024) -
Green AI: Exploring Carbon Footprints, Mitigation Strategies, and Trade Offs in Large Language Model Training
par: Liu, Vivian, et autres
Publié: (2024) -
AI Application Benchmarking: Power-Aware Performance Analysis for Vision and Language Models
par: Mayr, Martin, et autres
Publié: (2026) -
Profiling Large Language Model Inference on Apple Silicon: A Quantization Perspective
par: Benazir, Afsara, et autres
Publié: (2025) -
Impact of AI-Triage on Radiologist Report Turnaround Time: Real-World Time-Savings and Insights from Model Predictions
par: Thompson, Yee Lam Elim, et autres
Publié: (2025)