Benchmarking Compound AI Applications for Hardware-Software Co-Design
Fuente:
arXiv
Salvato in:
| Autori principali: | Samuthrsindh, Paramuth, Cervantes, Angel, Gohil, Varun, Chaudhry, Gohar Irfan, Delimitrou, Christina, Belay, Adam |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Towards Resource-Efficient Compound AI Systems
di: Chaudhry, Gohar Irfan, et al.
Pubblicazione: (2025)
di: Chaudhry, Gohar Irfan, et al.
Pubblicazione: (2025)
Junctiond: Extending FaaS Runtimes with Kernel-Bypass
di: Saurez, Enrique, et al.
Pubblicazione: (2024)
di: Saurez, Enrique, et al.
Pubblicazione: (2024)
The Sunk Carbon Fallacy: Rethinking Carbon Footprint Metrics for Effective Carbon-Aware Scheduling
di: Bashir, Noman, et al.
Pubblicazione: (2024)
di: Bashir, Noman, et al.
Pubblicazione: (2024)
Analytically-Driven Resource Management for Cloud-Native Microservices
di: Zhang, Yanqi, et al.
Pubblicazione: (2024)
di: Zhang, Yanqi, et al.
Pubblicazione: (2024)
Taming Serverless Cold Starts Through OS Co-Design
di: Holmes, Ben, et al.
Pubblicazione: (2025)
di: Holmes, Ben, et al.
Pubblicazione: (2025)
StreamWise: Serving Multi-Modal Generation in Real-Time at Scale
di: Qiu, Haoran, et al.
Pubblicazione: (2026)
di: Qiu, Haoran, et al.
Pubblicazione: (2026)
Fire-Flyer AI-HPC: A Cost-Effective Software-Hardware Co-Design for Deep Learning
di: An, Wei, et al.
Pubblicazione: (2024)
di: An, Wei, et al.
Pubblicazione: (2024)
Carbon- and Precedence-Aware Scheduling for Data Processing Clusters
di: Lechowicz, Adam, et al.
Pubblicazione: (2025)
di: Lechowicz, Adam, et al.
Pubblicazione: (2025)
TT-Edge: A Hardware-Software Co-Design for Energy-Efficient Tensor-Train Decomposition on Edge AI
di: Kwak, Hyunseok, et al.
Pubblicazione: (2025)
di: Kwak, Hyunseok, et al.
Pubblicazione: (2025)
Checkmate: Zero-Overhead Model Checkpointing via Network Gradient Replication
di: Bhardwaj, Ankit, et al.
Pubblicazione: (2025)
di: Bhardwaj, Ankit, et al.
Pubblicazione: (2025)
Lumos: Efficient Performance Modeling and Estimation for Large-scale LLM Training
di: Liang, Mingyu, et al.
Pubblicazione: (2025)
di: Liang, Mingyu, et al.
Pubblicazione: (2025)
The Case for Co-Designing Model Architectures with Hardware
di: Anthony, Quentin, et al.
Pubblicazione: (2024)
di: Anthony, Quentin, et al.
Pubblicazione: (2024)
TokenSim: Enabling Hardware and Software Exploration for Large Language Model Inference Systems
di: Wu, Feiyang, et al.
Pubblicazione: (2025)
di: Wu, Feiyang, et al.
Pubblicazione: (2025)
Coordinated Cooling and Compute Management for AI Datacenters
di: Abera, Nardos Belay, et al.
Pubblicazione: (2026)
di: Abera, Nardos Belay, et al.
Pubblicazione: (2026)
GPUArmor: A Hardware-Software Co-design for Efficient and Scalable Memory Safety on GPUs
di: Ziad, Mohamed Tarek Ibn, et al.
Pubblicazione: (2025)
di: Ziad, Mohamed Tarek Ibn, et al.
Pubblicazione: (2025)
LLMSched: Uncertainty-Aware Workload Scheduling for Compound LLM Applications
di: Zhu, Botao, et al.
Pubblicazione: (2025)
di: Zhu, Botao, et al.
Pubblicazione: (2025)
xNVMe: Unleashing Storage Hardware-Software Co-design
di: Lund, Simon A. F., et al.
Pubblicazione: (2024)
di: Lund, Simon A. F., et al.
Pubblicazione: (2024)
A Distributed Partitioning Software and its Applications
di: Sasidharan, Aparna
Pubblicazione: (2025)
di: Sasidharan, Aparna
Pubblicazione: (2025)
Gaia: Hybrid Hardware Acceleration for Serverless AI in the 3D Compute Continuum
di: Reisecker, Maximilian, et al.
Pubblicazione: (2025)
di: Reisecker, Maximilian, et al.
Pubblicazione: (2025)
AgentServe: Algorithm-System Co-Design for Efficient Agentic AI Serving on a Consumer-Grade GPU
di: Zhang, Yuning, et al.
Pubblicazione: (2026)
di: Zhang, Yuning, et al.
Pubblicazione: (2026)
DynaSplit: A Hardware-Software Co-Design Framework for Energy-Aware Inference on Edge
di: May, Daniel, et al.
Pubblicazione: (2024)
di: May, Daniel, et al.
Pubblicazione: (2024)
NasZip: Software and Hardware Co-Design to Accelerate Approximate Nearest Neighbor Search with DIMM-Based Near-Data Processing
di: Zou, Cheng, et al.
Pubblicazione: (2026)
di: Zou, Cheng, et al.
Pubblicazione: (2026)
Benchmarking Machine Learning Applications on Heterogeneous Architecture using Reframe
di: Rae, Christopher, et al.
Pubblicazione: (2024)
di: Rae, Christopher, et al.
Pubblicazione: (2024)
Increasing Efficiency and Result Reliability of Continuous Benchmarking for FaaS Applications
di: Rese, Tim C., et al.
Pubblicazione: (2024)
di: Rese, Tim C., et al.
Pubblicazione: (2024)
PVU: Design and Implementation of a Posit Vector Arithmetic Unit (PVU) for Enhanced Floating-Point Computing in Edge and AI Applications
di: Wu, Xinyu, et al.
Pubblicazione: (2025)
di: Wu, Xinyu, et al.
Pubblicazione: (2025)
LLM-CoOpt: A Co-Design and Optimization Framework for Efficient LLM Inference on Heterogeneous Platforms
di: Kong, Jie, et al.
Pubblicazione: (2026)
di: Kong, Jie, et al.
Pubblicazione: (2026)
Benchmarking Different Application Types across Heterogeneous Cloud Compute Services
di: Duggi, Nivedhitha, et al.
Pubblicazione: (2025)
di: Duggi, Nivedhitha, et al.
Pubblicazione: (2025)
Application-Centric Benchmarking of Distributed FaaS Platforms using BeFaaS
di: Grambow, Martin, et al.
Pubblicazione: (2023)
di: Grambow, Martin, et al.
Pubblicazione: (2023)
Towards Energy-Efficient Serverless Computing with Hardware Isolation
di: Carl, Natalie, et al.
Pubblicazione: (2025)
di: Carl, Natalie, et al.
Pubblicazione: (2025)
CARM Tool: Cache-Aware Roofline Model Automatic Benchmarking and Application Analysis
di: Morgado, José, et al.
Pubblicazione: (2026)
di: Morgado, José, et al.
Pubblicazione: (2026)
AI-coupled HPC Workflow Applications, Middleware and Performance
di: Brewer, Wes, et al.
Pubblicazione: (2024)
di: Brewer, Wes, et al.
Pubblicazione: (2024)
Sustainable AI Training via Hardware-Software Co-Design on NVIDIA, AMD, and Emerging GPU Architectures
di: Makin, Yashasvi, et al.
Pubblicazione: (2025)
di: Makin, Yashasvi, et al.
Pubblicazione: (2025)
PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers
di: Yeo, Gwangoo, et al.
Pubblicazione: (2024)
di: Yeo, Gwangoo, et al.
Pubblicazione: (2024)
Accelerating Compound LLM Training Workloads with Maestro
di: Yuan, Xiulong, et al.
Pubblicazione: (2026)
di: Yuan, Xiulong, et al.
Pubblicazione: (2026)
Serving Compound Inference Systems on Datacenter GPUs
di: Devata, Sriram, et al.
Pubblicazione: (2026)
di: Devata, Sriram, et al.
Pubblicazione: (2026)
Extracting the Potential of Emerging Hardware Accelerators for Symmetric Eigenvalue Decomposition
di: Wang, Hansheng, et al.
Pubblicazione: (2024)
di: Wang, Hansheng, et al.
Pubblicazione: (2024)
SIGMA: An AI-Empowered Training Stack on Early-Life Hardware
di: Qu, Lei, et al.
Pubblicazione: (2025)
di: Qu, Lei, et al.
Pubblicazione: (2025)
Designing Co-operation in Systems of Hierarchical, Multi-objective Schedulers for Stream Processing
di: Dangwal, Animesh, et al.
Pubblicazione: (2025)
di: Dangwal, Animesh, et al.
Pubblicazione: (2025)
FlashSketch: Sketch-Kernel Co-Design for Fast Sparse Sketching on GPUs
di: Dwaraknath, Rajat Vadiraj, et al.
Pubblicazione: (2026)
di: Dwaraknath, Rajat Vadiraj, et al.
Pubblicazione: (2026)
Edge System Design Using Containers and Unikernels for IoT Applications
di: Kaiser, Shahidullah, et al.
Pubblicazione: (2024)
di: Kaiser, Shahidullah, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Towards Resource-Efficient Compound AI Systems
di: Chaudhry, Gohar Irfan, et al.
Pubblicazione: (2025) -
Junctiond: Extending FaaS Runtimes with Kernel-Bypass
di: Saurez, Enrique, et al.
Pubblicazione: (2024) -
The Sunk Carbon Fallacy: Rethinking Carbon Footprint Metrics for Effective Carbon-Aware Scheduling
di: Bashir, Noman, et al.
Pubblicazione: (2024) -
Analytically-Driven Resource Management for Cloud-Native Microservices
di: Zhang, Yanqi, et al.
Pubblicazione: (2024) -
Taming Serverless Cold Starts Through OS Co-Design
di: Holmes, Ben, et al.
Pubblicazione: (2025)