Data-efficient Performance Modeling via Pre-training
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Chunting, Baghdadi, Riyadh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Agentic Auto-Scheduling: An Experimental Study of LLM-Guided Loop Optimization
by: Merouani, Massinissa, et al.
Published: (2025)
by: Merouani, Massinissa, et al.
Published: (2025)
LOOPer: A Learned Automatic Code Optimizer For Polyhedral Compilers
by: Merouani, Massinissa, et al.
Published: (2024)
by: Merouani, Massinissa, et al.
Published: (2024)
VTC: DNN Compilation with Virtual Tensors for Data Movement Elimination
by: Hu, Muyan, et al.
Published: (2026)
by: Hu, Muyan, et al.
Published: (2026)
Learned Cost Model for Placement on Reconfigurable Dataflow Hardware
by: Guha, Etash, et al.
Published: (2025)
by: Guha, Etash, et al.
Published: (2025)
Theoretical Foundations of GPU-Native Compilation for Rapid Code Iteration
by: Metinov, Adilet, et al.
Published: (2025)
by: Metinov, Adilet, et al.
Published: (2025)
Mirage Persistent Kernel: A Compiler and Runtime for Mega-Kernelizing Tensor Programs
by: Cheng, Xinhao, et al.
Published: (2025)
by: Cheng, Xinhao, et al.
Published: (2025)
Morphling: Fast, Fused, and Flexible GNN Training at Scale
by: Anubhab, et al.
Published: (2025)
by: Anubhab, et al.
Published: (2025)
veScale: Consistent and Efficient Tensor Programming with Eager-Mode SPMD
by: Li, Youjie, et al.
Published: (2025)
by: Li, Youjie, et al.
Published: (2025)
Scaling Deep Learning Training with MPMD Pipeline Parallelism
by: Xhebraj, Anxhelo, et al.
Published: (2024)
by: Xhebraj, Anxhelo, et al.
Published: (2024)
GPU-Accelerated Synthesis of Mixed-Boolean Arithmetic: Beyond Caching
by: Bathie, Gabriel, et al.
Published: (2026)
by: Bathie, Gabriel, et al.
Published: (2026)
PartIR: Composing SPMD Partitioning Strategies for Machine Learning
by: Alabed, Sami, et al.
Published: (2024)
by: Alabed, Sami, et al.
Published: (2024)
Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel
by: Jin, Hongyi, et al.
Published: (2026)
by: Jin, Hongyi, et al.
Published: (2026)
A Reinforcement Learning Environment for Automatic Code Optimization in the MLIR Compiler
by: Tirichine, Mohammed, et al.
Published: (2024)
by: Tirichine, Mohammed, et al.
Published: (2024)
FLMarket: Enabling Privacy-preserved Pre-training Data Pricing for Federated Learning
by: Wen, Zhenyu, et al.
Published: (2024)
by: Wen, Zhenyu, et al.
Published: (2024)
GPT-FL: Generative Pre-trained Model-Assisted Federated Learning
by: Zhang, Tuo, et al.
Published: (2023)
by: Zhang, Tuo, et al.
Published: (2023)
Vertical Federated Learning Hybrid Local Pre-training
by: Li, Wenguo, et al.
Published: (2024)
by: Li, Wenguo, et al.
Published: (2024)
Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-Flow
by: Mei, Yixuan, et al.
Published: (2024)
by: Mei, Yixuan, et al.
Published: (2024)
LLM-Pilot: Characterize and Optimize Performance of your LLM Inference Services
by: Łazuka, Małgorzata, et al.
Published: (2024)
by: Łazuka, Małgorzata, et al.
Published: (2024)
Distributed Locking as a Data Type
by: Haas, Julian, et al.
Published: (2024)
by: Haas, Julian, et al.
Published: (2024)
MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators
by: Zhang, Zheng, et al.
Published: (2025)
by: Zhang, Zheng, et al.
Published: (2025)
Efficient and Adaptable Overlapping for Computation and Communication via Signaling and Reordering
by: Hong, Ke, et al.
Published: (2025)
by: Hong, Ke, et al.
Published: (2025)
Towards Resiliency in Large Language Model Serving with KevlarFlow
by: Qian, Shangshu, et al.
Published: (2026)
by: Qian, Shangshu, et al.
Published: (2026)
Sal: Multi-modal Verification of Replicated Data Types
by: Ramesh, Pranav, et al.
Published: (2026)
by: Ramesh, Pranav, et al.
Published: (2026)
Concurrent Data Structures Made Easy (Extended Version)
by: Le, Callista, et al.
Published: (2024)
by: Le, Callista, et al.
Published: (2024)
Data Driven Optimization of GPU efficiency for Distributed LLM Adapter Serving
by: Agullo, Ferran, et al.
Published: (2026)
by: Agullo, Ferran, et al.
Published: (2026)
KPerfIR: Towards an Open and Compiler-centric Ecosystem for GPU Kernel Performance Tooling on Modern AI Workloads
by: Guan, Yue, et al.
Published: (2025)
by: Guan, Yue, et al.
Published: (2025)
The Future of Large Language Model Pre-training is Federated
by: Sani, Lorenzo, et al.
Published: (2024)
by: Sani, Lorenzo, et al.
Published: (2024)
PRDTs: Composable Knowledge-Based Consensus Protocols with Replicated Data Types
by: Haas, Julian, et al.
Published: (2025)
by: Haas, Julian, et al.
Published: (2025)
Verifying Properties of Index Arrays in a Purely-Functional Data-Parallel Language
by: Hinnerskov, Nikolaj Hey, et al.
Published: (2025)
by: Hinnerskov, Nikolaj Hey, et al.
Published: (2025)
Optimizing RLHF Training for Large Language Models with Stage Fusion
by: Zhong, Yinmin, et al.
Published: (2024)
by: Zhong, Yinmin, et al.
Published: (2024)
Federated Learning of Large Language Models with Parameter-Efficient Prompt Tuning and Adaptive Optimization
by: Che, Tianshi, et al.
Published: (2023)
by: Che, Tianshi, et al.
Published: (2023)
Scalable Training of Mixture-of-Experts Models with Megatron Core
by: Yan, Zijie, et al.
Published: (2026)
by: Yan, Zijie, et al.
Published: (2026)
Lobster: A GPU-Accelerated Framework for Neurosymbolic Programming
by: Biberstein, Paul, et al.
Published: (2025)
by: Biberstein, Paul, et al.
Published: (2025)
Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference
by: Yang, Mengtian, et al.
Published: (2026)
by: Yang, Mengtian, et al.
Published: (2026)
Axe: A Simple Unified Layout Abstraction for Machine Learning Compilers
by: Hou, Bohan, et al.
Published: (2026)
by: Hou, Bohan, et al.
Published: (2026)
Reward Augmentation in Reinforcement Learning for Testing Distributed Systems
by: Borgarelli, Andrea, et al.
Published: (2024)
by: Borgarelli, Andrea, et al.
Published: (2024)
Publish on Ping: A Better Way to Publish Reservations in Memory Reclamation for Concurrent Data Structures
by: Singh, Ajay, et al.
Published: (2025)
by: Singh, Ajay, et al.
Published: (2025)
An MLIR pipeline for offloading Fortran to FPGAs via OpenMP
by: Rodriguez-Canal, Gabriel, et al.
Published: (2025)
by: Rodriguez-Canal, Gabriel, et al.
Published: (2025)
P/D-Device: Disaggregated Large Language Model between Cloud and Devices
by: Jin, Yibo, et al.
Published: (2025)
by: Jin, Yibo, et al.
Published: (2025)
P/D-Serve: Serving Disaggregated Large Language Model at Scale
by: Jin, Yibo, et al.
Published: (2024)
by: Jin, Yibo, et al.
Published: (2024)
Similar Items
-
Agentic Auto-Scheduling: An Experimental Study of LLM-Guided Loop Optimization
by: Merouani, Massinissa, et al.
Published: (2025) -
LOOPer: A Learned Automatic Code Optimizer For Polyhedral Compilers
by: Merouani, Massinissa, et al.
Published: (2024) -
VTC: DNN Compilation with Virtual Tensors for Data Movement Elimination
by: Hu, Muyan, et al.
Published: (2026) -
Learned Cost Model for Placement on Reconfigurable Dataflow Hardware
by: Guha, Etash, et al.
Published: (2025) -
Theoretical Foundations of GPU-Native Compilation for Rapid Code Iteration
by: Metinov, Adilet, et al.
Published: (2025)