ReLATE: Learning Efficient Sparse Encoding for High-Performance Tensor Decomposition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Helal, Ahmed E., Checconi, Fabio, Laukemann, Jan, Soh, Yongseok, Tithi, Jesmin Jahan, Petrini, Fabrizio, Choi, Jee |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Accelerating Sparse Tensor Decomposition Using Adaptive Linearized Representation
von: Laukemann, Jan, et al.
Veröffentlicht: (2024)
von: Laukemann, Jan, et al.
Veröffentlicht: (2024)
Ridgeline: A 2D Roofline Model for Distributed Systems
von: Checconi, Fabio, et al.
Veröffentlicht: (2022)
von: Checconi, Fabio, et al.
Veröffentlicht: (2022)
Efficient Parallel Multi-Hop Reasoning: A Scalable Approach for Knowledge Graph Analysis
von: Tithi, Jesmin Jahan, et al.
Veröffentlicht: (2024)
von: Tithi, Jesmin Jahan, et al.
Veröffentlicht: (2024)
Scaling Intelligence: Designing Data Centers for Next-Gen Language Models
von: Tithi, Jesmin Jahan, et al.
Veröffentlicht: (2025)
von: Tithi, Jesmin Jahan, et al.
Veröffentlicht: (2025)
MAGNUS: Generating Data Locality to Accelerate Sparse Matrix-Matrix Multiplication on CPUs
von: Wolfson-Pou, Jordi, et al.
Veröffentlicht: (2025)
von: Wolfson-Pou, Jordi, et al.
Veröffentlicht: (2025)
Enhancing Scalability and Performance in Influence Maximization with Optimized Parallel Processing
von: Wu, Hanjiang, et al.
Veröffentlicht: (2024)
von: Wu, Hanjiang, et al.
Veröffentlicht: (2024)
Microarchitectural comparison and in-core modeling of state-of-the-art CPUs: Grace, Sapphire Rapids, and Genoa
von: Laukemann, Jan, et al.
Veröffentlicht: (2024)
von: Laukemann, Jan, et al.
Veröffentlicht: (2024)
CloverLeaf on Intel Multi-Core CPUs: A Case Study in Write-Allocate Evasion
von: Laukemann, Jan, et al.
Veröffentlicht: (2023)
von: Laukemann, Jan, et al.
Veröffentlicht: (2023)
Is Sparse Matrix Reordering Effective for Sparse Matrix-Vector Multiplication?
von: Asudeh, Omid, et al.
Veröffentlicht: (2025)
von: Asudeh, Omid, et al.
Veröffentlicht: (2025)
CoNST: Code Generator for Sparse Tensor Networks
von: Raje, Saurabh, et al.
Veröffentlicht: (2024)
von: Raje, Saurabh, et al.
Veröffentlicht: (2024)
Can Tensor Cores Benefit Memory-Bound Kernels? (No!)
von: Zhang, Lingqi, et al.
Veröffentlicht: (2025)
von: Zhang, Lingqi, et al.
Veröffentlicht: (2025)
Learning-Augmented Performance Model for Tensor Product Factorization in High-Order FEM
von: Ren, Xuanzhengbo, et al.
Veröffentlicht: (2026)
von: Ren, Xuanzhengbo, et al.
Veröffentlicht: (2026)
Staging Blocked Evaluation over Structured Sparse Matrices
von: Das, Pratyush, et al.
Veröffentlicht: (2024)
von: Das, Pratyush, et al.
Veröffentlicht: (2024)
Fast and Scalable Mixed Precision Euclidean Distance Calculations Using GPU Tensor Cores
von: Curless, Brian, et al.
Veröffentlicht: (2025)
von: Curless, Brian, et al.
Veröffentlicht: (2025)
Efficient GPU-Centered Singular Value Decomposition Using the Divide-and-Conquer Method
von: Liu, Shifang, et al.
Veröffentlicht: (2025)
von: Liu, Shifang, et al.
Veröffentlicht: (2025)
Cloud Performance Decomposition for Long-Term Performance Engineering: A Case Study
von: Debnath, Shimul, et al.
Veröffentlicht: (2026)
von: Debnath, Shimul, et al.
Veröffentlicht: (2026)
SHIRO: Near-Optimal Communication Strategies for Distributed Sparse Matrix Multiplication
von: Zhuang, Chen, et al.
Veröffentlicht: (2025)
von: Zhuang, Chen, et al.
Veröffentlicht: (2025)
Architecture Specific Generation of Large Scale Lattice Boltzmann Methods for Sparse Complex Geometries
von: Suffa, Philipp, et al.
Veröffentlicht: (2024)
von: Suffa, Philipp, et al.
Veröffentlicht: (2024)
An Auto-tuning Method for Run-time Data Transformation for Sparse Matrix-Vector Multiplication
von: Katagiri, Takahiro, et al.
Veröffentlicht: (2024)
von: Katagiri, Takahiro, et al.
Veröffentlicht: (2024)
Disaggregated Design for GPU-Based Volumetric Data Structures
von: Meneghin, Massimiliano, et al.
Veröffentlicht: (2025)
von: Meneghin, Massimiliano, et al.
Veröffentlicht: (2025)
Fine-Grained Energy Prediction For Parallellized LLM Inference With PIE-P
von: Dutt, Anurag, et al.
Veröffentlicht: (2025)
von: Dutt, Anurag, et al.
Veröffentlicht: (2025)
Introducing MareNostrum5: A European pre-exascale energy-efficient system designed to serve a broad spectrum of scientific workloads
von: Banchelli, Fabio, et al.
Veröffentlicht: (2025)
von: Banchelli, Fabio, et al.
Veröffentlicht: (2025)
EvalNet: A Practical Toolchain for Generation and Analysis of Extreme-Scale Interconnects
von: Besta, Maciej, et al.
Veröffentlicht: (2021)
von: Besta, Maciej, et al.
Veröffentlicht: (2021)
TaxBreak: Unmasking the Hidden Costs of LLM Inference Through Overhead Decomposition
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2026)
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2026)
Accelerating High-Order Finite Element Simulations at Extreme Scale with FP64 Tensor Cores
von: Tu, Jiqun, et al.
Veröffentlicht: (2026)
von: Tu, Jiqun, et al.
Veröffentlicht: (2026)
Xabclib:A Fully Auto-tuned Sparse Iterative Solver
von: Katagiri, Takahiro, et al.
Veröffentlicht: (2024)
von: Katagiri, Takahiro, et al.
Veröffentlicht: (2024)
Multi-DNN Inference of Sparse Models on Edge SoCs
von: Luo, Jiawei, et al.
Veröffentlicht: (2026)
von: Luo, Jiawei, et al.
Veröffentlicht: (2026)
Extracting Practical, Actionable Energy Insights from Supercomputer Telemetry and Logs
von: Cornelius, Melanie, et al.
Veröffentlicht: (2025)
von: Cornelius, Melanie, et al.
Veröffentlicht: (2025)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
Profiling and optimization of multi-card GPU machine learning jobs
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
Optimal Parallel Scheduling under Concave Speedup Functions
von: Li, Chengzhang, et al.
Veröffentlicht: (2025)
von: Li, Chengzhang, et al.
Veröffentlicht: (2025)
WebAssembly and Unikernels: A Comparative Study for Serverless at the Edge
von: Besozzi, Valerio, et al.
Veröffentlicht: (2025)
von: Besozzi, Valerio, et al.
Veröffentlicht: (2025)
Resource Management Schemes for Cloud-Native Platforms with Computing Containers of Docker and Kubernetes
von: Mao, Ying, et al.
Veröffentlicht: (2020)
von: Mao, Ying, et al.
Veröffentlicht: (2020)
Collaborative Processing for Multi-Tenant Inference on Memory-Constrained Edge TPUs
von: Ng, Nathan, et al.
Veröffentlicht: (2026)
von: Ng, Nathan, et al.
Veröffentlicht: (2026)
Preliminary report: Initial evaluation of StdPar implementations on AMD GPUs for HPC
von: Lin, Wei-Chen, et al.
Veröffentlicht: (2024)
von: Lin, Wei-Chen, et al.
Veröffentlicht: (2024)
Serving Chain-structured Jobs with Large Memory Footprints with Application to Large Foundation Model Serving
von: Sun, Tingyang, et al.
Veröffentlicht: (2026)
von: Sun, Tingyang, et al.
Veröffentlicht: (2026)
Reducing Tail Latencies Through Environment- and Neighbour-aware Thread Management
von: Jeffery, Andrew, et al.
Veröffentlicht: (2024)
von: Jeffery, Andrew, et al.
Veröffentlicht: (2024)
Dissecting the software-based measurement of CPU energy consumption: a comparative analysis
von: Raffin, Guillaume, et al.
Veröffentlicht: (2024)
von: Raffin, Guillaume, et al.
Veröffentlicht: (2024)
Bridding OT and PaaS in Edge-to-Cloud Continuum
von: Barrios, Carlos J, et al.
Veröffentlicht: (2025)
von: Barrios, Carlos J, et al.
Veröffentlicht: (2025)
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
von: Karfakis, George, et al.
Veröffentlicht: (2025)
von: Karfakis, George, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Accelerating Sparse Tensor Decomposition Using Adaptive Linearized Representation
von: Laukemann, Jan, et al.
Veröffentlicht: (2024) -
Ridgeline: A 2D Roofline Model for Distributed Systems
von: Checconi, Fabio, et al.
Veröffentlicht: (2022) -
Efficient Parallel Multi-Hop Reasoning: A Scalable Approach for Knowledge Graph Analysis
von: Tithi, Jesmin Jahan, et al.
Veröffentlicht: (2024) -
Scaling Intelligence: Designing Data Centers for Next-Gen Language Models
von: Tithi, Jesmin Jahan, et al.
Veröffentlicht: (2025) -
MAGNUS: Generating Data Locality to Accelerate Sparse Matrix-Matrix Multiplication on CPUs
von: Wolfson-Pou, Jordi, et al.
Veröffentlicht: (2025)