CoNST: Code Generator for Sparse Tensor Networks
Fuente:
arXiv
Guardado en:
| Autores principales: | Raje, Saurabh, Xu, Yufan, Rountev, Atanas, Valeev, Edward F., Sadayappan, Saday |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LEGO: A Layout Expression Language for Code Generation of Hierarchical Mapping
por: Tavakkoli, Amir Mohammad, et al.
Publicado: (2025)
por: Tavakkoli, Amir Mohammad, et al.
Publicado: (2025)
Is Sparse Matrix Reordering Effective for Sparse Matrix-Vector Multiplication?
por: Asudeh, Omid, et al.
Publicado: (2025)
por: Asudeh, Omid, et al.
Publicado: (2025)
Linear Layouts: Robust Code Generation of Efficient Tensor Computation Using $\mathbb{F}_2$
por: Zhou, Keren, et al.
Publicado: (2025)
por: Zhou, Keren, et al.
Publicado: (2025)
Comparing Parallel Functional Array Languages: Programming and Performance
por: van Balen, David, et al.
Publicado: (2025)
por: van Balen, David, et al.
Publicado: (2025)
pPython Performance Study
por: Byun, Chansup, et al.
Publicado: (2023)
por: Byun, Chansup, et al.
Publicado: (2023)
Simplicity Scales
por: Sampson, Andrew, et al.
Publicado: (2026)
por: Sampson, Andrew, et al.
Publicado: (2026)
pSTL-Bench: A Micro-Benchmark Suite for Assessing Scalability of C++ Parallel STL Implementations
por: Laso, Ruben, et al.
Publicado: (2024)
por: Laso, Ruben, et al.
Publicado: (2024)
Scheduling Languages: A Past, Present, and Future Taxonomy
por: Hall, Mary, et al.
Publicado: (2024)
por: Hall, Mary, et al.
Publicado: (2024)
Iterating Pointers: Enabling Static Analysis for Loop-based Pointers
por: Lepori, Andrea, et al.
Publicado: (2025)
por: Lepori, Andrea, et al.
Publicado: (2025)
Towards a Linear-Algebraic Hypervisor
por: Considine, Breandan
Publicado: (2026)
por: Considine, Breandan
Publicado: (2026)
TAPA: A Scalable Task-Parallel Dataflow Programming Framework for Modern FPGAs with Co-Optimization of HLS and Physical Design
por: Guo, Licheng, et al.
Publicado: (2022)
por: Guo, Licheng, et al.
Publicado: (2022)
Developing a Modular Compiler for a Subset of a C-like Language
por: Dutta, Debasish, et al.
Publicado: (2025)
por: Dutta, Debasish, et al.
Publicado: (2025)
Minimum Cost Loop Nests for Contraction of a Sparse Tensor with a Tensor Network
por: Kanakagiri, Raghavendra, et al.
Publicado: (2023)
por: Kanakagiri, Raghavendra, et al.
Publicado: (2023)
Agentic Auto-Scheduling: An Experimental Study of LLM-Guided Loop Optimization
por: Merouani, Massinissa, et al.
Publicado: (2025)
por: Merouani, Massinissa, et al.
Publicado: (2025)
LOOPRAG: Enhancing Loop Transformation Optimization with Retrieval-Augmented Large Language Models
por: Zhi, Yijie, et al.
Publicado: (2025)
por: Zhi, Yijie, et al.
Publicado: (2025)
Massimult: A Novel Parallel CPU Architecture Based on Combinator Reduction
por: Nicklisch-Franken, Jurgen, et al.
Publicado: (2024)
por: Nicklisch-Franken, Jurgen, et al.
Publicado: (2024)
eBPF-Based Instrumentation for Generalisable Diagnosis of Performance Degradation
por: Landau, Diogo, et al.
Publicado: (2025)
por: Landau, Diogo, et al.
Publicado: (2025)
Architecture Specific Generation of Large Scale Lattice Boltzmann Methods for Sparse Complex Geometries
por: Suffa, Philipp, et al.
Publicado: (2024)
por: Suffa, Philipp, et al.
Publicado: (2024)
Beyond Thread States: Diagnosing Performance Degradation with eBPF and Thread Dynamics
por: Landau, Diogo, et al.
Publicado: (2026)
por: Landau, Diogo, et al.
Publicado: (2026)
Can Tensor Cores Benefit Memory-Bound Kernels? (No!)
por: Zhang, Lingqi, et al.
Publicado: (2025)
por: Zhang, Lingqi, et al.
Publicado: (2025)
Learning-Augmented Performance Model for Tensor Product Factorization in High-Order FEM
por: Ren, Xuanzhengbo, et al.
Publicado: (2026)
por: Ren, Xuanzhengbo, et al.
Publicado: (2026)
Portable High-Performance Kernel Generation for a Computational Fluid Dynamics Code with DaCe
por: Andersson, Måns I., et al.
Publicado: (2025)
por: Andersson, Måns I., et al.
Publicado: (2025)
Staging Blocked Evaluation over Structured Sparse Matrices
por: Das, Pratyush, et al.
Publicado: (2024)
por: Das, Pratyush, et al.
Publicado: (2024)
Fast and Scalable Mixed Precision Euclidean Distance Calculations Using GPU Tensor Cores
por: Curless, Brian, et al.
Publicado: (2025)
por: Curless, Brian, et al.
Publicado: (2025)
ReLATE: Learning Efficient Sparse Encoding for High-Performance Tensor Decomposition
por: Helal, Ahmed E., et al.
Publicado: (2025)
por: Helal, Ahmed E., et al.
Publicado: (2025)
SHIRO: Near-Optimal Communication Strategies for Distributed Sparse Matrix Multiplication
por: Zhuang, Chen, et al.
Publicado: (2025)
por: Zhuang, Chen, et al.
Publicado: (2025)
An Auto-tuning Method for Run-time Data Transformation for Sparse Matrix-Vector Multiplication
por: Katagiri, Takahiro, et al.
Publicado: (2024)
por: Katagiri, Takahiro, et al.
Publicado: (2024)
Unified schemes for directive-based GPU offloading
por: Miki, Yohei, et al.
Publicado: (2024)
por: Miki, Yohei, et al.
Publicado: (2024)
Safe Memory Reclamation Techniques
por: Singh, Ajay
Publicado: (2025)
por: Singh, Ajay
Publicado: (2025)
Accelerating Sparse Tensor Decomposition Using Adaptive Linearized Representation
por: Laukemann, Jan, et al.
Publicado: (2024)
por: Laukemann, Jan, et al.
Publicado: (2024)
CARAT: Client-Side Adaptive RPC and Cache Co-Tuning for Parallel File Systems
por: Rashid, Md Hasanur, et al.
Publicado: (2026)
por: Rashid, Md Hasanur, et al.
Publicado: (2026)
PolyTOPS: Reconfigurable and Flexible Polyhedral Scheduler
por: Consolaro, Gianpietro, et al.
Publicado: (2024)
por: Consolaro, Gianpietro, et al.
Publicado: (2024)
Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
por: Zhang, Yaozheng, et al.
Publicado: (2025)
por: Zhang, Yaozheng, et al.
Publicado: (2025)
Towards Stream-Based Monitoring for EVM Networks
por: Onica, Emanuel, et al.
Publicado: (2025)
por: Onica, Emanuel, et al.
Publicado: (2025)
CodeRosetta: Pushing the Boundaries of Unsupervised Code Translation for Parallel Programming
por: TehraniJamsaz, Ali, et al.
Publicado: (2024)
por: TehraniJamsaz, Ali, et al.
Publicado: (2024)
Toward Scalable Docker-Based Emulations of Blockchain Networks for Research and Development
por: Pennino, Diego, et al.
Publicado: (2024)
por: Pennino, Diego, et al.
Publicado: (2024)
Cost-Performance Evaluation of General Compute Instances: AWS, Azure, GCP, and OCI
por: Tharwani, Jay, et al.
Publicado: (2024)
por: Tharwani, Jay, et al.
Publicado: (2024)
Performance Evaluation of a Next-Generation SX-Aurora TSUBASA Vector Supercomputer
por: Takahashi, Keichi, et al.
Publicado: (2023)
por: Takahashi, Keichi, et al.
Publicado: (2023)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
por: Lin, Mao, et al.
Publicado: (2026)
por: Lin, Mao, et al.
Publicado: (2026)
OMPILOT: Harnessing Transformer Models for Auto Parallelization to Shared Memory Computing Paradigms
por: Bhattacharjee, Arijit, et al.
Publicado: (2025)
por: Bhattacharjee, Arijit, et al.
Publicado: (2025)
Ejemplares similares
-
LEGO: A Layout Expression Language for Code Generation of Hierarchical Mapping
por: Tavakkoli, Amir Mohammad, et al.
Publicado: (2025) -
Is Sparse Matrix Reordering Effective for Sparse Matrix-Vector Multiplication?
por: Asudeh, Omid, et al.
Publicado: (2025) -
Linear Layouts: Robust Code Generation of Efficient Tensor Computation Using $\mathbb{F}_2$
por: Zhou, Keren, et al.
Publicado: (2025) -
Comparing Parallel Functional Array Languages: Programming and Performance
por: van Balen, David, et al.
Publicado: (2025) -
pPython Performance Study
por: Byun, Chansup, et al.
Publicado: (2023)