LEGO: A Layout Expression Language for Code Generation of Hierarchical Mapping
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tavakkoli, Amir Mohammad, Oancea, Cosmin, Hall, Mary |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scheduling Languages: A Past, Present, and Future Taxonomy
von: Hall, Mary, et al.
Veröffentlicht: (2024)
von: Hall, Mary, et al.
Veröffentlicht: (2024)
Comparing Parallel Functional Array Languages: Programming and Performance
von: van Balen, David, et al.
Veröffentlicht: (2025)
von: van Balen, David, et al.
Veröffentlicht: (2025)
Linear Layouts: Robust Code Generation of Efficient Tensor Computation Using $\mathbb{F}_2$
von: Zhou, Keren, et al.
Veröffentlicht: (2025)
von: Zhou, Keren, et al.
Veröffentlicht: (2025)
CoNST: Code Generator for Sparse Tensor Networks
von: Raje, Saurabh, et al.
Veröffentlicht: (2024)
von: Raje, Saurabh, et al.
Veröffentlicht: (2024)
Verifying Properties of Index Arrays in a Purely-Functional Data-Parallel Language
von: Hinnerskov, Nikolaj Hey, et al.
Veröffentlicht: (2025)
von: Hinnerskov, Nikolaj Hey, et al.
Veröffentlicht: (2025)
pSTL-Bench: A Micro-Benchmark Suite for Assessing Scalability of C++ Parallel STL Implementations
von: Laso, Ruben, et al.
Veröffentlicht: (2024)
von: Laso, Ruben, et al.
Veröffentlicht: (2024)
Developing a Modular Compiler for a Subset of a C-like Language
von: Dutta, Debasish, et al.
Veröffentlicht: (2025)
von: Dutta, Debasish, et al.
Veröffentlicht: (2025)
Iterating Pointers: Enabling Static Analysis for Loop-based Pointers
von: Lepori, Andrea, et al.
Veröffentlicht: (2025)
von: Lepori, Andrea, et al.
Veröffentlicht: (2025)
pPython Performance Study
von: Byun, Chansup, et al.
Veröffentlicht: (2023)
von: Byun, Chansup, et al.
Veröffentlicht: (2023)
Simplicity Scales
von: Sampson, Andrew, et al.
Veröffentlicht: (2026)
von: Sampson, Andrew, et al.
Veröffentlicht: (2026)
Towards a Linear-Algebraic Hypervisor
von: Considine, Breandan
Veröffentlicht: (2026)
von: Considine, Breandan
Veröffentlicht: (2026)
GPU Implementations for Midsize Integer Addition and Multiplication
von: Oancea, Cosmin E., et al.
Veröffentlicht: (2024)
von: Oancea, Cosmin E., et al.
Veröffentlicht: (2024)
LOOPRAG: Enhancing Loop Transformation Optimization with Retrieval-Augmented Large Language Models
von: Zhi, Yijie, et al.
Veröffentlicht: (2025)
von: Zhi, Yijie, et al.
Veröffentlicht: (2025)
Massimult: A Novel Parallel CPU Architecture Based on Combinator Reduction
von: Nicklisch-Franken, Jurgen, et al.
Veröffentlicht: (2024)
von: Nicklisch-Franken, Jurgen, et al.
Veröffentlicht: (2024)
TAPA: A Scalable Task-Parallel Dataflow Programming Framework for Modern FPGAs with Co-Optimization of HLS and Physical Design
von: Guo, Licheng, et al.
Veröffentlicht: (2022)
von: Guo, Licheng, et al.
Veröffentlicht: (2022)
Agentic Auto-Scheduling: An Experimental Study of LLM-Guided Loop Optimization
von: Merouani, Massinissa, et al.
Veröffentlicht: (2025)
von: Merouani, Massinissa, et al.
Veröffentlicht: (2025)
Portable High-Performance Kernel Generation for a Computational Fluid Dynamics Code with DaCe
von: Andersson, Måns I., et al.
Veröffentlicht: (2025)
von: Andersson, Måns I., et al.
Veröffentlicht: (2025)
CodeRosetta: Pushing the Boundaries of Unsupervised Code Translation for Parallel Programming
von: TehraniJamsaz, Ali, et al.
Veröffentlicht: (2024)
von: TehraniJamsaz, Ali, et al.
Veröffentlicht: (2024)
CUTHERMO: Understanding GPU Memory Inefficiencies with Heat Map Profiling
von: Zhao, Yanbo, et al.
Veröffentlicht: (2025)
von: Zhao, Yanbo, et al.
Veröffentlicht: (2025)
PolyTOPS: Reconfigurable and Flexible Polyhedral Scheduler
von: Consolaro, Gianpietro, et al.
Veröffentlicht: (2024)
von: Consolaro, Gianpietro, et al.
Veröffentlicht: (2024)
Safe Memory Reclamation Techniques
von: Singh, Ajay
Veröffentlicht: (2025)
von: Singh, Ajay
Veröffentlicht: (2025)
Unified schemes for directive-based GPU offloading
von: Miki, Yohei, et al.
Veröffentlicht: (2024)
von: Miki, Yohei, et al.
Veröffentlicht: (2024)
Optimizing Agentic Language Model Inference via Speculative Tool Calls
von: Nichols, Daniel, et al.
Veröffentlicht: (2025)
von: Nichols, Daniel, et al.
Veröffentlicht: (2025)
Data-Driven Analysis to Understand GPU Hardware Resource Usage of Optimizations
von: Islam, Tanzima Z., et al.
Veröffentlicht: (2024)
von: Islam, Tanzima Z., et al.
Veröffentlicht: (2024)
Orthrus: Accelerating Multi-BFT Consensus through Concurrent Partial Ordering of Transactions (Extended Version)
von: Lyu, Hanzheng, et al.
Veröffentlicht: (2024)
von: Lyu, Hanzheng, et al.
Veröffentlicht: (2024)
An Empirical Characterization of Outages and Incidents in Public Services for Large Language Models
von: Chu, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Chu, Xiaoyu, et al.
Veröffentlicht: (2025)
Mapple: A Domain-Specific Language for Mapping Distributed Programs
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
Efficient Serverless Cold Start: Reducing Library Loading Overhead by Profile-guided Optimization
von: Tariq, Syed Salauddin Mohammad, et al.
Veröffentlicht: (2025)
von: Tariq, Syed Salauddin Mohammad, et al.
Veröffentlicht: (2025)
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
Cost-Performance Evaluation of General Compute Instances: AWS, Azure, GCP, and OCI
von: Tharwani, Jay, et al.
Veröffentlicht: (2024)
von: Tharwani, Jay, et al.
Veröffentlicht: (2024)
Performance Evaluation of a Next-Generation SX-Aurora TSUBASA Vector Supercomputer
von: Takahashi, Keichi, et al.
Veröffentlicht: (2023)
von: Takahashi, Keichi, et al.
Veröffentlicht: (2023)
Architecture Specific Generation of Large Scale Lattice Boltzmann Methods for Sparse Complex Geometries
von: Suffa, Philipp, et al.
Veröffentlicht: (2024)
von: Suffa, Philipp, et al.
Veröffentlicht: (2024)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
von: Lin, Mao, et al.
Veröffentlicht: (2026)
von: Lin, Mao, et al.
Veröffentlicht: (2026)
EPIC: Efficient Position-Independent Caching for Serving Large Language Models
von: Hu, Junhao, et al.
Veröffentlicht: (2024)
von: Hu, Junhao, et al.
Veröffentlicht: (2024)
Can Large Language Models Predict Parallel Code Performance?
von: Bolet, Gregory, et al.
Veröffentlicht: (2025)
von: Bolet, Gregory, et al.
Veröffentlicht: (2025)
OMPILOT: Harnessing Transformer Models for Auto Parallelization to Shared Memory Computing Paradigms
von: Bhattacharjee, Arijit, et al.
Veröffentlicht: (2025)
von: Bhattacharjee, Arijit, et al.
Veröffentlicht: (2025)
A Performance Analysis of BFT Consensus for Blockchains
von: Chan, J. D., et al.
Veröffentlicht: (2024)
von: Chan, J. D., et al.
Veröffentlicht: (2024)
WebAssembly and Unikernels: A Comparative Study for Serverless at the Edge
von: Besozzi, Valerio, et al.
Veröffentlicht: (2025)
von: Besozzi, Valerio, et al.
Veröffentlicht: (2025)
Sampling in Cloud Benchmarking: A Critical Review and Methodological Guidelines
von: Akbari, Saman, et al.
Veröffentlicht: (2025)
von: Akbari, Saman, et al.
Veröffentlicht: (2025)
PASTA: A Modular Program Analysis Tool Framework for Accelerators
von: Lin, Mao, et al.
Veröffentlicht: (2026)
von: Lin, Mao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Scheduling Languages: A Past, Present, and Future Taxonomy
von: Hall, Mary, et al.
Veröffentlicht: (2024) -
Comparing Parallel Functional Array Languages: Programming and Performance
von: van Balen, David, et al.
Veröffentlicht: (2025) -
Linear Layouts: Robust Code Generation of Efficient Tensor Computation Using $\mathbb{F}_2$
von: Zhou, Keren, et al.
Veröffentlicht: (2025) -
CoNST: Code Generator for Sparse Tensor Networks
von: Raje, Saurabh, et al.
Veröffentlicht: (2024) -
Verifying Properties of Index Arrays in a Purely-Functional Data-Parallel Language
von: Hinnerskov, Nikolaj Hey, et al.
Veröffentlicht: (2025)