Salvato in:
| Autori principali: | Peccia, F. N., Bringmann, O. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2022
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2212.03034 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LOOPerSet: A Large-Scale Dataset for Data-Driven Polyhedral Compiler Optimization
di: Merouani, Massinissa, et al.
Pubblicazione: (2025)
di: Merouani, Massinissa, et al.
Pubblicazione: (2025)
DaCe AD: Unifying High-Performance Automatic Differentiation for Machine Learning and Scientific Computing
di: Boudaoud, Afif, et al.
Pubblicazione: (2025)
di: Boudaoud, Afif, et al.
Pubblicazione: (2025)
Flex Attention: A Programming Model for Generating Optimized Attention Kernels
di: Dong, Juechu, et al.
Pubblicazione: (2024)
di: Dong, Juechu, et al.
Pubblicazione: (2024)
The Next 700 ML-Enabled Compiler Optimizations
di: VenkataKeerthy, S., et al.
Pubblicazione: (2023)
di: VenkataKeerthy, S., et al.
Pubblicazione: (2023)
Who Wins the Race? (R Vs Python) - An Exploratory Study on Energy Consumption of Machine Learning Algorithms
di: Chattaraj, Rajrupa, et al.
Pubblicazione: (2025)
di: Chattaraj, Rajrupa, et al.
Pubblicazione: (2025)
Library Liberation: Competitive Performance Matmul Through Compiler-composed Nanokernels
di: Thangamani, Arun, et al.
Pubblicazione: (2025)
di: Thangamani, Arun, et al.
Pubblicazione: (2025)
ECCO: Evidence-Driven Causal Reasoning for Compiler Optimization
di: Pan, Haolin, et al.
Pubblicazione: (2026)
di: Pan, Haolin, et al.
Pubblicazione: (2026)
Agentic Auto-Scheduling: An Experimental Study of LLM-Guided Loop Optimization
di: Merouani, Massinissa, et al.
Pubblicazione: (2025)
di: Merouani, Massinissa, et al.
Pubblicazione: (2025)
Worst-Case Convergence Time of ML Algorithms via Extreme Value Theory
di: Tizpaz-Niari, Saeid, et al.
Pubblicazione: (2024)
di: Tizpaz-Niari, Saeid, et al.
Pubblicazione: (2024)
Automatic Generation of Fast and Accurate Performance Models for Deep Neural Network Accelerators
di: Lübeck, Konstantin, et al.
Pubblicazione: (2024)
di: Lübeck, Konstantin, et al.
Pubblicazione: (2024)
V-Seek: Accelerating LLM Reasoning on Open-hardware Server-class RISC-V Platforms
di: Rodrigo, Javier J. Poveda, et al.
Pubblicazione: (2025)
di: Rodrigo, Javier J. Poveda, et al.
Pubblicazione: (2025)
ModeSwitch-LLM: A Lightweight Phase-Aware Controller for Cross-Mode LLM Inference on a Single GPU
di: Sunesh, Aman, et al.
Pubblicazione: (2026)
di: Sunesh, Aman, et al.
Pubblicazione: (2026)
OMPILOT: Harnessing Transformer Models for Auto Parallelization to Shared Memory Computing Paradigms
di: Bhattacharjee, Arijit, et al.
Pubblicazione: (2025)
di: Bhattacharjee, Arijit, et al.
Pubblicazione: (2025)
oneDNN Graph Compiler: A Hybrid Approach for High-Performance Deep Learning Compilation
di: Li, Jianhui, et al.
Pubblicazione: (2023)
di: Li, Jianhui, et al.
Pubblicazione: (2023)
FLuRKA: Fast and accurate unified Low-Rank & Kernel Attention
di: Gupta, Ahan, et al.
Pubblicazione: (2023)
di: Gupta, Ahan, et al.
Pubblicazione: (2023)
Priority Sampling of Large Language Models for Compilers
di: Grubisic, Dejan, et al.
Pubblicazione: (2024)
di: Grubisic, Dejan, et al.
Pubblicazione: (2024)
Systematic Evaluation of Optimization Techniques for Long-Context Language Models
di: Ahmed, Ammar, et al.
Pubblicazione: (2025)
di: Ahmed, Ammar, et al.
Pubblicazione: (2025)
KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
di: Liu, Zirui, et al.
Pubblicazione: (2024)
di: Liu, Zirui, et al.
Pubblicazione: (2024)
Block Sparse Flash Attention
di: Ohayon, Daniel, et al.
Pubblicazione: (2025)
di: Ohayon, Daniel, et al.
Pubblicazione: (2025)
AttentionEngine: A Versatile Framework for Efficient Attention Mechanisms on Diverse Hardware Platforms
di: Chen, Feiyang, et al.
Pubblicazione: (2025)
di: Chen, Feiyang, et al.
Pubblicazione: (2025)
An energy-based comparative analysis of common approaches to text classification in the Legal domain
di: Gultekin, Sinan, et al.
Pubblicazione: (2023)
di: Gultekin, Sinan, et al.
Pubblicazione: (2023)
CodeRosetta: Pushing the Boundaries of Unsupervised Code Translation for Parallel Programming
di: TehraniJamsaz, Ali, et al.
Pubblicazione: (2024)
di: TehraniJamsaz, Ali, et al.
Pubblicazione: (2024)
Multi-DNN Inference of Sparse Models on Edge SoCs
di: Luo, Jiawei, et al.
Pubblicazione: (2026)
di: Luo, Jiawei, et al.
Pubblicazione: (2026)
DistZO2: High-Throughput and Memory-Efficient Zeroth-Order Fine-tuning LLMs with Distributed Parallel Computing
di: Wang, Liangyu, et al.
Pubblicazione: (2025)
di: Wang, Liangyu, et al.
Pubblicazione: (2025)
Optimizing Layout of Recursive Datatypes with Marmoset
di: Singhal, Vidush, et al.
Pubblicazione: (2024)
di: Singhal, Vidush, et al.
Pubblicazione: (2024)
Evaluating Compiler Optimization Impacts on zkVM Performance
di: Gassmann, Thomas, et al.
Pubblicazione: (2025)
di: Gassmann, Thomas, et al.
Pubblicazione: (2025)
Insum: Sparse GPU Kernels Simplified and Optimized with Indirect Einsums
di: Won, Jaeyeon, et al.
Pubblicazione: (2025)
di: Won, Jaeyeon, et al.
Pubblicazione: (2025)
It's Not Easy Being Green: On the Energy Efficiency of Programming Languages
di: van Kempen, Nicolas, et al.
Pubblicazione: (2024)
di: van Kempen, Nicolas, et al.
Pubblicazione: (2024)
Rule-Based Graph Programs Matching the Time Complexity of Imperative Algorithms
di: Alaoui, Ziad Ismaili, et al.
Pubblicazione: (2025)
di: Alaoui, Ziad Ismaili, et al.
Pubblicazione: (2025)
Quamba2: A Robust and Scalable Post-training Quantization Framework for Selective State Space Models
di: Chiang, Hung-Yueh, et al.
Pubblicazione: (2025)
di: Chiang, Hung-Yueh, et al.
Pubblicazione: (2025)
Energy-Aware LLMs: A step towards sustainable AI for downstream applications
di: Tran, Nguyen Phuc, et al.
Pubblicazione: (2025)
di: Tran, Nguyen Phuc, et al.
Pubblicazione: (2025)
QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
di: Lin, Yujun, et al.
Pubblicazione: (2024)
di: Lin, Yujun, et al.
Pubblicazione: (2024)
REAM: Merging Improves Pruning of Experts in LLMs
di: Jha, Saurav, et al.
Pubblicazione: (2026)
di: Jha, Saurav, et al.
Pubblicazione: (2026)
QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead
di: Zandieh, Amir, et al.
Pubblicazione: (2024)
di: Zandieh, Amir, et al.
Pubblicazione: (2024)
LowRA: Accurate and Efficient LoRA Fine-Tuning of LLMs under 2 Bits
di: Zhou, Zikai, et al.
Pubblicazione: (2025)
di: Zhou, Zikai, et al.
Pubblicazione: (2025)
Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes
di: Hendria, Willy Fitra
Pubblicazione: (2026)
di: Hendria, Willy Fitra
Pubblicazione: (2026)
Evaluating the Efficacy of Foundational Models: Advancing Benchmarking Practices to Enhance Fine-Tuning Decision-Making
di: Amujo, Oluyemi Enoch, et al.
Pubblicazione: (2024)
di: Amujo, Oluyemi Enoch, et al.
Pubblicazione: (2024)
Data Efficacy for Language Model Training
di: Dai, Yalun, et al.
Pubblicazione: (2025)
di: Dai, Yalun, et al.
Pubblicazione: (2025)
Bench360: Benchmarking Local LLM Inference from 360 Degrees
di: Stuhlmann, Linus, et al.
Pubblicazione: (2025)
di: Stuhlmann, Linus, et al.
Pubblicazione: (2025)
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing
di: Zheng, Wenhao, et al.
Pubblicazione: (2025)
di: Zheng, Wenhao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
LOOPerSet: A Large-Scale Dataset for Data-Driven Polyhedral Compiler Optimization
di: Merouani, Massinissa, et al.
Pubblicazione: (2025) -
DaCe AD: Unifying High-Performance Automatic Differentiation for Machine Learning and Scientific Computing
di: Boudaoud, Afif, et al.
Pubblicazione: (2025) -
Flex Attention: A Programming Model for Generating Optimized Attention Kernels
di: Dong, Juechu, et al.
Pubblicazione: (2024) -
The Next 700 ML-Enabled Compiler Optimizations
di: VenkataKeerthy, S., et al.
Pubblicazione: (2023) -
Who Wins the Race? (R Vs Python) - An Exploratory Study on Energy Consumption of Machine Learning Algorithms
di: Chattaraj, Rajrupa, et al.
Pubblicazione: (2025)