Approximating Uniform Random Rotations by Two-Block Structured Hadamard Rotations in High Dimensions
Fuente:
arXiv
Salvato in:
| Autori principali: | Zilca, Tomer, Mendelson, Gal |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Block Sparse Flash Attention
di: Ohayon, Daniel, et al.
Pubblicazione: (2025)
di: Ohayon, Daniel, et al.
Pubblicazione: (2025)
Feature Optimization for Time Series Forecasting via Novel Randomized Uphill Climbing
di: Van Thanh, Nguyen
Pubblicazione: (2025)
di: Van Thanh, Nguyen
Pubblicazione: (2025)
Search Your Block Floating Point Scales!
di: Gupta, Tanmaey, et al.
Pubblicazione: (2026)
di: Gupta, Tanmaey, et al.
Pubblicazione: (2026)
Leveraging Approximate Caching for Faster Retrieval-Augmented Generation
di: Bergman, Shai, et al.
Pubblicazione: (2025)
di: Bergman, Shai, et al.
Pubblicazione: (2025)
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization
di: Zhou, Zhongzhu, et al.
Pubblicazione: (2026)
di: Zhou, Zhongzhu, et al.
Pubblicazione: (2026)
USEFUSE: Uniform Stride for Enhanced Performance in Fused Layer Architecture of Deep Neural Networks
di: Ibrahim, Muhammad Sohail, et al.
Pubblicazione: (2024)
di: Ibrahim, Muhammad Sohail, et al.
Pubblicazione: (2024)
Approximating Heavy-Tailed Distributions with a Mixture of Bernstein Phase-Type and Hyperexponential Models
di: Ziani, Abdelhakim, et al.
Pubblicazione: (2025)
di: Ziani, Abdelhakim, et al.
Pubblicazione: (2025)
A Structure-Aware Framework for Learning Device Placements on Computation Graphs
di: Duan, Shukai, et al.
Pubblicazione: (2024)
di: Duan, Shukai, et al.
Pubblicazione: (2024)
oneDNN Graph Compiler: A Hybrid Approach for High-Performance Deep Learning Compilation
di: Li, Jianhui, et al.
Pubblicazione: (2023)
di: Li, Jianhui, et al.
Pubblicazione: (2023)
DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
di: Holmes, Connor, et al.
Pubblicazione: (2024)
di: Holmes, Connor, et al.
Pubblicazione: (2024)
DistZO2: High-Throughput and Memory-Efficient Zeroth-Order Fine-tuning LLMs with Distributed Parallel Computing
di: Wang, Liangyu, et al.
Pubblicazione: (2025)
di: Wang, Liangyu, et al.
Pubblicazione: (2025)
SLiM: One-shot Quantization and Sparsity with Low-rank Approximation for LLM Weight Compression
di: Mozaffari, Mohammad, et al.
Pubblicazione: (2024)
di: Mozaffari, Mohammad, et al.
Pubblicazione: (2024)
PoTAcc: A Pipeline for End-to-End Acceleration of Power-of-Two Quantized DNNs
di: Saha, Rappy, et al.
Pubblicazione: (2026)
di: Saha, Rappy, et al.
Pubblicazione: (2026)
GPU-Accelerated INT8 Quantization for KV Cache Compression in Large Language Models
di: Taneja, Maanas, et al.
Pubblicazione: (2026)
di: Taneja, Maanas, et al.
Pubblicazione: (2026)
MarginGate: Sparse Margin-Triggered Verification for Batch-Invariant LLM Inference
di: Chu, Kexin, et al.
Pubblicazione: (2026)
di: Chu, Kexin, et al.
Pubblicazione: (2026)
Reducing Compute Waste in LLMs through Kernel-Level DVFS
di: Spaan, Jeffrey, et al.
Pubblicazione: (2026)
di: Spaan, Jeffrey, et al.
Pubblicazione: (2026)
KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels
di: Wang, Han, et al.
Pubblicazione: (2026)
di: Wang, Han, et al.
Pubblicazione: (2026)
Single-Thread JPEG Decoder Benchmarks Mis-Evaluate ML Data Loaders
di: Iglovikov, Vladimir, et al.
Pubblicazione: (2026)
di: Iglovikov, Vladimir, et al.
Pubblicazione: (2026)
An Interpretable Latency Model for Speculative Decoding in LLM Serving
di: Kong, Linghao, et al.
Pubblicazione: (2026)
di: Kong, Linghao, et al.
Pubblicazione: (2026)
GreenServ: Energy-Efficient Context-Aware Dynamic Routing for Multi-Model LLM Inference
di: Ziller, Thomas, et al.
Pubblicazione: (2026)
di: Ziller, Thomas, et al.
Pubblicazione: (2026)
EARL: Energy-Aware Optimization of Liquid State Machines for Pervasive AI
di: Iqbal, Zain, et al.
Pubblicazione: (2026)
di: Iqbal, Zain, et al.
Pubblicazione: (2026)
Large-Scale Data Parallelization of Product Quantization and Inverted Indexing Using Dask
di: Abraham, Ashley N., et al.
Pubblicazione: (2026)
di: Abraham, Ashley N., et al.
Pubblicazione: (2026)
AutoKernel: Autonomous GPU Kernel Optimization via Iterative Agent-Driven Search
di: Jaber, Jaber, et al.
Pubblicazione: (2026)
di: Jaber, Jaber, et al.
Pubblicazione: (2026)
Heuristic-Based Merging of HPC Traces to Extend Hardware Counter Coverage
di: Aubach, Júlia Orteu, et al.
Pubblicazione: (2026)
di: Aubach, Júlia Orteu, et al.
Pubblicazione: (2026)
A Scalable k-Medoids Clustering via Whale Optimization Algorithm
di: Chenan, Huang, et al.
Pubblicazione: (2024)
di: Chenan, Huang, et al.
Pubblicazione: (2024)
CPINN-ABPI: Physics-Informed Neural Networks for Accurate Power Estimation in MPSoCs
di: Elshamy, Mohamed R., et al.
Pubblicazione: (2025)
di: Elshamy, Mohamed R., et al.
Pubblicazione: (2025)
PARD: Accelerating LLM Inference with Low-Cost PARallel Draft Model Adaptation
di: An, Zihao, et al.
Pubblicazione: (2025)
di: An, Zihao, et al.
Pubblicazione: (2025)
Flashlight: PyTorch Compiler Extensions to Accelerate Attention Variants
di: You, Bozhi, et al.
Pubblicazione: (2025)
di: You, Bozhi, et al.
Pubblicazione: (2025)
Parallel Implementations Assessment of a Spatial-Spectral Classifier for Hyperspectral Clinical Applications
di: Lazcano, Raquel, et al.
Pubblicazione: (2024)
di: Lazcano, Raquel, et al.
Pubblicazione: (2024)
Automating Energy-Efficient GPU Kernel Generation: A Fast Search-Based Compilation Approach
di: Zhang, Yijia, et al.
Pubblicazione: (2024)
di: Zhang, Yijia, et al.
Pubblicazione: (2024)
LiveTune: Dynamic Parameter Tuning for Feedback-Driven Optimization
di: Shabgahi, Soheil Zibakhsh, et al.
Pubblicazione: (2023)
di: Shabgahi, Soheil Zibakhsh, et al.
Pubblicazione: (2023)
Enhancing Tropical Cyclone Path Forecasting with an Improved Transformer Network
di: Van Thanh, Nguyen, et al.
Pubblicazione: (2025)
di: Van Thanh, Nguyen, et al.
Pubblicazione: (2025)
lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
di: Wang, Haoxin, et al.
Pubblicazione: (2025)
di: Wang, Haoxin, et al.
Pubblicazione: (2025)
WCDT: Systematic WCET Optimization for Decision Tree Implementations
di: Hölscher, Nils, et al.
Pubblicazione: (2025)
di: Hölscher, Nils, et al.
Pubblicazione: (2025)
BanditQ: Fair Bandits with Guaranteed Rewards
di: Sinha, Abhishek
Pubblicazione: (2023)
di: Sinha, Abhishek
Pubblicazione: (2023)
Plug-and-Play Performance Estimation for LLM Services without Relying on Labeled Data
di: Wang, Can, et al.
Pubblicazione: (2024)
di: Wang, Can, et al.
Pubblicazione: (2024)
MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models
di: Chitty-Venkata, Krishna Teja, et al.
Pubblicazione: (2025)
di: Chitty-Venkata, Krishna Teja, et al.
Pubblicazione: (2025)
MLPerf Automotive
di: Shojaei, Radoyeh, et al.
Pubblicazione: (2025)
di: Shojaei, Radoyeh, et al.
Pubblicazione: (2025)
SAfEPaTh: A System-Level Approach for Efficient Power and Thermal Estimation of Convolutional Neural Network Accelerator
di: Chen, Yukai, et al.
Pubblicazione: (2024)
di: Chen, Yukai, et al.
Pubblicazione: (2024)
Towards Computational Performance Engineering for Unsupervised Concept Drift Detection -- Complexities, Benchmarking, Performance Analysis
di: Werner, Elias, et al.
Pubblicazione: (2023)
di: Werner, Elias, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Block Sparse Flash Attention
di: Ohayon, Daniel, et al.
Pubblicazione: (2025) -
Feature Optimization for Time Series Forecasting via Novel Randomized Uphill Climbing
di: Van Thanh, Nguyen
Pubblicazione: (2025) -
Search Your Block Floating Point Scales!
di: Gupta, Tanmaey, et al.
Pubblicazione: (2026) -
Leveraging Approximate Caching for Faster Retrieval-Augmented Generation
di: Bergman, Shai, et al.
Pubblicazione: (2025) -
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization
di: Zhou, Zhongzhu, et al.
Pubblicazione: (2026)