Large-Scale Data Parallelization of Product Quantization and Inverted Indexing Using Dask
Fuente:
arXiv
Guardado en:
| Autores principales: | Abraham, Ashley N., Strelzoff, Andrew, Dozier, Haley R., Henslee, Althea C., Chappell, Mark A. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Product Quantization for Surface Soil Similarity
por: Dozier, Haley, et al.
Publicado: (2025)
por: Dozier, Haley, et al.
Publicado: (2025)
Characteristic Energy Behavior Profiling of Non-Residential Buildings
por: Dozier, Haley, et al.
Publicado: (2025)
por: Dozier, Haley, et al.
Publicado: (2025)
GPU-Accelerated INT8 Quantization for KV Cache Compression in Large Language Models
por: Taneja, Maanas, et al.
Publicado: (2026)
por: Taneja, Maanas, et al.
Publicado: (2026)
LOOPerSet: A Large-Scale Dataset for Data-Driven Polyhedral Compiler Optimization
por: Merouani, Massinissa, et al.
Publicado: (2025)
por: Merouani, Massinissa, et al.
Publicado: (2025)
Accelerating Sparse Ternary GEMM for Quantized ML on Apple Silicon
por: Lipshitz, Baraq, et al.
Publicado: (2025)
por: Lipshitz, Baraq, et al.
Publicado: (2025)
Tabular and Deep Reinforcement Learning for Gittins Index
por: Dhankhar, Harshit, et al.
Publicado: (2024)
por: Dhankhar, Harshit, et al.
Publicado: (2024)
P-MOSS: Scheduling Main-Memory Indexes Over NUMA Servers Using Next Token Prediction
por: Rayhan, Yeasir, et al.
Publicado: (2024)
por: Rayhan, Yeasir, et al.
Publicado: (2024)
Parallel Implementations Assessment of a Spatial-Spectral Classifier for Hyperspectral Clinical Applications
por: Lazcano, Raquel, et al.
Publicado: (2024)
por: Lazcano, Raquel, et al.
Publicado: (2024)
DistZO2: High-Throughput and Memory-Efficient Zeroth-Order Fine-tuning LLMs with Distributed Parallel Computing
por: Wang, Liangyu, et al.
Publicado: (2025)
por: Wang, Liangyu, et al.
Publicado: (2025)
EXAQ: Exponent Aware Quantization For LLMs Acceleration
por: Shkolnik, Moran, et al.
Publicado: (2024)
por: Shkolnik, Moran, et al.
Publicado: (2024)
Generalizing Scaling Laws for Dense and Sparse Large Language Models
por: Hossain, Md Arafat, et al.
Publicado: (2025)
por: Hossain, Md Arafat, et al.
Publicado: (2025)
Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
por: Chu, Kexin, et al.
Publicado: (2025)
por: Chu, Kexin, et al.
Publicado: (2025)
ParEVO: Synthesizing Code for Irregular Data: High-Performance Parallelism through Agentic Evolution
por: Yang, Liu, et al.
Publicado: (2026)
por: Yang, Liu, et al.
Publicado: (2026)
Sig2Model: A Boosting-Driven Model for Updatable Learned Indexes
por: Heidari, Alireza, et al.
Publicado: (2025)
por: Heidari, Alireza, et al.
Publicado: (2025)
A2Q+: Improving Accumulator-Aware Weight Quantization
por: Colbert, Ian, et al.
Publicado: (2024)
por: Colbert, Ian, et al.
Publicado: (2024)
Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error Reduction
por: Chhugani, Jatin, et al.
Publicado: (2026)
por: Chhugani, Jatin, et al.
Publicado: (2026)
KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
por: Liu, Zirui, et al.
Publicado: (2024)
por: Liu, Zirui, et al.
Publicado: (2024)
Quantitative Analysis of Deeply Quantized Tiny Neural Networks Robust to Adversarial Attacks
por: Zakariyya, Idris, et al.
Publicado: (2025)
por: Zakariyya, Idris, et al.
Publicado: (2025)
A Performance Evaluation of a Quantized Large Language Model on Various Smartphones
por: Çöplü, Tolga, et al.
Publicado: (2023)
por: Çöplü, Tolga, et al.
Publicado: (2023)
QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead
por: Zandieh, Amir, et al.
Publicado: (2024)
por: Zandieh, Amir, et al.
Publicado: (2024)
An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference
por: Yao, Feiyu, et al.
Publicado: (2026)
por: Yao, Feiyu, et al.
Publicado: (2026)
An Autotuning-based Optimization Framework for Mixed-kernel SVM Classifications in Smart Pixel Datasets and Heterojunction Transistors
por: Wu, Xingfu, et al.
Publicado: (2024)
por: Wu, Xingfu, et al.
Publicado: (2024)
PoTAcc: A Pipeline for End-to-End Acceleration of Power-of-Two Quantized DNNs
por: Saha, Rappy, et al.
Publicado: (2026)
por: Saha, Rappy, et al.
Publicado: (2026)
A 4D Hybrid Algorithm to Scale Parallel Training to Thousands of GPUs
por: Singh, Siddharth, et al.
Publicado: (2023)
por: Singh, Siddharth, et al.
Publicado: (2023)
SLiM: One-shot Quantization and Sparsity with Low-rank Approximation for LLM Weight Compression
por: Mozaffari, Mohammad, et al.
Publicado: (2024)
por: Mozaffari, Mohammad, et al.
Publicado: (2024)
SONIQ: System-Optimized Noise-Injected Ultra-Low-Precision Quantization with Full-Precision Parity
por: Zhou, Cyrus, et al.
Publicado: (2023)
por: Zhou, Cyrus, et al.
Publicado: (2023)
Throughput Optimization as a Strategic Lever in Large-Scale AI Systems: Evidence from Dataloader and Memory Profiling Innovations
por: Jha, Mayank
Publicado: (2026)
por: Jha, Mayank
Publicado: (2026)
Comment on paper: Position: Rethinking Post-Hoc Search-Based Neural Approaches for Solving Large-Scale Traveling Salesman Problems
por: Min, Yimeng
Publicado: (2024)
por: Min, Yimeng
Publicado: (2024)
Selective Parallel Loading of Large-Scale Compressed Graphs with ParaGrapher
por: Esfahani, Mohsen Koohi, et al.
Publicado: (2024)
por: Esfahani, Mohsen Koohi, et al.
Publicado: (2024)
Profiling Large Language Model Inference on Apple Silicon: A Quantization Perspective
por: Benazir, Afsara, et al.
Publicado: (2025)
por: Benazir, Afsara, et al.
Publicado: (2025)
Single-Thread JPEG Decoder Benchmarks Mis-Evaluate ML Data Loaders
por: Iglovikov, Vladimir, et al.
Publicado: (2026)
por: Iglovikov, Vladimir, et al.
Publicado: (2026)
Plug-and-Play Performance Estimation for LLM Services without Relying on Labeled Data
por: Wang, Can, et al.
Publicado: (2024)
por: Wang, Can, et al.
Publicado: (2024)
An Interpretable Latency Model for Speculative Decoding in LLM Serving
por: Kong, Linghao, et al.
Publicado: (2026)
por: Kong, Linghao, et al.
Publicado: (2026)
MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models
por: Chitty-Venkata, Krishna Teja, et al.
Publicado: (2025)
por: Chitty-Venkata, Krishna Teja, et al.
Publicado: (2025)
How Much Parallelism Is "Free"? A Principle of Near-Free Parallelism for Parallel Decoding
por: He, Minghua, et al.
Publicado: (2026)
por: He, Minghua, et al.
Publicado: (2026)
SynthEval: A Framework for Detailed Utility and Privacy Evaluation of Tabular Synthetic Data
por: Lautrup, Anton Danholt, et al.
Publicado: (2024)
por: Lautrup, Anton Danholt, et al.
Publicado: (2024)
PEVLM: Parallel Encoding for Vision-Language Models
por: Kang, Letian, et al.
Publicado: (2025)
por: Kang, Letian, et al.
Publicado: (2025)
ZO2: Scalable Zeroth-Order Fine-Tuning for Extremely Large Language Models with Limited GPU Memory
por: Wang, Liangyu, et al.
Publicado: (2025)
por: Wang, Liangyu, et al.
Publicado: (2025)
The Gittins Index: A Design Principle for Decision-Making Under Uncertainty
por: Scully, Ziv, et al.
Publicado: (2025)
por: Scully, Ziv, et al.
Publicado: (2025)
Accelerating Diffusion LLMs via Adaptive Parallel Decoding
por: Israel, Daniel, et al.
Publicado: (2025)
por: Israel, Daniel, et al.
Publicado: (2025)
Ejemplares similares
-
Product Quantization for Surface Soil Similarity
por: Dozier, Haley, et al.
Publicado: (2025) -
Characteristic Energy Behavior Profiling of Non-Residential Buildings
por: Dozier, Haley, et al.
Publicado: (2025) -
GPU-Accelerated INT8 Quantization for KV Cache Compression in Large Language Models
por: Taneja, Maanas, et al.
Publicado: (2026) -
LOOPerSet: A Large-Scale Dataset for Data-Driven Polyhedral Compiler Optimization
por: Merouani, Massinissa, et al.
Publicado: (2025) -
Accelerating Sparse Ternary GEMM for Quantized ML on Apple Silicon
por: Lipshitz, Baraq, et al.
Publicado: (2025)