Waltz: Temperature-Aware Cooperative Compression for High-Performance Compression-Based CSDs
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Dingcui, Song, Yunpeng, Huang, Yiyang, Zhao, Yumiao, Lv, Yina, Wang, Chundong, Zhang, Youtao, Shi, Liang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RARO: Reliability-aware Conversion with Enhanced Read Performance for QLC SSDs
by: Wang, Yanyun, et al.
Published: (2025)
by: Wang, Yanyun, et al.
Published: (2025)
Characterize LSM-tree Compaction Performance via On-Device LLM Inference
by: Ding, Jiabiao, et al.
Published: (2026)
by: Ding, Jiabiao, et al.
Published: (2026)
Accurate Performance Modeling And Uncertainty Analysis of Lossy Compression in Scientific Applications
by: Liu, Youyuan, et al.
Published: (2024)
by: Liu, Youyuan, et al.
Published: (2024)
Two Criteria for Performance Analysis of Optimization Algorithms
by: Jing, Yunpeng, et al.
Published: (2024)
by: Jing, Yunpeng, et al.
Published: (2024)
ConZone+: Practical Zoned Flash Storage Emulation for Consumer Devices
by: Yu, Dingcui, et al.
Published: (2025)
by: Yu, Dingcui, et al.
Published: (2025)
Hold Onto That Thought: Assessing KV Cache Compression On Reasoning
by: Liu, Minghui, et al.
Published: (2025)
by: Liu, Minghui, et al.
Published: (2025)
EDGC: Entropy-driven Dynamic Gradient Compression for Efficient LLM Training
by: Yi, Qingao, et al.
Published: (2025)
by: Yi, Qingao, et al.
Published: (2025)
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
by: Liu, Guangda, et al.
Published: (2024)
by: Liu, Guangda, et al.
Published: (2024)
RWKV-edge: Deeply Compressed RWKV for Resource-Constrained Devices
by: Choe, Wonkyo, et al.
Published: (2024)
by: Choe, Wonkyo, et al.
Published: (2024)
Toward Greener Matrix Operations by Lossless Compressed Formats
by: Tosoni, Francesco, et al.
Published: (2024)
by: Tosoni, Francesco, et al.
Published: (2024)
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
by: Fan, Ruibo, et al.
Published: (2026)
by: Fan, Ruibo, et al.
Published: (2026)
LoPace: A Lossless Optimized Prompt Accurate Compression Engine for Large Language Model Applications
by: Ulla, Aman
Published: (2026)
by: Ulla, Aman
Published: (2026)
FRSZ2 for In-Register Block Compression Inside GMRES on GPUs
by: Grützmacher, Thomas, et al.
Published: (2024)
by: Grützmacher, Thomas, et al.
Published: (2024)
GPU-Accelerated INT8 Quantization for KV Cache Compression in Large Language Models
by: Taneja, Maanas, et al.
Published: (2026)
by: Taneja, Maanas, et al.
Published: (2026)
High Performance Matrix Multiplication
by: Davis, Ethan
Published: (2025)
by: Davis, Ethan
Published: (2025)
Optimization of Armv9 architecture general large language model inference performance based on Llama.cpp
by: Chen, Longhao, et al.
Published: (2024)
by: Chen, Longhao, et al.
Published: (2024)
GVEL: Fast Graph Loading in Edgelist and Compressed Sparse Row (CSR) formats
by: Sahu, Subhajit
Published: (2023)
by: Sahu, Subhajit
Published: (2023)
On the Compression of Language Models for Code: An Empirical Study on CodeBERT
by: d'Aloisio, Giordano, et al.
Published: (2024)
by: d'Aloisio, Giordano, et al.
Published: (2024)
Is Quantum Optimization Ready? An Effort Towards Neural Network Compression using Adiabatic Quantum Computing
by: Wang, Zhehui, et al.
Published: (2025)
by: Wang, Zhehui, et al.
Published: (2025)
AI Application Benchmarking: Power-Aware Performance Analysis for Vision and Language Models
by: Mayr, Martin, et al.
Published: (2026)
by: Mayr, Martin, et al.
Published: (2026)
Non-Asymptotic Performance Analysis of DOA Estimation Based on Real-Valued Root-MUSIC
by: Liu, Junyang, et al.
Published: (2025)
by: Liu, Junyang, et al.
Published: (2025)
Selective Parallel Loading of Large-Scale Compressed Graphs with ParaGrapher
by: Esfahani, Mohsen Koohi, et al.
Published: (2024)
by: Esfahani, Mohsen Koohi, et al.
Published: (2024)
A Continuous Benchmarking Infrastructure for High-Performance Computing Applications
by: Alt, Christoph, et al.
Published: (2024)
by: Alt, Christoph, et al.
Published: (2024)
Model Compression and Efficient Inference for Large Language Models: A Survey
by: Wang, Wenxiao, et al.
Published: (2024)
by: Wang, Wenxiao, et al.
Published: (2024)
JSPIM: A Skew-Aware PIM Accelerator for High-Performance Databases Join and Select Operations
by: Tajdari, Sabiha, et al.
Published: (2025)
by: Tajdari, Sabiha, et al.
Published: (2025)
PerfSeer: An Efficient and Accurate Deep Learning Models Performance Predictor
by: Zhao, Xinlong, et al.
Published: (2025)
by: Zhao, Xinlong, et al.
Published: (2025)
oneDNN Graph Compiler: A Hybrid Approach for High-Performance Deep Learning Compilation
by: Li, Jianhui, et al.
Published: (2023)
by: Li, Jianhui, et al.
Published: (2023)
GreenLLM: SLO-Aware Dynamic Frequency Scaling for Energy-Efficient LLM Serving
by: Liu, Qunyou, et al.
Published: (2025)
by: Liu, Qunyou, et al.
Published: (2025)
CRAM: Large-scale Video Continual Learning with Bootstrapped Compression
by: Mall, Shivani, et al.
Published: (2025)
by: Mall, Shivani, et al.
Published: (2025)
SLiM: One-shot Quantization and Sparsity with Low-rank Approximation for LLM Weight Compression
by: Mozaffari, Mohammad, et al.
Published: (2024)
by: Mozaffari, Mohammad, et al.
Published: (2024)
SAHM: State-Aware Heterogeneous Multicore for Single-Thread Performance
by: Wadle, Shayne, et al.
Published: (2025)
by: Wadle, Shayne, et al.
Published: (2025)
ACALSim: A Scalable Parallel Simulation Framework for High-Performance System Design Space Exploration
by: Lin, Wei-Fen, et al.
Published: (2026)
by: Lin, Wei-Fen, et al.
Published: (2026)
CPMA: An Efficient Batch-Parallel Compressed Set Without Pointers
by: Wheatman, Brian, et al.
Published: (2023)
by: Wheatman, Brian, et al.
Published: (2023)
Pinching-Antenna Systems For Indoor Immersive Communications: A 3D-Modeling Based Performance Analysis
by: Wang, Yulei, et al.
Published: (2025)
by: Wang, Yulei, et al.
Published: (2025)
Adaptive Iterative Compression for High-Resolution Files: an Approach Focused on Preserving Visual Quality in Cinematic Workflows
by: Melo, Leonardo, et al.
Published: (2025)
by: Melo, Leonardo, et al.
Published: (2025)
SparAMX: Accelerating Compressed LLMs Token Generation on AMX-powered CPUs
by: AbouElhamayed, Ahmed F., et al.
Published: (2025)
by: AbouElhamayed, Ahmed F., et al.
Published: (2025)
Systematic Performance Evaluation Framework for LEO Mega-Constellation Satellite Networks
by: Wang, Yu, et al.
Published: (2024)
by: Wang, Yu, et al.
Published: (2024)
gpu tracker: Python Package for Tracking and Profiling GPU and Other Hardware Utilization in Both Desktop and High-Performance Computing Environments
by: Huckvale, Erik D., et al.
Published: (2024)
by: Huckvale, Erik D., et al.
Published: (2024)
Size-Aware Dispatching to Fluid Queues
by: Xie, Runhan, et al.
Published: (2025)
by: Xie, Runhan, et al.
Published: (2025)
From Talent to High Performance: The view of coaches, players and club coordinators on the relevant factors in the development of a Basketball player
by: L. Gonçalves
Published: (2017)
by: L. Gonçalves
Published: (2017)
Similar Items
-
RARO: Reliability-aware Conversion with Enhanced Read Performance for QLC SSDs
by: Wang, Yanyun, et al.
Published: (2025) -
Characterize LSM-tree Compaction Performance via On-Device LLM Inference
by: Ding, Jiabiao, et al.
Published: (2026) -
Accurate Performance Modeling And Uncertainty Analysis of Lossy Compression in Scientific Applications
by: Liu, Youyuan, et al.
Published: (2024) -
Two Criteria for Performance Analysis of Optimization Algorithms
by: Jing, Yunpeng, et al.
Published: (2024) -
ConZone+: Practical Zoned Flash Storage Emulation for Consumer Devices
by: Yu, Dingcui, et al.
Published: (2025)