DeepCQ: General-Purpose Deep-Surrogate Framework for Lossy Compression Quality Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912788996161536 |
|---|---|
| author | Mumenin, Khondoker Mirazul Underwood, Robert Dai, Dong Wang, Jinzhen Di, Sheng Lukić, Zarija Cappello, Franck |
| author_facet | Mumenin, Khondoker Mirazul Underwood, Robert Dai, Dong Wang, Jinzhen Di, Sheng Lukić, Zarija Cappello, Franck |
| contents | Error-bounded lossy compression techniques have become vital for scientific data management and analytics, given the ever-increasing volume of data generated by modern scientific simulations and instruments. Nevertheless, assessing data quality post-compression remains computationally expensive due to the intensive nature of metric calculations. In this work, we present a general-purpose deep-surrogate framework for lossy compression quality prediction (DeepCQ), with the following key contributions: 1) We develop a surrogate model for compression quality prediction that is generalizable to different error-bounded lossy compressors, quality metrics, and input datasets; 2) We adopt a novel two-stage design that decouples the computationally expensive feature-extraction stage from the light-weight metrics prediction, enabling efficient training and modular inference; 3) We optimize the model performance on time-evolving data using a mixture-of-experts design. Such a design enhances the robustness when predicting across simulation timesteps, especially when the training and test data exhibit significant variation. We validate the effectiveness of DeepCQ on four real-world scientific applications. Our results highlight the framework's exceptional predictive accuracy, with prediction errors generally under 10\% across most settings, significantly outperforming existing methods. Our framework empowers scientific users to make informed decisions about data compression based on their preferred data quality, thereby significantly reducing I/O and computational overhead in scientific data analysis. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_21433 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | DeepCQ: General-Purpose Deep-Surrogate Framework for Lossy Compression Quality Prediction Mumenin, Khondoker Mirazul Underwood, Robert Dai, Dong Wang, Jinzhen Di, Sheng Lukić, Zarija Cappello, Franck Machine Learning Distributed, Parallel, and Cluster Computing Performance Error-bounded lossy compression techniques have become vital for scientific data management and analytics, given the ever-increasing volume of data generated by modern scientific simulations and instruments. Nevertheless, assessing data quality post-compression remains computationally expensive due to the intensive nature of metric calculations. In this work, we present a general-purpose deep-surrogate framework for lossy compression quality prediction (DeepCQ), with the following key contributions: 1) We develop a surrogate model for compression quality prediction that is generalizable to different error-bounded lossy compressors, quality metrics, and input datasets; 2) We adopt a novel two-stage design that decouples the computationally expensive feature-extraction stage from the light-weight metrics prediction, enabling efficient training and modular inference; 3) We optimize the model performance on time-evolving data using a mixture-of-experts design. Such a design enhances the robustness when predicting across simulation timesteps, especially when the training and test data exhibit significant variation. We validate the effectiveness of DeepCQ on four real-world scientific applications. Our results highlight the framework's exceptional predictive accuracy, with prediction errors generally under 10\% across most settings, significantly outperforming existing methods. Our framework empowers scientific users to make informed decisions about data compression based on their preferred data quality, thereby significantly reducing I/O and computational overhead in scientific data analysis. |
| title | DeepCQ: General-Purpose Deep-Surrogate Framework for Lossy Compression Quality Prediction |
| topic | Machine Learning Distributed, Parallel, and Cluster Computing Performance |
| url | https://arxiv.org/abs/2512.21433 |