OneComp: One-Line Revolution for Generative AI Model Compression
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918418830065664 |
|---|---|
| author | Ichikawa, Yuma Kimura, Keiji Yoshida, Akihiro Fujimoto, Yudai Tokura, Hiroki Arai, Yamato Ishii, Yoshiyuki Kawakami, Yusei Shikada, Genki Jacquemond, Achille Fujisawa, Yoshihiko Fujisawa, Katsuki Honda, Takumi Sakai, Akira |
| author_facet | Ichikawa, Yuma Kimura, Keiji Yoshida, Akihiro Fujimoto, Yudai Tokura, Hiroki Arai, Yamato Ishii, Yoshiyuki Kawakami, Yusei Shikada, Genki Jacquemond, Achille Fujisawa, Yoshihiko Fujisawa, Katsuki Honda, Takumi Sakai, Akira |
| contents | Deploying foundation models is increasingly constrained by memory footprint, latency, and hardware costs. Post-training compression can mitigate these bottlenecks by reducing the precision of model parameters without significantly degrading performance; however, its practical implementation remains challenging as practitioners navigate a fragmented landscape of quantization algorithms, precision budgets, data-driven calibration strategies, and hardware-dependent execution regimes. We present OneComp, an open-source compression framework that transforms this expert workflow into a reproducible, resource-adaptive pipeline. Given a model identifier and available hardware, OneComp automatically inspects the model, plans mixed-precision assignments, and executes progressive quantization stages, ranging from layer-wise compression to block-wise refinement and global refinement. A key architectural choice is treating the first quantized checkpoint as a deployable pivot, ensuring that each subsequent stage improves the same model and that quality increases as more compute is invested. By converting state-of-the-art compression research into an extensible, open-source, hardware-aware pipeline, OneComp bridges the gap between algorithmic innovation and production-grade model deployment. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_28845 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | OneComp: One-Line Revolution for Generative AI Model Compression Ichikawa, Yuma Kimura, Keiji Yoshida, Akihiro Fujimoto, Yudai Tokura, Hiroki Arai, Yamato Ishii, Yoshiyuki Kawakami, Yusei Shikada, Genki Jacquemond, Achille Fujisawa, Yoshihiko Fujisawa, Katsuki Honda, Takumi Sakai, Akira Machine Learning Artificial Intelligence Computational Engineering, Finance, and Science Computation and Language Deploying foundation models is increasingly constrained by memory footprint, latency, and hardware costs. Post-training compression can mitigate these bottlenecks by reducing the precision of model parameters without significantly degrading performance; however, its practical implementation remains challenging as practitioners navigate a fragmented landscape of quantization algorithms, precision budgets, data-driven calibration strategies, and hardware-dependent execution regimes. We present OneComp, an open-source compression framework that transforms this expert workflow into a reproducible, resource-adaptive pipeline. Given a model identifier and available hardware, OneComp automatically inspects the model, plans mixed-precision assignments, and executes progressive quantization stages, ranging from layer-wise compression to block-wise refinement and global refinement. A key architectural choice is treating the first quantized checkpoint as a deployable pivot, ensuring that each subsequent stage improves the same model and that quality increases as more compute is invested. By converting state-of-the-art compression research into an extensible, open-source, hardware-aware pipeline, OneComp bridges the gap between algorithmic innovation and production-grade model deployment. |
| title | OneComp: One-Line Revolution for Generative AI Model Compression |
| topic | Machine Learning Artificial Intelligence Computational Engineering, Finance, and Science Computation and Language |
| url | https://arxiv.org/abs/2603.28845 |