OneComp: One-Line Revolution for Generative AI Model Compression

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ichikawa, Yuma, Kimura, Keiji, Yoshida, Akihiro, Fujimoto, Yudai, Tokura, Hiroki, Arai, Yamato, Ishii, Yoshiyuki, Kawakami, Yusei, Shikada, Genki, Jacquemond, Achille, Fujisawa, Yoshihiko, Fujisawa, Katsuki, Honda, Takumi, Sakai, Akira
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918418830065664
author Ichikawa, Yuma
Kimura, Keiji
Yoshida, Akihiro
Fujimoto, Yudai
Tokura, Hiroki
Arai, Yamato
Ishii, Yoshiyuki
Kawakami, Yusei
Shikada, Genki
Jacquemond, Achille
Fujisawa, Yoshihiko
Fujisawa, Katsuki
Honda, Takumi
Sakai, Akira
author_facet Ichikawa, Yuma
Kimura, Keiji
Yoshida, Akihiro
Fujimoto, Yudai
Tokura, Hiroki
Arai, Yamato
Ishii, Yoshiyuki
Kawakami, Yusei
Shikada, Genki
Jacquemond, Achille
Fujisawa, Yoshihiko
Fujisawa, Katsuki
Honda, Takumi
Sakai, Akira
contents Deploying foundation models is increasingly constrained by memory footprint, latency, and hardware costs. Post-training compression can mitigate these bottlenecks by reducing the precision of model parameters without significantly degrading performance; however, its practical implementation remains challenging as practitioners navigate a fragmented landscape of quantization algorithms, precision budgets, data-driven calibration strategies, and hardware-dependent execution regimes. We present OneComp, an open-source compression framework that transforms this expert workflow into a reproducible, resource-adaptive pipeline. Given a model identifier and available hardware, OneComp automatically inspects the model, plans mixed-precision assignments, and executes progressive quantization stages, ranging from layer-wise compression to block-wise refinement and global refinement. A key architectural choice is treating the first quantized checkpoint as a deployable pivot, ensuring that each subsequent stage improves the same model and that quality increases as more compute is invested. By converting state-of-the-art compression research into an extensible, open-source, hardware-aware pipeline, OneComp bridges the gap between algorithmic innovation and production-grade model deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2603_28845
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle OneComp: One-Line Revolution for Generative AI Model Compression
Ichikawa, Yuma
Kimura, Keiji
Yoshida, Akihiro
Fujimoto, Yudai
Tokura, Hiroki
Arai, Yamato
Ishii, Yoshiyuki
Kawakami, Yusei
Shikada, Genki
Jacquemond, Achille
Fujisawa, Yoshihiko
Fujisawa, Katsuki
Honda, Takumi
Sakai, Akira
Machine Learning
Artificial Intelligence
Computational Engineering, Finance, and Science
Computation and Language
Deploying foundation models is increasingly constrained by memory footprint, latency, and hardware costs. Post-training compression can mitigate these bottlenecks by reducing the precision of model parameters without significantly degrading performance; however, its practical implementation remains challenging as practitioners navigate a fragmented landscape of quantization algorithms, precision budgets, data-driven calibration strategies, and hardware-dependent execution regimes. We present OneComp, an open-source compression framework that transforms this expert workflow into a reproducible, resource-adaptive pipeline. Given a model identifier and available hardware, OneComp automatically inspects the model, plans mixed-precision assignments, and executes progressive quantization stages, ranging from layer-wise compression to block-wise refinement and global refinement. A key architectural choice is treating the first quantized checkpoint as a deployable pivot, ensuring that each subsequent stage improves the same model and that quality increases as more compute is invested. By converting state-of-the-art compression research into an extensible, open-source, hardware-aware pipeline, OneComp bridges the gap between algorithmic innovation and production-grade model deployment.
title OneComp: One-Line Revolution for Generative AI Model Compression
topic Machine Learning
Artificial Intelligence
Computational Engineering, Finance, and Science
Computation and Language
url https://arxiv.org/abs/2603.28845