Iceberg: Enhancing HLS Modeling with Synthetic Data

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ding, Zijian, Nguyen, Tung, Li, Weikai, Grover, Aditya, Sun, Yizhou, Cong, Jason
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915400385560576
author Ding, Zijian
Nguyen, Tung
Li, Weikai
Grover, Aditya
Sun, Yizhou
Cong, Jason
author_facet Ding, Zijian
Nguyen, Tung
Li, Weikai
Grover, Aditya
Sun, Yizhou
Cong, Jason
contents Deep learning-based prediction models for High-Level Synthesis (HLS) of hardware designs often struggle to generalize. In this paper, we study how to close the generalizability gap of these models through pretraining on synthetic data and introduce Iceberg, a synthetic data augmentation approach that expands both large language model (LLM)-generated programs and weak labels of unseen design configurations. Our weak label generation method is integrated with an in-context model architecture, enabling meta-learning from actual and proximate labels. Iceberg improves the geometric mean modeling accuracy by $86.4\%$ when adapt to six real-world applications with few-shot examples and achieves a $2.47\times$ and a $1.12\times$ better offline DSE performance when adapting to two different test datasets. Our open-sourced code is here: https://github.com/UCLA-VAST/iceberg
format Preprint
id arxiv_https___arxiv_org_abs_2507_09948
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Iceberg: Enhancing HLS Modeling with Synthetic Data
Ding, Zijian
Nguyen, Tung
Li, Weikai
Grover, Aditya
Sun, Yizhou
Cong, Jason
Machine Learning
Hardware Architecture
Deep learning-based prediction models for High-Level Synthesis (HLS) of hardware designs often struggle to generalize. In this paper, we study how to close the generalizability gap of these models through pretraining on synthetic data and introduce Iceberg, a synthetic data augmentation approach that expands both large language model (LLM)-generated programs and weak labels of unseen design configurations. Our weak label generation method is integrated with an in-context model architecture, enabling meta-learning from actual and proximate labels. Iceberg improves the geometric mean modeling accuracy by $86.4\%$ when adapt to six real-world applications with few-shot examples and achieves a $2.47\times$ and a $1.12\times$ better offline DSE performance when adapting to two different test datasets. Our open-sourced code is here: https://github.com/UCLA-VAST/iceberg
title Iceberg: Enhancing HLS Modeling with Synthetic Data
topic Machine Learning
Hardware Architecture
url https://arxiv.org/abs/2507.09948