Enregistré dans:
| Auteurs principaux: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2603.04476 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866918370847227904 |
|---|---|
| author | Xu, Ning Zhang, Zhaoyang Shu, Senlin Qi, Lei Lv, Jiaqi Wang, Wensuo Zhao, Tianhao Zhang, Chao Yang, Zhaoliang Li, Xiangyu Su, Zhaorui Li, Jingshan Geng, Xin |
| author_facet | Xu, Ning Zhang, Zhaoyang Shu, Senlin Qi, Lei Lv, Jiaqi Wang, Wensuo Zhao, Tianhao Zhang, Chao Yang, Zhaoliang Li, Xiangyu Su, Zhaorui Li, Jingshan Geng, Xin |
| contents | Modern EDA flows rely heavily on Tcl scripting, yet general LLMs perform poorly in this domain due to extreme data scarcity, domain-specific semantics, and the high reliability required in physical design. We present iScript, a domain-adapted Qwen3-8B model for Innovus Tcl script generation, and iScript-Bench, a comprehensive benchmark covering five task categories and three difficulty levels. To overcome the lack of training data, we introduce a multi-stage data synthesis pipeline that integrates command extraction, static linting, requirement back-inference, and Chain-of-Thought generation, producing a 10K-tuple (requirement, CoT, script) dataset. iScript is trained through a two-stage strategy combining domain-adaptive pretraining and supervised fine-tuning. To evaluate script correctness efficiently, we further propose a two-step verification framework consisting of static syntax verification and LLM-based functional evaluation. On our benchmark, iScript shows higher pass@k scores than currently state-of-the-art LLMs on average. These results demonstrate the effectiveness of domain adaptation and data synthesis for EDA scripting tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_04476 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | iScript: A Domain-Adapted Large Language Model and Benchmark for Physical Design Tcl Script Generation Xu, Ning Zhang, Zhaoyang Shu, Senlin Qi, Lei Lv, Jiaqi Wang, Wensuo Zhao, Tianhao Zhang, Chao Yang, Zhaoliang Li, Xiangyu Su, Zhaorui Li, Jingshan Geng, Xin Software Engineering Programming Languages Modern EDA flows rely heavily on Tcl scripting, yet general LLMs perform poorly in this domain due to extreme data scarcity, domain-specific semantics, and the high reliability required in physical design. We present iScript, a domain-adapted Qwen3-8B model for Innovus Tcl script generation, and iScript-Bench, a comprehensive benchmark covering five task categories and three difficulty levels. To overcome the lack of training data, we introduce a multi-stage data synthesis pipeline that integrates command extraction, static linting, requirement back-inference, and Chain-of-Thought generation, producing a 10K-tuple (requirement, CoT, script) dataset. iScript is trained through a two-stage strategy combining domain-adaptive pretraining and supervised fine-tuning. To evaluate script correctness efficiently, we further propose a two-step verification framework consisting of static syntax verification and LLM-based functional evaluation. On our benchmark, iScript shows higher pass@k scores than currently state-of-the-art LLMs on average. These results demonstrate the effectiveness of domain adaptation and data synthesis for EDA scripting tasks. |
| title | iScript: A Domain-Adapted Large Language Model and Benchmark for Physical Design Tcl Script Generation |
| topic | Software Engineering Programming Languages |
| url | https://arxiv.org/abs/2603.04476 |