A Reproducible, Scalable Pipeline for Synthesizing Autoregressive Model Literature
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916884244332544 |
|---|---|
| author | Alpay, Faruk Kilictas, Bugra Alakkad, Hamdi |
| author_facet | Alpay, Faruk Kilictas, Bugra Alakkad, Hamdi |
| contents | The accelerating pace of research on autoregressive generative models has produced thousands of papers, making manual literature surveys and reproduction studies increasingly impractical. We present a fully open-source, reproducible pipeline that automatically retrieves candidate documents from public repositories, filters them for relevance, extracts metadata, hyper-parameters and reported results, clusters topics, produces retrieval-augmented summaries and generates containerised scripts for re-running selected experiments. Quantitative evaluation on 50 manually-annotated papers shows F1 scores above 0.85 for relevance classification, hyper-parameter extraction and citation identification. Experiments on corpora of up to 1000 papers demonstrate near-linear scalability with eight CPU workers. Three case studies -- AWD-LSTM on WikiText-2, Transformer-XL on WikiText-103 and an autoregressive music model on the Lakh MIDI dataset -- confirm that the extracted settings support faithful reproduction, achieving test perplexities within 1--3% of the original reports. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_04612 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | A Reproducible, Scalable Pipeline for Synthesizing Autoregressive Model Literature Alpay, Faruk Kilictas, Bugra Alakkad, Hamdi Information Retrieval Digital Libraries Machine Learning 68P20, 68T05, 68T50 H.3.3; H.3.7; I.2.6; I.2.7 The accelerating pace of research on autoregressive generative models has produced thousands of papers, making manual literature surveys and reproduction studies increasingly impractical. We present a fully open-source, reproducible pipeline that automatically retrieves candidate documents from public repositories, filters them for relevance, extracts metadata, hyper-parameters and reported results, clusters topics, produces retrieval-augmented summaries and generates containerised scripts for re-running selected experiments. Quantitative evaluation on 50 manually-annotated papers shows F1 scores above 0.85 for relevance classification, hyper-parameter extraction and citation identification. Experiments on corpora of up to 1000 papers demonstrate near-linear scalability with eight CPU workers. Three case studies -- AWD-LSTM on WikiText-2, Transformer-XL on WikiText-103 and an autoregressive music model on the Lakh MIDI dataset -- confirm that the extracted settings support faithful reproduction, achieving test perplexities within 1--3% of the original reports. |
| title | A Reproducible, Scalable Pipeline for Synthesizing Autoregressive Model Literature |
| topic | Information Retrieval Digital Libraries Machine Learning 68P20, 68T05, 68T50 H.3.3; H.3.7; I.2.6; I.2.7 |
| url | https://arxiv.org/abs/2508.04612 |