A Reproducible, Scalable Pipeline for Synthesizing Autoregressive Model Literature

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Alpay, Faruk, Kilictas, Bugra, Alakkad, Hamdi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916884244332544
author Alpay, Faruk
Kilictas, Bugra
Alakkad, Hamdi
author_facet Alpay, Faruk
Kilictas, Bugra
Alakkad, Hamdi
contents The accelerating pace of research on autoregressive generative models has produced thousands of papers, making manual literature surveys and reproduction studies increasingly impractical. We present a fully open-source, reproducible pipeline that automatically retrieves candidate documents from public repositories, filters them for relevance, extracts metadata, hyper-parameters and reported results, clusters topics, produces retrieval-augmented summaries and generates containerised scripts for re-running selected experiments. Quantitative evaluation on 50 manually-annotated papers shows F1 scores above 0.85 for relevance classification, hyper-parameter extraction and citation identification. Experiments on corpora of up to 1000 papers demonstrate near-linear scalability with eight CPU workers. Three case studies -- AWD-LSTM on WikiText-2, Transformer-XL on WikiText-103 and an autoregressive music model on the Lakh MIDI dataset -- confirm that the extracted settings support faithful reproduction, achieving test perplexities within 1--3% of the original reports.
format Preprint
id arxiv_https___arxiv_org_abs_2508_04612
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Reproducible, Scalable Pipeline for Synthesizing Autoregressive Model Literature
Alpay, Faruk
Kilictas, Bugra
Alakkad, Hamdi
Information Retrieval
Digital Libraries
Machine Learning
68P20, 68T05, 68T50
H.3.3; H.3.7; I.2.6; I.2.7
The accelerating pace of research on autoregressive generative models has produced thousands of papers, making manual literature surveys and reproduction studies increasingly impractical. We present a fully open-source, reproducible pipeline that automatically retrieves candidate documents from public repositories, filters them for relevance, extracts metadata, hyper-parameters and reported results, clusters topics, produces retrieval-augmented summaries and generates containerised scripts for re-running selected experiments. Quantitative evaluation on 50 manually-annotated papers shows F1 scores above 0.85 for relevance classification, hyper-parameter extraction and citation identification. Experiments on corpora of up to 1000 papers demonstrate near-linear scalability with eight CPU workers. Three case studies -- AWD-LSTM on WikiText-2, Transformer-XL on WikiText-103 and an autoregressive music model on the Lakh MIDI dataset -- confirm that the extracted settings support faithful reproduction, achieving test perplexities within 1--3% of the original reports.
title A Reproducible, Scalable Pipeline for Synthesizing Autoregressive Model Literature
topic Information Retrieval
Digital Libraries
Machine Learning
68P20, 68T05, 68T50
H.3.3; H.3.7; I.2.6; I.2.7
url https://arxiv.org/abs/2508.04612