ReaderLM-v2: Small Language Model for HTML to Markdown and JSON

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Feng, Shi, Zesheng, Wang, Bo, Wang, Nan, Xiao, Han
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917943157194752
author Wang, Feng
Shi, Zesheng
Wang, Bo
Wang, Nan
Xiao, Han
author_facet Wang, Feng
Shi, Zesheng
Wang, Bo
Wang, Nan
Xiao, Han
contents We present ReaderLM-v2, a compact 1.5 billion parameter language model designed for efficient web content extraction. Our model processes documents up to 512K tokens, transforming messy HTML into clean Markdown or JSON formats with high accuracy -- making it an ideal tool for grounding large language models. The model's effectiveness results from two key innovations: (1) a three-stage data synthesis pipeline that generates high quality, diverse training data by iteratively drafting, refining, and critiquing web content extraction; and (2) a unified training framework combining continuous pre-training with multi-objective optimization. Intensive evaluation demonstrates that ReaderLM-v2 outperforms GPT-4o-2024-08-06 and other larger models by 15-20\% on carefully curated benchmarks, particularly excelling at documents exceeding 100K tokens, while maintaining significantly lower computational requirements.
format Preprint
id arxiv_https___arxiv_org_abs_2503_01151
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ReaderLM-v2: Small Language Model for HTML to Markdown and JSON
Wang, Feng
Shi, Zesheng
Wang, Bo
Wang, Nan
Xiao, Han
Computation and Language
Artificial Intelligence
Information Retrieval
68T50
I.2.7; I.2.10
We present ReaderLM-v2, a compact 1.5 billion parameter language model designed for efficient web content extraction. Our model processes documents up to 512K tokens, transforming messy HTML into clean Markdown or JSON formats with high accuracy -- making it an ideal tool for grounding large language models. The model's effectiveness results from two key innovations: (1) a three-stage data synthesis pipeline that generates high quality, diverse training data by iteratively drafting, refining, and critiquing web content extraction; and (2) a unified training framework combining continuous pre-training with multi-objective optimization. Intensive evaluation demonstrates that ReaderLM-v2 outperforms GPT-4o-2024-08-06 and other larger models by 15-20\% on carefully curated benchmarks, particularly excelling at documents exceeding 100K tokens, while maintaining significantly lower computational requirements.
title ReaderLM-v2: Small Language Model for HTML to Markdown and JSON
topic Computation and Language
Artificial Intelligence
Information Retrieval
68T50
I.2.7; I.2.10
url https://arxiv.org/abs/2503.01151