ReaderLM-v2: Small Language Model for HTML to Markdown and JSON
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866917943157194752 |
|---|---|
| author | Wang, Feng Shi, Zesheng Wang, Bo Wang, Nan Xiao, Han |
| author_facet | Wang, Feng Shi, Zesheng Wang, Bo Wang, Nan Xiao, Han |
| contents | We present ReaderLM-v2, a compact 1.5 billion parameter language model designed for efficient web content extraction. Our model processes documents up to 512K tokens, transforming messy HTML into clean Markdown or JSON formats with high accuracy -- making it an ideal tool for grounding large language models. The model's effectiveness results from two key innovations: (1) a three-stage data synthesis pipeline that generates high quality, diverse training data by iteratively drafting, refining, and critiquing web content extraction; and (2) a unified training framework combining continuous pre-training with multi-objective optimization. Intensive evaluation demonstrates that ReaderLM-v2 outperforms GPT-4o-2024-08-06 and other larger models by 15-20\% on carefully curated benchmarks, particularly excelling at documents exceeding 100K tokens, while maintaining significantly lower computational requirements. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_01151 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | ReaderLM-v2: Small Language Model for HTML to Markdown and JSON Wang, Feng Shi, Zesheng Wang, Bo Wang, Nan Xiao, Han Computation and Language Artificial Intelligence Information Retrieval 68T50 I.2.7; I.2.10 We present ReaderLM-v2, a compact 1.5 billion parameter language model designed for efficient web content extraction. Our model processes documents up to 512K tokens, transforming messy HTML into clean Markdown or JSON formats with high accuracy -- making it an ideal tool for grounding large language models. The model's effectiveness results from two key innovations: (1) a three-stage data synthesis pipeline that generates high quality, diverse training data by iteratively drafting, refining, and critiquing web content extraction; and (2) a unified training framework combining continuous pre-training with multi-objective optimization. Intensive evaluation demonstrates that ReaderLM-v2 outperforms GPT-4o-2024-08-06 and other larger models by 15-20\% on carefully curated benchmarks, particularly excelling at documents exceeding 100K tokens, while maintaining significantly lower computational requirements. |
| title | ReaderLM-v2: Small Language Model for HTML to Markdown and JSON |
| topic | Computation and Language Artificial Intelligence Information Retrieval 68T50 I.2.7; I.2.10 |
| url | https://arxiv.org/abs/2503.01151 |