Dripper: Token-Efficient Main HTML Extraction with a Lightweight LM
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Mengjie, Peng, Jiahui, Ning, Wenchang, Chu, Pei, Qiu, Jiantao, Ma, Ren, Zhu, He, Min, Rui, Lu, Lindong, Hou, Linfeng, Liu, Kaiwen, Qu, Yuan, Li, Zhenxiang, Xu, Chao, Tu, Zhongying, Zhang, Wentao, He, Conghui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AICC: Parse HTML Finer, Make Models Better -- A 7.3T AI-Ready Corpus Built by a Model-Based HTML Parser
von: Ma, Ren, et al.
Veröffentlicht: (2025)
von: Ma, Ren, et al.
Veröffentlicht: (2025)
Heterogeneous Adaptive Policy Optimization: Tailoring Optimization to Every Token's Nature
von: Liu, Zheng, et al.
Veröffentlicht: (2025)
von: Liu, Zheng, et al.
Veröffentlicht: (2025)
WanJuan-CC: A Safe and High-Quality Open-sourced English Webtext Dataset
von: Qiu, Jiantao, et al.
Veröffentlicht: (2024)
von: Qiu, Jiantao, et al.
Veröffentlicht: (2024)
Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification
von: Bai, Tianyi, et al.
Veröffentlicht: (2025)
von: Bai, Tianyi, et al.
Veröffentlicht: (2025)
Token Pruning in Multimodal Large Language Models: Are We Solving the Right Problem?
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding
von: Liu, Zheng, et al.
Veröffentlicht: (2025)
von: Liu, Zheng, et al.
Veröffentlicht: (2025)
Topic Over Source: The Key to Effective Data Mixing for Language Models Pre-training
von: Peng, Jiahui, et al.
Veröffentlicht: (2025)
von: Peng, Jiahui, et al.
Veröffentlicht: (2025)
Performance of ‘Sabiá’ Grass Irrigated with Drippers Installed in the Subsurface
von: Mayara Oliveira Rocha, et al.
Veröffentlicht: (2025)
von: Mayara Oliveira Rocha, et al.
Veröffentlicht: (2025)
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
von: Yu, Jiashuo, et al.
Veröffentlicht: (2025)
von: Yu, Jiashuo, et al.
Veröffentlicht: (2025)
ReaderLM-v2: Small Language Model for HTML to Markdown and JSON
von: Wang, Feng, et al.
Veröffentlicht: (2025)
von: Wang, Feng, et al.
Veröffentlicht: (2025)
Rethinking Token-wise Feature Caching: Accelerating Diffusion Transformers with Dual Feature Caching
von: Zou, Chang, et al.
Veröffentlicht: (2024)
von: Zou, Chang, et al.
Veröffentlicht: (2024)
DSDL: Data Set Description Language for Bridging Modalities and Tasks in AI Data
von: Wang, Bin, et al.
Veröffentlicht: (2024)
von: Wang, Bin, et al.
Veröffentlicht: (2024)
HTML-LSTM: Information Extraction from HTML Tables in Web Pages using Tree-Structured LSTM
von: Kawamura, Kazuki, et al.
Veröffentlicht: (2024)
von: Kawamura, Kazuki, et al.
Veröffentlicht: (2024)
OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations
von: Ouyang, Linke, et al.
Veröffentlicht: (2024)
von: Ouyang, Linke, et al.
Veröffentlicht: (2024)
Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
VADE: Variance-Aware Dynamic Sampling via Online Sample-Level Difficulty Estimation for Multimodal RL
von: Hu, Zengjie, et al.
Veröffentlicht: (2025)
von: Hu, Zengjie, et al.
Veröffentlicht: (2025)
PonderLM-3: Adaptive Token-Wise Pondering with Differentiable Masking
von: Li, He, et al.
Veröffentlicht: (2026)
von: Li, He, et al.
Veröffentlicht: (2026)
Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning
von: Bai, Tianyi, et al.
Veröffentlicht: (2025)
von: Bai, Tianyi, et al.
Veröffentlicht: (2025)
Respecting Self-Uncertainty in On-Policy Self-Distillation for Efficient LLM Reasoning
von: Ke, Junlong, et al.
Veröffentlicht: (2026)
von: Ke, Junlong, et al.
Veröffentlicht: (2026)
CEP162: A critical regulator of ciliary transition zone assembly and its implications in ciliopathies
von: Jun Yin, et al.
Veröffentlicht: (2025)
von: Jun Yin, et al.
Veröffentlicht: (2025)
AdaPonderLM: Gated Pondering Language Models with Token-Wise Adaptive Depth
von: Song, Shixiang, et al.
Veröffentlicht: (2026)
von: Song, Shixiang, et al.
Veröffentlicht: (2026)
On‐Time Meal Delivery Assisted by Drone Resupply
von: Wenqian Liu, et al.
Veröffentlicht: (2026)
von: Wenqian Liu, et al.
Veröffentlicht: (2026)
Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction
von: Zhang, Qintong, et al.
Veröffentlicht: (2024)
von: Zhang, Qintong, et al.
Veröffentlicht: (2024)
Convolution Identities of Stirling Numbers
von: Li, Nadia Na, et al.
Veröffentlicht: (2024)
von: Li, Nadia Na, et al.
Veröffentlicht: (2024)
Summation Formulae for Binomial Moments
von: Chen, Marta Na, et al.
Veröffentlicht: (2026)
von: Chen, Marta Na, et al.
Veröffentlicht: (2026)
WanJuanSiLu: A High-Quality Open-Source Webtext Dataset for Low-Resource Languages
von: Yu, Jia, et al.
Veröffentlicht: (2025)
von: Yu, Jia, et al.
Veröffentlicht: (2025)
Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration
von: Bai, Tianyi, et al.
Veröffentlicht: (2024)
von: Bai, Tianyi, et al.
Veröffentlicht: (2024)
Dataset Distillation with Neural Characteristic Function: A Minmax Perspective
von: Wang, Shaobo, et al.
Veröffentlicht: (2025)
von: Wang, Shaobo, et al.
Veröffentlicht: (2025)
LM-SPT: LM-Aligned Semantic Distillation for Speech Tokenization
von: Jo, Daejin, et al.
Veröffentlicht: (2025)
von: Jo, Daejin, et al.
Veröffentlicht: (2025)
The Palgrave Handbook of Popular Culture as PhilosophyBy David KyleJohnson, Dean A.Kowalski, ChrisLay and Kimberly S.Engels (Eds), Switzerland: Palgrave Macmillan. 2024. 2145 pp. $599.99 (hbk)
von: Mengjie Pei
Veröffentlicht: (2025)
von: Mengjie Pei
Veröffentlicht: (2025)
SynthVLM: Towards High-Quality and Efficient Synthesis of Image-Caption Datasets for Vision-Language Models
von: Liu, Zheng, et al.
Veröffentlicht: (2024)
von: Liu, Zheng, et al.
Veröffentlicht: (2024)
Lightweight yet Efficient: An External Attentive Graph Convolutional Network with Positional Prompts for Sequential Recommendation
von: Zhang, Jinyu, et al.
Veröffentlicht: (2025)
von: Zhang, Jinyu, et al.
Veröffentlicht: (2025)
ControlLM: Crafting Diverse Personalities for Language Models
von: Weng, Yixuan, et al.
Veröffentlicht: (2024)
von: Weng, Yixuan, et al.
Veröffentlicht: (2024)
Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models
von: Zhuang, Xinlin, et al.
Veröffentlicht: (2025)
von: Zhuang, Xinlin, et al.
Veröffentlicht: (2025)
Beyond a Single Extractor: Re-thinking HTML-to-Text Extraction for LLM Pretraining
von: Li, Jeffrey, et al.
Veröffentlicht: (2026)
von: Li, Jeffrey, et al.
Veröffentlicht: (2026)
VideoCompressa: Data-Efficient Video Understanding via Joint Temporal Compression and Spatial Reconstruction
von: Wang, Shaobo, et al.
Veröffentlicht: (2025)
von: Wang, Shaobo, et al.
Veröffentlicht: (2025)
PMA-Diffusion: A Physics-guided Mask-Aware Diffusion Framework for TSE from Sparse Observations
von: Liu, Lindong, et al.
Veröffentlicht: (2025)
von: Liu, Lindong, et al.
Veröffentlicht: (2025)
SpecTokenizer: A Lightweight Streaming Codec in the Compressed Spectrum Domain
von: Wan, Zixiang, et al.
Veröffentlicht: (2025)
von: Wan, Zixiang, et al.
Veröffentlicht: (2025)
Toward a Theory of Tokenization in LLMs
von: Rajaraman, Nived, et al.
Veröffentlicht: (2024)
von: Rajaraman, Nived, et al.
Veröffentlicht: (2024)
MinerU: An Open-Source Solution for Precise Document Content Extraction
von: Wang, Bin, et al.
Veröffentlicht: (2024)
von: Wang, Bin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
AICC: Parse HTML Finer, Make Models Better -- A 7.3T AI-Ready Corpus Built by a Model-Based HTML Parser
von: Ma, Ren, et al.
Veröffentlicht: (2025) -
Heterogeneous Adaptive Policy Optimization: Tailoring Optimization to Every Token's Nature
von: Liu, Zheng, et al.
Veröffentlicht: (2025) -
WanJuan-CC: A Safe and High-Quality Open-sourced English Webtext Dataset
von: Qiu, Jiantao, et al.
Veröffentlicht: (2024) -
Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification
von: Bai, Tianyi, et al.
Veröffentlicht: (2025) -
Token Pruning in Multimodal Large Language Models: Are We Solving the Right Problem?
von: Wen, Zichen, et al.
Veröffentlicht: (2025)