Cleaner Pretraining Corpus Curation with Neural Web Scraping
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Zhipeng, Liu, Zhenghao, Yan, Yukun, Liu, Zhiyuan, Yu, Ge, Xiong, Chenyan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Craw4LLM: Efficient Web Crawling for LLM Pretraining
by: Yu, Shi, et al.
Published: (2025)
by: Yu, Shi, et al.
Published: (2025)
ThinkNote: Enhancing Knowledge Integration and Utilization of Large Language Models via Constructivist Cognition Modeling
by: Xu, Zhipeng, et al.
Published: (2024)
by: Xu, Zhipeng, et al.
Published: (2024)
KG-Infused RAG: Augmenting Corpus-Based RAG with External Knowledge Graphs
by: Wu, Dingjun, et al.
Published: (2025)
by: Wu, Dingjun, et al.
Published: (2025)
RePro: Training Language Models to Faithfully Recycle the Web for Pretraining
by: Yu, Zichun, et al.
Published: (2025)
by: Yu, Zichun, et al.
Published: (2025)
Say More with Less: Understanding Prompt Learning Behaviors through Gist Compression
by: Li, Xinze, et al.
Published: (2024)
by: Li, Xinze, et al.
Published: (2024)
Toolink: Linking Toolkit Creation and Using through Chain-of-Solving on Open-Source Model
by: Qian, Cheng, et al.
Published: (2023)
by: Qian, Cheng, et al.
Published: (2023)
Midtraining Bridges Pretraining and Posttraining Distributions
by: Liu, Emmy, et al.
Published: (2025)
by: Liu, Emmy, et al.
Published: (2025)
RAG-DDR: Optimizing Retrieval-Augmented Generation Using Differentiable Data Rewards
by: Li, Xinze, et al.
Published: (2024)
by: Li, Xinze, et al.
Published: (2024)
ParamMute: Suppressing Knowledge-Critical FFNs for Faithful Retrieval-Augmented Generation
by: Huang, Pengcheng, et al.
Published: (2025)
by: Huang, Pengcheng, et al.
Published: (2025)
Generating Pretraining Tokens from Organic Data for Data-Bound Scaling
by: Yu, Zichun, et al.
Published: (2026)
by: Yu, Zichun, et al.
Published: (2026)
Mitigating Judgment Preference Bias in Large Language Models through Group-Based Polling
by: Liu, Shuliang, et al.
Published: (2025)
by: Liu, Shuliang, et al.
Published: (2025)
MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models
by: Yu, Zichun, et al.
Published: (2024)
by: Yu, Zichun, et al.
Published: (2024)
Judge as A Judge: Improving the Evaluation of Retrieval-Augmented Generation through the Judge-Consistency of Large Language Models
by: Liu, Shuliang, et al.
Published: (2025)
by: Liu, Shuliang, et al.
Published: (2025)
Mixture-of-Retrieval Experts for Reasoning-Guided Multimodal Knowledge Exploitation
by: Peng, Chunyi, et al.
Published: (2025)
by: Peng, Chunyi, et al.
Published: (2025)
Enhancing Long-Chain Reasoning Distillation through Error-Aware Self-Reflection
by: Wu, Zhuoyang, et al.
Published: (2025)
by: Wu, Zhuoyang, et al.
Published: (2025)
Legal$Δ$: Enhancing Legal Reasoning in LLMs via Reinforcement Learning with Chain-of-Thought Guided Information Gain
by: Dai, Xin, et al.
Published: (2025)
by: Dai, Xin, et al.
Published: (2025)
RankCoT: Refining Knowledge for Retrieval-Augmented Generation through Ranking Chain-of-Thoughts
by: Wu, Mingyan, et al.
Published: (2025)
by: Wu, Mingyan, et al.
Published: (2025)
Leveraging Large Language Models for Web Scraping
by: Ahluwalia, Aman, et al.
Published: (2024)
by: Ahluwalia, Aman, et al.
Published: (2024)
Group-Level Data Selection for Efficient Pretraining
by: Yu, Zichun, et al.
Published: (2025)
by: Yu, Zichun, et al.
Published: (2025)
VisRAG 2.0: Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation
by: Sun, Yubo, et al.
Published: (2025)
by: Sun, Yubo, et al.
Published: (2025)
Scientific Knowledge-driven Decoding Constraints Improving the Reliability of LLMs
by: Ma, Maotian, et al.
Published: (2026)
by: Ma, Maotian, et al.
Published: (2026)
Model-Generated Pretraining Signals Improves Zero-Shot Generalization of Text-to-Text Transformers
by: Gong, Linyuan, et al.
Published: (2023)
by: Gong, Linyuan, et al.
Published: (2023)
Building A Coding Assistant via the Retrieval-Augmented Language Model
by: Li, Xinze, et al.
Published: (2024)
by: Li, Xinze, et al.
Published: (2024)
Scraping the Shadows: Deep Learning Breakthroughs in Dark Web Intelligence
by: Bakermans, Ingmar, et al.
Published: (2025)
by: Bakermans, Ingmar, et al.
Published: (2025)
PersLLM: A Personified Training Approach for Large Language Models
by: Zeng, Zheni, et al.
Published: (2024)
by: Zeng, Zheni, et al.
Published: (2024)
LegalDuet: Learning Fine-grained Representations for Legal Judgment Prediction via a Dual-View Contrastive Learning
by: Xu, Buqiang, et al.
Published: (2024)
by: Xu, Buqiang, et al.
Published: (2024)
MetaMem: Evolving Meta-Memory for Knowledge Utilization through Self-Reflective Symbolic Optimization
by: Xin, Haidong, et al.
Published: (2026)
by: Xin, Haidong, et al.
Published: (2026)
Long-Chain Reasoning Distillation via Adaptive Prefix Alignment
by: Liu, Zhenghao, et al.
Published: (2026)
by: Liu, Zhenghao, et al.
Published: (2026)
KBAlign: Efficient Self Adaptation on Specific Knowledge Bases
by: Zeng, Zheni, et al.
Published: (2024)
by: Zeng, Zheni, et al.
Published: (2024)
OphIn-500K: Curating Web-Scale Visual Instructions for Scaling Ophthalmic Multimodal Large Language Models
by: Dong, Xuanzhao, et al.
Published: (2026)
by: Dong, Xuanzhao, et al.
Published: (2026)
KARE-RAG: Knowledge-Aware Refinement and Enhancement for RAG
by: Li, Yongjian, et al.
Published: (2025)
by: Li, Yongjian, et al.
Published: (2025)
Structured Knowledge Representation through Contextual Pages for Retrieval-Augmented Generation
by: Li, Xinze, et al.
Published: (2026)
by: Li, Xinze, et al.
Published: (2026)
ClueAnchor: Clue-Anchored Knowledge Reasoning Exploration and Optimization for Retrieval-Augmented Generation
by: Chen, Hao, et al.
Published: (2025)
by: Chen, Hao, et al.
Published: (2025)
DeepNote: Note-Centric Deep Retrieval-Augmented Generation
by: Wang, Ruobing, et al.
Published: (2024)
by: Wang, Ruobing, et al.
Published: (2024)
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models
by: Kang, Hao, et al.
Published: (2025)
by: Kang, Hao, et al.
Published: (2025)
HIPPO: Enhancing the Table Understanding Capability of LLMs through Hybrid-Modal Preference Optimization
by: Wang, Haolan, et al.
Published: (2025)
by: Wang, Haolan, et al.
Published: (2025)
MatPlotAgent: Method and Evaluation for LLM-Based Agentic Scientific Data Visualization
by: Yang, Zhiyu, et al.
Published: (2024)
by: Yang, Zhiyu, et al.
Published: (2024)
UltraLink: An Open-Source Knowledge-Enhanced Multilingual Supervised Fine-tuning Dataset
by: Wang, Haoyu, et al.
Published: (2024)
by: Wang, Haoyu, et al.
Published: (2024)
ReCUT: Balancing Reasoning Length and Accuracy in LLMs via Stepwise Trails and Preference Optimization
by: Jin, Zhensheng, et al.
Published: (2025)
by: Jin, Zhensheng, et al.
Published: (2025)
Mangosteen: An Open Thai Corpus for Language Model Pretraining
by: Phatthiyaphaibun, Wannaphong, et al.
Published: (2025)
by: Phatthiyaphaibun, Wannaphong, et al.
Published: (2025)
Similar Items
-
Craw4LLM: Efficient Web Crawling for LLM Pretraining
by: Yu, Shi, et al.
Published: (2025) -
ThinkNote: Enhancing Knowledge Integration and Utilization of Large Language Models via Constructivist Cognition Modeling
by: Xu, Zhipeng, et al.
Published: (2024) -
KG-Infused RAG: Augmenting Corpus-Based RAG with External Knowledge Graphs
by: Wu, Dingjun, et al.
Published: (2025) -
RePro: Training Language Models to Faithfully Recycle the Web for Pretraining
by: Yu, Zichun, et al.
Published: (2025) -
Say More with Less: Understanding Prompt Learning Behaviors through Gist Compression
by: Li, Xinze, et al.
Published: (2024)