Going over Fine Web with a Fine-Tooth Comb: Technical Report of Indexing Fine Web for Problematic Content Search and Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Marinas, Inés Altemir, Kucherenko, Anastasiia, Kucharavy, Andrei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Getting Your Indices in a Row: Full-Text Search for LLM Training Data for Real World
by: Marinas, Ines Altemir, et al.
Published: (2025)
by: Marinas, Ines Altemir, et al.
Published: (2025)
Is It Novel and Why? Fine-Grained Patent Novelty Prediction Based on Passage Retrieval
by: Knappich, Valentin, et al.
Published: (2026)
by: Knappich, Valentin, et al.
Published: (2026)
Fine-tuning Small Language Models as Efficient Enterprise Search Relevance Labelers
by: Kang, Yue, et al.
Published: (2026)
by: Kang, Yue, et al.
Published: (2026)
Fine-Grained Table Retrieval Through the Lens of Complex Queries
by: Kosiuk, Wojciech, et al.
Published: (2026)
by: Kosiuk, Wojciech, et al.
Published: (2026)
LEMUR: A Corpus for Robust Fine-Tuning of Multilingual Law Embedding Models for Retrieval
by: Ahmadi, Narges Baba, et al.
Published: (2026)
by: Ahmadi, Narges Baba, et al.
Published: (2026)
AWED-FiNER: Agents, Web applications, and Expert Detectors for Fine-grained Named Entity Recognition across 36 Languages for 6.6 Billion Speakers
by: Kaushik, Prachuryya, et al.
Published: (2026)
by: Kaushik, Prachuryya, et al.
Published: (2026)
Knowledge Compression via Question Generation: Enhancing Multihop Document Retrieval without Fine-tuning
by: Eponon, Anvi Alex, et al.
Published: (2025)
by: Eponon, Anvi Alex, et al.
Published: (2025)
RetroLLM: Empowering Large Language Models to Retrieve Fine-grained Evidence within Generation
by: Li, Xiaoxi, et al.
Published: (2024)
by: Li, Xiaoxi, et al.
Published: (2024)
NUDGE: Lightweight Non-Parametric Fine-Tuning of Embeddings for Retrieval
by: Zeighami, Sepanta, et al.
Published: (2024)
by: Zeighami, Sepanta, et al.
Published: (2024)
HierSearch: A Hierarchical Enterprise Deep Search Framework Integrating Local and Web Searches
by: Tan, Jiejun, et al.
Published: (2025)
by: Tan, Jiejun, et al.
Published: (2025)
GPT vs RETRO: Exploring the Intersection of Retrieval and Parameter-Efficient Fine-Tuning
by: Ficek, Aleksander, et al.
Published: (2024)
by: Ficek, Aleksander, et al.
Published: (2024)
Level-Navi Agent: A Framework and benchmark for Chinese Web Search Agents
by: Hu, Chuanrui, et al.
Published: (2024)
by: Hu, Chuanrui, et al.
Published: (2024)
EnterpriseEM: Fine-tuned Embeddings for Enterprise Semantic Search
by: Rathinasamy, Kamalkumar, et al.
Published: (2024)
by: Rathinasamy, Kamalkumar, et al.
Published: (2024)
UNH at CheckThat! 2025: Fine-tuning Vs Prompting in Claim Extraction
by: Wilder, Joe, et al.
Published: (2025)
by: Wilder, Joe, et al.
Published: (2025)
WebFAQ: A Multilingual Collection of Natural Q&A Datasets for Dense Retrieval
by: Dinzinger, Michael, et al.
Published: (2025)
by: Dinzinger, Michael, et al.
Published: (2025)
FineRec:Exploring Fine-grained Sequential Recommendation
by: Zhang, Xiaokun, et al.
Published: (2024)
by: Zhang, Xiaokun, et al.
Published: (2024)
Fine-tune the Entire RAG Architecture (including DPR retriever) for Question-Answering
by: Siriwardhana, Shamane, et al.
Published: (2021)
by: Siriwardhana, Shamane, et al.
Published: (2021)
Fine-Grained Guidance for Retrievers: Leveraging LLMs' Feedback in Retrieval-Augmented Generation
by: Liu, Yuhang, et al.
Published: (2024)
by: Liu, Yuhang, et al.
Published: (2024)
WebFAQ 2.0: A Multilingual QA Dataset with Mined Hard Negatives for Dense Retrieval
by: Dinzinger, Michael, et al.
Published: (2026)
by: Dinzinger, Michael, et al.
Published: (2026)
FG-RAG: Enhancing Query-Focused Summarization with Context-Aware Fine-Grained Graph RAG
by: Hong, Yubin, et al.
Published: (2025)
by: Hong, Yubin, et al.
Published: (2025)
Leveraging the Power of LLMs: A Fine-Tuning Approach for High-Quality Aspect-Based Summarization
by: Mullick, Ankan, et al.
Published: (2024)
by: Mullick, Ankan, et al.
Published: (2024)
WebNavigator: Global Web Navigation via Interaction Graph Retrieval
by: Zhang, Xuanwang, et al.
Published: (2026)
by: Zhang, Xuanwang, et al.
Published: (2026)
ConQRet: Benchmarking Fine-Grained Evaluation of Retrieval Augmented Argumentation with LLM Judges
by: Dhole, Kaustubh D., et al.
Published: (2024)
by: Dhole, Kaustubh D., et al.
Published: (2024)
FineTuneBench: How well do commercial fine-tuning APIs infuse knowledge into LLMs?
by: Wu, Eric, et al.
Published: (2024)
by: Wu, Eric, et al.
Published: (2024)
Cross-Domain Web Information Extraction at Pinterest
by: Farag, Michael, et al.
Published: (2025)
by: Farag, Michael, et al.
Published: (2025)
Leveraging Fine-Tuned Large Language Models for Interpretable Pancreatic Cystic Lesion Feature Extraction and Risk Categorization
by: Rasromani, Ebrahim, et al.
Published: (2025)
by: Rasromani, Ebrahim, et al.
Published: (2025)
FineRef: Fine-Grained Error Reflection and Correction for Long-Form Generation with Citations
by: Peng, Yixing, et al.
Published: (2025)
by: Peng, Yixing, et al.
Published: (2025)
Large Language Models Empowered Personalized Web Agents
by: Cai, Hongru, et al.
Published: (2024)
by: Cai, Hongru, et al.
Published: (2024)
Much of Geospatial Web Search Is Beyond Traditional GIS
by: Ilyankou, Ilya, et al.
Published: (2026)
by: Ilyankou, Ilya, et al.
Published: (2026)
Quantifying and Mitigating Selection Bias in LLMs: A Transferable LoRA Fine-Tuning and Efficient Majority Voting Approach
by: Guda, Blessed, et al.
Published: (2025)
by: Guda, Blessed, et al.
Published: (2025)
LFRAG: Layout-oriented Fine-grained Retrieval-Augmented Generation on Multimodal Document Understanding
by: Zhu, Yifan, et al.
Published: (2026)
by: Zhu, Yifan, et al.
Published: (2026)
IncompeBench: A Permissively Licensed, Fine-Grained Benchmark for Music Information Retrieval
by: Clavié, Benjamin, et al.
Published: (2026)
by: Clavié, Benjamin, et al.
Published: (2026)
Retrieval Collapses When AI Pollutes the Web
by: Yu, Hongyeon, et al.
Published: (2026)
by: Yu, Hongyeon, et al.
Published: (2026)
Characterizing Web Search in The Age of Generative AI
by: Kirsten, Elisabeth, et al.
Published: (2025)
by: Kirsten, Elisabeth, et al.
Published: (2025)
Can LLM Substitute Human Labeling? A Case Study of Fine-grained Chinese Address Entity Recognition Dataset for UAV Delivery
by: Yao, Yuxuan, et al.
Published: (2024)
by: Yao, Yuxuan, et al.
Published: (2024)
WebThinker: Empowering Large Reasoning Models with Deep Research Capability
by: Li, Xiaoxi, et al.
Published: (2025)
by: Li, Xiaoxi, et al.
Published: (2025)
The Synergy of Automated Pipelines with Prompt Engineering and Generative AI in Web Crawling
by: Huang, Chau-Jian
Published: (2024)
by: Huang, Chau-Jian
Published: (2024)
FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents
by: Thakur, Nandan, et al.
Published: (2025)
by: Thakur, Nandan, et al.
Published: (2025)
Neural Retrievers are Biased Towards LLM-Generated Content
by: Dai, Sunhao, et al.
Published: (2023)
by: Dai, Sunhao, et al.
Published: (2023)
CAPRAG: A Large Language Model Solution for Customer Service and Automatic Reporting using Vector and Graph Retrieval-Augmented Generation
by: Landolsi, Hamza, et al.
Published: (2025)
by: Landolsi, Hamza, et al.
Published: (2025)
Similar Items
-
Getting Your Indices in a Row: Full-Text Search for LLM Training Data for Real World
by: Marinas, Ines Altemir, et al.
Published: (2025) -
Is It Novel and Why? Fine-Grained Patent Novelty Prediction Based on Passage Retrieval
by: Knappich, Valentin, et al.
Published: (2026) -
Fine-tuning Small Language Models as Efficient Enterprise Search Relevance Labelers
by: Kang, Yue, et al.
Published: (2026) -
Fine-Grained Table Retrieval Through the Lens of Complex Queries
by: Kosiuk, Wojciech, et al.
Published: (2026) -
LEMUR: A Corpus for Robust Fine-Tuning of Multilingual Law Embedding Models for Retrieval
by: Ahmadi, Narges Baba, et al.
Published: (2026)