The Synergy of Automated Pipelines with Prompt Engineering and Generative AI in Web Crawling
Fuente:
arXiv
Saved in:
| Main Author: | Huang, Chau-Jian |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Document Quality Scoring for Web Crawling
by: Pezzuti, Francesca, et al.
Published: (2025)
by: Pezzuti, Francesca, et al.
Published: (2025)
SCP-116K: A High-Quality Problem-Solution Dataset and a Generalized Pipeline for Automated Extraction in the Higher Education Science Domain
by: Lu, Dakuan, et al.
Published: (2025)
by: Lu, Dakuan, et al.
Published: (2025)
AutoPureData: Automated Filtering of Undesirable Web Data to Update LLM Knowledge
by: Vadlapati, Praneeth
Published: (2024)
by: Vadlapati, Praneeth
Published: (2024)
Benchmarking Complex Multimodal Document Processing Pipelines: A Unified Evaluation Framework for Enterprise AI
by: Singh, Saurabh K., et al.
Published: (2026)
by: Singh, Saurabh K., et al.
Published: (2026)
Semi-Automated Knowledge Engineering and Process Mapping for Total Airport Management
by: Teo, Darryl, et al.
Published: (2026)
by: Teo, Darryl, et al.
Published: (2026)
Web Page Classification using LLMs for Crawling Support
by: Sasazawa, Yuichi, et al.
Published: (2025)
by: Sasazawa, Yuichi, et al.
Published: (2025)
On the impact of retrieved content representations in RAG Pipelines
by: Ross, Jonathan J, et al.
Published: (2026)
by: Ross, Jonathan J, et al.
Published: (2026)
Generative Language Models with Retrieval Augmented Generation for Automated Short Answer Scoring
by: Wang, Zifan, et al.
Published: (2024)
by: Wang, Zifan, et al.
Published: (2024)
Comparative Analysis of Neural Retriever-Reranker Pipelines for Retrieval-Augmented Generation over Knowledge Graphs in E-commerce Applications
by: Rumble, Teri, et al.
Published: (2025)
by: Rumble, Teri, et al.
Published: (2025)
The Text Classification Pipeline: Starting Shallow going Deeper
by: Siino, Marco, et al.
Published: (2024)
by: Siino, Marco, et al.
Published: (2024)
ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems
by: Saad-Falcon, Jon, et al.
Published: (2023)
by: Saad-Falcon, Jon, et al.
Published: (2023)
HierSearch: A Hierarchical Enterprise Deep Search Framework Integrating Local and Web Searches
by: Tan, Jiejun, et al.
Published: (2025)
by: Tan, Jiejun, et al.
Published: (2025)
GenQREnsemble: Zero-Shot LLM Ensemble Prompting for Generative Query Reformulation
by: Dhole, Kaustubh, et al.
Published: (2024)
by: Dhole, Kaustubh, et al.
Published: (2024)
Optimizing RAG Pipelines for Arabic: A Systematic Analysis of Core Components
by: Alsubhi, Jumana, et al.
Published: (2025)
by: Alsubhi, Jumana, et al.
Published: (2025)
A Semantic Search Pipeline for Causality-driven Adhoc Information Retrieval
by: Dalal, Dhairya, et al.
Published: (2025)
by: Dalal, Dhairya, et al.
Published: (2025)
RecGPT: Generative Personalized Prompts for Sequential Recommendation via ChatGPT Training Paradigm
by: Zhang, Yabin, et al.
Published: (2024)
by: Zhang, Yabin, et al.
Published: (2024)
ReGeS: Reciprocal Retrieval-Generation Synergy for Conversational Recommender Systems
by: Yang, Dayu, et al.
Published: (2025)
by: Yang, Dayu, et al.
Published: (2025)
Autofocus Retrieval: An Effective Pipeline for Multi-Hop Question Answering With Semi-Structured Knowledge
by: Boer, Derian, et al.
Published: (2025)
by: Boer, Derian, et al.
Published: (2025)
Take Care of Your Prompt Bias! Investigating and Mitigating Prompt Bias in Factual Knowledge Extraction
by: Xu, Ziyang, et al.
Published: (2024)
by: Xu, Ziyang, et al.
Published: (2024)
RuleRAG: Rule-Guided Retrieval-Augmented Generation with Language Models for Question Answering
by: Chen, Zhongwu, et al.
Published: (2024)
by: Chen, Zhongwu, et al.
Published: (2024)
Cross-Domain Web Information Extraction at Pinterest
by: Farag, Michael, et al.
Published: (2025)
by: Farag, Michael, et al.
Published: (2025)
Going over Fine Web with a Fine-Tooth Comb: Technical Report of Indexing Fine Web for Problematic Content Search and Retrieval
by: Marinas, Inés Altemir, et al.
Published: (2025)
by: Marinas, Inés Altemir, et al.
Published: (2025)
OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis
by: Li, Zhuofeng, et al.
Published: (2026)
by: Li, Zhuofeng, et al.
Published: (2026)
Large Language Models Empowered Personalized Web Agents
by: Cai, Hongru, et al.
Published: (2024)
by: Cai, Hongru, et al.
Published: (2024)
A Survey on Retrieval-Augmented Text Generation for Large Language Models
by: Huang, Yizheng, et al.
Published: (2024)
by: Huang, Yizheng, et al.
Published: (2024)
FinRetrieval: A Benchmark for Financial Data Retrieval by AI Agents
by: Kim, Eric Y., et al.
Published: (2026)
by: Kim, Eric Y., et al.
Published: (2026)
Benchmarking Prompt Sensitivity in Large Language Models
by: Razavi, Amirhossein, et al.
Published: (2025)
by: Razavi, Amirhossein, et al.
Published: (2025)
Haystack Engineering: Context Engineering for Heterogeneous and Agentic Long-Context Evaluation
by: Li, Mufei, et al.
Published: (2025)
by: Li, Mufei, et al.
Published: (2025)
WebThinker: Empowering Large Reasoning Models with Deep Research Capability
by: Li, Xiaoxi, et al.
Published: (2025)
by: Li, Xiaoxi, et al.
Published: (2025)
Chained Prompting for Better Systematic Review Search Strategies
by: Nasser, Fatima, et al.
Published: (2025)
by: Nasser, Fatima, et al.
Published: (2025)
ScrapeGraphAI-100k: Dataset for Schema-Constrained LLM Generation
by: Brach, William, et al.
Published: (2026)
by: Brach, William, et al.
Published: (2026)
Level-Navi Agent: A Framework and benchmark for Chinese Web Search Agents
by: Hu, Chuanrui, et al.
Published: (2024)
by: Hu, Chuanrui, et al.
Published: (2024)
C-Pack: Packed Resources For General Chinese Embeddings
by: Xiao, Shitao, et al.
Published: (2023)
by: Xiao, Shitao, et al.
Published: (2023)
WebFAQ: A Multilingual Collection of Natural Q&A Datasets for Dense Retrieval
by: Dinzinger, Michael, et al.
Published: (2025)
by: Dinzinger, Michael, et al.
Published: (2025)
Automated Neural Patent Landscaping in the Small Data Regime
by: Erana, Tisa Islam, et al.
Published: (2024)
by: Erana, Tisa Islam, et al.
Published: (2024)
LLM-Rec: Personalized Recommendation via Prompting Large Language Models
by: Lyu, Hanjia, et al.
Published: (2023)
by: Lyu, Hanjia, et al.
Published: (2023)
Better by Comparison: Retrieval-Augmented Contrastive Reasoning for Automatic Prompt Optimization
by: Lee, Juhyeon, et al.
Published: (2025)
by: Lee, Juhyeon, et al.
Published: (2025)
UNH at CheckThat! 2025: Fine-tuning Vs Prompting in Claim Extraction
by: Wilder, Joe, et al.
Published: (2025)
by: Wilder, Joe, et al.
Published: (2025)
SimpleDeepSearcher: Deep Information Seeking via Web-Powered Reasoning Trajectory Synthesis
by: Sun, Shuang, et al.
Published: (2025)
by: Sun, Shuang, et al.
Published: (2025)
PERSOMA: PERsonalized SOft ProMpt Adapter Architecture for Personalized Language Prompting
by: Hebert, Liam, et al.
Published: (2024)
by: Hebert, Liam, et al.
Published: (2024)
Similar Items
-
Document Quality Scoring for Web Crawling
by: Pezzuti, Francesca, et al.
Published: (2025) -
SCP-116K: A High-Quality Problem-Solution Dataset and a Generalized Pipeline for Automated Extraction in the Higher Education Science Domain
by: Lu, Dakuan, et al.
Published: (2025) -
AutoPureData: Automated Filtering of Undesirable Web Data to Update LLM Knowledge
by: Vadlapati, Praneeth
Published: (2024) -
Benchmarking Complex Multimodal Document Processing Pipelines: A Unified Evaluation Framework for Enterprise AI
by: Singh, Saurabh K., et al.
Published: (2026) -
Semi-Automated Knowledge Engineering and Process Mapping for Total Airport Management
by: Teo, Darryl, et al.
Published: (2026)