AutoPureData: Automated Filtering of Undesirable Web Data to Update LLM Knowledge
Fuente:
arXiv
Saved in:
| Main Author: | Vadlapati, Praneeth |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LML-DAP: Language Model Learning a Dataset for Data-Augmented Prediction
by: Vadlapati, Praneeth
Published: (2024)
by: Vadlapati, Praneeth
Published: (2024)
TrustDataFilter:Leveraging Trusted Knowledge Base Data for More Effective Filtering of Unknown Information
by: Zhang, Jinghong, et al.
Published: (2025)
by: Zhang, Jinghong, et al.
Published: (2025)
Automated Extraction and Creation of FBS Design Reasoning Knowledge Graphs from Structured Data in Product Catalogues Lacking Contextual Information
by: Sahadevan, Vijayalaxmi, et al.
Published: (2024)
by: Sahadevan, Vijayalaxmi, et al.
Published: (2024)
Auto-ARGUE: LLM-Based Report Generation Evaluation
by: Walden, William, et al.
Published: (2025)
by: Walden, William, et al.
Published: (2025)
Automated Neural Patent Landscaping in the Small Data Regime
by: Erana, Tisa Islam, et al.
Published: (2024)
by: Erana, Tisa Islam, et al.
Published: (2024)
The Synergy of Automated Pipelines with Prompt Engineering and Generative AI in Web Crawling
by: Huang, Chau-Jian
Published: (2024)
by: Huang, Chau-Jian
Published: (2024)
Self-Describing Structured Data with Dual-Layer Guidance: A Lightweight Alternative to RAG for Precision Retrieval in Large-Scale LLM Knowledge Navigation
by: Liu, Hung Ming
Published: (2026)
by: Liu, Hung Ming
Published: (2026)
SPARQL Query Generation with LLMs: Measuring the Impact of Training Data Memorization and Knowledge Injection
by: Gashkov, Aleksandr, et al.
Published: (2025)
by: Gashkov, Aleksandr, et al.
Published: (2025)
DR.EHR: Dense Retrieval for Electronic Health Record with Knowledge Injection and Synthetic Data
by: Zhao, Zhengyun, et al.
Published: (2025)
by: Zhao, Zhengyun, et al.
Published: (2025)
Semi-Automated Knowledge Engineering and Process Mapping for Total Airport Management
by: Teo, Darryl, et al.
Published: (2026)
by: Teo, Darryl, et al.
Published: (2026)
DALDALL: Data Augmentation for Lexical and Semantic Diverse in Legal Domain by leveraging LLM-Persona
by: Choi, Janghyeok, et al.
Published: (2026)
by: Choi, Janghyeok, et al.
Published: (2026)
TRAWL: External Knowledge-Enhanced Recommendation with LLM Assistance
by: Luo, Weiqing, et al.
Published: (2024)
by: Luo, Weiqing, et al.
Published: (2024)
Enhancing LLM Medical Coding with Structured External Knowledge
by: Gan, Yidong, et al.
Published: (2026)
by: Gan, Yidong, et al.
Published: (2026)
KG-LLM-Bench: A Scalable Benchmark for Evaluating LLM Reasoning on Textualized Knowledge Graphs
by: Markowitz, Elan, et al.
Published: (2025)
by: Markowitz, Elan, et al.
Published: (2025)
Enhancing LLM Generation with Knowledge Hypergraph for Evidence-Based Medicine
by: Dou, Chengfeng, et al.
Published: (2025)
by: Dou, Chengfeng, et al.
Published: (2025)
Knowledge Navigator: LLM-guided Browsing Framework for Exploratory Search in Scientific Literature
by: Katz, Uri, et al.
Published: (2024)
by: Katz, Uri, et al.
Published: (2024)
Table Meets LLM: Can Large Language Models Understand Structured Table Data? A Benchmark and Empirical Study
by: Sui, Yuan, et al.
Published: (2023)
by: Sui, Yuan, et al.
Published: (2023)
eCeLLM: Generalizing Large Language Models for E-commerce from Large-scale, High-quality Instruction Data
by: Peng, Bo, et al.
Published: (2024)
by: Peng, Bo, et al.
Published: (2024)
AutoData: A Multi-Agent System for Open Web Data Collection
by: Ma, Tianyi, et al.
Published: (2025)
by: Ma, Tianyi, et al.
Published: (2025)
Structure-R1: Dynamically Leveraging Structural Knowledge in LLM Reasoning through Reinforcement Learning
by: Wu, Junlin, et al.
Published: (2025)
by: Wu, Junlin, et al.
Published: (2025)
Query Attribute Modeling: Improving search relevance with Semantic Search and Meta Data Filtering
by: Menon, Karthik, et al.
Published: (2025)
by: Menon, Karthik, et al.
Published: (2025)
AutoSurvey: Large Language Models Can Automatically Write Surveys
by: Wang, Yidong, et al.
Published: (2024)
by: Wang, Yidong, et al.
Published: (2024)
LlamaRec-LKG-RAG: A Single-Pass, Learnable Knowledge Graph-RAG Framework for LLM-Based Ranking
by: Azizi, Vahid, et al.
Published: (2025)
by: Azizi, Vahid, et al.
Published: (2025)
From Conceptual Data Models to Multimodal Representation
by: Stockinger, Peter
Published: (2025)
by: Stockinger, Peter
Published: (2025)
Cross-Domain Web Information Extraction at Pinterest
by: Farag, Michael, et al.
Published: (2025)
by: Farag, Michael, et al.
Published: (2025)
DocQAC: Adaptive Trie-Guided Decoding for Effective In-Document Query Auto-Completion
by: Mehta, Rahul, et al.
Published: (2026)
by: Mehta, Rahul, et al.
Published: (2026)
Going over Fine Web with a Fine-Tooth Comb: Technical Report of Indexing Fine Web for Problematic Content Search and Retrieval
by: Marinas, Inés Altemir, et al.
Published: (2025)
by: Marinas, Inés Altemir, et al.
Published: (2025)
Efficient Evaluation of Large Language Models via Collaborative Filtering
by: Zhong, Xu-Xiang, et al.
Published: (2025)
by: Zhong, Xu-Xiang, et al.
Published: (2025)
The Effects of Hallucinations in Synthetic Training Data for Relation Extraction
by: Rogulsky, Steven, et al.
Published: (2024)
by: Rogulsky, Steven, et al.
Published: (2024)
A Survey on Recent Advances in Conversational Data Generation
by: Soudani, Heydar, et al.
Published: (2024)
by: Soudani, Heydar, et al.
Published: (2024)
Large Language Models Empowered Personalized Web Agents
by: Cai, Hongru, et al.
Published: (2024)
by: Cai, Hongru, et al.
Published: (2024)
Little Giants: Synthesizing High-Quality Embedding Data at Scale
by: Chen, Haonan, et al.
Published: (2024)
by: Chen, Haonan, et al.
Published: (2024)
$τ$-Knowledge: Evaluating Conversational Agents over Unstructured Knowledge
by: Shi, Quan, et al.
Published: (2026)
by: Shi, Quan, et al.
Published: (2026)
ClusterRAG: Cluster-Based Collaborative Filtering for Personalized Retrieval-Augmented Generation
by: Nkhata, Gibson, et al.
Published: (2026)
by: Nkhata, Gibson, et al.
Published: (2026)
FinRetrieval: A Benchmark for Financial Data Retrieval by AI Agents
by: Kim, Eric Y., et al.
Published: (2026)
by: Kim, Eric Y., et al.
Published: (2026)
MapQA: Open-domain Geospatial Question Answering on Map Data
by: Li, Zekun, et al.
Published: (2025)
by: Li, Zekun, et al.
Published: (2025)
WebThinker: Empowering Large Reasoning Models with Deep Research Capability
by: Li, Xiaoxi, et al.
Published: (2025)
by: Li, Xiaoxi, et al.
Published: (2025)
Auto-Cypher: Improving LLMs on Cypher generation via LLM-supervised generation-verification framework
by: Tiwari, Aman, et al.
Published: (2024)
by: Tiwari, Aman, et al.
Published: (2024)
RAG-KG-IL: A Multi-Agent Hybrid Framework for Reducing Hallucinations and Enhancing LLM Reasoning through RAG and Incremental Knowledge Graph Learning Integration
by: Yu, Hong Qing, et al.
Published: (2025)
by: Yu, Hong Qing, et al.
Published: (2025)
Exploring the Escalation of Source Bias in User, Data, and Recommender System Feedback Loop
by: Zhou, Yuqi, et al.
Published: (2024)
by: Zhou, Yuqi, et al.
Published: (2024)
Similar Items
-
LML-DAP: Language Model Learning a Dataset for Data-Augmented Prediction
by: Vadlapati, Praneeth
Published: (2024) -
TrustDataFilter:Leveraging Trusted Knowledge Base Data for More Effective Filtering of Unknown Information
by: Zhang, Jinghong, et al.
Published: (2025) -
Automated Extraction and Creation of FBS Design Reasoning Knowledge Graphs from Structured Data in Product Catalogues Lacking Contextual Information
by: Sahadevan, Vijayalaxmi, et al.
Published: (2024) -
Auto-ARGUE: LLM-Based Report Generation Evaluation
by: Walden, William, et al.
Published: (2025) -
Automated Neural Patent Landscaping in the Small Data Regime
by: Erana, Tisa Islam, et al.
Published: (2024)