Leveraging LLMs to Create Content Corpora for Niche Domains
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Franklin, Zhang, Sonya, Halevy, Alon |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AI-Friendly LaTeX: Using LaTeX Code as a Knowledge Source for Retrieval-Augmented Generation
by: Verhoeff, Tom
Published: (2026)
by: Verhoeff, Tom
Published: (2026)
Neuromem: A Granular Decomposition of the Streaming Lifecycle in External Memory for LLMs
by: Zhang, Ruicheng, et al.
Published: (2026)
by: Zhang, Ruicheng, et al.
Published: (2026)
DEUCE: Dual-diversity Enhancement and Uncertainty-awareness for Cold-start Active Learning
by: Guo, Jiaxin, et al.
Published: (2025)
by: Guo, Jiaxin, et al.
Published: (2025)
Generative AI-Based Virtual Assistant using Retrieval-Augmented Generation: An evaluation study for bachelor projects
by: Verşebeniuc, Dumitru, et al.
Published: (2026)
by: Verşebeniuc, Dumitru, et al.
Published: (2026)
RARE: Redundancy-Aware Retrieval Evaluation Framework for High-Similarity Corpora
by: Cho, Hanjun, et al.
Published: (2026)
by: Cho, Hanjun, et al.
Published: (2026)
Leveraging OpenFlamingo for Multimodal Embedding Analysis of C2C Car Parts Data
by: Rashid, Maisha Binte, et al.
Published: (2025)
by: Rashid, Maisha Binte, et al.
Published: (2025)
ARTAI: An Evaluation Platform to Assess Societal Risk of Recommender Algorithms
by: Ruan, Qin, et al.
Published: (2024)
by: Ruan, Qin, et al.
Published: (2024)
A ripple in time: a discontinuity in American history
by: Kolpakov, Alexander, et al.
Published: (2023)
by: Kolpakov, Alexander, et al.
Published: (2023)
Agentic Retrieval-Augmented Generation for Financial Document Question Answering
by: Shu, Yang, et al.
Published: (2026)
by: Shu, Yang, et al.
Published: (2026)
GISTBench: Evaluating LLM User Understanding via Evidence-Based Interest Verification
by: Fostiropoulos, Iordanis, et al.
Published: (2026)
by: Fostiropoulos, Iordanis, et al.
Published: (2026)
Comparison between the Structures of Word Co-occurrence and Word Similarity Networks for Ill-formed and Well-formed Texts in Taiwan Mandarin
by: Huang, Po-Hsuan, et al.
Published: (2024)
by: Huang, Po-Hsuan, et al.
Published: (2024)
RAGged Edges: The Double-Edged Sword of Retrieval-Augmented Chatbots
by: Feldman, Philip, et al.
Published: (2024)
by: Feldman, Philip, et al.
Published: (2024)
Enhancing Long-term RAG Chatbots with Psychological Models of Memory Importance and Forgetting
by: Sumida, Ryuichi, et al.
Published: (2024)
by: Sumida, Ryuichi, et al.
Published: (2024)
Can You Detect the Difference?
by: Tarım, İsmail, et al.
Published: (2025)
by: Tarım, İsmail, et al.
Published: (2025)
Token-Oriented Object Notation vs JSON: A Benchmark of Plain and Constrained Decoding Generation
by: Matveev, Ivan
Published: (2026)
by: Matveev, Ivan
Published: (2026)
When F1 Fails: Granularity-Aware Evaluation for Dialogue Topic Segmentation
by: Coen, Michael H.
Published: (2025)
by: Coen, Michael H.
Published: (2025)
Permanent Data Encoding (PDE): A Visual Language for Semantic Compression and Knowledge Preservation in 3-Character Units
by: Tsuyuki, Yoshiharu, et al.
Published: (2025)
by: Tsuyuki, Yoshiharu, et al.
Published: (2025)
PerSoMed: A Large-Scale Balanced Dataset for Persian Social Media Text Classification
by: Chehreh, Isun, et al.
Published: (2026)
by: Chehreh, Isun, et al.
Published: (2026)
ESGenius: Benchmarking LLMs on Environmental, Social, and Governance (ESG) and Sustainability Knowledge
by: He, Chaoyue, et al.
Published: (2025)
by: He, Chaoyue, et al.
Published: (2025)
NewsScope: Schema-Grounded Cross-Domain News Claim Extraction with Open Models
by: Pandya, Nidhi
Published: (2025)
by: Pandya, Nidhi
Published: (2025)
Agentic Framework for Political Biography Extraction
by: Zhu, Yifei, et al.
Published: (2026)
by: Zhu, Yifei, et al.
Published: (2026)
CIG: Measuring Conversational Information Gain in Deliberative Dialogues with Semantic Memory Dynamics
by: Chen, Ming-Bin, et al.
Published: (2026)
by: Chen, Ming-Bin, et al.
Published: (2026)
Extending AI for Research to the Humanities: A Multi-Agent Framework for Evidence-Grounded Scholarship
by: Pan, Yating, et al.
Published: (2026)
by: Pan, Yating, et al.
Published: (2026)
News Without Borders: Domain Adaptation of Multilingual Sentence Embeddings for Cross-lingual News Recommendation
by: Iana, Andreea, et al.
Published: (2024)
by: Iana, Andreea, et al.
Published: (2024)
BALI: Enhancing Biomedical Language Representations through Knowledge Graph and Language Model Alignment
by: Sakhovskiy, Andrey, et al.
Published: (2025)
by: Sakhovskiy, Andrey, et al.
Published: (2025)
Tokenization Disparities as Infrastructure Bias: How Subword Systems Create Inequities in LLM Access and Efficiency
by: Teklehaymanot, Hailay Kidu, et al.
Published: (2025)
by: Teklehaymanot, Hailay Kidu, et al.
Published: (2025)
Prompt-RAG: Pioneering Vector Embedding-Free Retrieval-Augmented Generation in Niche Domains, Exemplified by Korean Medicine
by: Kang, Bongsu, et al.
Published: (2024)
by: Kang, Bongsu, et al.
Published: (2024)
Leveraging LLMs to Enable Natural Language Search on Go-to-market Platforms
by: Yao, Jesse, et al.
Published: (2024)
by: Yao, Jesse, et al.
Published: (2024)
TrueReason: An Exemplar Personalised Learning System Integrating Reasoning with Foundational Models
by: Bulathwela, Sahan, et al.
Published: (2025)
by: Bulathwela, Sahan, et al.
Published: (2025)
Key Algorithms for Keyphrase Generation: Instruction-Based LLMs for Russian Scientific Keyphrases
by: Glazkova, Anna, et al.
Published: (2024)
by: Glazkova, Anna, et al.
Published: (2024)
LegalBench-BR: A Benchmark for Evaluating Large Language Models on Brazilian Legal Decision Classification
by: Neto, Pedro Barbosa de Carvalho
Published: (2026)
by: Neto, Pedro Barbosa de Carvalho
Published: (2026)
LLM-based Extraction of Contradictions from Patents
by: Trapp, Stefan, et al.
Published: (2024)
by: Trapp, Stefan, et al.
Published: (2024)
Data Processing for the OpenGPT-X Model Family
by: Brandizzi, Nicolo', et al.
Published: (2024)
by: Brandizzi, Nicolo', et al.
Published: (2024)
Supervised Semantic Differential for Cross-Cultural Concept Analysis: A Case Study of Human Affect
by: Sikora, Jan, et al.
Published: (2026)
by: Sikora, Jan, et al.
Published: (2026)
Using Instruction-Tuned Large Language Models to Identify Indicators of Vulnerability in Police Incident Narratives
by: Relins, Sam, et al.
Published: (2024)
by: Relins, Sam, et al.
Published: (2024)
Grounding Agent Memory in Contextual Intent
by: Yang, Ruozhen, et al.
Published: (2026)
by: Yang, Ruozhen, et al.
Published: (2026)
A Prompt-Aware Structuring Framework for Reliable Reuse of AI-Generated Content in the Agentic Web
by: Egami, Shusaku, et al.
Published: (2026)
by: Egami, Shusaku, et al.
Published: (2026)
Flippi: End To End GenAI Assistant for E-Commerce
by: Rajasekar, Anand A., et al.
Published: (2025)
by: Rajasekar, Anand A., et al.
Published: (2025)
Evaluation of Chunking Strategies for Effective Text Embedding in Low-Resource Language on Agricultural Documents
by: Chhoun, Sovandara, et al.
Published: (2026)
by: Chhoun, Sovandara, et al.
Published: (2026)
Stage-Audit: Auditable Source-Frontier Discovery for Cross-Wiki Tables
by: Shen, Chen
Published: (2026)
by: Shen, Chen
Published: (2026)
Similar Items
-
AI-Friendly LaTeX: Using LaTeX Code as a Knowledge Source for Retrieval-Augmented Generation
by: Verhoeff, Tom
Published: (2026) -
Neuromem: A Granular Decomposition of the Streaming Lifecycle in External Memory for LLMs
by: Zhang, Ruicheng, et al.
Published: (2026) -
DEUCE: Dual-diversity Enhancement and Uncertainty-awareness for Cold-start Active Learning
by: Guo, Jiaxin, et al.
Published: (2025) -
Generative AI-Based Virtual Assistant using Retrieval-Augmented Generation: An evaluation study for bachelor projects
by: Verşebeniuc, Dumitru, et al.
Published: (2026) -
RARE: Redundancy-Aware Retrieval Evaluation Framework for High-Similarity Corpora
by: Cho, Hanjun, et al.
Published: (2026)