Context-Driven Index Trimming: A Data Quality Perspective to Enhancing Precision of RALMs
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Kexin, Jin, Ruochun, Wang, Xi, Chen, Huan, Ren, Jing, Tang, Yuhua |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CAST: Character-and-Scene Episodic Memory for Agents
by: Ma, Kexin, et al.
Published: (2026)
by: Ma, Kexin, et al.
Published: (2026)
ContextCache: Context-Aware Semantic Cache for Multi-Turn Queries in Large Language Models
by: Yan, Jianxin, et al.
Published: (2025)
by: Yan, Jianxin, et al.
Published: (2025)
From Facts to Insights: A Persona-Driven Dual Memory Framework and Dataset for Role-Playing Agents
by: Zhang, Rongsheng, et al.
Published: (2026)
by: Zhang, Rongsheng, et al.
Published: (2026)
LLM and Agent-Driven Data Analysis: A Systematic Approach for Enterprise Applications and System-level Deployment
by: Wang, Xi, et al.
Published: (2025)
by: Wang, Xi, et al.
Published: (2025)
Grounding Natural Language to SQL Translation with Data-Based Self-Explanations
by: Fan, Yuankai, et al.
Published: (2024)
by: Fan, Yuankai, et al.
Published: (2024)
Enhancing Knowledge Graph Completion with Entity Neighborhood and Relation Context
by: Chen, Jianfang, et al.
Published: (2025)
by: Chen, Jianfang, et al.
Published: (2025)
OmniSQL: Synthesizing High-quality Text-to-SQL Data at Scale
by: Li, Haoyang, et al.
Published: (2025)
by: Li, Haoyang, et al.
Published: (2025)
LakeHopper: Cross Data Lakes Column Type Annotation through Model Adaptation
by: Sun, Yushi, et al.
Published: (2026)
by: Sun, Yushi, et al.
Published: (2026)
Prompt Engineering Techniques for Context-dependent Text-to-SQL in Arabic
by: Almohaimeed, Saleh, et al.
Published: (2025)
by: Almohaimeed, Saleh, et al.
Published: (2025)
Structured Prompt Language: Declarative Context Management for LLMs
by: Gong, Wen G.
Published: (2026)
by: Gong, Wen G.
Published: (2026)
Filling Memory Gaps: Enhancing Continual Semantic Parsing via SQL Syntax Variance-Guided LLMs without Real Data Replay
by: Liu, Ruiheng, et al.
Published: (2024)
by: Liu, Ruiheng, et al.
Published: (2024)
DataLab: A Unified Platform for LLM-Powered Business Intelligence
by: Weng, Luoxuan, et al.
Published: (2024)
by: Weng, Luoxuan, et al.
Published: (2024)
FineWeb-zhtw: Scalable Curation of Traditional Chinese Text Data from the Web
by: Lin, Cheng-Wei, et al.
Published: (2024)
by: Lin, Cheng-Wei, et al.
Published: (2024)
ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMs
by: Qi, Yanlin, et al.
Published: (2026)
by: Qi, Yanlin, et al.
Published: (2026)
Data-Driven Information Extraction and Enrichment of Molecular Profiling Data for Cancer Cell Lines
by: Smith, Ellery, et al.
Published: (2023)
by: Smith, Ellery, et al.
Published: (2023)
Enhancing Financial Market Predictions: Causality-Driven Feature Selection
by: Liang, Wenhao, et al.
Published: (2024)
by: Liang, Wenhao, et al.
Published: (2024)
Learning from Imperfect Data: Towards Efficient Knowledge Distillation of Autoregressive Language Models for Text-to-SQL
by: Zhong, Qihuang, et al.
Published: (2024)
by: Zhong, Qihuang, et al.
Published: (2024)
MAGIC: Generating Self-Correction Guideline for In-Context Text-to-SQL
by: Askari, Arian, et al.
Published: (2024)
by: Askari, Arian, et al.
Published: (2024)
LLM-R2: A Large Language Model Enhanced Rule-based Rewrite System for Boosting Query Efficiency
by: Li, Zhaodonghui, et al.
Published: (2024)
by: Li, Zhaodonghui, et al.
Published: (2024)
Replacing Multi-Step Assembly of Data Preparation Pipelines with One-Step LLM Pipeline Generation for Table QA
by: Li, Fengyu, et al.
Published: (2026)
by: Li, Fengyu, et al.
Published: (2026)
TCM-Ladder: A Benchmark for Multimodal Question Answering on Traditional Chinese Medicine
by: Xie, Jiacheng, et al.
Published: (2025)
by: Xie, Jiacheng, et al.
Published: (2025)
ACCIO: Table Understanding Enhanced via Contrastive Learning with Aggregations
by: Cho, Whanhee
Published: (2024)
by: Cho, Whanhee
Published: (2024)
PURPLE: Making a Large Language Model a Better SQL Writer
by: Ren, Tonghui, et al.
Published: (2024)
by: Ren, Tonghui, et al.
Published: (2024)
Uncovering the Impact of Chain-of-Thought Reasoning for Direct Preference Optimization: Lessons from Text-to-SQL
by: Liu, Hanbing, et al.
Published: (2025)
by: Liu, Hanbing, et al.
Published: (2025)
SQL-Encoder: Improving NL2SQL In-Context Learning Through a Context-Aware Encoder
by: Pourreza, Mohammadreza, et al.
Published: (2024)
by: Pourreza, Mohammadreza, et al.
Published: (2024)
ORANGE: An Online Reflection ANd GEneration framework with Domain Knowledge for Text-to-SQL
by: Jiao, Yiwen, et al.
Published: (2025)
by: Jiao, Yiwen, et al.
Published: (2025)
RADAR: Benchmarking Language Models on Imperfect Tabular Data
by: Gu, Ken, et al.
Published: (2025)
by: Gu, Ken, et al.
Published: (2025)
Rethinking Agentic Workflows: Evaluating Inference-Based Test-Time Scaling Strategies in Text2SQL Tasks
by: Guo, Jiajing, et al.
Published: (2025)
by: Guo, Jiajing, et al.
Published: (2025)
How Robust Are Router-LLMs? Analysis of the Fragility of LLM Routing Capabilities
by: Kassem, Aly M., et al.
Published: (2025)
by: Kassem, Aly M., et al.
Published: (2025)
Advancing the Database of Cross-Linguistic Colexifications with New Workflows and Data
by: Tjuka, Annika, et al.
Published: (2025)
by: Tjuka, Annika, et al.
Published: (2025)
Match, Compare, or Select? An Investigation of Large Language Models for Entity Matching
by: Wang, Tianshu, et al.
Published: (2024)
by: Wang, Tianshu, et al.
Published: (2024)
DP-Bench: A Benchmark for Evaluating Data Product Creation Systems
by: Chowdhury, Faisal, et al.
Published: (2025)
by: Chowdhury, Faisal, et al.
Published: (2025)
CodeS: Towards Building Open-source Language Models for Text-to-SQL
by: Li, Haoyang, et al.
Published: (2024)
by: Li, Haoyang, et al.
Published: (2024)
Datasets for Verb Alternations across Languages: BLM Templates and Data Augmentation Strategies
by: Samo, Giuseppe, et al.
Published: (2026)
by: Samo, Giuseppe, et al.
Published: (2026)
AutoDCWorkflow: LLM-based Data Cleaning Workflow Auto-Generation and Benchmark
by: Li, Lan, et al.
Published: (2024)
by: Li, Lan, et al.
Published: (2024)
SQLfuse: Enhancing Text-to-SQL Performance through Comprehensive LLM Synergy
by: Zhang, Tingkai, et al.
Published: (2024)
by: Zhang, Tingkai, et al.
Published: (2024)
Graph Learning in the Era of LLMs: A Survey from the Perspective of Data, Models, and Tasks
by: Li, Xunkai, et al.
Published: (2024)
by: Li, Xunkai, et al.
Published: (2024)
Columbo: Expanding Abbreviated Column Names for Tabular Data Using Large Language Models
by: Cai, Ting, et al.
Published: (2025)
by: Cai, Ting, et al.
Published: (2025)
LLM-Enhanced Data Management
by: Zhou, Xuanhe, et al.
Published: (2024)
by: Zhou, Xuanhe, et al.
Published: (2024)
Bidirectional Chinese and English Passive Sentences Dataset for Machine Translation
by: Ma, Xinyue, et al.
Published: (2026)
by: Ma, Xinyue, et al.
Published: (2026)
Similar Items
-
CAST: Character-and-Scene Episodic Memory for Agents
by: Ma, Kexin, et al.
Published: (2026) -
ContextCache: Context-Aware Semantic Cache for Multi-Turn Queries in Large Language Models
by: Yan, Jianxin, et al.
Published: (2025) -
From Facts to Insights: A Persona-Driven Dual Memory Framework and Dataset for Role-Playing Agents
by: Zhang, Rongsheng, et al.
Published: (2026) -
LLM and Agent-Driven Data Analysis: A Systematic Approach for Enterprise Applications and System-level Deployment
by: Wang, Xi, et al.
Published: (2025) -
Grounding Natural Language to SQL Translation with Data-Based Self-Explanations
by: Fan, Yuankai, et al.
Published: (2024)