MT-RAIG: Novel Benchmark and Evaluation Framework for Retrieval-Augmented Insight Generation over Multiple Tables
Fuente:
arXiv
Guardado en:
| Autores principales: | Seo, Kwangwook, Kwon, Donguk, Lee, Dongha |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Unveiling Implicit Table Knowledge with Question-Then-Pinpoint Reasoner for Insightful Table Summarization
por: Seo, Kwangwook, et al.
Publicado: (2024)
por: Seo, Kwangwook, et al.
Publicado: (2024)
P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist
por: Seo, Kwangwook, et al.
Publicado: (2026)
por: Seo, Kwangwook, et al.
Publicado: (2026)
BESPOKE: Benchmark for Search-Augmented Large Language Model Personalization via Diagnostic Feedback
por: Kim, Hyunseo, et al.
Publicado: (2025)
por: Kim, Hyunseo, et al.
Publicado: (2025)
In Their Own Words: Reasoning Traces Tailored for Small Models Make Them Better Reasoners
por: Kim, Jaehoon, et al.
Publicado: (2025)
por: Kim, Jaehoon, et al.
Publicado: (2025)
Region4Web: Rethinking Observation Space Granularity for Web Agents
por: Kwon, Donguk, et al.
Publicado: (2026)
por: Kwon, Donguk, et al.
Publicado: (2026)
VerifiNER: Verification-augmented NER via Knowledge-grounded Reasoning with Large Language Models
por: Kim, Seoyeon, et al.
Publicado: (2024)
por: Kim, Seoyeon, et al.
Publicado: (2024)
Personalizing Large Language Models using Retrieval Augmented Generation and Knowledge Graph
por: Prahlad, Deeksha, et al.
Publicado: (2025)
por: Prahlad, Deeksha, et al.
Publicado: (2025)
Evaluating Cultural Knowledge Processing in Large Language Models: A Cognitive Benchmarking Framework Integrating Retrieval-Augmented Generation
por: Lee, Hung-Shin, et al.
Publicado: (2025)
por: Lee, Hung-Shin, et al.
Publicado: (2025)
Can Large Language Models be Effective Online Opinion Miners?
por: Heo, Ryang, et al.
Publicado: (2025)
por: Heo, Ryang, et al.
Publicado: (2025)
kRAIG: A Natural Language-Driven Agent for Automated DataOps Pipeline Generation
por: Siva, Rohan, et al.
Publicado: (2026)
por: Siva, Rohan, et al.
Publicado: (2026)
On the Effectiveness of Integration Methods for Multimodal Dialogue Response Retrieval
por: Jang, Seongbo, et al.
Publicado: (2025)
por: Jang, Seongbo, et al.
Publicado: (2025)
YpathRAG:A Retrieval-Augmented Generation Framework and Benchmark for Pathology
por: Yu, Deshui, et al.
Publicado: (2025)
por: Yu, Deshui, et al.
Publicado: (2025)
Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History
por: Kim, Serin, et al.
Publicado: (2026)
por: Kim, Serin, et al.
Publicado: (2026)
Verified Misguidance: Measuring Structural Citation Failures in Search-Augmented LLMs
por: Seo, Yongsik, et al.
Publicado: (2026)
por: Seo, Yongsik, et al.
Publicado: (2026)
Quantifying Genuine Awareness in Hallucination Prediction Beyond Question-Side Shortcuts
por: Seo, Yeongbin, et al.
Publicado: (2025)
por: Seo, Yeongbin, et al.
Publicado: (2025)
MTRAG: A Multi-Turn Conversational Benchmark for Evaluating Retrieval-Augmented Generation Systems
por: Katsis, Yannis, et al.
Publicado: (2025)
por: Katsis, Yannis, et al.
Publicado: (2025)
Train-Attention: Meta-Learning Where to Focus in Continual Knowledge Learning
por: Seo, Yeongbin, et al.
Publicado: (2024)
por: Seo, Yeongbin, et al.
Publicado: (2024)
MIRAGE: A Metric-Intensive Benchmark for Retrieval-Augmented Generation Evaluation
por: Park, Chanhee, et al.
Publicado: (2025)
por: Park, Chanhee, et al.
Publicado: (2025)
Benchmarking Retrieval-Augmented Generation for Medicine
por: Xiong, Guangzhi, et al.
Publicado: (2024)
por: Xiong, Guangzhi, et al.
Publicado: (2024)
Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents
por: Kim, Wonjoong, et al.
Publicado: (2025)
por: Kim, Wonjoong, et al.
Publicado: (2025)
Unanswerability Evaluation for Retrieval Augmented Generation
por: Peng, Xiangyu, et al.
Publicado: (2024)
por: Peng, Xiangyu, et al.
Publicado: (2024)
MAGIC: A Multi-Hop and Graph-Based Benchmark for Inter-Context Conflicts in Retrieval-Augmented Generation
por: Lee, Jungyeon, et al.
Publicado: (2025)
por: Lee, Jungyeon, et al.
Publicado: (2025)
CReSt: A Comprehensive Benchmark for Retrieval-Augmented Generation with Complex Reasoning over Structured Documents
por: Khang, Minsoo, et al.
Publicado: (2025)
por: Khang, Minsoo, et al.
Publicado: (2025)
IPQA: A Benchmark for Core Intent Identification in Personalized Question Answering
por: Kim, Jieyong, et al.
Publicado: (2025)
por: Kim, Jieyong, et al.
Publicado: (2025)
Evaluating Retrieval Quality in Retrieval-Augmented Generation
por: Salemi, Alireza, et al.
Publicado: (2024)
por: Salemi, Alireza, et al.
Publicado: (2024)
Ragas: Automated Evaluation of Retrieval Augmented Generation
por: Es, Shahul, et al.
Publicado: (2023)
por: Es, Shahul, et al.
Publicado: (2023)
Not All Languages are Equal: Insights into Multilingual Retrieval-Augmented Generation
por: Wu, Suhang, et al.
Publicado: (2024)
por: Wu, Suhang, et al.
Publicado: (2024)
Make Compound Sentences Simple to Analyze: Learning to Split Sentences for Aspect-based Sentiment Analysis
por: Seo, Yongsik, et al.
Publicado: (2024)
por: Seo, Yongsik, et al.
Publicado: (2024)
Legal-DC: Benchmarking Retrieval-Augmented Generation for Legal Documents
por: Li, Yaocong, et al.
Publicado: (2026)
por: Li, Yaocong, et al.
Publicado: (2026)
Benchmarking Retrieval-Augmented Generation for Chemistry
por: Zhong, Xianrui, et al.
Publicado: (2025)
por: Zhong, Xianrui, et al.
Publicado: (2025)
CRAB: A Benchmark for Evaluating Curation of Retrieval-Augmented LLMs in Biomedicine
por: Zhong, Hanmeng, et al.
Publicado: (2025)
por: Zhong, Hanmeng, et al.
Publicado: (2025)
Stop Playing the Guessing Game! Target-free User Simulation for Evaluating Conversational Recommender Systems
por: Kim, Sunghwan, et al.
Publicado: (2024)
por: Kim, Sunghwan, et al.
Publicado: (2024)
Rethinking Retrieval-Augmented Generation for Medicine: A Large-Scale, Systematic Expert Evaluation and Practical Insights
por: Kim, Hyunjae, et al.
Publicado: (2025)
por: Kim, Hyunjae, et al.
Publicado: (2025)
RARE: Retrieval-Aware Robustness Evaluation for Retrieval-Augmented Generation Systems
por: Zeng, Yixiao, et al.
Publicado: (2025)
por: Zeng, Yixiao, et al.
Publicado: (2025)
MEMERAG: A Multilingual End-to-End Meta-Evaluation Benchmark for Retrieval Augmented Generation
por: Blandón, María Andrea Cruz, et al.
Publicado: (2025)
por: Blandón, María Andrea Cruz, et al.
Publicado: (2025)
DomainRAG: A Chinese Benchmark for Evaluating Domain-specific Retrieval-Augmented Generation
por: Wang, Shuting, et al.
Publicado: (2024)
por: Wang, Shuting, et al.
Publicado: (2024)
ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems
por: Saad-Falcon, Jon, et al.
Publicado: (2023)
por: Saad-Falcon, Jon, et al.
Publicado: (2023)
Multiple Abstraction Level Retrieve Augment Generation
por: Zheng, Zheng, et al.
Publicado: (2025)
por: Zheng, Zheng, et al.
Publicado: (2025)
PRGB Benchmark: A Robust Placeholder-Assisted Algorithm for Benchmarking Retrieval-Augmented Generation
por: Tan, Zhehao, et al.
Publicado: (2025)
por: Tan, Zhehao, et al.
Publicado: (2025)
ReTAG: Retrieval-Enhanced, Topic-Augmented Graph-Based Global Sensemaking
por: Kim, Boyoung, et al.
Publicado: (2025)
por: Kim, Boyoung, et al.
Publicado: (2025)
Ejemplares similares
-
Unveiling Implicit Table Knowledge with Question-Then-Pinpoint Reasoner for Insightful Table Summarization
por: Seo, Kwangwook, et al.
Publicado: (2024) -
P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist
por: Seo, Kwangwook, et al.
Publicado: (2026) -
BESPOKE: Benchmark for Search-Augmented Large Language Model Personalization via Diagnostic Feedback
por: Kim, Hyunseo, et al.
Publicado: (2025) -
In Their Own Words: Reasoning Traces Tailored for Small Models Make Them Better Reasoners
por: Kim, Jaehoon, et al.
Publicado: (2025) -
Region4Web: Rethinking Observation Space Granularity for Web Agents
por: Kwon, Donguk, et al.
Publicado: (2026)