Unanswerability Evaluation for Retrieval Augmented Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Peng, Xiangyu, Choubey, Prafulla Kumar, Xiong, Caiming, Wu, Chien-Sheng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Turning Conversations into Workflows: A Framework to Extract and Evaluate Dialog Workflows for Service AI Agents
by: Choubey, Prafulla Kumar, et al.
Published: (2025)
by: Choubey, Prafulla Kumar, et al.
Published: (2025)
Do RAG Systems Cover What Matters? Evaluating and Optimizing Responses with Sub-Question Coverage
by: Xie, Kaige, et al.
Published: (2024)
by: Xie, Kaige, et al.
Published: (2024)
Benchmarking Deep Search over Heterogeneous Enterprise Data
by: Choubey, Prafulla Kumar, et al.
Published: (2025)
by: Choubey, Prafulla Kumar, et al.
Published: (2025)
Agentic Uncertainty Quantification
by: Zhang, Jiaxin, et al.
Published: (2026)
by: Zhang, Jiaxin, et al.
Published: (2026)
Embrace Divergence for Richer Insights: A Multi-document Summarization Benchmark and a Case Study on Summarizing Diverse Information from News Articles
by: Huang, Kung-Hsiang, et al.
Published: (2023)
by: Huang, Kung-Hsiang, et al.
Published: (2023)
Scaling Knowledge Graph Construction through Synthetic Data Generation and Distillation
by: Choubey, Prafulla Kumar, et al.
Published: (2024)
by: Choubey, Prafulla Kumar, et al.
Published: (2024)
SiReRAG: Indexing Similar and Related Information for Multihop Reasoning
by: Zhang, Nan, et al.
Published: (2024)
by: Zhang, Nan, et al.
Published: (2024)
CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions
by: Huang, Kung-Hsiang, et al.
Published: (2025)
by: Huang, Kung-Hsiang, et al.
Published: (2025)
Dont Stop Early: Scalable Enterprise Deep Research with Controlled Information Flow and Evidence-Aware Termination
by: Choubey, Prafulla Kumar, et al.
Published: (2026)
by: Choubey, Prafulla Kumar, et al.
Published: (2026)
Nudging the Boundaries of LLM Reasoning
by: Chen, Justin Chih-Yao, et al.
Published: (2025)
by: Chen, Justin Chih-Yao, et al.
Published: (2025)
GTA: Generating Long-Horizon Tasks for Web Agents at Scale
by: Huang, Tenghao, et al.
Published: (2026)
by: Huang, Tenghao, et al.
Published: (2026)
ReGenesis: LLMs can Grow into Reasoning Generalists via Self-Improvement
by: Peng, Xiangyu, et al.
Published: (2024)
by: Peng, Xiangyu, et al.
Published: (2024)
UNIDOC-BENCH: A Unified Benchmark for Document-Centric Multimodal RAG
by: Peng, Xiangyu, et al.
Published: (2025)
by: Peng, Xiangyu, et al.
Published: (2025)
Agentic Confidence Calibration
by: Zhang, Jiaxin, et al.
Published: (2026)
by: Zhang, Jiaxin, et al.
Published: (2026)
Shared Imagination: LLMs Hallucinate Alike
by: Zhou, Yilun, et al.
Published: (2024)
by: Zhou, Yilun, et al.
Published: (2024)
Investigating Retrieval-Augmented Generation Systems on Unanswerable, Uncheatable, Realistic, Multi-hop Queries
by: Liu, Gabrielle Kaili-May, et al.
Published: (2025)
by: Liu, Gabrielle Kaili-May, et al.
Published: (2025)
Summary of a Haystack: A Challenge to Long-Context LLMs and RAG Systems
by: Laban, Philippe, et al.
Published: (2024)
by: Laban, Philippe, et al.
Published: (2024)
Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment
by: Laban, Philippe, et al.
Published: (2023)
by: Laban, Philippe, et al.
Published: (2023)
Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
by: Xu, Austin, et al.
Published: (2025)
by: Xu, Austin, et al.
Published: (2025)
UAQFact: Evaluating Factual Knowledge Utilization of LLMs on Unanswerable Questions
by: Tan, Chuanyuan, et al.
Published: (2025)
by: Tan, Chuanyuan, et al.
Published: (2025)
Breaking It Down: Domain-Aware Semantic Segmentation for Retrieval Augmented Generation
by: Allamraju, Aparajitha, et al.
Published: (2025)
by: Allamraju, Aparajitha, et al.
Published: (2025)
GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness
by: Huang, Kung-Hsiang, et al.
Published: (2025)
by: Huang, Kung-Hsiang, et al.
Published: (2025)
Harnessing RLHF for Robust Unanswerability Recognition and Trustworthy Response Generation in LLMs
by: Lin, Shuyuan, et al.
Published: (2025)
by: Lin, Shuyuan, et al.
Published: (2025)
BingoGuard: LLM Content Moderation Tools with Risk Levels
by: Yin, Fan, et al.
Published: (2025)
by: Yin, Fan, et al.
Published: (2025)
LLMs are Biased Evaluators But Not Biased for Retrieval Augmented Generation
by: Chen, Yen-Shan, et al.
Published: (2024)
by: Chen, Yen-Shan, et al.
Published: (2024)
RARE: Retrieval-Aware Robustness Evaluation for Retrieval-Augmented Generation Systems
by: Zeng, Yixiao, et al.
Published: (2025)
by: Zeng, Yixiao, et al.
Published: (2025)
Answering the Unanswerable Is to Err Knowingly: Analyzing and Mitigating Abstention Failures in Large Reasoning Models
by: Liu, Yi, et al.
Published: (2025)
by: Liu, Yi, et al.
Published: (2025)
Towards Unbiased Evaluation of Detecting Unanswerable Questions in EHRSQL
by: Yang, Yongjin, et al.
Published: (2024)
by: Yang, Yongjin, et al.
Published: (2024)
Towards Automated Smart Contract Generation: Evaluation, Benchmarking, and Retrieval-Augmented Repair
by: Chen, Zaoyu, et al.
Published: (2025)
by: Chen, Zaoyu, et al.
Published: (2025)
Comprehensive and Practical Evaluation of Retrieval-Augmented Generation Systems for Medical Question Answering
by: Ngo, Nghia Trung, et al.
Published: (2024)
by: Ngo, Nghia Trung, et al.
Published: (2024)
Evaluating Retrieval Quality in Retrieval-Augmented Generation
by: Salemi, Alireza, et al.
Published: (2024)
by: Salemi, Alireza, et al.
Published: (2024)
Ragas: Automated Evaluation of Retrieval Augmented Generation
by: Es, Shahul, et al.
Published: (2023)
by: Es, Shahul, et al.
Published: (2023)
Synchronous Faithfulness Monitoring for Trustworthy Retrieval-Augmented Generation
by: Wu, Di, et al.
Published: (2024)
by: Wu, Di, et al.
Published: (2024)
Benchmarking Retrieval-Augmented Generation for Medicine
by: Xiong, Guangzhi, et al.
Published: (2024)
by: Xiong, Guangzhi, et al.
Published: (2024)
RAGViz: Diagnose and Visualize Retrieval-Augmented Generation
by: Wang, Tevin, et al.
Published: (2024)
by: Wang, Tevin, et al.
Published: (2024)
Efficient Dynamic Clustering-Based Document Compression for Retrieval-Augmented-Generation
by: Li, Weitao, et al.
Published: (2025)
by: Li, Weitao, et al.
Published: (2025)
HyperRAG: Reasoning N-ary Facts over Hypergraphs for Retrieval Augmented Generation
by: Lien, Wen-Sheng, et al.
Published: (2026)
by: Lien, Wen-Sheng, et al.
Published: (2026)
Does RAG Introduce Unfairness in LLMs? Evaluating Fairness in Retrieval-Augmented Generation Systems
by: Wu, Xuyang, et al.
Published: (2024)
by: Wu, Xuyang, et al.
Published: (2024)
Retrieval-Augmented Generation for Natural Language Processing: A Survey
by: Wu, Shangyu, et al.
Published: (2024)
by: Wu, Shangyu, et al.
Published: (2024)
LRAGE: Legal Retrieval Augmented Generation Evaluation Tool
by: Park, Minhu, et al.
Published: (2025)
by: Park, Minhu, et al.
Published: (2025)
Similar Items
-
Turning Conversations into Workflows: A Framework to Extract and Evaluate Dialog Workflows for Service AI Agents
by: Choubey, Prafulla Kumar, et al.
Published: (2025) -
Do RAG Systems Cover What Matters? Evaluating and Optimizing Responses with Sub-Question Coverage
by: Xie, Kaige, et al.
Published: (2024) -
Benchmarking Deep Search over Heterogeneous Enterprise Data
by: Choubey, Prafulla Kumar, et al.
Published: (2025) -
Agentic Uncertainty Quantification
by: Zhang, Jiaxin, et al.
Published: (2026) -
Embrace Divergence for Richer Insights: A Multi-document Summarization Benchmark and a Case Study on Summarizing Diverse Information from News Articles
by: Huang, Kung-Hsiang, et al.
Published: (2023)