Intrinsic Evaluation of RAG Systems for Deep-Logic Questions
Fuente:
arXiv
Salvato in:
| Autori principali: | Hu, Junyi, Zhou, You, Wang, Jie |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Detecting AI-Generated Texts in Cross-Domains
di: Zhou, You, et al.
Pubblicazione: (2024)
di: Zhou, You, et al.
Pubblicazione: (2024)
Constructing Cloze Questions Generatively
di: Sun, Yicheng, et al.
Pubblicazione: (2024)
di: Sun, Yicheng, et al.
Pubblicazione: (2024)
TrafficRAG: A Multimodal RAG Framework for Traffic Accident Liability Determination
di: Li, Xu, et al.
Pubblicazione: (2026)
di: Li, Xu, et al.
Pubblicazione: (2026)
A Library of LLM Intrinsics for Retrieval-Augmented Generation
di: Danilevsky, Marina, et al.
Pubblicazione: (2025)
di: Danilevsky, Marina, et al.
Pubblicazione: (2025)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
di: Oketunji, Abiodun Finbarrs
Pubblicazione: (2023)
di: Oketunji, Abiodun Finbarrs
Pubblicazione: (2023)
Comparing the Performance of LLMs in RAG-based Question-Answering: A Case Study in Computer Science Literature
di: Dayarathne, Ranul, et al.
Pubblicazione: (2025)
di: Dayarathne, Ranul, et al.
Pubblicazione: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
di: Peters, Sydney, et al.
Pubblicazione: (2025)
di: Peters, Sydney, et al.
Pubblicazione: (2025)
A Fuzzy Logic Prompting Framework for Large Language Models in Adaptive and Uncertain Tasks
di: Figueiredo, Vanessa
Pubblicazione: (2025)
di: Figueiredo, Vanessa
Pubblicazione: (2025)
Pharos-ESG: A Framework for Multimodal Parsing, Contextual Narration, and Hierarchical Labeling of ESG Report
di: Chen, Yan, et al.
Pubblicazione: (2025)
di: Chen, Yan, et al.
Pubblicazione: (2025)
FATHOMS-RAG: A Framework for the Assessment of Thinking and Observation in Multimodal Systems that use Retrieval Augmented Generation
di: Hildebrand, Samuel, et al.
Pubblicazione: (2025)
di: Hildebrand, Samuel, et al.
Pubblicazione: (2025)
The Paradox of Robustness: Decoupling Rule-Based Logic from Affective Noise in High-Stakes Decision-Making
di: Chun, Jon, et al.
Pubblicazione: (2026)
di: Chun, Jon, et al.
Pubblicazione: (2026)
Understanding LLM Evaluator Behavior: A Structured Multi-Evaluator Framework for Merchant Risk Assessment
di: Wang, Liang, et al.
Pubblicazione: (2026)
di: Wang, Liang, et al.
Pubblicazione: (2026)
Question Answering Over Spatio-Temporal Knowledge Graph
di: Dai, Xinbang, et al.
Pubblicazione: (2024)
di: Dai, Xinbang, et al.
Pubblicazione: (2024)
FlexStructRAG: Flexible Structure-Aware Multi-Granular Relational Retrieval for RAG
di: Chen, Mengzhu, et al.
Pubblicazione: (2026)
di: Chen, Mengzhu, et al.
Pubblicazione: (2026)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
di: Saji, Alan, et al.
Pubblicazione: (2025)
di: Saji, Alan, et al.
Pubblicazione: (2025)
Evaluating Relational Reasoning in LLMs with REL
di: Fesser, Lukas, et al.
Pubblicazione: (2026)
di: Fesser, Lukas, et al.
Pubblicazione: (2026)
Evaluating Steering Techniques using Human Similarity Judgments
di: Studdiford, Zach, et al.
Pubblicazione: (2025)
di: Studdiford, Zach, et al.
Pubblicazione: (2025)
Evaluating LLM Metrics Through Real-World Capabilities
di: Miller, Justin K, et al.
Pubblicazione: (2025)
di: Miller, Justin K, et al.
Pubblicazione: (2025)
Temporal Knowledge Question Answering via Abstract Reasoning Induction
di: Chen, Ziyang, et al.
Pubblicazione: (2023)
di: Chen, Ziyang, et al.
Pubblicazione: (2023)
RoleRAG: Enhancing LLM Role-Playing via Graph Guided Retrieval
di: Wang, Yongjie, et al.
Pubblicazione: (2025)
di: Wang, Yongjie, et al.
Pubblicazione: (2025)
Graph Guided Question Answer Generation for Procedural Question-Answering
di: Pham, Hai X., et al.
Pubblicazione: (2024)
di: Pham, Hai X., et al.
Pubblicazione: (2024)
EduAdapt: A Question Answer Benchmark Dataset for Evaluating Grade-Level Adaptability in LLMs
di: Naeem, Numaan, et al.
Pubblicazione: (2025)
di: Naeem, Numaan, et al.
Pubblicazione: (2025)
Fine-Tuned Large Language Models for Logical Translation: Reducing Hallucinations with Lang2Logic
di: Pan, Muyu, et al.
Pubblicazione: (2025)
di: Pan, Muyu, et al.
Pubblicazione: (2025)
Reasoning-Based AI for Startup Evaluation (R.A.I.S.E.): A Memory-Augmented, Multi-Step Decision Framework
di: Preuveneers, Jack, et al.
Pubblicazione: (2025)
di: Preuveneers, Jack, et al.
Pubblicazione: (2025)
Teaching Probabilistic Logical Reasoning to Transformers
di: Nafar, Aliakbar, et al.
Pubblicazione: (2023)
di: Nafar, Aliakbar, et al.
Pubblicazione: (2023)
Do We Always Need Query-Level Workflows? Rethinking Agentic Workflow Generation for Multi-Agent Systems
di: Wang, Zixu, et al.
Pubblicazione: (2026)
di: Wang, Zixu, et al.
Pubblicazione: (2026)
DQA: Diagnostic Question Answering for IT Support
di: Kapoor, Vishaal, et al.
Pubblicazione: (2026)
di: Kapoor, Vishaal, et al.
Pubblicazione: (2026)
A Graph-based RAG for Energy Efficiency Question Answering
di: Campi, Riccardo, et al.
Pubblicazione: (2025)
di: Campi, Riccardo, et al.
Pubblicazione: (2025)
Observations on Building RAG Systems for Technical Documents
di: Soman, Sumit, et al.
Pubblicazione: (2024)
di: Soman, Sumit, et al.
Pubblicazione: (2024)
Low-Resource Court Judgment Summarization for Common Law Systems
di: Liu, Shuaiqi, et al.
Pubblicazione: (2024)
di: Liu, Shuaiqi, et al.
Pubblicazione: (2024)
Exploring Graph Representations of Logical Forms for Language Modeling
di: Sullivan, Michael
Pubblicazione: (2025)
di: Sullivan, Michael
Pubblicazione: (2025)
Xinyu: An Efficient LLM-based System for Commentary Generation
di: Wu, Yiquan, et al.
Pubblicazione: (2024)
di: Wu, Yiquan, et al.
Pubblicazione: (2024)
VERA: Validation and Evaluation of Retrieval-Augmented Systems
di: Ding, Tianyu, et al.
Pubblicazione: (2024)
di: Ding, Tianyu, et al.
Pubblicazione: (2024)
ROZA Graphs: Self-Improving Near-Deterministic RAG through Evidence-Centric Feedback
di: Penaroza, Matthew
Pubblicazione: (2026)
di: Penaroza, Matthew
Pubblicazione: (2026)
Bidirectional RAG: Safe Self-Improving Retrieval-Augmented Generation Through Multi-Stage Validation
di: Chinthala, Teja
Pubblicazione: (2025)
di: Chinthala, Teja
Pubblicazione: (2025)
Introducing Brain-like Concepts to Embodied Hand-crafted Dialog Management System
di: Joublin, Frank, et al.
Pubblicazione: (2024)
di: Joublin, Frank, et al.
Pubblicazione: (2024)
CoE: Collaborative Entropy for Uncertainty Quantification in Agentic Multi-LLM Systems
di: Sun, Kangkang, et al.
Pubblicazione: (2026)
di: Sun, Kangkang, et al.
Pubblicazione: (2026)
AskSport: Web Application for Sports Question-Answering
di: Onofre, Enzo B, et al.
Pubblicazione: (2025)
di: Onofre, Enzo B, et al.
Pubblicazione: (2025)
How LLMs Are Persuaded: A Few Attention Heads, Rerouted
di: Sun, Xiangkun, et al.
Pubblicazione: (2026)
di: Sun, Xiangkun, et al.
Pubblicazione: (2026)
From PDF to RAG-Ready: Evaluating Document Conversion Frameworks for Domain-Specific Question Answering
di: Santos, José Guilherme Marques dos, et al.
Pubblicazione: (2026)
di: Santos, José Guilherme Marques dos, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Detecting AI-Generated Texts in Cross-Domains
di: Zhou, You, et al.
Pubblicazione: (2024) -
Constructing Cloze Questions Generatively
di: Sun, Yicheng, et al.
Pubblicazione: (2024) -
TrafficRAG: A Multimodal RAG Framework for Traffic Accident Liability Determination
di: Li, Xu, et al.
Pubblicazione: (2026) -
A Library of LLM Intrinsics for Retrieval-Augmented Generation
di: Danilevsky, Marina, et al.
Pubblicazione: (2025) -
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
di: Oketunji, Abiodun Finbarrs
Pubblicazione: (2023)