RAGalyst: Automated Human-Aligned Agentic Evaluation for Domain-Specific RAG
Fuente:
arXiv
Salvato in:
| Autori principali: | Gao, Joshua, Pham, Quoc Huy, Varghese, Subin, Saurav, Silwal, Hoskere, Vedhus |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
View-Invariant Pixelwise Anomaly Detection in Multi-object Scenes with Adaptive View Synthesis
di: Varghese, Subin, et al.
Pubblicazione: (2024)
di: Varghese, Subin, et al.
Pubblicazione: (2024)
BridgeEQA: Virtual Embodied Agents for Real Bridge Inspections
di: Varghese, Subin, et al.
Pubblicazione: (2025)
di: Varghese, Subin, et al.
Pubblicazione: (2025)
ViewDelta: Scaling Scene Change Detection through Text-Conditioning
di: Varghese, Subin, et al.
Pubblicazione: (2024)
di: Varghese, Subin, et al.
Pubblicazione: (2024)
Multiclass Post-Earthquake Building Assessment Integrating High-Resolution Optical and SAR Satellite Imagery, Ground Motion, and Soil Data with Transformers
di: Singh, Deepank, et al.
Pubblicazione: (2024)
di: Singh, Deepank, et al.
Pubblicazione: (2024)
Domain-Specific Data Generation Framework for RAG Adaptation
di: Tian, Chris Xing, et al.
Pubblicazione: (2025)
di: Tian, Chris Xing, et al.
Pubblicazione: (2025)
RAFT: Adapting Language Model to Domain Specific RAG
di: Zhang, Tianjun, et al.
Pubblicazione: (2024)
di: Zhang, Tianjun, et al.
Pubblicazione: (2024)
RAVine: Reality-Aligned Evaluation for Agentic Search
di: Xu, Yilong, et al.
Pubblicazione: (2025)
di: Xu, Yilong, et al.
Pubblicazione: (2025)
GraphRAG-Bench: Challenging Domain-Specific Reasoning for Evaluating Graph Retrieval-Augmented Generation
di: Xiao, Yilin, et al.
Pubblicazione: (2025)
di: Xiao, Yilin, et al.
Pubblicazione: (2025)
Transparent Reference-free Automated Evaluation of Open-Ended User Survey Responses
di: An, Subin, et al.
Pubblicazione: (2025)
di: An, Subin, et al.
Pubblicazione: (2025)
From RAG to Agentic RAG for Faithful Islamic Question Answering
di: Bhatia, Gagan, et al.
Pubblicazione: (2026)
di: Bhatia, Gagan, et al.
Pubblicazione: (2026)
Towards Efficient Large Language Models for Scientific Text: A Review
di: To, Huy Quoc, et al.
Pubblicazione: (2024)
di: To, Huy Quoc, et al.
Pubblicazione: (2024)
Chain-of-Rank: Enhancing Large Language Models for Domain-Specific RAG in Edge Device
di: Lee, Juntae, et al.
Pubblicazione: (2025)
di: Lee, Juntae, et al.
Pubblicazione: (2025)
Agentic Adversarial QA for Improving Domain-Specific LLMs
di: Grari, Vincent, et al.
Pubblicazione: (2026)
di: Grari, Vincent, et al.
Pubblicazione: (2026)
OASES: Outcome-Aligned Search-Evaluation Co-Training for Agentic Search
di: Zhang, Erhan, et al.
Pubblicazione: (2026)
di: Zhang, Erhan, et al.
Pubblicazione: (2026)
LLMs Judge Themselves: A Game-Theoretic Framework for Human-Aligned Evaluation
di: Yang, Gao, et al.
Pubblicazione: (2025)
di: Yang, Gao, et al.
Pubblicazione: (2025)
Automated Benchmark Generation from Domain Guidelines Informed by Bloom's Taxonomy
di: Chen, Si, et al.
Pubblicazione: (2026)
di: Chen, Si, et al.
Pubblicazione: (2026)
Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents
di: Yehudai, Asaf, et al.
Pubblicazione: (2026)
di: Yehudai, Asaf, et al.
Pubblicazione: (2026)
Towards Agentic RAG with Deep Reasoning: A Survey of RAG-Reasoning Systems in LLMs
di: Li, Yangning, et al.
Pubblicazione: (2025)
di: Li, Yangning, et al.
Pubblicazione: (2025)
Towards AI Evaluation in Domain-Specific RAG Systems: The AgriHubi Case Study
di: Hasan, Md. Toufique, et al.
Pubblicazione: (2026)
di: Hasan, Md. Toufique, et al.
Pubblicazione: (2026)
DO-RAG: A Domain-Specific QA Framework Using Knowledge Graph-Enhanced Retrieval-Augmented Generation
di: Opoku, David Osei, et al.
Pubblicazione: (2025)
di: Opoku, David Osei, et al.
Pubblicazione: (2025)
Evaluating ChatGPT on Nuclear Domain-Specific Data
di: Anwar, Muhammad, et al.
Pubblicazione: (2024)
di: Anwar, Muhammad, et al.
Pubblicazione: (2024)
CORAL: Adaptive Retrieval Loop for Culturally-Aligned Multilingual RAG
di: Lee, Nayeon, et al.
Pubblicazione: (2026)
di: Lee, Nayeon, et al.
Pubblicazione: (2026)
Verus-SpecGym: An Agentic Environment for Evaluating Specification Autoformalization
di: Agarwal, Anmol, et al.
Pubblicazione: (2026)
di: Agarwal, Anmol, et al.
Pubblicazione: (2026)
Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG
di: Singh, Aditi, et al.
Pubblicazione: (2025)
di: Singh, Aditi, et al.
Pubblicazione: (2025)
ASTRID -- An Automated and Scalable TRIaD for the Evaluation of RAG-based Clinical Question Answering Systems
di: Chowdhury, Mohita, et al.
Pubblicazione: (2025)
di: Chowdhury, Mohita, et al.
Pubblicazione: (2025)
Efficient and Transferable Agentic Knowledge Graph RAG via Reinforcement Learning
di: Lin, Junhong, et al.
Pubblicazione: (2025)
di: Lin, Junhong, et al.
Pubblicazione: (2025)
Reinforcement Learning for Optimizing RAG for Domain Chatbots
di: Kulkarni, Mandar, et al.
Pubblicazione: (2024)
di: Kulkarni, Mandar, et al.
Pubblicazione: (2024)
RAG-QA Arena: Evaluating Domain Robustness for Long-form Retrieval Augmented Question Answering
di: Han, Rujun, et al.
Pubblicazione: (2024)
di: Han, Rujun, et al.
Pubblicazione: (2024)
LalaEval: A Holistic Human Evaluation Framework for Domain-Specific Large Language Models
di: Sun, Chongyan, et al.
Pubblicazione: (2024)
di: Sun, Chongyan, et al.
Pubblicazione: (2024)
JADE: Bridging the Strategic-Operational Gap in Dynamic Agentic RAG
di: Chen, Yiqun, et al.
Pubblicazione: (2026)
di: Chen, Yiqun, et al.
Pubblicazione: (2026)
From Guidelines to Guarantees: A Graph-Based Evaluation Harness for Domain-Specific Evaluation of LLMs
di: Lundin, Jessica M., et al.
Pubblicazione: (2025)
di: Lundin, Jessica M., et al.
Pubblicazione: (2025)
Vision-Based Adaptive Robotics for Autonomous Surface Crack Repair
di: Genova, Joshua, et al.
Pubblicazione: (2024)
di: Genova, Joshua, et al.
Pubblicazione: (2024)
FormalAlign: Automated Alignment Evaluation for Autoformalization
di: Lu, Jianqiao, et al.
Pubblicazione: (2024)
di: Lu, Jianqiao, et al.
Pubblicazione: (2024)
Evaluating Causal Explanation in Medical Reports with LLM-Based and Human-Aligned Metrics
di: Cho, Yousang, et al.
Pubblicazione: (2025)
di: Cho, Yousang, et al.
Pubblicazione: (2025)
SLMEval: Entropy-Based Calibration for Human-Aligned Evaluation of Large Language Models
di: Daynauth, Roland, et al.
Pubblicazione: (2025)
di: Daynauth, Roland, et al.
Pubblicazione: (2025)
BoRP: Bootstrapped Regression Probing for Scalable and Human-Aligned LLM Evaluation
di: Sun, Peng, et al.
Pubblicazione: (2026)
di: Sun, Peng, et al.
Pubblicazione: (2026)
Toward Subtrait-Level Model Explainability in Automated Writing Evaluation
di: Andrade-Lotero, Alejandro, et al.
Pubblicazione: (2025)
di: Andrade-Lotero, Alejandro, et al.
Pubblicazione: (2025)
Less is More for RAG: Information Gain Pruning for Generator-Aligned Reranking and Evidence Selection
di: Song, Zhipeng, et al.
Pubblicazione: (2026)
di: Song, Zhipeng, et al.
Pubblicazione: (2026)
Hallucination-Resistant, Domain-Specific Research Assistant with Self-Evaluation and Vector-Grounded Retrieval
di: Bhavsar, Vivek, et al.
Pubblicazione: (2025)
di: Bhavsar, Vivek, et al.
Pubblicazione: (2025)
Anomaly Detection in Human Language via Meta-Learning: A Few-Shot Approach
di: Singla, Saurav, et al.
Pubblicazione: (2025)
di: Singla, Saurav, et al.
Pubblicazione: (2025)
Documenti analoghi
-
View-Invariant Pixelwise Anomaly Detection in Multi-object Scenes with Adaptive View Synthesis
di: Varghese, Subin, et al.
Pubblicazione: (2024) -
BridgeEQA: Virtual Embodied Agents for Real Bridge Inspections
di: Varghese, Subin, et al.
Pubblicazione: (2025) -
ViewDelta: Scaling Scene Change Detection through Text-Conditioning
di: Varghese, Subin, et al.
Pubblicazione: (2024) -
Multiclass Post-Earthquake Building Assessment Integrating High-Resolution Optical and SAR Satellite Imagery, Ground Motion, and Soil Data with Transformers
di: Singh, Deepank, et al.
Pubblicazione: (2024) -
Domain-Specific Data Generation Framework for RAG Adaptation
di: Tian, Chris Xing, et al.
Pubblicazione: (2025)