Towards Lighter and Robust Evaluation for Retrieval Augmented Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Ispas, Alex-Razvan, Simon, Charles-Elie, Caspani, Fabien, Guigue, Vincent |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Multi-Task, Multi-Modal Approach for Predicting Categorical and Dimensional Emotions
di: Ispas, Alex-Răzvan, et al.
Pubblicazione: (2023)
di: Ispas, Alex-Răzvan, et al.
Pubblicazione: (2023)
Enhancing Long-Term Memory using Hierarchical Aggregate Tree for Retrieval Augmented Generation
di: A, Aadharsh Aadhithya, et al.
Pubblicazione: (2024)
di: A, Aadharsh Aadhithya, et al.
Pubblicazione: (2024)
Detecting Hallucinations in Graph Retrieval-Augmented Generation via Attention Patterns and Semantic Alignment
di: Li, Shanghao, et al.
Pubblicazione: (2025)
di: Li, Shanghao, et al.
Pubblicazione: (2025)
Considering Length Diversity in Retrieval-Augmented Summarization
di: Juseon-Do, et al.
Pubblicazione: (2025)
di: Juseon-Do, et al.
Pubblicazione: (2025)
Hallucination or Creativity: How to Evaluate AI-Generated Scientific Stories?
di: Argese, Alex, et al.
Pubblicazione: (2026)
di: Argese, Alex, et al.
Pubblicazione: (2026)
Accurate and Energy Efficient: Local Retrieval-Augmented Generation Models Outperform Commercial Large Language Models in Medical Tasks
di: Vrettos, Konstantinos, et al.
Pubblicazione: (2025)
di: Vrettos, Konstantinos, et al.
Pubblicazione: (2025)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
di: Oketunji, Abiodun Finbarrs
Pubblicazione: (2023)
di: Oketunji, Abiodun Finbarrs
Pubblicazione: (2023)
NeuralNexus at BEA 2025 Shared Task: Retrieval-Augmented Prompting for Mistake Identification in AI Tutors
di: Naeem, Numaan, et al.
Pubblicazione: (2025)
di: Naeem, Numaan, et al.
Pubblicazione: (2025)
Statistical Measures for Explainable Aspect-Based Sentiment Analysis: A Case Study on Environmental Discourse in Reddit
di: Stracqualursi, Luisa, et al.
Pubblicazione: (2026)
di: Stracqualursi, Luisa, et al.
Pubblicazione: (2026)
A Library of LLM Intrinsics for Retrieval-Augmented Generation
di: Danilevsky, Marina, et al.
Pubblicazione: (2025)
di: Danilevsky, Marina, et al.
Pubblicazione: (2025)
TALE: A Tool-Augmented Framework for Reference-Free Evaluation of Large Language Models
di: Badshah, Sher, et al.
Pubblicazione: (2025)
di: Badshah, Sher, et al.
Pubblicazione: (2025)
ConQRet: Benchmarking Fine-Grained Evaluation of Retrieval Augmented Argumentation with LLM Judges
di: Dhole, Kaustubh D., et al.
Pubblicazione: (2024)
di: Dhole, Kaustubh D., et al.
Pubblicazione: (2024)
PARAPHRASUS : A Comprehensive Benchmark for Evaluating Paraphrase Detection Models
di: Michail, Andrianos, et al.
Pubblicazione: (2024)
di: Michail, Andrianos, et al.
Pubblicazione: (2024)
MedPI: Evaluating AI Systems in Medical Patient-facing Interactions
di: V., Diego Fajardo, et al.
Pubblicazione: (2025)
di: V., Diego Fajardo, et al.
Pubblicazione: (2025)
PlanRAG: A Plan-then-Retrieval Augmented Generation for Generative Large Language Models as Decision Makers
di: Lee, Myeonghwa, et al.
Pubblicazione: (2024)
di: Lee, Myeonghwa, et al.
Pubblicazione: (2024)
Knowledge-Augmented Multimodal Clinical Rationale Generation for Disease Diagnosis with Small Language Models
di: Niu, Shuai, et al.
Pubblicazione: (2024)
di: Niu, Shuai, et al.
Pubblicazione: (2024)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
di: Saji, Alan, et al.
Pubblicazione: (2025)
di: Saji, Alan, et al.
Pubblicazione: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
di: Peters, Sydney, et al.
Pubblicazione: (2025)
di: Peters, Sydney, et al.
Pubblicazione: (2025)
SPRInG: Continual LLM Personalization via Selective Parametric Adaptation and Retrieval-Interpolated Generation
di: Kim, Seoyeon, et al.
Pubblicazione: (2026)
di: Kim, Seoyeon, et al.
Pubblicazione: (2026)
Spotlights and Blindspots: Evaluating Machine-Generated Text Detection
di: Stowe, Kevin, et al.
Pubblicazione: (2026)
di: Stowe, Kevin, et al.
Pubblicazione: (2026)
Agentic Retrieval-Augmented Generation for Financial Document Question Answering
di: Shu, Yang, et al.
Pubblicazione: (2026)
di: Shu, Yang, et al.
Pubblicazione: (2026)
RAC: Efficient LLM Factuality Correction with Retrieval Augmentation
di: Li, Changmao, et al.
Pubblicazione: (2024)
di: Li, Changmao, et al.
Pubblicazione: (2024)
FIN-bench-v2: A Unified and Robust Benchmark Suite for Evaluating Finnish Large Language Models
di: Kytöniemi, Joona, et al.
Pubblicazione: (2025)
di: Kytöniemi, Joona, et al.
Pubblicazione: (2025)
Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models
di: Hawkins, John
Pubblicazione: (2025)
di: Hawkins, John
Pubblicazione: (2025)
A Comparative Analysis of Retrieval-Augmented Generation Techniques for Bengali Standard-to-Dialect Machine Translation Using LLMs
di: Sami, K. M. Jubair, et al.
Pubblicazione: (2025)
di: Sami, K. M. Jubair, et al.
Pubblicazione: (2025)
Generative Active Testing: Efficient LLM Evaluation via Proxy Task Adaptation
di: Ramakrishnan, Aashish Anantha, et al.
Pubblicazione: (2026)
di: Ramakrishnan, Aashish Anantha, et al.
Pubblicazione: (2026)
Cetvel: A Unified Benchmark for Evaluating Language Understanding, Generation and Cultural Capacity of LLMs for Turkish
di: Er, Yakup Abrek, et al.
Pubblicazione: (2025)
di: Er, Yakup Abrek, et al.
Pubblicazione: (2025)
A methodological analysis of prompt perturbations and their effect on attack success rates
di: Machado, Tiago, et al.
Pubblicazione: (2025)
di: Machado, Tiago, et al.
Pubblicazione: (2025)
VERA: Validation and Evaluation of Retrieval-Augmented Systems
di: Ding, Tianyu, et al.
Pubblicazione: (2024)
di: Ding, Tianyu, et al.
Pubblicazione: (2024)
A Comparative Study of DSL Code Generation: Fine-Tuning vs. Optimized Retrieval Augmentation
di: Bassamzadeh, Nastaran, et al.
Pubblicazione: (2024)
di: Bassamzadeh, Nastaran, et al.
Pubblicazione: (2024)
Learning vs Retrieval: The Role of In-Context Examples in Regression with Large Language Models
di: Nafar, Aliakbar, et al.
Pubblicazione: (2024)
di: Nafar, Aliakbar, et al.
Pubblicazione: (2024)
Robustness of Large Language Models to Perturbations in Text
di: Singh, Ayush, et al.
Pubblicazione: (2024)
di: Singh, Ayush, et al.
Pubblicazione: (2024)
TAGS: A Test-Time Generalist-Specialist Framework with Retrieval-Augmented Reasoning and Verification
di: Wu, Jianghao, et al.
Pubblicazione: (2025)
di: Wu, Jianghao, et al.
Pubblicazione: (2025)
Neural Machine Translation for Malayalam Paraphrase Generation
di: Varghese, Christeena, et al.
Pubblicazione: (2024)
di: Varghese, Christeena, et al.
Pubblicazione: (2024)
UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning
di: Ovcharov, Volodymyr
Pubblicazione: (2026)
di: Ovcharov, Volodymyr
Pubblicazione: (2026)
Towards Safer Chatbots: Automated Policy Compliance Evaluation of Custom GPTs
di: Rodriguez, David, et al.
Pubblicazione: (2025)
di: Rodriguez, David, et al.
Pubblicazione: (2025)
Automated MCQA Benchmarking at Scale: Evaluating Reasoning Traces as Retrieval Sources for Domain Adaptation of Small Language Models
di: Gokdemir, Ozan, et al.
Pubblicazione: (2025)
di: Gokdemir, Ozan, et al.
Pubblicazione: (2025)
RADD: Retrieval-Augmented Discrete Diffusion for Multi-Modal Knowledge Graph Completion
di: Niu, Guanglin, et al.
Pubblicazione: (2026)
di: Niu, Guanglin, et al.
Pubblicazione: (2026)
To Retrieve or Not to Retrieve? Uncertainty Detection for Dynamic Retrieval Augmented Generation
di: Dhole, Kaustubh D.
Pubblicazione: (2025)
di: Dhole, Kaustubh D.
Pubblicazione: (2025)
Partially Recentralization Softmax Loss for Vision-Language Models Robustness
di: Wang, Hao, et al.
Pubblicazione: (2024)
di: Wang, Hao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
A Multi-Task, Multi-Modal Approach for Predicting Categorical and Dimensional Emotions
di: Ispas, Alex-Răzvan, et al.
Pubblicazione: (2023) -
Enhancing Long-Term Memory using Hierarchical Aggregate Tree for Retrieval Augmented Generation
di: A, Aadharsh Aadhithya, et al.
Pubblicazione: (2024) -
Detecting Hallucinations in Graph Retrieval-Augmented Generation via Attention Patterns and Semantic Alignment
di: Li, Shanghao, et al.
Pubblicazione: (2025) -
Considering Length Diversity in Retrieval-Augmented Summarization
di: Juseon-Do, et al.
Pubblicazione: (2025) -
Hallucination or Creativity: How to Evaluate AI-Generated Scientific Stories?
di: Argese, Alex, et al.
Pubblicazione: (2026)