Real-Time Evaluation Models for RAG: Who Detects Hallucinations Best?
Fuente:
arXiv
Salvato in:
| Autore principale: | Sardana, Ashish |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The HalluRAG Dataset: Detecting Closed-Domain Hallucinations in RAG Applications Using an LLM's Internal States
di: Ridder, Fabian, et al.
Pubblicazione: (2024)
di: Ridder, Fabian, et al.
Pubblicazione: (2024)
Real-Time Detection of Hallucinated Entities in Long-Form Generation
di: Obeso, Oscar, et al.
Pubblicazione: (2025)
di: Obeso, Oscar, et al.
Pubblicazione: (2025)
Feedback-Enhanced Hallucination-Resistant Vision-Language Model for Real-Time Scene Understanding
di: Alsulaimawi, Zahir
Pubblicazione: (2025)
di: Alsulaimawi, Zahir
Pubblicazione: (2025)
The Semantic Illusion: Certified Limits of Embedding-Based Hallucination Detection in RAG Systems
di: Sinha, Debu
Pubblicazione: (2025)
di: Sinha, Debu
Pubblicazione: (2025)
Evaluating Evaluation Metrics -- The Mirage of Hallucination Detection
di: Kulkarni, Atharva, et al.
Pubblicazione: (2025)
di: Kulkarni, Atharva, et al.
Pubblicazione: (2025)
Hallucination Detection and Mitigation with Diffusion in Multi-Variate Time-Series Foundation Models
di: Wichitwechkarn, Vijja, et al.
Pubblicazione: (2025)
di: Wichitwechkarn, Vijja, et al.
Pubblicazione: (2025)
HalluGraph: Auditable Hallucination Detection for Legal RAG Systems via Knowledge Graph Alignment
di: Noël, Valentin, et al.
Pubblicazione: (2025)
di: Noël, Valentin, et al.
Pubblicazione: (2025)
Exploring the Dynamic Scheduling Space of Real-Time Generative AI Applications on Emerging Heterogeneous Systems
di: Karami, Rachid, et al.
Pubblicazione: (2025)
di: Karami, Rachid, et al.
Pubblicazione: (2025)
STEB: In Search of the Best Evaluation Approach for Synthetic Time Series
di: Stenger, Michael, et al.
Pubblicazione: (2025)
di: Stenger, Michael, et al.
Pubblicazione: (2025)
Deploying Tiny LVLM Judges for Real-World Evaluation of Chart Models: Lessons Learned and Best Practices
di: Laskar, Md Tahmid Rahman, et al.
Pubblicazione: (2025)
di: Laskar, Md Tahmid Rahman, et al.
Pubblicazione: (2025)
Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws
di: Sardana, Nikhil, et al.
Pubblicazione: (2023)
di: Sardana, Nikhil, et al.
Pubblicazione: (2023)
Ever: Mitigating Hallucination in Large Language Models through Real-Time Verification and Rectification
di: Kang, Haoqiang, et al.
Pubblicazione: (2023)
di: Kang, Haoqiang, et al.
Pubblicazione: (2023)
Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation
di: Mündler, Niels, et al.
Pubblicazione: (2023)
di: Mündler, Niels, et al.
Pubblicazione: (2023)
Sparse Upcycling: Inference Inefficient Finetuning
di: Doubov, Sasha, et al.
Pubblicazione: (2024)
di: Doubov, Sasha, et al.
Pubblicazione: (2024)
Teaming LLMs to Detect and Mitigate Hallucinations
di: Till, Demian, et al.
Pubblicazione: (2025)
di: Till, Demian, et al.
Pubblicazione: (2025)
Mixture-of-PageRanks: Replacing Long-Context with Real-Time, Sparse GraphRAG
di: Alonso, Nicholas, et al.
Pubblicazione: (2024)
di: Alonso, Nicholas, et al.
Pubblicazione: (2024)
DynamicBench: Evaluating Real-Time Report Generation in Large Language Models
di: Li, Jingyao, et al.
Pubblicazione: (2025)
di: Li, Jingyao, et al.
Pubblicazione: (2025)
An Efficient Real Time DDoS Detection Model Using Machine Learning Algorithms
di: Suvra, Debashis Kar
Pubblicazione: (2025)
di: Suvra, Debashis Kar
Pubblicazione: (2025)
Real-Time Proactive Anomaly Detection via Forward and Backward Forecast Modeling
di: Olmos, Luis, et al.
Pubblicazione: (2026)
di: Olmos, Luis, et al.
Pubblicazione: (2026)
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization
di: Tolochinsky, Elad, et al.
Pubblicazione: (2026)
di: Tolochinsky, Elad, et al.
Pubblicazione: (2026)
A Training-Time Diagnostic for Generalization via the Log-Alignment Ratio
di: Shehper, Ali, et al.
Pubblicazione: (2026)
di: Shehper, Ali, et al.
Pubblicazione: (2026)
Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment
di: Huang, Audrey, et al.
Pubblicazione: (2025)
di: Huang, Audrey, et al.
Pubblicazione: (2025)
Real-Time Machine Learning for Embedded Anomaly Detection
di: Benmachiche, Abdelmadjid, et al.
Pubblicazione: (2025)
di: Benmachiche, Abdelmadjid, et al.
Pubblicazione: (2025)
Leveraging Graph Structures to Detect Hallucinations in Large Language Models
di: Nonkes, Noa, et al.
Pubblicazione: (2024)
di: Nonkes, Noa, et al.
Pubblicazione: (2024)
The Best of Both Worlds: On the Dilemma of Out-of-distribution Detection
di: Zhang, Qingyang, et al.
Pubblicazione: (2024)
di: Zhang, Qingyang, et al.
Pubblicazione: (2024)
Don't Lag, RAG: Training-Free Adversarial Detection Using RAG
di: Kazoom, Roie, et al.
Pubblicazione: (2025)
di: Kazoom, Roie, et al.
Pubblicazione: (2025)
Semantic Energy: Detecting LLM Hallucination Beyond Entropy
di: Ma, Huan, et al.
Pubblicazione: (2025)
di: Ma, Huan, et al.
Pubblicazione: (2025)
Neural Message-Passing on Attention Graphs for Hallucination Detection
di: Frasca, Fabrizio, et al.
Pubblicazione: (2025)
di: Frasca, Fabrizio, et al.
Pubblicazione: (2025)
Detecting Hallucinations in SpeechLLMs at Inference Time Using Attention Maps
di: Waldendorf, Jonas, et al.
Pubblicazione: (2026)
di: Waldendorf, Jonas, et al.
Pubblicazione: (2026)
RADAR: Mechanistic Pathways for Detecting Data Contamination in LLM Evaluation
di: Kattamuri, Ashish, et al.
Pubblicazione: (2025)
di: Kattamuri, Ashish, et al.
Pubblicazione: (2025)
Real-Time Moving Flock Detection in Pedestrian Trajectories Using Sequential Deep Learning Models
di: Sanjjamts, Amartaivan, et al.
Pubblicazione: (2025)
di: Sanjjamts, Amartaivan, et al.
Pubblicazione: (2025)
Know Your RAG: Dataset Taxonomy and Generation Strategies for Evaluating RAG Systems
di: de Lima, Rafael Teixeira, et al.
Pubblicazione: (2024)
di: de Lima, Rafael Teixeira, et al.
Pubblicazione: (2024)
Real-Time Adaptive Anomaly Detection in Industrial IoT Environments
di: Raeiszadeh, Mahsa, et al.
Pubblicazione: (2026)
di: Raeiszadeh, Mahsa, et al.
Pubblicazione: (2026)
HalluField: Detecting LLM Hallucinations via Field-Theoretic Modeling
di: Vu, Minh, et al.
Pubblicazione: (2025)
di: Vu, Minh, et al.
Pubblicazione: (2025)
Attention Sinks as Internal Signals for Hallucination Detection in Large Language Models
di: Binkowski, Jakub, et al.
Pubblicazione: (2026)
di: Binkowski, Jakub, et al.
Pubblicazione: (2026)
Detecting and Preventing Hallucinations in Large Vision Language Models
di: Gunjal, Anisha, et al.
Pubblicazione: (2023)
di: Gunjal, Anisha, et al.
Pubblicazione: (2023)
Who Guards the Guardians? The Challenges of Evaluating Identifiability of Learned Representations
di: Joshi, Shruti, et al.
Pubblicazione: (2026)
di: Joshi, Shruti, et al.
Pubblicazione: (2026)
Robust Hallucination Detection in LLMs via Adaptive Token Selection
di: Niu, Mengjia, et al.
Pubblicazione: (2025)
di: Niu, Mengjia, et al.
Pubblicazione: (2025)
ORION Grounded in Context: Retrieval-Based Method for Hallucination Detection
di: Gerner, Assaf, et al.
Pubblicazione: (2025)
di: Gerner, Assaf, et al.
Pubblicazione: (2025)
FLaG: Fine-Grained Latent Grouping for Hallucination Detection
di: Ye, Wentao, et al.
Pubblicazione: (2026)
di: Ye, Wentao, et al.
Pubblicazione: (2026)
Documenti analoghi
-
The HalluRAG Dataset: Detecting Closed-Domain Hallucinations in RAG Applications Using an LLM's Internal States
di: Ridder, Fabian, et al.
Pubblicazione: (2024) -
Real-Time Detection of Hallucinated Entities in Long-Form Generation
di: Obeso, Oscar, et al.
Pubblicazione: (2025) -
Feedback-Enhanced Hallucination-Resistant Vision-Language Model for Real-Time Scene Understanding
di: Alsulaimawi, Zahir
Pubblicazione: (2025) -
The Semantic Illusion: Certified Limits of Embedding-Based Hallucination Detection in RAG Systems
di: Sinha, Debu
Pubblicazione: (2025) -
Evaluating Evaluation Metrics -- The Mirage of Hallucination Detection
di: Kulkarni, Atharva, et al.
Pubblicazione: (2025)