Truth or Mirage? Towards End-to-End Factuality Evaluation with LLM-Oasis
Fuente:
arXiv
Salvato in:
| Autori principali: | Scirè, Alessandro, Bejgu, Andrei Stefan, Tedeschi, Simone, Ghonim, Karim, Martelli, Federico, Navigli, Roberto |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FENICE: Factuality Evaluation of summarization based on Natural language Inference and Claim Extraction
di: Scirè, Alessandro, et al.
Pubblicazione: (2024)
di: Scirè, Alessandro, et al.
Pubblicazione: (2024)
Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering
di: Molfese, Francesco Maria, et al.
Pubblicazione: (2025)
di: Molfese, Francesco Maria, et al.
Pubblicazione: (2025)
Word Sense Linking: Disambiguating Outside the Sandbox
di: Bejgu, Andrei Stefan, et al.
Pubblicazione: (2024)
di: Bejgu, Andrei Stefan, et al.
Pubblicazione: (2024)
Guardians of the Machine Translation Meta-Evaluation: Sentinel Metrics Fall In!
di: Perrella, Stefano, et al.
Pubblicazione: (2024)
di: Perrella, Stefano, et al.
Pubblicazione: (2024)
Do Large Language Models Understand Word Senses?
di: Meconi, Domenico, et al.
Pubblicazione: (2025)
di: Meconi, Domenico, et al.
Pubblicazione: (2025)
Optimizing LLMs for Italian: Reducing Token Fertility and Enhancing Efficiency Through Vocabulary Adaptation
di: Moroni, Luca, et al.
Pubblicazione: (2025)
di: Moroni, Luca, et al.
Pubblicazione: (2025)
LiteraryQA: Towards Effective Evaluation of Long-document Narrative QA
di: Bonomo, Tommaso, et al.
Pubblicazione: (2025)
di: Bonomo, Tommaso, et al.
Pubblicazione: (2025)
Detecting Errors through Ensembling Prompts (DEEP): An End-to-End LLM Framework for Detecting Factual Errors
di: Chandler, Alex, et al.
Pubblicazione: (2024)
di: Chandler, Alex, et al.
Pubblicazione: (2024)
LLMs Lost in Translation: M-ALERT uncovers Cross-Linguistic Safety Inconsistencies
di: Friedrich, Felix, et al.
Pubblicazione: (2024)
di: Friedrich, Felix, et al.
Pubblicazione: (2024)
ALERT: A Comprehensive Benchmark for Assessing Large Language Models' Safety through Red Teaming
di: Tedeschi, Simone, et al.
Pubblicazione: (2024)
di: Tedeschi, Simone, et al.
Pubblicazione: (2024)
Process Reward Models Meet Planning: Generating Precise and Scalable Datasets for Step-Level Rewards
di: Pisano, Raffaele, et al.
Pubblicazione: (2026)
di: Pisano, Raffaele, et al.
Pubblicazione: (2026)
Interpretable Coreference Resolution Evaluation Using Explicit Semantics
di: Gatti, Bruno, et al.
Pubblicazione: (2026)
di: Gatti, Bruno, et al.
Pubblicazione: (2026)
ReTraceQA: Evaluating Reasoning Traces of Small Language Models in Commonsense Question Answering
di: Molfese, Francesco Maria, et al.
Pubblicazione: (2025)
di: Molfese, Francesco Maria, et al.
Pubblicazione: (2025)
Has Machine Translation Evaluation Achieved Human Parity? The Human Reference and the Limits of Progress
di: Proietti, Lorenzo, et al.
Pubblicazione: (2025)
di: Proietti, Lorenzo, et al.
Pubblicazione: (2025)
DRIFT: Detecting Representational Inconsistencies for Factual Truthfulness
di: Bhatnagar, Rohan, et al.
Pubblicazione: (2026)
di: Bhatnagar, Rohan, et al.
Pubblicazione: (2026)
ZEBRA: Zero-Shot Example-Based Retrieval Augmentation for Commonsense Question Answering
di: Molfese, Francesco Maria, et al.
Pubblicazione: (2024)
di: Molfese, Francesco Maria, et al.
Pubblicazione: (2024)
The End of Manual Decoding: Towards Truly End-to-End Language Models
di: Wang, Zhichao, et al.
Pubblicazione: (2025)
di: Wang, Zhichao, et al.
Pubblicazione: (2025)
URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models
di: Yan, Ruiqi, et al.
Pubblicazione: (2025)
di: Yan, Ruiqi, et al.
Pubblicazione: (2025)
MEMERAG: A Multilingual End-to-End Meta-Evaluation Benchmark for Retrieval Augmented Generation
di: Blandón, María Andrea Cruz, et al.
Pubblicazione: (2025)
di: Blandón, María Andrea Cruz, et al.
Pubblicazione: (2025)
End-to-End Evaluation for Low-Latency Simultaneous Speech Translation
di: Huber, Christian, et al.
Pubblicazione: (2023)
di: Huber, Christian, et al.
Pubblicazione: (2023)
RAG-Zeval: Towards Robust and Interpretable Evaluation on RAG Responses through End-to-End Rule-Guided Reasoning
di: Li, Kun, et al.
Pubblicazione: (2025)
di: Li, Kun, et al.
Pubblicazione: (2025)
AutoML-guided Fusion of Entity and LLM-based Representations for Document Classification
di: Koloski, Boshko, et al.
Pubblicazione: (2024)
di: Koloski, Boshko, et al.
Pubblicazione: (2024)
End-to-End Chatbot Evaluation with Adaptive Reasoning and Uncertainty Filtering
di: Dang, Nhi, et al.
Pubblicazione: (2026)
di: Dang, Nhi, et al.
Pubblicazione: (2026)
City-LEO: Toward Transparent City Management Using LLM with End-to-End Optimization
di: Jiao, Zihao, et al.
Pubblicazione: (2024)
di: Jiao, Zihao, et al.
Pubblicazione: (2024)
Towards Effective Extraction and Evaluation of Factual Claims
di: Metropolitansky, Dasha, et al.
Pubblicazione: (2025)
di: Metropolitansky, Dasha, et al.
Pubblicazione: (2025)
The Mirage of Model Editing: Revisiting Evaluation in the Wild
di: Yang, Wanli, et al.
Pubblicazione: (2025)
di: Yang, Wanli, et al.
Pubblicazione: (2025)
ChatCFD: An LLM-Driven Agent for End-to-End CFD Automation with Structured Knowledge and Reasoning
di: Fan, E, et al.
Pubblicazione: (2025)
di: Fan, E, et al.
Pubblicazione: (2025)
Maverick: Efficient and Accurate Coreference Resolution Defying Recent Trends
di: Martinelli, Giuliano, et al.
Pubblicazione: (2024)
di: Martinelli, Giuliano, et al.
Pubblicazione: (2024)
THaMES: An End-to-End Tool for Hallucination Mitigation and Evaluation in Large Language Models
di: Liang, Mengfei, et al.
Pubblicazione: (2024)
di: Liang, Mengfei, et al.
Pubblicazione: (2024)
EvoScientist: Towards Multi-Agent Evolving AI Scientists for End-to-End Scientific Discovery
di: Lyu, Yougang, et al.
Pubblicazione: (2026)
di: Lyu, Yougang, et al.
Pubblicazione: (2026)
Chasing Progress, Not Perfection: Revisiting Strategies for End-to-End LLM Plan Generation
di: Huang, Sukai, et al.
Pubblicazione: (2024)
di: Huang, Sukai, et al.
Pubblicazione: (2024)
VeriLocc: End-to-End Cross-Architecture Register Allocation via LLM
di: Jin, Lesheng, et al.
Pubblicazione: (2025)
di: Jin, Lesheng, et al.
Pubblicazione: (2025)
Towards End-to-End Spoken Grammatical Error Correction
di: Bannò, Stefano, et al.
Pubblicazione: (2023)
di: Bannò, Stefano, et al.
Pubblicazione: (2023)
TruthEval: A Dataset to Evaluate LLM Truthfulness and Reliability
di: Khatun, Aisha, et al.
Pubblicazione: (2024)
di: Khatun, Aisha, et al.
Pubblicazione: (2024)
LOGOS: LLM-driven End-to-End Grounded Theory Development and Schema Induction for Qualitative Research
di: Pi, Xinyu, et al.
Pubblicazione: (2025)
di: Pi, Xinyu, et al.
Pubblicazione: (2025)
Towards End-to-End Open Conversational Machine Reading
di: Zhou, Sizhe, et al.
Pubblicazione: (2022)
di: Zhou, Sizhe, et al.
Pubblicazione: (2022)
Grounded Visual Factualization: Factual Anchor-Based Finetuning for Enhancing MLLM Factual Consistency
di: Morbiato, Filippo, et al.
Pubblicazione: (2025)
di: Morbiato, Filippo, et al.
Pubblicazione: (2025)
Hallucination to Truth: A Review of Fact-Checking and Factuality Evaluation in Large Language Models
di: Rahman, Subhey Sadi, et al.
Pubblicazione: (2025)
di: Rahman, Subhey Sadi, et al.
Pubblicazione: (2025)
FuDoBa: Fusing Document and Knowledge Graph-based Representations with Bayesian Optimisation
di: Koloski, Boshko, et al.
Pubblicazione: (2025)
di: Koloski, Boshko, et al.
Pubblicazione: (2025)
Towards Building an End-to-End Multilingual Automatic Lyrics Transcription Model
di: Huang, Jiawen, et al.
Pubblicazione: (2024)
di: Huang, Jiawen, et al.
Pubblicazione: (2024)
Documenti analoghi
-
FENICE: Factuality Evaluation of summarization based on Natural language Inference and Claim Extraction
di: Scirè, Alessandro, et al.
Pubblicazione: (2024) -
Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering
di: Molfese, Francesco Maria, et al.
Pubblicazione: (2025) -
Word Sense Linking: Disambiguating Outside the Sandbox
di: Bejgu, Andrei Stefan, et al.
Pubblicazione: (2024) -
Guardians of the Machine Translation Meta-Evaluation: Sentinel Metrics Fall In!
di: Perrella, Stefano, et al.
Pubblicazione: (2024) -
Do Large Language Models Understand Word Senses?
di: Meconi, Domenico, et al.
Pubblicazione: (2025)