Argument-Based Comparative Question Answering Evaluation Benchmark
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nikishina, Irina, Anwar, Saba, Dolgov, Nikolay, Manina, Maria, Ignatenko, Daria, Moskvoretskii, Viktor, Shelmanov, Artem, Baldwin, Tim, Biemann, Chris |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Self-Taught Self-Correction for Small Language Models
von: Moskvoretskii, Viktor, et al.
Veröffentlicht: (2025)
von: Moskvoretskii, Viktor, et al.
Veröffentlicht: (2025)
Do I look like a `cat.n.01` to you? A Taxonomy Image Generation Benchmark
von: Moskvoretskii, Viktor, et al.
Veröffentlicht: (2025)
von: Moskvoretskii, Viktor, et al.
Veröffentlicht: (2025)
Low-Resource Machine Translation through the Lens of Personalized Federated Learning
von: Moskvoretskii, Viktor, et al.
Veröffentlicht: (2024)
von: Moskvoretskii, Viktor, et al.
Veröffentlicht: (2024)
TaxoLLaMA: WordNet-based Model for Solving Multiple Lexical Semantic Tasks
von: Moskvoretskii, Viktor, et al.
Veröffentlicht: (2024)
von: Moskvoretskii, Viktor, et al.
Veröffentlicht: (2024)
Adaptive Retrieval Without Self-Knowledge? Bringing Uncertainty Back Home
von: Moskvoretskii, Viktor, et al.
Veröffentlicht: (2025)
von: Moskvoretskii, Viktor, et al.
Veröffentlicht: (2025)
Will It Still Be True Tomorrow? Multilingual Evergreen Question Classification to Improve Trustworthy QA
von: Pletenev, Sergey, et al.
Veröffentlicht: (2025)
von: Pletenev, Sergey, et al.
Veröffentlicht: (2025)
Perspectives - Interactive Document Clustering in the Discourse Analysis Tool Suite
von: Fischer, Tim, et al.
Veröffentlicht: (2026)
von: Fischer, Tim, et al.
Veröffentlicht: (2026)
LLM-Independent Adaptive RAG: Let the Question Speak for Itself
von: Marina, Maria, et al.
Veröffentlicht: (2025)
von: Marina, Maria, et al.
Veröffentlicht: (2025)
Leveraging LLM Parametric Knowledge for Fact Checking without Retrieval
von: Vazhentsev, Artem, et al.
Veröffentlicht: (2026)
von: Vazhentsev, Artem, et al.
Veröffentlicht: (2026)
Token-Level Density-Based Uncertainty Quantification Methods for Eliciting Truthfulness of Large Language Models
von: Vazhentsev, Artem, et al.
Veröffentlicht: (2025)
von: Vazhentsev, Artem, et al.
Veröffentlicht: (2025)
Uncertainty Quantification for Large Language Diffusion Models
von: Vazhentsev, Artem, et al.
Veröffentlicht: (2026)
von: Vazhentsev, Artem, et al.
Veröffentlicht: (2026)
Inference-Time Selective Debiasing to Enhance Fairness in Text Classification Models
von: Kuzmin, Gleb, et al.
Veröffentlicht: (2024)
von: Kuzmin, Gleb, et al.
Veröffentlicht: (2024)
Asking and Answering Questions to Extract Event-Argument Structures
von: Uddin, Md Nayem, et al.
Veröffentlicht: (2024)
von: Uddin, Md Nayem, et al.
Veröffentlicht: (2024)
Location Aware Modular Biencoder for Tourism Question Answering
von: Li, Haonan, et al.
Veröffentlicht: (2024)
von: Li, Haonan, et al.
Veröffentlicht: (2024)
GIMMICK -- Globally Inclusive Multimodal Multitask Cultural Knowledge Benchmarking
von: Schneider, Florian, et al.
Veröffentlicht: (2025)
von: Schneider, Florian, et al.
Veröffentlicht: (2025)
Question-Answering Based Summarization of Electronic Health Records using Retrieval Augmented Generation
von: Saba, Walid, et al.
Veröffentlicht: (2024)
von: Saba, Walid, et al.
Veröffentlicht: (2024)
A Head to Predict and a Head to Question: Pre-trained Uncertainty Quantification Heads for Hallucination Detection in LLM Outputs
von: Shelmanov, Artem, et al.
Veröffentlicht: (2025)
von: Shelmanov, Artem, et al.
Veröffentlicht: (2025)
Creating a Taxonomy for Retrieval Augmented Generation Applications
von: Nikishina, Irina, et al.
Veröffentlicht: (2024)
von: Nikishina, Irina, et al.
Veröffentlicht: (2024)
Large Language Models Are Overparameterized Text Encoders
von: K, Thennal D, et al.
Veröffentlicht: (2024)
von: K, Thennal D, et al.
Veröffentlicht: (2024)
Dataset of Quotation Attribution in German News Articles
von: Petersen-Frey, Fynn, et al.
Veröffentlicht: (2024)
von: Petersen-Frey, Fynn, et al.
Veröffentlicht: (2024)
MVL-SIB: A Massively Multilingual Vision-Language Benchmark for Cross-Modal Topical Matching
von: Schmidt, Fabian David, et al.
Veröffentlicht: (2025)
von: Schmidt, Fabian David, et al.
Veröffentlicht: (2025)
Unconditional Truthfulness: Learning Unconditional Uncertainty of Large Language Models
von: Vazhentsev, Artem, et al.
Veröffentlicht: (2024)
von: Vazhentsev, Artem, et al.
Veröffentlicht: (2024)
GroUSE: A Benchmark to Evaluate Evaluators in Grounded Question Answering
von: Muller, Sacha, et al.
Veröffentlicht: (2024)
von: Muller, Sacha, et al.
Veröffentlicht: (2024)
EncouRAGe: Evaluating RAG Local, Fast, and Reliable
von: Strich, Jan, et al.
Veröffentlicht: (2025)
von: Strich, Jan, et al.
Veröffentlicht: (2025)
DisastQA: A Comprehensive Benchmark for Evaluating Question Answering in Disaster Management
von: Chen, Zhitong, et al.
Veröffentlicht: (2026)
von: Chen, Zhitong, et al.
Veröffentlicht: (2026)
IMB: An Italian Medical Benchmark for Question Answering
von: Romano, Antonio, et al.
Veröffentlicht: (2025)
von: Romano, Antonio, et al.
Veröffentlicht: (2025)
Uncertainty-Aware Attention Heads: Efficient Unsupervised Uncertainty Quantification for LLMs
von: Vazhentsev, Artem, et al.
Veröffentlicht: (2025)
von: Vazhentsev, Artem, et al.
Veröffentlicht: (2025)
FREB-TQA: A Fine-Grained Robustness Evaluation Benchmark for Table Question Answering
von: Zhou, Wei, et al.
Veröffentlicht: (2024)
von: Zhou, Wei, et al.
Veröffentlicht: (2024)
Benchmarking Uncertainty Quantification Methods for Large Language Models with LM-Polygraph
von: Vashurin, Roman, et al.
Veröffentlicht: (2024)
von: Vashurin, Roman, et al.
Veröffentlicht: (2024)
MULTITAT: Benchmarking Multilingual Table-and-Text Question Answering
von: Zhang, Xuanliang, et al.
Veröffentlicht: (2025)
von: Zhang, Xuanliang, et al.
Veröffentlicht: (2025)
CasiMedicos-Arg: A Medical Question Answering Dataset Annotated with Explanatory Argumentative Structures
von: Sviridova, Ekaterina, et al.
Veröffentlicht: (2024)
von: Sviridova, Ekaterina, et al.
Veröffentlicht: (2024)
Evaluating Large Language Models for Evidence-Based Clinical Question Answering
von: Wang, Can, et al.
Veröffentlicht: (2025)
von: Wang, Can, et al.
Veröffentlicht: (2025)
Combining LLMs and Knowledge Graphs to Reduce Hallucinations in Question Answering
von: Pusch, Larissa, et al.
Veröffentlicht: (2024)
von: Pusch, Larissa, et al.
Veröffentlicht: (2024)
Evaluating the Retrieval Component in LLM-Based Question Answering Systems
von: Alinejad, Ashkan, et al.
Veröffentlicht: (2024)
von: Alinejad, Ashkan, et al.
Veröffentlicht: (2024)
Vikhr: The Family of Open-Source Instruction-Tuned Large Language Models for Russian
von: Nikolich, Aleksandr, et al.
Veröffentlicht: (2024)
von: Nikolich, Aleksandr, et al.
Veröffentlicht: (2024)
Pre-trained Transformer-Based Approach for Arabic Question Answering : A Comparative Study
von: Alsubhi, Kholoud, et al.
Veröffentlicht: (2021)
von: Alsubhi, Kholoud, et al.
Veröffentlicht: (2021)
Uhura: A Benchmark for Evaluating Scientific Question Answering and Truthfulness in Low-Resource African Languages
von: Bayes, Edward, et al.
Veröffentlicht: (2024)
von: Bayes, Edward, et al.
Veröffentlicht: (2024)
MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge
von: He, Jie, et al.
Veröffentlicht: (2024)
von: He, Jie, et al.
Veröffentlicht: (2024)
Spider4SPARQL: A Complex Benchmark for Evaluating Knowledge Graph Question Answering Systems
von: Kosten, Catherine, et al.
Veröffentlicht: (2023)
von: Kosten, Catherine, et al.
Veröffentlicht: (2023)
ESQA: Event Sequences Question Answering
von: Abdullaeva, Irina, et al.
Veröffentlicht: (2024)
von: Abdullaeva, Irina, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Self-Taught Self-Correction for Small Language Models
von: Moskvoretskii, Viktor, et al.
Veröffentlicht: (2025) -
Do I look like a `cat.n.01` to you? A Taxonomy Image Generation Benchmark
von: Moskvoretskii, Viktor, et al.
Veröffentlicht: (2025) -
Low-Resource Machine Translation through the Lens of Personalized Federated Learning
von: Moskvoretskii, Viktor, et al.
Veröffentlicht: (2024) -
TaxoLLaMA: WordNet-based Model for Solving Multiple Lexical Semantic Tasks
von: Moskvoretskii, Viktor, et al.
Veröffentlicht: (2024) -
Adaptive Retrieval Without Self-Knowledge? Bringing Uncertainty Back Home
von: Moskvoretskii, Viktor, et al.
Veröffentlicht: (2025)