Evaluating Consistencies in LLM responses through a Semantic Clustering of Question Answering
Fuente:
arXiv
Salvato in:
| Autori principali: | Lee, Yanggyu, Kim, Jihie |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Improving LLM Classification of Logical Errors by Integrating Error Relationship into Prompts
di: Lee, Yanggyu, et al.
Pubblicazione: (2024)
di: Lee, Yanggyu, et al.
Pubblicazione: (2024)
Piece of Table: A Divide-and-Conquer Approach for Selecting Subtables in Table Question Answering
di: Lee, Wonjin, et al.
Pubblicazione: (2024)
di: Lee, Wonjin, et al.
Pubblicazione: (2024)
KoBBQ: Korean Bias Benchmark for Question Answering
di: Jin, Jiho, et al.
Pubblicazione: (2023)
di: Jin, Jiho, et al.
Pubblicazione: (2023)
Discrepancy Detection at the Data Level: Toward Consistent Multilingual Question Answering
di: Calvo-Bartolomé, Lorena, et al.
Pubblicazione: (2025)
di: Calvo-Bartolomé, Lorena, et al.
Pubblicazione: (2025)
CIC: A Framework for Culturally-Aware Image Captioning
di: Yun, Youngsik, et al.
Pubblicazione: (2024)
di: Yun, Youngsik, et al.
Pubblicazione: (2024)
Denoising Table-Text Retrieval for Open-Domain Question Answering
di: Kang, Deokhyung, et al.
Pubblicazione: (2024)
di: Kang, Deokhyung, et al.
Pubblicazione: (2024)
FHIR-AgentBench: Benchmarking LLM Agents for Realistic Interoperable EHR Question Answering
di: Lee, Gyubok, et al.
Pubblicazione: (2025)
di: Lee, Gyubok, et al.
Pubblicazione: (2025)
LLM Robustness Against Misinformation in Biomedical Question Answering
di: Bondarenko, Alexander, et al.
Pubblicazione: (2024)
di: Bondarenko, Alexander, et al.
Pubblicazione: (2024)
Improved LLM Agents for Financial Document Question Answering
di: Tan, Nelvin, et al.
Pubblicazione: (2025)
di: Tan, Nelvin, et al.
Pubblicazione: (2025)
QPaug: Question and Passage Augmentation for Open-Domain Question Answering of LLMs
di: Kim, Minsang, et al.
Pubblicazione: (2024)
di: Kim, Minsang, et al.
Pubblicazione: (2024)
SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation
di: Kendre, Shrikant, et al.
Pubblicazione: (2025)
di: Kendre, Shrikant, et al.
Pubblicazione: (2025)
A Semantic-Sampling Framework for Evaluating Calibration in Open-Ended Question Answering
di: Wang, Zhanliang, et al.
Pubblicazione: (2026)
di: Wang, Zhanliang, et al.
Pubblicazione: (2026)
YESciEval: Robust LLM-as-a-Judge for Scientific Question Answering
di: D'Souza, Jennifer, et al.
Pubblicazione: (2025)
di: D'Souza, Jennifer, et al.
Pubblicazione: (2025)
Efficient Multi-Hop Question Answering over Knowledge Graphs via LLM Planning and Embedding-Guided Search
di: Shrestha, Manil, et al.
Pubblicazione: (2025)
di: Shrestha, Manil, et al.
Pubblicazione: (2025)
Automatic Evaluation of Healthcare LLMs Beyond Question-Answering
di: Arias-Duart, Anna, et al.
Pubblicazione: (2025)
di: Arias-Duart, Anna, et al.
Pubblicazione: (2025)
Understanding Network Behaviors through Natural Language Question-Answering
di: Xing, Mingzhe, et al.
Pubblicazione: (2025)
di: Xing, Mingzhe, et al.
Pubblicazione: (2025)
Optimized Biomedical Question-Answering Services with LLM and Multi-BERT Integration
di: Qian, Cheng, et al.
Pubblicazione: (2024)
di: Qian, Cheng, et al.
Pubblicazione: (2024)
Improving LLM Reliability with RAG in Religious Question-Answering: MufassirQAS
di: Alan, Ahmet Yusuf, et al.
Pubblicazione: (2024)
di: Alan, Ahmet Yusuf, et al.
Pubblicazione: (2024)
Compositional Consistency-Guided Decoding for Three-Way Logical Question Answering
di: Huang, Tianyi, et al.
Pubblicazione: (2026)
di: Huang, Tianyi, et al.
Pubblicazione: (2026)
Evaluating Correctness and Faithfulness of Instruction-Following Models for Question Answering
di: Adlakha, Vaibhav, et al.
Pubblicazione: (2023)
di: Adlakha, Vaibhav, et al.
Pubblicazione: (2023)
CoE-Ops: Collaboration of LLM-based Experts for AIOps Question-Answering
di: Zhao, Jinkun, et al.
Pubblicazione: (2025)
di: Zhao, Jinkun, et al.
Pubblicazione: (2025)
Prompt-Based Clarity Evaluation and Topic Detection in Political Question Answering
di: Prahallad, Lavanya, et al.
Pubblicazione: (2026)
di: Prahallad, Lavanya, et al.
Pubblicazione: (2026)
OLAPH: Improving Factuality in Biomedical Long-form Question Answering
di: Jeong, Minbyul, et al.
Pubblicazione: (2024)
di: Jeong, Minbyul, et al.
Pubblicazione: (2024)
SemViQA: A Semantic Question Answering System for Vietnamese Information Fact-Checking
di: Tran, Dien X., et al.
Pubblicazione: (2025)
di: Tran, Dien X., et al.
Pubblicazione: (2025)
Improving Commonsense Bias Classification by Mitigating the Influence of Demographic Terms
di: Lee, JinKyu, et al.
Pubblicazione: (2024)
di: Lee, JinKyu, et al.
Pubblicazione: (2024)
AQAScore: Evaluating Semantic Alignment in Text-to-Audio Generation via Audio Question Answering
di: Kuan, Chun-Yi, et al.
Pubblicazione: (2026)
di: Kuan, Chun-Yi, et al.
Pubblicazione: (2026)
MedExQA: Medical Question Answering Benchmark with Multiple Explanations
di: Kim, Yunsoo, et al.
Pubblicazione: (2024)
di: Kim, Yunsoo, et al.
Pubblicazione: (2024)
Confidence-guided Refinement Reasoning for Zero-shot Question Answering
di: Jang, Youwon, et al.
Pubblicazione: (2025)
di: Jang, Youwon, et al.
Pubblicazione: (2025)
CompactRAG: Reducing LLM Calls and Token Overhead in Multi-Hop Question Answering
di: Yang, Hao, et al.
Pubblicazione: (2026)
di: Yang, Hao, et al.
Pubblicazione: (2026)
An Analysis of Decoding Methods for LLM-based Agents for Faithful Multi-Hop Question Answering
di: Murphy, Alexander, et al.
Pubblicazione: (2025)
di: Murphy, Alexander, et al.
Pubblicazione: (2025)
TRAQ: Trustworthy Retrieval Augmented Question Answering via Conformal Prediction
di: Li, Shuo, et al.
Pubblicazione: (2023)
di: Li, Shuo, et al.
Pubblicazione: (2023)
Long-Span Question-Answering: Automatic Question Generation and QA-System Ranking via Side-by-Side Evaluation
di: Bohnet, Bernd, et al.
Pubblicazione: (2024)
di: Bohnet, Bernd, et al.
Pubblicazione: (2024)
ResearchQA: Evaluating Scholarly Question Answering at Scale Across 75 Fields with Survey-Mined Questions and Rubrics
di: Yifei, Li S., et al.
Pubblicazione: (2025)
di: Yifei, Li S., et al.
Pubblicazione: (2025)
Evaluating the Meta- and Object-Level Reasoning of Large Language Models for Question Answering
di: Ferguson, Nick, et al.
Pubblicazione: (2025)
di: Ferguson, Nick, et al.
Pubblicazione: (2025)
Case-Based Reasoning Approach for Solving Financial Question Answering
di: Kim, Yikyung, et al.
Pubblicazione: (2024)
di: Kim, Yikyung, et al.
Pubblicazione: (2024)
Efficient Medical Question Answering with Knowledge-Augmented Question Generation
di: Khlaut, Julien, et al.
Pubblicazione: (2024)
di: Khlaut, Julien, et al.
Pubblicazione: (2024)
When Language Shapes Thought: Cross-Lingual Transfer of Factual Knowledge in Question Answering
di: Kang, Eojin, et al.
Pubblicazione: (2025)
di: Kang, Eojin, et al.
Pubblicazione: (2025)
StableSketcher: Enhancing Diffusion Model for Pixel-based Sketch Generation via Visual Question Answering Feedback
di: Park, Jiho, et al.
Pubblicazione: (2025)
di: Park, Jiho, et al.
Pubblicazione: (2025)
SemRAG: Semantic Knowledge-Augmented RAG for Improved Question-Answering
di: Zhong, Kezhen, et al.
Pubblicazione: (2025)
di: Zhong, Kezhen, et al.
Pubblicazione: (2025)
Training LLMs with Reinforcement Learning for Intent-Aware Personalized Question Answering
di: Amirizaniani, Maryam, et al.
Pubblicazione: (2026)
di: Amirizaniani, Maryam, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Improving LLM Classification of Logical Errors by Integrating Error Relationship into Prompts
di: Lee, Yanggyu, et al.
Pubblicazione: (2024) -
Piece of Table: A Divide-and-Conquer Approach for Selecting Subtables in Table Question Answering
di: Lee, Wonjin, et al.
Pubblicazione: (2024) -
KoBBQ: Korean Bias Benchmark for Question Answering
di: Jin, Jiho, et al.
Pubblicazione: (2023) -
Discrepancy Detection at the Data Level: Toward Consistent Multilingual Question Answering
di: Calvo-Bartolomé, Lorena, et al.
Pubblicazione: (2025) -
CIC: A Framework for Culturally-Aware Image Captioning
di: Yun, Youngsik, et al.
Pubblicazione: (2024)