Evaluating Consistencies in LLM responses through a Semantic Clustering of Question Answering
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Yanggyu, Kim, Jihie |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving LLM Classification of Logical Errors by Integrating Error Relationship into Prompts
by: Lee, Yanggyu, et al.
Published: (2024)
by: Lee, Yanggyu, et al.
Published: (2024)
Piece of Table: A Divide-and-Conquer Approach for Selecting Subtables in Table Question Answering
by: Lee, Wonjin, et al.
Published: (2024)
by: Lee, Wonjin, et al.
Published: (2024)
KoBBQ: Korean Bias Benchmark for Question Answering
by: Jin, Jiho, et al.
Published: (2023)
by: Jin, Jiho, et al.
Published: (2023)
Discrepancy Detection at the Data Level: Toward Consistent Multilingual Question Answering
by: Calvo-Bartolomé, Lorena, et al.
Published: (2025)
by: Calvo-Bartolomé, Lorena, et al.
Published: (2025)
CIC: A Framework for Culturally-Aware Image Captioning
by: Yun, Youngsik, et al.
Published: (2024)
by: Yun, Youngsik, et al.
Published: (2024)
Denoising Table-Text Retrieval for Open-Domain Question Answering
by: Kang, Deokhyung, et al.
Published: (2024)
by: Kang, Deokhyung, et al.
Published: (2024)
FHIR-AgentBench: Benchmarking LLM Agents for Realistic Interoperable EHR Question Answering
by: Lee, Gyubok, et al.
Published: (2025)
by: Lee, Gyubok, et al.
Published: (2025)
LLM Robustness Against Misinformation in Biomedical Question Answering
by: Bondarenko, Alexander, et al.
Published: (2024)
by: Bondarenko, Alexander, et al.
Published: (2024)
Improved LLM Agents for Financial Document Question Answering
by: Tan, Nelvin, et al.
Published: (2025)
by: Tan, Nelvin, et al.
Published: (2025)
QPaug: Question and Passage Augmentation for Open-Domain Question Answering of LLMs
by: Kim, Minsang, et al.
Published: (2024)
by: Kim, Minsang, et al.
Published: (2024)
SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation
by: Kendre, Shrikant, et al.
Published: (2025)
by: Kendre, Shrikant, et al.
Published: (2025)
A Semantic-Sampling Framework for Evaluating Calibration in Open-Ended Question Answering
by: Wang, Zhanliang, et al.
Published: (2026)
by: Wang, Zhanliang, et al.
Published: (2026)
YESciEval: Robust LLM-as-a-Judge for Scientific Question Answering
by: D'Souza, Jennifer, et al.
Published: (2025)
by: D'Souza, Jennifer, et al.
Published: (2025)
Efficient Multi-Hop Question Answering over Knowledge Graphs via LLM Planning and Embedding-Guided Search
by: Shrestha, Manil, et al.
Published: (2025)
by: Shrestha, Manil, et al.
Published: (2025)
Automatic Evaluation of Healthcare LLMs Beyond Question-Answering
by: Arias-Duart, Anna, et al.
Published: (2025)
by: Arias-Duart, Anna, et al.
Published: (2025)
Understanding Network Behaviors through Natural Language Question-Answering
by: Xing, Mingzhe, et al.
Published: (2025)
by: Xing, Mingzhe, et al.
Published: (2025)
Optimized Biomedical Question-Answering Services with LLM and Multi-BERT Integration
by: Qian, Cheng, et al.
Published: (2024)
by: Qian, Cheng, et al.
Published: (2024)
Improving LLM Reliability with RAG in Religious Question-Answering: MufassirQAS
by: Alan, Ahmet Yusuf, et al.
Published: (2024)
by: Alan, Ahmet Yusuf, et al.
Published: (2024)
Compositional Consistency-Guided Decoding for Three-Way Logical Question Answering
by: Huang, Tianyi, et al.
Published: (2026)
by: Huang, Tianyi, et al.
Published: (2026)
Evaluating Correctness and Faithfulness of Instruction-Following Models for Question Answering
by: Adlakha, Vaibhav, et al.
Published: (2023)
by: Adlakha, Vaibhav, et al.
Published: (2023)
CoE-Ops: Collaboration of LLM-based Experts for AIOps Question-Answering
by: Zhao, Jinkun, et al.
Published: (2025)
by: Zhao, Jinkun, et al.
Published: (2025)
Prompt-Based Clarity Evaluation and Topic Detection in Political Question Answering
by: Prahallad, Lavanya, et al.
Published: (2026)
by: Prahallad, Lavanya, et al.
Published: (2026)
OLAPH: Improving Factuality in Biomedical Long-form Question Answering
by: Jeong, Minbyul, et al.
Published: (2024)
by: Jeong, Minbyul, et al.
Published: (2024)
SemViQA: A Semantic Question Answering System for Vietnamese Information Fact-Checking
by: Tran, Dien X., et al.
Published: (2025)
by: Tran, Dien X., et al.
Published: (2025)
Improving Commonsense Bias Classification by Mitigating the Influence of Demographic Terms
by: Lee, JinKyu, et al.
Published: (2024)
by: Lee, JinKyu, et al.
Published: (2024)
AQAScore: Evaluating Semantic Alignment in Text-to-Audio Generation via Audio Question Answering
by: Kuan, Chun-Yi, et al.
Published: (2026)
by: Kuan, Chun-Yi, et al.
Published: (2026)
MedExQA: Medical Question Answering Benchmark with Multiple Explanations
by: Kim, Yunsoo, et al.
Published: (2024)
by: Kim, Yunsoo, et al.
Published: (2024)
Confidence-guided Refinement Reasoning for Zero-shot Question Answering
by: Jang, Youwon, et al.
Published: (2025)
by: Jang, Youwon, et al.
Published: (2025)
CompactRAG: Reducing LLM Calls and Token Overhead in Multi-Hop Question Answering
by: Yang, Hao, et al.
Published: (2026)
by: Yang, Hao, et al.
Published: (2026)
An Analysis of Decoding Methods for LLM-based Agents for Faithful Multi-Hop Question Answering
by: Murphy, Alexander, et al.
Published: (2025)
by: Murphy, Alexander, et al.
Published: (2025)
TRAQ: Trustworthy Retrieval Augmented Question Answering via Conformal Prediction
by: Li, Shuo, et al.
Published: (2023)
by: Li, Shuo, et al.
Published: (2023)
Long-Span Question-Answering: Automatic Question Generation and QA-System Ranking via Side-by-Side Evaluation
by: Bohnet, Bernd, et al.
Published: (2024)
by: Bohnet, Bernd, et al.
Published: (2024)
ResearchQA: Evaluating Scholarly Question Answering at Scale Across 75 Fields with Survey-Mined Questions and Rubrics
by: Yifei, Li S., et al.
Published: (2025)
by: Yifei, Li S., et al.
Published: (2025)
Evaluating the Meta- and Object-Level Reasoning of Large Language Models for Question Answering
by: Ferguson, Nick, et al.
Published: (2025)
by: Ferguson, Nick, et al.
Published: (2025)
Case-Based Reasoning Approach for Solving Financial Question Answering
by: Kim, Yikyung, et al.
Published: (2024)
by: Kim, Yikyung, et al.
Published: (2024)
Efficient Medical Question Answering with Knowledge-Augmented Question Generation
by: Khlaut, Julien, et al.
Published: (2024)
by: Khlaut, Julien, et al.
Published: (2024)
When Language Shapes Thought: Cross-Lingual Transfer of Factual Knowledge in Question Answering
by: Kang, Eojin, et al.
Published: (2025)
by: Kang, Eojin, et al.
Published: (2025)
StableSketcher: Enhancing Diffusion Model for Pixel-based Sketch Generation via Visual Question Answering Feedback
by: Park, Jiho, et al.
Published: (2025)
by: Park, Jiho, et al.
Published: (2025)
SemRAG: Semantic Knowledge-Augmented RAG for Improved Question-Answering
by: Zhong, Kezhen, et al.
Published: (2025)
by: Zhong, Kezhen, et al.
Published: (2025)
Training LLMs with Reinforcement Learning for Intent-Aware Personalized Question Answering
by: Amirizaniani, Maryam, et al.
Published: (2026)
by: Amirizaniani, Maryam, et al.
Published: (2026)
Similar Items
-
Improving LLM Classification of Logical Errors by Integrating Error Relationship into Prompts
by: Lee, Yanggyu, et al.
Published: (2024) -
Piece of Table: A Divide-and-Conquer Approach for Selecting Subtables in Table Question Answering
by: Lee, Wonjin, et al.
Published: (2024) -
KoBBQ: Korean Bias Benchmark for Question Answering
by: Jin, Jiho, et al.
Published: (2023) -
Discrepancy Detection at the Data Level: Toward Consistent Multilingual Question Answering
by: Calvo-Bartolomé, Lorena, et al.
Published: (2025) -
CIC: A Framework for Culturally-Aware Image Captioning
by: Yun, Youngsik, et al.
Published: (2024)