A Semantic-Sampling Framework for Evaluating Calibration in Open-Ended Question Answering
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Zhanliang, Xiao, Jiancong, Jin, Ruochen, Yang, Shu, Hou, Bojian, Shen, Li |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Restoring Calibration for Aligned Large Language Models: A Calibration-Aware Fine-Tuning Approach
by: Xiao, Jiancong, et al.
Published: (2025)
by: Xiao, Jiancong, et al.
Published: (2025)
Beyond Prompting: An Efficient Embedding Framework for Open-Domain Question Answering
by: Hu, Zhanghao, et al.
Published: (2025)
by: Hu, Zhanghao, et al.
Published: (2025)
Fine-Tuning Attention Modules Only: Enhancing Weight Disentanglement in Task Arithmetic
by: Jin, Ruochen, et al.
Published: (2024)
by: Jin, Ruochen, et al.
Published: (2024)
MCU: An Evaluation Framework for Open-Ended Game Agents
by: Zheng, Xinyue, et al.
Published: (2023)
by: Zheng, Xinyue, et al.
Published: (2023)
MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research
by: Chen, Hui, et al.
Published: (2025)
by: Chen, Hui, et al.
Published: (2025)
ASTRA: Adaptive Semantic Tree Reasoning Architecture for Complex Table Question Answering
by: Guo, Xiaoke, et al.
Published: (2026)
by: Guo, Xiaoke, et al.
Published: (2026)
Temporal Knowledge Graph Question Answering: A Survey
by: Su, Miao, et al.
Published: (2024)
by: Su, Miao, et al.
Published: (2024)
MentalChat16K: A Benchmark Dataset for Conversational Mental Health Assistance
by: Xu, Jia, et al.
Published: (2025)
by: Xu, Jia, et al.
Published: (2025)
XL-Suite: Cross-Lingual Synthetic Training and Evaluation Data for Open-Ended Generation
by: Iyer, Vivek, et al.
Published: (2025)
by: Iyer, Vivek, et al.
Published: (2025)
Dreaming in Code for Curriculum Learning in Open-Ended Worlds
by: Mitsides, Konstantinos, et al.
Published: (2026)
by: Mitsides, Konstantinos, et al.
Published: (2026)
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
by: Lee, Jaehyeok, et al.
Published: (2026)
by: Lee, Jaehyeok, et al.
Published: (2026)
AutoLibra: Agent Metric Induction from Open-Ended Human Feedback
by: Zhu, Hao, et al.
Published: (2025)
by: Zhu, Hao, et al.
Published: (2025)
Self-Rewarding Rubric-Based Reinforcement Learning for Open-Ended Reasoning
by: Ye, Zhiling, et al.
Published: (2025)
by: Ye, Zhiling, et al.
Published: (2025)
DIVERGE: Diversity-Enhanced RAG for Open-Ended Information Seeking
by: Hu, Tianyi, et al.
Published: (2026)
by: Hu, Tianyi, et al.
Published: (2026)
Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts
by: Samvelyan, Mikayel, et al.
Published: (2024)
by: Samvelyan, Mikayel, et al.
Published: (2024)
Time-MQA: Time Series Multi-Task Question Answering with Context Enhancement
by: Kong, Yaxuan, et al.
Published: (2025)
by: Kong, Yaxuan, et al.
Published: (2025)
Beyond Semantic Similarity: A Component-Wise Evaluation Framework for Medical Question Answering Systems with Health Equity Implications
by: Sakib, Abu Noman Md, et al.
Published: (2026)
by: Sakib, Abu Noman Md, et al.
Published: (2026)
Interpretable Question Answering with Knowledge Graphs
by: Aneja, Kartikeya, et al.
Published: (2025)
by: Aneja, Kartikeya, et al.
Published: (2025)
The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
by: Lu, Chris, et al.
Published: (2024)
by: Lu, Chris, et al.
Published: (2024)
Beyond Independent Passages: Adaptive Passage Combination Retrieval for Retrieval Augmented Open-Domain Question Answering
by: Ko, Ting-Wen, et al.
Published: (2025)
by: Ko, Ting-Wen, et al.
Published: (2025)
Can LLMs Reliably Simulate Human Learner Actions? A Simulation Authoring Framework for Open-Ended Learning Environments
by: Mannekote, Amogh, et al.
Published: (2024)
by: Mannekote, Amogh, et al.
Published: (2024)
RAMQA: A Unified Framework for Retrieval-Augmented Multi-Modal Question Answering
by: Bai, Yang, et al.
Published: (2025)
by: Bai, Yang, et al.
Published: (2025)
GRLO: Towards Generalizable Reinforcement Learning in Open-Ended Environments from Zero
by: Yin, Shangjian, et al.
Published: (2026)
by: Yin, Shangjian, et al.
Published: (2026)
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models
by: Couturier, Camille, et al.
Published: (2025)
by: Couturier, Camille, et al.
Published: (2025)
M-QUEST -- Meme Question-Understanding Evaluation on Semantics and Toxicity
by: De Giorgis, Stefano, et al.
Published: (2026)
by: De Giorgis, Stefano, et al.
Published: (2026)
A Large-Scale Benchmark for Evaluating Large Language Models on Medical Question Answering in Romanian
by: Rogoz, Ana-Cristina, et al.
Published: (2025)
by: Rogoz, Ana-Cristina, et al.
Published: (2025)
AQAScore: Evaluating Semantic Alignment in Text-to-Audio Generation via Audio Question Answering
by: Kuan, Chun-Yi, et al.
Published: (2026)
by: Kuan, Chun-Yi, et al.
Published: (2026)
Biomedical Entity Linking as Multiple Choice Question Answering
by: Lin, Zhenxi, et al.
Published: (2024)
by: Lin, Zhenxi, et al.
Published: (2024)
SEMQA: Semi-Extractive Multi-Source Question Answering
by: Schuster, Tal, et al.
Published: (2023)
by: Schuster, Tal, et al.
Published: (2023)
COMPKE: Complex Question Answering under Knowledge Editing
by: Cheng, Keyuan, et al.
Published: (2025)
by: Cheng, Keyuan, et al.
Published: (2025)
Assessing Modality Bias in Video Question Answering Benchmarks with Multimodal Large Language Models
by: Park, Jean, et al.
Published: (2024)
by: Park, Jean, et al.
Published: (2024)
LinkQ: An LLM-Assisted Visual Interface for Knowledge Graph Question-Answering
by: Li, Harry, et al.
Published: (2024)
by: Li, Harry, et al.
Published: (2024)
Flexible Agent Alignment with Goal Inference from Open-Ended Dialog
by: Ma, Rachel, et al.
Published: (2025)
by: Ma, Rachel, et al.
Published: (2025)
Deceiving Question-Answering Models: A Hybrid Word-Level Adversarial Approach
by: Li, Jiyao, et al.
Published: (2024)
by: Li, Jiyao, et al.
Published: (2024)
W-RAG: Weakly Supervised Dense Retrieval in RAG for Open-domain Question Answering
by: Nian, Jinming, et al.
Published: (2024)
by: Nian, Jinming, et al.
Published: (2024)
Towards Robust Evaluation: A Comprehensive Taxonomy of Datasets and Metrics for Open Domain Question Answering in the Era of Large Language Models
by: Srivastava, Akchay, et al.
Published: (2024)
by: Srivastava, Akchay, et al.
Published: (2024)
Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math
by: Pandit, Shrey, et al.
Published: (2025)
by: Pandit, Shrey, et al.
Published: (2025)
TopBench: A Benchmark for Implicit Prediction and Reasoning over Tabular Question Answering
by: Ji, An-Yang, et al.
Published: (2026)
by: Ji, An-Yang, et al.
Published: (2026)
Multi-hop Question Answering under Temporal Knowledge Editing
by: Cheng, Keyuan, et al.
Published: (2024)
by: Cheng, Keyuan, et al.
Published: (2024)
From Chat Logs to Collective Insights: Aggregative Question Answering
by: Zhang, Wentao, et al.
Published: (2025)
by: Zhang, Wentao, et al.
Published: (2025)
Similar Items
-
Restoring Calibration for Aligned Large Language Models: A Calibration-Aware Fine-Tuning Approach
by: Xiao, Jiancong, et al.
Published: (2025) -
Beyond Prompting: An Efficient Embedding Framework for Open-Domain Question Answering
by: Hu, Zhanghao, et al.
Published: (2025) -
Fine-Tuning Attention Modules Only: Enhancing Weight Disentanglement in Task Arithmetic
by: Jin, Ruochen, et al.
Published: (2024) -
MCU: An Evaluation Framework for Open-Ended Game Agents
by: Zheng, Xinyue, et al.
Published: (2023) -
MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research
by: Chen, Hui, et al.
Published: (2025)