CFMatch: Aligning Automated Answer Equivalence Evaluation with Expert Judgments For Open-Domain Question Answering
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Li, Zongxia, Mondal, Ishani, Liang, Yijun, Nghiem, Huy, Boyd-Graber, Jordan |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
PEDANTS: Cheap but Effective and Interpretable Answer Equivalence
par: Li, Zongxia, et autres
Publié: (2024)
par: Li, Zongxia, et autres
Publié: (2024)
How the Advent of Ubiquitous Large Language Models both Stymie and Turbocharge Dynamic Adversarial Question Generation
par: Sung, Yoo Yeon, et autres
Publié: (2024)
par: Sung, Yoo Yeon, et autres
Publié: (2024)
SciDoc2Diagrammer-MAF: Towards Generation of Scientific Diagrams from Documents guided by Multi-Aspect Feedback Refinement
par: Mondal, Ishani, et autres
Publié: (2024)
par: Mondal, Ishani, et autres
Publié: (2024)
Reverse Question Answering: Can an LLM Write a Question so Hard (or Bad) that it Can't Answer?
par: Balepur, Nishant, et autres
Publié: (2024)
par: Balepur, Nishant, et autres
Publié: (2024)
Large Language Models Are Effective Human Annotation Assistants, But Not Good Independent Annotators
par: Gu, Feng, et autres
Publié: (2025)
par: Gu, Feng, et autres
Publié: (2025)
DiscoTrace: Representing and Comparing Answering Strategies of Humans and LLMs in Information-Seeking Question Answering
par: Srikanth, Neha, et autres
Publié: (2026)
par: Srikanth, Neha, et autres
Publié: (2026)
Is your benchmark truly adversarial? AdvScore: Evaluating Human-Grounded Adversarialness
par: Sung, Yoo Yeon, et autres
Publié: (2024)
par: Sung, Yoo Yeon, et autres
Publié: (2024)
SMART-Editor: A Multi-Agent Framework for Human-Like Design Editing with Structural Integrity
par: Mondal, Ishani, et autres
Publié: (2025)
par: Mondal, Ishani, et autres
Publié: (2025)
Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers
par: Yona, Gal, et autres
Publié: (2024)
par: Yona, Gal, et autres
Publié: (2024)
CANVAS: Continuity-Aware Narratives via Visual Agentic Storyboarding
par: Mondal, Ishani, et autres
Publié: (2026)
par: Mondal, Ishani, et autres
Publié: (2026)
Pregnant Questions: The Importance of Pragmatic Awareness in Maternal Health Question Answering
par: Srikanth, Neha, et autres
Publié: (2023)
par: Srikanth, Neha, et autres
Publié: (2023)
Do great minds think alike? Investigating Human-AI Complementarity in Question Answering with CAIMIRA
par: Gor, Maharshi, et autres
Publié: (2024)
par: Gor, Maharshi, et autres
Publié: (2024)
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation
par: Li, Zongxia, et autres
Publié: (2025)
par: Li, Zongxia, et autres
Publié: (2025)
A Dataset of Open-Domain Question Answering with Multiple-Span Answers
par: Luo, Zhiyi, et autres
Publié: (2024)
par: Luo, Zhiyi, et autres
Publié: (2024)
Discrepancy Detection at the Data Level: Toward Consistent Multilingual Question Answering
par: Calvo-Bartolomé, Lorena, et autres
Publié: (2025)
par: Calvo-Bartolomé, Lorena, et autres
Publié: (2025)
Large Language Models Struggle to Describe the Haystack without Human Help: Human-in-the-loop Evaluation of Topic Models
par: Li, Zongxia, et autres
Publié: (2025)
par: Li, Zongxia, et autres
Publié: (2025)
AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?
par: Gor, Maharshi, et autres
Publié: (2026)
par: Gor, Maharshi, et autres
Publié: (2026)
Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above
par: Balepur, Nishant, et autres
Publié: (2025)
par: Balepur, Nishant, et autres
Publié: (2025)
A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges
par: Li, Zongxia, et autres
Publié: (2025)
par: Li, Zongxia, et autres
Publié: (2025)
Benchmarked Yet Not Measured -- Generative AI Should be Evaluated Against Real-World Utility
par: Mondal, Ishani, et autres
Publié: (2026)
par: Mondal, Ishani, et autres
Publié: (2026)
ProxAnn: Use-Oriented Evaluations of Topic Models and Document Clustering
par: Hoyle, Alexander, et autres
Publié: (2025)
par: Hoyle, Alexander, et autres
Publié: (2025)
Consensus or Conflict? Fine-Grained Evaluation of Conflicting Answers in Question-Answering
par: Nachshoni, Eviatar, et autres
Publié: (2025)
par: Nachshoni, Eviatar, et autres
Publié: (2025)
Evaluating Answer Reranking Strategies in Time-sensitive Question Answering
par: Kardan, Mehmet, et autres
Publié: (2025)
par: Kardan, Mehmet, et autres
Publié: (2025)
RAGalyst: Automated Human-Aligned Agentic Evaluation for Domain-Specific RAG
par: Gao, Joshua, et autres
Publié: (2025)
par: Gao, Joshua, et autres
Publié: (2025)
Answerability in Retrieval-Augmented Open-Domain Question Answering
par: Abdumalikov, Rustam, et autres
Publié: (2024)
par: Abdumalikov, Rustam, et autres
Publié: (2024)
Dr3: Ask Large Language Models Not to Give Off-Topic Answers in Open Domain Multi-Hop Question Answering
par: Gao, Yuan, et autres
Publié: (2024)
par: Gao, Yuan, et autres
Publié: (2024)
Improving the TENOR of Labeling: Re-evaluating Topic Models for Content Analysis
par: Li, Zongxia, et autres
Publié: (2024)
par: Li, Zongxia, et autres
Publié: (2024)
How much reliable is ChatGPT's prediction on Information Extraction under Input Perturbations?
par: Mondal, Ishani, et autres
Publié: (2024)
par: Mondal, Ishani, et autres
Publié: (2024)
Open Domain Question Answering with Conflicting Contexts
par: Liu, Siyi, et autres
Publié: (2024)
par: Liu, Siyi, et autres
Publié: (2024)
KazQAD: Kazakh Open-Domain Question Answering Dataset
par: Yeshpanov, Rustem, et autres
Publié: (2024)
par: Yeshpanov, Rustem, et autres
Publié: (2024)
Generator-Retriever-Generator Approach for Open-Domain Question Answering
par: Abdallah, Abdelrahman, et autres
Publié: (2023)
par: Abdallah, Abdelrahman, et autres
Publié: (2023)
Gaming the Answer Matcher: Examining the Impact of Text Manipulation on Automated Judgment
par: Khatore, Manas, et autres
Publié: (2025)
par: Khatore, Manas, et autres
Publié: (2025)
QPaug: Question and Passage Augmentation for Open-Domain Question Answering of LLMs
par: Kim, Minsang, et autres
Publié: (2024)
par: Kim, Minsang, et autres
Publié: (2024)
ExpertQA: Expert-Curated Questions and Attributed Answers
par: Malaviya, Chaitanya, et autres
Publié: (2023)
par: Malaviya, Chaitanya, et autres
Publié: (2023)
QAEncoder: Towards Aligned Representation Learning in Question Answering Systems
par: Wang, Zhengren, et autres
Publié: (2024)
par: Wang, Zhengren, et autres
Publié: (2024)
NAVIG: Natural Language-guided Analysis with Vision Language Models for Image Geo-localization
par: Zhang, Zheyuan, et autres
Publié: (2025)
par: Zhang, Zheyuan, et autres
Publié: (2025)
Labeled Interactive Topic Models
par: Seelman, Kyle, et autres
Publié: (2023)
par: Seelman, Kyle, et autres
Publié: (2023)
KARL: Knowledge-Aware Retrieval and Representations aid Retention and Learning in Students
par: Shu, Matthew, et autres
Publié: (2024)
par: Shu, Matthew, et autres
Publié: (2024)
Improving Retrieval Augmented Open-Domain Question-Answering with Vectorized Contexts
par: Chen, Zhuo, et autres
Publié: (2024)
par: Chen, Zhuo, et autres
Publié: (2024)
Ask, Answer, and Detect: Role-Playing LLMs for Personality Detection with Question-Conditioned Mixture-of-Experts
par: Lyu, Yifan, et autres
Publié: (2025)
par: Lyu, Yifan, et autres
Publié: (2025)
Documents similaires
-
PEDANTS: Cheap but Effective and Interpretable Answer Equivalence
par: Li, Zongxia, et autres
Publié: (2024) -
How the Advent of Ubiquitous Large Language Models both Stymie and Turbocharge Dynamic Adversarial Question Generation
par: Sung, Yoo Yeon, et autres
Publié: (2024) -
SciDoc2Diagrammer-MAF: Towards Generation of Scientific Diagrams from Documents guided by Multi-Aspect Feedback Refinement
par: Mondal, Ishani, et autres
Publié: (2024) -
Reverse Question Answering: Can an LLM Write a Question so Hard (or Bad) that it Can't Answer?
par: Balepur, Nishant, et autres
Publié: (2024) -
Large Language Models Are Effective Human Annotation Assistants, But Not Good Independent Annotators
par: Gu, Feng, et autres
Publié: (2025)