PEDANTS: Cheap but Effective and Interpretable Answer Equivalence
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Zongxia, Mondal, Ishani, Liang, Yijun, Nghiem, Huy, Boyd-Graber, Jordan Lee |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CFMatch: Aligning Automated Answer Equivalence Evaluation with Expert Judgments For Open-Domain Question Answering
di: Li, Zongxia, et al.
Pubblicazione: (2024)
di: Li, Zongxia, et al.
Pubblicazione: (2024)
Large Language Models Are Effective Human Annotation Assistants, But Not Good Independent Annotators
di: Gu, Feng, et al.
Pubblicazione: (2025)
di: Gu, Feng, et al.
Pubblicazione: (2025)
SMART-Editor: A Multi-Agent Framework for Human-Like Design Editing with Structural Integrity
di: Mondal, Ishani, et al.
Pubblicazione: (2025)
di: Mondal, Ishani, et al.
Pubblicazione: (2025)
SciDoc2Diagrammer-MAF: Towards Generation of Scientific Diagrams from Documents guided by Multi-Aspect Feedback Refinement
di: Mondal, Ishani, et al.
Pubblicazione: (2024)
di: Mondal, Ishani, et al.
Pubblicazione: (2024)
How the Advent of Ubiquitous Large Language Models both Stymie and Turbocharge Dynamic Adversarial Question Generation
di: Sung, Yoo Yeon, et al.
Pubblicazione: (2024)
di: Sung, Yoo Yeon, et al.
Pubblicazione: (2024)
A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges
di: Li, Zongxia, et al.
Pubblicazione: (2025)
di: Li, Zongxia, et al.
Pubblicazione: (2025)
Is your benchmark truly adversarial? AdvScore: Evaluating Human-Grounded Adversarialness
di: Sung, Yoo Yeon, et al.
Pubblicazione: (2024)
di: Sung, Yoo Yeon, et al.
Pubblicazione: (2024)
SMARTER: A Data-efficient Framework to Improve Toxicity Detection with Explanation via Self-augmenting Large Language Models
di: Nghiem, Huy, et al.
Pubblicazione: (2025)
di: Nghiem, Huy, et al.
Pubblicazione: (2025)
Do great minds think alike? Investigating Human-AI Complementarity in Question Answering with CAIMIRA
di: Gor, Maharshi, et al.
Pubblicazione: (2024)
di: Gor, Maharshi, et al.
Pubblicazione: (2024)
HateCOT: An Explanation-Enhanced Dataset for Generalizable Offensive Speech Detection via Large Language Models
di: Nghiem, Huy, et al.
Pubblicazione: (2024)
di: Nghiem, Huy, et al.
Pubblicazione: (2024)
"You Gotta be a Doctor, Lin": An Investigation of Name-Based Bias of Large Language Models in Employment Recommendations
di: Nghiem, Huy, et al.
Pubblicazione: (2024)
di: Nghiem, Huy, et al.
Pubblicazione: (2024)
Discrepancy Detection at the Data Level: Toward Consistent Multilingual Question Answering
di: Calvo-Bartolomé, Lorena, et al.
Pubblicazione: (2025)
di: Calvo-Bartolomé, Lorena, et al.
Pubblicazione: (2025)
Balancing Safety and Helpfulness in Healthcare AI Assistants through Iterative Preference Alignment
di: Nghiem, Huy, et al.
Pubblicazione: (2025)
di: Nghiem, Huy, et al.
Pubblicazione: (2025)
CANVAS: Continuity-Aware Narratives via Visual Agentic Storyboarding
di: Mondal, Ishani, et al.
Pubblicazione: (2026)
di: Mondal, Ishani, et al.
Pubblicazione: (2026)
AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?
di: Gor, Maharshi, et al.
Pubblicazione: (2026)
di: Gor, Maharshi, et al.
Pubblicazione: (2026)
Bias in the Tails: How Name-conditioned Evaluative Framing in Resume Summaries Destabilizes LLM-based Hiring
di: Nghiem, Huy, et al.
Pubblicazione: (2026)
di: Nghiem, Huy, et al.
Pubblicazione: (2026)
Implicit Probabilistic Reasoning Does Not Reflect Explicit Answers in Large Language Models
di: Mondal, Manuel, et al.
Pubblicazione: (2024)
di: Mondal, Manuel, et al.
Pubblicazione: (2024)
I've got the "Answer"! Interpretation of LLMs Hidden States in Question Answering
di: Goloviznina, Valeriya, et al.
Pubblicazione: (2024)
di: Goloviznina, Valeriya, et al.
Pubblicazione: (2024)
Integrated Framework for LLM Evaluation with Answer Generation
di: Lee, Sujeong, et al.
Pubblicazione: (2025)
di: Lee, Sujeong, et al.
Pubblicazione: (2025)
Reasoning or Retrieval? A Study of Answer Attribution on Large Reasoning Models
di: Wang, Yuhui, et al.
Pubblicazione: (2025)
di: Wang, Yuhui, et al.
Pubblicazione: (2025)
Towards Understanding In-Context Learning with Contrastive Demonstrations and Saliency Maps
di: Liu, Fuxiao, et al.
Pubblicazione: (2023)
di: Liu, Fuxiao, et al.
Pubblicazione: (2023)
CORG: Generating Answers from Complex, Interrelated Contexts
di: Lee, Hyunji, et al.
Pubblicazione: (2025)
di: Lee, Hyunji, et al.
Pubblicazione: (2025)
Finding Answers in Thought Matters: Revisiting Evaluation on Large Language Models with Reasoning
di: Jo, Hwiyeol, et al.
Pubblicazione: (2025)
di: Jo, Hwiyeol, et al.
Pubblicazione: (2025)
Beyond Answers: Transferring Reasoning Capabilities to Smaller LLMs Using Multi-Teacher Knowledge Distillation
di: Tian, Yijun, et al.
Pubblicazione: (2024)
di: Tian, Yijun, et al.
Pubblicazione: (2024)
Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs
di: Kossen, Jannik, et al.
Pubblicazione: (2024)
di: Kossen, Jannik, et al.
Pubblicazione: (2024)
ExpertQA: Expert-Curated Questions and Attributed Answers
di: Malaviya, Chaitanya, et al.
Pubblicazione: (2023)
di: Malaviya, Chaitanya, et al.
Pubblicazione: (2023)
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation
di: Li, Zongxia, et al.
Pubblicazione: (2025)
di: Li, Zongxia, et al.
Pubblicazione: (2025)
Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions
di: Wiegreffe, Sarah, et al.
Pubblicazione: (2024)
di: Wiegreffe, Sarah, et al.
Pubblicazione: (2024)
Sandwich Reasoning: An Answer-Reasoning-Answer Approach for Low-Latency Query Correction
di: Zhang, Chen, et al.
Pubblicazione: (2026)
di: Zhang, Chen, et al.
Pubblicazione: (2026)
VietMix: A Naturally-Occurring Parallel Corpus and Augmentation Framework for Vietnamese-English Code-Mixed Machine Translation
di: Tran, Hieu, et al.
Pubblicazione: (2025)
di: Tran, Hieu, et al.
Pubblicazione: (2025)
Latent Self-Consistency for Reliable Majority-Set Selection in Short- and Long-Answer Reasoning
di: Oh, Jungsuk, et al.
Pubblicazione: (2025)
di: Oh, Jungsuk, et al.
Pubblicazione: (2025)
OpenSeal: Good, Fast, and Cheap Construction of an Open-Source Southeast Asian LLM via Parallel Data
di: Nguyen, Tan Sang, et al.
Pubblicazione: (2026)
di: Nguyen, Tan Sang, et al.
Pubblicazione: (2026)
Counterfactual Probing for Hallucination Detection and Mitigation in Large Language Models
di: Feng, Yijun
Pubblicazione: (2025)
di: Feng, Yijun
Pubblicazione: (2025)
Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs through a Global Scale Prompt Hacking Competition
di: Schulhoff, Sander, et al.
Pubblicazione: (2023)
di: Schulhoff, Sander, et al.
Pubblicazione: (2023)
No Answer Needed: Predicting LLM Answer Accuracy from Question-Only Linear Probes
di: Cencerrado, Iván Vicente Moreno, et al.
Pubblicazione: (2025)
di: Cencerrado, Iván Vicente Moreno, et al.
Pubblicazione: (2025)
Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents
di: Kim, Wonjoong, et al.
Pubblicazione: (2025)
di: Kim, Wonjoong, et al.
Pubblicazione: (2025)
Diffusion Language Models Know the Answer Before Decoding
di: Li, Pengxiang, et al.
Pubblicazione: (2025)
di: Li, Pengxiang, et al.
Pubblicazione: (2025)
Quality of Answers of Generative Large Language Models vs Peer Patients for Interpreting Lab Test Results for Lay Patients: Evaluation Study
di: He, Zhe, et al.
Pubblicazione: (2024)
di: He, Zhe, et al.
Pubblicazione: (2024)
Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above
di: Balepur, Nishant, et al.
Pubblicazione: (2025)
di: Balepur, Nishant, et al.
Pubblicazione: (2025)
Explicit Diversity Conditions for Effective Question Answer Generation with Large Language Models
di: Yadav, Vikas, et al.
Pubblicazione: (2024)
di: Yadav, Vikas, et al.
Pubblicazione: (2024)
Documenti analoghi
-
CFMatch: Aligning Automated Answer Equivalence Evaluation with Expert Judgments For Open-Domain Question Answering
di: Li, Zongxia, et al.
Pubblicazione: (2024) -
Large Language Models Are Effective Human Annotation Assistants, But Not Good Independent Annotators
di: Gu, Feng, et al.
Pubblicazione: (2025) -
SMART-Editor: A Multi-Agent Framework for Human-Like Design Editing with Structural Integrity
di: Mondal, Ishani, et al.
Pubblicazione: (2025) -
SciDoc2Diagrammer-MAF: Towards Generation of Scientific Diagrams from Documents guided by Multi-Aspect Feedback Refinement
di: Mondal, Ishani, et al.
Pubblicazione: (2024) -
How the Advent of Ubiquitous Large Language Models both Stymie and Turbocharge Dynamic Adversarial Question Generation
di: Sung, Yoo Yeon, et al.
Pubblicazione: (2024)