Answer Matching Outperforms Multiple Choice for Language Model Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chandak, Nikhil, Goel, Shashwat, Prabhu, Ameya, Hardt, Moritz, Geiping, Jonas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scaling Open-Ended Reasoning to Predict the Future
von: Chandak, Nikhil, et al.
Veröffentlicht: (2025)
von: Chandak, Nikhil, et al.
Veröffentlicht: (2025)
FutureSim: Replaying World Events to Evaluate Adaptive Agents
von: Goel, Shashwat, et al.
Veröffentlicht: (2026)
von: Goel, Shashwat, et al.
Veröffentlicht: (2026)
Great Models Think Alike and this Undermines AI Oversight
von: Goel, Shashwat, et al.
Veröffentlicht: (2025)
von: Goel, Shashwat, et al.
Veröffentlicht: (2025)
Pitfalls in Evaluating Language Model Forecasters
von: Paleka, Daniel, et al.
Veröffentlicht: (2025)
von: Paleka, Daniel, et al.
Veröffentlicht: (2025)
Wu's Method can Boost Symbolic AI to Rival Silver Medalists and AlphaGeometry to Outperform Gold Medalists at IMO Geometry
von: Sinha, Shiven, et al.
Veröffentlicht: (2024)
von: Sinha, Shiven, et al.
Veröffentlicht: (2024)
Mapping Post-Training Forgetting in Language Models at Scale
von: Harmon, Jackson, et al.
Veröffentlicht: (2025)
von: Harmon, Jackson, et al.
Veröffentlicht: (2025)
Efficiently Dispatching Flash Attention For Partially Filled Attention Masks
von: Sharma, Agniv, et al.
Veröffentlicht: (2024)
von: Sharma, Agniv, et al.
Veröffentlicht: (2024)
Can Language Models Falsify? Evaluating Algorithmic Reasoning with Counterexample Creation
von: Sinha, Shiven, et al.
Veröffentlicht: (2025)
von: Sinha, Shiven, et al.
Veröffentlicht: (2025)
Training on the Test Task Confounds Evaluation and Emergence
von: Dominguez-Olmedo, Ricardo, et al.
Veröffentlicht: (2024)
von: Dominguez-Olmedo, Ricardo, et al.
Veröffentlicht: (2024)
Anchored Answers: Unravelling Positional Bias in GPT-2's Multiple-Choice Questions
von: Li, Ruizhe, et al.
Veröffentlicht: (2024)
von: Li, Ruizhe, et al.
Veröffentlicht: (2024)
Luna-2: Scalable Single-Token Evaluation with Small Language Models
von: Goel, Vatsal, et al.
Veröffentlicht: (2026)
von: Goel, Vatsal, et al.
Veröffentlicht: (2026)
An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks
von: Boreiko, Valentyn, et al.
Veröffentlicht: (2024)
von: Boreiko, Valentyn, et al.
Veröffentlicht: (2024)
Corrective Machine Unlearning
von: Goel, Shashwat, et al.
Veröffentlicht: (2024)
von: Goel, Shashwat, et al.
Veröffentlicht: (2024)
Multiple Choice Learning of Low-Rank Adapters for Language Modeling
von: Letzelter, Victor, et al.
Veröffentlicht: (2025)
von: Letzelter, Victor, et al.
Veröffentlicht: (2025)
Scaling Embeddings Outperforms Scaling Experts in Language Models
von: Liu, Hong, et al.
Veröffentlicht: (2026)
von: Liu, Hong, et al.
Veröffentlicht: (2026)
A Study on Large Language Models' Limitations in Multiple-Choice Question Answering
von: Khatun, Aisha, et al.
Veröffentlicht: (2024)
von: Khatun, Aisha, et al.
Veröffentlicht: (2024)
LLM generation novelty through the lens of semantic similarity
von: Davydov, Philipp, et al.
Veröffentlicht: (2025)
von: Davydov, Philipp, et al.
Veröffentlicht: (2025)
Proportional Aggregation of Preferences for Sequential Decision Making
von: Chandak, Nikhil, et al.
Veröffentlicht: (2023)
von: Chandak, Nikhil, et al.
Veröffentlicht: (2023)
Capability-Based Scaling Trends for LLM-Based Red-Teaming
von: Panfilov, Alexander, et al.
Veröffentlicht: (2025)
von: Panfilov, Alexander, et al.
Veröffentlicht: (2025)
Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence
von: McLeish, Sean, et al.
Veröffentlicht: (2025)
von: McLeish, Sean, et al.
Veröffentlicht: (2025)
LiveOIBench: Can Large Language Models Outperform Human Contestants in Informatics Olympiads?
von: Zou, Kaijian, et al.
Veröffentlicht: (2025)
von: Zou, Kaijian, et al.
Veröffentlicht: (2025)
Generating Plausible Distractors for Multiple-Choice Questions via Student Choice Prediction
von: Lee, Yooseop, et al.
Veröffentlicht: (2025)
von: Lee, Yooseop, et al.
Veröffentlicht: (2025)
Generating Multiple-Choice Knowledge Questions with Interpretable Difficulty Estimation using Knowledge Graphs and Large Language Models
von: Şakiroğlu, Mehmet Can, et al.
Veröffentlicht: (2026)
von: Şakiroğlu, Mehmet Can, et al.
Veröffentlicht: (2026)
Intrinsic Credit Assignment for Long Horizon Interaction
von: Auzina, Ilze Amanda, et al.
Veröffentlicht: (2026)
von: Auzina, Ilze Amanda, et al.
Veröffentlicht: (2026)
Towards Acyclic Preference Evaluation of Language Models via Multiple Evaluators
von: Hu, Zhengyu, et al.
Veröffentlicht: (2024)
von: Hu, Zhengyu, et al.
Veröffentlicht: (2024)
Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation
von: Zhang, Yuhui, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhui, et al.
Veröffentlicht: (2025)
Option-ID Based Elimination For Multiple Choice Questions
von: Zhu, Zhenhao, et al.
Veröffentlicht: (2025)
von: Zhu, Zhenhao, et al.
Veröffentlicht: (2025)
Biomedical Entity Linking as Multiple Choice Question Answering
von: Lin, Zhenxi, et al.
Veröffentlicht: (2024)
von: Lin, Zhenxi, et al.
Veröffentlicht: (2024)
Rewarding Intellectual Humility Learning When Not To Answer In Large Language Models
von: Jha, Abha, et al.
Veröffentlicht: (2026)
von: Jha, Abha, et al.
Veröffentlicht: (2026)
Specialised or Generic? Tokenization Choices for Radiology Language Models
von: Warr, Hermione, et al.
Veröffentlicht: (2025)
von: Warr, Hermione, et al.
Veröffentlicht: (2025)
Learning to Interpret Weight Differences in Language Models
von: Goel, Avichal, et al.
Veröffentlicht: (2025)
von: Goel, Avichal, et al.
Veröffentlicht: (2025)
Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
von: Li, Kenneth, et al.
Veröffentlicht: (2023)
von: Li, Kenneth, et al.
Veröffentlicht: (2023)
Explicit Diversity Conditions for Effective Question Answer Generation with Large Language Models
von: Yadav, Vikas, et al.
Veröffentlicht: (2024)
von: Yadav, Vikas, et al.
Veröffentlicht: (2024)
Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text
von: Hans, Abhimanyu, et al.
Veröffentlicht: (2024)
von: Hans, Abhimanyu, et al.
Veröffentlicht: (2024)
On The Truthfulness of 'Surprisingly Likely' Responses of Large Language Models
von: Goel, Naman
Veröffentlicht: (2023)
von: Goel, Naman
Veröffentlicht: (2023)
First is Not Really Better Than Last: Evaluating Layer Choice and Aggregation Strategies in Language Model Data Influence Estimation
von: Vitel, Dmytro, et al.
Veröffentlicht: (2025)
von: Vitel, Dmytro, et al.
Veröffentlicht: (2025)
Pattern Recognition or Medical Knowledge? The Problem with Multiple-Choice Questions in Medicine
von: Griot, Maxime, et al.
Veröffentlicht: (2024)
von: Griot, Maxime, et al.
Veröffentlicht: (2024)
Promote, Suppress, Iterate: How Language Models Answer One-to-Many Factual Queries
von: Yan, Tianyi Lorena, et al.
Veröffentlicht: (2025)
von: Yan, Tianyi Lorena, et al.
Veröffentlicht: (2025)
Test-Time Training on Nearest Neighbors for Large Language Models
von: Hardt, Moritz, et al.
Veröffentlicht: (2023)
von: Hardt, Moritz, et al.
Veröffentlicht: (2023)
Improving Score Reliability of Multiple Choice Benchmarks with Consistency Evaluation and Altered Answer Choices
von: Cavalin, Paulo, et al.
Veröffentlicht: (2025)
von: Cavalin, Paulo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Scaling Open-Ended Reasoning to Predict the Future
von: Chandak, Nikhil, et al.
Veröffentlicht: (2025) -
FutureSim: Replaying World Events to Evaluate Adaptive Agents
von: Goel, Shashwat, et al.
Veröffentlicht: (2026) -
Great Models Think Alike and this Undermines AI Oversight
von: Goel, Shashwat, et al.
Veröffentlicht: (2025) -
Pitfalls in Evaluating Language Model Forecasters
von: Paleka, Daniel, et al.
Veröffentlicht: (2025) -
Wu's Method can Boost Symbolic AI to Rival Silver Medalists and AlphaGeometry to Outperform Gold Medalists at IMO Geometry
von: Sinha, Shiven, et al.
Veröffentlicht: (2024)