Are We Done with MMLU?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gema, Aryo Pradipta, Leang, Joshua Ong Jun, Hong, Giwon, Devoto, Alessio, Mancino, Alberto Carlo Maria, Saxena, Rohit, He, Xuanli, Zhao, Yu, Du, Xiaotang, Madani, Mohammad Reza Ghasemi, Barale, Claire, McHardy, Robert, Harris, Joshua, Kaddour, Jean, van Krieken, Emile, Minervini, Pasquale |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs
von: Saxena, Rohit, et al.
Veröffentlicht: (2025)
von: Saxena, Rohit, et al.
Veröffentlicht: (2025)
SCOPE: Self-Play via Co-Evolving Policies for Open-Ended Tasks
von: Kwan, Wai-Chung, et al.
Veröffentlicht: (2026)
von: Kwan, Wai-Chung, et al.
Veröffentlicht: (2026)
Self-Training Large Language Models for Tool-Use Without Demonstrations
von: Luo, Ne, et al.
Veröffentlicht: (2025)
von: Luo, Ne, et al.
Veröffentlicht: (2025)
Analysing the Residual Stream of Language Models Under Knowledge Conflicts
von: Zhao, Yu, et al.
Veröffentlicht: (2024)
von: Zhao, Yu, et al.
Veröffentlicht: (2024)
Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering
von: Zhao, Yu, et al.
Veröffentlicht: (2024)
von: Zhao, Yu, et al.
Veröffentlicht: (2024)
GRADA: Graph-based Reranking against Adversarial Documents Attack
von: Zheng, Jingjie, et al.
Veröffentlicht: (2025)
von: Zheng, Jingjie, et al.
Veröffentlicht: (2025)
CoMAT: Chain of Mathematically Annotated Thought Improves Mathematical Reasoning
von: Leang, Joshua Ong Jun, et al.
Veröffentlicht: (2024)
von: Leang, Joshua Ong Jun, et al.
Veröffentlicht: (2024)
The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models
von: Hong, Giwon, et al.
Veröffentlicht: (2024)
von: Hong, Giwon, et al.
Veröffentlicht: (2024)
Edinburgh Clinical NLP at SemEval-2024 Task 2: Fine-tune your model unless you have access to GPT-4
von: Gema, Aryo Pradipta, et al.
Veröffentlicht: (2024)
von: Gema, Aryo Pradipta, et al.
Veröffentlicht: (2024)
Noiser: Bounded Input Perturbations for Attributing Large Language Models
von: Madani, Mohammad Reza Ghasemi, et al.
Veröffentlicht: (2025)
von: Madani, Mohammad Reza Ghasemi, et al.
Veröffentlicht: (2025)
PiCSAR: Probabilistic Confidence Selection And Ranking for Reasoning Chains
von: Leang, Joshua Ong Jun, et al.
Veröffentlicht: (2025)
von: Leang, Joshua Ong Jun, et al.
Veröffentlicht: (2025)
Mixtures of In-Context Learners
von: Hong, Giwon, et al.
Veröffentlicht: (2024)
von: Hong, Giwon, et al.
Veröffentlicht: (2024)
Parameter-Efficient Fine-Tuning of LLaMA for the Clinical Domain
von: Gema, Aryo Pradipta, et al.
Veröffentlicht: (2023)
von: Gema, Aryo Pradipta, et al.
Veröffentlicht: (2023)
Analyzing LLM Instruction Optimization for Tabular Fact Verification
von: Du, Xiaotang, et al.
Veröffentlicht: (2026)
von: Du, Xiaotang, et al.
Veröffentlicht: (2026)
Edinburgh Clinical NLP at MEDIQA-CORR 2024: Guiding Large Language Models with Hints
von: Gema, Aryo Pradipta, et al.
Veröffentlicht: (2024)
von: Gema, Aryo Pradipta, et al.
Veröffentlicht: (2024)
Neurosymbolic Diffusion Models
von: van Krieken, Emile, et al.
Veröffentlicht: (2025)
von: van Krieken, Emile, et al.
Veröffentlicht: (2025)
Neurosymbolic Reasoning Shortcuts under the Independence Assumption
von: van Krieken, Emile, et al.
Veröffentlicht: (2025)
von: van Krieken, Emile, et al.
Veröffentlicht: (2025)
Enhancing Long Document Long Form Summarisation with Self-Planning
von: Du, Xiaotang, et al.
Veröffentlicht: (2025)
von: Du, Xiaotang, et al.
Veröffentlicht: (2025)
Theorem Prover as a Judge for Synthetic Data Generation
von: Leang, Joshua Ong Jun, et al.
Veröffentlicht: (2025)
von: Leang, Joshua Ong Jun, et al.
Veröffentlicht: (2025)
On the Independence Assumption in Neurosymbolic Learning
von: van Krieken, Emile, et al.
Veröffentlicht: (2024)
von: van Krieken, Emile, et al.
Veröffentlicht: (2024)
Same Answer, Different Representations: Hidden instability in VLMs
von: Wani, Farooq Ahmad, et al.
Veröffentlicht: (2026)
von: Wani, Farooq Ahmad, et al.
Veröffentlicht: (2026)
VLM-RobustBench: A Comprehensive Benchmark for Robustness of Vision-Language Models
von: Saxena, Rohit, et al.
Veröffentlicht: (2026)
von: Saxena, Rohit, et al.
Veröffentlicht: (2026)
PosterSum: A Multimodal Benchmark for Scientific Poster Summarization
von: Saxena, Rohit, et al.
Veröffentlicht: (2025)
von: Saxena, Rohit, et al.
Veröffentlicht: (2025)
DeCoRe: Decoding by Contrasting Retrieval Heads to Mitigate Hallucinations
von: Gema, Aryo Pradipta, et al.
Veröffentlicht: (2024)
von: Gema, Aryo Pradipta, et al.
Veröffentlicht: (2024)
An Analysis of Decoding Methods for LLM-based Agents for Faithful Multi-Hop Question Answering
von: Murphy, Alexander, et al.
Veröffentlicht: (2025)
von: Murphy, Alexander, et al.
Veröffentlicht: (2025)
OpenSIR: Open-Ended Self-Improving Reasoner
von: Kwan, Wai-Chung, et al.
Veröffentlicht: (2025)
von: Kwan, Wai-Chung, et al.
Veröffentlicht: (2025)
Do Composed Image Retrieval Benchmarks Require Multimodal Composition?
von: Attimonelli, Matteo, et al.
Veröffentlicht: (2026)
von: Attimonelli, Matteo, et al.
Veröffentlicht: (2026)
Optimisation in Neurosymbolic Learning Systems
von: van Krieken, Emile
Veröffentlicht: (2024)
von: van Krieken, Emile
Veröffentlicht: (2024)
Attention Is All You Need But You Don't Need All Of It For Inference of Large Language Models
von: Tyukin, Georgy, et al.
Veröffentlicht: (2024)
von: Tyukin, Georgy, et al.
Veröffentlicht: (2024)
An Auditing Test To Detect Behavioral Shift in Language Models
von: Richter, Leo, et al.
Veröffentlicht: (2024)
von: Richter, Leo, et al.
Veröffentlicht: (2024)
A Simple and Effective $L_2$ Norm-Based Strategy for KV Cache Compression
von: Devoto, Alessio, et al.
Veröffentlicht: (2024)
von: Devoto, Alessio, et al.
Veröffentlicht: (2024)
Scalpel vs. Hammer: GRPO Amplifies Existing Capabilities, SFT Replaces Them
von: Rajani, Neel, et al.
Veröffentlicht: (2025)
von: Rajani, Neel, et al.
Veröffentlicht: (2025)
NAFTA, Mexico and the China factor / David McHardy Reid, Alethia Jimenez and Peter Rahmer
von: McHardy Reid, David
von: McHardy Reid, David
Intellectual property rights : a comparative perspective on Asia, the EU, and North America / David McHardy Reid
von: McHardy Reid, David
Veröffentlicht: (2012)
von: McHardy Reid, David
Veröffentlicht: (2012)
Intellectual Property Rights: A Comparative Perspective on Asia, the EU, and North America
von: David McHardy Reid
Veröffentlicht: (2012)
von: David McHardy Reid
Veröffentlicht: (2012)
Gradient-Based Optimization on Gödel Logic as Discrete Local Search
von: Daniele, Alessandro, et al.
Veröffentlicht: (2025)
von: Daniele, Alessandro, et al.
Veröffentlicht: (2025)
Using Natural Language Explanations to Improve Robustness of In-context Learning
von: He, Xuanli, et al.
Veröffentlicht: (2023)
von: He, Xuanli, et al.
Veröffentlicht: (2023)
Dot Product is All You Need: Bridging the Gap Between Item Recommendation and Link Prediction
von: Malitesta, Daniele, et al.
Veröffentlicht: (2024)
von: Malitesta, Daniele, et al.
Veröffentlicht: (2024)
The Hot Mess of AI: How Does Misalignment Scale With Model Intelligence and Task Complexity?
von: Hägele, Alexander, et al.
Veröffentlicht: (2026)
von: Hägele, Alexander, et al.
Veröffentlicht: (2026)
Adaptive Computation Modules: Granular Conditional Computation For Efficient Inference
von: Wójcik, Bartosz, et al.
Veröffentlicht: (2023)
von: Wójcik, Bartosz, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs
von: Saxena, Rohit, et al.
Veröffentlicht: (2025) -
SCOPE: Self-Play via Co-Evolving Policies for Open-Ended Tasks
von: Kwan, Wai-Chung, et al.
Veröffentlicht: (2026) -
Self-Training Large Language Models for Tool-Use Without Demonstrations
von: Luo, Ne, et al.
Veröffentlicht: (2025) -
Analysing the Residual Stream of Language Models Under Knowledge Conflicts
von: Zhao, Yu, et al.
Veröffentlicht: (2024) -
Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering
von: Zhao, Yu, et al.
Veröffentlicht: (2024)