DOVE: A Large-Scale Multi-Dimensional Predictions Dataset Towards Meaningful LLM Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Habba, Eliya, Arviv, Ofir, Itzhak, Itay, Perlitz, Yotam, Bandel, Elron, Choshen, Leshem, Shmueli-Scheuer, Michal, Stanovsky, Gabriel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration
von: Habba, Eliya, et al.
Veröffentlicht: (2026)
von: Habba, Eliya, et al.
Veröffentlicht: (2026)
Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench
von: Perlitz, Yotam, et al.
Veröffentlicht: (2024)
von: Perlitz, Yotam, et al.
Veröffentlicht: (2024)
Efficient Benchmarking of Language Models
von: Perlitz, Yotam, et al.
Veröffentlicht: (2023)
von: Perlitz, Yotam, et al.
Veröffentlicht: (2023)
ErrorMap and ErrorAtlas: Charting the Failure Landscape of Large Language Models
von: Ashury-Tahan, Shir, et al.
Veröffentlicht: (2026)
von: Ashury-Tahan, Shir, et al.
Veröffentlicht: (2026)
Robustness as an Emergent Property of Task Performance
von: Ashury-Tahan, Shir, et al.
Veröffentlicht: (2026)
von: Ashury-Tahan, Shir, et al.
Veröffentlicht: (2026)
Unitxt: Flexible, Shareable and Reusable Data Preparation and Evaluation for Generative AI
von: Bandel, Elron, et al.
Veröffentlicht: (2024)
von: Bandel, Elron, et al.
Veröffentlicht: (2024)
From Feelings to Metrics: Understanding and Formalizing How Users Vibe-Test LLMs
von: Itzhak, Itay, et al.
Veröffentlicht: (2026)
von: Itzhak, Itay, et al.
Veröffentlicht: (2026)
The Mighty ToRR: A Benchmark for Table Reasoning and Robustness
von: Ashury-Tahan, Shir, et al.
Veröffentlicht: (2025)
von: Ashury-Tahan, Shir, et al.
Veröffentlicht: (2025)
Instructions Shape Production of Language, not Processing
von: Waldis, Andreas, et al.
Veröffentlicht: (2026)
von: Waldis, Andreas, et al.
Veröffentlicht: (2026)
CLEAR: Error Analysis via LLM-as-a-Judge Made Easy
von: Yehudai, Asaf, et al.
Veröffentlicht: (2025)
von: Yehudai, Asaf, et al.
Veröffentlicht: (2025)
General Agent Evaluation
von: Bandel, Elron, et al.
Veröffentlicht: (2026)
von: Bandel, Elron, et al.
Veröffentlicht: (2026)
Beyond Benchmarks: On The False Promise of AI Regulation
von: Stanovsky, Gabriel, et al.
Veröffentlicht: (2025)
von: Stanovsky, Gabriel, et al.
Veröffentlicht: (2025)
PromptSuite: A Task-Agnostic Framework for Multi-Prompt Generation
von: Habba, Eliya, et al.
Veröffentlicht: (2025)
von: Habba, Eliya, et al.
Veröffentlicht: (2025)
ReliableEval: A Recipe for Stochastic LLM Evaluation via Method of Moments
von: Lior, Gili, et al.
Veröffentlicht: (2025)
von: Lior, Gili, et al.
Veröffentlicht: (2025)
Holmes: A Benchmark to Assess the Linguistic Competence of Language Models
von: Waldis, Andreas, et al.
Veröffentlicht: (2024)
von: Waldis, Andreas, et al.
Veröffentlicht: (2024)
A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks
von: Keren, Tomer, et al.
Veröffentlicht: (2026)
von: Keren, Tomer, et al.
Veröffentlicht: (2026)
Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs
von: Itzhak, Itay, et al.
Veröffentlicht: (2025)
von: Itzhak, Itay, et al.
Veröffentlicht: (2025)
Genie: Achieving Human Parity in Content-Grounded Datasets Generation
von: Yehudai, Asaf, et al.
Veröffentlicht: (2024)
von: Yehudai, Asaf, et al.
Veröffentlicht: (2024)
How Safe is Your Safety Metric? Automatic Concatenation Tests for Metric Reliability
von: Fandina, Ora Nova, et al.
Veröffentlicht: (2024)
von: Fandina, Ora Nova, et al.
Veröffentlicht: (2024)
ScheMatiQ: From Research Question to Structured Data through Interactive Schema Discovery
von: Levy, Shahar, et al.
Veröffentlicht: (2026)
von: Levy, Shahar, et al.
Veröffentlicht: (2026)
Instructed to Bias: Instruction-Tuned Language Models Exhibit Emergent Cognitive Bias
von: Itzhak, Itay, et al.
Veröffentlicht: (2023)
von: Itzhak, Itay, et al.
Veröffentlicht: (2023)
Stay Tuned: An Empirical Study of the Impact of Hyperparameters on LLM Tuning in Real-World Applications
von: Halfon, Alon, et al.
Veröffentlicht: (2024)
von: Halfon, Alon, et al.
Veröffentlicht: (2024)
Think Again! The Effect of Test-Time Compute on Preferences, Opinions, and Beliefs of Large Language Models
von: Kour, George, et al.
Veröffentlicht: (2025)
von: Kour, George, et al.
Veröffentlicht: (2025)
Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents
von: Yehudai, Asaf, et al.
Veröffentlicht: (2026)
von: Yehudai, Asaf, et al.
Veröffentlicht: (2026)
Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
A Hitchhiker's Guide to Scaling Law Estimation
von: Choshen, Leshem, et al.
Veröffentlicht: (2024)
von: Choshen, Leshem, et al.
Veröffentlicht: (2024)
JSON Whisperer: Efficient JSON Editing with LLMs
von: Duanis, Sarel, et al.
Veröffentlicht: (2025)
von: Duanis, Sarel, et al.
Veröffentlicht: (2025)
Comparing Humans and Models on a Similar Scale: Towards Cognitive Gender Bias Evaluation in Coreference Resolution
von: Lior, Gili, et al.
Veröffentlicht: (2023)
von: Lior, Gili, et al.
Veröffentlicht: (2023)
Mediocrity is the key for LLM as a Judge Anchor Selection
von: Don-Yehiya, Shachar, et al.
Veröffentlicht: (2026)
von: Don-Yehiya, Shachar, et al.
Veröffentlicht: (2026)
Can Gradient Descent Simulate Prompting?
von: Zhang, Eric, et al.
Veröffentlicht: (2025)
von: Zhang, Eric, et al.
Veröffentlicht: (2025)
Fuse to Forget: Bias Reduction and Selective Memorization through Model Fusion
von: Zaman, Kerem, et al.
Veröffentlicht: (2023)
von: Zaman, Kerem, et al.
Veröffentlicht: (2023)
The State and Fate of Summarization Datasets: A Survey
von: Dahan, Noam, et al.
Veröffentlicht: (2024)
von: Dahan, Noam, et al.
Veröffentlicht: (2024)
NeurIPS 2023 LLM Efficiency Fine-tuning Competition
von: Saroufim, Mark, et al.
Veröffentlicht: (2025)
von: Saroufim, Mark, et al.
Veröffentlicht: (2025)
PolySQL: Scaling Text-to-SQL Evaluation Across SQL Dialects via Automated Backend Isomorphism
von: Perlitz, Yotam, et al.
Veröffentlicht: (2026)
von: Perlitz, Yotam, et al.
Veröffentlicht: (2026)
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users
von: Hilel, Almog, et al.
Veröffentlicht: (2025)
von: Hilel, Almog, et al.
Veröffentlicht: (2025)
The ShareLM Collection and Plugin: Contributing Human-Model Chats for the Benefit of the Community
von: Don-Yehiya, Shachar, et al.
Veröffentlicht: (2024)
von: Don-Yehiya, Shachar, et al.
Veröffentlicht: (2024)
Naturally Occurring Feedback is Common, Extractable and Useful
von: Don-Yehiya, Shachar, et al.
Veröffentlicht: (2024)
von: Don-Yehiya, Shachar, et al.
Veröffentlicht: (2024)
Neural Minimum Weight Perfect Matching for Quantum Error Codes
von: Peled, Yotam, et al.
Veröffentlicht: (2026)
von: Peled, Yotam, et al.
Veröffentlicht: (2026)
Active Speech Enhancement: Active Speech Denoising Decliping and Deveraberation
von: Yaish, Ofir, et al.
Veröffentlicht: (2025)
von: Yaish, Ofir, et al.
Veröffentlicht: (2025)
Sloth: scaling laws for LLM skills to predict multi-benchmark performance across families
von: Polo, Felipe Maia, et al.
Veröffentlicht: (2024)
von: Polo, Felipe Maia, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration
von: Habba, Eliya, et al.
Veröffentlicht: (2026) -
Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench
von: Perlitz, Yotam, et al.
Veröffentlicht: (2024) -
Efficient Benchmarking of Language Models
von: Perlitz, Yotam, et al.
Veröffentlicht: (2023) -
ErrorMap and ErrorAtlas: Charting the Failure Landscape of Large Language Models
von: Ashury-Tahan, Shir, et al.
Veröffentlicht: (2026) -
Robustness as an Emergent Property of Task Performance
von: Ashury-Tahan, Shir, et al.
Veröffentlicht: (2026)