A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Keren, Tomer, Calderon, Nitay, Yehudai, Asaf, Perlitz, Yotam, Shmueli-Scheuer, Michal, Reichert, Roi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CLEAR: Error Analysis via LLM-as-a-Judge Made Easy
von: Yehudai, Asaf, et al.
Veröffentlicht: (2025)
von: Yehudai, Asaf, et al.
Veröffentlicht: (2025)
Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents
von: Yehudai, Asaf, et al.
Veröffentlicht: (2026)
von: Yehudai, Asaf, et al.
Veröffentlicht: (2026)
General Agent Evaluation
von: Bandel, Elron, et al.
Veröffentlicht: (2026)
von: Bandel, Elron, et al.
Veröffentlicht: (2026)
JuStRank: Benchmarking LLM Judges for System Ranking
von: Gera, Ariel, et al.
Veröffentlicht: (2024)
von: Gera, Ariel, et al.
Veröffentlicht: (2024)
Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench
von: Perlitz, Yotam, et al.
Veröffentlicht: (2024)
von: Perlitz, Yotam, et al.
Veröffentlicht: (2024)
Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration
von: Habba, Eliya, et al.
Veröffentlicht: (2026)
von: Habba, Eliya, et al.
Veröffentlicht: (2026)
Survey on Evaluation of LLM-based Agents
von: Yehudai, Asaf, et al.
Veröffentlicht: (2025)
von: Yehudai, Asaf, et al.
Veröffentlicht: (2025)
On Behalf of the Stakeholders: Trends in NLP Model Interpretability in the Era of LLMs
von: Calderon, Nitay, et al.
Veröffentlicht: (2024)
von: Calderon, Nitay, et al.
Veröffentlicht: (2024)
Efficient Benchmarking of Language Models
von: Perlitz, Yotam, et al.
Veröffentlicht: (2023)
von: Perlitz, Yotam, et al.
Veröffentlicht: (2023)
LIBERTy: A Causal Framework for Benchmarking Concept-Based Explanations of LLMs with Structural Counterfactuals
von: Toker, Gilat, et al.
Veröffentlicht: (2026)
von: Toker, Gilat, et al.
Veröffentlicht: (2026)
ELT-Bench-Verified: Benchmark Quality Issues Underestimate AI Agent Capabilities
von: Zanoli, Christopher, et al.
Veröffentlicht: (2026)
von: Zanoli, Christopher, et al.
Veröffentlicht: (2026)
The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs
von: Calderon, Nitay, et al.
Veröffentlicht: (2025)
von: Calderon, Nitay, et al.
Veröffentlicht: (2025)
The Mighty ToRR: A Benchmark for Table Reasoning and Robustness
von: Ashury-Tahan, Shir, et al.
Veröffentlicht: (2025)
von: Ashury-Tahan, Shir, et al.
Veröffentlicht: (2025)
Unitxt: Flexible, Shareable and Reusable Data Preparation and Evaluation for Generative AI
von: Bandel, Elron, et al.
Veröffentlicht: (2024)
von: Bandel, Elron, et al.
Veröffentlicht: (2024)
Think Again! The Effect of Test-Time Compute on Preferences, Opinions, and Beliefs of Large Language Models
von: Kour, George, et al.
Veröffentlicht: (2025)
von: Kour, George, et al.
Veröffentlicht: (2025)
CUBE: A Standard for Unifying Agent Benchmarks
von: Lacoste, Alexandre, et al.
Veröffentlicht: (2026)
von: Lacoste, Alexandre, et al.
Veröffentlicht: (2026)
When LLMs are Unfit Use FastFit: Fast and Effective Text Classification with Many Classes
von: Yehudai, Asaf, et al.
Veröffentlicht: (2024)
von: Yehudai, Asaf, et al.
Veröffentlicht: (2024)
Teaching Values to Machines: Simulating Human-Like Behavior in LLMs
von: Yehudai, Asaf, et al.
Veröffentlicht: (2026)
von: Yehudai, Asaf, et al.
Veröffentlicht: (2026)
The Colorful Future of LLMs: Evaluating and Improving LLMs as Emotional Supporters for Queer Youth
von: Lissak, Shir, et al.
Veröffentlicht: (2024)
von: Lissak, Shir, et al.
Veröffentlicht: (2024)
ErrorMap and ErrorAtlas: Charting the Failure Landscape of Large Language Models
von: Ashury-Tahan, Shir, et al.
Veröffentlicht: (2026)
von: Ashury-Tahan, Shir, et al.
Veröffentlicht: (2026)
Robustness as an Emergent Property of Task Performance
von: Ashury-Tahan, Shir, et al.
Veröffentlicht: (2026)
von: Ashury-Tahan, Shir, et al.
Veröffentlicht: (2026)
NL-Eye: Abductive NLI for Images
von: Ventura, Mor, et al.
Veröffentlicht: (2024)
von: Ventura, Mor, et al.
Veröffentlicht: (2024)
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?
von: Wang, Leyao, et al.
Veröffentlicht: (2026)
von: Wang, Leyao, et al.
Veröffentlicht: (2026)
WildIFEval: Instruction Following in the Wild
von: Lior, Gili, et al.
Veröffentlicht: (2025)
von: Lior, Gili, et al.
Veröffentlicht: (2025)
DOVE: A Large-Scale Multi-Dimensional Predictions Dataset Towards Meaningful LLM Evaluation
von: Habba, Eliya, et al.
Veröffentlicht: (2025)
von: Habba, Eliya, et al.
Veröffentlicht: (2025)
How Safe is Your Safety Metric? Automatic Concatenation Tests for Metric Reliability
von: Fandina, Ora Nova, et al.
Veröffentlicht: (2024)
von: Fandina, Ora Nova, et al.
Veröffentlicht: (2024)
Selective Self-Rehearsal: A Fine-Tuning Approach to Improve Generalization in Large Language Models
von: Gupta, Sonam, et al.
Veröffentlicht: (2024)
von: Gupta, Sonam, et al.
Veröffentlicht: (2024)
RSRCC: A Remote Sensing Regional Change Comprehension Benchmark Constructed via Retrieval-Augmented Best-of-N Ranking
von: Kazoom, Roie, et al.
Veröffentlicht: (2026)
von: Kazoom, Roie, et al.
Veröffentlicht: (2026)
A Nurse is Blue and Elephant is Rugby: Cross Domain Alignment in Large Language Models Reveal Human-like Patterns
von: Yehudai, Asaf, et al.
Veröffentlicht: (2024)
von: Yehudai, Asaf, et al.
Veröffentlicht: (2024)
Interactive Explanations for Reinforcement-Learning Agents
von: Amitai, Yotam, et al.
Veröffentlicht: (2025)
von: Amitai, Yotam, et al.
Veröffentlicht: (2025)
Can We Govern the Agent-to-Agent Economy?
von: Chaffer, Tomer Jordi
Veröffentlicht: (2025)
von: Chaffer, Tomer Jordi
Veröffentlicht: (2025)
ParaCodex: A Profiling-Guided Autonomous Coding Agent for Reliable Parallel Code Generation and Translation
von: Kaplan, Erel, et al.
Veröffentlicht: (2026)
von: Kaplan, Erel, et al.
Veröffentlicht: (2026)
MorphoBench: A Benchmark with Difficulty Adaptive to Model Reasoning
von: Wang, Xukai, et al.
Veröffentlicht: (2025)
von: Wang, Xukai, et al.
Veröffentlicht: (2025)
Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality
von: Calderon, Nitay, et al.
Veröffentlicht: (2026)
von: Calderon, Nitay, et al.
Veröffentlicht: (2026)
Stay Tuned: An Empirical Study of the Impact of Hyperparameters on LLM Tuning in Real-World Applications
von: Halfon, Alon, et al.
Veröffentlicht: (2024)
von: Halfon, Alon, et al.
Veröffentlicht: (2024)
Selectively Sharing Experiences Improves Multi-Agent Reinforcement Learning
von: Gerstgrasser, Matthias, et al.
Veröffentlicht: (2023)
von: Gerstgrasser, Matthias, et al.
Veröffentlicht: (2023)
Workflows vs Agents for Code Translation
von: Gray, Henry, et al.
Veröffentlicht: (2025)
von: Gray, Henry, et al.
Veröffentlicht: (2025)
Repository Intelligence Graph: Deterministic Architectural Map for LLM Code Assistants
von: Cherny-Shahar, Tsvi, et al.
Veröffentlicht: (2026)
von: Cherny-Shahar, Tsvi, et al.
Veröffentlicht: (2026)
All You Need is Sally-Anne: ToM in AI Strongly Supported After Surpassing Tests for 3-Year-Olds
von: Alon, Nitay, et al.
Veröffentlicht: (2025)
von: Alon, Nitay, et al.
Veröffentlicht: (2025)
TASTE: A Designer-Annotated Multi-Dimensional Preference Dataset for AI-Generated Graphic Design
von: Zhu, Haonan, et al.
Veröffentlicht: (2026)
von: Zhu, Haonan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
CLEAR: Error Analysis via LLM-as-a-Judge Made Easy
von: Yehudai, Asaf, et al.
Veröffentlicht: (2025) -
Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents
von: Yehudai, Asaf, et al.
Veröffentlicht: (2026) -
General Agent Evaluation
von: Bandel, Elron, et al.
Veröffentlicht: (2026) -
JuStRank: Benchmarking LLM Judges for System Ranking
von: Gera, Ariel, et al.
Veröffentlicht: (2024) -
Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench
von: Perlitz, Yotam, et al.
Veröffentlicht: (2024)