State of What Art? A Call for Multi-Prompt LLM Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mizrahi, Moran, Kaplan, Guy, Malkin, Dan, Dror, Rotem, Shahaf, Dafna, Stanovsky, Gabriel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Cooking Up Creativity: Enhancing LLM Creativity through Structured Recombination
von: Mizrahi, Moran, et al.
Veröffentlicht: (2025)
von: Mizrahi, Moran, et al.
Veröffentlicht: (2025)
AI Planning Framework for LLM-Based Web Agents
von: Shahnovsky, Orit, et al.
Veröffentlicht: (2026)
von: Shahnovsky, Orit, et al.
Veröffentlicht: (2026)
Visual Editing with LLM-based Tool Chaining: An Efficient Distillation Approach for Real-Time Applications
von: Sultan, Oren, et al.
Veröffentlicht: (2024)
von: Sultan, Oren, et al.
Veröffentlicht: (2024)
PromptSuite: A Task-Agnostic Framework for Multi-Prompt Generation
von: Habba, Eliya, et al.
Veröffentlicht: (2025)
von: Habba, Eliya, et al.
Veröffentlicht: (2025)
The State and Fate of Summarization Datasets: A Survey
von: Dahan, Noam, et al.
Veröffentlicht: (2024)
von: Dahan, Noam, et al.
Veröffentlicht: (2024)
One Joke to Rule them All? On the (Im)possibility of Generalizing Humor
von: Turgeman, Mor, et al.
Veröffentlicht: (2025)
von: Turgeman, Mor, et al.
Veröffentlicht: (2025)
Comparing Humans and Models on a Similar Scale: Towards Cognitive Gender Bias Evaluation in Coreference Resolution
von: Lior, Gili, et al.
Veröffentlicht: (2023)
von: Lior, Gili, et al.
Veröffentlicht: (2023)
The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs
von: Calderon, Nitay, et al.
Veröffentlicht: (2025)
von: Calderon, Nitay, et al.
Veröffentlicht: (2025)
InterFeat: A Pipeline for Finding Interesting Scientific Features
von: Ofer, Dan, et al.
Veröffentlicht: (2025)
von: Ofer, Dan, et al.
Veröffentlicht: (2025)
ParallelPARC: A Scalable Pipeline for Generating Natural-Language Analogies
von: Sultan, Oren, et al.
Veröffentlicht: (2024)
von: Sultan, Oren, et al.
Veröffentlicht: (2024)
ReliableEval: A Recipe for Stochastic LLM Evaluation via Method of Moments
von: Lior, Gili, et al.
Veröffentlicht: (2025)
von: Lior, Gili, et al.
Veröffentlicht: (2025)
Do Zombies Understand? A Choose-Your-Own-Adventure Exploration of Machine Cognition
von: Goldstein, Ariel, et al.
Veröffentlicht: (2024)
von: Goldstein, Ariel, et al.
Veröffentlicht: (2024)
DOVE: A Large-Scale Multi-Dimensional Predictions Dataset Towards Meaningful LLM Evaluation
von: Habba, Eliya, et al.
Veröffentlicht: (2025)
von: Habba, Eliya, et al.
Veröffentlicht: (2025)
Finding your MUSE: Mining Unexpected Solutions Engine
von: Sweed, Nir, et al.
Veröffentlicht: (2025)
von: Sweed, Nir, et al.
Veröffentlicht: (2025)
Time to Talk: LLM Agents for Asynchronous Group Communication in Mafia Games
von: Eckhaus, Niv, et al.
Veröffentlicht: (2025)
von: Eckhaus, Niv, et al.
Veröffentlicht: (2025)
In-Context Learning on a Budget: A Case Study in Token Classification
von: Berger, Uri, et al.
Veröffentlicht: (2024)
von: Berger, Uri, et al.
Veröffentlicht: (2024)
SAUCE: Synchronous and Asynchronous User-Customizable Environment for Multi-Agent LLM Interaction
von: Neuberger, Shlomo, et al.
Veröffentlicht: (2024)
von: Neuberger, Shlomo, et al.
Veröffentlicht: (2024)
Surveying the Landscape of Image Captioning Evaluation: A Comprehensive Taxonomy, Trends and Metrics Analysis
von: Berger, Uri, et al.
Veröffentlicht: (2024)
von: Berger, Uri, et al.
Veröffentlicht: (2024)
Leveraging Collection-Wide Similarities for Unsupervised Document Structure Extraction
von: Lior, Gili, et al.
Veröffentlicht: (2024)
von: Lior, Gili, et al.
Veröffentlicht: (2024)
Leveraging Digitized Newspapers to Collect Summarization Data in Low-Resource Languages
von: Dahan, Noam, et al.
Veröffentlicht: (2025)
von: Dahan, Noam, et al.
Veröffentlicht: (2025)
Beyond Memorization: Distinguishing between Reductive and Epistemic Reasoning in LLMs using Classic Logic Puzzles
von: Gabay, Adi, et al.
Veröffentlicht: (2026)
von: Gabay, Adi, et al.
Veröffentlicht: (2026)
Anticipatory Evaluation of Language Models
von: Park, Jungsoo, et al.
Veröffentlicht: (2025)
von: Park, Jungsoo, et al.
Veröffentlicht: (2025)
SEAM: A Stochastic Benchmark for Multi-Document Tasks
von: Lior, Gili, et al.
Veröffentlicht: (2024)
von: Lior, Gili, et al.
Veröffentlicht: (2024)
Comparing the Framing Effect in Humans and LLMs on Naturally Occurring Texts
von: Lior, Gili, et al.
Veröffentlicht: (2025)
von: Lior, Gili, et al.
Veröffentlicht: (2025)
Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs
von: Itzhak, Itay, et al.
Veröffentlicht: (2025)
von: Itzhak, Itay, et al.
Veröffentlicht: (2025)
Applying Intrinsic Debiasing on Downstream Tasks: Challenges and Considerations for Machine Translation
von: Iluz, Bar, et al.
Veröffentlicht: (2024)
von: Iluz, Bar, et al.
Veröffentlicht: (2024)
LLMs versus the Halting Problem: Characterizing Program Termination Reasoning
von: Sultan, Oren, et al.
Veröffentlicht: (2026)
von: Sultan, Oren, et al.
Veröffentlicht: (2026)
Mind the Gap... or Not? How Translation Errors and Evaluation Details Skew Multilingual Results
von: Peter, Jan-Thorsten, et al.
Veröffentlicht: (2025)
von: Peter, Jan-Thorsten, et al.
Veröffentlicht: (2025)
Towards Reliable Benchmarking: A Contamination Free, Controllable Evaluation Framework for Multi-step LLM Function Calling
von: Maekawa, Seiji, et al.
Veröffentlicht: (2025)
von: Maekawa, Seiji, et al.
Veröffentlicht: (2025)
Improving Image Captioning by Mimicking Human Reformulation Feedback at Inference-time
von: Berger, Uri, et al.
Veröffentlicht: (2025)
von: Berger, Uri, et al.
Veröffentlicht: (2025)
Looking Beyond The Top-1: Transformers Determine Top Tokens In Order
von: Lioubashevski, Daria, et al.
Veröffentlicht: (2024)
von: Lioubashevski, Daria, et al.
Veröffentlicht: (2024)
More Documents, Same Length: Isolating the Challenge of Multiple Documents in RAG
von: Levy, Shahar, et al.
Veröffentlicht: (2025)
von: Levy, Shahar, et al.
Veröffentlicht: (2025)
Can LLMs Help Uncover Insights about LLMs? A Large-Scale, Evolving Literature Analysis of Frontier LLMs
von: Park, Jungsoo, et al.
Veröffentlicht: (2025)
von: Park, Jungsoo, et al.
Veröffentlicht: (2025)
Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration
von: Habba, Eliya, et al.
Veröffentlicht: (2026)
von: Habba, Eliya, et al.
Veröffentlicht: (2026)
SHAP Meets Tensor Networks: Provably Tractable Explanations with Parallelism
von: Marzouk, Reda, et al.
Veröffentlicht: (2025)
von: Marzouk, Reda, et al.
Veröffentlicht: (2025)
Vocab Diet: Reshaping the Vocabulary of LLMs via Vector Arithmetic
von: Reif, Yuval, et al.
Veröffentlicht: (2025)
von: Reif, Yuval, et al.
Veröffentlicht: (2025)
Beyond Benchmarks: On The False Promise of AI Regulation
von: Stanovsky, Gabriel, et al.
Veröffentlicht: (2025)
von: Stanovsky, Gabriel, et al.
Veröffentlicht: (2025)
From Feelings to Metrics: Understanding and Formalizing How Users Vibe-Test LLMs
von: Itzhak, Itay, et al.
Veröffentlicht: (2026)
von: Itzhak, Itay, et al.
Veröffentlicht: (2026)
What Prompts Don't Say: Understanding and Managing Underspecification in LLM Prompts
von: Yang, Chenyang, et al.
Veröffentlicht: (2025)
von: Yang, Chenyang, et al.
Veröffentlicht: (2025)
signwriting-evaluation: Effective Sign Language Evaluation via SignWriting
von: Moryossef, Amit, et al.
Veröffentlicht: (2024)
von: Moryossef, Amit, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Cooking Up Creativity: Enhancing LLM Creativity through Structured Recombination
von: Mizrahi, Moran, et al.
Veröffentlicht: (2025) -
AI Planning Framework for LLM-Based Web Agents
von: Shahnovsky, Orit, et al.
Veröffentlicht: (2026) -
Visual Editing with LLM-based Tool Chaining: An Efficient Distillation Approach for Real-Time Applications
von: Sultan, Oren, et al.
Veröffentlicht: (2024) -
PromptSuite: A Task-Agnostic Framework for Multi-Prompt Generation
von: Habba, Eliya, et al.
Veröffentlicht: (2025) -
The State and Fate of Summarization Datasets: A Survey
von: Dahan, Noam, et al.
Veröffentlicht: (2024)