ReliableEval: A Recipe for Stochastic LLM Evaluation via Method of Moments
Fuente:
arXiv
Saved in:
| Main Authors: | Lior, Gili, Habba, Eliya, Levy, Shahar, Caciularu, Avi, Stanovsky, Gabriel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SEAM: A Stochastic Benchmark for Multi-Document Tasks
by: Lior, Gili, et al.
Published: (2024)
by: Lior, Gili, et al.
Published: (2024)
PromptSuite: A Task-Agnostic Framework for Multi-Prompt Generation
by: Habba, Eliya, et al.
Published: (2025)
by: Habba, Eliya, et al.
Published: (2025)
ScheMatiQ: From Research Question to Structured Data through Interactive Schema Discovery
by: Levy, Shahar, et al.
Published: (2026)
by: Levy, Shahar, et al.
Published: (2026)
Comparing Humans and Models on a Similar Scale: Towards Cognitive Gender Bias Evaluation in Coreference Resolution
by: Lior, Gili, et al.
Published: (2023)
by: Lior, Gili, et al.
Published: (2023)
Beyond Benchmarks: On The False Promise of AI Regulation
by: Stanovsky, Gabriel, et al.
Published: (2025)
by: Stanovsky, Gabriel, et al.
Published: (2025)
From Feelings to Metrics: Understanding and Formalizing How Users Vibe-Test LLMs
by: Itzhak, Itay, et al.
Published: (2026)
by: Itzhak, Itay, et al.
Published: (2026)
Leveraging Collection-Wide Similarities for Unsupervised Document Structure Extraction
by: Lior, Gili, et al.
Published: (2024)
by: Lior, Gili, et al.
Published: (2024)
DOVE: A Large-Scale Multi-Dimensional Predictions Dataset Towards Meaningful LLM Evaluation
by: Habba, Eliya, et al.
Published: (2025)
by: Habba, Eliya, et al.
Published: (2025)
Comparing the Framing Effect in Humans and LLMs on Naturally Occurring Texts
by: Lior, Gili, et al.
Published: (2025)
by: Lior, Gili, et al.
Published: (2025)
Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration
by: Habba, Eliya, et al.
Published: (2026)
by: Habba, Eliya, et al.
Published: (2026)
JSON Whisperer: Efficient JSON Editing with LLMs
by: Duanis, Sarel, et al.
Published: (2025)
by: Duanis, Sarel, et al.
Published: (2025)
More Documents, Same Length: Isolating the Challenge of Multiple Documents in RAG
by: Levy, Shahar, et al.
Published: (2025)
by: Levy, Shahar, et al.
Published: (2025)
Dont Add, dont Miss: Effective Content Preserving Generation from Pre-Selected Text Spans
by: Slobodkin, Aviv, et al.
Published: (2023)
by: Slobodkin, Aviv, et al.
Published: (2023)
Latent Reasoning with Supervised Thinking States
by: Amos, Ido, et al.
Published: (2026)
by: Amos, Ido, et al.
Published: (2026)
Reversed Attention: On The Gradient Descent Of Attention Layers In GPT
by: Katz, Shahar, et al.
Published: (2024)
by: Katz, Shahar, et al.
Published: (2024)
The State and Fate of Summarization Datasets: A Survey
by: Dahan, Noam, et al.
Published: (2024)
by: Dahan, Noam, et al.
Published: (2024)
Do Zombies Understand? A Choose-Your-Own-Adventure Exploration of Machine Cognition
by: Goldstein, Ariel, et al.
Published: (2024)
by: Goldstein, Ariel, et al.
Published: (2024)
TruthEval: A Dataset to Evaluate LLM Truthfulness and Reliability
by: Khatun, Aisha, et al.
Published: (2024)
by: Khatun, Aisha, et al.
Published: (2024)
State of What Art? A Call for Multi-Prompt LLM Evaluation
by: Mizrahi, Moran, et al.
Published: (2023)
by: Mizrahi, Moran, et al.
Published: (2023)
Time to Talk: LLM Agents for Asynchronous Group Communication in Mafia Games
by: Eckhaus, Niv, et al.
Published: (2025)
by: Eckhaus, Niv, et al.
Published: (2025)
Unpacking Tokenization: Evaluating Text Compression and its Correlation with Model Performance
by: Goldman, Omer, et al.
Published: (2024)
by: Goldman, Omer, et al.
Published: (2024)
In-Context Learning on a Budget: A Case Study in Token Classification
by: Berger, Uri, et al.
Published: (2024)
by: Berger, Uri, et al.
Published: (2024)
Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models
by: Ghandeharioun, Asma, et al.
Published: (2024)
by: Ghandeharioun, Asma, et al.
Published: (2024)
Surveying the Landscape of Image Captioning Evaluation: A Comprehensive Taxonomy, Trends and Metrics Analysis
by: Berger, Uri, et al.
Published: (2024)
by: Berger, Uri, et al.
Published: (2024)
MDCure: A Scalable Pipeline for Multi-Document Instruction-Following
by: Liu, Gabrielle Kaili-May, et al.
Published: (2024)
by: Liu, Gabrielle Kaili-May, et al.
Published: (2024)
Leveraging Digitized Newspapers to Collect Summarization Data in Low-Resource Languages
by: Dahan, Noam, et al.
Published: (2025)
by: Dahan, Noam, et al.
Published: (2025)
Beyond Memorization: Distinguishing between Reductive and Epistemic Reasoning in LLMs using Classic Logic Puzzles
by: Gabay, Adi, et al.
Published: (2026)
by: Gabay, Adi, et al.
Published: (2026)
DiagramEval: Evaluating LLM-Generated Diagrams via Graphs
by: Liang, Chumeng, et al.
Published: (2025)
by: Liang, Chumeng, et al.
Published: (2025)
Teach2Eval: An Indirect Evaluation Method for LLM by Judging How It Teaches
by: Zhou, Yuhang, et al.
Published: (2025)
by: Zhou, Yuhang, et al.
Published: (2025)
Anticipatory Evaluation of Language Models
by: Park, Jungsoo, et al.
Published: (2025)
by: Park, Jungsoo, et al.
Published: (2025)
Identifying User Goals from UI Trajectories
by: Berkovitch, Omri, et al.
Published: (2024)
by: Berkovitch, Omri, et al.
Published: (2024)
Segment-Based Attention Masking for GPTs
by: Katz, Shahar, et al.
Published: (2024)
by: Katz, Shahar, et al.
Published: (2024)
RocketEval: Efficient Automated LLM Evaluation via Grading Checklist
by: Wei, Tianjun, et al.
Published: (2025)
by: Wei, Tianjun, et al.
Published: (2025)
ILRR: Inference-Time Steering Method for Masked Diffusion Language Models
by: Avrahami, Eden, et al.
Published: (2026)
by: Avrahami, Eden, et al.
Published: (2026)
ConvApparel: A Benchmark Dataset and Validation Framework for User Simulators in Conversational Recommenders
by: Meshi, Ofer, et al.
Published: (2026)
by: Meshi, Ofer, et al.
Published: (2026)
Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs
by: Itzhak, Itay, et al.
Published: (2025)
by: Itzhak, Itay, et al.
Published: (2025)
Visual Riddles: a Commonsense and World Knowledge Challenge for Large Vision and Language Models
by: Bitton-Guetta, Nitzan, et al.
Published: (2024)
by: Bitton-Guetta, Nitzan, et al.
Published: (2024)
SAUCE: Synchronous and Asynchronous User-Customizable Environment for Multi-Agent LLM Interaction
by: Neuberger, Shlomo, et al.
Published: (2024)
by: Neuberger, Shlomo, et al.
Published: (2024)
Applying Intrinsic Debiasing on Downstream Tasks: Challenges and Considerations for Machine Translation
by: Iluz, Bar, et al.
Published: (2024)
by: Iluz, Bar, et al.
Published: (2024)
RepEval: Effective Text Evaluation with LLM Representation
by: Sheng, Shuqian, et al.
Published: (2024)
by: Sheng, Shuqian, et al.
Published: (2024)
Similar Items
-
SEAM: A Stochastic Benchmark for Multi-Document Tasks
by: Lior, Gili, et al.
Published: (2024) -
PromptSuite: A Task-Agnostic Framework for Multi-Prompt Generation
by: Habba, Eliya, et al.
Published: (2025) -
ScheMatiQ: From Research Question to Structured Data through Interactive Schema Discovery
by: Levy, Shahar, et al.
Published: (2026) -
Comparing Humans and Models on a Similar Scale: Towards Cognitive Gender Bias Evaluation in Coreference Resolution
by: Lior, Gili, et al.
Published: (2023) -
Beyond Benchmarks: On The False Promise of AI Regulation
by: Stanovsky, Gabriel, et al.
Published: (2025)