Saved in:
| Main Authors: | Lior, Gili, Caciularu, Avi, Cattan, Arie, Levy, Shahar, Shapira, Ori, Stanovsky, Gabriel |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2406.16086 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ReliableEval: A Recipe for Stochastic LLM Evaluation via Method of Moments
by: Lior, Gili, et al.
Published: (2025)
by: Lior, Gili, et al.
Published: (2025)
Comparing Humans and Models on a Similar Scale: Towards Cognitive Gender Bias Evaluation in Coreference Resolution
by: Lior, Gili, et al.
Published: (2023)
by: Lior, Gili, et al.
Published: (2023)
PromptSuite: A Task-Agnostic Framework for Multi-Prompt Generation
by: Habba, Eliya, et al.
Published: (2025)
by: Habba, Eliya, et al.
Published: (2025)
Leveraging Collection-Wide Similarities for Unsupervised Document Structure Extraction
by: Lior, Gili, et al.
Published: (2024)
by: Lior, Gili, et al.
Published: (2024)
Consensus or Conflict? Fine-Grained Evaluation of Conflicting Answers in Question-Answering
by: Nachshoni, Eviatar, et al.
Published: (2025)
by: Nachshoni, Eviatar, et al.
Published: (2025)
Comparing the Framing Effect in Humans and LLMs on Naturally Occurring Texts
by: Lior, Gili, et al.
Published: (2025)
by: Lior, Gili, et al.
Published: (2025)
DRAGged into Conflicts: Detecting and Addressing Conflicting Sources in Search-Augmented LLMs
by: Cattan, Arie, et al.
Published: (2025)
by: Cattan, Arie, et al.
Published: (2025)
More Documents, Same Length: Isolating the Challenge of Multiple Documents in RAG
by: Levy, Shahar, et al.
Published: (2025)
by: Levy, Shahar, et al.
Published: (2025)
Multi-Review Fusion-in-Context
by: Slobodkin, Aviv, et al.
Published: (2024)
by: Slobodkin, Aviv, et al.
Published: (2024)
DoubleDipper: Improving Long-Context LLMs via Context Recycling
by: Cattan, Arie, et al.
Published: (2024)
by: Cattan, Arie, et al.
Published: (2024)
ScheMatiQ: From Research Question to Structured Data through Interactive Schema Discovery
by: Levy, Shahar, et al.
Published: (2026)
by: Levy, Shahar, et al.
Published: (2026)
MDCure: A Scalable Pipeline for Multi-Document Instruction-Following
by: Liu, Gabrielle Kaili-May, et al.
Published: (2024)
by: Liu, Gabrielle Kaili-May, et al.
Published: (2024)
A Unifying Scheme for Extractive Content Selection Tasks
by: Amar, Shmuel, et al.
Published: (2025)
by: Amar, Shmuel, et al.
Published: (2025)
Information Types in Product Reviews
by: Shapira, Ori, et al.
Published: (2025)
by: Shapira, Ori, et al.
Published: (2025)
Measuring the Effect of Transcription Noise on Downstream Language Understanding Tasks
by: Shapira, Ori, et al.
Published: (2025)
by: Shapira, Ori, et al.
Published: (2025)
Latent Reasoning with Supervised Thinking States
by: Amos, Ido, et al.
Published: (2026)
by: Amos, Ido, et al.
Published: (2026)
Dont Add, dont Miss: Effective Content Preserving Generation from Pre-Selected Text Spans
by: Slobodkin, Aviv, et al.
Published: (2023)
by: Slobodkin, Aviv, et al.
Published: (2023)
Reversed Attention: On The Gradient Descent Of Attention Layers In GPT
by: Katz, Shahar, et al.
Published: (2024)
by: Katz, Shahar, et al.
Published: (2024)
The State and Fate of Summarization Datasets: A Survey
by: Dahan, Noam, et al.
Published: (2024)
by: Dahan, Noam, et al.
Published: (2024)
The Power of Summary-Source Alignments
by: Ernst, Ori, et al.
Published: (2024)
by: Ernst, Ori, et al.
Published: (2024)
Do Zombies Understand? A Choose-Your-Own-Adventure Exploration of Machine Cognition
by: Goldstein, Ariel, et al.
Published: (2024)
by: Goldstein, Ariel, et al.
Published: (2024)
Attribute First, then Generate: Locally-attributable Grounded Text Generation
by: Slobodkin, Aviv, et al.
Published: (2024)
by: Slobodkin, Aviv, et al.
Published: (2024)
TabAgent: A Framework for Replacing Agentic Generative Components with Tabular-Textual Classifiers
by: Levy, Ido, et al.
Published: (2026)
by: Levy, Ido, et al.
Published: (2026)
Applying Intrinsic Debiasing on Downstream Tasks: Challenges and Considerations for Machine Translation
by: Iluz, Bar, et al.
Published: (2024)
by: Iluz, Bar, et al.
Published: (2024)
ConvApparel: A Benchmark Dataset and Validation Framework for User Simulators in Conversational Recommenders
by: Meshi, Ofer, et al.
Published: (2026)
by: Meshi, Ofer, et al.
Published: (2026)
CoverBench: A Challenging Benchmark for Complex Claim Verification
by: Jacovi, Alon, et al.
Published: (2024)
by: Jacovi, Alon, et al.
Published: (2024)
Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models
by: Ghandeharioun, Asma, et al.
Published: (2024)
by: Ghandeharioun, Asma, et al.
Published: (2024)
Beyond Benchmarks: On The False Promise of AI Regulation
by: Stanovsky, Gabriel, et al.
Published: (2025)
by: Stanovsky, Gabriel, et al.
Published: (2025)
In-Context Learning on a Budget: A Case Study in Token Classification
by: Berger, Uri, et al.
Published: (2024)
by: Berger, Uri, et al.
Published: (2024)
Leveraging Digitized Newspapers to Collect Summarization Data in Low-Resource Languages
by: Dahan, Noam, et al.
Published: (2025)
by: Dahan, Noam, et al.
Published: (2025)
Beyond Memorization: Distinguishing between Reductive and Epistemic Reasoning in LLMs using Classic Logic Puzzles
by: Gabay, Adi, et al.
Published: (2026)
by: Gabay, Adi, et al.
Published: (2026)
QAPyramid: Fine-grained Evaluation of Content Selection for Text Summarization
by: Zhang, Shiyue, et al.
Published: (2024)
by: Zhang, Shiyue, et al.
Published: (2024)
Identifying User Goals from UI Trajectories
by: Berkovitch, Omri, et al.
Published: (2024)
by: Berkovitch, Omri, et al.
Published: (2024)
Explicating the Implicit: Argument Detection Beyond Sentence Boundaries
by: Roit, Paul, et al.
Published: (2024)
by: Roit, Paul, et al.
Published: (2024)
Unpacking Tokenization: Evaluating Text Compression and its Correlation with Model Performance
by: Goldman, Omer, et al.
Published: (2024)
by: Goldman, Omer, et al.
Published: (2024)
Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs
by: Itzhak, Itay, et al.
Published: (2025)
by: Itzhak, Itay, et al.
Published: (2025)
Segment-Based Attention Masking for GPTs
by: Katz, Shahar, et al.
Published: (2024)
by: Katz, Shahar, et al.
Published: (2024)
CLATTER: Comprehensive Entailment Reasoning for Hallucination Detection
by: Eliav, Ron, et al.
Published: (2025)
by: Eliav, Ron, et al.
Published: (2025)
State of What Art? A Call for Multi-Prompt LLM Evaluation
by: Mizrahi, Moran, et al.
Published: (2023)
by: Mizrahi, Moran, et al.
Published: (2023)
Time to Talk: LLM Agents for Asynchronous Group Communication in Mafia Games
by: Eckhaus, Niv, et al.
Published: (2025)
by: Eckhaus, Niv, et al.
Published: (2025)
Similar Items
-
ReliableEval: A Recipe for Stochastic LLM Evaluation via Method of Moments
by: Lior, Gili, et al.
Published: (2025) -
Comparing Humans and Models on a Similar Scale: Towards Cognitive Gender Bias Evaluation in Coreference Resolution
by: Lior, Gili, et al.
Published: (2023) -
PromptSuite: A Task-Agnostic Framework for Multi-Prompt Generation
by: Habba, Eliya, et al.
Published: (2025) -
Leveraging Collection-Wide Similarities for Unsupervised Document Structure Extraction
by: Lior, Gili, et al.
Published: (2024) -
Consensus or Conflict? Fine-Grained Evaluation of Conflicting Answers in Question-Answering
by: Nachshoni, Eviatar, et al.
Published: (2025)