HeQ: a Large and Diverse Hebrew Reading Comprehension Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | Cohen, Amir DN, Merhav, Hilla, Goldberg, Yoav, Tsarfaty, Reut |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Simple Linguistic Inferences of Large Language Models (LLMs): Blind Spots and Blinds
by: Basmov, Victoria, et al.
Published: (2023)
by: Basmov, Victoria, et al.
Published: (2023)
Beyond Word Boundaries: A Hebrew Coreference Benchmark and an Evaluation Protocol for Morphologically Complex Text
by: Greenfeld, Refael Shaked, et al.
Published: (2026)
by: Greenfeld, Refael Shaked, et al.
Published: (2026)
HeSum: a Novel Dataset for Abstractive Text Summarization in Hebrew
by: Paz-Argaman, Tzuf, et al.
Published: (2024)
by: Paz-Argaman, Tzuf, et al.
Published: (2024)
LLMs' Reading Comprehension Is Affected by Parametric Knowledge and Struggles with Hypothetical Statements
by: Basmov, Victoria, et al.
Published: (2024)
by: Basmov, Victoria, et al.
Published: (2024)
A Novel Computational and Modeling Foundation for Automatic Coherence Assessment
by: Maimon, Aviya, et al.
Published: (2023)
by: Maimon, Aviya, et al.
Published: (2023)
Mitigating Hallucinations in Large Vision-Language Models (LVLMs) via Language-Contrastive Decoding (LCD)
by: Manevich, Avshalom, et al.
Published: (2024)
by: Manevich, Avshalom, et al.
Published: (2024)
NoviCode: Generating Programs from Natural Language Utterances by Novices
by: Mordechai, Asaf Achi, et al.
Published: (2024)
by: Mordechai, Asaf Achi, et al.
Published: (2024)
A Truly Joint Neural Architecture for Segmentation and Parsing
by: Levi, Danit Yshaayahu, et al.
Published: (2024)
by: Levi, Danit Yshaayahu, et al.
Published: (2024)
MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of Documents
by: Wolfson, Tomer, et al.
Published: (2025)
by: Wolfson, Tomer, et al.
Published: (2025)
IQ Test for LLMs: An Evaluation Framework for Uncovering Core Skills in LLMs
by: Maimon, Aviya, et al.
Published: (2025)
by: Maimon, Aviya, et al.
Published: (2025)
Diversity Over Quantity: A Lesson From Few Shot Relation Classification
by: Cohen, Amir DN, et al.
Published: (2024)
by: Cohen, Amir DN, et al.
Published: (2024)
MRL Parsing Without Tears: The Case of Hebrew
by: Shmidman, Shaltiel, et al.
Published: (2024)
by: Shmidman, Shaltiel, et al.
Published: (2024)
Into the Unknown: Generating Geospatial Descriptions for New Environments
by: Paz-Argaman, Tzuf, et al.
Published: (2024)
by: Paz-Argaman, Tzuf, et al.
Published: (2024)
Is It Really Long Context if All You Need Is Retrieval? Towards Genuinely Difficult Long Context NLP
by: Goldman, Omer, et al.
Published: (2024)
by: Goldman, Omer, et al.
Published: (2024)
Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models
by: Levy, Mosh, et al.
Published: (2024)
by: Levy, Mosh, et al.
Published: (2024)
Multilingual Instruction Tuning With Just a Pinch of Multilinguality
by: Shaham, Uri, et al.
Published: (2024)
by: Shaham, Uri, et al.
Published: (2024)
Dicta-LM 3.0: Advancing The Frontier of Hebrew Sovereign LLMs
by: Shmidman, Shaltiel, et al.
Published: (2026)
by: Shmidman, Shaltiel, et al.
Published: (2026)
Unpacking Tokenization: Evaluating Text Compression and its Correlation with Model Performance
by: Goldman, Omer, et al.
Published: (2024)
by: Goldman, Omer, et al.
Published: (2024)
Do Pretrained Contextual Language Models Distinguish between Hebrew Homograph Analyses?
by: Shmidman, Avi, et al.
Published: (2024)
by: Shmidman, Avi, et al.
Published: (2024)
Beyond Facts: Benchmarking Distributional Reading Comprehension in Large Language Models
by: Guo, Pei-Fu, et al.
Published: (2026)
by: Guo, Pei-Fu, et al.
Published: (2026)
Adapting LLMs to Hebrew: Unveiling DictaLM 2.0 with Enhanced Vocabulary and Instruction Capabilities
by: Shmidman, Shaltiel, et al.
Published: (2024)
by: Shmidman, Shaltiel, et al.
Published: (2024)
HeTGB: A Comprehensive Benchmark for Heterophilic Text-Attributed Graphs
by: Li, Shujie, et al.
Published: (2025)
by: Li, Shujie, et al.
Published: (2025)
MRCEval: A Comprehensive, Challenging and Accessible Machine Reading Comprehension Benchmark
by: Ma, Shengkun, et al.
Published: (2025)
by: Ma, Shengkun, et al.
Published: (2025)
Humans Perceive Wrong Narratives from AI Reasoning Texts
by: Levy, Mosh, et al.
Published: (2025)
by: Levy, Mosh, et al.
Published: (2025)
State over Tokens: Characterizing the Role of Reasoning Tokens
by: Levy, Mosh, et al.
Published: (2025)
by: Levy, Mosh, et al.
Published: (2025)
Description-Based Text Similarity
by: Ravfogel, Shauli, et al.
Published: (2023)
by: Ravfogel, Shauli, et al.
Published: (2023)
Knowledge Navigator: LLM-guided Browsing Framework for Exploratory Search in Scientific Literature
by: Katz, Uri, et al.
Published: (2024)
by: Katz, Uri, et al.
Published: (2024)
Location Not Found: Exposing Implicit Local and Global Biases in Multilingual LLMs
by: Mor-Lan, Guy, et al.
Published: (2026)
by: Mor-Lan, Guy, et al.
Published: (2026)
NER Retriever: Zero-Shot Named Entity Retrieval with Type-Aware Embeddings
by: Shachar, Or, et al.
Published: (2025)
by: Shachar, Or, et al.
Published: (2025)
Can Large Language Model Summarizers Adapt to Diverse Scientific Communication Goals?
by: Fonseca, Marcio, et al.
Published: (2024)
by: Fonseca, Marcio, et al.
Published: (2024)
Generating Reading Comprehension Exercises with Large Language Models for Educational Applications
by: Huang, Xingyu, et al.
Published: (2025)
by: Huang, Xingyu, et al.
Published: (2025)
Mevaker: Conclusion Extraction and Allocation Resources for the Hebrew Language
by: Shalumov, Vitaly, et al.
Published: (2024)
by: Shalumov, Vitaly, et al.
Published: (2024)
Large Language Models in the Clinic: A Comprehensive Benchmark
by: Liu, Fenglin, et al.
Published: (2024)
by: Liu, Fenglin, et al.
Published: (2024)
Decoding Open-Ended Information Seeking Goals from Eye Movements in Reading
by: Hadar, Cfir Avraham, et al.
Published: (2025)
by: Hadar, Cfir Avraham, et al.
Published: (2025)
Generating Benchmarks for Factuality Evaluation of Language Models
by: Muhlgay, Dor, et al.
Published: (2023)
by: Muhlgay, Dor, et al.
Published: (2023)
The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality
by: Cheng, Aileen, et al.
Published: (2025)
by: Cheng, Aileen, et al.
Published: (2025)
ECLeKTic: a Novel Challenge Set for Evaluation of Cross-Lingual Knowledge Transfer
by: Goldman, Omer, et al.
Published: (2025)
by: Goldman, Omer, et al.
Published: (2025)
MTQE.en-he: Machine Translation Quality Estimation for English-Hebrew
by: Rosenbaum, Andy, et al.
Published: (2026)
by: Rosenbaum, Andy, et al.
Published: (2026)
Read Before You Think: Mitigating LLM Comprehension Failures with Step-by-Step Reading
by: Han, Feijiang, et al.
Published: (2025)
by: Han, Feijiang, et al.
Published: (2025)
ElectriQ: A Benchmark for Assessing the Response Capability of Large Language Models in Power Marketing
by: Wang, Jinzhi, et al.
Published: (2025)
by: Wang, Jinzhi, et al.
Published: (2025)
Similar Items
-
Simple Linguistic Inferences of Large Language Models (LLMs): Blind Spots and Blinds
by: Basmov, Victoria, et al.
Published: (2023) -
Beyond Word Boundaries: A Hebrew Coreference Benchmark and an Evaluation Protocol for Morphologically Complex Text
by: Greenfeld, Refael Shaked, et al.
Published: (2026) -
HeSum: a Novel Dataset for Abstractive Text Summarization in Hebrew
by: Paz-Argaman, Tzuf, et al.
Published: (2024) -
LLMs' Reading Comprehension Is Affected by Parametric Knowledge and Struggles with Hypothetical Statements
by: Basmov, Victoria, et al.
Published: (2024) -
A Novel Computational and Modeling Foundation for Automatic Coherence Assessment
by: Maimon, Aviya, et al.
Published: (2023)