MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of Documents
Fuente:
arXiv
Saved in:
| Main Authors: | Wolfson, Tomer, Trivedi, Harsh, Geva, Mor, Goldberg, Yoav, Roth, Dan, Khot, Tushar, Sabharwal, Ashish, Tsarfaty, Reut |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
NoviCode: Generating Programs from Natural Language Utterances by Novices
by: Mordechai, Asaf Achi, et al.
Published: (2024)
by: Mordechai, Asaf Achi, et al.
Published: (2024)
LLMs' Reading Comprehension Is Affected by Parametric Knowledge and Struggles with Hypothetical Statements
by: Basmov, Victoria, et al.
Published: (2024)
by: Basmov, Victoria, et al.
Published: (2024)
Simple Linguistic Inferences of Large Language Models (LLMs): Blind Spots and Blinds
by: Basmov, Victoria, et al.
Published: (2023)
by: Basmov, Victoria, et al.
Published: (2023)
HeQ: a Large and Diverse Hebrew Reading Comprehension Benchmark
by: Cohen, Amir DN, et al.
Published: (2025)
by: Cohen, Amir DN, et al.
Published: (2025)
AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents
by: Trivedi, Harsh, et al.
Published: (2024)
by: Trivedi, Harsh, et al.
Published: (2024)
Leveraging In-Context Learning for Language Model Agents
by: Gupta, Shivanshu, et al.
Published: (2025)
by: Gupta, Shivanshu, et al.
Published: (2025)
Beyond Word Boundaries: A Hebrew Coreference Benchmark and an Evaluation Protocol for Morphologically Complex Text
by: Greenfeld, Refael Shaked, et al.
Published: (2026)
by: Greenfeld, Refael Shaked, et al.
Published: (2026)
Mitigating Hallucinations in Large Vision-Language Models (LVLMs) via Language-Contrastive Decoding (LCD)
by: Manevich, Avshalom, et al.
Published: (2024)
by: Manevich, Avshalom, et al.
Published: (2024)
A Novel Computational and Modeling Foundation for Automatic Coherence Assessment
by: Maimon, Aviya, et al.
Published: (2023)
by: Maimon, Aviya, et al.
Published: (2023)
A Truly Joint Neural Architecture for Segmentation and Parsing
by: Levi, Danit Yshaayahu, et al.
Published: (2024)
by: Levi, Danit Yshaayahu, et al.
Published: (2024)
Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context
by: Gur-Arieh, Yoav, et al.
Published: (2025)
by: Gur-Arieh, Yoav, et al.
Published: (2025)
Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
by: Gur-Arieh, Yoav, et al.
Published: (2026)
by: Gur-Arieh, Yoav, et al.
Published: (2026)
Disentangling MLP Neuron Weights in Vocabulary Space
by: Avrahamy, Asaf, et al.
Published: (2026)
by: Avrahamy, Asaf, et al.
Published: (2026)
ADaPT: As-Needed Decomposition and Planning with Language Models
by: Prasad, Archiki, et al.
Published: (2023)
by: Prasad, Archiki, et al.
Published: (2023)
Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMs
by: Gupta, Shashank, et al.
Published: (2023)
by: Gupta, Shashank, et al.
Published: (2023)
EnrichIndex: Using LLMs to Enrich Retrieval Indices Offline
by: Chen, Peter Baile, et al.
Published: (2025)
by: Chen, Peter Baile, et al.
Published: (2025)
Weakly Supervised Text-to-SQL Parsing through Question Decomposition
by: Wolfson, Tomer, et al.
Published: (2021)
by: Wolfson, Tomer, et al.
Published: (2021)
Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers
by: Yona, Gal, et al.
Published: (2024)
by: Yona, Gal, et al.
Published: (2024)
SAGE: Structure Aware Graph Expansion for Retrieval of Heterogeneous Data
by: Titiya, Prasham, et al.
Published: (2026)
by: Titiya, Prasham, et al.
Published: (2026)
Beyond N-Grams: Rethinking Evaluation Metrics and Strategies for Multilingual Abstractive Summarization
by: Mondshine, Itai, et al.
Published: (2025)
by: Mondshine, Itai, et al.
Published: (2025)
Beyond English: The Impact of Prompt Translation Strategies across Languages and Tasks in Multilingual LLMs
by: Mondshine, Itai, et al.
Published: (2025)
by: Mondshine, Itai, et al.
Published: (2025)
SUPER: Evaluating Agents on Setting Up and Executing Tasks from Research Repositories
by: Bogin, Ben, et al.
Published: (2024)
by: Bogin, Ben, et al.
Published: (2024)
Effective QA-driven Annotation of Predicate-Argument Relations Across Languages
by: Davidov, Jonathan, et al.
Published: (2026)
by: Davidov, Jonathan, et al.
Published: (2026)
Decomposed Prompting Does Not Fix Knowledge Gaps, But Helps Models Say "I Don't Know"
by: Madhwal, Dhruv, et al.
Published: (2026)
by: Madhwal, Dhruv, et al.
Published: (2026)
Estimating Knowledge in Large Language Models Without Generating a Single Token
by: Gottesman, Daniela, et al.
Published: (2024)
by: Gottesman, Daniela, et al.
Published: (2024)
Inferring Functionality of Attention Heads from their Parameters
by: Elhelo, Amit, et al.
Published: (2024)
by: Elhelo, Amit, et al.
Published: (2024)
Where Do We Go from Here? Multi-scale Allocentric Relational Inference from Natural Spatial Descriptions
by: Paz-Argaman, Tzuf, et al.
Published: (2024)
by: Paz-Argaman, Tzuf, et al.
Published: (2024)
Performance Gap in Entity Knowledge Extraction Across Modalities in Vision Language Models
by: Cohen, Ido, et al.
Published: (2024)
by: Cohen, Ido, et al.
Published: (2024)
MRL Parsing Without Tears: The Case of Hebrew
by: Shmidman, Shaltiel, et al.
Published: (2024)
by: Shmidman, Shaltiel, et al.
Published: (2024)
Superlatives in Context: Modeling the Implicit Semantics of Superlatives
by: Pyatkin, Valentina, et al.
Published: (2024)
by: Pyatkin, Valentina, et al.
Published: (2024)
Do Pretrained Contextual Language Models Distinguish between Hebrew Homograph Analyses?
by: Shmidman, Avi, et al.
Published: (2024)
by: Shmidman, Avi, et al.
Published: (2024)
Enhancing Automated Interpretability with Output-Centric Feature Descriptions
by: Gur-Arieh, Yoav, et al.
Published: (2025)
by: Gur-Arieh, Yoav, et al.
Published: (2025)
Precise In-Parameter Concept Erasure in Large Language Models
by: Gur-Arieh, Yoav, et al.
Published: (2025)
by: Gur-Arieh, Yoav, et al.
Published: (2025)
Friends and Grandmothers in Silico: Localizing Entity Cells in Language Models
by: Yona, Itay, et al.
Published: (2026)
by: Yona, Itay, et al.
Published: (2026)
Constructing Interpretable Features from Compositional Neuron Groups
by: Shafran, Or, et al.
Published: (2025)
by: Shafran, Or, et al.
Published: (2025)
Language Models Encode Numbers Using Digit Representations in Base 10
by: Levy, Amit Arnold, et al.
Published: (2024)
by: Levy, Amit Arnold, et al.
Published: (2024)
Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models
by: Levy, Mosh, et al.
Published: (2024)
by: Levy, Mosh, et al.
Published: (2024)
Why Are Linear RNNs More Parallelizable?
by: Merrill, William, et al.
Published: (2026)
by: Merrill, William, et al.
Published: (2026)
TransientTables: Evaluating LLMs' Reasoning on Temporally Evolving Semi-structured Tables
by: Shankarampeta, Abhilash, et al.
Published: (2025)
by: Shankarampeta, Abhilash, et al.
Published: (2025)
The Daring Dozen
Published: (2004)
Published: (2004)
Similar Items
-
NoviCode: Generating Programs from Natural Language Utterances by Novices
by: Mordechai, Asaf Achi, et al.
Published: (2024) -
LLMs' Reading Comprehension Is Affected by Parametric Knowledge and Struggles with Hypothetical Statements
by: Basmov, Victoria, et al.
Published: (2024) -
Simple Linguistic Inferences of Large Language Models (LLMs): Blind Spots and Blinds
by: Basmov, Victoria, et al.
Published: (2023) -
HeQ: a Large and Diverse Hebrew Reading Comprehension Benchmark
by: Cohen, Amir DN, et al.
Published: (2025) -
AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents
by: Trivedi, Harsh, et al.
Published: (2024)