Archer: A Human-Labeled Text-to-SQL Dataset with Arithmetic, Commonsense and Hypothetical Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Zheng, Danna, Lapata, Mirella, Pan, Jeff Z. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Reliable are LLMs as Knowledge Bases? Re-thinking Facutality and Consistency
by: Zheng, Danna, et al.
Published: (2024)
by: Zheng, Danna, et al.
Published: (2024)
Long-Form Information Alignment Evaluation Beyond Atomic Facts
by: Zheng, Danna, et al.
Published: (2025)
by: Zheng, Danna, et al.
Published: (2025)
TrustScore: Reference-Free Evaluation of LLM Response Trustworthiness
by: Zheng, Danna, et al.
Published: (2024)
by: Zheng, Danna, et al.
Published: (2024)
Masking in Multi-hop QA: An Analysis of How Language Models Perform with Context Permutation
by: Huang, Wenyu, et al.
Published: (2025)
by: Huang, Wenyu, et al.
Published: (2025)
Learning to Reason for Long-Form Story Generation
by: Gurung, Alexander, et al.
Published: (2025)
by: Gurung, Alexander, et al.
Published: (2025)
Rethinking Memory in LLM based Agents: Representations, Operations, and Emerging Topics
by: Du, Yiming, et al.
Published: (2025)
by: Du, Yiming, et al.
Published: (2025)
Integrating Large Language Models with Graph-based Reasoning for Conversational Question Answering
by: Jain, Parag, et al.
Published: (2024)
by: Jain, Parag, et al.
Published: (2024)
Reasoning about Intent for Ambiguous Requests
by: Saparina, Irina, et al.
Published: (2025)
by: Saparina, Irina, et al.
Published: (2025)
Debating for Better Reasoning: An Unsupervised Multimodal Approach
by: Adhikari, Ashutosh, et al.
Published: (2025)
by: Adhikari, Ashutosh, et al.
Published: (2025)
Lightweight Latent Reasoning for Narrative Tasks
by: Gurung, Alexander, et al.
Published: (2025)
by: Gurung, Alexander, et al.
Published: (2025)
PixT3: Pixel-based Table-To-Text Generation
by: Alonso, Iñigo, et al.
Published: (2023)
by: Alonso, Iñigo, et al.
Published: (2023)
A Modular Approach for Multimodal Summarization of TV Shows
by: Mahon, Louis, et al.
Published: (2024)
by: Mahon, Louis, et al.
Published: (2024)
AMBROSIA: A Benchmark for Parsing Ambiguous Questions into Database Queries
by: Saparina, Irina, et al.
Published: (2024)
by: Saparina, Irina, et al.
Published: (2024)
CHIRON: Rich Character Representations in Long-Form Narratives
by: Gurung, Alexander, et al.
Published: (2024)
by: Gurung, Alexander, et al.
Published: (2024)
Improving Generalization in Semantic Parsing by Increasing Natural Language Variation
by: Saparina, Irina, et al.
Published: (2024)
by: Saparina, Irina, et al.
Published: (2024)
Context-Aware Hierarchical Merging for Long Document Summarization
by: Ou, Litu, et al.
Published: (2025)
by: Ou, Litu, et al.
Published: (2025)
Prompting Large Language Models with Knowledge Graphs for Question Answering Involving Long-tail Facts
by: Huang, Wenyu, et al.
Published: (2024)
by: Huang, Wenyu, et al.
Published: (2024)
Less is More: Making Smaller Language Models Competent Subgraph Retrievers for Multi-hop KGQA
by: Huang, Wenyu, et al.
Published: (2024)
by: Huang, Wenyu, et al.
Published: (2024)
BookWorm: A Dataset for Character Description and Analysis
by: Papoudakis, Argyrios, et al.
Published: (2024)
by: Papoudakis, Argyrios, et al.
Published: (2024)
Disambiguate First, Parse Later: Generating Interpretations for Ambiguity Resolution in Semantic Parsing
by: Saparina, Irina, et al.
Published: (2025)
by: Saparina, Irina, et al.
Published: (2025)
TABLET: A Large-Scale Dataset for Robust Visual Table Understanding
by: Alonso, Iñigo, et al.
Published: (2025)
by: Alonso, Iñigo, et al.
Published: (2025)
Uncertainty Quantification in Retrieval Augmented Question Answering
by: Perez-Beltrachini, Laura, et al.
Published: (2025)
by: Perez-Beltrachini, Laura, et al.
Published: (2025)
Think Before you Write: QA-Guided Reasoning for Character Descriptions in Books
by: Papoudakis, Argyrios, et al.
Published: (2026)
by: Papoudakis, Argyrios, et al.
Published: (2026)
Hierarchical Indexing for Retrieval-Augmented Opinion Summarization
by: Hosking, Tom, et al.
Published: (2024)
by: Hosking, Tom, et al.
Published: (2024)
Evaluating LLMs for Targeted Concept Simplification for Domain-Specific Texts
by: Asthana, Sumit, et al.
Published: (2024)
by: Asthana, Sumit, et al.
Published: (2024)
Long-Context Reasoning Through Proxy-Based Chain-of-Thought Tuning
by: Li, Miao, et al.
Published: (2026)
by: Li, Miao, et al.
Published: (2026)
What Is That Talk About? A Video-to-Text Summarization Dataset for Scientific Presentations
by: Liu, Dongqi, et al.
Published: (2025)
by: Liu, Dongqi, et al.
Published: (2025)
Generating Visual Stories with Grounded and Coreferent Characters
by: Liu, Danyang, et al.
Published: (2024)
by: Liu, Danyang, et al.
Published: (2024)
BUCA: A Binary Classification Approach to Unsupervised Commonsense Question Answering
by: He, Jie, et al.
Published: (2023)
by: He, Jie, et al.
Published: (2023)
GraphLit: Learning Text-Enriched Dynamic Character Network Representations for Literary Study
by: Michel, Gaspard, et al.
Published: (2026)
by: Michel, Gaspard, et al.
Published: (2026)
Compositional Generalisation for Explainable Hate Speech Detection
by: Calabrese, Agostina, et al.
Published: (2025)
by: Calabrese, Agostina, et al.
Published: (2025)
Help Me Write a Story: Evaluating LLMs' Ability to Generate Writing Feedback
by: Rashkin, Hannah, et al.
Published: (2025)
by: Rashkin, Hannah, et al.
Published: (2025)
Meta-Adaptive Prompt Distillation for Few-Shot Visual Question Answering
by: Gupta, Akash, et al.
Published: (2025)
by: Gupta, Akash, et al.
Published: (2025)
Improving Retrieval-augmented Text-to-SQL with AST-based Ranking and Schema Pruning
by: Shen, Zhili, et al.
Published: (2024)
by: Shen, Zhili, et al.
Published: (2024)
STaR-SQL: Self-Taught Reasoner for Text-to-SQL
by: He, Mingqian, et al.
Published: (2025)
by: He, Mingqian, et al.
Published: (2025)
SimLM: Can Language Models Infer Parameters of Physical Systems?
by: Memery, Sean, et al.
Published: (2023)
by: Memery, Sean, et al.
Published: (2023)
Learning to Plan and Generate Text with Citations
by: Fierro, Constanza, et al.
Published: (2024)
by: Fierro, Constanza, et al.
Published: (2024)
Instances and Labels: Hierarchy-aware Joint Supervised Contrastive Learning for Hierarchical Multi-Label Text Classification
by: Yu, Simon, et al.
Published: (2023)
by: Yu, Simon, et al.
Published: (2023)
Extended Japanese Commonsense Morality Dataset with Masked Token and Label Enhancement
by: Ohashi, Takumi, et al.
Published: (2024)
by: Ohashi, Takumi, et al.
Published: (2024)
Ko-PIQA: A Korean Physical Commonsense Reasoning Dataset with Cultural Context
by: Choi, Dasol, et al.
Published: (2025)
by: Choi, Dasol, et al.
Published: (2025)
Similar Items
-
How Reliable are LLMs as Knowledge Bases? Re-thinking Facutality and Consistency
by: Zheng, Danna, et al.
Published: (2024) -
Long-Form Information Alignment Evaluation Beyond Atomic Facts
by: Zheng, Danna, et al.
Published: (2025) -
TrustScore: Reference-Free Evaluation of LLM Response Trustworthiness
by: Zheng, Danna, et al.
Published: (2024) -
Masking in Multi-hop QA: An Analysis of How Language Models Perform with Context Permutation
by: Huang, Wenyu, et al.
Published: (2025) -
Learning to Reason for Long-Form Story Generation
by: Gurung, Alexander, et al.
Published: (2025)