BenchCLAMP: A Benchmark for Evaluating Language Models on Syntactic and Semantic Parsing
Fuente:
arXiv
Saved in:
| Main Authors: | Roy, Subhro, Thomson, Sam, Chen, Tongfei, Shin, Richard, Pauls, Adam, Eisner, Jason, Van Durme, Benjamin |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning to Retrieve Iteratively for In-Context Learning
by: Chen, Yunmo, et al.
Published: (2024)
by: Chen, Yunmo, et al.
Published: (2024)
MICE for CATs: Model-Internal Confidence Estimation for Calibrating Agents with Tools
by: Subramani, Nishant, et al.
Published: (2025)
by: Subramani, Nishant, et al.
Published: (2025)
Zero and Few-shot Semantic Parsing with Ambiguous Inputs
by: Stengel-Eskin, Elias, et al.
Published: (2023)
by: Stengel-Eskin, Elias, et al.
Published: (2023)
Hierarchical corpus encoder: Fusing generative retrieval and dense indices
by: Chen, Tongfei, et al.
Published: (2025)
by: Chen, Tongfei, et al.
Published: (2025)
LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts
by: Hashemi, Helia, et al.
Published: (2024)
by: Hashemi, Helia, et al.
Published: (2024)
Multi-Field Adaptive Retrieval
by: Li, Millicent, et al.
Published: (2024)
by: Li, Millicent, et al.
Published: (2024)
Do Androids Know They're Only Dreaming of Electric Sheep?
by: CH-Wang, Sky, et al.
Published: (2023)
by: CH-Wang, Sky, et al.
Published: (2023)
Interpreting User Requests in the Context of Natural Language Standing Instructions
by: Moghe, Nikita, et al.
Published: (2023)
by: Moghe, Nikita, et al.
Published: (2023)
LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error
by: Wang, Boshi, et al.
Published: (2024)
by: Wang, Boshi, et al.
Published: (2024)
Syntactic and Semantic Control of Large Language Models via Sequential Monte Carlo
by: Loula, João, et al.
Published: (2025)
by: Loula, João, et al.
Published: (2025)
Automating the Analysis of Parsing Algorithms (and other Dynamic Programs)
by: Vieira, Tim, et al.
Published: (2025)
by: Vieira, Tim, et al.
Published: (2025)
DeonticBench: A Benchmark for Reasoning over Rules
by: Dou, Guangyao, et al.
Published: (2026)
by: Dou, Guangyao, et al.
Published: (2026)
RE-Adapt: Reverse Engineered Adaptation of Large Language Models
by: Fleshman, William, et al.
Published: (2024)
by: Fleshman, William, et al.
Published: (2024)
A Joint Multitask Model for Morpho-Syntactic Parsing
by: Inostroza, Demian, et al.
Published: (2025)
by: Inostroza, Demian, et al.
Published: (2025)
RenoBench: A Citation Parsing Benchmark
by: Sarin, Parth, et al.
Published: (2026)
by: Sarin, Parth, et al.
Published: (2026)
Modelling Child Learning and Parsing of Long-range Syntactic Dependencies
by: Mahon, Louis, et al.
Published: (2025)
by: Mahon, Louis, et al.
Published: (2025)
A Study on How Attention Scores in the BERT Model are Aware of Lexical Categories in Syntactic and Semantic Tasks on the GLUE Benchmark
by: Jang, Dongjun, et al.
Published: (2024)
by: Jang, Dongjun, et al.
Published: (2024)
LoRA-Augmented Generation (LAG) for Knowledge-Intensive Language Tasks
by: Fleshman, William, et al.
Published: (2025)
by: Fleshman, William, et al.
Published: (2025)
A Systematic Comparison of Syntactic Representations of Dependency Parsing
by: Wisniewski, Guillaume, et al.
Published: (2025)
by: Wisniewski, Guillaume, et al.
Published: (2025)
Evaluating Semantic and Syntactic Understanding in Large Language Models for Payroll Systems
by: Maclean, Hendrika, et al.
Published: (2026)
by: Maclean, Hendrika, et al.
Published: (2026)
Language Models and Logic Programs for Trustworthy Tax Reasoning
by: Jurayj, William, et al.
Published: (2025)
by: Jurayj, William, et al.
Published: (2025)
Localization vs. Semantics: Visual Representations in Unimodal and Multimodal Models
by: Li, Zhuowan, et al.
Published: (2022)
by: Li, Zhuowan, et al.
Published: (2022)
SCOPE: Tree-based Self-Correcting Online Log Parsing via Syntactic-Semantic Collaboration
by: Fan, Dongyi, et al.
Published: (2026)
by: Fan, Dongyi, et al.
Published: (2026)
arXiv2Table: Toward Realistic Benchmarking and Evaluation for LLM-Based Literature-Review Table Generation
by: Wang, Weiqi, et al.
Published: (2025)
by: Wang, Weiqi, et al.
Published: (2025)
Compressed Chain of Thought: Efficient Reasoning Through Dense Representations
by: Cheng, Jeffrey, et al.
Published: (2024)
by: Cheng, Jeffrey, et al.
Published: (2024)
Let's Think Var-by-Var: Large Language Models Enable Ad Hoc Probabilistic Reasoning
by: Xia, Shepard, et al.
Published: (2024)
by: Xia, Shepard, et al.
Published: (2024)
CAIT: A Syntactic Parsing Toolkit for Child-Adult InTeractions
by: Padovani, Francesca, et al.
Published: (2026)
by: Padovani, Francesca, et al.
Published: (2026)
Compactor: Calibrated Query-Agnostic KV Cache Compression with Approximate Leverage Scores
by: Chari, Vivek, et al.
Published: (2025)
by: Chari, Vivek, et al.
Published: (2025)
Accelerating Language Model Workflows with Prompt Choreography
by: Bai, TJ, et al.
Published: (2025)
by: Bai, TJ, et al.
Published: (2025)
Tur[k]ingBench: A Challenge Benchmark for Web Agents
by: Xu, Kevin, et al.
Published: (2024)
by: Xu, Kevin, et al.
Published: (2024)
Controlled Evaluation of Syntactic Knowledge in Multilingual Language Models
by: Kryvosheieva, Daria, et al.
Published: (2024)
by: Kryvosheieva, Daria, et al.
Published: (2024)
SEQR: Secure and Efficient QR-based LoRA Routing
by: Fleshman, William, et al.
Published: (2025)
by: Fleshman, William, et al.
Published: (2025)
SpectR: Dynamically Composing LM Experts with Spectral Routing
by: Fleshman, William, et al.
Published: (2025)
by: Fleshman, William, et al.
Published: (2025)
Urdu Dependency Parsing and Treebank Development: A Syntactic and Morphological Perspective
by: Habib, Nudrat
Published: (2024)
by: Habib, Nudrat
Published: (2024)
Verifiable by Design: Aligning Language Models to Quote from Pre-Training Data
by: Zhang, Jingyu, et al.
Published: (2024)
by: Zhang, Jingyu, et al.
Published: (2024)
Targeted Syntactic Evaluation of Language Models on Georgian Case Alignment
by: Gallagher, Daniel, et al.
Published: (2026)
by: Gallagher, Daniel, et al.
Published: (2026)
LLMs Provide Unstable Answers to Legal Questions
by: Blair-Stanek, Andrew, et al.
Published: (2025)
by: Blair-Stanek, Andrew, et al.
Published: (2025)
The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models
by: Kim, Seungone, et al.
Published: (2024)
by: Kim, Seungone, et al.
Published: (2024)
RE-AdaptIR: Improving Information Retrieval through Reverse Engineered Adaptation
by: Fleshman, William, et al.
Published: (2024)
by: Fleshman, William, et al.
Published: (2024)
BLT: Can Large Language Models Handle Basic Legal Text?
by: Blair-Stanek, Andrew, et al.
Published: (2023)
by: Blair-Stanek, Andrew, et al.
Published: (2023)
Similar Items
-
Learning to Retrieve Iteratively for In-Context Learning
by: Chen, Yunmo, et al.
Published: (2024) -
MICE for CATs: Model-Internal Confidence Estimation for Calibrating Agents with Tools
by: Subramani, Nishant, et al.
Published: (2025) -
Zero and Few-shot Semantic Parsing with Ambiguous Inputs
by: Stengel-Eskin, Elias, et al.
Published: (2023) -
Hierarchical corpus encoder: Fusing generative retrieval and dense indices
by: Chen, Tongfei, et al.
Published: (2025) -
LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts
by: Hashemi, Helia, et al.
Published: (2024)