Hybrid Human-LLM Corpus Construction and LLM Evaluation for Rare Linguistic Phenomena
Fuente:
arXiv
Saved in:
| Main Authors: | Weissweiler, Leonie, Köksal, Abdullatif, Schütze, Hinrich |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SYNTHEVAL: Hybrid Behavioral Testing of NLP Models with Synthetic CheckLists
by: Zhao, Raoyuan, et al.
Published: (2024)
by: Zhao, Raoyuan, et al.
Published: (2024)
Consistent Document-Level Relation Extraction via Counterfactuals
by: Modarressi, Ali, et al.
Published: (2024)
by: Modarressi, Ali, et al.
Published: (2024)
CRAFT Your Dataset: Task-Specific Synthetic Dataset Generation Through Corpus Retrieval and Augmentation
by: Ziegler, Ingo, et al.
Published: (2024)
by: Ziegler, Ingo, et al.
Published: (2024)
MemLLM: Finetuning LLMs to Use An Explicit Read-Write Memory
by: Modarressi, Ali, et al.
Published: (2024)
by: Modarressi, Ali, et al.
Published: (2024)
LongForm: Effective Instruction Tuning with Reverse Instructions
by: Köksal, Abdullatif, et al.
Published: (2023)
by: Köksal, Abdullatif, et al.
Published: (2023)
Verbing Weirds Language (Models): Evaluation of English Zero-Derivation in Five LLMs
by: Mortensen, David R., et al.
Published: (2024)
by: Mortensen, David R., et al.
Published: (2024)
Linguistic Generalizations are not Rules: Impacts on Evaluation of LMs
by: Weissweiler, Leonie, et al.
Published: (2025)
by: Weissweiler, Leonie, et al.
Published: (2025)
How far can bias go? Tracing bias from pretraining data to alignment
by: Thaler, Marion, et al.
Published: (2024)
by: Thaler, Marion, et al.
Published: (2024)
Constructions Are So Difficult That Even Large Language Models Get Them Right for the Wrong Reasons
by: Zhou, Shijia, et al.
Published: (2024)
by: Zhou, Shijia, et al.
Published: (2024)
TurkishMMLU: Measuring Massive Multitask Language Understanding in Turkish
by: Yüksel, Arda, et al.
Published: (2024)
by: Yüksel, Arda, et al.
Published: (2024)
Do We Know What LLMs Don't Know? A Study of Consistency in Knowledge Probing
by: Zhao, Raoyuan, et al.
Published: (2025)
by: Zhao, Raoyuan, et al.
Published: (2025)
Derivational Morphology Reveals Analogical Generalization in Large Language Models
by: Hofmann, Valentin, et al.
Published: (2024)
by: Hofmann, Valentin, et al.
Published: (2024)
MURI: High-Quality Instruction Tuning Datasets for Low-Resource Languages via Reverse Instructions
by: Köksal, Abdullatif, et al.
Published: (2024)
by: Köksal, Abdullatif, et al.
Published: (2024)
BabyLM's First Constructions: Causal probing provides a signal of learning
by: Rozner, Joshua, et al.
Published: (2025)
by: Rozner, Joshua, et al.
Published: (2025)
MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs
by: Jumelet, Jaap, et al.
Published: (2025)
by: Jumelet, Jaap, et al.
Published: (2025)
Constructions are Revealed in Word Distributions
by: Rozner, Joshua, et al.
Published: (2025)
by: Rozner, Joshua, et al.
Published: (2025)
Left, Right, or Center? Evaluating LLM Framing in News Classification and Generation
by: Kennedy, Molly, et al.
Published: (2026)
by: Kennedy, Molly, et al.
Published: (2026)
Do We Still Need Humans in the Loop? Comparing Human and LLM Annotation in Active Learning for Hostility Detection
by: Hakimi, Ahmad Dawar, et al.
Published: (2026)
by: Hakimi, Ahmad Dawar, et al.
Published: (2026)
Models Can and Should Embrace the Communicative Nature of Human-Generated Math
by: Boguraev, Sasha, et al.
Published: (2024)
by: Boguraev, Sasha, et al.
Published: (2024)
Language Models Learn Constructional Semantics, Not To Mention Syntax: Investigating LM Understanding of Paired-Focus Constructions
by: Scivetti, Wesley, et al.
Published: (2026)
by: Scivetti, Wesley, et al.
Published: (2026)
GlotCC: An Open Broad-Coverage CommonCrawl Corpus and Pipeline for Minority Languages
by: Kargaran, Amir Hossein, et al.
Published: (2024)
by: Kargaran, Amir Hossein, et al.
Published: (2024)
RET-LLM: Towards a General Read-Write Memory for Large Language Models
by: Modarressi, Ali, et al.
Published: (2023)
by: Modarressi, Ali, et al.
Published: (2023)
UCxn: Typologically Informed Annotation of Constructions Atop Universal Dependencies
by: Weissweiler, Leonie, et al.
Published: (2024)
by: Weissweiler, Leonie, et al.
Published: (2024)
HYPEROFA: Expanding LLM Vocabulary to New Languages via Hypernetwork-Based Embedding Initialization
by: Özeren, Enes, et al.
Published: (2025)
by: Özeren, Enes, et al.
Published: (2025)
Through the LLM Looking Glass: A Socratic Probing of Donkeys, Elephants, and Markets
by: Kennedy, Molly, et al.
Published: (2025)
by: Kennedy, Molly, et al.
Published: (2025)
Evaluating Contextually Mediated Factual Recall in Multilingual Large Language Models
by: Liu, Yihong, et al.
Published: (2026)
by: Liu, Yihong, et al.
Published: (2026)
From Rosetta to Match-Up: A Paired Corpus of Linguistic Puzzles with Human and LLM Benchmarks
by: Majmudar, Neh, et al.
Published: (2026)
by: Majmudar, Neh, et al.
Published: (2026)
GKnow: Measuring the Entanglement of Gender Bias and Factual Gender
by: Veloso, Leonor, et al.
Published: (2026)
by: Veloso, Leonor, et al.
Published: (2026)
Both Direct and Indirect Evidence Contribute to Dative Alternation Preferences in Language Models
by: Yao, Qing, et al.
Published: (2025)
by: Yao, Qing, et al.
Published: (2025)
Bring Your Own Knowledge: A Survey of Methods for LLM Knowledge Expansion
by: Wang, Mingyang, et al.
Published: (2025)
by: Wang, Mingyang, et al.
Published: (2025)
GLUScope: A Tool for Analyzing GLU Neurons in Transformer Language Models
by: Gerstner, Sebastian, et al.
Published: (2026)
by: Gerstner, Sebastian, et al.
Published: (2026)
Understanding Gated Neurons in Transformers from Their Input-Output Functionality
by: Gerstner, Sebastian, et al.
Published: (2025)
by: Gerstner, Sebastian, et al.
Published: (2025)
Decomposed Prompting: Probing Multilingual Linguistic Structure Knowledge in Large Language Models
by: Nie, Ercong, et al.
Published: (2024)
by: Nie, Ercong, et al.
Published: (2024)
Evaluate What You Can't Evaluate: Unassessable Quality for Generated Response
by: Liu, Yongkang, et al.
Published: (2023)
by: Liu, Yongkang, et al.
Published: (2023)
LLM in the Loop: Creating the ParaDeHate Dataset for Hate Speech Detoxification
by: Yuan, Shuzhou, et al.
Published: (2025)
by: Yuan, Shuzhou, et al.
Published: (2025)
Breaking the Script Barrier in Multilingual Pre-Trained Language Models with Transliteration-Based Post-Training Alignment
by: Xhelili, Orgest, et al.
Published: (2024)
by: Xhelili, Orgest, et al.
Published: (2024)
Mechanistic Understanding and Mitigation of Language Confusion in English-Centric Large Language Models
by: Nie, Ercong, et al.
Published: (2025)
by: Nie, Ercong, et al.
Published: (2025)
Why Better Cross-Lingual Alignment Fails for Better Cross-Lingual Transfer: Case of Encoders
by: Veitsman, Yana, et al.
Published: (2026)
by: Veitsman, Yana, et al.
Published: (2026)
The Anatomy of an Edit: Mechanism-Guided Activation Steering for Knowledge Editing
by: Cao, Yuan, et al.
Published: (2026)
by: Cao, Yuan, et al.
Published: (2026)
A Comprehensive Evaluation of Multilingual Chain-of-Thought Reasoning: Performance, Consistency, and Faithfulness Across Languages
by: Zhao, Raoyuan, et al.
Published: (2025)
by: Zhao, Raoyuan, et al.
Published: (2025)
Similar Items
-
SYNTHEVAL: Hybrid Behavioral Testing of NLP Models with Synthetic CheckLists
by: Zhao, Raoyuan, et al.
Published: (2024) -
Consistent Document-Level Relation Extraction via Counterfactuals
by: Modarressi, Ali, et al.
Published: (2024) -
CRAFT Your Dataset: Task-Specific Synthetic Dataset Generation Through Corpus Retrieval and Augmentation
by: Ziegler, Ingo, et al.
Published: (2024) -
MemLLM: Finetuning LLMs to Use An Explicit Read-Write Memory
by: Modarressi, Ali, et al.
Published: (2024) -
LongForm: Effective Instruction Tuning with Reverse Instructions
by: Köksal, Abdullatif, et al.
Published: (2023)