Evaluating Morphological Compositional Generalization in Large Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Ismayilzada, Mete, Circi, Defne, Sälevä, Jonne, Sirin, Hale, Köksal, Abdullatif, Dhingra, Bhuwan, Bosselut, Antoine, Ataman, Duygu, van der Plas, Lonneke |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Creativity in AI: Progresses and Challenges
por: Ismayilzada, Mete, et al.
Publicado: (2024)
por: Ismayilzada, Mete, et al.
Publicado: (2024)
Evaluating Creative Short Story Generation in Humans and Large Language Models
por: Ismayilzada, Mete, et al.
Publicado: (2024)
por: Ismayilzada, Mete, et al.
Publicado: (2024)
Beyond statistical significance: Quantifying uncertainty and statistical variability in multilingual and multitask NLP evaluation
por: Sälevä, Jonne, et al.
Publicado: (2025)
por: Sälevä, Jonne, et al.
Publicado: (2025)
CresOWLve: Benchmarking Creative Problem-Solving Over Real-World Knowledge
por: Ismayilzada, Mete, et al.
Publicado: (2026)
por: Ismayilzada, Mete, et al.
Publicado: (2026)
Extracting Polymer Nanocomposite Samples from Full-Length Documents
por: Khalighinejad, Ghazal, et al.
Publicado: (2024)
por: Khalighinejad, Ghazal, et al.
Publicado: (2024)
Creative Preference Optimization
por: Ismayilzada, Mete, et al.
Publicado: (2025)
por: Ismayilzada, Mete, et al.
Publicado: (2025)
Large Language Models Align with the Human Brain during Creative Thinking
por: Ismayilzada, Mete, et al.
Publicado: (2026)
por: Ismayilzada, Mete, et al.
Publicado: (2026)
The Effectiveness of Morphology-aware Segmentation in Low-Resource Neural Machine Translation
por: Sälevä, Jonne, et al.
Publicado: (2021)
por: Sälevä, Jonne, et al.
Publicado: (2021)
ParaNames 1.0: Creating an Entity Name Corpus for 400+ Languages using Wikidata
por: Sälevä, Jonne, et al.
Publicado: (2024)
por: Sälevä, Jonne, et al.
Publicado: (2024)
From 124 Million Tokens to 1,021 Neologisms: A Large-Scale Pipeline for Automatic Neologism Detection
por: Rossini, Diego, et al.
Publicado: (2026)
por: Rossini, Diego, et al.
Publicado: (2026)
Understanding the effects of language-specific class imbalance in multilingual fine-tuning
por: Jung, Vincent, et al.
Publicado: (2024)
por: Jung, Vincent, et al.
Publicado: (2024)
Binary Token-Level Classification with DeBERTa for All-Type MWE Identification: A Lightweight Approach with Linguistic Enhancement
por: Rossini, Diego, et al.
Publicado: (2026)
por: Rossini, Diego, et al.
Publicado: (2026)
DiffuCOMET: Contextual Commonsense Knowledge Diffusion
por: Gao, Silin, et al.
Publicado: (2024)
por: Gao, Silin, et al.
Publicado: (2024)
Can language models learn analogical reasoning? Investigating training objectives and comparisons to human performance
por: Petersen, Molly R., et al.
Publicado: (2023)
por: Petersen, Molly R., et al.
Publicado: (2023)
Loose and Tight: Creative Formation but Rigid Use of Nominal Compounds in Conspiracist Texts
por: Alessandro Miani, et al.
Publicado: (2024)
por: Alessandro Miani, et al.
Publicado: (2024)
REFINER: Reasoning Feedback on Intermediate Representations
por: Paul, Debjit, et al.
Publicado: (2023)
por: Paul, Debjit, et al.
Publicado: (2023)
Exploring Defeasibility in Causal Reasoning
por: Cui, Shaobo, et al.
Publicado: (2024)
por: Cui, Shaobo, et al.
Publicado: (2024)
The Ouroboros of Benchmarking: Reasoning Evaluation in an Era of Saturation
por: Deveci, İbrahim Ethem, et al.
Publicado: (2025)
por: Deveci, İbrahim Ethem, et al.
Publicado: (2025)
HintsOfTruth: A Multimodal Checkworthiness Detection Dataset with Real and Synthetic Claims
por: van der Meer, Michiel, et al.
Publicado: (2025)
por: van der Meer, Michiel, et al.
Publicado: (2025)
Hybrid Human-LLM Corpus Construction and LLM Evaluation for Rare Linguistic Phenomena
por: Weissweiler, Leonie, et al.
Publicado: (2024)
por: Weissweiler, Leonie, et al.
Publicado: (2024)
GenEOL: Harnessing the Generative Power of LLMs for Training-Free Sentence Embeddings
por: Thirukovalluru, Raghuveer, et al.
Publicado: (2024)
por: Thirukovalluru, Raghuveer, et al.
Publicado: (2024)
Modelling Analogies and Analogical Reasoning: Connecting Cognitive Science Theory and NLP Research
por: Petersen, Molly R, et al.
Publicado: (2025)
por: Petersen, Molly R, et al.
Publicado: (2025)
ChatShop: Interactive Information Seeking with Language Agents
por: Chen, Sanxing, et al.
Publicado: (2024)
por: Chen, Sanxing, et al.
Publicado: (2024)
To Trust or Not to Trust? Enhancing Large Language Models' Situated Faithfulness to External Contexts
por: Huang, Yukun, et al.
Publicado: (2024)
por: Huang, Yukun, et al.
Publicado: (2024)
OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ Languages
por: Palen-Michel, Chester, et al.
Publicado: (2024)
por: Palen-Michel, Chester, et al.
Publicado: (2024)
Vision2Code: A Multi-Domain Benchmark for Evaluating Image-to-Code Generation
por: Periasami, Ajay Vikram, et al.
Publicado: (2026)
por: Periasami, Ajay Vikram, et al.
Publicado: (2026)
Calibrating Long-form Generations from Large Language Models
por: Huang, Yukun, et al.
Publicado: (2024)
por: Huang, Yukun, et al.
Publicado: (2024)
Consistent Document-Level Relation Extraction via Counterfactuals
por: Modarressi, Ali, et al.
Publicado: (2024)
por: Modarressi, Ali, et al.
Publicado: (2024)
Cite Pretrain: Retrieval-Free Knowledge Attribution for Large Language Models
por: Huang, Yukun, et al.
Publicado: (2025)
por: Huang, Yukun, et al.
Publicado: (2025)
Real-time Factuality Assessment from Adversarial Feedback
por: Chen, Sanxing, et al.
Publicado: (2024)
por: Chen, Sanxing, et al.
Publicado: (2024)
Atomic Self-Consistency for Better Long Form Generations
por: Thirukovalluru, Raghuveer, et al.
Publicado: (2024)
por: Thirukovalluru, Raghuveer, et al.
Publicado: (2024)
RVPO: Risk-Sensitive Alignment via Variance Regularization
por: Montero, Ivan, et al.
Publicado: (2026)
por: Montero, Ivan, et al.
Publicado: (2026)
Fuzzy Speculative Decoding for a Tunable Accuracy-Runtime Tradeoff
por: Holsman, Maximilian, et al.
Publicado: (2025)
por: Holsman, Maximilian, et al.
Publicado: (2025)
Detecting Structured Language Alternations in Historical Documents by Combining Language Identification with Fourier Analysis
por: Sirin, Hale, et al.
Publicado: (2024)
por: Sirin, Hale, et al.
Publicado: (2024)
Generalization Measures for Zero-Shot Cross-Lingual Transfer
por: Bassi, Saksham, et al.
Publicado: (2024)
por: Bassi, Saksham, et al.
Publicado: (2024)
Dynamic embedded topic models and change-point detection for exploring literary-historical hypotheses
por: Sirin, Hale, et al.
Publicado: (2024)
por: Sirin, Hale, et al.
Publicado: (2024)
Generalizability of Large Language Model-Based Agents: A Comprehensive Survey
por: Zhang, Minxing, et al.
Publicado: (2025)
por: Zhang, Minxing, et al.
Publicado: (2025)
The Challenges of Evaluating LLM Applications: An Analysis of Automated, Human, and LLM-Based Approaches
por: Abeysinghe, Bhashithe, et al.
Publicado: (2024)
por: Abeysinghe, Bhashithe, et al.
Publicado: (2024)
Hierarchical Multi-Label Classification of Online Vaccine Concerns
por: Zhu, Chloe Qinyu, et al.
Publicado: (2024)
por: Zhu, Chloe Qinyu, et al.
Publicado: (2024)
TUMLU: A Unified and Native Language Understanding Benchmark for Turkic Languages
por: Isbarov, Jafar, et al.
Publicado: (2025)
por: Isbarov, Jafar, et al.
Publicado: (2025)
Ejemplares similares
-
Creativity in AI: Progresses and Challenges
por: Ismayilzada, Mete, et al.
Publicado: (2024) -
Evaluating Creative Short Story Generation in Humans and Large Language Models
por: Ismayilzada, Mete, et al.
Publicado: (2024) -
Beyond statistical significance: Quantifying uncertainty and statistical variability in multilingual and multitask NLP evaluation
por: Sälevä, Jonne, et al.
Publicado: (2025) -
CresOWLve: Benchmarking Creative Problem-Solving Over Real-World Knowledge
por: Ismayilzada, Mete, et al.
Publicado: (2026) -
Extracting Polymer Nanocomposite Samples from Full-Length Documents
por: Khalighinejad, Ghazal, et al.
Publicado: (2024)