What is "Typological Diversity" in NLP?
Fuente:
arXiv
Saved in:
| Main Authors: | Ploeger, Esther, Poelman, Wessel, de Lhoneux, Miryam, Bjerva, Johannes |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Principled Framework for Evaluating on Typologically Diverse Languages
by: Ploeger, Esther, et al.
Published: (2024)
by: Ploeger, Esther, et al.
Published: (2024)
How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLP
by: Tatariya, Kushal, et al.
Published: (2024)
by: Tatariya, Kushal, et al.
Published: (2024)
Form and Meaning in Intrinsic Multilingual Evaluations
by: Poelman, Wessel, et al.
Published: (2026)
by: Poelman, Wessel, et al.
Published: (2026)
The Roles of English in Evaluating Multilingual Language Models
by: Poelman, Wessel, et al.
Published: (2024)
by: Poelman, Wessel, et al.
Published: (2024)
Multilingual Gradient Word-Order Typology from Universal Dependencies
by: Baylor, Emi, et al.
Published: (2024)
by: Baylor, Emi, et al.
Published: (2024)
QQ: A Toolkit for Language Identifiers and Metadata
by: Poelman, Wessel, et al.
Published: (2026)
by: Poelman, Wessel, et al.
Published: (2026)
On the Interplay between Positional Encodings, Morphological Complexity, and Word Order Flexibility
by: Tatariya, Kushal, et al.
Published: (2025)
by: Tatariya, Kushal, et al.
Published: (2025)
Confounding Factors in Relating Model Performance to Morphology
by: Poelman, Wessel, et al.
Published: (2025)
by: Poelman, Wessel, et al.
Published: (2025)
Typologically Informed Parameter Aggregation
by: Accou, Stef, et al.
Published: (2026)
by: Accou, Stef, et al.
Published: (2026)
Sociolinguistically Informed Interpretability: A Case Study on Hinglish Emotion Classification
by: Tatariya, Kushal, et al.
Published: (2024)
by: Tatariya, Kushal, et al.
Published: (2024)
We Need to Measure Data Diversity in NLP -- Better and Broader
by: Nguyen, Dong, et al.
Published: (2025)
by: Nguyen, Dong, et al.
Published: (2025)
Type and Complexity Signals in Multilingual Question Representations
by: Kokot, Robin, et al.
Published: (2025)
by: Kokot, Robin, et al.
Published: (2025)
Recipe for Zero-shot POS Tagging: Is It Useful in Realistic Scenarios?
by: Vandenbulcke, Zeno, et al.
Published: (2024)
by: Vandenbulcke, Zeno, et al.
Published: (2024)
Trans-Tokenization and Cross-lingual Vocabulary Transfers: Language Adaptation of LLMs for Low-Resource NLP
by: Remy, François, et al.
Published: (2024)
by: Remy, François, et al.
Published: (2024)
Against All Odds: Overcoming Typology, Script, and Language Confusion in Multilingual Embedding Inversion Attacks
by: Chen, Yiyi, et al.
Published: (2024)
by: Chen, Yiyi, et al.
Published: (2024)
Pixology: Probing the Linguistic and Visual Capabilities of Pixel-based Language Models
by: Tatariya, Kushal, et al.
Published: (2024)
by: Tatariya, Kushal, et al.
Published: (2024)
Towards Tailored Recovery of Lexical Diversity in Literary Machine Translation
by: Ploeger, Esther, et al.
Published: (2024)
by: Ploeger, Esther, et al.
Published: (2024)
NLP Security and Ethics, in the Wild
by: Lent, Heather, et al.
Published: (2025)
by: Lent, Heather, et al.
Published: (2025)
Knowledge Graphs, Large Language Models, and Hallucinations: An NLP Perspective
by: Lavrinovics, Ernests, et al.
Published: (2024)
by: Lavrinovics, Ernests, et al.
Published: (2024)
Large Language Models are Easily Confused: A Quantitative Metric, Security Implications and Typological Analysis
by: Chen, Yiyi, et al.
Published: (2024)
by: Chen, Yiyi, et al.
Published: (2024)
Engineering Conversational Search Systems: A Review of Applications, Architectures, and Functional Components
by: Schneider, Phillip, et al.
Published: (2024)
by: Schneider, Phillip, et al.
Published: (2024)
Patterns of Persistence and Diffusibility across the World's Languages
by: Chen, Yiyi, et al.
Published: (2024)
by: Chen, Yiyi, et al.
Published: (2024)
Linguistically Grounded Analysis of Language Models using Shapley Head Values
by: Fekete, Marcell, et al.
Published: (2024)
by: Fekete, Marcell, et al.
Published: (2024)
CreoleVal: Multilingual Multitask Benchmarks for Creoles
by: Lent, Heather, et al.
Published: (2023)
by: Lent, Heather, et al.
Published: (2023)
When Discourse Pressures Conflict: Information Structure in Vision-Language Model Outputs
by: Fekete, Marcell, et al.
Published: (2026)
by: Fekete, Marcell, et al.
Published: (2026)
Multi-perspective Alignment for Increasing Naturalness in Neural Machine Translation
by: Lai, Huiyuan, et al.
Published: (2024)
by: Lai, Huiyuan, et al.
Published: (2024)
Follow the Path: Reasoning over Knowledge Graph Paths to Improve Large Language Model Factuality
by: Zhang, Mike, et al.
Published: (2025)
by: Zhang, Mike, et al.
Published: (2025)
NLP needs Diversity outside of 'Diversity'
by: Tint, Joshua
Published: (2026)
by: Tint, Joshua
Published: (2026)
Characterizing Memorization in Diffusion Language Models: Generalized Extraction and Sampling Effects
by: Luo, Xiaoyu, et al.
Published: (2026)
by: Luo, Xiaoyu, et al.
Published: (2026)
MultiHal: Multilingual Dataset for Knowledge-Graph Grounded Evaluation of LLM Hallucinations
by: Lavrinovics, Ernests, et al.
Published: (2025)
by: Lavrinovics, Ernests, et al.
Published: (2025)
Text Embedding Inversion Security for Multilingual Language Models
by: Chen, Yiyi, et al.
Published: (2024)
by: Chen, Yiyi, et al.
Published: (2024)
Do LLMs Really Memorize Personally Identifiable Information? Revisiting PII Leakage with a Cue-Controlled Memorization Framework
by: Luo, Xiaoyu, et al.
Published: (2026)
by: Luo, Xiaoyu, et al.
Published: (2026)
Shared Path: Unraveling Memorization in Multilingual LLMs through Language Similarities
by: Luo, Xiaoyu, et al.
Published: (2025)
by: Luo, Xiaoyu, et al.
Published: (2025)
The Responsible Development of Automated Student Feedback with Generative AI
by: Lindsay, Euan D, et al.
Published: (2023)
by: Lindsay, Euan D, et al.
Published: (2023)
Designing NLP Systems That Adapt to Diverse Worldviews
by: Creanga, Claudiu, et al.
Published: (2024)
by: Creanga, Claudiu, et al.
Published: (2024)
ALGEN: Few-shot Inversion Attacks on Textual Embeddings using Alignment and Generation
by: Chen, Yiyi, et al.
Published: (2025)
by: Chen, Yiyi, et al.
Published: (2025)
The Fragility of Chain-of-Thought Monitoring Across Typologically Diverse Languages
by: Onyame, Eric, et al.
Published: (2026)
by: Onyame, Eric, et al.
Published: (2026)
On-Device LLMs for Home Assistant: Dual Role in Intent Detection and Response Generation
by: Birkmose, Rune, et al.
Published: (2025)
by: Birkmose, Rune, et al.
Published: (2025)
Large Language Models Share Representations of Latent Grammatical Concepts Across Typologically Diverse Languages
by: Brinkmann, Jannik, et al.
Published: (2025)
by: Brinkmann, Jannik, et al.
Published: (2025)
Computational Typology
by: Jäger, Gerhard
Published: (2025)
by: Jäger, Gerhard
Published: (2025)
Similar Items
-
A Principled Framework for Evaluating on Typologically Diverse Languages
by: Ploeger, Esther, et al.
Published: (2024) -
How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLP
by: Tatariya, Kushal, et al.
Published: (2024) -
Form and Meaning in Intrinsic Multilingual Evaluations
by: Poelman, Wessel, et al.
Published: (2026) -
The Roles of English in Evaluating Multilingual Language Models
by: Poelman, Wessel, et al.
Published: (2024) -
Multilingual Gradient Word-Order Typology from Universal Dependencies
by: Baylor, Emi, et al.
Published: (2024)