Using Correspondence Patterns to Identify Irregular Words in Cognate sets Through Leave-One-Out Validation
Fuente:
arXiv
Saved in:
| Main Authors: | Blum, Frederic, List, Johann-Mattis |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Trimming Phonetic Alignments Improves the Inference of Sound Correspondence Patterns from Multilingual Wordlists
by: Blum, Frederic, et al.
Published: (2023)
by: Blum, Frederic, et al.
Published: (2023)
From Isolates to Families: Using Neural Networks for Automated Language Affiliation
by: Blum, Frederic, et al.
Published: (2025)
by: Blum, Frederic, et al.
Published: (2025)
Everybody Likes to Sleep: A Computer-Assisted Comparison of Object Naming Data from 30 Languages
by: Kučerová, Alžběta, et al.
Published: (2025)
by: Kučerová, Alžběta, et al.
Published: (2025)
A Computational Model for the Assessment of Mutual Intelligibility Among Closely Related Languages
by: Nieder, Jessica, et al.
Published: (2024)
by: Nieder, Jessica, et al.
Published: (2024)
Partial Colexifications Improve Concept Embeddings
by: Rubehn, Arne, et al.
Published: (2025)
by: Rubehn, Arne, et al.
Published: (2025)
From Psycholinguistics to Computer Vision. A Comprehensive Review of Object Naming Data and Studies
by: Alžběta Kučerová, et al.
Published: (2026)
by: Alžběta Kučerová, et al.
Published: (2026)
Generating Feature Vectors from Phonetic Transcriptions in Cross-Linguistic Data Formats
by: Rubehn, Arne, et al.
Published: (2024)
by: Rubehn, Arne, et al.
Published: (2024)
Advancing the Database of Cross-Linguistic Colexifications with New Workflows and Data
by: Tjuka, Annika, et al.
Published: (2025)
by: Tjuka, Annika, et al.
Published: (2025)
Unstable Grounds for Beautiful Trees? Testing the Robustness of Concept Translations in the Compilation of Multilingual Wordlists
by: Snee, David, et al.
Published: (2025)
by: Snee, David, et al.
Published: (2025)
Are Sounds Sound for Phylogenetic Reconstruction?
by: Häuser, Luise, et al.
Published: (2024)
by: Häuser, Luise, et al.
Published: (2024)
The Cognate Data Bottleneck in Language Phylogenetics
by: Häuser, Luise, et al.
Published: (2025)
by: Häuser, Luise, et al.
Published: (2025)
Automated Cognate Detection as a Supervised Link Prediction Task with Cognate Transformer
by: Akavarapu, V. S. D. S. Mahesh, et al.
Published: (2024)
by: Akavarapu, V. S. D. S. Mahesh, et al.
Published: (2024)
Data and Code Accompanying the Study "Unstable Grounds for Beautiful Trees? Testing the Robustness of Concept Translations in the Compilation of Multilingual Wordlists"
by: Snee, David, et al.
Published: (2025)
by: Snee, David, et al.
Published: (2025)
LooComp: Leverage Leave-One-Out Strategy to Encoder-only Transformer for Efficient Query-aware Context Compression
by: Do, Thao, et al.
Published: (2026)
by: Do, Thao, et al.
Published: (2026)
Computational Approaches for Integrating out Subjectivity in Cognate Synonym Selection
by: Häuser, Luise, et al.
Published: (2024)
by: Häuser, Luise, et al.
Published: (2024)
Annotating and Inferring Compositional Structures in Numeral Systems Across Languages
by: Rubehn, Arne, et al.
Published: (2025)
by: Rubehn, Arne, et al.
Published: (2025)
Over-representation of phonological features in basic vocabulary doesn't replicate when controlling for spatial and phylogenetic effects
by: Blum, Frederic
Published: (2025)
by: Blum, Frederic
Published: (2025)
How communicatively optimal are exact numeral systems? Once more on lexicon size and morphosyntactic complexity
by: Cathcart, Chundra, et al.
Published: (2026)
by: Cathcart, Chundra, et al.
Published: (2026)
Correspondence Analysis and PMI-Based Word Embeddings: A Comparative Study
by: Qi, Qianqian, et al.
Published: (2024)
by: Qi, Qianqian, et al.
Published: (2024)
One Word Is Not Enough: Simple Prompts Improve Word Embeddings
by: Ranjan, Rajeev
Published: (2025)
by: Ranjan, Rajeev
Published: (2025)
Benchmarks Are Not That Out of Distribution: Word Overlap Predicts Performance
by: Chung, Woojin, et al.
Published: (2026)
by: Chung, Woojin, et al.
Published: (2026)
Handling Korean Out-of-Vocabulary Words with Phoneme Representation Learning
by: Kim, Nayeon, et al.
Published: (2025)
by: Kim, Nayeon, et al.
Published: (2025)
Character-aware Transformers Learn an Irregular Morphological Pattern Yet None Generalize Like Humans
by: Ramarao, Akhilesh Kakolu, et al.
Published: (2026)
by: Ramarao, Akhilesh Kakolu, et al.
Published: (2026)
Out of One, Many: Using Language Models to Simulate Human Samples
by: Argyle, Lisa P., et al.
Published: (2022)
by: Argyle, Lisa P., et al.
Published: (2022)
Parsing Through Boundaries in Chinese Word Segmentation
by: Chen, Yige, et al.
Published: (2025)
by: Chen, Yige, et al.
Published: (2025)
Identifying Narrative Patterns and Outliers in Holocaust Testimonies Using Topic Modeling
by: Ifergan, Maxim, et al.
Published: (2024)
by: Ifergan, Maxim, et al.
Published: (2024)
Be Cognative: Cognates in the Rehabilitation of Cochlear Implant Users with German as a Second Language – A Computer‐Based Experiment
by: Susann Thyson, et al.
Published: (2025)
by: Susann Thyson, et al.
Published: (2025)
What's the Difference? Supporting Users in Identifying the Effects of Prompt and Model Changes Through Token Patterns
by: Hedderich, Michael A., et al.
Published: (2025)
by: Hedderich, Michael A., et al.
Published: (2025)
Fast Computation of Leave-One-Out Cross-Validation for $k$-NN Regression
by: Kanagawa, Motonobu
Published: (2024)
by: Kanagawa, Motonobu
Published: (2024)
Using Context to Improve Word Segmentation
by: Hu, Stephanie, et al.
Published: (2025)
by: Hu, Stephanie, et al.
Published: (2025)
Improving Acoustic Word Embeddings through Correspondence Training of Self-supervised Speech Representations
by: Meghanani, Amit, et al.
Published: (2024)
by: Meghanani, Amit, et al.
Published: (2024)
Cutting Through the Noise: Boosting LLM Performance on Math Word Problems
by: Anantheswaran, Ujjwala, et al.
Published: (2024)
by: Anantheswaran, Ujjwala, et al.
Published: (2024)
Can You Learn Semantics Through Next-Word Prediction? The Case of Entailment
by: Merrill, William, et al.
Published: (2024)
by: Merrill, William, et al.
Published: (2024)
Weighted Leave-One-Out Cross Validation
by: Pronzato, Luc, et al.
Published: (2025)
by: Pronzato, Luc, et al.
Published: (2025)
Author-Specific Linguistic Patterns Unveiled: A Deep Learning Study on Word Class Distributions
by: Krauss, Patrick, et al.
Published: (2025)
by: Krauss, Patrick, et al.
Published: (2025)
A Sea of Words: An In-Depth Analysis of Anchors for Text Data
by: Lopardo, Gianluigi, et al.
Published: (2022)
by: Lopardo, Gianluigi, et al.
Published: (2022)
From Where Words Come: Efficient Regularization of Code Tokenizers Through Source Attribution
by: Chizhov, Pavel, et al.
Published: (2026)
by: Chizhov, Pavel, et al.
Published: (2026)
Jailbreaking Large Language Models Through Alignment Vulnerabilities in Out-of-Distribution Settings
by: Huang, Yue, et al.
Published: (2024)
by: Huang, Yue, et al.
Published: (2024)
Classifying Graphemes in English Words Through the Application of a Fuzzy Inference System
by: Rose, Samuel, et al.
Published: (2024)
by: Rose, Samuel, et al.
Published: (2024)
Correlation Does Not Imply Compensation: Complexity and Irregularity in the Lexicon
by: Doucette, Amanda, et al.
Published: (2024)
by: Doucette, Amanda, et al.
Published: (2024)
Similar Items
-
Trimming Phonetic Alignments Improves the Inference of Sound Correspondence Patterns from Multilingual Wordlists
by: Blum, Frederic, et al.
Published: (2023) -
From Isolates to Families: Using Neural Networks for Automated Language Affiliation
by: Blum, Frederic, et al.
Published: (2025) -
Everybody Likes to Sleep: A Computer-Assisted Comparison of Object Naming Data from 30 Languages
by: Kučerová, Alžběta, et al.
Published: (2025) -
A Computational Model for the Assessment of Mutual Intelligibility Among Closely Related Languages
by: Nieder, Jessica, et al.
Published: (2024) -
Partial Colexifications Improve Concept Embeddings
by: Rubehn, Arne, et al.
Published: (2025)