Subword Tokenization Strategies for Kurdish Word Embeddings
Fuente:
arXiv
Guardado en:
| Autores principales: | Salehi, Ali, Jacobs, Cassandra L. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Evaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge
por: Batsuren, Khuyagbaatar, et al.
Publicado: (2024)
por: Batsuren, Khuyagbaatar, et al.
Publicado: (2024)
Tokenization Falling Short: On Subword Robustness in Large Language Models
por: Chai, Yekun, et al.
Publicado: (2024)
por: Chai, Yekun, et al.
Publicado: (2024)
Token Alignment via Character Matching for Subword Completion
por: Athiwaratkun, Ben, et al.
Publicado: (2024)
por: Athiwaratkun, Ben, et al.
Publicado: (2024)
Statistical Uncertainty in Word Embeddings: GloVe-V
por: Vallebueno, Andrea, et al.
Publicado: (2024)
por: Vallebueno, Andrea, et al.
Publicado: (2024)
Enhancing Sindhi Word Segmentation using Subword Representation Learning and Position-aware Self-attention
por: Ali, Wazir, et al.
Publicado: (2020)
por: Ali, Wazir, et al.
Publicado: (2020)
T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient Embeddings
por: Deiseroth, Björn, et al.
Publicado: (2024)
por: Deiseroth, Björn, et al.
Publicado: (2024)
Assessing the Importance of Frequency versus Compositionality for Subword-based Tokenization in NMT
por: Wolleb, Benoist, et al.
Publicado: (2023)
por: Wolleb, Benoist, et al.
Publicado: (2023)
On the scaling relationship between cloze probabilities and language model next-token prediction
por: Jacobs, Cassandra L., et al.
Publicado: (2026)
por: Jacobs, Cassandra L., et al.
Publicado: (2026)
Evaluating Subword Tokenization Techniques for Bengali: A Benchmark Study with BengaliBPE
por: Patwary, Firoj Ahmmed, et al.
Publicado: (2025)
por: Patwary, Firoj Ahmmed, et al.
Publicado: (2025)
A Subword Embedding Approach for Variation Detection in Luxembourgish User Comments
por: Lutgen, Anne-Marie, et al.
Publicado: (2026)
por: Lutgen, Anne-Marie, et al.
Publicado: (2026)
Leading Whitespaces of Language Models' Subword Vocabulary Pose a Confound for Calculating Word Probabilities
por: Oh, Byung-Doh, et al.
Publicado: (2024)
por: Oh, Byung-Doh, et al.
Publicado: (2024)
Team Ryu's Submission to SIGMORPHON 2024 Shared Task on Subword Tokenization
por: Li, Zilong
Publicado: (2024)
por: Li, Zilong
Publicado: (2024)
The distribution of discourse relations within and across turns in spontaneous conversation
por: Cortez, S. Magalí López, et al.
Publicado: (2023)
por: Cortez, S. Magalí López, et al.
Publicado: (2023)
Evaluating Morphological Plausibility of Subword Tokenization via Statistical Alignment with Morpho-Syntactic Features
por: Stephen, Abishek, et al.
Publicado: (2026)
por: Stephen, Abishek, et al.
Publicado: (2026)
Decoupling the Benefits of Subword Tokenization for Language Model Training via Byte-level Simulation
por: Gigant, Théo, et al.
Publicado: (2026)
por: Gigant, Théo, et al.
Publicado: (2026)
A Bayesian account of pronoun and neopronoun acquisition
por: Jacobs, Cassandra L., et al.
Publicado: (2025)
por: Jacobs, Cassandra L., et al.
Publicado: (2025)
Distributional Properties of Subword Regularization
por: Cognetta, Marco, et al.
Publicado: (2024)
por: Cognetta, Marco, et al.
Publicado: (2024)
Lexically Grounded Subword Segmentation
por: Libovický, Jindřich, et al.
Publicado: (2024)
por: Libovický, Jindřich, et al.
Publicado: (2024)
Linguistic Laws Meet Protein Sequences: A Comparative Analysis of Subword Tokenization Methods
por: Suyunu, Burak, et al.
Publicado: (2024)
por: Suyunu, Burak, et al.
Publicado: (2024)
LGSE: Lexically Grounded Subword Embedding Initialization for Low-Resource Language Adaptation
por: Teklehaymanot, Hailay, et al.
Publicado: (2026)
por: Teklehaymanot, Hailay, et al.
Publicado: (2026)
OFA: A Framework of Initializing Unseen Subword Embeddings for Efficient Large-scale Multilingual Continued Pretraining
por: Liu, Yihong, et al.
Publicado: (2023)
por: Liu, Yihong, et al.
Publicado: (2023)
Idiom Detection in Sorani Kurdish Texts
por: Omer, Skala Kamaran, et al.
Publicado: (2025)
por: Omer, Skala Kamaran, et al.
Publicado: (2025)
ByteSpan: Information-Driven Subword Tokenisation
por: Goriely, Zébulon, et al.
Publicado: (2025)
por: Goriely, Zébulon, et al.
Publicado: (2025)
Language and Speech Technology for Central Kurdish Varieties
por: Ahmadi, Sina, et al.
Publicado: (2024)
por: Ahmadi, Sina, et al.
Publicado: (2024)
Optimal Turkish Subword Strategies at Scale: Systematic Evaluation of Data, Vocabulary, Morphology Interplay
por: Altinok, Duygu
Publicado: (2026)
por: Altinok, Duygu
Publicado: (2026)
An Evaluation of Sindhi Word Embedding in Semantic Analogies and Downstream Tasks
por: Ali, Wazir, et al.
Publicado: (2024)
por: Ali, Wazir, et al.
Publicado: (2024)
Subword models struggle with word learning, but surprisal hides it
por: Bunzeck, Bastian, et al.
Publicado: (2025)
por: Bunzeck, Bastian, et al.
Publicado: (2025)
The Learning Dynamics of Subword Segmentation for Morphologically Diverse Languages
por: Meyer, Francois, et al.
Publicado: (2025)
por: Meyer, Francois, et al.
Publicado: (2025)
Morphological Typology in BPE Subword Productivity and Language Modeling
por: Parra, Iñigo
Publicado: (2024)
por: Parra, Iñigo
Publicado: (2024)
FLEURS-Kobani: Extending the FLEURS Dataset for Northern Kurdish
por: Jaff, Daban Q., et al.
Publicado: (2026)
por: Jaff, Daban Q., et al.
Publicado: (2026)
One Word Is Not Enough: Simple Prompts Improve Word Embeddings
por: Ranjan, Rajeev
Publicado: (2025)
por: Ranjan, Rajeev
Publicado: (2025)
Large-scale cloze evaluation reveals that token prediction tasks are neither lexically nor semantically aligned
por: Jacobs, Cassandra L., et al.
Publicado: (2024)
por: Jacobs, Cassandra L., et al.
Publicado: (2024)
Tokenization Strategies for Low-Resource Agglutinative Languages in Word2Vec: Case Study on Turkish and Finnish
por: Hu, Jinfan Frank
Publicado: (2025)
por: Hu, Jinfan Frank
Publicado: (2025)
Learning Mutually Informed Representations for Characters and Subwords
por: Wang, Yilin, et al.
Publicado: (2023)
por: Wang, Yilin, et al.
Publicado: (2023)
StochasTok: Improving Fine-Grained Subword Understanding in LLMs
por: Sims, Anya, et al.
Publicado: (2025)
por: Sims, Anya, et al.
Publicado: (2025)
Stolen Subwords: Importance of Vocabularies for Machine Translation Model Stealing
por: Zouhar, Vilém
Publicado: (2024)
por: Zouhar, Vilém
Publicado: (2024)
Automatic Text Summarization (ATS) for Research Documents in Sorani Kurdish
por: Abdulrahman, Rondik Hadi, et al.
Publicado: (2025)
por: Abdulrahman, Rondik Hadi, et al.
Publicado: (2025)
KurdSTS: The Kurdish Semantic Textual Similarity
por: Abdullah, Abdulhady Abas, et al.
Publicado: (2025)
por: Abdullah, Abdulhady Abas, et al.
Publicado: (2025)
Revisiting Word Embeddings in the LLM Era
por: Mahajan, Yash, et al.
Publicado: (2025)
por: Mahajan, Yash, et al.
Publicado: (2025)
Evaluating Metrics for Bias in Word Embeddings
por: Schröder, Sarah, et al.
Publicado: (2021)
por: Schröder, Sarah, et al.
Publicado: (2021)
Ejemplares similares
-
Evaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge
por: Batsuren, Khuyagbaatar, et al.
Publicado: (2024) -
Tokenization Falling Short: On Subword Robustness in Large Language Models
por: Chai, Yekun, et al.
Publicado: (2024) -
Token Alignment via Character Matching for Subword Completion
por: Athiwaratkun, Ben, et al.
Publicado: (2024) -
Statistical Uncertainty in Word Embeddings: GloVe-V
por: Vallebueno, Andrea, et al.
Publicado: (2024) -
Enhancing Sindhi Word Segmentation using Subword Representation Learning and Position-aware Self-attention
por: Ali, Wazir, et al.
Publicado: (2020)