Enhancing Sindhi Word Segmentation using Subword Representation Learning and Position-aware Self-attention
Fuente:
arXiv
Saved in:
| Main Authors: | Ali, Wazir, Kumar, Jay, Tumrani, Saifullah, Nour, Redhwan, Noor, Adeeb, Xu, Zenglin |
|---|---|
| Format: | Preprint |
| Published: |
2020
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SiNFluD: Creating and Evaluating Figurative Language Dataset for Sindhi
by: Ali, Wazir, et al.
Published: (2026)
by: Ali, Wazir, et al.
Published: (2026)
An Evaluation of Sindhi Word Embedding in Semantic Analogies and Downstream Tasks
by: Ali, Wazir, et al.
Published: (2024)
by: Ali, Wazir, et al.
Published: (2024)
BioUNER: A Benchmark Dataset for Clinical Urdu Named Entity Recognition
by: Ali, Wazir, et al.
Published: (2026)
by: Ali, Wazir, et al.
Published: (2026)
Subword Tokenization Strategies for Kurdish Word Embeddings
by: Salehi, Ali, et al.
Published: (2025)
by: Salehi, Ali, et al.
Published: (2025)
Lexically Grounded Subword Segmentation
by: Libovický, Jindřich, et al.
Published: (2024)
by: Libovický, Jindřich, et al.
Published: (2024)
The Learning Dynamics of Subword Segmentation for Morphologically Diverse Languages
by: Meyer, Francois, et al.
Published: (2025)
by: Meyer, Francois, et al.
Published: (2025)
Learning Mutually Informed Representations for Characters and Subwords
by: Wang, Yilin, et al.
Published: (2023)
by: Wang, Yilin, et al.
Published: (2023)
A Survey of Large Language Models for European Languages
by: Ali, Wazir, et al.
Published: (2024)
by: Ali, Wazir, et al.
Published: (2024)
Integrating a Heterogeneous Graph with Entity-aware Self-attention using Relative Position Labels for Reading Comprehension Model
by: Foolad, Shima, et al.
Published: (2023)
by: Foolad, Shima, et al.
Published: (2023)
Leading Whitespaces of Language Models' Subword Vocabulary Pose a Confound for Calculating Word Probabilities
by: Oh, Byung-Doh, et al.
Published: (2024)
by: Oh, Byung-Doh, et al.
Published: (2024)
Evaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge
by: Batsuren, Khuyagbaatar, et al.
Published: (2024)
by: Batsuren, Khuyagbaatar, et al.
Published: (2024)
Language Maintenance or Shift: A Case Study of Sketches Soulful Sindhi Songs
by: Muhammad Hassan Abbasi, et al.
Published: (2025)
by: Muhammad Hassan Abbasi, et al.
Published: (2025)
Distributional Properties of Subword Regularization
by: Cognetta, Marco, et al.
Published: (2024)
by: Cognetta, Marco, et al.
Published: (2024)
Language Shift and Ethnic Identity: Focus on Malaysian Sindhis
by: David Maya Khemlani
Published: (2020)
by: David Maya Khemlani
Published: (2020)
Context-aware Rotary Position Embedding
by: Veisi, Ali, et al.
Published: (2025)
by: Veisi, Ali, et al.
Published: (2025)
Token Alignment via Character Matching for Subword Completion
by: Athiwaratkun, Ben, et al.
Published: (2024)
by: Athiwaratkun, Ben, et al.
Published: (2024)
Linguistic Trends among Young Sindhi Community Members in Karachi
by: Muhammad Hassan Abbasi
Published: (2020)
by: Muhammad Hassan Abbasi
Published: (2020)
ByteSpan: Information-Driven Subword Tokenisation
by: Goriely, Zébulon, et al.
Published: (2025)
by: Goriely, Zébulon, et al.
Published: (2025)
Subword models struggle with word learning, but surprisal hides it
by: Bunzeck, Bastian, et al.
Published: (2025)
by: Bunzeck, Bastian, et al.
Published: (2025)
Morphological Typology in BPE Subword Productivity and Language Modeling
by: Parra, Iñigo
Published: (2024)
by: Parra, Iñigo
Published: (2024)
StylusAI: Stylistic Adaptation for Robust German Handwritten Text Generation
by: Riaz, Nauman, et al.
Published: (2024)
by: Riaz, Nauman, et al.
Published: (2024)
T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient Embeddings
by: Deiseroth, Björn, et al.
Published: (2024)
by: Deiseroth, Björn, et al.
Published: (2024)
Existential Definability over the Subword Ordering
by: Baumann, Pascal, et al.
Published: (2022)
by: Baumann, Pascal, et al.
Published: (2022)
Stolen Subwords: Importance of Vocabularies for Machine Translation Model Stealing
by: Zouhar, Vilém
Published: (2024)
by: Zouhar, Vilém
Published: (2024)
Tokenization Falling Short: On Subword Robustness in Large Language Models
by: Chai, Yekun, et al.
Published: (2024)
by: Chai, Yekun, et al.
Published: (2024)
StochasTok: Improving Fine-Grained Subword Understanding in LLMs
by: Sims, Anya, et al.
Published: (2025)
by: Sims, Anya, et al.
Published: (2025)
Understanding Subword Compositionality of Large Language Models
by: Peng, Qiwei, et al.
Published: (2025)
by: Peng, Qiwei, et al.
Published: (2025)
Word-Representable Graphs and Locality of Words
by: Böll, Philipp, et al.
Published: (2025)
by: Böll, Philipp, et al.
Published: (2025)
Two-step Automated Cybercrime Coded Word Detection using Multi-level Representation Learning
by: Kim, Yongyeon, et al.
Published: (2024)
by: Kim, Yongyeon, et al.
Published: (2024)
On the Effect of (Near) Duplicate Subwords in Language Modelling
by: Schäfer, Anton, et al.
Published: (2024)
by: Schäfer, Anton, et al.
Published: (2024)
Contrastive Learning with Enhanced Abstract Representations using Grouped Loss of Abstract Semantic Supervision
by: Suissa, Omri, et al.
Published: (2025)
by: Suissa, Omri, et al.
Published: (2025)
A Systematic Analysis of Subwords and Cross-Lingual Transfer in Multilingual Translation
by: Meyer, Francois, et al.
Published: (2024)
by: Meyer, Francois, et al.
Published: (2024)
SubRegWeigh: Effective and Efficient Annotation Weighing with Subword Regularization
by: Tsuji, Kohei, et al.
Published: (2024)
by: Tsuji, Kohei, et al.
Published: (2024)
A Subword Embedding Approach for Variation Detection in Luxembourgish User Comments
by: Lutgen, Anne-Marie, et al.
Published: (2026)
by: Lutgen, Anne-Marie, et al.
Published: (2026)
Assessing the Importance of Frequency versus Compositionality for Subword-based Tokenization in NMT
by: Wolleb, Benoist, et al.
Published: (2023)
by: Wolleb, Benoist, et al.
Published: (2023)
The Impact of Word Splitting on the Semantic Content of Contextualized Word Representations
by: Soler, Aina Garí, et al.
Published: (2024)
by: Soler, Aina Garí, et al.
Published: (2024)
TokDrift: When LLM Speaks in Subwords but Code Speaks in Grammar
by: Li, Yinxi, et al.
Published: (2025)
by: Li, Yinxi, et al.
Published: (2025)
Handling Korean Out-of-Vocabulary Words with Phoneme Representation Learning
by: Kim, Nayeon, et al.
Published: (2025)
by: Kim, Nayeon, et al.
Published: (2025)
Self-attention vector output similarities reveal how machines pay attention
by: Halevi, Tal, et al.
Published: (2025)
by: Halevi, Tal, et al.
Published: (2025)
Subword enumeration up to stack-sorting equivalence
by: Campbell, John M., et al.
Published: (2026)
by: Campbell, John M., et al.
Published: (2026)
Similar Items
-
SiNFluD: Creating and Evaluating Figurative Language Dataset for Sindhi
by: Ali, Wazir, et al.
Published: (2026) -
An Evaluation of Sindhi Word Embedding in Semantic Analogies and Downstream Tasks
by: Ali, Wazir, et al.
Published: (2024) -
BioUNER: A Benchmark Dataset for Clinical Urdu Named Entity Recognition
by: Ali, Wazir, et al.
Published: (2026) -
Subword Tokenization Strategies for Kurdish Word Embeddings
by: Salehi, Ali, et al.
Published: (2025) -
Lexically Grounded Subword Segmentation
by: Libovický, Jindřich, et al.
Published: (2024)