Enhancing Sindhi Word Segmentation using Subword Representation Learning and Position-aware Self-attention
Fuente:
arXiv
Salvato in:
| Autori principali: | Ali, Wazir, Kumar, Jay, Tumrani, Saifullah, Nour, Redhwan, Noor, Adeeb, Xu, Zenglin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2020
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SiNFluD: Creating and Evaluating Figurative Language Dataset for Sindhi
di: Ali, Wazir, et al.
Pubblicazione: (2026)
di: Ali, Wazir, et al.
Pubblicazione: (2026)
An Evaluation of Sindhi Word Embedding in Semantic Analogies and Downstream Tasks
di: Ali, Wazir, et al.
Pubblicazione: (2024)
di: Ali, Wazir, et al.
Pubblicazione: (2024)
BioUNER: A Benchmark Dataset for Clinical Urdu Named Entity Recognition
di: Ali, Wazir, et al.
Pubblicazione: (2026)
di: Ali, Wazir, et al.
Pubblicazione: (2026)
Subword Tokenization Strategies for Kurdish Word Embeddings
di: Salehi, Ali, et al.
Pubblicazione: (2025)
di: Salehi, Ali, et al.
Pubblicazione: (2025)
Lexically Grounded Subword Segmentation
di: Libovický, Jindřich, et al.
Pubblicazione: (2024)
di: Libovický, Jindřich, et al.
Pubblicazione: (2024)
The Learning Dynamics of Subword Segmentation for Morphologically Diverse Languages
di: Meyer, Francois, et al.
Pubblicazione: (2025)
di: Meyer, Francois, et al.
Pubblicazione: (2025)
Learning Mutually Informed Representations for Characters and Subwords
di: Wang, Yilin, et al.
Pubblicazione: (2023)
di: Wang, Yilin, et al.
Pubblicazione: (2023)
A Survey of Large Language Models for European Languages
di: Ali, Wazir, et al.
Pubblicazione: (2024)
di: Ali, Wazir, et al.
Pubblicazione: (2024)
Integrating a Heterogeneous Graph with Entity-aware Self-attention using Relative Position Labels for Reading Comprehension Model
di: Foolad, Shima, et al.
Pubblicazione: (2023)
di: Foolad, Shima, et al.
Pubblicazione: (2023)
Leading Whitespaces of Language Models' Subword Vocabulary Pose a Confound for Calculating Word Probabilities
di: Oh, Byung-Doh, et al.
Pubblicazione: (2024)
di: Oh, Byung-Doh, et al.
Pubblicazione: (2024)
Evaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge
di: Batsuren, Khuyagbaatar, et al.
Pubblicazione: (2024)
di: Batsuren, Khuyagbaatar, et al.
Pubblicazione: (2024)
Language Maintenance or Shift: A Case Study of Sketches Soulful Sindhi Songs
di: Muhammad Hassan Abbasi, et al.
Pubblicazione: (2025)
di: Muhammad Hassan Abbasi, et al.
Pubblicazione: (2025)
Distributional Properties of Subword Regularization
di: Cognetta, Marco, et al.
Pubblicazione: (2024)
di: Cognetta, Marco, et al.
Pubblicazione: (2024)
Language Shift and Ethnic Identity: Focus on Malaysian Sindhis
di: David Maya Khemlani
Pubblicazione: (2020)
di: David Maya Khemlani
Pubblicazione: (2020)
Context-aware Rotary Position Embedding
di: Veisi, Ali, et al.
Pubblicazione: (2025)
di: Veisi, Ali, et al.
Pubblicazione: (2025)
Token Alignment via Character Matching for Subword Completion
di: Athiwaratkun, Ben, et al.
Pubblicazione: (2024)
di: Athiwaratkun, Ben, et al.
Pubblicazione: (2024)
Linguistic Trends among Young Sindhi Community Members in Karachi
di: Muhammad Hassan Abbasi
Pubblicazione: (2020)
di: Muhammad Hassan Abbasi
Pubblicazione: (2020)
ByteSpan: Information-Driven Subword Tokenisation
di: Goriely, Zébulon, et al.
Pubblicazione: (2025)
di: Goriely, Zébulon, et al.
Pubblicazione: (2025)
Subword models struggle with word learning, but surprisal hides it
di: Bunzeck, Bastian, et al.
Pubblicazione: (2025)
di: Bunzeck, Bastian, et al.
Pubblicazione: (2025)
Morphological Typology in BPE Subword Productivity and Language Modeling
di: Parra, Iñigo
Pubblicazione: (2024)
di: Parra, Iñigo
Pubblicazione: (2024)
StylusAI: Stylistic Adaptation for Robust German Handwritten Text Generation
di: Riaz, Nauman, et al.
Pubblicazione: (2024)
di: Riaz, Nauman, et al.
Pubblicazione: (2024)
T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient Embeddings
di: Deiseroth, Björn, et al.
Pubblicazione: (2024)
di: Deiseroth, Björn, et al.
Pubblicazione: (2024)
Existential Definability over the Subword Ordering
di: Baumann, Pascal, et al.
Pubblicazione: (2022)
di: Baumann, Pascal, et al.
Pubblicazione: (2022)
Stolen Subwords: Importance of Vocabularies for Machine Translation Model Stealing
di: Zouhar, Vilém
Pubblicazione: (2024)
di: Zouhar, Vilém
Pubblicazione: (2024)
Tokenization Falling Short: On Subword Robustness in Large Language Models
di: Chai, Yekun, et al.
Pubblicazione: (2024)
di: Chai, Yekun, et al.
Pubblicazione: (2024)
StochasTok: Improving Fine-Grained Subword Understanding in LLMs
di: Sims, Anya, et al.
Pubblicazione: (2025)
di: Sims, Anya, et al.
Pubblicazione: (2025)
Understanding Subword Compositionality of Large Language Models
di: Peng, Qiwei, et al.
Pubblicazione: (2025)
di: Peng, Qiwei, et al.
Pubblicazione: (2025)
Word-Representable Graphs and Locality of Words
di: Böll, Philipp, et al.
Pubblicazione: (2025)
di: Böll, Philipp, et al.
Pubblicazione: (2025)
Two-step Automated Cybercrime Coded Word Detection using Multi-level Representation Learning
di: Kim, Yongyeon, et al.
Pubblicazione: (2024)
di: Kim, Yongyeon, et al.
Pubblicazione: (2024)
On the Effect of (Near) Duplicate Subwords in Language Modelling
di: Schäfer, Anton, et al.
Pubblicazione: (2024)
di: Schäfer, Anton, et al.
Pubblicazione: (2024)
Contrastive Learning with Enhanced Abstract Representations using Grouped Loss of Abstract Semantic Supervision
di: Suissa, Omri, et al.
Pubblicazione: (2025)
di: Suissa, Omri, et al.
Pubblicazione: (2025)
A Systematic Analysis of Subwords and Cross-Lingual Transfer in Multilingual Translation
di: Meyer, Francois, et al.
Pubblicazione: (2024)
di: Meyer, Francois, et al.
Pubblicazione: (2024)
SubRegWeigh: Effective and Efficient Annotation Weighing with Subword Regularization
di: Tsuji, Kohei, et al.
Pubblicazione: (2024)
di: Tsuji, Kohei, et al.
Pubblicazione: (2024)
A Subword Embedding Approach for Variation Detection in Luxembourgish User Comments
di: Lutgen, Anne-Marie, et al.
Pubblicazione: (2026)
di: Lutgen, Anne-Marie, et al.
Pubblicazione: (2026)
Assessing the Importance of Frequency versus Compositionality for Subword-based Tokenization in NMT
di: Wolleb, Benoist, et al.
Pubblicazione: (2023)
di: Wolleb, Benoist, et al.
Pubblicazione: (2023)
The Impact of Word Splitting on the Semantic Content of Contextualized Word Representations
di: Soler, Aina Garí, et al.
Pubblicazione: (2024)
di: Soler, Aina Garí, et al.
Pubblicazione: (2024)
TokDrift: When LLM Speaks in Subwords but Code Speaks in Grammar
di: Li, Yinxi, et al.
Pubblicazione: (2025)
di: Li, Yinxi, et al.
Pubblicazione: (2025)
Handling Korean Out-of-Vocabulary Words with Phoneme Representation Learning
di: Kim, Nayeon, et al.
Pubblicazione: (2025)
di: Kim, Nayeon, et al.
Pubblicazione: (2025)
Self-attention vector output similarities reveal how machines pay attention
di: Halevi, Tal, et al.
Pubblicazione: (2025)
di: Halevi, Tal, et al.
Pubblicazione: (2025)
Subword enumeration up to stack-sorting equivalence
di: Campbell, John M., et al.
Pubblicazione: (2026)
di: Campbell, John M., et al.
Pubblicazione: (2026)
Documenti analoghi
-
SiNFluD: Creating and Evaluating Figurative Language Dataset for Sindhi
di: Ali, Wazir, et al.
Pubblicazione: (2026) -
An Evaluation of Sindhi Word Embedding in Semantic Analogies and Downstream Tasks
di: Ali, Wazir, et al.
Pubblicazione: (2024) -
BioUNER: A Benchmark Dataset for Clinical Urdu Named Entity Recognition
di: Ali, Wazir, et al.
Pubblicazione: (2026) -
Subword Tokenization Strategies for Kurdish Word Embeddings
di: Salehi, Ali, et al.
Pubblicazione: (2025) -
Lexically Grounded Subword Segmentation
di: Libovický, Jindřich, et al.
Pubblicazione: (2024)