Optimal Turkish Subword Strategies at Scale: Systematic Evaluation of Data, Vocabulary, Morphology Interplay
Fuente:
arXiv
Salvato in:
| Autore principale: | Altinok, Duygu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Introducing TrGLUE and SentiTurca: A Comprehensive Benchmark for Turkish General Language Understanding and Sentiment Analysis
di: Altinok, Duygu
Pubblicazione: (2025)
di: Altinok, Duygu
Pubblicazione: (2025)
Whispering Context: Distilling Syntax and Semantics for Long Speech Transcripts
di: Altinok, Duygu
Pubblicazione: (2025)
di: Altinok, Duygu
Pubblicazione: (2025)
Smooth Operators: LLMs Translating Imperfect Hints into Disfluency-Rich Transcripts
di: Altinok, Duygu
Pubblicazione: (2025)
di: Altinok, Duygu
Pubblicazione: (2025)
Evaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge
di: Batsuren, Khuyagbaatar, et al.
Pubblicazione: (2024)
di: Batsuren, Khuyagbaatar, et al.
Pubblicazione: (2024)
D-NLP at SemEval-2024 Task 2: Evaluating Clinical Inference Capabilities of Large Language Models
di: Altinok, Duygu
Pubblicazione: (2024)
di: Altinok, Duygu
Pubblicazione: (2024)
Boosting CTC-Based ASR Using LLM-Based Intermediate Loss Regularization
di: Altinok, Duygu
Pubblicazione: (2025)
di: Altinok, Duygu
Pubblicazione: (2025)
Mind the Gap: Entity-Preserved Context-Aware ASR Structured Transcriptions
di: Altinok, Duygu
Pubblicazione: (2025)
di: Altinok, Duygu
Pubblicazione: (2025)
Token Alignment via Character Matching for Subword Completion
di: Athiwaratkun, Ben, et al.
Pubblicazione: (2024)
di: Athiwaratkun, Ben, et al.
Pubblicazione: (2024)
Evaluating Morphological Compositional Generalization in Large Language Models
di: Ismayilzada, Mete, et al.
Pubblicazione: (2024)
di: Ismayilzada, Mete, et al.
Pubblicazione: (2024)
Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies
di: Tao, Chaofan, et al.
Pubblicazione: (2024)
di: Tao, Chaofan, et al.
Pubblicazione: (2024)
Team Ryu's Submission to SIGMORPHON 2024 Shared Task on Subword Tokenization
di: Li, Zilong
Pubblicazione: (2024)
di: Li, Zilong
Pubblicazione: (2024)
Understanding Subword Compositionality of Large Language Models
di: Peng, Qiwei, et al.
Pubblicazione: (2025)
di: Peng, Qiwei, et al.
Pubblicazione: (2025)
LGSE: Lexically Grounded Subword Embedding Initialization for Low-Resource Language Adaptation
di: Teklehaymanot, Hailay, et al.
Pubblicazione: (2026)
di: Teklehaymanot, Hailay, et al.
Pubblicazione: (2026)
Understanding the Interplay of Scale, Data, and Bias in Language Models: A Case Study with BERT
di: Ali, Muhammad, et al.
Pubblicazione: (2024)
di: Ali, Muhammad, et al.
Pubblicazione: (2024)
Evaluating Modern Large Language Models on Low-Resource and Morphologically Rich Languages:A Cross-Lingual Benchmark Across Cantonese, Japanese, and Turkish
di: Xia, Chengxuan, et al.
Pubblicazione: (2025)
di: Xia, Chengxuan, et al.
Pubblicazione: (2025)
Scaling LLM Pre-training with Vocabulary Curriculum
di: Yu, Fangyuan
Pubblicazione: (2025)
di: Yu, Fangyuan
Pubblicazione: (2025)
A Large-Scale Dataset and Citation Intent Classification in Turkish with LLMs
di: Karaca, Kemal Sami, et al.
Pubblicazione: (2025)
di: Karaca, Kemal Sami, et al.
Pubblicazione: (2025)
TurkBench: A Benchmark for Evaluating Turkish Large Language Models
di: Toraman, Çağrı, et al.
Pubblicazione: (2026)
di: Toraman, Çağrı, et al.
Pubblicazione: (2026)
Organic Data-Driven Approach for Turkish Grammatical Error Correction and LLMs
di: Ersoy, Asım, et al.
Pubblicazione: (2024)
di: Ersoy, Asım, et al.
Pubblicazione: (2024)
Evaluating the Quality of Benchmark Datasets for Low-Resource Languages: A Case Study on Turkish
di: Cengiz, Ayşe Aysu, et al.
Pubblicazione: (2025)
di: Cengiz, Ayşe Aysu, et al.
Pubblicazione: (2025)
Scaling BERT Models for Turkish Automatic Punctuation and Capitalization Correction
di: Saoud, Abdulkader, et al.
Pubblicazione: (2024)
di: Saoud, Abdulkader, et al.
Pubblicazione: (2024)
MoVoC: Morphology-Aware Subword Construction for Geez Script Languages
di: Teklehaymanot, Hailay Kidu, et al.
Pubblicazione: (2025)
di: Teklehaymanot, Hailay Kidu, et al.
Pubblicazione: (2025)
TokDrift: When LLM Speaks in Subwords but Code Speaks in Grammar
di: Li, Yinxi, et al.
Pubblicazione: (2025)
di: Li, Yinxi, et al.
Pubblicazione: (2025)
Systematic Evaluation of the Quality of Synthetic Clinical Notes Rephrased by LLMs at Million-Note Scale
di: Liu, Jinghui, et al.
Pubblicazione: (2026)
di: Liu, Jinghui, et al.
Pubblicazione: (2026)
Overcoming Vocabulary Mismatch: Vocabulary-agnostic Teacher Guided Language Modeling
di: Shin, Haebin, et al.
Pubblicazione: (2025)
di: Shin, Haebin, et al.
Pubblicazione: (2025)
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size
di: Hayou, Soufiane, et al.
Pubblicazione: (2025)
di: Hayou, Soufiane, et al.
Pubblicazione: (2025)
Evaluating LLMs for Zeolite Synthesis Event Extraction (ZSEE): A Systematic Analysis of Prompting Strategies
di: Rathore, Charan Prakash, et al.
Pubblicazione: (2025)
di: Rathore, Charan Prakash, et al.
Pubblicazione: (2025)
T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient Embeddings
di: Deiseroth, Björn, et al.
Pubblicazione: (2024)
di: Deiseroth, Björn, et al.
Pubblicazione: (2024)
OCRTurk: A Comprehensive OCR Benchmark for Turkish
di: Yılmaz, Deniz, et al.
Pubblicazione: (2026)
di: Yılmaz, Deniz, et al.
Pubblicazione: (2026)
A Comprehensive Analysis of Static Word Embeddings for Turkish
di: Sarıtaş, Karahan, et al.
Pubblicazione: (2024)
di: Sarıtaş, Karahan, et al.
Pubblicazione: (2024)
Evaluating Prompting Strategies and Large Language Models in Systematic Literature Review Screening: Relevance and Task-Stage Classification
di: Han, Binglan, et al.
Pubblicazione: (2025)
di: Han, Binglan, et al.
Pubblicazione: (2025)
Context Aware Lemmatization and Morphological Tagging Method in Turkish
di: Sayallar, Cagri
Pubblicazione: (2025)
di: Sayallar, Cagri
Pubblicazione: (2025)
VBART: The Turkish LLM
di: Turker, Meliksah, et al.
Pubblicazione: (2024)
di: Turker, Meliksah, et al.
Pubblicazione: (2024)
Diffutron: A Masked Diffusion Language Model for Turkish Language
di: Kocabay, Şuayp Talha, et al.
Pubblicazione: (2026)
di: Kocabay, Şuayp Talha, et al.
Pubblicazione: (2026)
Introducing cosmosGPT: Monolingual Training for Turkish Language Models
di: Kesgin, H. Toprak, et al.
Pubblicazione: (2024)
di: Kesgin, H. Toprak, et al.
Pubblicazione: (2024)
Interplay of Machine Translation, Diacritics, and Diacritization
di: Chen, Wei-Rui, et al.
Pubblicazione: (2024)
di: Chen, Wei-Rui, et al.
Pubblicazione: (2024)
DVAGen: Dynamic Vocabulary Augmented Generation
di: Du, Wei, et al.
Pubblicazione: (2025)
di: Du, Wei, et al.
Pubblicazione: (2025)
Benchmarking GPT-4 on Algorithmic Problems: A Systematic Evaluation of Prompting Strategies
di: Petruzzellis, Flavio, et al.
Pubblicazione: (2024)
di: Petruzzellis, Flavio, et al.
Pubblicazione: (2024)
Un-considering Contextual Information: Assessing LLMs' Understanding of Indexical Elements
di: Oguz, Metehan, et al.
Pubblicazione: (2025)
di: Oguz, Metehan, et al.
Pubblicazione: (2025)
Fine-tuning Transformer-based Encoder for Turkish Language Understanding Tasks
di: Yildirim, Savas
Pubblicazione: (2024)
di: Yildirim, Savas
Pubblicazione: (2024)
Documenti analoghi
-
Introducing TrGLUE and SentiTurca: A Comprehensive Benchmark for Turkish General Language Understanding and Sentiment Analysis
di: Altinok, Duygu
Pubblicazione: (2025) -
Whispering Context: Distilling Syntax and Semantics for Long Speech Transcripts
di: Altinok, Duygu
Pubblicazione: (2025) -
Smooth Operators: LLMs Translating Imperfect Hints into Disfluency-Rich Transcripts
di: Altinok, Duygu
Pubblicazione: (2025) -
Evaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge
di: Batsuren, Khuyagbaatar, et al.
Pubblicazione: (2024) -
D-NLP at SemEval-2024 Task 2: Evaluating Clinical Inference Capabilities of Large Language Models
di: Altinok, Duygu
Pubblicazione: (2024)