Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bayram, M. Ali, Fincan, Ali Arda, Gümüş, Ahmet Semih, Karakaş, Sercan, Diri, Banu, Yıldırım, Savaş |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Tokens with Meaning: A Hybrid Tokenization Approach for Turkish
von: Bayram, M. Ali, et al.
Veröffentlicht: (2025)
von: Bayram, M. Ali, et al.
Veröffentlicht: (2025)
Doğal Dil İşlemede Tokenizasyon Standartları ve Ölçümü: Türkçe Üzerinden Büyük Dil Modellerinin Karşılaştırmalı Analizi
von: Bayram, M. Ali, et al.
Veröffentlicht: (2025)
von: Bayram, M. Ali, et al.
Veröffentlicht: (2025)
Setting Standards in Turkish NLP: TR-MMLU for Large Language Model Evaluation
von: Bayram, M. Ali, et al.
Veröffentlicht: (2024)
von: Bayram, M. Ali, et al.
Veröffentlicht: (2024)
Büyük Dil Modelleri için TR-MMLU Benchmarkı: Performans Değerlendirmesi, Zorluklar ve İyileştirme Fırsatları
von: Bayram, M. Ali, et al.
Veröffentlicht: (2025)
von: Bayram, M. Ali, et al.
Veröffentlicht: (2025)
NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles
von: Jia, Xiao
Veröffentlicht: (2026)
von: Jia, Xiao
Veröffentlicht: (2026)
Low-Resource English-Tigrinya MT: Leveraging Multilingual Models, Custom Tokenizers, and Clean Evaluation Benchmarks
von: Teklehaymanot, Hailay Kidu, et al.
Veröffentlicht: (2025)
von: Teklehaymanot, Hailay Kidu, et al.
Veröffentlicht: (2025)
The Syntactic Acceptability Dataset (Preview): A Resource for Machine Learning and Linguistic Analysis of English
von: Juzek, Tom S
Veröffentlicht: (2025)
von: Juzek, Tom S
Veröffentlicht: (2025)
CogniLoad: A Synthetic Natural Language Reasoning Benchmark With Tunable Length, Intrinsic Difficulty, and Distractor Density
von: Kaiser, Daniel, et al.
Veröffentlicht: (2025)
von: Kaiser, Daniel, et al.
Veröffentlicht: (2025)
LLM-supported document separation for printed reviews from zbMATH Open
von: Pluzhnikov, Ivan, et al.
Veröffentlicht: (2026)
von: Pluzhnikov, Ivan, et al.
Veröffentlicht: (2026)
From Prompting to Preference Optimization: A Comparative Study of LLM-based Automated Essay Scoring
von: Nguyen, Minh Hoang, et al.
Veröffentlicht: (2026)
von: Nguyen, Minh Hoang, et al.
Veröffentlicht: (2026)
Experimentation in Content Moderation using RWKV
von: Yildirim, Umut, et al.
Veröffentlicht: (2024)
von: Yildirim, Umut, et al.
Veröffentlicht: (2024)
MedMemoryBench: Benchmarking Agent Memory in Personalized Healthcare
von: Wang, Yihao, et al.
Veröffentlicht: (2026)
von: Wang, Yihao, et al.
Veröffentlicht: (2026)
How Human-Like Are Large Language Models? A Register-Aware Linguistic Evaluation Framework
von: Nieth, Björn, et al.
Veröffentlicht: (2026)
von: Nieth, Björn, et al.
Veröffentlicht: (2026)
Mitigating LLM Hallucinations through Domain-Grounded Tiered Retrieval
von: Haque, Md. Asraful, et al.
Veröffentlicht: (2026)
von: Haque, Md. Asraful, et al.
Veröffentlicht: (2026)
Contrasting Linguistic Patterns in Human and LLM-Generated News Text
von: Muñoz-Ortiz, Alberto, et al.
Veröffentlicht: (2023)
von: Muñoz-Ortiz, Alberto, et al.
Veröffentlicht: (2023)
A Stochastic Analysis of the Linguistic Provenance of English Place Names
von: Dalvean, Michael
Veröffentlicht: (2023)
von: Dalvean, Michael
Veröffentlicht: (2023)
Charting a Decade of Computational Linguistics in Italy: The CLiC-it Corpus
von: Alzetta, Chiara, et al.
Veröffentlicht: (2025)
von: Alzetta, Chiara, et al.
Veröffentlicht: (2025)
Computational Social Linguistics for Telugu Cultural Preservation: Novel Algorithms for Chandassu Metrical Pattern Recognition
von: Pavan, Boddu Sri, et al.
Veröffentlicht: (2025)
von: Pavan, Boddu Sri, et al.
Veröffentlicht: (2025)
Morphological Synthesizer for Ge'ez Language: Addressing Morphological Complexity and Resource Limitations
von: Gebremariam, Gebrearegawi, et al.
Veröffentlicht: (2025)
von: Gebremariam, Gebrearegawi, et al.
Veröffentlicht: (2025)
CLMN: Concept based Language Models via Neural Symbolic Reasoning
von: Yang, Yibo
Veröffentlicht: (2025)
von: Yang, Yibo
Veröffentlicht: (2025)
Argument Quality Annotation and Gender Bias Detection in Financial Communication through Large Language Models
von: Alhamzeh, Alaa, et al.
Veröffentlicht: (2025)
von: Alhamzeh, Alaa, et al.
Veröffentlicht: (2025)
IFMTBench: A Comprehensive Benchmark for Multilingual Translation Instruction Following
von: Sun, Mingrui, et al.
Veröffentlicht: (2026)
von: Sun, Mingrui, et al.
Veröffentlicht: (2026)
Evaluating Pixel Language Models on Non-Standardized Languages
von: Muñoz-Ortiz, Alberto, et al.
Veröffentlicht: (2024)
von: Muñoz-Ortiz, Alberto, et al.
Veröffentlicht: (2024)
Fast Quiet-STaR: Thinking Without Thought Tokens
von: Huang, Wei, et al.
Veröffentlicht: (2025)
von: Huang, Wei, et al.
Veröffentlicht: (2025)
Examining Linguistic Shifts in Academic Writing Before and After the Launch of ChatGPT: A Study on Preprint Papers
von: Bao, Tong, et al.
Veröffentlicht: (2025)
von: Bao, Tong, et al.
Veröffentlicht: (2025)
Harnessing non-adversarial robustness in large language models
von: Zhou, Qinghua, et al.
Veröffentlicht: (2026)
von: Zhou, Qinghua, et al.
Veröffentlicht: (2026)
Context Aware Lemmatization and Morphological Tagging Method in Turkish
von: Sayallar, Cagri
Veröffentlicht: (2025)
von: Sayallar, Cagri
Veröffentlicht: (2025)
How much do LLMs learn from negative examples?
von: Hamdan, Shadi, et al.
Veröffentlicht: (2025)
von: Hamdan, Shadi, et al.
Veröffentlicht: (2025)
Communicative Agents for Slideshow Storytelling Video Generation based on LLMs
von: Fan, Jingxing, et al.
Veröffentlicht: (2025)
von: Fan, Jingxing, et al.
Veröffentlicht: (2025)
An Epidemiological Knowledge Graph extracted from the World Health Organization's Disease Outbreak News
von: Consoli, Sergio, et al.
Veröffentlicht: (2025)
von: Consoli, Sergio, et al.
Veröffentlicht: (2025)
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
von: Basu, Abhinaba
Veröffentlicht: (2026)
von: Basu, Abhinaba
Veröffentlicht: (2026)
D-SMART: Enhancing LLM Dialogue Consistency via Dynamic Structured Memory And Reasoning Tree
von: Lei, Xiang, et al.
Veröffentlicht: (2025)
von: Lei, Xiang, et al.
Veröffentlicht: (2025)
From Noise to Diversity: Random Embedding Injection in LLM Reasoning
von: Kim, Heejun, et al.
Veröffentlicht: (2026)
von: Kim, Heejun, et al.
Veröffentlicht: (2026)
NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution
von: Breneur, Oleksandr Marchenko, et al.
Veröffentlicht: (2026)
von: Breneur, Oleksandr Marchenko, et al.
Veröffentlicht: (2026)
Rethinking the Multilingual Reasoning Gap with Layer Swap
von: Lasbordes, Maxence, et al.
Veröffentlicht: (2026)
von: Lasbordes, Maxence, et al.
Veröffentlicht: (2026)
TIS-DPO: Token-level Importance Sampling for Direct Preference Optimization With Estimated Weights
von: Liu, Aiwei, et al.
Veröffentlicht: (2024)
von: Liu, Aiwei, et al.
Veröffentlicht: (2024)
WeDLM: Reconciling Diffusion Language Models with Standard Causal Attention for Fast Inference
von: Liu, Aiwei, et al.
Veröffentlicht: (2025)
von: Liu, Aiwei, et al.
Veröffentlicht: (2025)
Constitution or Collapse? Exploring Constitutional AI with Llama 3-8B
von: Zhang, Xue
Veröffentlicht: (2025)
von: Zhang, Xue
Veröffentlicht: (2025)
XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
von: Estevanell-Valladares, Ernesto L., et al.
Veröffentlicht: (2025)
von: Estevanell-Valladares, Ernesto L., et al.
Veröffentlicht: (2025)
MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
von: Yim, Wen-wai, et al.
Veröffentlicht: (2025)
von: Yim, Wen-wai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Tokens with Meaning: A Hybrid Tokenization Approach for Turkish
von: Bayram, M. Ali, et al.
Veröffentlicht: (2025) -
Doğal Dil İşlemede Tokenizasyon Standartları ve Ölçümü: Türkçe Üzerinden Büyük Dil Modellerinin Karşılaştırmalı Analizi
von: Bayram, M. Ali, et al.
Veröffentlicht: (2025) -
Setting Standards in Turkish NLP: TR-MMLU for Large Language Model Evaluation
von: Bayram, M. Ali, et al.
Veröffentlicht: (2024) -
Büyük Dil Modelleri için TR-MMLU Benchmarkı: Performans Değerlendirmesi, Zorluklar ve İyileştirme Fırsatları
von: Bayram, M. Ali, et al.
Veröffentlicht: (2025) -
NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles
von: Jia, Xiao
Veröffentlicht: (2026)