Qtok: A Comprehensive Framework for Evaluating Multilingual Tokenizer Quality in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Chelombitko, Iaroslav, Safronov, Egor, Komissarov, Aleksey |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SampoNLP: A Self-Referential Toolkit for Morphological Analysis of Subword Tokenizers
by: Chelombitko, Iaroslav, et al.
Published: (2026)
by: Chelombitko, Iaroslav, et al.
Published: (2026)
LLM-as-a-Judge: Rapid Evaluation of Legal Document Recommendation for Retrieval-Augmented Generation
by: Pradhan, Anu, et al.
Published: (2025)
by: Pradhan, Anu, et al.
Published: (2025)
Subword-Based Comparative Linguistics across 242 Languages Using Wikipedia Glottosets
by: Chelombitko, Iaroslav, et al.
Published: (2026)
by: Chelombitko, Iaroslav, et al.
Published: (2026)
EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta
by: Bernard, Raymond, et al.
Published: (2024)
by: Bernard, Raymond, et al.
Published: (2024)
Evolve: A Persistent Knowledge Lifecycle for Small Language Models
by: Hovagimian, Dikran
Published: (2026)
by: Hovagimian, Dikran
Published: (2026)
MoVoC: Morphology-Aware Subword Construction for Geez Script Languages
by: Teklehaymanot, Hailay Kidu, et al.
Published: (2025)
by: Teklehaymanot, Hailay Kidu, et al.
Published: (2025)
BudgetMem: Learning Selective Memory Policies for Cost-Efficient Long-Context Processing in Language Models
by: Alla, Chandra Vamsi Krishna, et al.
Published: (2025)
by: Alla, Chandra Vamsi Krishna, et al.
Published: (2025)
Optimizing Retrieval-Augmented Generation (RAG) for Colloquial Cantonese: A LoRA-Based Systematic Review
by: Calonge, David Santandreu, et al.
Published: (2025)
by: Calonge, David Santandreu, et al.
Published: (2025)
What Matters in LLM-Based Feature Extractor for Recommender? A Systematic Analysis of Prompts, Models, and Adaptation
by: Shi, Kainan, et al.
Published: (2025)
by: Shi, Kainan, et al.
Published: (2025)
Improving the Performance of Sequential Recommendation Systems with an Extended Large Language Model
by: Choi, Sinnyum, et al.
Published: (2025)
by: Choi, Sinnyum, et al.
Published: (2025)
RACE-Align: Retrieval-Augmented and Chain-of-Thought Enhanced Preference Alignment for Large Language Models
by: Yan, Qihang, et al.
Published: (2025)
by: Yan, Qihang, et al.
Published: (2025)
PatenTEB: A Comprehensive Benchmark and Model Family for Patent Text Embedding
by: Ayaou, Iliass, et al.
Published: (2025)
by: Ayaou, Iliass, et al.
Published: (2025)
DeformAr: Rethinking NER Evaluation through Component Analysis and Visual Analytics
by: Younes, Ahmed Mustafa
Published: (2025)
by: Younes, Ahmed Mustafa
Published: (2025)
Utilizing Large Language Models to Synthesize Product Desirability Datasets
by: Hastings, John D., et al.
Published: (2024)
by: Hastings, John D., et al.
Published: (2024)
Compressed code: the hidden effects of quantization and distillation on programming tokens
by: Siniaev, Viacheslav, et al.
Published: (2026)
by: Siniaev, Viacheslav, et al.
Published: (2026)
ComplianceNLP: Knowledge-Graph-Augmented RAG for Multi-Framework Regulatory Gap Detection
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
RouteNLP: Closed-Loop LLM Routing with Conformal Cascading and Distillation Co-Optimization
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
Learning to Detect Relevant Contexts and Knowledge for Response Selection in Retrieval-based Dialogue Systems
by: Hua, Kai, et al.
Published: (2025)
by: Hua, Kai, et al.
Published: (2025)
Real-Time RAG for the Identification of Supply Chain Vulnerabilities
by: Ponnock, Jesse, et al.
Published: (2025)
by: Ponnock, Jesse, et al.
Published: (2025)
Mind the Gap: A Generalized Approach for Cross-Modal Embedding Alignment
by: Yadav, Arihan, et al.
Published: (2024)
by: Yadav, Arihan, et al.
Published: (2024)
LLM Reasoning for Cold-Start Item Recommendation
by: Li, Shijun, et al.
Published: (2025)
by: Li, Shijun, et al.
Published: (2025)
EdgeJury: Cross-Reviewed Small-Model Ensembles for Truthful Question Answering on Serverless Edge Inference
by: Kumar, Aayush
Published: (2025)
by: Kumar, Aayush
Published: (2025)
Experimentation Accelerator: Interpretable Insights and Creative Recommendations for A/B Testing with Content-Aware ranking
by: Hu, Zhengmian, et al.
Published: (2026)
by: Hu, Zhengmian, et al.
Published: (2026)
MODP: Multi Objective Directional Prompting
by: Nema, Aashutosh, et al.
Published: (2025)
by: Nema, Aashutosh, et al.
Published: (2025)
Tokens with Meaning: A Hybrid Tokenization Approach for Turkish
by: Bayram, M. Ali, et al.
Published: (2025)
by: Bayram, M. Ali, et al.
Published: (2025)
NSFL: A Post-Training Neuro-Symbolic Fuzzy Logic Framework for Boolean Operators in Neural Embeddings
by: Vexler, Vladi, et al.
Published: (2026)
by: Vexler, Vladi, et al.
Published: (2026)
DEUCE: Dual-diversity Enhancement and Uncertainty-awareness for Cold-start Active Learning
by: Guo, Jiaxin, et al.
Published: (2025)
by: Guo, Jiaxin, et al.
Published: (2025)
HySemRAG: A Hybrid Semantic Retrieval-Augmented Generation Framework for Automated Literature Synthesis and Methodological Gap Analysis
by: Godinez, Alejandro
Published: (2025)
by: Godinez, Alejandro
Published: (2025)
Hard Negative Mining for Domain-Specific Retrieval in Enterprise Systems
by: Meghwani, Hansa, et al.
Published: (2025)
by: Meghwani, Hansa, et al.
Published: (2025)
Combating data scarcity in recommendation services: Integrating cognitive types of VARK and neural network technologies (LLM)
by: Zmanovskii, Nikita
Published: (2026)
by: Zmanovskii, Nikita
Published: (2026)
FinGround: Detecting and Grounding Financial Hallucinations via Atomic Claim Verification
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
LLMs as Architects and Critics for Multi-Source Opinion Summarization
by: Attri, Anuj, et al.
Published: (2025)
by: Attri, Anuj, et al.
Published: (2025)
Why We Feel What We Feel: Joint Detection of Emotions and Their Opinion Triggers in E-commerce
by: Attri, Arnav, et al.
Published: (2025)
by: Attri, Arnav, et al.
Published: (2025)
Enhancing Document AI Data Generation Through Graph-Based Synthetic Layouts
by: Agarwal, Amit, et al.
Published: (2024)
by: Agarwal, Amit, et al.
Published: (2024)
Challenges and Applications of Large Language Models: A Comparison of GPT and DeepSeek family of models
by: Sharma, Shubham, et al.
Published: (2025)
by: Sharma, Shubham, et al.
Published: (2025)
Learning When to Remember: Risk-Sensitive Contextual Bandits for Abstention-Aware Memory Retrieval in LLM-Based Coding Agents
by: Iscan, Mehmet
Published: (2026)
by: Iscan, Mehmet
Published: (2026)
FinQAPT: Empowering Financial Decisions with End-to-End LLM-driven Question Answering Pipeline
by: Singh, Kuldeep, et al.
Published: (2024)
by: Singh, Kuldeep, et al.
Published: (2024)
ArcheType: A Novel Framework for Open-Source Column Type Annotation using Large Language Models
by: Feuer, Benjamin, et al.
Published: (2023)
by: Feuer, Benjamin, et al.
Published: (2023)
Benchmarking Google Embeddings 2 against Open-Source Models for Multilingual Dense Retrieval and RAG Systems
by: Cirillo, Stefano, et al.
Published: (2026)
by: Cirillo, Stefano, et al.
Published: (2026)
Retrieval Augmented Thought Process for Private Data Handling in Healthcare
by: Pouplin, Thomas, et al.
Published: (2024)
by: Pouplin, Thomas, et al.
Published: (2024)
Similar Items
-
SampoNLP: A Self-Referential Toolkit for Morphological Analysis of Subword Tokenizers
by: Chelombitko, Iaroslav, et al.
Published: (2026) -
LLM-as-a-Judge: Rapid Evaluation of Legal Document Recommendation for Retrieval-Augmented Generation
by: Pradhan, Anu, et al.
Published: (2025) -
Subword-Based Comparative Linguistics across 242 Languages Using Wikipedia Glottosets
by: Chelombitko, Iaroslav, et al.
Published: (2026) -
EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta
by: Bernard, Raymond, et al.
Published: (2024) -
Evolve: A Persistent Knowledge Lifecycle for Small Language Models
by: Hovagimian, Dikran
Published: (2026)