Tokenizer Fertility and Zero-Shot Performance of Foundation Models on Ukrainian Legal Text: A Comparative Study
Fuente:
arXiv
Saved in:
| Main Author: | Ovcharov, Volodymyr |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning
by: Ovcharov, Volodymyr
Published: (2026)
by: Ovcharov, Volodymyr
Published: (2026)
Temporal Concept Drift in Legal Judgment Prediction: Neural Baselines Across Three Epochs of Ukrainian Court Decisions
by: Ovcharov, Volodymyr
Published: (2026)
by: Ovcharov, Volodymyr
Published: (2026)
The Tokenizer Tax Across 25 European Languages: Domain Invariance, Cross-Lingual Few-Shot Effects, and the Ukrainian Penalty
by: Ovcharov, Volodymyr
Published: (2026)
by: Ovcharov, Volodymyr
Published: (2026)
Automatic Construction of a Legal Citation Graph from 100 Million Ukrainian Court Decisions: Large-Scale Extraction, Topological Analysis, and Ontology-Driven Clustering
by: Ovcharov, Volodymyr
Published: (2026)
by: Ovcharov, Volodymyr
Published: (2026)
Temporal Decay of Co-Citation Predictability: A 20-Year Statute Retrieval Benchmark from 396M Ukrainian Court Citations
by: Ovcharov, Volodymyr
Published: (2026)
by: Ovcharov, Volodymyr
Published: (2026)
Citation Grounding: Detecting and Reducing LLM Citation Hallucinations via Legal Citation Graphs
by: Ovcharov, Volodymyr
Published: (2026)
by: Ovcharov, Volodymyr
Published: (2026)
Tokenization Matters: Improving Zero-Shot NER for Indic Languages
by: Pattnayak, Priyaranjan, et al.
Published: (2025)
by: Pattnayak, Priyaranjan, et al.
Published: (2025)
Zero-Shot Detection of LLM-Generated Text via Implicit Reward Model
by: Liu, Runheng, et al.
Published: (2026)
by: Liu, Runheng, et al.
Published: (2026)
Cross-lingual Text Classification Transfer: The Case of Ukrainian
by: Dementieva, Daryna, et al.
Published: (2024)
by: Dementieva, Daryna, et al.
Published: (2024)
Detecting AI-Generated Paraphrases in Bengali: A Comparative Study of Zero-Shot and Fine-Tuned Transformers
by: Islam, Md. Rakibul, et al.
Published: (2025)
by: Islam, Md. Rakibul, et al.
Published: (2025)
Can Small Models Reason About Legal Documents? A Comparative Study
by: Vaddi, Snehit
Published: (2026)
by: Vaddi, Snehit
Published: (2026)
Benchmarking Zero-Shot Robustness of Multimodal Foundation Models: A Pilot Study
by: Wang, Chenguang, et al.
Published: (2024)
by: Wang, Chenguang, et al.
Published: (2024)
Bias in Text Embedding Models
by: Rakivnenko, Vasyl, et al.
Published: (2024)
by: Rakivnenko, Vasyl, et al.
Published: (2024)
Exploiting Text Semantics for Few and Zero Shot Node Classification on Text-attributed Graph
by: Wang, Yuxiang, et al.
Published: (2025)
by: Wang, Yuxiang, et al.
Published: (2025)
Optimizing Medical Question-Answering Systems: A Comparative Study of Fine-Tuned and Zero-Shot Large Language Models with RAG Framework
by: Hassan, Tasnimul, et al.
Published: (2025)
by: Hassan, Tasnimul, et al.
Published: (2025)
DivScore: Zero-Shot Detection of LLM-Generated Text in Specialized Domains
by: Chen, Zhihui, et al.
Published: (2025)
by: Chen, Zhihui, et al.
Published: (2025)
Flex-Judge: Text-Only Reasoning Unleashes Zero-Shot Multimodal Evaluators
by: Ko, Jongwoo, et al.
Published: (2025)
by: Ko, Jongwoo, et al.
Published: (2025)
Enhancing Text-based Knowledge Graph Completion with Zero-Shot Large Language Models: A Focus on Semantic Enhancement
by: Yang, Rui, et al.
Published: (2023)
by: Yang, Rui, et al.
Published: (2023)
Regional Tiny Stories: Using Small Models to Compare Language Learning and Tokenizer Performance
by: Patil, Nirvan, et al.
Published: (2025)
by: Patil, Nirvan, et al.
Published: (2025)
From Legal Text to Executable Decision Models: Evaluating Structured Representations for Legal Decision Model Generation
by: Graus, David
Published: (2026)
by: Graus, David
Published: (2026)
Glimpse: Enabling White-Box Methods to Use Proprietary Models for Zero-Shot LLM-Generated Text Detection
by: Bao, Guangsheng, et al.
Published: (2024)
by: Bao, Guangsheng, et al.
Published: (2024)
Unpacking Tokenization: Evaluating Text Compression and its Correlation with Model Performance
by: Goldman, Omer, et al.
Published: (2024)
by: Goldman, Omer, et al.
Published: (2024)
Bridging National and International Legal Data: Two Projects Based on the Japanese Legal Standard XML Schema for Comparative Law Studies
by: Nakamura, Makoto
Published: (2026)
by: Nakamura, Makoto
Published: (2026)
Language Complexity Measurement as a Noisy Zero-Shot Proxy for Evaluating LLM Performance
by: Moell, Birger, et al.
Published: (2025)
by: Moell, Birger, et al.
Published: (2025)
Benchmarking Open-Source Large Language Models for Persian in Zero-Shot and Few-Shot Learning
by: Cherakhloo, Mahdi, et al.
Published: (2025)
by: Cherakhloo, Mahdi, et al.
Published: (2025)
MIO: A Foundation Model on Multimodal Tokens
by: Wang, Zekun, et al.
Published: (2024)
by: Wang, Zekun, et al.
Published: (2024)
Training Text-to-Molecule Models with Context-Aware Tokenization
by: Kim, Seojin, et al.
Published: (2025)
by: Kim, Seojin, et al.
Published: (2025)
KL3M Tokenizers: A Family of Domain-Specific and Character-Level Tokenizers for Legal, Financial, and Preprocessing Applications
by: Bommarito, Michael J, et al.
Published: (2025)
by: Bommarito, Michael J, et al.
Published: (2025)
No Text Needed: Forecasting MT Quality and Inequity from Fertility and Metadata
by: Lundin, Jessica M., et al.
Published: (2025)
by: Lundin, Jessica M., et al.
Published: (2025)
Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text
by: Hans, Abhimanyu, et al.
Published: (2024)
by: Hans, Abhimanyu, et al.
Published: (2024)
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models
by: Cho, Gyeongje, et al.
Published: (2025)
by: Cho, Gyeongje, et al.
Published: (2025)
Boosting Zero-Shot Crosslingual Performance using LLM-Based Augmentations with Effective Data Selection
by: Fazili, Barah, et al.
Published: (2024)
by: Fazili, Barah, et al.
Published: (2024)
Zero-Shot End-to-End Relation Extraction in Chinese: A Comparative Study of Gemini, LLaMA and ChatGPT
by: Du, Shaoshuai, et al.
Published: (2025)
by: Du, Shaoshuai, et al.
Published: (2025)
A Comparative Study of Decoding Strategies in Medical Text Generation
by: Presacan, Oriana, et al.
Published: (2025)
by: Presacan, Oriana, et al.
Published: (2025)
A Comparative Study of Quality Evaluation Methods for Text Summarization
by: Nguyen, Huyen, et al.
Published: (2024)
by: Nguyen, Huyen, et al.
Published: (2024)
Sticking to the Mean: Detecting Sticky Tokens in Text Embedding Models
by: Chen, Kexin, et al.
Published: (2025)
by: Chen, Kexin, et al.
Published: (2025)
Evaluating the Performance of AI Text Detectors, Few-Shot and Chain-of-Thought Prompting Using DeepSeek Generated Text
by: Alshammari, Hulayyil, et al.
Published: (2025)
by: Alshammari, Hulayyil, et al.
Published: (2025)
Beyond Fertility: Analyzing STRR as a Metric for Multilingual Tokenization Evaluation
by: Nayeem, Mir Tafseer, et al.
Published: (2025)
by: Nayeem, Mir Tafseer, et al.
Published: (2025)
Zero-Shot Verification-guided Chain of Thoughts
by: Chowdhury, Jishnu Ray, et al.
Published: (2025)
by: Chowdhury, Jishnu Ray, et al.
Published: (2025)
Large Language Models are Zero-Shot Next Location Predictors
by: Beneduce, Ciro, et al.
Published: (2024)
by: Beneduce, Ciro, et al.
Published: (2024)
Similar Items
-
UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning
by: Ovcharov, Volodymyr
Published: (2026) -
Temporal Concept Drift in Legal Judgment Prediction: Neural Baselines Across Three Epochs of Ukrainian Court Decisions
by: Ovcharov, Volodymyr
Published: (2026) -
The Tokenizer Tax Across 25 European Languages: Domain Invariance, Cross-Lingual Few-Shot Effects, and the Ukrainian Penalty
by: Ovcharov, Volodymyr
Published: (2026) -
Automatic Construction of a Legal Citation Graph from 100 Million Ukrainian Court Decisions: Large-Scale Extraction, Topological Analysis, and Ontology-Driven Clustering
by: Ovcharov, Volodymyr
Published: (2026) -
Temporal Decay of Co-Citation Predictability: A 20-Year Statute Retrieval Benchmark from 396M Ukrainian Court Citations
by: Ovcharov, Volodymyr
Published: (2026)