Conditional Unigram Tokenization with Parallel Data
Fuente:
arXiv
Saved in:
| Main Authors: | Vico, Gianluca, Libovický, Jindřinch |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Crowdsourcing Piedmontese to Test LLMs on Non-Standard Orthography
by: Vico, Gianluca, et al.
Published: (2026)
by: Vico, Gianluca, et al.
Published: (2026)
CUS-QA: Local-Knowledge-Oriented Open-Ended Question Answering Dataset
by: Libovický, Jindřich, et al.
Published: (2025)
by: Libovický, Jindřich, et al.
Published: (2025)
Which Pieces Does Unigram Tokenization Really Need?
by: Land, Sander, et al.
Published: (2025)
by: Land, Sander, et al.
Published: (2025)
Investigating the Effect of Parallel Data in the Cross-Lingual Transfer for Vision-Language Encoders
by: Manea, Andrei-Alexandru, et al.
Published: (2025)
by: Manea, Andrei-Alexandru, et al.
Published: (2025)
Evaluating Morphological Plausibility of Subword Tokenization via Statistical Alignment with Morpho-Syntactic Features
by: Stephen, Abishek, et al.
Published: (2026)
by: Stephen, Abishek, et al.
Published: (2026)
Rethinking Tokenization for Rich Morphology: The Dominance of Unigram over BPE and Morphological Alignment
by: Vemula, Saketh Reddy, et al.
Published: (2025)
by: Vemula, Saketh Reddy, et al.
Published: (2025)
Beyond Literal Token Overlap: Token Alignability for Multilinguality
by: Hämmerl, Katharina, et al.
Published: (2025)
by: Hämmerl, Katharina, et al.
Published: (2025)
On the Credibility of Evaluating LLMs using Survey Questions
by: Libovický, Jindřich
Published: (2026)
by: Libovický, Jindřich
Published: (2026)
LLM as a Meta-Judge: Synthetic Data for NLP Evaluation Metric Validation
by: Eigler, Lukáš, et al.
Published: (2026)
by: Eigler, Lukáš, et al.
Published: (2026)
Lexically Grounded Subword Segmentation
by: Libovický, Jindřich, et al.
Published: (2024)
by: Libovický, Jindřich, et al.
Published: (2024)
Multilingual Vision-Language Models, A Survey
by: Manea, Andrei-Alexandru, et al.
Published: (2025)
by: Manea, Andrei-Alexandru, et al.
Published: (2025)
How Gender Interacts with Political Values: A Case Study on Czech BERT Models
by: Ali, Adnan Al, et al.
Published: (2024)
by: Ali, Adnan Al, et al.
Published: (2024)
Understanding Cross-Lingual Alignment -- A Survey
by: Hämmerl, Katharina, et al.
Published: (2024)
by: Hämmerl, Katharina, et al.
Published: (2024)
Different Time, Different Language: Revisiting the Bias Against Non-Native Speakers in GPT Detectors
by: Ali, Adnan Al, et al.
Published: (2026)
by: Ali, Adnan Al, et al.
Published: (2026)
BlockBPE: Parallel BPE Tokenization
by: You, Amos
Published: (2025)
by: You, Amos
Published: (2025)
Parallel Token Prediction for Language Models
by: Draxler, Felix, et al.
Published: (2025)
by: Draxler, Felix, et al.
Published: (2025)
Parallel Tokenizers: Rethinking Vocabulary Design for Cross-Lingual Transfer
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
Efficient Document Parsing via Parallel Token Prediction
by: Li, Lei, et al.
Published: (2026)
by: Li, Lei, et al.
Published: (2026)
How Important is `Perfect' English for Machine Translation Prompts?
by: Schmidtová, Patrícia, et al.
Published: (2025)
by: Schmidtová, Patrícia, et al.
Published: (2025)
Scaling Reasoning Tokens via RL and Parallel Thinking: Evidence From Competitive Programming
by: Zhang, Qianfan, et al.
Published: (2026)
by: Zhang, Qianfan, et al.
Published: (2026)
Teaching LLMs at Charles University: Assignments and Activities
by: Helcl, Jindřich, et al.
Published: (2024)
by: Helcl, Jindřich, et al.
Published: (2024)
LoPT: Lossless Parallel Tokenization Acceleration for Long Context Inference of Large Language Model
by: Shao, Wei, et al.
Published: (2025)
by: Shao, Wei, et al.
Published: (2025)
Enhancing Conceptual Understanding in Multimodal Contrastive Learning through Hard Negative Samples
by: Rösch, Philipp J., et al.
Published: (2024)
by: Rösch, Philipp J., et al.
Published: (2024)
ProPD: Dynamic Token Tree Pruning and Generation for LLM Parallel Decoding
by: Zhong, Shuzhang, et al.
Published: (2024)
by: Zhong, Shuzhang, et al.
Published: (2024)
BehaviorSFT: Behavioral Token Conditioning for Clinical Agents Across the Proactivity Spectrum
by: Kim, Yubin, et al.
Published: (2025)
by: Kim, Yubin, et al.
Published: (2025)
Multilingual Text-to-Image Generation Magnifies Gender Stereotypes and Prompt Engineering May Not Help You
by: Friedrich, Felix, et al.
Published: (2024)
by: Friedrich, Felix, et al.
Published: (2024)
ACADATA: Parallel Dataset of Academic Data for Machine Translation
by: Lacunza, Iñaki, et al.
Published: (2025)
by: Lacunza, Iñaki, et al.
Published: (2025)
Training Large Language Models To Reason In Parallel With Global Forking Tokens
by: Jia, Sheng, et al.
Published: (2025)
by: Jia, Sheng, et al.
Published: (2025)
Are Large Language Models Good Data Preprocessors?
by: Meguellati, Elyas, et al.
Published: (2025)
by: Meguellati, Elyas, et al.
Published: (2025)
Personas with Attitudes: Controlling LLMs for Diverse Data Annotation
by: Fröhling, Leon, et al.
Published: (2024)
by: Fröhling, Leon, et al.
Published: (2024)
Model-Based Quality Assessment for Massively Multilingual Parallel Data
by: Ibrahim, Abdelaziz M. A., et al.
Published: (2026)
by: Ibrahim, Abdelaziz M. A., et al.
Published: (2026)
Speculating LLMs' Chinese Training Data Pollution from Their Tokens
by: Zhang, Qingjie, et al.
Published: (2025)
by: Zhang, Qingjie, et al.
Published: (2025)
Parallel Sampling from Masked Diffusion Models via Conditional Independence Testing
by: Azangulov, Iskander, et al.
Published: (2025)
by: Azangulov, Iskander, et al.
Published: (2025)
Tokenization of Gaze Data
by: Rolff, Tim, et al.
Published: (2025)
by: Rolff, Tim, et al.
Published: (2025)
Charles Translator: A Machine Translation System between Ukrainian and Czech
by: Popel, Martin, et al.
Published: (2024)
by: Popel, Martin, et al.
Published: (2024)
Response-Conditioned Parallel-to-Sequential Orchestration for Multi-Agent Systems
by: Tastan, Nurbek, et al.
Published: (2026)
by: Tastan, Nurbek, et al.
Published: (2026)
dParallel: Learnable Parallel Decoding for dLLMs
by: Chen, Zigeng, et al.
Published: (2025)
by: Chen, Zigeng, et al.
Published: (2025)
The Mediomatix Corpus: Parallel Data for Romansh Language Varieties via Comparable Schoolbooks
by: Hopton, Zachary, et al.
Published: (2025)
by: Hopton, Zachary, et al.
Published: (2025)
Pretraining Strategies using Monolingual and Parallel Data for Low-Resource Machine Translation
by: Nguefack, Idriss Nguepi, et al.
Published: (2025)
by: Nguefack, Idriss Nguepi, et al.
Published: (2025)
SubData: Bridging Heterogeneous Datasets to Enable Theory-Driven Evaluation of Political and Demographic Perspectives in LLMs
by: Bernardelle, Pietro, et al.
Published: (2024)
by: Bernardelle, Pietro, et al.
Published: (2024)
Similar Items
-
Crowdsourcing Piedmontese to Test LLMs on Non-Standard Orthography
by: Vico, Gianluca, et al.
Published: (2026) -
CUS-QA: Local-Knowledge-Oriented Open-Ended Question Answering Dataset
by: Libovický, Jindřich, et al.
Published: (2025) -
Which Pieces Does Unigram Tokenization Really Need?
by: Land, Sander, et al.
Published: (2025) -
Investigating the Effect of Parallel Data in the Cross-Lingual Transfer for Vision-Language Encoders
by: Manea, Andrei-Alexandru, et al.
Published: (2025) -
Evaluating Morphological Plausibility of Subword Tokenization via Statistical Alignment with Morpho-Syntactic Features
by: Stephen, Abishek, et al.
Published: (2026)