What Language is This? Ask Your Tokenizer
Fuente:
arXiv
Saved in:
| Main Authors: | Meister, Clara, Yavuz, Ahmetcan, Lesci, Pietro, Pimentel, Tiago |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Causal Estimation of Tokenisation Bias
by: Lesci, Pietro, et al.
Published: (2025)
by: Lesci, Pietro, et al.
Published: (2025)
How to Compute the Probability of a Word
by: Pimentel, Tiago, et al.
Published: (2024)
by: Pimentel, Tiago, et al.
Published: (2024)
Towards a Similarity-adjusted Surprisal Theory
by: Meister, Clara, et al.
Published: (2024)
by: Meister, Clara, et al.
Published: (2024)
Locally Typical Sampling
by: Meister, Clara, et al.
Published: (2022)
by: Meister, Clara, et al.
Published: (2022)
Testing the Predictions of Surprisal Theory in 11 Languages
by: Wilcox, Ethan Gotlieb, et al.
Published: (2023)
by: Wilcox, Ethan Gotlieb, et al.
Published: (2023)
On the Efficacy of Sampling Adapters
by: Meister, Clara, et al.
Published: (2023)
by: Meister, Clara, et al.
Published: (2023)
Analyzing Wrap-Up Effects through an Information-Theoretic Lens
by: Meister, Clara, et al.
Published: (2022)
by: Meister, Clara, et al.
Published: (2022)
Causal Estimation of Memorisation Profiles
by: Lesci, Pietro, et al.
Published: (2024)
by: Lesci, Pietro, et al.
Published: (2024)
Back to Bytes: Revisiting Tokenization Through UTF-8
by: Moryossef, Amit, et al.
Published: (2025)
by: Moryossef, Amit, et al.
Published: (2025)
Tending Towards Stability: Convergence Challenges in Small Language Models
by: Martinez, Richard Diehl, et al.
Published: (2024)
by: Martinez, Richard Diehl, et al.
Published: (2024)
AnchorAL: Computationally Efficient Active Learning for Large and Imbalanced Datasets
by: Lesci, Pietro, et al.
Published: (2024)
by: Lesci, Pietro, et al.
Published: (2024)
When and What to Ask: AskBench and Rubric-Guided RLVR for LLM Clarification
by: Zhao, Jiale, et al.
Published: (2026)
by: Zhao, Jiale, et al.
Published: (2026)
Formal Aspects of Language Modeling
by: Cotterell, Ryan, et al.
Published: (2023)
by: Cotterell, Ryan, et al.
Published: (2023)
Investigating Critical Period Effects in Language Acquisition through Neural Language Models
by: Constantinescu, Ionut, et al.
Published: (2024)
by: Constantinescu, Ionut, et al.
Published: (2024)
Broken Tokens? Your Language Model can Secretly Handle Non-Canonical Tokenizations
by: Zheng, Brian Siyuan, et al.
Published: (2025)
by: Zheng, Brian Siyuan, et al.
Published: (2025)
Convergence and Divergence of Language Models under Different Random Seeds
by: Fehlauer, Finlay, et al.
Published: (2025)
by: Fehlauer, Finlay, et al.
Published: (2025)
Mind What You Ask For: Emotional and Rational Faces of Persuasion by Large Language Models
by: Mieleszczenko-Kowszewicz, Wiktoria, et al.
Published: (2025)
by: Mieleszczenko-Kowszewicz, Wiktoria, et al.
Published: (2025)
Self-Training Large Language Models for Tool-Use Without Demonstrations
by: Luo, Ne, et al.
Published: (2025)
by: Luo, Ne, et al.
Published: (2025)
What You See is What You Ask: Evaluating Audio Descriptions
by: Kala, Divy, et al.
Published: (2025)
by: Kala, Divy, et al.
Published: (2025)
What Do Prosody and Text Convey? Characterizing How Meaningful Information is Distributed Across Multiple Channels
by: Yadavalli, Aditya, et al.
Published: (2025)
by: Yadavalli, Aditya, et al.
Published: (2025)
ByteSpan: Information-Driven Subword Tokenisation
by: Goriely, Zébulon, et al.
Published: (2025)
by: Goriely, Zébulon, et al.
Published: (2025)
Ask Early, Ask Late, Ask Right: When Does Clarification Timing Matter for Long-Horizon Agents?
by: Gulati, Anmol, et al.
Published: (2026)
by: Gulati, Anmol, et al.
Published: (2026)
Ask LLMs Directly, "What shapes your bias?": Measuring Social Bias in Large Language Models
by: Shin, Jisu, et al.
Published: (2024)
by: Shin, Jisu, et al.
Published: (2024)
What Patients Really Ask: Exploring the Effect of False Assumptions in Patient Information Seeking
by: Xiong, Raymond, et al.
Published: (2026)
by: Xiong, Raymond, et al.
Published: (2026)
Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization
by: Foroutan, Negar, et al.
Published: (2025)
by: Foroutan, Negar, et al.
Published: (2025)
Ask Good Questions for Large Language Models
by: Wu, Qi, et al.
Published: (2025)
by: Wu, Qi, et al.
Published: (2025)
RULER: What's the Real Context Size of Your Long-Context Language Models?
by: Hsieh, Cheng-Ping, et al.
Published: (2024)
by: Hsieh, Cheng-Ping, et al.
Published: (2024)
Turkish Native Language Identification V2
by: Uluslu, Ahmet Yavuz, et al.
Published: (2023)
by: Uluslu, Ahmet Yavuz, et al.
Published: (2023)
Tokenisation is NP-Complete
by: Whittington, Philip, et al.
Published: (2024)
by: Whittington, Philip, et al.
Published: (2024)
Unlearning What Matters: Token-Level Attribution for Precise Language Model Unlearning
by: Wu, Jiawei, et al.
Published: (2026)
by: Wu, Jiawei, et al.
Published: (2026)
What's In Your Field? Mapping Scientific Research with Knowledge Graphs and Large Language Models
by: Das, Abhipsha, et al.
Published: (2025)
by: Das, Abhipsha, et al.
Published: (2025)
On the Effect of (Near) Duplicate Subwords in Language Modelling
by: Schäfer, Anton, et al.
Published: (2024)
by: Schäfer, Anton, et al.
Published: (2024)
The Language You Ask In: Language-Conditioned Ideological Divergence in LLM Analysis of Contested Political Documents
by: Smirnov, Oleg
Published: (2026)
by: Smirnov, Oleg
Published: (2026)
Local and Global Decoding in Text Generation
by: Gareev, Daniel, et al.
Published: (2024)
by: Gareev, Daniel, et al.
Published: (2024)
Knowing When to Ask -- Bridging Large Language Models and Data
by: Radhakrishnan, Prashanth, et al.
Published: (2024)
by: Radhakrishnan, Prashanth, et al.
Published: (2024)
IntelliAsk: Learning to Ask High-Quality Research Questions via RLVR
by: Sharma, Karun, et al.
Published: (2026)
by: Sharma, Karun, et al.
Published: (2026)
Benchmarking Distributional Alignment of Large Language Models
by: Meister, Nicole, et al.
Published: (2024)
by: Meister, Nicole, et al.
Published: (2024)
Do Pre-Trained Language Models Detect and Understand Semantic Underspecification? Ask the DUST!
by: Wildenburg, Frank, et al.
Published: (2024)
by: Wildenburg, Frank, et al.
Published: (2024)
What is Your Favorite Gender, MLM? Gender Bias Evaluation in Multilingual Masked Language Models
by: Yu, Jeongrok, et al.
Published: (2024)
by: Yu, Jeongrok, et al.
Published: (2024)
Teaching Vision-Language Models to Ask: Resolving Ambiguity in Visual Questions
by: Jian, Pu, et al.
Published: (2025)
by: Jian, Pu, et al.
Published: (2025)
Similar Items
-
Causal Estimation of Tokenisation Bias
by: Lesci, Pietro, et al.
Published: (2025) -
How to Compute the Probability of a Word
by: Pimentel, Tiago, et al.
Published: (2024) -
Towards a Similarity-adjusted Surprisal Theory
by: Meister, Clara, et al.
Published: (2024) -
Locally Typical Sampling
by: Meister, Clara, et al.
Published: (2022) -
Testing the Predictions of Surprisal Theory in 11 Languages
by: Wilcox, Ethan Gotlieb, et al.
Published: (2023)