Tokenization is Sensitive to Language Variation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wegmann, Anna, Nguyen, Dong, Jurgens, David |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Neurobiber: Fast and Interpretable Stylistic Feature Extraction
von: Alkiek, Kenan, et al.
Veröffentlicht: (2025)
von: Alkiek, Kenan, et al.
Veröffentlicht: (2025)
What's Mine becomes Yours: Defining, Annotating and Detecting Context-Dependent Paraphrases in News Interview Dialogs
von: Wegmann, Anna, et al.
Veröffentlicht: (2024)
von: Wegmann, Anna, et al.
Veröffentlicht: (2024)
Leveraging Multilingual Training for Authorship Representation: Enhancing Generalization across Languages and Domains
von: Kim, Junghwan, et al.
Veröffentlicht: (2025)
von: Kim, Junghwan, et al.
Veröffentlicht: (2025)
Are Rules Meant to be Broken? Understanding Multilingual Moral Reasoning as a Computational Pipeline with UniMoral
von: Kumar, Shivani, et al.
Veröffentlicht: (2025)
von: Kumar, Shivani, et al.
Veröffentlicht: (2025)
Beyond Consensus: Perspectivist Modeling and Evaluation of Annotator Disagreement in NLP
von: Xu, Yinuo, et al.
Veröffentlicht: (2026)
von: Xu, Yinuo, et al.
Veröffentlicht: (2026)
Modeling Empathetic Alignment in Conversation
von: Yang, Jiamin, et al.
Veröffentlicht: (2024)
von: Yang, Jiamin, et al.
Veröffentlicht: (2024)
The Call for Socially Aware Language Technologies
von: Yang, Diyi, et al.
Veröffentlicht: (2024)
von: Yang, Diyi, et al.
Veröffentlicht: (2024)
The Muddy Waters of Modeling Empathy in Language: The Practical Impacts of Theoretical Constructs
von: Lahnala, Allison, et al.
Veröffentlicht: (2025)
von: Lahnala, Allison, et al.
Veröffentlicht: (2025)
A Critical Reflection and Forward Perspective on Empathy and Natural Language Processing
von: Lahnala, Allison, et al.
Veröffentlicht: (2022)
von: Lahnala, Allison, et al.
Veröffentlicht: (2022)
Think Multilingual, Not Harder: A Data-Efficient Framework for Teaching Reasoning Models to Code-Switch
von: Lin, Eleanor M., et al.
Veröffentlicht: (2026)
von: Lin, Eleanor M., et al.
Veröffentlicht: (2026)
NUTMEG: Separating Signal From Noise in Annotator Disagreement
von: Ivey, Jonathan, et al.
Veröffentlicht: (2025)
von: Ivey, Jonathan, et al.
Veröffentlicht: (2025)
The Noisy Path from Source to Citation: Measuring How Scholars Engage with Past Research
von: Chen, Hong, et al.
Veröffentlicht: (2025)
von: Chen, Hong, et al.
Veröffentlicht: (2025)
Are Economists Always More Introverted? Analyzing Consistency in Persona-Assigned LLMs
von: Reusens, Manon, et al.
Veröffentlicht: (2025)
von: Reusens, Manon, et al.
Veröffentlicht: (2025)
Mapping the Podcast Ecosystem with the Structured Podcast Research Corpus
von: Litterer, Benjamin, et al.
Veröffentlicht: (2024)
von: Litterer, Benjamin, et al.
Veröffentlicht: (2024)
The Language of Trauma: Modeling Traumatic Event Descriptions Across Domains with Explainable AI
von: Schirmer, Miriam, et al.
Veröffentlicht: (2024)
von: Schirmer, Miriam, et al.
Veröffentlicht: (2024)
Cooperative Profiles Predict Multi-Agent LLM Team Performance in AI for Science Workflows
von: Kumar, Shivani, et al.
Veröffentlicht: (2026)
von: Kumar, Shivani, et al.
Veröffentlicht: (2026)
Social Perceptions of English Spelling Variation on Twitter: A Comparative Analysis of Human and LLM Responses
von: Nguyen, Dong, et al.
Veröffentlicht: (2025)
von: Nguyen, Dong, et al.
Veröffentlicht: (2025)
Big Reasoning with Small Models: Instruction Retrieval at Inference Time
von: Alkiek, Kenan, et al.
Veröffentlicht: (2025)
von: Alkiek, Kenan, et al.
Veröffentlicht: (2025)
SemToken: Semantic-Aware Tokenization for Efficient Long-Context Language Modeling
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
Modeling Annotator Disagreement with Demographic-Aware Experts and Synthetic Perspectives
von: Xu, Yinuo, et al.
Veröffentlicht: (2025)
von: Xu, Yinuo, et al.
Veröffentlicht: (2025)
A Test of Time: Predicting the Sustainable Success of Online Collaboration in Wikipedia
von: Israeli, Abraham, et al.
Veröffentlicht: (2024)
von: Israeli, Abraham, et al.
Veröffentlicht: (2024)
Length-MAX Tokenizer for Language Models
von: Dong, Dong, et al.
Veröffentlicht: (2025)
von: Dong, Dong, et al.
Veröffentlicht: (2025)
Do LLMs Understand Social Knowledge? Evaluating the Sociability of Large Language Models with SocKET Benchmark
von: Choi, Minje, et al.
Veröffentlicht: (2023)
von: Choi, Minje, et al.
Veröffentlicht: (2023)
Moving Beyond Next-Token Prediction: Transformers are Context-Sensitive Language Generators
von: Rhee, Phill Kyu
Veröffentlicht: (2025)
von: Rhee, Phill Kyu
Veröffentlicht: (2025)
Unveiling Selection Biases: Exploring Order and Token Sensitivity in Large Language Models
von: Wei, Sheng-Lun, et al.
Veröffentlicht: (2024)
von: Wei, Sheng-Lun, et al.
Veröffentlicht: (2024)
Token-Sensitive Enclosure Semantics for Measurement-Bearing Expressions
von: Hulak, David B., et al.
Veröffentlicht: (2026)
von: Hulak, David B., et al.
Veröffentlicht: (2026)
CTPD: Cross Tokenizer Preference Distillation
von: Nguyen, Truong, et al.
Veröffentlicht: (2026)
von: Nguyen, Truong, et al.
Veröffentlicht: (2026)
Thinking Tokens for Language Modeling
von: Herel, David, et al.
Veröffentlicht: (2024)
von: Herel, David, et al.
Veröffentlicht: (2024)
SPRIG: Improving Large Language Model Performance by System Prompt Optimization
von: Zhang, Lechen, et al.
Veröffentlicht: (2024)
von: Zhang, Lechen, et al.
Veröffentlicht: (2024)
Beyond the Explicit: A Bilingual Dataset for Dehumanization Detection in Social Media
von: Assenmacher, Dennis, et al.
Veröffentlicht: (2025)
von: Assenmacher, Dennis, et al.
Veröffentlicht: (2025)
Problematic Tokens: Tokenizer Bias in Large Language Models
von: Yang, Jin, et al.
Veröffentlicht: (2024)
von: Yang, Jin, et al.
Veröffentlicht: (2024)
Phonetic Modeling of Dialectal Variation in Vietnamese Speech
von: Hoang, Quan Ngoc, et al.
Veröffentlicht: (2026)
von: Hoang, Quan Ngoc, et al.
Veröffentlicht: (2026)
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching
von: Nguyen, Truong, et al.
Veröffentlicht: (2026)
von: Nguyen, Truong, et al.
Veröffentlicht: (2026)
When "A Helpful Assistant" Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models
von: Zheng, Mingqian, et al.
Veröffentlicht: (2023)
von: Zheng, Mingqian, et al.
Veröffentlicht: (2023)
FTP: A Fine-grained Token-wise Pruner for Large Language Models via Token Routing
von: Li, Zekai, et al.
Veröffentlicht: (2024)
von: Li, Zekai, et al.
Veröffentlicht: (2024)
Rethinking Tokenization: Crafting Better Tokenizers for Large Language Models
von: Yang, Jinbiao
Veröffentlicht: (2024)
von: Yang, Jinbiao
Veröffentlicht: (2024)
Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals' Subjective Text Perceptions
von: Orlikowski, Matthias, et al.
Veröffentlicht: (2025)
von: Orlikowski, Matthias, et al.
Veröffentlicht: (2025)
Modeling Public Perceptions of Science in Media
von: Pei, Jiaxin, et al.
Veröffentlicht: (2025)
von: Pei, Jiaxin, et al.
Veröffentlicht: (2025)
An Encoder-Integrated PhoBERT with Graph Attention for Vietnamese Token-Level Classification
von: Nguyen, Ba-Quang
Veröffentlicht: (2025)
von: Nguyen, Ba-Quang
Veröffentlicht: (2025)
SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models
von: Zhang, Xin, et al.
Veröffentlicht: (2023)
von: Zhang, Xin, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Neurobiber: Fast and Interpretable Stylistic Feature Extraction
von: Alkiek, Kenan, et al.
Veröffentlicht: (2025) -
What's Mine becomes Yours: Defining, Annotating and Detecting Context-Dependent Paraphrases in News Interview Dialogs
von: Wegmann, Anna, et al.
Veröffentlicht: (2024) -
Leveraging Multilingual Training for Authorship Representation: Enhancing Generalization across Languages and Domains
von: Kim, Junghwan, et al.
Veröffentlicht: (2025) -
Are Rules Meant to be Broken? Understanding Multilingual Moral Reasoning as a Computational Pipeline with UniMoral
von: Kumar, Shivani, et al.
Veröffentlicht: (2025) -
Beyond Consensus: Perspectivist Modeling and Evaluation of Annotator Disagreement in NLP
von: Xu, Yinuo, et al.
Veröffentlicht: (2026)