The Effect of Scripts and Formats on LLM Numeracy
Fuente:
arXiv
Saved in:
| Main Authors: | Reddy, Varshini, Schmidt, Craig W., Ebner, Seth, Wiemerslage, Adam, Pinter, Yuval, Tanner, Chris |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Tokenization with Split Trees
by: Schmidt, Craig W., et al.
Published: (2026)
by: Schmidt, Craig W., et al.
Published: (2026)
How Much is Enough? The Diminishing Returns of Tokenization Training Data
by: Reddy, Varshini, et al.
Published: (2025)
by: Reddy, Varshini, et al.
Published: (2025)
Boundless Byte Pair Encoding: Breaking the Pre-tokenization Barrier
by: Schmidt, Craig W., et al.
Published: (2025)
by: Schmidt, Craig W., et al.
Published: (2025)
No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding
by: Krumdick, Michael, et al.
Published: (2025)
by: Krumdick, Michael, et al.
Published: (2025)
Faster Superword Tokenization
by: Schmidt, Craig W., et al.
Published: (2026)
by: Schmidt, Craig W., et al.
Published: (2026)
Cost-Efficient Estimation of General Abilities Across Benchmarks
by: Krumdick, Michael, et al.
Published: (2026)
by: Krumdick, Michael, et al.
Published: (2026)
Greed is All You Need: An Evaluation of Tokenizer Inference Methods
by: Uzan, Omri, et al.
Published: (2024)
by: Uzan, Omri, et al.
Published: (2024)
Tokenization Is More Than Compression
by: Schmidt, Craig W., et al.
Published: (2024)
by: Schmidt, Craig W., et al.
Published: (2024)
On Finding Inconsistencies in Documents
by: Lovering, Charles J., et al.
Published: (2025)
by: Lovering, Charles J., et al.
Published: (2025)
SEC-QA: A Systematic Evaluation Corpus for Financial QA
by: Lai, Viet Dac, et al.
Published: (2024)
by: Lai, Viet Dac, et al.
Published: (2024)
Entropy-Driven Pre-Tokenization for Byte-Pair Encoding
by: Hu, Yifan, et al.
Published: (2025)
by: Hu, Yifan, et al.
Published: (2025)
The Degree of Language Diacriticity and Its Effect on Tasks
by: Cohen, Adi, et al.
Published: (2026)
by: Cohen, Adi, et al.
Published: (2026)
From Algebraic Word Problem to Program: A Formalized Approach
by: Wiemerslage, Adam, et al.
Published: (2020)
by: Wiemerslage, Adam, et al.
Published: (2020)
BizBench: A Quantitative Reasoning Benchmark for Business and Finance
by: Koncel-Kedziorski, Rik, et al.
Published: (2023)
by: Koncel-Kedziorski, Rik, et al.
Published: (2023)
Probing Subphonemes in Morphology Models
by: Astrach, Gal, et al.
Published: (2025)
by: Astrach, Gal, et al.
Published: (2025)
Hebrew Diacritics Restoration using Visual Representation
by: Elboher, Yair, et al.
Published: (2025)
by: Elboher, Yair, et al.
Published: (2025)
Don't Touch My Diacritics
by: Gorman, Kyle, et al.
Published: (2024)
by: Gorman, Kyle, et al.
Published: (2024)
BiVert: Bidirectional Vocabulary Evaluation using Relations for Machine Translation
by: Cherf, Carinne, et al.
Published: (2024)
by: Cherf, Carinne, et al.
Published: (2024)
Information Types in Product Reviews
by: Shapira, Ori, et al.
Published: (2025)
by: Shapira, Ori, et al.
Published: (2025)
CharBench: Evaluating the Role of Tokenization in Character-Level Tasks
by: Uzan, Omri, et al.
Published: (2025)
by: Uzan, Omri, et al.
Published: (2025)
Improving Low-Resource Morphological Inflection via Self-Supervised Objectives
by: Wiemerslage, Adam, et al.
Published: (2025)
by: Wiemerslage, Adam, et al.
Published: (2025)
Which Pieces Does Unigram Tokenization Really Need?
by: Land, Sander, et al.
Published: (2025)
by: Land, Sander, et al.
Published: (2025)
DocFinQA: A Long-Context Financial Reasoning Dataset
by: Reddy, Varshini, et al.
Published: (2024)
by: Reddy, Varshini, et al.
Published: (2024)
Model-Based Ranking of Source Languages for Zero-Shot Cross-Lingual Transfer
by: Ebrahimi, Abteen, et al.
Published: (2025)
by: Ebrahimi, Abteen, et al.
Published: (2025)
Splintering Nonconcatenative Languages for Better Tokenization
by: Gazit, Bar, et al.
Published: (2025)
by: Gazit, Bar, et al.
Published: (2025)
Protecting Privacy in Classifiers by Token Manipulation
by: Harel, Re'em, et al.
Published: (2024)
by: Harel, Re'em, et al.
Published: (2024)
Token-Level Privacy in Large Language Models
by: Harel, Re'em, et al.
Published: (2025)
by: Harel, Re'em, et al.
Published: (2025)
FrontierFinance: A Long-Horizon Computer-Use Benchmark of Real-World Financial Tasks
by: Krumdick, Michael, et al.
Published: (2026)
by: Krumdick, Michael, et al.
Published: (2026)
An Analysis of BPE Vocabulary Trimming in Neural Machine Translation
by: Cognetta, Marco, et al.
Published: (2024)
by: Cognetta, Marco, et al.
Published: (2024)
One Language, Two Scripts: Probing Script-Invariance in LLM Concept Representations
by: Karne, Sripad
Published: (2026)
by: Karne, Sripad
Published: (2026)
Exploring Internal Numeracy in Language Models: A Case Study on ALBERT
by: Wennberg, Ulme, et al.
Published: (2024)
by: Wennberg, Ulme, et al.
Published: (2024)
Revealing the Numeracy Gap: An Empirical Investigation of Text Embedding Models
by: Deng, Ningyuan, et al.
Published: (2025)
by: Deng, Ningyuan, et al.
Published: (2025)
An Analysis of Multilingual FActScore
by: Vu, Kim Trong, et al.
Published: (2024)
by: Vu, Kim Trong, et al.
Published: (2024)
A Closer Look at Claim Decomposition
by: Wanner, Miriam, et al.
Published: (2024)
by: Wanner, Miriam, et al.
Published: (2024)
Pre-Calc: Learning to Use the Calculator Improves Numeracy in Language Models
by: Veerendranath, Vishruth, et al.
Published: (2024)
by: Veerendranath, Vishruth, et al.
Published: (2024)
Unknown Script: Impact of Script on Cross-Lingual Transfer
by: Tufa, Wondimagegnhue Tsegaye, et al.
Published: (2024)
by: Tufa, Wondimagegnhue Tsegaye, et al.
Published: (2024)
How Robust Are Large Language Models for Clinical Numeracy? An Empirical Study on Numerical Reasoning Abilities in Clinical Contexts
by: Nguyen, Minh-Vuong, et al.
Published: (2026)
by: Nguyen, Minh-Vuong, et al.
Published: (2026)
Language Model Probabilities are Not Calibrated in Numeric Contexts
by: Lovering, Charles, et al.
Published: (2024)
by: Lovering, Charles, et al.
Published: (2024)
SkyScript-100M: 1,000,000,000 Pairs of Scripts and Shooting Scripts for Short Drama
by: Tang, Jing, et al.
Published: (2024)
by: Tang, Jing, et al.
Published: (2024)
Exposing Numeracy Gaps: A Benchmark to Evaluate Fundamental Numerical Abilities in Large Language Models
by: Li, Haoyang, et al.
Published: (2025)
by: Li, Haoyang, et al.
Published: (2025)
Similar Items
-
Tokenization with Split Trees
by: Schmidt, Craig W., et al.
Published: (2026) -
How Much is Enough? The Diminishing Returns of Tokenization Training Data
by: Reddy, Varshini, et al.
Published: (2025) -
Boundless Byte Pair Encoding: Breaking the Pre-tokenization Barrier
by: Schmidt, Craig W., et al.
Published: (2025) -
No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding
by: Krumdick, Michael, et al.
Published: (2025) -
Faster Superword Tokenization
by: Schmidt, Craig W., et al.
Published: (2026)