Is Sanskrit the most token-efficient language? A quantitative study using GPT, Gemini, and SentencePiece
Fuente:
arXiv
Saved in:
| Main Author: | Kumar, Anshul |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Pragya: An AI-Based Semantic Recommendation System for Sanskrit Subhasitas
by: Raorane, Tanisha, et al.
Published: (2026)
by: Raorane, Tanisha, et al.
Published: (2026)
Zero-Shot End-to-End Relation Extraction in Chinese: A Comparative Study of Gemini, LLaMA and ChatGPT
by: Du, Shaoshuai, et al.
Published: (2025)
by: Du, Shaoshuai, et al.
Published: (2025)
Looking beyond the next token
by: Thankaraj, Abitha, et al.
Published: (2025)
by: Thankaraj, Abitha, et al.
Published: (2025)
The pitfalls of next-token prediction
by: Bachmann, Gregor, et al.
Published: (2024)
by: Bachmann, Gregor, et al.
Published: (2024)
ChatGPT vs Human-authored Text: Insights into Controllable Text Summarization and Sentence Style Transfer
by: Liu, Dongqi, et al.
Published: (2023)
by: Liu, Dongqi, et al.
Published: (2023)
Building Production-Ready Probes For Gemini
by: Kramár, János, et al.
Published: (2026)
by: Kramár, János, et al.
Published: (2026)
Shaping capabilities with token-level data filtering
by: Rathi, Neil, et al.
Published: (2026)
by: Rathi, Neil, et al.
Published: (2026)
Scaling Transformer to 1M tokens and beyond with RMT
by: Bulatov, Aydar, et al.
Published: (2023)
by: Bulatov, Aydar, et al.
Published: (2023)
Interpretable Next-token Prediction via the Generalized Induction Head
by: Kim, Eunji, et al.
Published: (2024)
by: Kim, Eunji, et al.
Published: (2024)
Language models are better than humans at next-token prediction
by: Shlegeris, Buck, et al.
Published: (2022)
by: Shlegeris, Buck, et al.
Published: (2022)
BgGPT 1.0: Extending English-centric LLMs to other languages
by: Alexandrov, Anton, et al.
Published: (2024)
by: Alexandrov, Anton, et al.
Published: (2024)
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
by: Zhu, Yuxuan, et al.
Published: (2025)
by: Zhu, Yuxuan, et al.
Published: (2025)
Perturbation: A simple and efficient adversarial tracer for representation learning in language models
by: Rozner, Joshua, et al.
Published: (2026)
by: Rozner, Joshua, et al.
Published: (2026)
COMPACT: Common-token Optimized Model Pruning Across Channels and Tokens
by: Kwek, Eugene, et al.
Published: (2025)
by: Kwek, Eugene, et al.
Published: (2025)
All or None: Identifiable Linear Properties of Next-token Predictors in Language Modeling
by: Marconato, Emanuele, et al.
Published: (2024)
by: Marconato, Emanuele, et al.
Published: (2024)
Essential-Web v1.0: 24T tokens of organized web data
by: AI, Essential, et al.
Published: (2025)
by: AI, Essential, et al.
Published: (2025)
You only need 4 extra tokens: Synergistic Test-time Adaptation for LLMs
by: Xu, Yijie, et al.
Published: (2025)
by: Xu, Yijie, et al.
Published: (2025)
A Fuzzy Evaluation of Sentence Encoders on Grooming Risk Classification
by: Bihani, Geetanjali, et al.
Published: (2025)
by: Bihani, Geetanjali, et al.
Published: (2025)
Static Word Embeddings for Sentence Semantic Representation
by: Wada, Takashi, et al.
Published: (2025)
by: Wada, Takashi, et al.
Published: (2025)
Low-Cost Generation and Evaluation of Dictionary Example Sentences
by: Cai, Bill, et al.
Published: (2024)
by: Cai, Bill, et al.
Published: (2024)
Can large language models replace humans in the systematic review process? Evaluating GPT-4's efficacy in screening and extracting data from peer-reviewed and grey literature in multiple languages
by: Khraisha, Qusai, et al.
Published: (2023)
by: Khraisha, Qusai, et al.
Published: (2023)
Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
by: Nagarajan, Vaishnavh, et al.
Published: (2025)
by: Nagarajan, Vaishnavh, et al.
Published: (2025)
Capabilities of Gemini Models in Medicine
by: Saab, Khaled, et al.
Published: (2024)
by: Saab, Khaled, et al.
Published: (2024)
Towards Nepali-language LLMs: Efficient GPT training with a Nepali BPE tokenizer
by: Shrestha, Adarsha, et al.
Published: (2025)
by: Shrestha, Adarsha, et al.
Published: (2025)
Architectural Flaw Detection in Civil Engineering Using GPT-4
by: Kumar, Saket, et al.
Published: (2024)
by: Kumar, Saket, et al.
Published: (2024)
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
by: Gemini Team, et al.
Published: (2024)
by: Gemini Team, et al.
Published: (2024)
Segment Any Text: A Universal Approach for Robust, Efficient and Adaptable Sentence Segmentation
by: Frohmann, Markus, et al.
Published: (2024)
by: Frohmann, Markus, et al.
Published: (2024)
Improving Faithfulness of Abstractive Summarization by Controlling Confounding Effect of Irrelevant Sentences
by: Ghoshal, Asish, et al.
Published: (2022)
by: Ghoshal, Asish, et al.
Published: (2022)
Stream Aligner: Efficient Sentence-Level Alignment via Distribution Induction
by: Lou, Hantao, et al.
Published: (2025)
by: Lou, Hantao, et al.
Published: (2025)
2-Tier SimCSE: Elevating BERT for Robust Sentence Embeddings
by: Wang, Yumeng, et al.
Published: (2025)
by: Wang, Yumeng, et al.
Published: (2025)
Advancing Multimodal Medical Capabilities of Gemini
by: Yang, Lin, et al.
Published: (2024)
by: Yang, Lin, et al.
Published: (2024)
Towards Linguistic Neural Representation Learning and Sentence Retrieval from Electroencephalogram Recordings
by: Zhou, Jinzhao, et al.
Published: (2024)
by: Zhou, Jinzhao, et al.
Published: (2024)
BERT-ASC: Auxiliary-Sentence Construction for Implicit Aspect Learning in Sentiment Analysis
by: Ahmed, Murtadha, et al.
Published: (2022)
by: Ahmed, Murtadha, et al.
Published: (2022)
Memory Tokens: Large Language Models Can Generate Reversible Sentence Embeddings
by: Sastre, Ignacio, et al.
Published: (2025)
by: Sastre, Ignacio, et al.
Published: (2025)
A general tensor-structured compression scheme for efficient large language models
by: Lu, Ying, et al.
Published: (2026)
by: Lu, Ying, et al.
Published: (2026)
On the generalization of language models from in-context learning and finetuning: a controlled study
by: Lampinen, Andrew K., et al.
Published: (2025)
by: Lampinen, Andrew K., et al.
Published: (2025)
ArabianGPT: Native Arabic GPT-based Large Language Model
by: Koubaa, Anis, et al.
Published: (2024)
by: Koubaa, Anis, et al.
Published: (2024)
Persona-Coded Poly-Encoder: Persona-Guided Multi-Stream Conversational Sentence Scoring
by: Liu, Junfeng, et al.
Published: (2023)
by: Liu, Junfeng, et al.
Published: (2023)
Single layer tiny Co$^4$ outpaces GPT-2 and GPT-BERT
by: Zain, Noor Ul, et al.
Published: (2025)
by: Zain, Noor Ul, et al.
Published: (2025)
RobustSentEmbed: Robust Sentence Embeddings Using Adversarial Self-Supervised Contrastive Learning
by: Asl, Javad Rafiei, et al.
Published: (2024)
by: Asl, Javad Rafiei, et al.
Published: (2024)
Similar Items
-
Pragya: An AI-Based Semantic Recommendation System for Sanskrit Subhasitas
by: Raorane, Tanisha, et al.
Published: (2026) -
Zero-Shot End-to-End Relation Extraction in Chinese: A Comparative Study of Gemini, LLaMA and ChatGPT
by: Du, Shaoshuai, et al.
Published: (2025) -
Looking beyond the next token
by: Thankaraj, Abitha, et al.
Published: (2025) -
The pitfalls of next-token prediction
by: Bachmann, Gregor, et al.
Published: (2024) -
ChatGPT vs Human-authored Text: Insights into Controllable Text Summarization and Sentence Style Transfer
by: Liu, Dongqi, et al.
Published: (2023)