ks-pret-5m: a 5 million word, 12 million token kashmiri pretraining dataset
Fuente:
arXiv
Guardado en:
| Autores principales: | Malik, Haq Nawaz, Nissar, Nahfid |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ks-lit-3m: A 3.1 million word kashmiri text dataset for large language model pretraining
por: Malik, Haq Nawaz
Publicado: (2026)
por: Malik, Haq Nawaz
Publicado: (2026)
600k-ks-ocr: a large-scale synthetic dataset for optical character recognition in kashmiri script
por: Malik, Haq Nawaz
Publicado: (2026)
por: Malik, Haq Nawaz
Publicado: (2026)
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
por: Gemini Team, et al.
Publicado: (2024)
por: Gemini Team, et al.
Publicado: (2024)
BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs
por: Zhang, Sheng, et al.
Publicado: (2023)
por: Zhang, Sheng, et al.
Publicado: (2023)
synthocr-gen: A synthetic ocr dataset generator for low-resource languages- breaking the data barrier
por: Malik, Haq Nawaz, et al.
Publicado: (2026)
por: Malik, Haq Nawaz, et al.
Publicado: (2026)
MIRIAD: Augmenting LLMs with millions of medical query-response pairs
por: Zheng, Qinyue, et al.
Publicado: (2025)
por: Zheng, Qinyue, et al.
Publicado: (2025)
FSboard: Over 3 million characters of ASL fingerspelling collected via smartphones
por: Georg, Manfred, et al.
Publicado: (2024)
por: Georg, Manfred, et al.
Publicado: (2024)
Tgea: An error-annotated dataset and benchmark tasks for text generation from pretrained language models
por: He, Jie, et al.
Publicado: (2025)
por: He, Jie, et al.
Publicado: (2025)
Synthetic bootstrapped pretraining
por: Yang, Zitong, et al.
Publicado: (2025)
por: Yang, Zitong, et al.
Publicado: (2025)
Does mBERT understand Romansh? Evaluating word embeddings using word alignment
por: Dolev, Eyal Liron
Publicado: (2023)
por: Dolev, Eyal Liron
Publicado: (2023)
Synthetic continued pretraining
por: Yang, Zitong, et al.
Publicado: (2024)
por: Yang, Zitong, et al.
Publicado: (2024)
Let a million entrepreneurs grow!
por: Kumar, Mrityunjay
Publicado: (2024)
por: Kumar, Mrityunjay
Publicado: (2024)
Emergent inabilities? Inverse scaling over the course of pretraining
por: Michaelov, James A., et al.
Publicado: (2023)
por: Michaelov, James A., et al.
Publicado: (2023)
Why do LLMs attend to the first token?
por: Barbero, Federico, et al.
Publicado: (2025)
por: Barbero, Federico, et al.
Publicado: (2025)
Where is the signal in tokenization space?
por: Geh, Renato Lui, et al.
Publicado: (2024)
por: Geh, Renato Lui, et al.
Publicado: (2024)
TextGram: Towards a better domain-adaptive pretraining
por: Hiwarkhedkar, Sharayu, et al.
Publicado: (2024)
por: Hiwarkhedkar, Sharayu, et al.
Publicado: (2024)
What is a word?
por: Murphy, Elliot
Publicado: (2024)
por: Murphy, Elliot
Publicado: (2024)
Institutional Books 1.0: A 242B token dataset from Harvard Library's collections, refined for accuracy and usability
por: Cargnelutti, Matteo, et al.
Publicado: (2025)
por: Cargnelutti, Matteo, et al.
Publicado: (2025)
Comparative analysis of subword tokenization approaches for Indian languages
por: Das, Sudhansu Bala, et al.
Publicado: (2025)
por: Das, Sudhansu Bala, et al.
Publicado: (2025)
Interpreting token compositionality in LLMs: A robustness analysis
por: Aljaafari, Nura, et al.
Publicado: (2024)
por: Aljaafari, Nura, et al.
Publicado: (2024)
Contextual morphologically-guided tokenization for Latin encoder models
por: Hudspeth, Marisa, et al.
Publicado: (2025)
por: Hudspeth, Marisa, et al.
Publicado: (2025)
AnomaLLMy -- Detecting anomalous tokens in black-box LLMs through low-confidence single-token predictions
por: Witold, Waligóra
Publicado: (2024)
por: Witold, Waligóra
Publicado: (2024)
How far can bias go? Tracing bias from pretraining data to alignment
por: Thaler, Marion, et al.
Publicado: (2024)
por: Thaler, Marion, et al.
Publicado: (2024)
Looking beyond the next token
por: Thankaraj, Abitha, et al.
Publicado: (2025)
por: Thankaraj, Abitha, et al.
Publicado: (2025)
The pitfalls of next-token prediction
por: Bachmann, Gregor, et al.
Publicado: (2024)
por: Bachmann, Gregor, et al.
Publicado: (2024)
Mapping 12 million lakes globally at 10 m resolution
por: Hu, Beihui, et al.
Publicado: (2025)
por: Hu, Beihui, et al.
Publicado: (2025)
On multi-token prediction for efficient LLM inference
por: Mehra, Somesh, et al.
Publicado: (2025)
por: Mehra, Somesh, et al.
Publicado: (2025)
BPDec: Unveiling the Potential of Masked Language Modeling Decoder in BERT pretraining
por: Liang, Wen, et al.
Publicado: (2024)
por: Liang, Wen, et al.
Publicado: (2024)
Do pretrained Transformers Learn In-Context by Gradient Descent?
por: Shen, Lingfeng, et al.
Publicado: (2023)
por: Shen, Lingfeng, et al.
Publicado: (2023)
Finetuning LLMs for EvaCun 2025 token prediction shared task
por: Jon, Josef, et al.
Publicado: (2025)
por: Jon, Josef, et al.
Publicado: (2025)
Better & Faster Large Language Models via Multi-token Prediction
por: Gloeckle, Fabian, et al.
Publicado: (2024)
por: Gloeckle, Fabian, et al.
Publicado: (2024)
Collaborative decoding of critical tokens for boosting factuality of large language models
por: Jin, Lifeng, et al.
Publicado: (2024)
por: Jin, Lifeng, et al.
Publicado: (2024)
LBPE: Long-token-first Tokenization to Improve Large Language Models
por: Lian, Haoran, et al.
Publicado: (2024)
por: Lian, Haoran, et al.
Publicado: (2024)
Getting the most out of your tokenizer for pre-training and domain adaptation
por: Dagan, Gautier, et al.
Publicado: (2024)
por: Dagan, Gautier, et al.
Publicado: (2024)
Let your LLM generate a few tokens and you will reduce the need for retrieval
por: Déjean, Hervé
Publicado: (2024)
por: Déjean, Hervé
Publicado: (2024)
[Call for Papers] The 2nd BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus
por: Choshen, Leshem, et al.
Publicado: (2024)
por: Choshen, Leshem, et al.
Publicado: (2024)
Jacobian Scopes: token-level causal attributions in LLMs
por: Liu, Toni J. B., et al.
Publicado: (2026)
por: Liu, Toni J. B., et al.
Publicado: (2026)
Prediction hubs are context-informed frequent tokens in LLMs
por: Nielsen, Beatrix M. G., et al.
Publicado: (2025)
por: Nielsen, Beatrix M. G., et al.
Publicado: (2025)
Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs
por: Singh, Aaditya K., et al.
Publicado: (2024)
por: Singh, Aaditya K., et al.
Publicado: (2024)
Do language models plan ahead for future tokens?
por: Wu, Wilson, et al.
Publicado: (2024)
por: Wu, Wilson, et al.
Publicado: (2024)
Ejemplares similares
-
ks-lit-3m: A 3.1 million word kashmiri text dataset for large language model pretraining
por: Malik, Haq Nawaz
Publicado: (2026) -
600k-ks-ocr: a large-scale synthetic dataset for optical character recognition in kashmiri script
por: Malik, Haq Nawaz
Publicado: (2026) -
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
por: Gemini Team, et al.
Publicado: (2024) -
BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs
por: Zhang, Sheng, et al.
Publicado: (2023) -
synthocr-gen: A synthetic ocr dataset generator for low-resource languages- breaking the data barrier
por: Malik, Haq Nawaz, et al.
Publicado: (2026)