Developer Utilities Reference Data: Token Costs, Model Pricing, and Text Processing Benchmarks

Fuente: Zenodo
Saved in:
Bibliographic Details
Main Author: Khare, Mohit
Format: Recurso digital
Published: Zenodo 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866901112744837120
author Khare, Mohit
author_facet Khare, Mohit
contents <p>Reference datasets compiled by <a href="https://mohitkhare.me">Mohit Khare</a> for developer utilities and tooling. This dataset provides three categories of structured data useful for software engineers working with large language models and text processing systems.</p><p>Contents include: (1) Comparative pricing data for major LLM APIs including Anthropic Claude, OpenAI GPT, Google Gemini, Meta Llama, Mistral, and DeepSeek, with per-million-token costs and context window sizes. (2) Empirical token estimation benchmarks measuring character-to-token and word-to-token ratios across tokenizers, plus language-specific multipliers for 10 languages. (3) Performance benchmarks for common text processing operations (tokenization, NER, regex, JSON serialization) across popular Python libraries.</p><p>Related resources:</p><ul><li><a href="https://mohitkhare.me">mohitkhare.me</a> - Developer portfolio and utilities</li><li><a href="https://mohitkhare.me/blog">Blog</a> - Technical writing on development topics</li></ul>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_19269465
institution Zenodo
language
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle Developer Utilities Reference Data: Token Costs, Model Pricing, and Text Processing Benchmarks
Khare, Mohit
developer tools
token estimation
LLM pricing
text processing
benchmarks
API costs
<p>Reference datasets compiled by <a href="https://mohitkhare.me">Mohit Khare</a> for developer utilities and tooling. This dataset provides three categories of structured data useful for software engineers working with large language models and text processing systems.</p><p>Contents include: (1) Comparative pricing data for major LLM APIs including Anthropic Claude, OpenAI GPT, Google Gemini, Meta Llama, Mistral, and DeepSeek, with per-million-token costs and context window sizes. (2) Empirical token estimation benchmarks measuring character-to-token and word-to-token ratios across tokenizers, plus language-specific multipliers for 10 languages. (3) Performance benchmarks for common text processing operations (tokenization, NER, regex, JSON serialization) across popular Python libraries.</p><p>Related resources:</p><ul><li><a href="https://mohitkhare.me">mohitkhare.me</a> - Developer portfolio and utilities</li><li><a href="https://mohitkhare.me/blog">Blog</a> - Technical writing on development topics</li></ul>
title Developer Utilities Reference Data: Token Costs, Model Pricing, and Text Processing Benchmarks
topic developer tools
token estimation
LLM pricing
text processing
benchmarks
API costs
url https://doi.org/10.5281/zenodo.19269465