| _version_ | 1866901112744837120 |
|---|---|
| author | Khare, Mohit |
| author_facet | Khare, Mohit |
| contents | <p>Reference datasets compiled by <a href="https://mohitkhare.me">Mohit Khare</a> for developer utilities and tooling. This dataset provides three categories of structured data useful for software engineers working with large language models and text processing systems.</p><p>Contents include: (1) Comparative pricing data for major LLM APIs including Anthropic Claude, OpenAI GPT, Google Gemini, Meta Llama, Mistral, and DeepSeek, with per-million-token costs and context window sizes. (2) Empirical token estimation benchmarks measuring character-to-token and word-to-token ratios across tokenizers, plus language-specific multipliers for 10 languages. (3) Performance benchmarks for common text processing operations (tokenization, NER, regex, JSON serialization) across popular Python libraries.</p><p>Related resources:</p><ul><li><a href="https://mohitkhare.me">mohitkhare.me</a> - Developer portfolio and utilities</li><li><a href="https://mohitkhare.me/blog">Blog</a> - Technical writing on development topics</li></ul> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_19269465 |
| institution | Zenodo |
| language | |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Developer Utilities Reference Data: Token Costs, Model Pricing, and Text Processing Benchmarks Khare, Mohit developer tools token estimation LLM pricing text processing benchmarks API costs <p>Reference datasets compiled by <a href="https://mohitkhare.me">Mohit Khare</a> for developer utilities and tooling. This dataset provides three categories of structured data useful for software engineers working with large language models and text processing systems.</p><p>Contents include: (1) Comparative pricing data for major LLM APIs including Anthropic Claude, OpenAI GPT, Google Gemini, Meta Llama, Mistral, and DeepSeek, with per-million-token costs and context window sizes. (2) Empirical token estimation benchmarks measuring character-to-token and word-to-token ratios across tokenizers, plus language-specific multipliers for 10 languages. (3) Performance benchmarks for common text processing operations (tokenization, NER, regex, JSON serialization) across popular Python libraries.</p><p>Related resources:</p><ul><li><a href="https://mohitkhare.me">mohitkhare.me</a> - Developer portfolio and utilities</li><li><a href="https://mohitkhare.me/blog">Blog</a> - Technical writing on development topics</li></ul> |
| title | Developer Utilities Reference Data: Token Costs, Model Pricing, and Text Processing Benchmarks |
| topic | developer tools token estimation LLM pricing text processing benchmarks API costs |
| url | https://doi.org/10.5281/zenodo.19269465 |