Speed and Conversational Large Language Models: Not All Is About Tokens per Second

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Conde, Javier, González, Miguel, Reviriego, Pedro, Gao, Zhen, Liu, Shanshan, Lombardi, Fabrizio
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916626210750464
author Conde, Javier
González, Miguel
Reviriego, Pedro
Gao, Zhen
Liu, Shanshan
Lombardi, Fabrizio
author_facet Conde, Javier
González, Miguel
Reviriego, Pedro
Gao, Zhen
Liu, Shanshan
Lombardi, Fabrizio
contents The speed of open-weights large language models (LLMs) and its dependency on the task at hand, when run on GPUs, is studied to present a comparative analysis of the speed of the most popular open LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2502_16721
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Speed and Conversational Large Language Models: Not All Is About Tokens per Second
Conde, Javier
González, Miguel
Reviriego, Pedro
Gao, Zhen
Liu, Shanshan
Lombardi, Fabrizio
Computation and Language
Artificial Intelligence
The speed of open-weights large language models (LLMs) and its dependency on the task at hand, when run on GPUs, is studied to present a comparative analysis of the speed of the most popular open LLMs.
title Speed and Conversational Large Language Models: Not All Is About Tokens per Second
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2502.16721