tokens per second

/TOH-kunz per SEK-und/ · noun · AI & Machine Learning · Origin: 2023

Definitions

  1. The standard throughput metric for language model inference, measuring how many tokens a model can generate per second. Higher TPS enables more responsive user experiences and lower serving costs. Varies dramatically based on model size, hardware, quantization, and batch size.

    In plain English: How fast an AI can spit out words — more tokens per second means you get your answer faster.

    Example: After switching to the quantized model, we went from 15 tokens per second to 80, and users stopped complaining about the loading spinner.

Related Terms