tokens per second
/TOH-kunz per SEK-und/ · noun · AI & Machine Learning · Origin: 2023
Definitions
The standard throughput metric for language model inference, measuring how many tokens a model can generate per second. Higher TPS enables more responsive user experiences and lower serving costs. Varies dramatically based on model size, hardware, quantization, and batch size.
In plain English: How fast an AI can spit out words — more tokens per second means you get your answer faster.
Example: After switching to the quantized model, we went from 15 tokens per second to 80, and users stopped complaining about the loading spinner.