Inference Speed
Noun · AI & Machine Learning
Definitions
How quickly a model can process inputs and generate outputs during inference, often measured in latency or tokens per second. Inference speed influences user experience directly, especially in interactive applications.
In plain English: The speed at which an AI model produces outputs.
Example: "The team accepted slightly shorter answers in exchange for better inference speed during live customer chats."