Inference Time
Noun · AI & Machine Learning
Definitions
The time taken by a model or AI pipeline to compute an output for a specific input at runtime. The phrase is often used interchangeably with inference latency, though it can also refer more narrowly to model compute time itself.
In plain English: The amount of time an AI system takes to generate a result.
Example: "Inference time doubled after they added a second retrieval pass, even though the model settings stayed the same."