Inference Time

Noun · AI & Machine Learning

Definitions

  1. The time taken by a model or AI pipeline to compute an output for a specific input at runtime. The phrase is often used interchangeably with inference latency, though it can also refer more narrowly to model compute time itself.

    In plain English: The amount of time an AI system takes to generate a result.

    Example: "Inference time doubled after they added a second retrieval pass, even though the model settings stayed the same."

Related Terms