Inference Speed

Noun · AI & Machine Learning

Definitions

  1. How quickly a model can process inputs and generate outputs during inference, often measured in latency or tokens per second. Inference speed influences user experience directly, especially in interactive applications.

    In plain English: The speed at which an AI model produces outputs.

    Example: "The team accepted slightly shorter answers in exchange for better inference speed during live customer chats."

Related Terms