Inference Cost

Noun · AI & Machine Learning

Definitions

  1. The cost incurred when running a model to produce outputs, including compute usage, provider charges, and operational overhead. Inference cost is a central product constraint because it scales directly with user activity.

    In plain English: What it costs to run an AI model for each request or job.

    Example: "Caching and prompt trimming cut inference cost enough to make the drafting feature viable for all customers, not just enterprise plans."

Related Terms