Inference Cost
Noun · AI & Machine Learning
Definitions
The cost incurred when running a model to produce outputs, including compute usage, provider charges, and operational overhead. Inference cost is a central product constraint because it scales directly with user activity.
In plain English: What it costs to run an AI model for each request or job.
Example: "Caching and prompt trimming cut inference cost enough to make the drafting feature viable for all customers, not just enterprise plans."