Inference Budget

Noun · AI & Machine Learning

Definitions

  1. The amount of compute, latency, tokens, or money available for running inference on a task or product. Inference budgets force tradeoffs among model quality, context size, tool use, and responsiveness.

    In plain English: The amount of resources available for running an AI model on a request.

    Example: "The assistant could call tools only twice per request because the team needed to stay within the inference budget for free-tier accounts."

Related Terms