Inference Budget
Noun · AI & Machine Learning
Definitions
The amount of compute, latency, tokens, or money available for running inference on a task or product. Inference budgets force tradeoffs among model quality, context size, tool use, and responsiveness.
In plain English: The amount of resources available for running an AI model on a request.
Example: "The assistant could call tools only twice per request because the team needed to stay within the inference budget for free-tier accounts."