Serving Glossary
Browse 3 serving terms defined in plain English, from the cultural dictionary of computing.
3 Serving Terms
- AI Inference Cost
- The compute and infrastructure cost of running a trained AI model to produce outputs for real users.
- Model Throughput
- The rate at which a model-serving system can process requests or generate tokens over time.
- vLLM
- An open-source inference engine for large language models focused on high-throughput serving and efficient memory usage.
Related Topics
- Inference (2 terms in common)
- Open Source (1 terms in common)
- Cost (1 terms in common)
- Infrastructure (1 terms in common)
- Performance (1 terms in common)