Auto-scaling
Noun · Development
Definitions
The automatic adjustment of compute resources — adding or removing server instances, containers, or function invocations — in response to real-time demand metrics such as CPU utilization, request rate, or queue depth.
In plain English: Automatically adding more servers when traffic spikes and removing them when things quiet down, so you only pay for what you need.
Example: "Our auto-scaling policy spins up extra pods when CPU crosses 70% and scales back down after 5 minutes of calm."
Related Terms
- infrastructure as code
- Cloud
- Container
- Deployment
- IaaS
- PaaS
- Supervisor
- ConfigMap
- Environment Variable
- Object Storage
- Block Storage
- Availability Zone
- Region
- Spot Instance
- Auto-scaling Group
- Cold Start
- LLVM
- Artifact Registry
- Elastic Beanstalk
- Elastic Search
- Google Cloud Run
- Launch Configuration
- Package Registry
- Scale Out
- Scale Up
- Self-Hosting
- Terraform Provider
- Test Environment
- Toolchain
- Vagrant
- Virtual Device
- Virtualization
- Hyperscaler
- OpenStack
- Coreweave
- Cloud Security
- DevOps Engineer