Auto-scaling a cloud server means your infrastructure automatically adds or removes resources — CPU, RAM, instances — in response to real demand, without manual intervention.
This guide covers what auto-scaling is, when it makes sense to enable it, and how to configure it both on a managed cloud platform and on your own VPS.
Vertical vs. Horizontal Scaling
Before configuring anything, it's important to understand the difference between the two scaling models:
| Type | What it does | When to use it | Limit |
|---|---|---|---|
| Vertical (scale up/down) | Increases CPU/RAM on the same instance | Monolithic apps, databases | Maximum plan size |
| Horizontal (scale out/in) | Adds or removes identical instances | Stateless apps, APIs, microservices | Practically unlimited |
The most common and powerful auto-scaling model is horizontal: a group of instances behind a load balancer that grows or shrinks based on defined metrics (CPU, requests per second, latency).
Key Concepts Before Configuring
Regardless of the provider, you'll need three components:
- Instance template (Launch Template / Image): the exact machine configuration that will be cloned when scaling. It must be immutable and idempotent.
- Auto Scaling Group: defines the minimum, maximum, and desired number of instances.
- Scaling policies: the rules that trigger growth or reduction (based on metrics or a schedule).
Your application must also be stateless: user sessions, uploaded files, and any state must live outside the instances (database, Redis, object storage like S3 or equivalent). If an instance is terminated when the group shrinks, it cannot lose data.
How to Set Up Auto Scaling on AWS (EC2 Auto Scaling)
AWS has the most mature auto-scaling system. The basic process:
1. Create a Launch Template
In the AWS console → EC2 → Launch Templates. Define the AMI, instance type, SSH key, security groups, and a startup script (User Data) that automatically installs your application.
2. Create the Auto Scaling Group
- EC2 → Auto Scaling Groups → Create.
- Select the Launch Template you created.
- Choose availability zones (use at least two for high availability).
- Attach an existing Application Load Balancer or create a new one.
- Set: minimum = 1, desired = 2, maximum = 10 (adjust to your budget).
3. Add Scaling Policies
The simplest and most recommended policy is Target Tracking: the group tries to keep a metric at a target value.
# Example with AWS CLI
aws autoscaling put-scaling-policy \
--auto-scaling-group-name my-group \
--policy-name scale-by-cpu \
--policy-type TargetTrackingScaling \
--target-tracking-configuration '{
"PredefinedMetricSpecification": {
"PredefinedMetricType": "ASGAverageCPUUtilization"
},
"TargetValue": 60.0
}'
With this, if average CPU exceeds 60%, AWS adds instances; if it drops below 60% for a sustained period, it removes them.
Auto-Scaling on DigitalOcean and Hetzner
Not every project needs the complexity of AWS. Simpler providers also offer auto-scaling options:
DigitalOcean
DigitalOcean App Platform includes native auto-scaling for containerized applications. For Droplets (VMs), horizontal scaling is managed via Load Balancers + Droplet backups / Snapshots with provisioning scripts, or through Terraform using the official DO provider.
Hetzner Cloud
Hetzner doesn't have native group auto-scaling, but its REST API is straightforward and you can automate it with tools like hcloud-autoscaler (open source) or a custom script that queries metrics and calls the API to create or delete servers.
For mid-sized projects on a tight budget, Hetzner plus a lightweight orchestrator is a competitive option. If you need a ready-to-use managed solution, the team at elenlace.com can guide you on which provider best fits your workload.
Best Practices for Robust Auto-Scaling
- Test cold starts: simulate a new instance joining the group from scratch. Does it take less than 2 minutes to be ready? That's the target.
- Cooldown period: set a cooldown period (300 seconds by default in AWS) to prevent the system from adding and removing instances in rapid loops.
- Connection draining: before an instance is removed, the load balancer must stop sending new traffic to it and wait for in-progress requests to finish.
- Billing alerts: a bug in a scaling policy can spin up hundreds of instances. Set up spending alerts with your provider.
- Resource tagging: use tags (
env=prod,project=myapp) to identify group instances in your bill and logs.
Browse more cloud architecture guides in our cloud hosting section, where you'll find provider comparisons, security tutorials, and real-world use cases.
Key Takeaways
- Horizontal auto-scaling adds or removes instances based on real metrics, eliminating bottlenecks from traffic spikes.
- Your application must be stateless to benefit from auto-scaling: sessions and files must live in services external to the server.
- AWS EC2 Auto Scaling with Target Tracking is the most mature solution; DigitalOcean App Platform and Hetzner with scripting are more cost-effective alternatives for mid-sized projects.
- Always configure cooldown periods, connection draining, and billing alerts to prevent unexpected behavior.
- A well-defined instance template (immutable, with automated provisioning) is the foundation of any reliable auto-scaling group.
Is your application already experiencing traffic spikes and you're not sure where to start? Talk to us at elenlace.com and we'll design a scalable cloud architecture tailored to your budget.
FAQ
Does auto-scaling terminate instances even if they're processing requests?
No, if properly configured. The load balancer stops sending new traffic to the instance being removed and waits for active connections to finish (connection draining). Only then is the instance terminated.
How much does auto-scaling cost on AWS?
The Auto Scaling Groups service itself has no additional cost; you pay only for the EC2 instances that are running. The cost varies by instance type and region.
Can I use auto-scaling with databases?
For relational databases, horizontal scaling is complex (it involves replication and sharding). The most practical option is to use managed services like Amazon RDS with Multi-AZ or Aurora Serverless, which handle read scaling and, in the case of Serverless, write scaling as well.
Useful resources
Other providers and guides worth comparing: