Automatically retry launching EC2 instances (e.g., p5.4xlarge) until a target count is reached. Designed for GPU instance types where on-demand capacity is frequently unavailable.
- Smart retry with exponential backoff and jitter
- Error classification: retryable errors keep retrying, fatal errors stop immediately
- Tracks running instances by tag, only requests the deficit
- Health check: waits for instance status checks, terminates failed instances
- SNS notification on success or fatal error
- Dual logging: file (
launch.log) + console
python3 --version
pip install -r requirements.txtEnsure AWS CLI is configured with credentials that have EC2 and SNS permissions:
aws configure
# Or use environment variables:
# export AWS_ACCESS_KEY_ID=...
# export AWS_SECRET_ACCESS_KEY=...
# export AWS_DEFAULT_REGION=us-west-2Required IAM permissions:
ec2:RunInstancesec2:DescribeInstancesec2:DescribeInstanceStatusec2:TerminateInstancesec2:DescribeSubnetsec2:DescribeSecurityGroupsec2:DescribeImagesec2:DescribeKeyPairsec2:CreateTagssns:Publish
# Create topic
aws sns create-topic --name ec2-launcher-notify --region us-west-2
# Subscribe your email (you'll receive a confirmation email - click the link)
aws sns subscribe \
--topic-arn arn:aws:sns:us-west-2:123456789012:ec2-launcher-notify \
--protocol email \
--notification-endpoint your@email.com \
--region us-west-2Copy the Topic ARN into config.json → notification.sns_topic_arn.
If you don't have one:
aws ec2 create-key-pair \
--key-name my-key-pair \
--query 'KeyMaterial' \
--output text \
--region us-west-2 > my-key-pair.pem
chmod 400 my-key-pair.pemEnsure the security group allows SSH access (port 22) from your IP:
aws ec2 authorize-security-group-ingress \
--group-id sg-xxxxxxxxx \
--protocol tcp \
--port 22 \
--cidr YOUR_IP/32 \
--region us-west-2If you want to use SSM Session Manager instead of SSH:
- Create an IAM Role with the
AmazonSSMManagedInstanceCorepolicy - Create an Instance Profile and attach the role
- Set the ARN in
config.json→iam_instance_profile_arn
Edit config.json with your actual values:
| Field | Required | Description |
|---|---|---|
region |
Yes | AWS Region (e.g., us-west-2, us-east-1) |
instance_type |
Yes | Instance type (e.g., p5.4xlarge) |
ami_id |
Yes | AMI ID |
subnet_id |
Yes | Subnet ID (determines AZ) |
security_group_ids |
Yes | List of Security Group IDs |
key_name |
Yes | EC2 Key Pair name |
iam_instance_profile_arn |
No | IAM Instance Profile ARN (for SSM) |
target_count |
Yes | Number of instances to launch |
root_volume_size_gb |
Yes | Root EBS volume size in GB |
root_volume_type |
Yes | Root EBS volume type (gp3 recommended) |
tags.Name |
Yes | Instance Name tag (used for tracking) |
tags.Project |
Yes | Project tag (used for tracking) |
retry.initial_interval_seconds |
Yes | Initial retry interval (default: 60) |
retry.max_interval_seconds |
Yes | Max retry interval cap (default: 300) |
retry.backoff_multiplier |
Yes | Backoff multiplier (default: 1.5) |
retry.jitter |
Yes | Add random jitter to avoid thundering herd |
notification.sns_topic_arn |
Yes | SNS Topic ARN for notifications |
python3 launch.pynohup python3 launch.py > /dev/null 2>&1 &
echo $! # save the PIDscreen -S launcher
python3 launch.py
# Ctrl+A, D to detach
# screen -r launcher to reattachtmux new -s launcher
python3 launch.py
# Ctrl+B, D to detach
# tmux attach -t launcher to reattachtail -f launch.log| Error Type | Behavior |
|---|---|
InsufficientInstanceCapacity |
Exponential backoff, keep retrying |
RequestLimitExceeded / Throttling |
Double backoff, keep retrying |
InvalidParameterValue, InvalidSubnetID, UnauthorizedOperation, etc. |
Stop immediately, send SNS alert |
"Config file not found: config.json"
Run the script from the ec2-retry-launcher/ directory, or ensure config.json is in the current working directory.
Validation fails with "Invalid subnet_id" Check that the subnet exists in the configured region and your credentials have access.
"Failed to send SNS notification"
- Verify the SNS Topic ARN is correct
- Verify your email subscription is confirmed
- Check IAM permissions include
sns:Publish
Script keeps retrying with InsufficientInstanceCapacity This is normal - it means the region/AZ doesn't have capacity yet. The script will keep trying. Consider trying a different subnet (different AZ) if available.