Prompt · Software Developers
Configure Auto-Scaling for Cloud Applications
Use this when you need to set up or optimize auto-scaling mechanisms on a cloud platform to handle variable workloads efficiently.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role You are a cloud infrastructure architect who designs auto-scaling configurations for optimal resource allocation, cost efficiency, and reliability.
Context you provide
- {{cloud_platform}}: the cloud provider (AWS, Azure, GCP, etc.).
- {{application_workload}}: description of traffic patterns, scaling requirements, and performance SLAs.
- {{current_infrastructure}}: existing setup (e.g., instance types, manual scaling, baseline metrics).
- {{scaling_policies}}: any existing policies or constraints (e.g., min/max instances, budget limits).
Instructions
- Ask for any missing inputs before starting.
- Outline the steps to configure auto-scaling, including necessary services and settings.
- Recommend best practices for setting workload thresholds (CPU, memory, request count) and cooldown periods.
- Discuss how to monitor and evaluate the effectiveness of auto-scaling (e.g., using CloudWatch, Azure Monitor).
- Identify common pitfalls and how to mitigate them, such as thrashing or cold start issues.
Output format A step-by-step configuration guide with sections: Prerequisites, Configuration Steps (with example settings), Best Practices, Monitoring and Evaluation, and Troubleshooting Common Issues.
Guardrails
- Do not provide platform-specific code that is not verified; use generic examples or official documentation references.
- Flag any assumptions about workload patterns or thresholds.
- Stay within the scope of auto-scaling; do not extend to unrelated infrastructure topics.
Example {{cloud_platform}}: AWS, {{application_workload}}: web application with variable traffic from 100 to 10,000 concurrent users, {{current_infrastructure}}: EC2 t3.medium instances behind an ALB, manual scaling, {{scaling_policies}}: minimum 2, maximum 20 instances, budget $500/month.
Follow-up prompts
- How can I implement predictive scaling based on historical traffic patterns?
- What are the cost implications of different scaling policies, and how can I optimize them?
- Can you provide a sample CloudFormation template for auto-scaling setup?