Complete AI Training

Prompt course · 8 lessons · 25 prompts · 1 hour · Beginner

AI for Site Reliability Engineers

A prompt course built around the work site reliability engineers actually do. Eight lessons take you from incident updates to postmortems, capacity planning, and design reviews.

Share

What you'll learn

  • Clear Incident Updates: Turn raw incident details into updates, handoffs, and stakeholder explanations that people can follow under pressure.
  • Runbooks And Scripts: Convert what you know into runbooks, script skeletons, and plain explanations of infrastructure changes.
  • Quieter Alerts: Write better queries, cut alert noise, and shape dashboards around what users actually feel.
  • Reading Logs Faster: Interpret logs, connect symptoms across services, and choose the next debugging step with more confidence.
  • Blameless Postmortems: Draft postmortems, pull out action items, and turn each incident into something the team learns from.
  • Capacity And Cost: Forecast capacity, review cloud spend, and plan rightsizing changes without overprovisioning.
  • Design Reviews And Rollbacks: Stress-test designs, plan chaos experiments, and prepare rollback steps before changes go live.
  • Incident Preparedness: Build incident simulations, check runbook coverage, and define metrics that improve the next response.

What's inside

8 lessons · 25 prompts
  1. Before you start · framework course Context Engineering and Structured PromptsContext engineering helps you build structured prompts for SRE runbooks, like an agent that diagnoses a latency spike from logs and dashboards.
  2. Start here Priya's Thursday, two waysA day in the life of a Site Reliability Engineer, before and after these prompts.
  3. 01 Lesson 1 · 3 prompts Incident Communication Basics
  4. 02 Lesson 2 · 3 prompts Runbooks And Automation
  5. 03 Lesson 3 · 3 prompts Monitoring And Alerting
  6. 04 Lesson 4 · 3 prompts Troubleshooting From Logs
  7. 05 Lesson 5 · 4 prompts Postmortems And Learning
  8. 06 Lesson 6 · 3 prompts Capacity And Cost Planning
  9. 07 Lesson 7 · 3 prompts Reliability Design Reviews
  10. 08 Lesson 8 · 3 prompts Advanced Incident Preparedness

About this course

7 topics

Prompts That Fit Site Reliability Work

This course is a set of prompts you can use the same day you learn them. Each lesson starts with a real reliability task and shows you how to hand the first draft to an AI assistant.

You will write incident updates, build runbooks, quiet noisy alerts, read logs, draft postmortems, plan capacity, review designs, and prepare for the incidents you have not had yet.

  1. The lessons
    1. Incident Communication Basics: Use AI to turn raw incident details into clear updates, handoffs, and stakeholder explanations.
    2. Runbooks And Automation: Use AI to convert operational knowledge into runbooks, script skeletons, and explanations of infrastructure changes.
    3. Monitoring And Alerting: Use AI to write better queries, cut alert noise, and design SLO dashboards that reflect user experience.
    4. Troubleshooting From Logs: Use AI to interpret logs, connect symptoms across services, and decide the next debugging step.
    5. Postmortems And Learning: Use AI to draft blameless postmortems, pull out action items, and turn incidents into team learning.
    6. Capacity And Cost Planning: Use AI to forecast capacity, review cloud spend, and plan rightsizing changes without overprovisioning.
    7. Reliability Design Reviews: Use AI to stress-test designs, plan chaos experiments, and prepare rollback steps before changes go live.
    8. Advanced Incident Preparedness: Use AI to build incident simulations, check runbook coverage, and define metrics that improve future response.
  2. What The Course Covers

    The eight lessons follow the shape of a reliability engineer's week. You start with the writing that happens during an incident, then move to the quieter work: runbooks, alert tuning, log reading, and postmortems.

    The later lessons look forward. Capacity and cost planning, design reviews, and incident simulations are the work that prevents the next outage rather than describing it.

  3. How The Lessons Connect

    Each lesson builds on the last. The incident update prompt teaches you to compress messy detail into clear sentences, and that same habit carries into postmortems and design reviews.

    The log reading lesson feeds the troubleshooting you do in a review. The runbook lesson feeds the preparedness lesson, where you check whether your runbooks actually cover the failures you might face.

  4. Using The Prompts Well

    Give the AI real context: the service name, what changed recently, who is affected, and what you already know. Short, specific inputs beat long, vague ones every time.

    Then treat the answer as a draft. Read it out loud, cut anything that sounds like a template, and add the detail only you would know. The prompts are a starting point, not a script.

  5. Who This Course Is For

    It is for site reliability engineers, platform engineers, and anyone on an on-call rotation who writes about systems. It suits people who are comfortable with the job and new to using AI in it.

    If you have tried an AI assistant and found the answers generic, this course is the fix. The prompts give it enough shape to be useful.

  6. Safety And Privacy At Work

    Check what your company allows before you paste anything into an AI tool. Customer data, credentials, internal hostnames, and security details usually should not go in.

    A safe habit is to describe the situation rather than copy it. Say what kind of service failed and how it behaved, and keep the identifying details out. When in doubt, ask your security team.

  7. Your Next Step

    Start with the incident communication lesson and use it on your next alert, even a small one. Notice how much editing the draft needs and how much time it saves.

    From there, work through the lessons in order or jump to the one that matches your week. The goal is a set of prompts you reach for without thinking.