Course overview
Start hereAI for Site Reliability Engineers
START HERE

Priya's Thursday, two ways

8 lessons · 25 prompts

A day in the life of a Site Reliability Engineer: what changes with these prompts.

Track progress as a member

Priya, a site reliability engineer at a payments company

Priya's Thursday starts with an alert on the checkout service. Response times are climbing and a handful of payments have failed. She pastes the alert details and a few log lines into Claude and asks for a plain-language summary and three likely causes, ranked. The answer is not perfect, but it points her at a connection pool that filled up overnight, and she confirms it in ten minutes.

Mid-morning she opens ChatGPT and works on the runbook she has been putting off for a database failover. She feeds it her rough notes from the last drill and asks for a step-by-step version a new hire could follow, with the rollback steps spelled out. She edits it down, adds the two warnings only she knows about, and shares it with the team.

After lunch there is a design review for a new caching layer. She asks Gemini to stress-test the plan: what happens if the cache is cold, what a partial network problem looks like, and how a rollback would go. The questions it raises become the first half of her review notes.

By four o'clock she has drafted the postmortem for last week's incident, blameless and with clear action items. The time she won back goes to walking a new hire through the failover runbook, and she leaves on time for once.

Before

  • Alerts fire and nobody knows who owns them
  • Incident updates written at midnight, half finished
  • Runbooks live in one person's head
  • Postmortems slip to next quarter

After this course

  • Updates drafted in minutes, reviewed in five
  • Runbooks anyone on the team can follow
  • Postmortems written the week they happen
  • Capacity plans ready before the budget talk

What you'll learn

  • Clear Incident Updates: Turn raw incident details into updates, handoffs, and stakeholder explanations that people can follow under pressure.
  • Runbooks And Scripts: Convert what you know into runbooks, script skeletons, and plain explanations of infrastructure changes.
  • Quieter Alerts: Write better queries, cut alert noise, and shape dashboards around what users actually feel.
  • Reading Logs Faster: Interpret logs, connect symptoms across services, and choose the next debugging step with more confidence.
  • Blameless Postmortems: Draft postmortems, pull out action items, and turn each incident into something the team learns from.
  • Capacity And Cost: Forecast capacity, review cloud spend, and plan rightsizing changes without overprovisioning.
  • Design Reviews And Rollbacks: Stress-test designs, plan chaos experiments, and prepare rollback steps before changes go live.
  • Incident Preparedness: Build incident simulations, check runbook coverage, and define metrics that improve the next response.

How this course works

  1. 8 lessonsOne task of your job each, from incident communication basics to advanced incident preparedness.
  2. Ready-to-paste promptsCopy, fill in the parts in {{brackets}}, paste into ChatGPT, Claude or Gemini.
  3. Tick and completeTick the prompts you tried and mark each lesson complete.
  4. Get certifiedFinish and keep the prompts as your own library.