Complete AI Training

All-in-one AI plugin for site reliability engineers

Site Reliability Engineer: the all-in-one AI plugin

17 skills, 25 prompts, 7 apps and 12 bot agents for incidents, alerts, service targets and runbooks, added to the AI you already use.

Works inChatGPTClaudeMicrosoft 365 CopilotGeminiGrokGrok BotCursorKimi
Site Reliability Engineer at work
Inside the boxv1.0.0
  • 17skills
  • 25prompts
  • 7apps
  • 12Grok bots

What's inside

Your week is alerts at odd hours, dashboards nobody agrees on, and an incident call where five people search for the same answer. Writing a postmortem takes an afternoon. The runbook is out of date. You know the fix. You need the facts, the timeline and the words, fast.

  • 17 reliability skills

    Skills for reading logs, spotting error patterns, setting service targets, building runbooks, tuning alerts and running blameless postmortems. Each one answers a real question from your week.

  • 25 ready-made prompts

    Prompts grouped in eight topics, from incident updates to capacity and cost planning. Open one, add your service name, and get a usable draft in seconds.

  • 7 working apps

    Apps for triaging an incident, coordinating root cause, checking runbook prerequisites, and building dashboards where every number traces back to its source.

  • 12 bot agents

    A Grok Bot team: incident responder, observability engineer, postmortem writer, deployment engineer, Kubernetes platform specialist, load tester and cloud cost optimizer, ready to work alongside you.

  • Works in eight AIs

    Use it in ChatGPT, Claude, Microsoft 365 Copilot, Gemini, Grok, Grok Bot, Cursor and Kimi. The same toolkit, inside the assistant you already trust.

  • Daily news and ideas

    Members get a live connection for daily AI news for this job, one new idea each day with the prompt to run it, and automatic updates to every skill.

A day with your plugin

This plugin drops a reliability toolkit into the AI you already open. Ask about a noisy alert, a slow query or a failing deploy, and it works from the skills, prompts and apps built for this job. You get a plan, a timeline and a checklist you can act on.

  1. Morning

    You open the alert queue. The system monitoring skill reads the overnight logs and metrics, points to the two alerts that matter and suggests what to check first.

  2. Late morning

    An incident call starts. You use the incident response coordinator to sort roles, draft the first stakeholder update and keep the timeline clean while people work.

  3. Afternoon

    Your service has no clear target. The reliability target skill helps you pick the right indicator and objective, and shows what the error budget means for this week's releases.

  4. End of day

    The postmortem is due. The postmortem writing app turns your notes into a blameless draft, with the timeline, the parts still uncertain, and action items with owners.

Install it in your AI

Choose the AI you use. No technical knowledge needed: it's a few clicks and some copy and paste.

ChatGPT

Your own GPT on chatgpt.com, plus a plugin for Codex and ChatGPT workspaces. About 5 minutes.

  1. Download the ChatGPT version and unzip it (double-click the file).
  2. Go to chatgpt.com, click GPTs in the left menu, then Create, then the Configure tab.
  3. Name it Site Reliability Engineer Assistant. Open custom-gpt/instructions.txt, copy everything and paste it into Instructions.
  4. Under Knowledge, upload all the files from the custom-gpt/knowledge folder.
  5. Paste the lines from conversation-starters.txt as conversation starters and click Create.
  6. For the daily news, ideas and skill updates: in ChatGPT go to Settings → Apps & Connectors, add the Complete AI Training connector and sign in with your membership. The README shows each click.

Download for members

Members download the plugin for every job, with daily news, ideas and skill updates. Or buy this plugin alone for $49, once.

Become a member

Or buy this plugin for $49 · Sign in

Everything in the plugin

Your AI picks the right part by itself. You can also ask for one by name.

Skills for site reliability engineers 17

  • system monitoring assistantDigs through logs, metrics, and alerts to spot anomalies and name next steps, like a memory spike on one node.
  • devops incident responderDiagnoses a live outage, drafts the postmortem, tunes the alert, and flags runbook gaps, like checkout returning 500s.
  • sre engineerSets SLIs, SLOs, and error budgets, cuts toil, and designs fault tolerance, like a 99.9 percent API target.
  • monitoring specialistBuilds metrics collection, alert rules, dashboards, log aggregation, tracing, and SLA reporting for a new service.
  • it operationsGives frameworks for incidents, change management, capacity planning, and alert tuning, like sizing next quarter's database capacity.
  • engineering runbookTurns topology, alerts, dashboards, procedures, and on-call schedules into one copyable page for the person holding the pager.
Show all 14 skills
  • incident response coordinatorRuns triage, stakeholder updates, resourcing, and the postmortem, like keeping a 30-minute outage's communications on schedule.
  • runbook creationWrites operational runbooks and SOPs with troubleshooting and recovery steps, like restarting a stuck message queue safely.
  • error detectiveSearches logs and code for error patterns and stack traces across systems, like one bad deploy showing up in three services.
  • agency observability designerDesigns metrics, logs, and traces with golden signals and SLOs, like fixing alerting that pages the team nightly.
  • agency engineering sreActs as a senior SRE voice on SLOs, error budgets, chaos engineering, and toil reduction for systems at scale.
  • incident managementSets up incident processes, escalation paths, on-call schedules, and post-incident reviews, like defining who gets paged at 3am.
  • observability monitoring slo implementBuilds the full SLO framework with meaningful SLIs and error budget monitoring, like balancing reliability against release speed.
  • slo implementationGives you the framework for SLIs, SLOs, and error budgets, like choosing the right window for a latency objective.

Plus a start-here assistant, the prompt library and the app builder, and for members the daily brief.

Ready-made prompts 25

Show all 8 topics
  • Advanced Incident Preparedness3 prompts

Apps you can build with AI 7

Your Grok Bot team 12

  • Incident ResponderAssess severity, stabilize systems, coordinate communication, and produce blameless post-incident reports.
  • Sre EngineerDefine SLOs, manage error budgets, and reduce toil for system reliability.
  • Observability EngineerDesigns and maintains production monitoring, logging, and tracing systems for reliability.
  • Postmortem WritingGuide blameless postmortems from incident data to action items.
  • Deployment EngineerDesigns and optimizes CI/CD pipelines for faster, safer deployments with automated rollbacks and monitoring, including GitOps and progressive delivery
  • Platform Sre KubernetesManages production Kubernetes deployments with safe rollouts, rollbacks, and security defaults.
  • Load Testing SpecialistDesigns and executes load tests to find system bottlenecks and capacity limits.
  • Aws Cost OptimizerAnalyze AWS spending and recommend cost savings using CLI and Cost Explorer.
  • Monitoring SpecialistMonitors infrastructure health, collects metrics, and alerts on symptoms to keep systems reliable.
  • Devops TroubleshooterDiagnoses production incidents using logs, metrics, and traces with systematic root cause analysis.
  • Incident Runbook TemplatesGenerate incident response runbooks with detection, triage, and mitigation steps.
  • Datadog AutomationAutomate Datadog monitoring, metrics, logs, monitors, dashboards, events, and downtimes via Rube MCP.

Get the plugin

Two ways to get it. Most people choose the membership: it includes every plugin and keeps them up to date.

Recommended

Membership

$9 a month, billed yearly

  • The plugin for site reliability engineers, and for 500 other jobs
  • Daily AI news and a new idea every day, inside your AI
  • Skills update by themselves
  • 900+ video courses and certificates
  • Download up to 2 plugins a day
Become a member

This plugin only

Standalone plugin

$49 once

  • The plugin for site reliability engineers, for all eight AIs
  • All 17 skills, 25 prompts and 7 apps
  • Yours to keep, no subscription
  • No daily news, ideas or skill updates

Secure payment by Stripe. Bought it already? Use the link in your email.

Questions

What is a plugin, and how do I install it if I am not technical?

A plugin is a set of instructions and files that teaches your AI assistant how to do this job. You download it once, follow the short install steps for your AI, then just ask questions in plain words.

Which AI assistants does it work with?

Eight: ChatGPT, Claude, Microsoft 365 Copilot, Gemini, Grok, Grok Bot, Cursor and Kimi. Install it in the one you already use at work. Nothing changes about how you open or chat with that assistant.

What is the difference between the member download and the $49 one-time plugin?

The $49 one-time plugin is the standalone download with all the skills, prompts and apps. Members also get the live connection to daily AI news for this job, a new idea every day with its prompt, and automatic updates when skills change.

Is my data safe?

The plugin is instructions and files. It does not send your data anywhere on its own. Your chats stay with the AI provider you chose, under that provider's privacy terms and your company rules.

How many downloads can I get?

You can download two plugins a day. That covers this one plus another job or team plugin, so you can set up a colleague or cover two roles from the same account.

Learn to use AI as a site reliability engineer

The plugin does the work; the learning path shows you how to get the most out of it, step by step.

Open the learning path

Plugins for related jobs