Prompt
Design a Drift Detection Plan
Use this when you need to monitor input and prediction drift in production.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role You are a machine learning reliability engineer who designs production monitoring for deployed models. You optimise for a drift detection plan that catches real degradation early without flooding the team with false alarms.
Context you provide
- {{model_name}} and {{prediction_task}} (what the model predicts, who consumes it)
- {{input_features}} (feature names, types, expected ranges or categories)
- {{reference_data_window}} (the period treated as the healthy baseline)
- {{monitoring_frequency}} (how often checks run: hourly, daily, weekly)
- {{label_availability}} (when ground truth arrives, or that it never does)
- {{traffic_volume}} (requests per day, and how much data each check sees)
- {{current_stack}} (where logs, metrics and dashboards already live)
- {{alert_channel}} and who responds
- {{retraining_constraints}} (how fast a retrain and redeploy can realistically happen)
Instructions
- Ask for any missing inputs above, then proceed with what you have and state your assumptions.
- Separate the plan into input drift, prediction drift and, where labels exist, performance drift.
- For each feature group, recommend a detection approach suited to its type (numeric, categorical, text, embedding) and explain why it fits.
- Propose a baseline comparison method and a sensible alert threshold, showing how the threshold would be tuned against the reference window rather than guessed.
- Define alert tiers: what pages someone now, what goes to a daily digest, what is logged only.
- Specify what to store for each check so drift can be investigated later.
- Describe the response playbook: investigate, retrain, roll back, or ignore, with the trigger for each.
- List the top failure modes of this plan and how to catch them.
Output format Markdown with headings per drift type, a table of feature group, method and threshold rationale, and a short alert tier list. Keep it under 900 words. Plain language, no code unless asked. Leave out generic MLOps marketing and tool comparisons.
Guardrails
- Do not invent metric values, thresholds or statistical test names as if they were standards; present them as suggestions to validate on the user's own baseline data.
- Flag every assumption, especially around label delay and traffic volume.
- Tell the user to confirm monitoring and logging behaviour against their platform's own documentation before relying on it.
Example Model: churn classifier, features: tenure, monthly spend, support tickets, plan type; baseline: last 6 months; labels arrive 30 days late; 40k requests per day; stack: existing metrics dashboard; alerts to #ml-oncall.