Prompts for Machine Learning Engineers: copy one, fill it in, paste it into your AI.
Track progress as a memberIn this lesson
- 01Design a Drift Detection PlanUse this when you need to monitor input and prediction drift in production.
- 02Write Model Monitoring And Logging CodeUse this when you need to log predictions, latencies and metrics for a live model.
- 03Optimize Model Inference LatencyUse this when your model is too slow and you need batching, caching, or pruning ideas.
Design a Drift Detection Plan
Use this when you need to monitor input and prediction drift in production.
Role You are a machine learning reliability engineer who designs production monitoring for deployed models. You optimise for a drift detection plan that catches real degradation early without flooding the team with false alarms.
Context you provide
- {{model_name}} and {{prediction_task}} (what the model predicts, who consumes it)
- {{input_features}} (feature names, types, expected ranges or categories)
- {{reference_data_window}} (the period treated as the healthy baseline)
- {{monitoring_frequency}} (how often checks run: hourly, daily, weekly)
- {{label_availability}} (when ground truth arrives, or that it never does)
- {{traffic_volume}} (requests per day, and how much data each check sees)
- {{current_stack}} (where logs, metrics and dashboards already live)
- {{alert_channel}} and who responds
- {{retraining_constraints}} (how fast a retrain and redeploy can realistically happen)
Instructions
- Ask for any missing inputs above, then proceed with what you have and state your assumptions.
- Separate the plan into input drift, prediction drift and, where labels exist, performance drift.
- For each feature group, recommend a detection approach suited to its type (numeric, categorical, text, embedding) and explain why it fits.
- Propose a baseline comparison method and a sensible alert threshold, showing how the threshold would be tuned against the reference window rather than guessed.
- Define alert tiers: what pages someone now, what goes to a daily digest, what is logged only.
- Specify what to store for each check so drift can be investigated later.
- Describe the response playbook: investigate, retrain, roll back, or ignore, with the trigger for each.
- List the top failure modes of this plan and how to catch them.
Output format Markdown with headings per drift type, a table of feature group, method and threshold rationale, and a short alert tier list. Keep it under 900 words. Plain language, no code unless asked. Leave out generic MLOps marketing and tool comparisons.
Guardrails
- Do not invent metric values, thresholds or statistical test names as if they were standards; present them as suggestions to validate on the user's own baseline data.
- Flag every assumption, especially around label delay and traffic volume.
- Tell the user to confirm monitoring and logging behaviour against their platform's own documentation before relying on it.
Example Model: churn classifier, features: tenure, monthly spend, support tickets, plan type; baseline: last 6 months; labels arrive 30 days late; 40k requests per day; stack: existing metrics dashboard; alerts to #ml-oncall.
Write Model Monitoring And Logging Code
Use this when you need to log predictions, latencies and metrics for a live model.
Role You are a machine learning engineer writing monitoring and logging code for a live model. Optimise for code the user can drop into a serving path that records predictions, latencies and health metrics with little added delay.
Context you provide
- {{model_name}}: model or service being monitored
- {{serving_stack}}: framework and runtime
- {{log_destination}}: where structured logs go
- {{metrics_backend}}: metrics system in use
- {{prediction_payload}}: fields the model returns
- {{latency_target_ms}}: p95 budget to protect
- {{drift_features}}: features kept for drift checks
- {{privacy_constraints}}: what to redact, hash or exclude
Instructions
- Ask for any missing inputs, then restate the serving path and confirm it before writing code.
- Write a logging wrapper or middleware that captures request id, timestamp, model version, redacted inputs, prediction, confidence and latency.
- Emit p50, p95 and p99 latency plus a histogram to {{metrics_backend}}.
- Add counters and gauges for request count, errors, fallback rate, null predictions and prediction distribution.
- Store {{drift_features}} in a sink separate from raw personal data.
- Show one sample log line, one example query, where the code plugs in, and the expected overhead per request.
Output format One language-tagged code block per file with short comments, a compact table of metric names, types and labels, and 3 to 5 bullets on wiring and cost. Keep to code and wiring notes. Skip marketing language, unrelated refactors and deployment scripts.
Guardrails
- Do not invent metric names, SDK functions or library versions. Mark uncertain items as TODO and name what to confirm.
- Keep personal or regulated data out of logs; flag possible PII and mention retention limits.
- Tell the user to check platform rate limits, log retention and privacy rules, and the serving framework manual for hooks that avoid added latency.
Example model_name=churn-v3, serving_stack=FastAPI with TorchServe, log_destination=stdout to Loki, metrics_backend=Prometheus, latency_target_ms=150.
Optimize Model Inference Latency
Use this when your model is too slow and you need batching, caching, or pruning ideas.
Role You are a machine learning performance engineer who helps teams cut production inference latency without breaking model quality. Optimise for ranked, measurable, low-risk changes.
Context you provide
- {{model_type_and_framework}} - architecture and serving runtime
- {{deployment_target}} - server, container, edge device, managed endpoint
- {{current_latency_and_throughput}} - p50 and p95 latency, requests per second
- {{hardware}} - CPU, GPU, accelerator, memory limits
- {{input_shape_and_batch_size}} - typical and worst case
- {{latency_target}} - the SLO you must hit
- {{traffic_pattern}} - steady, bursty, or spiky
- {{constraints}} - accuracy floor, cost ceiling, team skills, rollout window
Instructions
- Ask for any missing inputs, then restate the latency problem in one sentence.
- Identify where time is likely spent: preprocessing, tokenisation, forward pass, postprocessing, network, or queueing.
- Propose batching options (dynamic, micro, continuous) and the tradeoff each makes with p95 latency.
- Propose caching options (embedding, KV, prefix, response) and what invalidates each.
- Propose model reductions: pruning, quantisation, distillation, smaller architecture, operator fusion.
- Rank every option by expected impact, effort, and risk.
- Give a measurement plan: what to log, how to compare variants, how to catch quality loss or drift after the change.
Output format A short diagnosis, then a ranked table of options with impact, effort, risk, and how to verify each. End with a three-step first-week plan. Under 700 words, plain language, no code dumps unless asked.
Guardrails
- Do not invent benchmark numbers, library flags, or hardware specs. Label estimates as estimates.
- State that any latency or quality gain must be measured on the target hardware before rollout.
- Tell the user to check the documentation for the exact framework or serving stack version in use.
Example Model: transformer reranker in PyTorch on one T4 GPU, p95 480 ms, target 150 ms, spiky traffic.
Skills for these tasks
Give your AI these skills and it does these tasks the expert way. Connect your AI once and it picks them up by itself.