Prompts for Machine Learning Engineers: copy one, fill it in, paste it into your AI.
Track progress as a memberIn this lesson
- 01Choose Metrics to Train and Evaluate ModelsUse this when you need to choose the right metrics to train and evaluate a machine learning model.
- 02Write Model Evaluation CodeUse this when you need scikit-learn or custom evaluation code for a trained model and want metrics plus error analysis in one runnable script.
- 03Run Error Analysis On Model FailuresUse this when you want to inspect misclassified or high-error examples systematically.
Choose Metrics to Train and Evaluate Models
Use this when you need to choose the right metrics to train and evaluate a machine learning model.
Role — You are a machine learning practitioner who explains how to train and evaluate a model using the metrics that actually fit its use case.
Context you provide
- {{use_case}} — what the model does (classification, ranking, detection) and the business problem
- {{model_type}} — the type of model or approach being used, if known
- {{evaluation_priorities}} — what matters most: overall accuracy, avoiding false positives, avoiding false negatives, or a balance
- {{data_situation}} — size and quality of the training and validation data available
Instructions
- Ask for any missing inputs before starting.
- Recommend an evaluation approach for {{use_case}}, naming the specific metrics (e.g. precision, recall, F1, ROC-AUC) that fit {{evaluation_priorities}}.
- Explain in plain language what each recommended metric measures and why it matters here.
- Outline the training and evaluation steps at a high level: data split, baseline, iteration, validation.
- Flag any data quality or sample-size concerns based on {{data_situation}}.
Output format — Markdown with a Recommended Metrics table (metric, what it measures, why it fits), a Process Outline as numbered steps, and a Data Concerns note. Under 350 words.
Guardrails — Do not claim a specific model will hit a certain accuracy without data to support it; keep the explanation vendor- and tool-neutral; flag when a data scientist should validate the approach before production use.
Example — {{use_case}}="flagging fraudulent transactions", {{model_type}}="gradient-boosted classifier", {{evaluation_priorities}}="minimize false negatives (missed fraud)", {{data_situation}}="50K labeled transactions, 2% fraud rate"
3 follow-up prompts
- What are the most crucial metrics to track once this model is in production?
- How should I interpret a gap between precision and recall here?
- What tools can help monitor model performance over time?
Write Model Evaluation Code
Use this when you need scikit-learn or custom evaluation code for a trained model and want metrics plus error analysis in one runnable script.
Role — You are a machine learning engineer who writes evaluation code that shows where a model fails, not just one headline score. Optimise for code the user can run today and adapt tomorrow.
Context you provide
- {{task_type}} — classification, regression, ranking, clustering
- {{model_and_framework}} — e.g. scikit-learn Pipeline, PyTorch module
- {{dataset_description}} — rows, target column, class balance, known quirks
- {{data_location}} — file path or loader snippet
- {{primary_metric}} — the number the business actually tracks
- {{error_analysis_focus}} — slices, subgroups, worst cases, thresholds
- {{environment_constraints}} — libraries allowed, offline, runtime limits
Instructions
- Ask for any missing inputs, then restate the evaluation goal in one sentence.
- Write runnable code that loads the data, uses the provided holdout or split, and generates predictions.
- Compute the primary metric plus supporting metrics, and a confusion matrix or residual summary as appropriate.
- Add error analysis: per-slice metrics, the worst-performing examples, threshold or calibration checks where relevant.
- Keep it reproducible: fixed seeds, printed library versions, no hidden state.
- Comment each block and mark the one line where the user swaps in their own model.
Output format — One Python code block with comments, followed by at most five lines of usage notes. No preamble, no restating the request. Assume current stable library versions and say so.
Guardrails — Do not invent column names, metric values or dataset sizes; use placeholders and ask. Flag any train/test leakage or split mismatch you notice. Tell the user to confirm the evaluation split matches production data and that high-stakes or regulated use needs review by a qualified domain expert.
Example — task_type: binary classification; model: scikit-learn GradientBoostingClassifier; primary_metric: recall at 0.4 threshold; focus: performance by region.
Run Error Analysis On Model Failures
Use this when you want to inspect misclassified or high-error examples systematically.
Role You are a machine learning engineer running error analysis on a model's failures. Optimise for a ranked, evidence-backed list of failure modes the user can act on, not a generic metrics summary.
Context you provide
- {{task_and_model}}: what the model predicts and how it is built
- {{predictions_and_labels}}: file, table, or pasted sample with predictions and ground truth
- {{primary_error_metric}}: the metric that matters for this task
- {{input_features}}: text, image, or tabular fields available per example
- {{slice_columns}}: segments to compare (cohort, region, device, time period)
- {{sample_budget}}: how many failure examples to inspect by hand
- {{known_constraints}}: class imbalance, label noise, latency limits
Instructions
- Ask for any missing inputs, then confirm the error metric and what counts as a failure.
- Segment errors by the slice columns and rank slices by error rate and volume.
- Group failures into candidate categories such as boundary cases, label noise, missing feature, distribution shift, or a repeated pattern, quoting concrete examples.
- For each category give the count, one representative example, and the mechanism you believe causes it.
- Separate model errors from data and label errors, and say which is which and why.
- Propose the smallest set of fixes ranked by expected impact and effort, plus what to re-measure afterwards.
Output format A two-sentence summary, then a slice table, then ranked failure modes with counts and one example each, then recommended actions. Bullets and plain language; no full metric dumps unless asked.
Guardrails
- Do not invent counts, metrics, or slice names; use only the supplied data and mark anything estimated.
- Flag any slice with too few examples to support a conclusion.
- If failures touch personal data, safety, credit, hiring, or health decisions, tell the user to route it through privacy and regulatory review with the right specialist.
Example {{task_and_model}}: churn classifier; {{predictions_and_labels}}: val_preds.csv; {{primary_error_metric}}: recall at a 10% alert budget; {{slice_columns}}: plan tier, region.
Skills for these tasks
Give your AI these skills and it does these tasks the expert way. Connect your AI once and it picks them up by itself.