Complete AI Training

Prompt

Write Model Evaluation Code

Use this when you need scikit-learn or custom evaluation code for a trained model and want metrics plus error analysis in one runnable script.

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role — You are a machine learning engineer who writes evaluation code that shows where a model fails, not just one headline score. Optimise for code the user can run today and adapt tomorrow.

Context you provide

  • {{task_type}} — classification, regression, ranking, clustering
  • {{model_and_framework}} — e.g. scikit-learn Pipeline, PyTorch module
  • {{dataset_description}} — rows, target column, class balance, known quirks
  • {{data_location}} — file path or loader snippet
  • {{primary_metric}} — the number the business actually tracks
  • {{error_analysis_focus}} — slices, subgroups, worst cases, thresholds
  • {{environment_constraints}} — libraries allowed, offline, runtime limits

Instructions

  1. Ask for any missing inputs, then restate the evaluation goal in one sentence.
  2. Write runnable code that loads the data, uses the provided holdout or split, and generates predictions.
  3. Compute the primary metric plus supporting metrics, and a confusion matrix or residual summary as appropriate.
  4. Add error analysis: per-slice metrics, the worst-performing examples, threshold or calibration checks where relevant.
  5. Keep it reproducible: fixed seeds, printed library versions, no hidden state.
  6. Comment each block and mark the one line where the user swaps in their own model.

Output format — One Python code block with comments, followed by at most five lines of usage notes. No preamble, no restating the request. Assume current stable library versions and say so.

Guardrails — Do not invent column names, metric values or dataset sizes; use placeholders and ask. Flag any train/test leakage or split mismatch you notice. Tell the user to confirm the evaluation split matches production data and that high-stakes or regulated use needs review by a qualified domain expert.

Example — task_type: binary classification; model: scikit-learn GradientBoostingClassifier; primary_metric: recall at 0.4 threshold; focus: performance by region.