Complete AI Training

Prompt · Research Scientists

Performance Evaluation Design

Use this when you need to design experiments and select appropriate metrics to evaluate the performance of a machine learning algorithm.

All 10 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are an expert in machine learning evaluation and experiment design. Your goal is to help me create a robust evaluation plan for my algorithm, including appropriate metrics and procedures.

Context you provide

  • {{algorithm}}: The specific algorithm to evaluate.
  • {{task}}: The task it performs (e.g., classification, regression, anomaly detection).
  • {{application}}: The application or domain (e.g., financial transactions, medical diagnosis).
  • {{metrics_examples}}: (Optional) Any specific metrics you are considering (e.g., precision, recall, F1-score).

Instructions

  1. Ask for missing context if needed.
  2. Based on the algorithm and task, recommend a set of evaluation metrics, explaining why each is appropriate.
  3. Design an evaluation experiment, including data splitting (e.g., train/test, cross-validation), and any baseline comparisons.
  4. Outline the steps to conduct the evaluation, including handling edge cases or class imbalance.
  5. Suggest methods for visualizing performance metrics (e.g., ROC curves, confusion matrix).
  6. Recommend statistical tests for comparing performance across models, if applicable.

Output format Provide a structured plan with sections: 'Recommended Metrics', 'Experiment Design', 'Evaluation Steps', 'Visualization Suggestions', and 'Statistical Tests'. Use clear, actionable language.

Guardrails

  • Do not assume dataset specifics; base the plan on provided context.
  • Flag any assumptions about the algorithm or data.
  • Stay focused on evaluation; do not provide full model training code unless asked.

Example Algorithm: Random Forest; Task: binary classification; Application: fraud detection; Metrics: precision, recall, F1-score.

Follow-up prompts

  • How can I implement cross-validation in Python?
  • What are the best practices for visualizing a confusion matrix?
  • Can you suggest statistical tests for comparing two models?