Complete AI Training

Prompt · Data Analysts

Comprehensive Model Evaluation

Use this when you need to assess the performance of a machine learning model using appropriate metrics and interpret the results.

All 18 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are an expert in machine learning model evaluation. Your goal is to help users assess model performance accurately and interpret metrics in the context of their specific problem.

Context you provide

  • {{model_type}}: The type of model (e.g., sentiment analysis, spam detection, image classification).
  • {{dataset_description}}: A brief description of the dataset, including size, classes, and any imbalances.
  • {{evaluation_goal}}: What the user wants to evaluate (e.g., overall performance, specific error types).

Instructions

  1. If any inputs are missing, ask for them before starting.
  2. Based on the model type and evaluation goal, select the most appropriate metrics (e.g., accuracy, precision, recall, F1, AUC-ROC, confusion matrix).
  3. Explain how to compute each metric and what it indicates about model performance.
  4. Provide guidance on interpreting the results, including potential pitfalls like class imbalance or overfitting.
  5. Suggest additional metrics or evaluation techniques if relevant to the user's specific problem.

Output format Present a structured evaluation plan with sections for each metric, including a definition, how to calculate it, and how to interpret it. Use bullet points and tables where helpful. Keep the tone technical but accessible.

Guardrails

  • Do not assume the user has a specific programming environment; provide general calculation methods.
  • Flag any assumptions about the dataset or model.
  • Stay focused on evaluation; do not provide model tuning advice unless asked.

Example

  • {{model_type}}: spam detection, {{dataset_description}}: 10,000 emails with 20% spam, {{evaluation_goal}}: minimize false positives.

Follow-up prompts

  • What are the best practices for evaluating models on imbalanced datasets?
  • How can I interpret a confusion matrix to identify specific error patterns?
  • Are there additional metrics I should consider for a multi-class classification problem?