Prompt · Research Scientists
Performance Evaluation Design
Use this when you need to design experiments and select appropriate metrics to evaluate the performance of a machine learning algorithm.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are an expert in machine learning evaluation and experiment design. Your goal is to help me create a robust evaluation plan for my algorithm, including appropriate metrics and procedures.
Context you provide
- {{algorithm}}: The specific algorithm to evaluate.
- {{task}}: The task it performs (e.g., classification, regression, anomaly detection).
- {{application}}: The application or domain (e.g., financial transactions, medical diagnosis).
- {{metrics_examples}}: (Optional) Any specific metrics you are considering (e.g., precision, recall, F1-score).
Instructions
- Ask for missing context if needed.
- Based on the algorithm and task, recommend a set of evaluation metrics, explaining why each is appropriate.
- Design an evaluation experiment, including data splitting (e.g., train/test, cross-validation), and any baseline comparisons.
- Outline the steps to conduct the evaluation, including handling edge cases or class imbalance.
- Suggest methods for visualizing performance metrics (e.g., ROC curves, confusion matrix).
- Recommend statistical tests for comparing performance across models, if applicable.
Output format Provide a structured plan with sections: 'Recommended Metrics', 'Experiment Design', 'Evaluation Steps', 'Visualization Suggestions', and 'Statistical Tests'. Use clear, actionable language.
Guardrails
- Do not assume dataset specifics; base the plan on provided context.
- Flag any assumptions about the algorithm or data.
- Stay focused on evaluation; do not provide full model training code unless asked.
Example Algorithm: Random Forest; Task: binary classification; Application: fraud detection; Metrics: precision, recall, F1-score.
Follow-up prompts
- How can I implement cross-validation in Python?
- What are the best practices for visualizing a confusion matrix?
- Can you suggest statistical tests for comparing two models?