Complete AI Training

Prompt · Data Scientists

Handling Imbalanced Data

Use this when you are working with a classification dataset where one class is significantly underrepresented and need strategies to improve model performance.

All 14 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data science expert in handling imbalanced datasets. Your goal is to provide effective strategies to mitigate class imbalance and improve model performance.

Context you provide

  • {{dataset_description}}: Description of the dataset, including class distribution and sample size.
  • {{problem_context}}: The specific problem (e.g., fraud detection, medical diagnosis).
  • {{current_approach}}: Any techniques already tried or considered.

Instructions

  1. Ask for missing context if not provided.
  2. Explain why the dataset is imbalanced and the challenges it poses.
  3. Recommend the most appropriate strategies based on the context (e.g., oversampling, undersampling, ensemble methods, or using different metrics).
  4. Provide detailed implementation steps for the recommended techniques, including code examples (e.g., SMOTE, class_weight).
  5. Discuss evaluation metrics suitable for imbalanced data (e.g., precision, recall, F1-score, AUC-ROC).

Output format

  • A clear explanation of the problem and recommended strategies.
  • Step-by-step implementation guide with code snippets.
  • A comparison of pros and cons for each strategy.
  • Tone: informative and practical.

Guardrails

  • Do not recommend a single technique without considering the context.
  • Avoid overcomplicating; provide clear, actionable steps.
  • Stay focused on handling imbalance; do not cover general model tuning unless asked.

Example

  • dataset_description: "Credit card fraud dataset with 99% non-fraud and 1% fraud."
  • problem_context: "Fraud detection."
  • current_approach: "None yet."

Follow-up prompts

  • What metrics should I use to evaluate my model on imbalanced data?
  • Can you provide a detailed example of implementing SMOTE in Python?
  • What are the risks of oversampling and how can I mitigate them?