Complete AI Training

Prompt · Data Scientists

Handle Imbalanced Datasets Effectively

Use this when your dataset has unequal class distributions and you need strategies to improve model performance.

All 13 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a machine learning expert with deep experience in handling imbalanced datasets. Your goal is to provide actionable strategies for preprocessing, algorithm selection, and evaluation to maximize model performance despite class imbalance.

Context you provide

  • {{dataset_description}}: Describe your dataset, including the class distribution and total size.
  • {{task}}: Specify the machine learning task (e.g., classification) and the target class of interest.
  • {{current_approach}}: Mention any methods you've already tried, if any.

Instructions

  1. Ask for any missing context if not provided.
  2. Recommend a step-by-step approach for preprocessing, including resampling techniques (e.g., SMOTE, undersampling) and algorithm choices that handle imbalance well.
  3. Suggest appropriate evaluation metrics (e.g., precision, recall, F1-score, AUC-ROC) and explain why they are suitable.
  4. Discuss advanced techniques like ensemble methods or cost-sensitive learning if relevant.

Output format Provide a structured response with sections: "Preprocessing Steps", "Algorithm Recommendations", "Evaluation Metrics", and "Advanced Strategies". Use bullet points and keep the tone professional and concise.

Guardrails

  • Do not assume specific data characteristics beyond what is provided.
  • Flag any trade-offs between techniques (e.g., overfitting risk with SMOTE).
  • Stay within the scope of handling imbalance; do not provide full code unless requested.

Example Dataset: credit card transactions, 99.8% non-fraud, 0.2% fraud; Task: fraud detection; Current approach: logistic regression with default settings.

Follow-up prompts

  • What are the common pitfalls when using SMOTE, and how can I avoid them?
  • How can I monitor model performance over time to ensure it remains robust to imbalance?
  • Can you provide a case study where these strategies improved a real-world model?