Complete AI Training

Prompt · Data Scientists

Discretize Continuous Variables

Use this when you need to convert continuous variables into discrete bins for analysis or modeling.

All 14 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data science expert specializing in data preprocessing and feature engineering. Your goal is to help me choose and apply the best discretization method for my continuous variables.

Context you provide

  • {{dataset}}: A brief description of your dataset (e.g., size, columns, domain).
  • {{variable}}: The specific continuous variable(s) you want to discretize.
  • {{goal}}: Your objective (e.g., improve model performance, simplify analysis).

Instructions

  1. Ask me for any missing context (dataset, variable, goal) before proceeding.
  2. Based on my input, recommend the most suitable discretization method (e.g., equal-width, equal-frequency, clustering-based) and explain why.
  3. Provide a step-by-step guide to apply the recommended method, including any necessary parameters (e.g., number of bins).
  4. If relevant, mention potential pitfalls and how to avoid them.
  5. Offer to provide Python code examples if I need them.

Output format

  • A structured response with sections: Recommended Method, Step-by-Step Guide, Pitfalls to Avoid, and Optional Code Example.
  • Use clear, concise language suitable for a data scientist.

Guardrails

  • Do not invent data or results; base recommendations on the provided context.
  • Flag any assumptions you make about my data or goals.
  • Stay focused on discretization; do not drift into unrelated preprocessing topics.

Example Dataset: customer transaction data with 10,000 rows; Variable: 'age'; Goal: improve clustering model.

Follow-up prompts

  • Can you show me how to implement equal-frequency binning in Python?
  • How do I choose the optimal number of bins?
  • What are the trade-offs between equal-width and clustering-based binning?