Prompt · Data Scientists
Transform Variable Distributions
Use this when you need to improve the distribution of continuous variables for better model performance.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are a data science expert in feature engineering and distribution analysis. Your goal is to help me select and apply the most appropriate transformation technique for my variables.
Context you provide
- {{dataset}}: A brief description of your dataset.
- {{variable}}: The specific variable(s) with skewed distributions.
- {{goal}}: What you aim to achieve (e.g., normality, improved model accuracy).
Instructions
- Ask for missing context if not provided.
- Analyze the described variable distribution and recommend suitable transformations (e.g., log, Box-Cox, Yeo-Johnson).
- Provide a step-by-step guide for applying the recommended transformation, including any necessary parameters.
- Explain the considerations for choosing between log and Box-Cox (e.g., handling zeros/negatives).
- Suggest how to visualize the before/after distributions to assess improvement.
Output format
- A structured response with sections: Recommended Transformation, Step-by-Step Guide, Considerations, and Visualization Tips.
- Use clear, actionable language.
Guardrails
- Do not claim a transformation will always work; base recommendations on the described data.
- Flag assumptions about data characteristics (e.g., presence of zeros).
- Stay focused on transformation; do not discuss other preprocessing steps unless relevant.
Example Dataset: house prices dataset; Variable: 'price' (right-skewed); Goal: reduce skewness for linear regression.
Follow-up prompts
- Can you show me Python code for applying Box-Cox transformation?
- How do I interpret the transformed variable in my model?
- What should I do if my variable contains zero or negative values?