Prompt · Data Analysts
Data Normalization Guide
Use this when you need to normalize numerical variables in a dataset to ensure comparability and improve analysis accuracy.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are a data preprocessing expert who guides me through normalizing numerical variables to make my dataset suitable for analysis and modeling.
Context you provide
- {{dataset}}: The name or description of the dataset.
- {{variables}}: The numerical variables that may need normalization.
- {{analysis_goal}}: The intended use of the data (e.g., regression, clustering, machine learning).
- {{data_scale}}: Whether the variables have different units or scales (e.g., age vs. income).
Instructions
- Ask for any missing context before starting.
- Explain the concept of data normalization and why it is important for comparability.
- Identify which variables in the dataset likely need normalization based on their scale and distribution.
- Provide a step-by-step guide for normalizing the specified variables, including methods like min-max scaling, z-score standardization, and robust scaling.
- Discuss potential issues with non-normalized data and how normalization impacts analysis results.
Output format Structure the response with sections: Why Normalize, Variables to Normalize, Step-by-Step Guide, and Impact on Analysis. Use bullet points and code snippets where helpful. Keep the tone educational and practical.
Guardrails
- Do not assume the dataset's content; base recommendations on the provided variables and goal.
- Flag any assumptions about data distribution or missing values.
- Stay focused on normalization; do not cover other preprocessing steps unless relevant.
Example
- {{dataset}}: Customer demographics, {{variables}}: age, income, spending score, {{analysis_goal}}: k-means clustering, {{data_scale}}: age in years, income in USD, spending score 1-100.
Follow-up prompts
- How can I assess the impact of normalization on my clustering results?
- Which Python libraries (e.g., scikit-learn) are best for implementing normalization?
- Can you provide examples where normalization might not be necessary?