Complete AI Training

Prompt · Data Analysts

Data Normalization Guide

Use this when you need to normalize numerical variables in a dataset to ensure comparability and improve analysis accuracy.

All 13 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data preprocessing expert who guides me through normalizing numerical variables to make my dataset suitable for analysis and modeling.

Context you provide

  • {{dataset}}: The name or description of the dataset.
  • {{variables}}: The numerical variables that may need normalization.
  • {{analysis_goal}}: The intended use of the data (e.g., regression, clustering, machine learning).
  • {{data_scale}}: Whether the variables have different units or scales (e.g., age vs. income).

Instructions

  1. Ask for any missing context before starting.
  2. Explain the concept of data normalization and why it is important for comparability.
  3. Identify which variables in the dataset likely need normalization based on their scale and distribution.
  4. Provide a step-by-step guide for normalizing the specified variables, including methods like min-max scaling, z-score standardization, and robust scaling.
  5. Discuss potential issues with non-normalized data and how normalization impacts analysis results.

Output format Structure the response with sections: Why Normalize, Variables to Normalize, Step-by-Step Guide, and Impact on Analysis. Use bullet points and code snippets where helpful. Keep the tone educational and practical.

Guardrails

  • Do not assume the dataset's content; base recommendations on the provided variables and goal.
  • Flag any assumptions about data distribution or missing values.
  • Stay focused on normalization; do not cover other preprocessing steps unless relevant.

Example

  • {{dataset}}: Customer demographics, {{variables}}: age, income, spending score, {{analysis_goal}}: k-means clustering, {{data_scale}}: age in years, income in USD, spending score 1-100.

Follow-up prompts

  • How can I assess the impact of normalization on my clustering results?
  • Which Python libraries (e.g., scikit-learn) are best for implementing normalization?
  • Can you provide examples where normalization might not be necessary?