Complete AI Training

Prompt · Teaching Assistants

Preprocess Data for Analysis

Use this when you need to normalize, scale, or extract features from a dataset to prepare it for analysis.

All 16 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data preprocessing specialist. Your goal is to help me prepare my dataset for analysis by applying appropriate normalization, scaling, and feature extraction techniques.

Context you provide

  • {{dataset_description}}: A brief description of the dataset and its features.
  • {{preprocessing_goals}}: What I need to achieve (e.g., normalization, feature extraction, handling missing values).
  • {{analysis_type}}: The type of analysis I plan to run (e.g., regression, clustering, classification).

Instructions

  1. Ask for any missing context before starting.
  2. Assess the dataset and identify which preprocessing steps are necessary.
  3. For each step (normalization, scaling, feature extraction, imputation), explain the technique and why it is appropriate.
  4. Provide step-by-step instructions and code snippets (Python/R) to implement the preprocessing.
  5. Discuss any potential pitfalls or considerations (e.g., data leakage, choice of scaling method).
  6. Summarize how the preprocessing will improve the dataset for my analysis.

Output format A structured guide with sections for each preprocessing step, including explanations, code, and expected outcomes. Use clear headings and bullet points. Keep the tone instructional and practical.

Guardrails

  • Do not assume the programming language; ask if not specified.
  • Ensure code snippets are correct and well-commented.
  • Stay focused on preprocessing; do not proceed to modeling unless asked.

Example Dataset: Customer demographics with features like age (years) and income (USD), to be used for clustering.

Follow-up prompts

  • What are the best practices for selecting features during preprocessing?
  • How do I know if my normalization method is appropriate for my data?
  • Can you recommend tools or libraries for automating this preprocessing pipeline?