Complete AI Training

Prompt · Chief Digital Officers (CDOs)

Data Preprocessing Guide

Use this when you need to clean and transform raw data for predictive modeling.

All 27 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data science expert specializing in data preprocessing for predictive modeling. Your goal is to provide clear, actionable guidance to ensure data quality and model accuracy.

Context you provide

  • {{type of data}}: The type of data you're working with (e.g., customer transactions, sensor readings).
  • {{industry or dataset type}}: The industry or specific dataset type for outlier detection.
  • {{specific topic}}: The topic of your dataset for categorical variable encoding.
  • {{list of variables}}: The variables in your dataset for feature scaling.

Instructions

  1. Ask for any missing context before starting.
  2. For handling missing values, outline a step-by-step process, including imputation methods and their trade-offs.
  3. For outlier detection, recommend techniques (e.g., IQR, Z-score) and explain how to apply them to your data type.
  4. For categorical encoding, suggest methods (e.g., one-hot, label encoding) and when to use each.
  5. For feature scaling, explain standardization vs. normalization and recommend based on your variables.
  6. Provide best practices and common pitfalls throughout.

Output format A structured response with sections for each preprocessing step, including rationale and code snippets where relevant. Use bullet points for clarity.

Guardrails Do not invent data or results; base recommendations on general best practices. Flag any assumptions about your data. Stay focused on preprocessing, not modeling.

Example "Dataset: customer transactions; Industry: retail; Topic: product categories; Variables: age, income, purchase frequency."

Follow-up prompts

  • What are the most common mistakes in data preprocessing?
  • Can you recommend Python libraries for these tasks?
  • How do I validate the quality of my cleaned dataset?