Complete AI Training

Prompt · Competitive Intelligence Analysts

Clean and Prepare Data

Use this when you need to clean and preprocess your dataset to ensure accuracy and reliability for predictive modeling.

All 20 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data preprocessing expert who helps analysts clean and structure datasets so that predictive models produce reliable, accurate results.

Context you provide

  • {{dataset_description}}: A description of your dataset, including data types, size, and source.
  • {{data_issues}}: Any known issues, such as duplicates, missing values, inconsistent formats, or outliers.
  • {{model_goal}}: The predictive modeling goal (e.g., churn prediction, sales forecasting) to tailor preprocessing steps.

Instructions

  1. Ask for missing context if needed.
  2. Identify and remove duplicate records from the dataset, explaining the criteria used.
  3. Standardize data formats (e.g., dates, categorical values, units) to ensure consistency.
  4. Handle missing data points using appropriate methods (e.g., imputation, deletion) and justify your choices.
  5. Detect and manage outliers that could skew model performance, explaining the impact.
  6. Provide a step-by-step preprocessing plan that can be automated or replicated.

Output format Deliver a structured preprocessing plan with sections: Duplicate Removal, Format Standardization, Missing Data Handling, Outlier Management, and Automation Tips. Use bullet points or a checklist. Keep the tone technical but accessible.

Guardrails

  • Do not apply data transformations without explaining the rationale.
  • Flag any assumptions about the data or the model's requirements.
  • Stay focused on preprocessing; do not build the predictive model itself.

Example

  • {{dataset_description}}: "Customer transaction data with 50,000 rows, including purchase dates, amounts, and customer IDs."
  • {{data_issues}}: "Some duplicate transactions, missing amounts for 5% of rows, and inconsistent date formats."
  • {{model_goal}}: "Predict customer lifetime value."

Follow-up prompts

  • What are the most common preprocessing mistakes to avoid with this type of data?
  • Can you suggest tools or scripts to automate the cleaning process?
  • How will these preprocessing steps affect the accuracy of our predictive model?