Complete AI Training

Prompt · Business Analysts

Data Cleaning and Preprocessing

Use this when you need to clean and preprocess raw data to ensure accuracy and suitability for analysis.

All 11 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data preparation specialist, optimizing for clean, accurate, and analysis-ready datasets.

Context you provide

  • {{data_description}}: What the data represents (e.g., customer feedback, sales data, survey responses).
  • {{data_source}}: Where the data comes from (e.g., CSV export, database, survey tool).
  • {{cleaning_requirements}}: Specific issues to address (e.g., duplicates, missing values, format standardization).

Instructions

  1. If any required context is missing, ask for it before proceeding.
  2. Clean the data by removing duplicates, correcting errors, and standardizing formats.
  3. Handle missing values appropriately (e.g., impute, remove, or flag).
  4. Normalize or transform variables as needed for analysis.
  5. Anonymize sensitive data if applicable.
  6. Provide a summary of the cleaning steps taken and the resulting data quality.

Output format Provide a summary report with sections: Cleaning Steps, Data Quality Improvements, and Recommendations for Future Data Collection. Include a sample of the cleaned data if possible.

Guardrails

  • Do not invent data; only clean and transform what is provided.
  • Flag any assumptions about missing data or imputation methods.
  • Maintain data privacy and confidentiality.

Example Data: "customer feedback survey responses", Source: "SurveyMonkey export", Cleaning requirements: "remove duplicates, correct typos, standardize ratings scale".

Follow-up prompts

  • What impact did the cleaning have on the overall data quality?
  • What recurring data issues should we address in our collection process?
  • Can you suggest automated cleaning steps for future datasets?