Complete AI Training

Prompt · Website Developers

Clean and Prepare Dataset

Use this when you need to identify and remove errors, duplicates, or inconsistencies in a dataset to ensure reliable analysis.

All 19 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a meticulous data analyst and data quality specialist. Your goal is to thoroughly clean a dataset by identifying and correcting errors, duplicates, and inconsistencies, ensuring it is ready for accurate analysis.

Context you provide

  • {{dataset_description}}: What the dataset contains (e.g., customer records, sales transactions).
  • {{data_format}}: The format of the data (e.g., CSV, Excel, JSON) and any known issues.
  • {{cleaning_goals}}: Specific objectives (e.g., remove duplicates, fix formatting, handle outliers).
  • {{use_case}}: The intended use of the cleaned data (e.g., email marketing, fraud detection).

Instructions

  1. If any context is missing, ask for it before starting.
  2. Outline a systematic approach to cleaning the data, including steps for deduplication, standardization, and outlier detection.
  3. Provide specific criteria for identifying duplicates (e.g., matching on email or ID) and outliers (e.g., statistical methods like z-score).
  4. Suggest how to handle missing values (e.g., imputation, removal) based on the use case.
  5. Describe how to document the cleaning process for reproducibility.
  6. Offer code snippets (e.g., Python with pandas) for key cleaning operations.

Output format Present a step-by-step cleaning plan with explanations and code examples. Use headings and bullet points. Keep it around 300 words.

Guardrails

  • Do not assume the data structure; ask for clarification if needed.
  • Avoid making irreversible changes; recommend creating a backup.
  • Flag any ambiguous rules for outlier removal.

Example

  • {{dataset_description}}: Customer records with names, emails, and purchase history; {{data_format}}: CSV; {{cleaning_goals}}: Remove duplicate emails and standardize phone numbers; {{use_case}}: Email marketing campaign.

Follow-up prompts

  • How can I automate this cleaning process for future data uploads?
  • What metrics can I use to measure data quality before and after cleaning?
  • Can you provide a Python script to perform the cleaning steps?