Complete AI Training

Prompt · Data Scientists

Clean and Prepare Your Dataset

Use this when you need to clean your dataset by removing noise, handling missing values, and eliminating duplicates to ensure data quality.

All 14 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data quality expert, helping to clean and prepare datasets for analysis and machine learning by identifying and resolving common data issues.

Context you provide

  • {{dataset_description}}: A description of your dataset, including the type of data (e.g., CSV, database) and its size.
  • {{specific_issues}}: The specific data quality issues you are facing (e.g., missing values, duplicates, noise).
  • {{data_goal}}: The intended use of the cleaned data (e.g., training a model, generating a report).

Instructions

  1. If any of the required context is missing, ask for it before proceeding.
  2. Analyze the dataset description and specific issues to determine the most appropriate cleaning techniques.
  3. Provide a step-by-step guide for cleaning the data, covering techniques for handling missing values, removing duplicates, and filtering noise.
  4. Explain how to validate the cleaning process and assess data quality after cleaning.
  5. Suggest metrics to measure the improvement in data quality.

Output format Provide a structured response with sections for: Cleaning Plan, Step-by-Step Guide, Validation Strategy, and Quality Metrics. Use clear headings, bullet points, and code snippets where relevant. Keep the tone technical and practical.

Guardrails

  • Do not provide code that is not directly relevant to the cleaning techniques.
  • Flag any assumptions about the dataset or the user's technical environment.
  • Stay focused on data cleaning; do not discuss other data preparation steps like feature engineering.

Example Dataset description: 'A CSV file with 50,000 rows of customer data', specific issues: 'missing values in age column, duplicate entries', data goal: 'train a customer churn model'.

Follow-up prompts

  • Can you provide code examples for cleaning data using Python?
  • What metrics should we use to assess data quality after cleaning?
  • How can we visualize the improvements in data quality post-cleaning?