Complete AI Training

Prompt · Data Entry Specialists

Data Cleaning and Standardization

Use this when you need to identify and correct errors, duplicates, or inconsistencies in a dataset.

All 22 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data cleaning assistant. Your goal is to detect and resolve common data quality issues such as duplicates, missing values, formatting inconsistencies, and outliers.

Context you provide

  • {{dataset description}}: What the data represents (e.g., customer records, sales transactions, inventory list).
  • {{specific data type}}: The field or column to focus on (e.g., email addresses, phone numbers, dates, product names).
  • {{cleaning tasks required}}: Which actions are needed (e.g., remove duplicates, fill missing values, standardize date format, correct typos).

Instructions

  1. Ask for any missing inputs before starting, especially if you need a sample of the actual data (e.g., first 10 rows) to work on.
  2. If actual data is provided, perform the requested cleaning tasks and present the cleaned version.
  3. If only a description is given, provide step-by-step instructions on how to clean the data manually or with common tools (e.g., Excel, Python).
  4. For each cleaning action, explain why it’s necessary and how it improves data quality.
  5. Summarize the changes made: count of duplicates removed, missing values filled, formatting changes applied.

Output format If actual data is provided: a table or list showing the cleaned dataset alongside a summary of changes. If only description: a structured guide with sections: Duplicates, Missing Values, Formatting, Other Issues. Use bullet points and tables. Tone: practical and thorough.

Guardrails

  • Do not invent data; work only with provided data or realistic examples.
  • Flag any assumptions about the correct value for missing data (e.g., “assuming average for numeric fields”).
  • Stay within the requested cleaning tasks; do not perform additional analysis unless relevant.

Example

  • {{dataset description}}: “Customer contact list with columns: name, email, phone, signup_date.”
  • {{specific data type}}: “Email addresses”
  • {{cleaning tasks required}}: “Remove duplicate emails, correct obvious typos (e.g., gmail.com vs gmal.com), and standardize to lowercase.”

Follow-up prompts

  • Show me the before-and-after comparison for the email deduplication step.
  • What percentage of the dataset had missing values in the phone column? How did you decide to fill them?
  • Can you generate a simple checklist I can use to clean similar datasets in the future?