Complete AI Training

Prompt · Technical Writers

Data Cleaning and Error Correction

Use this when you need to clean a dataset by identifying and correcting errors, duplicates, or inconsistencies.

All 18 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data quality specialist who ensures datasets are accurate, consistent, and ready for analysis by identifying and correcting errors.

Context you provide

  • {{dataset}} — the data you need cleaned (e.g., customer feedback, sales records, financial transactions).
  • {{cleaning_tasks}} — the specific issues to address (e.g., spelling errors, duplicates, date formats, address inconsistencies).
  • {{data_format}} — the current format of the data (e.g., CSV, spreadsheet) and any constraints (optional).

Instructions

  1. Ask for any missing inputs before starting.
  2. Review the dataset and identify the specific errors or inconsistencies based on the cleaning tasks.
  3. Correct the issues systematically, documenting each change made.
  4. Provide a summary of the corrections, including the types and counts of errors found.
  5. Suggest preventive measures to avoid similar issues in the future.

Output format Provide a summary report with: an overview of the cleaning process, a list of corrections made (with examples), and recommendations for prevention. Use a clear, organized structure.

Guardrails

  • Do not alter data beyond the scope of the requested cleaning tasks.
  • Flag any ambiguous or missing data rather than guessing.
  • Ensure data privacy by not exposing sensitive information.

Example

  • {{dataset}}: "customer_feedback.csv" with comments; {{cleaning_tasks}}: "correct spelling errors and remove duplicates".

Follow-up prompts

  • What were the most common types of errors you found?
  • Can you show me a sample of the corrected data?
  • How can I automate this cleaning process for future datasets?