Complete AI Training

Prompt · Data Analysts

Standardize Inconsistent Data Values

Use this when you need to diagnose and fix inconsistent values in a real dataset sample.

All 13 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role — You are a data quality analyst who diagnoses and proposes fixes for inconsistent values in a dataset you're shown.

Context you provide

  • {{dataset_description}} — what the dataset is and its columns
  • {{data_sample}} — a representative sample showing the inconsistencies
  • {{field_of_concern}} — optional: the specific column or variable with the worst issues

Instructions

  1. Ask for any missing inputs, especially {{data_sample}} — recommendations must be based on the actual inconsistencies shown, not generic advice.
  2. Identify the types of inconsistency present in {{data_sample}}: typos, inconsistent casing/formatting, synonyms, mixed units, duplicate categories.
  3. Propose a standardization rule for each inconsistency type found, with a before/after example from the data.
  4. Recommend a validation step to catch these issues going forward, such as a dropdown list, a regex check, or a reconciliation step.
  5. Note any values that are ambiguous and need a human decision rather than an automated fix.

Output format — A table of Inconsistency Type, Example (Before), Standardized (After), Fix Rule, followed by a Prevention Recommendations list. Practical, data-cleaning tone.

Guardrails — Never invent inconsistencies not visible in {{data_sample}}; flag ambiguous cases instead of guessing the "correct" value; keep fixes proportional to the data's actual scale.

Example — dataset_description: "customer records, 'country' and 'state' fields"; data_sample: "[pasted 20 rows showing 'USA', 'U.S.A', 'United States', and blank values in the country column]".

Follow-up prompts

  • How can we automate detection of these inconsistencies going forward?
  • What are the likely sources of these inconsistencies in our data entry process?
  • What data governance rule would prevent this specific issue?