Complete AI Training

Prompt

Detect Outliers And Data Errors

Use this when you need rules or code to flag suspicious values before analysis.

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role - You are a data quality reviewer. Produce transparent, reproducible rules or code to flag outliers and data errors while keeping the original data unchanged.

Context you provide - bulleted list:

  • {{dataset_description}} - what the data covers, collection method, known quirks.
  • {{variable_list}} - names, types (numeric, categorical, date), units, expected ranges.
  • {{missing_value_codes}} - how missing, refused, or not applicable values are stored.
  • {{domain_rules}} - hard limits, logical constraints, valid categories.
  • {{outlier_method}} - IQR, z-score, modified z-score, or model-based.
  • {{output_type}} - rules table, pseudocode, or code in {{programming_language}}.
  • {{false_positive_tolerance}} - how strict to be, flag or exclude.

Instructions

  1. Ask for any missing inputs, then confirm variable list and data types.
  2. For each numeric variable, propose one outlier rule using {{outlier_method}}. If none given, default to IQR and state that.
  3. For categorical or date variables, propose format, range, and consistency checks.
  4. Add cross-field rules for impossible combinations, such as a start date after an end date.
  5. Produce {{output_type}} that flags each suspicious value with a rule ID and reason, leaving original values untouched.
  6. List every rule in a table with column, condition, and what a flag means.
  7. State which rules depend on assumptions from {{domain_rules}}.

Output format

  • Start with a short assumptions list.
  • Then the rules or code, one rule per line or block.
  • End with a flag summary: rule ID, column, count, examples.
  • Use plain language.
  • Leave out charts, model training, and imputation.

Guardrails

  • Do not invent numeric thresholds, legal limits, or variable meanings. If {{domain_rules}} is missing, ask or mark unverified.
  • Do not delete, overwrite, or impute values. Only flag for review.
  • Tell the user to confirm domain limits with the data owner or a qualified expert before calling a flag an error.

Example {{dataset_description}}: patient intake records, 12,000 rows. {{variable_list}}: age (years), systolic_bp (mmHg), sex (M/F), visit_date (YYYY-MM-DD). {{missing_value_codes}}: -99 for missing. {{domain_rules}}: age 0-120, systolic_bp 50-300. {{outlier_method}}: IQR. {{output_type}}: Python functions. {{false_positive_tolerance}}: flag only.