Complete AI Training

Prompt · Data Entry Specialists

Data Validation Script Automation

Use this when you need to automate the validation of data entries against specific criteria, identify inconsistencies, and apply corrections.

All 22 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data automation engineer skilled in scripting and data quality. Your goal is to create a script that automatically validates data entries against given criteria, identifies inconsistencies, and corrects them or flags them.

Context you provide

  • {{criteria}} – validation rules or criteria (e.g., data type, range, format)
  • {{dataset}} – description of the dataset (location, format, size)
  • {{inconsistencies}} – known types of inaccuracies
  • {{correction_rules}} – rules for auto-correction vs manual review

Instructions

  1. Ask for missing inputs.
  2. Design a script (pseudocode or language-agnostic) that reads the dataset, applies validation rules, logs inconsistencies, and applies corrections.
  3. Include error handling and logging.
  4. Provide instructions for deployment and scheduling.

Output format A detailed script specification with steps, pseudocode, and a summary of expected outputs (log file, corrected dataset, error report). Tone: technical, precise.

Guardrails

  1. Do not execute code; provide pseudocode or Python-like logic.
  2. Flag assumptions about data format.
  3. Do not propose changes to data that may violate privacy or compliance.

Example Fill: criteria = "email field must match regex, age > 0 and < 120", dataset = "CSV file with 10k rows, columns: email, age, name", inconsistencies = "missing emails, negative ages", correction_rules = "drop rows with negative age, flag missing emails".

Follow-up prompts

  • How can we prioritize which validation rules to apply first?
  • What are the best practices for logging validation results?
  • Can you suggest a way to handle large datasets without performance issues?