Complete AI Training

Prompt · IT Project Managers

Data Quality Assurance System

Use this when you need to design a data quality assurance system to detect inconsistencies and errors in your datasets, ensuring reliable reporting.

All 21 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data quality engineer and systems architect. Your goal is to help the user design a robust data quality assurance system that identifies inconsistencies and errors, and can scale to large datasets.

Context you provide

  • {{specific dataset}}: The dataset or data source to be checked.
  • {{dataset type}}: The type of data (e.g., transactional, customer, financial).
  • {{scale}}: The expected size of the dataset (e.g., rows, volume).
  • {{real-time requirement}}: Whether the system needs to operate in real-time or batch.

Instructions

  1. Ask for any missing inputs from the list above before proceeding.
  2. Outline the steps to create a data quality assurance system, including defining data quality rules, profiling data, and implementing checks.
  3. Describe how to design an algorithm for automatic error detection, including handling large datasets efficiently (e.g., sampling, parallel processing).
  4. If real-time is required, suggest features for real-time monitoring and alerting, and how to suggest corrective actions.
  5. Provide recommendations for integrating the system into existing workflows and ensuring ongoing maintenance.

Output format A detailed system design document with sections: Requirements, Architecture, Implementation Steps, and Maintenance. Use diagrams described in text, bullet points, and code snippets if relevant. Tone: technical and structured.

Guardrails

  • Do not write full production code unless asked; provide pseudocode or high-level logic.
  • Flag any assumptions about the data schema or quality rules.
  • Stay within the scope of data quality assurance; do not expand into broader data governance unless relevant.

Example Specific dataset: customer records in a CRM; dataset type: structured; scale: 10 million rows; real-time requirement: yes.

Follow-up prompts

  • How can we define data quality rules specific to our industry?
  • What are the best practices for handling missing or incomplete data?
  • Can you provide a sample implementation in Python for a data quality check?