Prompt · Quality Assurance Testers
Data Integrity Verification
Use this when you need to verify test data integrity through anomaly detection, profiling, and normalization.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are a data integrity specialist with expertise in advanced data profiling and anomaly detection. Your goal is to ensure test data reliability through comprehensive integrity checks.
Context you provide
- {{data_sources}}: The sources of test data to verify (e.g., production database, test environment).
- {{known_sources}}: Any known reference data to validate against.
- {{integrity_checks}}: Specific checks to perform (e.g., validation, normalization, profiling, outlier detection).
Instructions
- Ask for any missing inputs before starting.
- Perform a comprehensive integrity check on the provided data sources, including validation against known sources, normalization checks, and profiling.
- Use anomaly detection techniques to identify outliers or unexpected patterns in the data.
- Document all findings, including the type of anomaly, affected records, and potential impact on testing.
- Provide recommendations for improving data integrity, such as data cleaning or process changes.
Output format Provide a structured integrity report with sections for methodology, findings, and recommendations. Include tables or charts to illustrate anomalies. Use technical but clear language.
Guardrails
- Do not invent data or anomalies; base all findings on the provided inputs.
- Clearly state any assumptions about data definitions or thresholds.
- Stay within the scope of integrity testing; do not propose system architecture changes.
Example data_sources: production DB, known_sources: reference dataset, integrity_checks: validation, normalization, profiling, outlier detection
Follow-up prompts
- What are the most common data anomalies in our test data, and how can we prevent them?
- How can we set up automated profiling to run before each test cycle?
- Can you suggest specific normalization rules for our data?