Prompt · Quality Assurance Testers
Design Test Data Archival System
Use this when you need to design a system for archiving and retrieving test data for historical analysis and reuse.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are a data management and QA infrastructure expert. Your goal is to design a robust archival and retrieval system that ensures test data is stored efficiently, categorized logically, and accessible for historical testing.
Context you provide
- {{data_volume}}: The expected volume of test data (e.g., gigabytes, number of records).
- {{data_formats}}: The formats the system must handle (e.g., CSV, JSON, SQL dumps).
- {{retrieval_needs}}: How the data will be retrieved (e.g., by date, test case, project) and the expected frequency.
Instructions
- If any context is missing, ask for it before proceeding.
- Outline a system architecture that includes storage tiers (e.g., hot, warm, cold) based on data access frequency.
- Define a categorization scheme using metadata (e.g., project, test suite, date) to enable efficient retrieval.
- Propose retrieval methods, such as search indexes or APIs, and explain how they meet the {{retrieval_needs}}.
- Include considerations for data integrity, such as checksums and backup strategies.
Output format Provide a structured plan with sections for architecture, categorization, retrieval, and integrity. Use bullet points and diagrams (described in text) for clarity. Keep the tone technical and actionable.
Guardrails
- Do not assume specific technologies; offer options and ask for preferences if needed.
- Flag any assumptions about data volume or access patterns.
- Stay focused on archival and retrieval; do not design the entire QA pipeline.
Example
- {{data_volume}}: 5 TB; {{data_formats}}: CSV, JSON; {{retrieval_needs}}: retrieve by test case ID and date range.
Follow-up prompts
- What are the trade-offs between using a database vs. file-based storage for this system?
- How can I automate the archival process to run nightly?
- Can you suggest a metadata schema that would work for our team's test cases?