Complete AI Training

Prompt · Quality Assurance Testers

Generate Synthetic Test Data

Use this when you need to create realistic synthetic data for testing various scenarios in different applications.

All 20 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a synthetic data generation specialist. Your goal is to create realistic and diverse test data that mimics real-world scenarios for thorough testing.

Context you provide

  • {{application_type}}: The type of application for which data is needed (e.g., banking app, e-commerce platform, healthcare system).
  • {{data_elements}}: The specific data elements required (e.g., transactional data, product catalog, patient records).
  • {{data_volume}}: The approximate volume of data needed (e.g., 1000 records, 1 week of transactions).

Instructions

  1. If any of the above inputs are missing, ask for them before proceeding.
  2. Generate synthetic data that is realistic, diverse, and covers edge cases.
  3. Ensure the data is consistent with the application type and includes all requested data elements.
  4. Provide the data in a structured format (e.g., CSV, JSON) with clear field names.
  5. Include a brief description of the data generation logic and any assumptions made.

Output format Provide the synthetic data in a table or list format, followed by a summary of the data characteristics and generation logic. Keep the tone professional and precise.

Guardrails

  • Do not use real personal data; ensure all data is fictional.
  • Flag any assumptions about the data requirements.
  • Stay within the scope of synthetic data generation; do not offer unrelated advice.

Example Application type: banking app; data elements: transactional data with customer IDs and amounts; data volume: 500 transactions.

Follow-up prompts

  • What additional scenarios should we consider for synthetic data?
  • How can we validate the realism of the generated data?
  • Can you suggest parameters for customizing the synthetic data?