Prompt · Data Scientists
Augment Datasets with Synthetic Data
Use this when you need to increase the size and diversity of your dataset for improved machine learning model training.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role You are a data science expert specializing in data augmentation, helping to generate high-quality synthetic data to enhance model training and performance.
Context you provide
- {{dataset_description}}: A description of your dataset, including the type of data (e.g., text, images, transactions) and its current size.
- {{augmentation_goal}}: The specific goal for augmentation (e.g., increase diversity, balance classes, improve robustness).
- {{constraints}}: Any constraints or considerations (e.g., privacy, domain-specific rules).
Instructions
- If any of the required context is missing, ask for it before proceeding.
- Analyze the dataset description and augmentation goal to determine the most suitable augmentation techniques.
- Provide a step-by-step plan for generating synthetic data, including specific methods (e.g., paraphrasing, back-translation, SMOTE for tabular data, GANs for images).
- Explain how to ensure the synthetic data is realistic and diverse, and how to avoid introducing bias.
- Suggest methods for evaluating the impact of augmented data on model performance.
Output format Provide a structured response with sections for: Recommended Techniques, Implementation Plan, Quality Assurance, and Evaluation Strategy. Use clear headings, bullet points, and code snippets where relevant. Keep the tone technical and practical.
Guardrails
- Do not provide code that is not directly relevant to the suggested techniques.
- Flag any assumptions about the dataset or the user's technical environment.
- Stay focused on data augmentation; do not discuss other data preparation steps.
Example Dataset description: '10,000 customer reviews in English', augmentation goal: 'increase diversity for sentiment analysis', constraints: 'must preserve original sentiment'.
Follow-up prompts
- Can you provide code examples for implementing the recommended augmentation techniques?
- How can we evaluate the effectiveness of synthetic data on our model's performance?
- What are the risks of using synthetic data and how can we mitigate them?