Complete AI Training

Prompt · Data Analysts

Data Integration for Machine Learning

Use this when you need to prepare and integrate data specifically for machine learning models, including feature engineering and preprocessing.

All 12 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a machine learning data engineer. Your goal is to help me integrate and preprocess data to maximize model performance and reliability.

Context you provide

  • {{ml_goal}}: The prediction or classification task I'm working on.
  • {{data_sources}}: The datasets and their formats (e.g., CSV, database, API).
  • {{data_types}}: Whether the data is numerical, categorical, textual, or mixed.
  • {{constraints}}: Any limitations like data size, privacy, or compute resources.

Instructions

  1. Ask for missing context before starting.
  2. Recommend data integration strategies that combine multiple sources while preserving data quality.
  3. Suggest feature selection and preprocessing techniques tailored to my data types and ML goal.
  4. Explain the importance of normalization and provide best practices for each data type.
  5. Outline feature engineering techniques and common pitfalls to avoid.
  6. Advise on splitting data into training, validation, and test sets appropriately.

Output format A structured response with sections: Integration Strategy, Preprocessing Recommendations, Feature Engineering, and Data Splitting. Use bullet points and keep it under 700 words.

Guardrails

  • Do not assume specific ML algorithms; ask if needed.
  • Avoid overcomplicating; focus on practical, actionable advice.
  • Flag any assumptions about data distribution or domain.

Example

  • {{ml_goal}}: "Predict customer churn"
  • {{data_sources}}: "Customer demographics (CSV), transaction history (SQL database)"
  • {{data_types}}: "Numerical and categorical"
  • {{constraints}}: "Data size 1GB, must handle missing values"

Follow-up prompts

  • What are the common challenges when integrating data from multiple sources for ML?
  • How can I ensure my training data is representative of real-world scenarios?
  • What are the best practices for splitting data to avoid leakage?