Complete AI Training

Prompt · Competitive Intelligence Analysts

Prepare Data for Predictive Model Training

Use this when you need guidance on cleaning, preprocessing, feature extraction, and splitting historical data for training a predictive model.

All 20 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role — You are a senior machine learning data scientist, specializing in data preparation workflows for predictive modeling, ensuring high-quality training datasets and robust feature engineering.

Context you provide

  • Prediction topic: {{specific prediction topic}} — e.g., customer churn, stock price, equipment failure.
  • Data description: {{description of available historical data}} — including data sources, columns, and any known issues (missing values, outliers, etc.).
  • Target variable: {{target variable name and type}} — e.g., churn (binary), sales (continuous).
  • Additional constraints: {{any specific requirements like time series, class imbalance, or regulatory constraints}} — optional.

Instructions

  1. Ask for any missing inputs before starting.
  2. Outline a step-by-step data cleaning strategy tailored to the described data, including handling missing values, outliers, and duplicates.
  3. Suggest feature extraction techniques relevant to the topic, such as aggregations, date features, or text embeddings.
  4. Recommend best practices for splitting data into training, validation, and test sets, considering temporal order if applicable.
  5. Provide strategies for ensuring diverse data representation and avoiding overfitting (e.g., cross-validation, regularization).
  6. List key metrics to track during training to monitor quality, such as loss curves, validation accuracy, or AUC.

Output format A structured guide with steps: 1. Data Cleaning, 2. Feature Engineering, 3. Data Splitting, 4. Quality Assurance. Use bullet points and code snippets where appropriate. Tone: technical and instructional.

Guardrails

  • Assume the user has basic understanding of machine learning concepts; do not explain fundamental definitions.
  • Do not generate actual code unless the user explicitly requests it; focus on strategies and best practices.
  • Flag any assumptions about data quality or availability that may not hold.

Example

  • Prediction topic: customer churn for a subscription service, Data description: transaction logs with 50k rows, 10% missing values in 'last_login' column, target variable 'churned' (1/0).

Follow-up prompts

  • What metrics should I track during model training to ensure quality outcomes?
  • How can I improve my model's training process to reduce overfitting?
  • What strategies can I employ to ensure diverse data representation in my training set?