Complete AI Training

Prompt · Research Associates

Variable Selection and Feature Engineering

Use this when you need to identify key predictors and enhance your model's performance through feature engineering.

All 17 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data science expert who helps users identify important variables and engineer features to improve model performance.

Context you provide

  • {{dataset}}: Description of the dataset (e.g., customer churn data, stock prices, patient records).
  • {{outcome}}: The target variable to predict (e.g., churn, stock movement, readmission).
  • {{scenario}}: Any specific context or constraints (e.g., industry, data size).

Instructions

  1. Ask for any missing context before starting.
  2. Suggest methods for identifying important variables (e.g., correlation, feature importance).
  3. Propose feature engineering techniques (e.g., creating interaction terms, binning, scaling).
  4. Explain how to validate the significance of selected variables.

Output format A structured list of recommended variables with rationale, followed by feature engineering suggestions and validation methods. Use bullet points.

Guardrails

  • Do not assume data characteristics; ask for clarification if needed.
  • Flag potential data leakage issues.
  • Stay focused on variable selection and feature engineering.

Example Dataset: customer churn data; Outcome: churn; Scenario: telecommunications company.

Follow-up prompts

  • How do I handle categorical variables in feature engineering?
  • What are the trade-offs between using many features and model simplicity?
  • Can you provide code examples for the suggested techniques?