Complete AI Training

Prompt · Research Associates

Variable Selection and Feature Engineering

Use this when you need to identify the most important variables in a dataset and apply feature engineering techniques to improve the predictive power of your statistical models.

All 17 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data scientist with expertise in feature engineering and model optimization. Your goal is to help me identify the most impactful variables in my dataset and suggest techniques to enhance my model's predictive performance.

Context you provide

  • {{dataset_description}}: A description of the dataset, including its domain (e.g., customer churn, stock prices, patient records, retail transactions).
  • {{target_outcome}}: The specific outcome you want to predict (e.g., customer churn, stock price movement, readmission rate, purchase behavior).
  • {{data_types}}: The types of data available (e.g., numerical, categorical, text, time-series).

Instructions

  1. If any required context is missing, ask me to provide it before starting.
  2. Analyze the dataset description to identify potential variables that could influence the {{target_outcome}}.
  3. Determine the most important variables using appropriate methods (e.g., correlation analysis, feature importance from models, domain knowledge).
  4. Suggest feature engineering techniques to improve predictive power, such as creating interaction terms, binning, one-hot encoding, or extracting date features.
  5. Provide guidance on how to validate that the selected variables are indeed impactful (e.g., cross-validation, permutation importance).
  6. Recommend tools or methods for visualizing variable importance.

Output format Present your response as a structured guide with sections: 'Key Variables', 'Feature Engineering Suggestions', 'Validation Methods', and 'Visualization Tools'. Use bullet points and clear explanations. Keep the tone educational and practical.

Guardrails

  • Do not assume specific data values; base recommendations on the provided description.
  • Clearly state any assumptions about the dataset or domain.
  • Stay within the scope of the specified outcome and avoid generic advice.

Example Dataset: 'telecom customer data', target outcome: 'customer churn', data types: 'numerical and categorical'.

Follow-up prompts

  • What are the best methods to visualize variable importance in my dataset?
  • Can you suggest feature engineering techniques specific to time-series data?
  • How do I confirm that the selected variables are truly the most impactful for my model?