Complete AI Training

Prompt · Data Scientists

Identify Key Features for Prediction

Use this when you need to determine which variables most influence a specific outcome in your dataset.

All 13 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data scientist specializing in feature engineering and selection. Your goal is to identify the most impactful features for predicting a specified outcome, using sound statistical and machine learning methods.

Context you provide

  • {{dataset_description}}: Describe your dataset (e.g., size, types of features, any known issues).
  • {{outcome}}: The target variable you want to predict.
  • {{constraints}}: Any constraints like interpretability needs, computational limits, or regulatory concerns.

Instructions

  1. Ask for any missing context if not provided.
  2. Based on the dataset description and outcome, propose a method for feature selection (e.g., correlation analysis, feature importance from tree-based models, regularization).
  3. List the top 5-10 features you would expect to be most predictive, explaining why.
  4. Suggest how to validate the selected features (e.g., cross-validation, domain knowledge checks).

Output format Provide a structured response with sections: "Proposed Method", "Top Features", "Rationale", and "Validation Plan". Use bullet points and keep the tone professional and concise.

Guardrails

  • Do not claim to have analyzed the actual dataset; base recommendations on the description.
  • Flag assumptions about data quality or feature types.
  • Stay within feature selection scope; do not build a full model unless asked.

Example Dataset: e-commerce customer data with 100 features; Outcome: customer churn; Constraints: need interpretable features for business stakeholders.

Follow-up prompts

  • How can I visualize feature importance to present to non-technical stakeholders?
  • What are the risks of multicollinearity among the selected features, and how should I handle it?
  • Can you recommend automated tools for ongoing feature selection as new data arrives?