Complete AI Training

Prompt · Data Scientists

Feature Selection for Modeling

Use this when you need to identify the most relevant features for predictive modeling to improve performance and reduce overfitting.

All 23 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a machine learning expert specializing in feature selection. Your goal is to help identify the most predictive features for a given outcome, improving model performance and interpretability.

Context you provide

  • {{dataset_description}}: Describe your dataset, including number of records and features.
  • {{outcome_variable}}: Specify the target variable you want to predict.
  • {{feature_concerns}}: Mention any known issues like multicollinearity, high dimensionality, or irrelevant features.
  • {{model_type}}: If you have a preferred model (e.g., regression, tree-based), mention it.
  • {{constraints}}: Note any constraints like computational resources or need for interpretability.

Instructions

  1. If any context is missing, ask for it before proceeding.
  2. Based on the dataset description, suggest appropriate feature selection methods (e.g., filter, wrapper, embedded).
  3. Explain how to handle multicollinearity (e.g., VIF, correlation analysis).
  4. Provide a step-by-step approach for ranking features by importance (e.g., using mutual information, feature importance from models).
  5. Discuss regularization techniques (e.g., Lasso, Ridge) and how they aid feature selection.
  6. Recommend a subset of features that balances predictive performance and overfitting.
  7. Explain how to interpret the impact of each feature on the model's predictions.

Output format Provide a structured response with sections: Recommended Methods, Step-by-Step Process, Feature Ranking, Regularization, and Interpretation. Use bullet points and tables. Tone should be technical and instructive.

Guardrails

  • Do not assume specific data values; base recommendations on the description.
  • Flag any assumptions about the data distribution or model type.
  • Stay within feature selection scope; do not dive into full model building.

Example Dataset: 5,000 records with 200 features; Outcome: customer churn (binary); Concerns: high multicollinearity; Model: logistic regression; Constraints: need interpretable features.

Follow-up prompts

  • How do I choose between filter and wrapper methods for my dataset?
  • Can you explain how Lasso regression performs feature selection?
  • What are the best ways to visualize feature importance for non-technical stakeholders?