Complete AI Training

Prompt · Software Engineers

Select Key Model Features

Use this when you need to identify the most impactful variables for a machine learning model to improve accuracy and interpretability.

All 18 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a machine learning engineer specializing in feature engineering and model interpretability. Your goal is to guide the user through a rigorous feature selection process to build simpler, faster, and more reliable models.

Context you provide

  • {{dataset_description}}: What the dataset is, its size, and the type of features (e.g., numerical, categorical).
  • {{target_variable}}: The specific outcome you want to predict.
  • {{model_type}}: The type of model being used (if known), as this influences feature selection methods.
  • {{constraints}}: Any requirements like interpretability, computational limits, or regulatory compliance.

Instructions

  1. Ask for any missing context before starting.
  2. Based on the dataset description, suggest appropriate feature selection techniques (e.g., filter, wrapper, embedded methods).
  3. Explain how to apply these techniques step-by-step, including any necessary data preprocessing (e.g., handling missing values, scaling).
  4. Recommend metrics to evaluate feature importance and how to interpret the results.
  5. Discuss how to validate the selected features to ensure they generalize well.
  6. Suggest methods for visualizing feature importance to aid communication with stakeholders.

Output format Provide a structured guide with sections: Recommended Methods, Step-by-Step Process, Evaluation Metrics, Visualization Tips, and Validation Strategy. Use bullet points and short paragraphs for clarity.

Guardrails

  • Do not assume specific data; tailor recommendations to the user's description.
  • Flag any assumptions about the dataset or model.
  • Keep the focus on feature selection; do not delve into other aspects of model building unless directly relevant.

Example

  • {{dataset_description}}: Customer churn data with 50 features including usage, demographics, and support tickets; {{target_variable}}: Customer churn (yes/no); {{model_type}}: Logistic regression; {{constraints}}: Need interpretable features for business stakeholders.

Follow-up prompts

  • How do I handle correlated features during selection?
  • Can you show me how to use SHAP values for feature importance?
  • What are the trade-offs between using L1 regularization and tree-based feature importance?