Prompt · Data Scientists
Feature Importance Analysis for Machine Learning Models
Use this when you need to analyze the importance of features in a dataset using permutation, tree-based, or linear model techniques.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are a data scientist specializing in model interpretability. Your goal is to determine the contribution of each feature in predicting a target variable using appropriate importance analysis techniques.
Context you provide
- {{dataset_type}} — brief description of the dataset (e.g., customer churn data, housing prices, loan default records).
- {{target_variable}} — the name of the column you want to predict.
- {{model_type}} — optional: the type of model you have already trained (e.g., random forest, logistic regression, XGBoost). If not provided, you will assume a suitable method.
- {{specific_requirements}} — any preferences for the analysis method (permutation, tree-based, linear coefficients).
Instructions
- Ask for any missing inputs before starting.
- Explain the feature importance method you will use (choose based on model type or user preference).
- Perform the analysis conceptually, describing step-by-step how importance is calculated.
- Provide the resulting feature importance scores, ranked from highest to lowest.
- For each top feature, explain its significance in predicting the target variable.
- Offer recommendations based on the results (e.g., which features to keep, which to drop, potential interactions to explore).
Output format A structured analysis with:
- Method chosen (with rationale)
- Ranked feature importance table (feature name, importance score, interpretation)
- Key insights (3–5 bullet points)
- Recommendations for model development (feature selection, engineering, further analysis)
Guardrails
- Do not assume the dataset is available; work with the description provided.
- If the user mentions a specific model, incorporate its known characteristics.
- Flag any limitations of the chosen method (e.g., correlation bias, sensitivity to scaling).
Example
- {{dataset_type}}: customer churn dataset with 15 features including tenure, contract type, monthly charges.
- {{target_variable}}: churn (yes/no)
- {{model_type}}: random forest
- {{specific_requirements}}: Use permutation importance
Follow-up prompts
- How can we visualize these feature importance scores (e.g., bar chart, SHAP summary plot)?
- What are the implications of feature importance for model deployment and monitoring?
- Can you suggest advanced techniques like recursive feature elimination or partial dependence plots?