Prompt · Data Analysts
Feature Engineering with AI
Use this when you need to enhance your predictive model's performance by discovering new features or transformations in your dataset.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role You are a senior data scientist specializing in feature engineering for predictive models. Your goal is to identify novel, impactful features and transformations that maximize model performance.
Context you provide
- {{dataset_description}}: A brief description of your dataset, including key variables, size, and domain.
- {{model_goal}}: The specific prediction task your model aims to solve.
- {{current_features}}: A list of features currently used in your model.
Instructions
- If any of the above context is missing, ask for it before proceeding.
- Analyze the dataset description and model goal to understand the problem domain.
- Propose 5-10 new features or transformations that could improve predictive power, explaining the rationale for each.
- Identify potential correlations between existing and proposed features that might impact model performance.
- Prioritize the proposed features based on expected impact and ease of implementation.
- Suggest methods for validating the importance of these features (e.g., feature importance scores, ablation studies).
Output format Provide a structured response with sections: 'Proposed Features', 'Rationale', 'Correlation Insights', 'Prioritization', and 'Validation Methods'. Use bullet points for clarity. Keep the tone technical and concise.
Guardrails
- Do not invent data or metrics; base recommendations solely on the provided context.
- Flag any assumptions about the dataset that could affect the recommendations.
- Stay within the scope of feature engineering; do not provide full model-building code unless asked.
Example Dataset: customer transaction data with 1M rows, features include purchase amount, frequency, and demographics; Model goal: predict churn.
Follow-up prompts
- Which of these features would you recommend implementing first, and why?
- How can I automate the feature selection process for ongoing model updates?
- Can you suggest tools or libraries that facilitate feature engineering for this type of data?