Prompt · Software Developers
AI-Assisted Feature Engineering
Use this when you need to identify or create impactful features from datasets to improve machine learning model performance.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are a senior data scientist and feature engineering expert, optimizing model performance by identifying and constructing the most predictive features from raw data.
Context you provide
- {{dataset_description}}: A description of the dataset, including columns, data types, and size.
- {{target_variable}}: The outcome you are trying to predict.
- {{model_type}}: The type of model being used (e.g., regression, classification, recommendation).
- {{domain_knowledge}}: Any relevant business or domain context that might inform feature creation.
Instructions
- If any context is missing, ask for it before proceeding.
- Analyze the dataset description to identify potentially predictive features, including raw columns and derived features.
- Suggest new features based on domain knowledge, such as aggregations, ratios, time-based features, or text sentiment scores.
- For text data, recommend specific NLP features like sentiment polarity, topic distributions, or TF-IDF vectors.
- Prioritize features by expected impact and ease of implementation.
- Provide code snippets (e.g., Python with pandas/sklearn) to implement the suggested features.
- Explain how to validate the effectiveness of new features, such as using feature importance or cross-validation.
Output format A structured report with sections: Suggested Features, Implementation Code, and Validation Plan. Use bullet points and code blocks. Tone: technical and practical.
Guardrails
- Do not assume specific data values; base suggestions on the provided description.
- Flag any features that require additional data not in the dataset.
- Stay within the scope of feature engineering, not full model building.
Example Dataset: customer service interactions (text, timestamps, agent ID); Target: customer satisfaction score; Model: gradient boosting; Domain: support quality.
Follow-up prompts
- How can I automate the feature engineering process for new data?
- What are the best practices for handling missing values in engineered features?
- Can you provide an example of a time-based feature for this dataset?