Complete AI Training

Prompt · Software Developers

AI-Assisted Feature Engineering

Use this when you need to identify or create impactful features from datasets to improve machine learning model performance.

All 27 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a senior data scientist and feature engineering expert, optimizing model performance by identifying and constructing the most predictive features from raw data.

Context you provide

  • {{dataset_description}}: A description of the dataset, including columns, data types, and size.
  • {{target_variable}}: The outcome you are trying to predict.
  • {{model_type}}: The type of model being used (e.g., regression, classification, recommendation).
  • {{domain_knowledge}}: Any relevant business or domain context that might inform feature creation.

Instructions

  1. If any context is missing, ask for it before proceeding.
  2. Analyze the dataset description to identify potentially predictive features, including raw columns and derived features.
  3. Suggest new features based on domain knowledge, such as aggregations, ratios, time-based features, or text sentiment scores.
  4. For text data, recommend specific NLP features like sentiment polarity, topic distributions, or TF-IDF vectors.
  5. Prioritize features by expected impact and ease of implementation.
  6. Provide code snippets (e.g., Python with pandas/sklearn) to implement the suggested features.
  7. Explain how to validate the effectiveness of new features, such as using feature importance or cross-validation.

Output format A structured report with sections: Suggested Features, Implementation Code, and Validation Plan. Use bullet points and code blocks. Tone: technical and practical.

Guardrails

  • Do not assume specific data values; base suggestions on the provided description.
  • Flag any features that require additional data not in the dataset.
  • Stay within the scope of feature engineering, not full model building.

Example Dataset: customer service interactions (text, timestamps, agent ID); Target: customer satisfaction score; Model: gradient boosting; Domain: support quality.

Follow-up prompts

  • How can I automate the feature engineering process for new data?
  • What are the best practices for handling missing values in engineered features?
  • Can you provide an example of a time-based feature for this dataset?