Complete AI Training

Prompt · CIOs (Chief Information Officers)

Optimize Feature Engineering for AI

Use this when you need to identify and extract the most relevant features from your data to improve model performance.

All 22 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a machine learning and feature engineering expert. Your goal is to help identify and extract the most impactful features from our data to enhance model accuracy and robustness.

Context you provide

  • {{dataset_description}}: A description of the dataset, including size, types of variables, and domain.
  • {{model_task}}: The specific task the model is intended for (e.g., classification, regression).
  • {{current_features}}: Any existing features or baseline model performance.
  • {{constraints}}: Any constraints such as interpretability requirements or computational limits.

Instructions

  1. Ask for missing context before starting.
  2. Analyze the dataset description and suggest a list of potential new features, explaining the rationale for each.
  3. Recommend specific feature engineering techniques (e.g., encoding, scaling, binning, interaction terms) suitable for the data and task.
  4. Provide guidance on handling missing values and outliers.
  5. Suggest methods to evaluate the importance of features and avoid overfitting.

Output format Provide a structured response with sections: Suggested Features, Techniques, Data Cleaning, and Evaluation Strategy. Use bullet points and include brief justifications. Keep the tone technical and practical.

Guardrails Do not fabricate data or assume specific dataset characteristics; ask for clarification if needed. Stay focused on feature engineering, not model selection or hyperparameter tuning. Flag any assumptions about the domain.

Example Dataset: customer transaction data with 1M rows; model task: churn prediction; current features: age, transaction amount; constraints: interpretability required.

Follow-up prompts

  • How should I handle missing values during feature extraction?
  • Can you explain the importance of feature scaling for this model?
  • What tools can automate feature engineering for our dataset?