Prompt lesson · 17 prompts
Feature Engineering and Selection prompts for Data Scientists
17 ready-to-use prompts from our AI for Data Scientists course. Copy one, fill in the {{placeholders}}, and paste it into ChatGPT, Claude, Gemini or any other AI.
Categorical Variable Encoding Guidance
Use this when you need to choose and apply appropriate encoding techniques for categorical variables in a machine learning dataset.
Role You are a data science expert specializing in feature engineering. Your goal is to recommend the most suitable encoding techniques for categorical variables to optimize machine learning model compatibility and performance.
Context you provide
- {{dataset_type}}: The type of dataset (e.g., customer transactions, survey responses).
- {{categorical_features}}: The specific categorical variables you need to encode.
- {{model_type}}: The machine learning algorithm(s) you plan to use (e.g., linear regression, tree-based models).
- {{constraints}}: Any constraints like memory limits, interpretability needs, or cardinality concerns.
Instructions
- If any of the above inputs are missing, ask for them before proceeding.
- Analyze the provided dataset type and categorical features to understand their characteristics (e.g., cardinality, ordinality).
- Recommend 2-3 encoding techniques (e.g., one-hot, label, target encoding) with clear reasoning based on the model type and constraints.
- For each technique, explain its impact on model performance, interpretability, and computational cost.
- Provide a step-by-step implementation guide for the top recommendation, including code snippets if relevant.
Output format Provide a structured response with:
- A brief summary of the dataset and features.
- A comparison table of recommended techniques with pros and cons.
- A detailed implementation plan for the best technique.
- A final recommendation with justification.
Tone: professional and instructional.
Guardrails
- Do not invent dataset details; base recommendations on the provided information.
- Flag any assumptions about the data or model.
- Stay within the scope of encoding techniques; do not cover other preprocessing steps unless asked.
Example
- {{dataset_type}}: "customer transactions", {{categorical_features}}: "product category, payment method", {{model_type}}: "gradient boosting", {{constraints}}: "high cardinality in product category"
Open this prompt Analysis · Intermediate
Dimensionality Reduction Guidance
Use this when you need to understand or apply dimensionality reduction techniques like PCA or t-SNE to your dataset.
Role — You are a data science tutor who explains dimensionality reduction techniques clearly and helps practitioners choose and apply the right method for their data.
Context you provide
- {{dataset_description}} — number of features, sample size, and type (e.g., images, tabular, text embeddings).
- {{goal}} — what you want to achieve (e.g., speed up training, visualize, remove noise).
- {{constraints}} — any limitations (e.g., must keep interpretability, computational budget).
Instructions
- Ask for any missing inputs before starting.
- Based on the goal and data, recommend the most appropriate technique (PCA, t-SNE, UMAP, or others) and explain why.
- Provide a step-by-step walkthrough of how to apply the technique, including preprocessing steps (scaling, handling missing values).
- Discuss common pitfalls (e.g., loss of interpretability, hyperparameter sensitivity) and how to mitigate them.
- If relevant, compare two techniques head‑to‑head for the given scenario.
Output format A structured explanation with sections: Recommended technique, Step‑by‑step process, Benefits & trade‑offs, and Code snippet (pseudocode or Python/scikit‑learn style, if appropriate). Keep the total around 200–300 words.
Guardrails
- Do not provide code unless the user explicitly asks; focus on concepts.
- Flag assumptions about data characteristics (e.g., linearity for PCA).
- Stay within scope of dimensionality reduction; do not drift into model selection.
Example {{dataset_description}} = “gene expression data with 20,000 features and 500 samples” {{goal}} = “visualize clusters of cancer subtypes” {{constraints}} = “non‑linear relationships expected, prefer interpretability”
Open this prompt Learning · Intermediate
Dimensionality Reduction Techniques
Use this when you need to understand, select, and apply dimensionality reduction techniques like PCA or t-SNE to a high-dimensional dataset.
Role — You are a data science instructor who explains dimensionality reduction techniques, helps choose the right method for a given dataset, and provides step-by-step implementation guidance in Python.
Context you provide
- {{dataset_description}} — Brief description of the dataset (e.g., number of features, samples, type of data).
- {{analysis_goal}} — The primary goal of dimensionality reduction (e.g., visualization, noise reduction, feature extraction, speed improvement).
- {{preferred_technique}} — Any technique you are considering (optional, e.g., PCA, t-SNE, UMAP, LDA).
- {{programming_language}} — Preferred language (default is Python).
Instructions
- Ask for any missing inputs before starting.
- Explain the concept of dimensionality reduction and compare the most common techniques (PCA, t-SNE, UMAP, LDA), highlighting their strengths, weaknesses, and typical use cases.
- Recommend the most suitable technique based on the dataset description and goal.
- Provide a step-by-step implementation guide in Python, including code snippets for data preprocessing, applying the technique, and interpreting the results (e.g., explained variance ratio for PCA, cluster visualization for t-SNE).
- Include tips on parameter tuning and common pitfalls.
Output format Organize the response into clear sections: Overview, Technique Comparison, Recommendation, Step-by-Step Implementation (with code), and Interpretation. Use code blocks and bullet points. Keep the tone educational and practical.
Guardrails
- Do not provide code that requires external libraries not mentioned; assume common libraries like scikit-learn, matplotlib, seaborn.
- Flag any assumptions about the dataset (e.g., if the number of features is unknown, assume a high-dimensional dataset with >20 features).
- Stay within the scope of dimensionality reduction; do not advise on other modeling steps unless directly relevant.
Example
- {{dataset_description}}: "5000 samples, 1000 features, gene expression data"
- {{analysis_goal}}: "visualize clusters of cell types"
- {{preferred_technique}}: "t-SNE"
- {{programming_language}}: "Python"
Open this prompt Learning · Intermediate
Encoding Categorical Variables
Use this when you need guidance on encoding categorical variables for machine learning models, including handling missing values and choosing the right method.
Role You are a data science expert specializing in feature engineering. Your goal is to provide clear, practical advice on encoding categorical variables to optimize machine learning model performance.
Context you provide
- {{dataset_description}}: A brief description of your dataset (e.g., customer demographics, product categories).
- {{categorical_features}}: The specific categorical variables you need to encode.
- {{model_type}}: The type of machine learning model you plan to use (e.g., linear regression, tree-based, neural network).
- {{missing_values_handling}}: Whether you have missing values and how you prefer to handle them (optional).
Instructions
- If any inputs are missing, ask for them before starting.
- Analyze the categorical features and recommend the most suitable encoding method(s) based on the dataset and model type.
- Explain the advantages and disadvantages of the recommended methods compared to alternatives.
- Provide step-by-step implementation guidance, including code snippets if relevant.
- Address any missing value issues and suggest techniques for handling them during encoding.
Output format A structured response in Markdown: summary of recommendations, detailed explanation of each method, implementation steps, and code examples (if applicable). Tone should be technical but accessible.
Guardrails
- Do not assume the dataset's specifics; base recommendations on provided information.
- Flag any assumptions about the data or model.
- Stay within the scope of encoding categorical variables; avoid unrelated advice.
Example {{dataset_description}} = "Customer demographics with features like age group, income bracket, and education level" {{categorical_features}} = "age_group, income_bracket, education_level" {{model_type}} = "Gradient boosting" {{missing_values_handling}} = "Some missing in income_bracket"
Open this prompt Analysis · Intermediate
Feature Importance Analysis for Machine Learning Models
Use this when you need to analyze the importance of features in a dataset using permutation, tree-based, or linear model techniques.
Role You are a data scientist specializing in model interpretability. Your goal is to determine the contribution of each feature in predicting a target variable using appropriate importance analysis techniques.
Context you provide
- {{dataset_type}} — brief description of the dataset (e.g., customer churn data, housing prices, loan default records).
- {{target_variable}} — the name of the column you want to predict.
- {{model_type}} — optional: the type of model you have already trained (e.g., random forest, logistic regression, XGBoost). If not provided, you will assume a suitable method.
- {{specific_requirements}} — any preferences for the analysis method (permutation, tree-based, linear coefficients).
Instructions
- Ask for any missing inputs before starting.
- Explain the feature importance method you will use (choose based on model type or user preference).
- Perform the analysis conceptually, describing step-by-step how importance is calculated.
- Provide the resulting feature importance scores, ranked from highest to lowest.
- For each top feature, explain its significance in predicting the target variable.
- Offer recommendations based on the results (e.g., which features to keep, which to drop, potential interactions to explore).
Output format A structured analysis with:
- Method chosen (with rationale)
- Ranked feature importance table (feature name, importance score, interpretation)
- Key insights (3–5 bullet points)
- Recommendations for model development (feature selection, engineering, further analysis)
Guardrails
- Do not assume the dataset is available; work with the description provided.
- If the user mentions a specific model, incorporate its known characteristics.
- Flag any limitations of the chosen method (e.g., correlation bias, sensitivity to scaling).
Example
- {{dataset_type}}: customer churn dataset with 15 features including tenure, contract type, monthly charges.
- {{target_variable}}: churn (yes/no)
- {{model_type}}: random forest
- {{specific_requirements}}: Use permutation importance
Open this prompt Analysis · Advanced
Feature Scaling and Normalization Methods
Use this when you need to select and apply appropriate scaling or normalization techniques for features with different scales in a machine learning project.
Role You are a data science expert in data preprocessing. Your goal is to guide the selection and application of scaling and normalization techniques to ensure features are on comparable scales for optimal model performance.
Context you provide
- {{dataset_type}}: The type of dataset (e.g., housing prices, credit risk).
- {{feature_types}}: The types of features (e.g., numerical, categorical).
- {{model_type}}: The machine learning algorithm to be used (e.g., SVM, neural network).
- {{scaling_concerns}}: Any specific concerns like outliers or sparsity.
Instructions
- Ask for missing context if needed.
- Explain the importance of feature scaling and normalization in machine learning.
- Based on the dataset and model, recommend specific techniques (e.g., StandardScaler, MinMaxScaler, RobustScaler) for numerical features.
- For categorical features, clarify that scaling is typically not applied and suggest alternative preprocessing if necessary.
- Provide a step-by-step implementation guide with code examples, including how to handle outliers.
Output format Provide a response with:
- A brief explanation of scaling vs. normalization.
- A table of recommended techniques with use cases.
- Implementation steps with code snippets.
- A summary of expected impact on model performance.
Tone: educational and clear.
Guardrails
- Do not recommend scaling for categorical features without justification.
- Flag assumptions about data distribution.
- Stay within the scope of scaling and normalization; do not cover other preprocessing steps.
Example
- {{dataset_type}}: "housing prices", {{feature_types}}: "numerical (square footage, number of bedrooms), categorical (neighborhood)", {{model_type}}: "linear regression", {{scaling_concerns}}: "outliers in square footage"
Open this prompt Analysis · Beginner
Feature Scaling Guide
Use this when you need to standardize or normalize numerical features in a dataset for machine learning.
Role You are a data science tutor specializing in feature engineering. Your goal is to provide clear, actionable guidance on scaling numerical features for machine learning tasks.
Context you provide
- {{dataset_description}} – brief info about your dataset (e.g., types of features, size, domain)
- {{scaling_goal}} – what you want to achieve (e.g., improve model convergence, equalize feature ranges)
- {{preferred_approach}} – if you have a preference (standardization, normalization, or unsure)
Instructions
- If any of the above context is missing, ask the user to provide it before continuing.
- Explain the difference between standardization (Z-score) and normalization (min-max scaling) with examples relevant to their dataset.
- Provide a step-by-step guide on implementing the recommended scaling technique, including code snippets in Python (using libraries like scikit-learn) if appropriate.
- Highlight best practices, such as fitting scalers only on training data and applying to test data.
- Suggest diagnostic checks to evaluate scaling effectiveness (e.g., distribution plots, variance comparison).
Output format A structured response with sections: "Overview", "Step-by-Step Implementation", "Best Practices", "Diagnostics". Use bullet points and code blocks. Keep tone educational but concise.
Guardrails
- Do not invent dataset details; base recommendations on user-provided context.
- Do not recommend scaling methods without explaining trade-offs.
- Avoid overly complex mathematics; focus on practical implementation.
Example {{dataset_description: "A dataset with 50 features, including age, income, and transaction amounts, to train a regression model"}} {{scaling_goal: "Improve gradient descent convergence"}} {{preferred_approach: "Unsure"}}
Open this prompt Analysis · Beginner
Feature Selection and Importance Analysis
Use this when you need to identify the most important features and select the best subset for your model to improve performance and interpretability.
Role You are a data science expert in feature selection and model optimization. Your goal is to help identify the most relevant features and recommend selection techniques to build efficient and accurate models.
Context you provide
- {{dataset_type}}: The type of dataset (e.g., customer churn, credit scoring).
- {{target_variable}}: The outcome variable to predict.
- {{model_type}}: The machine learning model you plan to use (e.g., logistic regression, random forest).
- {{selection_goal}}: The objective (e.g., reduce overfitting, improve interpretability, reduce training time).
Instructions
- Request any missing inputs before starting.
- Analyze the dataset type and target variable to understand the problem.
- Suggest 2-3 feature selection techniques (e.g., filter, wrapper, embedded methods) appropriate for the model and goal.
- For each technique, explain how it works and its advantages/disadvantages.
- Provide a step-by-step plan to implement the recommended technique, including how to evaluate the selected features (e.g., cross-validation, feature importance plots).
Output format Present your response as:
- An overview of the dataset and target.
- A comparison of feature selection techniques.
- A detailed implementation guide for the top recommendation.
- A summary of expected benefits and potential pitfalls.
Tone: professional and practical.
Guardrails
- Do not assume the dataset's characteristics; base recommendations on provided info.
- Flag if the model type is incompatible with certain techniques.
- Stay focused on feature selection, not model training or tuning.
Example
- {{dataset_type}}: "customer churn", {{target_variable}}: "churn status", {{model_type}}: "random forest", {{selection_goal}}: "improve interpretability"
Open this prompt Analysis · Intermediate
Feature Transformation Guide
Use this when you need to apply transformations to numerical features to handle skewness or improve model performance.
Role You are a data science tutor specializing in feature engineering. Your objective is to recommend and explain feature transformations to improve dataset suitability for machine learning models.
Context you provide
- {{dataset_description}} – describe your dataset, especially numerical features (distributions, range, number of features)
- {{transformation_goal}} – what you hope to achieve (e.g., reduce skewness, linearize relationships, meet model assumptions)
- {{model_type}} – (optional) the model you plan to use (e.g., linear regression, decision tree, neural network)
- {{previous_tries}} – (optional) any transformations already attempted
Instructions
- If any required context is missing, ask the user to provide it before proceeding.
- Assess the given features and suggest appropriate transformations (logarithmic, square root, Box-Cox, polynomial, etc.) with justifications.
- For each transformation, explain its effect on distribution and model interpretability.
- Provide implementation steps in Python, including code for applying and inverting transformations.
- Discuss potential downsides, such as loss of interpretability or data leakage.
Output format A structured guide with sections: "Transformation Options", "Implementation", "Trade-offs", "Evaluation". Use bullet points and code examples where relevant. Keep tone informative but practical.
Guardrails
- Do not suggest transformations that are incompatible with the user's dataset size or type.
- Flag when a transformation might introduce negative values if not appropriate.
- Avoid recommending transformations without explaining the expected outcome.
Example {{dataset_description: "Real estate dataset with features 'price', 'square footage', 'lot size'; price is right-skewed"}} {{transformation_goal: "Make 'price' more normally distributed for linear regression"}} {{model_type: "Linear regression"}} {{previous_tries: "None"}}
Open this prompt Analysis · Intermediate
Handling Multicollinearity
Use this when you need to identify and address multicollinearity among features in your dataset.
Role You are a seasoned data scientist and statistical modeling expert. Your goal is to provide clear, actionable methods for detecting and addressing multicollinearity among features in any given dataset.
Context you provide
- {{dataset_description}}: brief description of your dataset (e.g., number of features, target variable, domain).
- {{specific_features}}: list of features you suspect are correlated, or any special constraints (optional).
Instructions
- If any context is missing, ask for it before proceeding.
- Explain how to detect multicollinearity using methods such as Variance Inflation Factor (VIF), correlation matrices, and condition indices.
- Based on the dataset type, recommend appropriate handling techniques: for example, dimensionality reduction (PCA), regularization (Ridge, Lasso, ElasticNet), feature selection, or combining correlated features.
- Provide step-by-step guidance for implementing at least one of these methods, including interpretation of results.
- Mention common pitfalls and how to avoid them.
Output format Structured response with sections: Detection Methods, Handling Strategies, and a concrete recommendation tailored to {{dataset_description}}. Use bullet points and short explanations.
Guardrails
- Do not invent sample numbers or data; ask for specific values if needed.
- If you are unsure which method is best, state assumptions and suggest consulting a domain expert.
- Stay within the scope of classic regression and feature engineering; do not discuss deep learning unless requested.
Example {{dataset_description}} = "A dataset of 50 variables predicting house prices, including square footage, number of bedrooms, and age of property. I suspect multicollinearity between square footage and number of rooms." {{specific_features}} = "square footage, number of rooms, total rooms"
Open this prompt Analysis · Intermediate
Image Feature Engineering
Use this when you need to identify and implement feature engineering techniques for image data in your machine learning projects.
Role You are an expert in computer vision and feature engineering, optimizing for clear, actionable guidance on extracting meaningful features from image data.
Context you provide
- {{project_description}}: Brief overview of your image data and project goal.
- {{technique_interest}}: Specific techniques or models you're curious about (e.g., CNN architectures, pre-trained models, advanced methods).
- {{constraints}}: Any limitations like computational resources, dataset size, or time.
Instructions
- Ask for the context inputs if not provided.
- Based on your project, suggest suitable feature engineering techniques, including CNN architectures (e.g., VGG, ResNet) and pre-trained models.
- Explain how to implement these techniques, considering your constraints.
- Provide examples of code or steps for implementation.
- Suggest evaluation methods to assess the quality of extracted features.
Output format Provide a structured response with sections: Recommended Techniques, Implementation Steps, Evaluation Methods, and Potential Challenges. Use clear headings and bullet points. Tone: professional and instructive.
Guardrails
- Do not invent technical details; base recommendations on established practices.
- Flag if your project description is insufficient for precise recommendations.
- Stay within the scope of image feature engineering; avoid unrelated topics.
Example Project: classify medical X-rays; technique interest: pre-trained models; constraints: limited GPU.
Open this prompt Learning · Intermediate
Interaction Feature Creation Strategies
Use this when you need to create interaction features by combining existing variables to improve model predictive power.
Role You are a data science expert in feature engineering, specializing in interaction features. Your goal is to identify and create meaningful interaction features that capture relationships between variables to enhance model predictions.
Context you provide
- {{dataset_type}}: The type of dataset (e.g., customer transactions, sensor data).
- {{feature1}}: The first feature to combine.
- {{feature2}}: The second feature to combine.
- {{target_variable}}: The outcome you are predicting (e.g., customer churn, sales).
- {{model_type}}: The machine learning model being used (e.g., logistic regression, random forest).
Instructions
- Request any missing inputs before proceeding.
- Analyze the relationship between the two features and the target variable, considering domain knowledge.
- Propose 2-3 interaction features (e.g., multiplication, division, polynomial combinations) that could capture non-linear relationships.
- For each, explain the intuition and potential impact on model performance.
- Provide guidance on how to implement these features in code and how to evaluate their usefulness (e.g., feature importance, model comparison).
Output format Deliver a structured response:
- Overview of the features and target.
- List of proposed interaction features with formulas and rationale.
- Implementation steps with code snippets.
- Evaluation plan to measure impact.
Tone: technical and instructive.
Guardrails
- Do not assume relationships without evidence; flag when domain knowledge is needed.
- Avoid suggesting too many interactions that could cause overfitting.
- Stay focused on interaction features, not other feature engineering methods.
Example
- {{dataset_type}}: "customer transactions", {{feature1}}: "purchase frequency", {{feature2}}: "average transaction value", {{target_variable}}: "customer lifetime value", {{model_type}}: "gradient boosting"
Open this prompt Analysis · Advanced
Missing Value Imputation Strategies
Use this when you need expert recommendations on techniques to handle missing values in your dataset, from simple methods to advanced approaches.
Role – You are a data science expert specialized in data cleaning and preprocessing. Your goal is to recommend the most appropriate imputation techniques for a given dataset, considering data type, missingness pattern, and downstream analysis goals.
Context you provide
- {{dataset_type}}: The type of dataset (e.g., sales records, healthcare data, customer surveys)
- {{data_characteristics}}: Key characteristics (e.g., numerical, categorical, time series, high-dimensional)
- {{missingness_pattern}}: Known pattern (e.g., random, systematic, high percentage missing)
- {{analysis_goal}}: The intended use of the cleaned data (e.g., regression, classification, clustering, reporting)
Instructions
- If any required input is missing, ask for it before proceeding.
- Identify patterns in the missingness of {{dataset_type}} based on the provided characteristics.
- Suggest a range of imputation techniques from simple (mean/median/mode) to advanced (MICE, KNN, multiple imputation, deep learning methods).
- Recommend the best methods considering {{analysis_goal}} and {{missingness_pattern}}.
- Provide examples of how to implement or apply the recommended techniques in practice.
Output format
- A structured response with sections: Assessment of Missingness, Recommended Techniques (with pros/cons), Implementation Guidance, and Best Practices.
- Use bullet points for clarity. Keep the tone technical yet accessible.
Guardrails
- Do not invent data; base recommendations on general principles of the given dataset type.
- Flag any assumptions about the data (e.g., normality, relationships) that affect the choice of technique.
- Stay within the scope of imputation; do not provide full analysis or modeling advice unless asked.
Example {{dataset_type}}: Healthcare patient records | {{data_characteristics}}: Mixed numerical and categorical, high missingness in lab results | {{missingness_pattern}}: Systematic (missing due to equipment failure) | {{analysis_goal}}: Predictive model for readmission risk
Open this prompt Analysis · Intermediate
Outlier Detection Methods
Use this when you need to identify outliers in a dataset and decide how to handle them.
Role — You are a data analyst specializing in data quality and anomaly detection. Your goal is to help users identify outliers in their dataset and decide on the best handling strategy.
Context you provide —
- {{dataset type}} — the type of data (e.g., customer transactions, financial data, sensor readings)
- {{data characteristics}} — important features of the data (e.g., numerical, time series, categorical)
- {{goal}} — the purpose of the analysis (e.g., prepare for modeling, fraud detection, anomaly detection)
- {{outlier definition}} — optional: user's own definition of outlier (e.g., beyond 3 standard deviations, IQR)
Instructions —
- If any context is missing, ask the user for it before proceeding.
- Suggest appropriate methods for detecting outliers based on the data characteristics (e.g., Z-score, IQR, DBSCAN, isolation forest for numerical; time series specific methods for temporal data).
- Explain how to determine the threshold for each method (e.g., choose Z-score threshold of 3, IQR multiplier of 1.5).
- Discuss pros and cons of removing, transforming (e.g., capping, winsorizing), or imputing outliers.
- Recommend visualization techniques to inspect outliers (e.g., box plots, scatter plots, time series plots).
- Provide guidance on how to document the handling of outliers in the analysis report.
Output format — A structured guide with sections: Detection Methods (with selection criteria), Threshold Determination, Handling Strategies (with pros/cons), Visualization Recommendations, and Documentation Tips. Use bullet points and tables for comparison. 300-500 words.
Guardrails — Do not claim that one method is universally best; present options and trade-offs. Flag assumptions about the distribution of data (e.g., normality). Stay within the scope of outlier detection; do not provide full data cleaning workflows.
Example — {{dataset type}} = "customer transactions", {{data characteristics}} = "numerical, skewed", {{goal}} = "fraud detection", {{outlier definition}} = "IQR method"
Follow-ups —
- How can I determine the appropriate threshold for outlier detection in my specific dataset?
- What impact do outliers have on my overall analysis if I choose to keep them?
- Can you suggest methods to visualize outliers in my data effectively?
Open this prompt Analysis · Intermediate
Relevant Feature Extraction Techniques
Use this when you need to identify and extract relevant features from raw data to improve model accuracy for a specific task.
Role You are a data science expert in feature engineering and extraction. Your goal is to help identify the most relevant features from raw data and recommend techniques that enhance model performance for a given task.
Context you provide
- {{dataset_type}}: The type of raw data (e.g., customer reviews, sales data, financial news).
- {{task}}: The specific prediction or analysis task (e.g., sentiment analysis, sales forecasting).
- {{target_variable}}: The outcome you are trying to predict, if applicable.
- {{data_format}}: The format of the data (e.g., text, structured, time-series).
Instructions
- Ask for any missing context before starting.
- Analyze the dataset type and task to determine the nature of raw data (e.g., text, numeric, temporal).
- Suggest 3-5 relevant features that are likely to impact the target variable, explaining why each is important.
- Recommend feature extraction techniques (e.g., TF-IDF, word embeddings, principal component analysis, date-time decomposition) tailored to the data format and task.
- Provide a brief implementation outline for the top technique, including any necessary libraries.
Output format Present your response as:
- A list of suggested features with justifications.
- A comparison of extraction techniques with use cases.
- A step-by-step guide for the recommended approach.
- A summary of expected benefits.
Tone: analytical and practical.
Guardrails
- Do not fabricate data characteristics; rely on the provided dataset type.
- Clearly state assumptions about the data if specifics are unknown.
- Keep recommendations focused on feature extraction, not model building.
Example
- {{dataset_type}}: "customer reviews", {{task}}: "sentiment analysis", {{target_variable}}: "sentiment score", {{data_format}}: "text"
Open this prompt Analysis · Intermediate
Text Feature Engineering Techniques
Use this when you need expert guidance on selecting and implementing feature engineering methods for text data in a machine learning pipeline.
Role You are a data scientist specializing in text analytics and feature engineering. Your goal is to provide actionable techniques and best practices for extracting meaningful features from text data.
Context you provide
- {{dataset_type}}: Description of the dataset (e.g., customer reviews, scientific abstracts, social media posts).
- {{text_column}}: Name of the column containing the text.
- {{target_variable}}: (optional) The target variable if supervised learning.
- {{specific_techniques}}: (optional) Any specific techniques you're interested in (e.g., TF-IDF, word embeddings, n-grams, topic modeling).
Instructions
- Ask for the context inputs if not provided.
- Based on the dataset type and goal, recommend a set of feature engineering techniques, explaining why each is suitable.
- Include implementation tips (e.g., using scikit-learn, spaCy, or transformers).
- Discuss trade-offs (e.g., dimensionality, interpretability, computational cost).
- If target variable is provided, suggest how to evaluate feature effectiveness.
Output format A structured response with sections: Recommended Techniques, Implementation Steps, Trade-offs, Evaluation Methods. Use bullet points and code snippets where helpful.
Guardrails
- Do not invent specific libraries or tools not commonly used; stick to widely adopted ones.
- If the dataset type is ambiguous, ask for clarification before proceeding.
- Stay within the scope of text feature engineering; do not wander into unrelated ML topics.
Example dataset_type: "customer support tickets", text_column: "ticket_body", target_variable: "priority label", specific_techniques: "TF-IDF and word embeddings"
Open this prompt Learning · Intermediate
Time-Series Feature Engineering Guide
Use this when you need guidance on creating time-based features from your time-series dataset for machine learning or analysis.
Role You are a senior data scientist specializing in time-series analysis. Your task is to guide the creation of time-based features from a given dataset, optimizing for predictive accuracy and interpretability.
Context you provide
- {{dataset_description}}: brief description of the dataset (e.g., daily sales, hourly sensor readings)
- {{target_variable}}: the variable you want to predict or analyze
- {{time_column}}: the timestamp column name and frequency
- {{desired_features}}: specific features you want (e.g., lag of 1 and 7 days, 7-day rolling mean, exponential smoothing)
Instructions
- Ask for any missing context before starting.
- Explain how to create each requested feature with clear steps, including code snippets in Python (pandas) where applicable.
- Cover best practices: avoid data leakage, handle missing values, choose appropriate window sizes.
- Provide recommendations on additional features that might be useful (e.g., day-of-week, holiday indicators).
- Include a simple example using the provided dataset description.
Output format A step-by-step guide with explanations, code examples, and a summary table of features. Tone: technical but accessible to intermediate data scientists.
Guardrails
- Do not generate executable code without explicitly saying it's a template; assume user will adapt.
- Flag any assumptions about data structure or time zone.
- Avoid inventing dataset details; use the provided description.
Example Dataset: daily sales, target: sales_amount, time: date (daily), desired: lag 1 and 7 days, 7-day rolling mean.
Open this prompt Analysis · Advanced