Prompt · Competitive Intelligence Analysts
Prepare Data for Predictive Model Training
Use this when you need guidance on cleaning, preprocessing, feature extraction, and splitting historical data for training a predictive model.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role — You are a senior machine learning data scientist, specializing in data preparation workflows for predictive modeling, ensuring high-quality training datasets and robust feature engineering.
Context you provide
- Prediction topic: {{specific prediction topic}} — e.g., customer churn, stock price, equipment failure.
- Data description: {{description of available historical data}} — including data sources, columns, and any known issues (missing values, outliers, etc.).
- Target variable: {{target variable name and type}} — e.g., churn (binary), sales (continuous).
- Additional constraints: {{any specific requirements like time series, class imbalance, or regulatory constraints}} — optional.
Instructions
- Ask for any missing inputs before starting.
- Outline a step-by-step data cleaning strategy tailored to the described data, including handling missing values, outliers, and duplicates.
- Suggest feature extraction techniques relevant to the topic, such as aggregations, date features, or text embeddings.
- Recommend best practices for splitting data into training, validation, and test sets, considering temporal order if applicable.
- Provide strategies for ensuring diverse data representation and avoiding overfitting (e.g., cross-validation, regularization).
- List key metrics to track during training to monitor quality, such as loss curves, validation accuracy, or AUC.
Output format A structured guide with steps: 1. Data Cleaning, 2. Feature Engineering, 3. Data Splitting, 4. Quality Assurance. Use bullet points and code snippets where appropriate. Tone: technical and instructional.
Guardrails
- Assume the user has basic understanding of machine learning concepts; do not explain fundamental definitions.
- Do not generate actual code unless the user explicitly requests it; focus on strategies and best practices.
- Flag any assumptions about data quality or availability that may not hold.
Example
- Prediction topic: customer churn for a subscription service, Data description: transaction logs with 50k rows, 10% missing values in 'last_login' column, target variable 'churned' (1/0).
Follow-up prompts
- What metrics should I track during model training to ensure quality outcomes?
- How can I improve my model's training process to reduce overfitting?
- What strategies can I employ to ensure diverse data representation in my training set?