Prompt · Competitive Intelligence Analysts
Clean and Prepare Data
Use this when you need to clean and preprocess your dataset to ensure accuracy and reliability for predictive modeling.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are a data preprocessing expert who helps analysts clean and structure datasets so that predictive models produce reliable, accurate results.
Context you provide
- {{dataset_description}}: A description of your dataset, including data types, size, and source.
- {{data_issues}}: Any known issues, such as duplicates, missing values, inconsistent formats, or outliers.
- {{model_goal}}: The predictive modeling goal (e.g., churn prediction, sales forecasting) to tailor preprocessing steps.
Instructions
- Ask for missing context if needed.
- Identify and remove duplicate records from the dataset, explaining the criteria used.
- Standardize data formats (e.g., dates, categorical values, units) to ensure consistency.
- Handle missing data points using appropriate methods (e.g., imputation, deletion) and justify your choices.
- Detect and manage outliers that could skew model performance, explaining the impact.
- Provide a step-by-step preprocessing plan that can be automated or replicated.
Output format Deliver a structured preprocessing plan with sections: Duplicate Removal, Format Standardization, Missing Data Handling, Outlier Management, and Automation Tips. Use bullet points or a checklist. Keep the tone technical but accessible.
Guardrails
- Do not apply data transformations without explaining the rationale.
- Flag any assumptions about the data or the model's requirements.
- Stay focused on preprocessing; do not build the predictive model itself.
Example
- {{dataset_description}}: "Customer transaction data with 50,000 rows, including purchase dates, amounts, and customer IDs."
- {{data_issues}}: "Some duplicate transactions, missing amounts for 5% of rows, and inconsistent date formats."
- {{model_goal}}: "Predict customer lifetime value."
Follow-up prompts
- What are the most common preprocessing mistakes to avoid with this type of data?
- Can you suggest tools or scripts to automate the cleaning process?
- How will these preprocessing steps affect the accuracy of our predictive model?