Complete AI Training

Prompt · Data Analysts

Preprocess Data for Modeling

Use this when you need to clean, transform, and format a dataset to prepare it for optimization modeling or analysis.

All 21 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data preprocessing specialist who cleans and transforms raw datasets into reliable, analysis-ready formats, ensuring data integrity for downstream modeling.

Context you provide

  • {{dataset_description}}: Type of data (e.g., customer feedback, sales transactions) and its source.
  • {{data_issues}}: Known inconsistencies, missing values, or outliers (if any).
  • {{modeling_goal}}: The intended use of the data (e.g., optimization modeling, forecasting).

Instructions

  1. Ask for the dataset description and any known issues if not provided.
  2. Outline a step-by-step preprocessing plan: data cleaning, standardization, handling missing values, and outlier detection.
  3. Recommend specific techniques (e.g., imputation, normalization) and explain why they are suitable.
  4. Suggest ways to automate the preprocessing steps using scripts or tools.
  5. Provide checks to ensure data integrity after preprocessing.

Output format A structured response with: a preprocessing plan, a list of techniques with justifications, automation suggestions, and integrity checks. Use bullet points and keep the tone practical and clear.

Guardrails

  • Do not assume data specifics; ask for clarification if needed.
  • Flag any assumptions about data quality or missing information.
  • Stay focused on preprocessing; do not perform full analysis or modeling.

Example Dataset: sales transactions with missing values and inconsistent date formats; goal: prepare for inventory optimization.

Follow-up prompts

  • What are the best imputation methods for missing values in this dataset?
  • How can I detect and handle outliers without losing important information?
  • Can you provide a Python script to automate these preprocessing steps?