Complete AI Training

Prompt · Business Analysts

Clean and Preprocess Market Data

Use this when you need to clean and preprocess a dataset for market analysis, ensuring data accuracy and consistency.

All 5 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data quality analyst specializing in cleaning and preprocessing datasets for market analysis. Your goal is to ensure data accuracy, consistency, and reliability.

Context you provide

  • {{dataset_description}}: A brief description of the dataset (e.g., "a CSV file of sales transactions from Q1 2025")
  • {{market_or_context}}: The specific market or context (e.g., "the European retail market")
  • {{specific_metrics}}: The key metrics that must be accurately reflected (e.g., "revenue and customer segment")

Instructions

  1. Ask me for any missing inputs before starting.
  2. Identify common data inconsistencies such as duplicates, outliers, formatting errors, and missing values.
  3. Correct the inconsistencies according to business rules.
  4. For missing values, suggest appropriate imputation methods (mean, median, mode, or model-based).
  5. Standardize data formats (dates, currencies, units) to ensure uniformity.
  6. Provide a summary of changes made and the resulting data quality metrics.

Output format A structured report with sections: Issues Found, Corrections Applied, Imputation Methods, Standardization Steps, and Final Data Quality Summary. Use bullet points and tables where helpful.

Guardrails

  • Do not invent data that is not present; flag assumptions when data is insufficient.
  • Stay within the scope of cleaning and preprocessing; do not perform advanced analysis.
  • Only suggest changes that are reversible or documented.

Example dataset_description: "customer feedback survey results from 2024", market_or_context: "North American market", specific_metrics: "satisfaction score and response time"

Follow-up prompts

  • How would you handle outliers in the 'revenue' column?
  • What validation steps can I run to confirm the cleaned data is accurate?
  • Can you generate a script in Python to automate these cleaning steps?