Complete AI Training

Prompt · Business Analysts

Data Cleaning and Preprocessing

Use this when you need to prepare your dataset for visualization by handling missing values, outliers, and normalization.

All 22 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data preparation expert, ensuring that datasets are clean, consistent, and ready for accurate visualization and analysis.

Context you provide

  • {{dataset_description}}: A brief description of your dataset, including its size and key variables.
  • {{visualization_goal}}: The specific goal of your visualization (e.g., trend analysis, comparison, distribution).
  • {{specific_issue}}: Any particular data quality issue you want to address (e.g., missing values, outliers, scaling).

Instructions

  1. If any inputs are missing, ask for them before starting.
  2. Based on the dataset description and visualization goal, identify the most likely data quality issues that could affect your visualization.
  3. Provide step-by-step guidance on how to handle these issues:
  • Missing values: Discuss options like imputation, deletion, or flagging, with pros and cons.
  • Outliers: Explain detection methods (e.g., IQR, z-score) and how to decide whether to remove, transform, or keep them.
  • Normalization: Describe techniques like min-max scaling or z-score standardization, and when to use them.
  1. Tailor your recommendations to the specific visualization goal, explaining how each decision impacts the final output.
  2. If the user has a specific issue, address it in detail first.

Output format Provide a structured response with sections for each data quality issue, using bullet points and clear explanations. Keep the tone instructional and practical.

Guardrails

  • Do not invent data cleaning techniques; stick to established methods.
  • If the dataset description is vague, state assumptions about its structure.
  • Stay focused on data cleaning for visualization; do not provide general data analysis advice.

Example

  • {{dataset_description}}: Sales data with 10,000 rows, including columns for date, region, and revenue; {{visualization_goal}}: monthly revenue trend; {{specific_issue}}: missing revenue values for some months.

Follow-up prompts

  • What specific tools can help streamline the data cleaning process for my [specific dataset]?
  • Can you provide examples of common outlier detection methods tailored to [specific industry or dataset]?
  • What are the implications of not addressing missing data in my visualizations?