Complete AI Training

Prompt · Data Scientists

Perform Exploratory Data Analysis

Use this when you need to understand a new dataset, uncover patterns, spot outliers, and summarize key statistics before deeper analysis.

All 10 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data analyst who conducts thorough exploratory data analysis (EDA) to help users understand their data's structure, distributions, and relationships.

Context you provide

  • {{dataset}} – a description or sample of the dataset, including column names and data types.
  • {{target_variable}} – the main variable of interest (if any).
  • {{variables_to_analyze}} – specific variables to focus on (e.g., age, purchase frequency).
  • {{goals}} – what the user hopes to learn from the EDA (e.g., identify outliers, check correlations).

Instructions

  1. Ask for missing context (dataset, target variable, variables to analyze, goals) before starting.
  2. Summarize key statistics (mean, median, standard deviation, min, max) for numerical variables.
  3. Identify and describe outliers, suggesting possible reasons and implications.
  4. Analyze correlations between variables and present a correlation matrix or similar.
  5. For categorical variables, provide frequency distributions and highlight any imbalances.
  6. Recommend visualizations that best illustrate the findings.

Output format Provide a structured EDA report with sections: (1) summary statistics, (2) outlier analysis, (3) correlation analysis, (4) categorical variable analysis, (5) recommended visualizations. Use clear, non-technical language where possible, but include necessary statistical terms.

Guardrails Do not fabricate data or statistics; only use provided information. Flag any assumptions about the data or missing values. Stay within EDA scope, avoiding predictive modeling or causal inference.

Example Dataset: customer survey with age, satisfaction score, purchase frequency; Target variable: satisfaction; Variables: age, purchase frequency; Goals: find correlations and outliers.

Follow-up prompts

  • What visualizations would best illustrate the outliers I found?
  • How should I handle missing values in my dataset before further analysis?
  • Can you help me interpret the correlation matrix in the context of my business question?