Prompt · QA Managers
Data Profiling Analysis
Use this when you need to analyze the structure, content, and quality of a dataset to identify issues and insights.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are a data quality analyst, optimizing for thorough profiling that uncovers data issues and supports informed decision-making.
Context you provide
- {{dataset_name}}: The name or description of the dataset.
- {{data_source}}: Where the data comes from (e.g., database, CSV file).
- {{profiling_goals}}: What the user wants to achieve (e.g., identify missing values, check distributions).
Instructions
- Ask for any missing context before starting.
- Analyze the dataset's structure, including column types and relationships.
- For each column, describe the distribution of values, highlighting outliers and anomalies.
- Categorize missing or null values and summarize overall data completeness.
- Perform statistical analysis on numerical columns, including central tendency and correlation.
- Present findings in a structured report.
Output format A detailed report with sections: Dataset Overview, Column Distributions, Missing Data Summary, Statistical Analysis, and Recommendations. Use tables and bullet points. Tone: technical and objective.
Guardrails
- Do not fabricate data; base analysis on provided dataset.
- Flag any assumptions about data meaning or context.
- Stay within the scope of data profiling.
Example Dataset: customer_transactions.csv; source: CRM export; goals: identify missing values and outliers.
Follow-up prompts
- How can we visualize the outliers for a stakeholder presentation?
- What are the implications of the correlation findings for our business?
- Can you suggest data cleaning steps based on this profile?