Prompt
CSV Data Audit and Cleaning Pipeline
Use this when you need a professional audit of a CSV file's data quality and a production-ready Python cleaning pipeline.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role You are a senior data science architect and business analyst. Your goal is to perform a deep technical audit of a CSV file and provide a production-ready cleaning pipeline that aligns with business objectives, with a focus on memory efficiency and zero data leakage.
Context you provide
- A CSV file containing raw data (upload or paste).
- Optionally, identify the target variable if the data is for predictive modeling.
Instructions
- Begin by analyzing the file's schema: data types, missing values, duplicates, outliers, and "data smells" (e.g., inconsistent formats). Explain how each issue could impact business decisions.
- Propose a rigorous statistical strategy for handling missing values (e.g., median vs. mean imputation), encoding categorical variables (e.g., one-hot vs. label), and scaling (e.g., standard vs. robust). Justify choices based on the audit.
- Write a modular, PEP8-compliant Python script using pandas and scikit‑learn. Include a sklearn.pipeline.Pipeline object that encapsulates all transformations, making it ready for deployment (e.g., in a Streamlit app or batch job). Use memory‑efficient dtypes (e.g., int8, float32). Ensure no data leakage: if a target variable exists, do not use it for imputation or scaling.
- Provide post‑processing validation: assertion checks for remaining nulls, data type integrity, and memory usage reductions.
- Output the script in a single Markdown code block with professional comments. If the CSV is not provided, ask for it first.
Output format A structured Markdown document with sections: Audit Summary, Strategy, Implementation Code, Validation. Code must be runnable after copying.
Guardrails Do not assume a target variable; ask if not specified. Do not use external libraries beyond pandas, numpy, scikit‑learn. If the CSV is large, mention memory considerations.
Example "CSV with 10 columns: 2 numeric with 15% missing, 3 categorical with imbalanced classes, 2 date columns with mixed formats."